Physiological and emotional state recognition method for diversified people

By combining spectral adaptive enhancement and multimodal signal pattern recognition technologies with a population stratification verification network, the problem of skin color bias and environmental interference in remote photoplethysmography technology for diverse populations has been solved. Robust, fair, and accurate emotion recognition for diverse populations has been achieved, making it suitable for scenarios such as smart cockpits, security monitoring, and telemedicine.

CN122320548APending Publication Date: 2026-07-03BEIJING QINGSI INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING QINGSI INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-04-14
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing remote photoplethysmography technology suffers from skin color bias, environmental interference, and age differences when applied to diverse populations, resulting in unfair recognition performance, poor robustness, and difficulty in achieving accurate emotional state recognition.

Method used

Near-infrared simulation channels are constructed using spectral adaptive enhancement technology. Combined with multimodal signal pattern recognition and population stratification verification networks, skin color bias is eliminated, modal complementarity and fault tolerance are achieved, and the differences in autonomic nervous system responses among different populations are adapted to.

Benefits of technology

It achieves robust, fair, and accurate emotion recognition for diverse populations, improves recognition accuracy in complex scenarios, and provides a reliable non-contact emotion state monitoring solution for scenarios such as smart cockpits, security monitoring, and telemedicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122320548A_ABST
    Figure CN122320548A_ABST
Patent Text Reader

Abstract

This invention relates to the fields of physiological signal processing and emotion recognition technology, and discloses a method for recognizing physiological and emotional states of diverse populations. The method includes: acquiring facial video streams, infrared thermal imaging, and electrodermal signals; performing spectral adaptive enhancement based on skin color to obtain a spectrally enhanced video stream; inputting the spectrally enhanced video stream into a spatially modulated attention network to obtain a remote photoplethysmography waveform; extracting temperature distribution features from infrared thermal imaging; extracting electrodermal response features from electrodermal signals; aligning the three signals spatiotemporally and constructing alternative features for the failed signal based on the signal-to-noise ratio; inputting the three signals and alternative features into a multimodal signal pattern recognition network for weighted fusion, outputting a fused physiological feature; inputting the fused physiological feature into a population-level verification network for population adaptive matching and individual baseline correction, outputting the final emotional state. This invention achieves robust, fair, and accurate emotion recognition for diverse populations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of physiological signal processing and emotion recognition technology, and in particular to a method for recognizing the physiological and emotional states of diverse populations. Background Technology

[0002] With the rapid development of public safety, smart cockpits, and telemedicine, non-contact and seamless real-time monitoring of people's physiological and emotional states has become a core technological requirement. Remote photoplethysmography (TPM) technology extracts physiological indicators such as heart rate and heart rate variability by analyzing periodic changes in skin color in facial videos, providing a non-contact solution for emotion recognition. This technology can directly utilize existing camera equipment, offering advantages such as low deployment costs and wide applicability, demonstrating broad application prospects in scenarios such as driver status monitoring, security checkpoint risk warning, and online psychological counseling.

[0003] However, existing remote photoplethysmography (TP) technology suffers from fundamental shortcomings in practical applications targeting diverse populations. First, due to the strong light absorption characteristics of melanin in dark-skinned individuals, the signal-to-noise ratio of TP signals for these individuals is only one-third that of those for light-skinned individuals, leading to significant skin color bias in recognition performance. Second, single TP signals are easily affected by environmental interference such as changes in lighting, head movements, and facial occlusion, and are prone to failure in complex scenarios. Furthermore, single physiological indicators such as heart rate variability are insufficient to distinguish between different emotional states corresponding to similar physiological responses like tension, excitement, and fatigue, and the autonomic nervous system response patterns differ significantly across age groups, making it impossible for general models to achieve personalized and accurate recognition. These shortcomings make it difficult for existing technologies to achieve fair, robust, and accurate emotional state recognition in multi-ethnic settings such as airports and train stations, as well as among different age groups such as the elderly and youth. A technological solution that can simultaneously address population differences and environmental robustness is urgently needed.

[0004] Therefore, this invention proposes a method for recognizing the physiological and emotional states of diverse populations. Summary of the Invention

[0005] This invention provides a method for recognizing the physiological and emotional states of diverse populations. It eliminates skin color bias through spectral adaptive enhancement, achieves modal complementarity and fault tolerance through multimodal signal pattern recognition, and realizes personalized and accurate recognition through population stratification verification. It solves the technical defects of existing technologies, such as unfair recognition of people of different skin colors and ages, easy failure of signals in complex scenarios, and inability of single physiological indicators to accurately distinguish emotional states. It achieves robust, fair, and accurate emotion recognition for diverse populations.

[0006] This invention provides a method for recognizing the physiological and emotional states of diverse populations, including: Acquire facial video streams, infrared thermal imaging sequences, and skin electrical activity signals of the target object; The facial video stream is subjected to spectral adaptive enhancement based on the skin color type of the target object to obtain a spectral enhanced video stream. Spectral adaptive enhancement includes adjusting the weight of each color channel according to the skin color type. When the skin color type is dark, a near-infrared simulation channel is also constructed. The spectral-enhanced video stream is input into a spatially modulated attention network to obtain a remote photoplethysmography waveform; Temperature distribution feature vectors are extracted from infrared thermal imaging sequences. These feature vectors include the forehead temperature change rate, the nose tip temperature fluctuation amplitude, and the cheek temperature asymmetry index. The skin conductance response feature vector is extracted from the skin conductance activity signal. The skin conductance response feature vector includes the baseline value of skin conductance level, skin conductance response frequency, and skin conductance recovery time. Spatiotemporal alignment of remote photoplethysmography waveforms, temperature distribution feature vectors, and skin conductance response feature vectors; The validity of a signal is determined by the real-time signal-to-noise ratio of the remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector. For failed signals with a real-time signal-to-noise ratio lower than a preset threshold, alternative features of the failed signal are constructed based on the physiological correlation characteristics of other valid signals. The remote photoplethysmography waveform, temperature distribution feature vector, skin conductance response feature vector, and alternative features are input into a multimodal signal pattern recognition network. The multimodal signal pattern recognition network dynamically determines the fusion weights based on the effective signal-to-noise ratio of the remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector, and performs weighted fusion to output a fused physiological feature vector. The fused physiological feature vector is input into the population stratification validation network. The population stratification validation network matches the initial candidate emotional state from multiple preset population physiological feature templates based on the target object's skin color type and age range, and corrects the initial candidate emotional state based on the target object's individual baseline physiological characteristics, and outputs the final emotional state.

[0007] Furthermore, constructing the near-infrared simulation channel includes: The average pixel value of multiple skin regions in the facial video stream of the target object is identified, and the skin color type of the target object is determined based on the average pixel value. The skin color types include light skin color, medium skin color and dark skin color. When the skin tone type is light skin tone, the green channel weight of the facial video stream is set as the first weight value, and the red and blue channel weights are set as the second weight values, with the first weight value being greater than the second weight value; When the skin tone type is dark skin tone, the weighted difference between the pixel values ​​of the red channel and the pixel values ​​of the blue channel in the facial video stream is used as the pixel value of the near-infrared analog channel. The weight of the near-infrared analog channel in the facial video stream is set as the third weight value, and the weight of the green channel in the facial video stream is set as the fourth weight value. The third weight value is greater than the fourth weight value. The pixel values ​​of each channel of the facial video stream are multiplied by their corresponding channel weights and then summed to generate a spectrally enhanced video stream.

[0008] Furthermore, the spatially modulated attention network obtains a long-range photoplethysmography waveform, including: The spectral-enhanced video stream is divided into multiple time windows, each containing a preset number of consecutive video frames; Spatiotemporal feature extraction is performed on video frames within each waveform extraction time window. Spatiotemporal feature extraction includes extracting spatial features through three-dimensional convolution and extracting color change difference features between adjacent frames through temporal center difference convolution. Temporal center difference convolution enhances the ability to capture periodic changes in skin color by calculating the feature difference between the current frame and adjacent frames. The extracted spatiotemporal features are segmented into multiple spatiotemporal channels, and each spatiotemporal channel corresponds to the feature sequence of a local region of the face within the waveform extraction time window; A spatial importance map is generated based on the feature sequences of all spatiotemporal pipelines. Regions with high weight values ​​in the spatial importance map correspond to skin areas in the spectral enhanced video stream that are fully exposed, have small motion amplitude, and are unobstructed. Attention weighting is applied to the feature sequences of the spatiotemporal pipeline based on the spatial importance map, so that regions with high weight values ​​in the spatial importance map receive higher attention weights. The attention-weighted spatiotemporal pipeline feature sequence is decoded into a one-dimensional remote photoplethysmography waveform.

[0009] Further, the real-time signal-to-noise ratio of the remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector is calculated, including: The spatiotemporally aligned remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector are used as inputs; Within a preset time window in which the target object's face remains still, the signal amplitude sequence of the remote photoplethysmography waveform is acquired. The average value of the preset ratio range with the lowest amplitude value in the signal amplitude sequence is taken as the baseline noise amplitude. The ratio of the peak signal amplitude to the baseline noise amplitude within the current time window is taken as the first real-time signal-to-noise ratio of the remote photoplethysmography waveform. Multiple sets of facial temperature distribution data were collected when the target object was in a resting state. Temperature fluctuation components unrelated to emotional response were extracted from each set of facial temperature distribution data. A preset noise template was constructed, and the matching degree between the temperature change curve in the current time window and the preset noise template was used as the second real-time signal-to-noise ratio of the temperature distribution feature vector. During the resting period when there is no response to the skin conductance signal, the low-frequency fluctuation amplitude of the skin conductance level signal is calculated as the baseline drift amplitude, and the ratio of the skin conductance response amplitude to the baseline drift amplitude within the current time window is used as the third real-time signal-to-noise ratio of the skin conductance response feature vector.

[0010] Furthermore, the multimodal signal pattern recognition network dynamically determines the fusion weights and performs weighted fusion based on the effective signal-to-noise ratio of the remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector, including: The first real-time signal-to-noise ratio of the remote photoplethysmography waveform is converted into a first fusion weight through a first mapping function. The second real-time signal-to-noise ratio of the temperature distribution feature vector is converted into a second fusion weight through a second mapping function. The third real-time signal-to-noise ratio of the skin conductance response feature vector is converted into a third fusion weight through a third mapping function. The first, second, and third mapping functions are monotonically increasing functions, so that signals with high signal-to-noise ratios receive higher fusion weights. The remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector are weighted and concatenated according to the first fusion weight, the second fusion weight, and the third fusion weight to generate a fused physiological feature vector.

[0011] Furthermore, it also includes determining that the remote photoplethysmography waveform is in a failed state when the first real-time signal-to-noise ratio of the remote photoplethysmography waveform is lower than the first preset threshold, constructing alternative features of the remote photoplethysmography waveform based on the temperature distribution feature vector and the skin conductance response feature vector, and using a preset alternative weight as the fusion weight of the remote photoplethysmography waveform. The alternative weight is greater than the fusion weight transformed by the mapping function of the first real-time signal-to-noise ratio, and does not exceed the maximum fusion weight of the effective signal. When the second real-time signal-to-noise ratio of the temperature distribution feature vector is lower than the second preset threshold, the temperature distribution feature vector is determined to be in a failed state, and alternative features of the temperature distribution feature vector are constructed based on the remote photoplethysmography waveform and the skin conductance response feature vector. When the third real-time signal-to-noise ratio of the skin electric response feature vector is lower than the third preset threshold, the skin electric response feature vector is determined to be in a failed state. Alternative features of the skin electric response feature vector are constructed based on the remote photoplethysmography waveform and temperature distribution feature vector. The alternative features are constructed by extracting feature components that are physiologically related to the failure signal from the effective signal, and mapping the feature components to the estimated value of the failure signal through a pre-trained regression model.

[0012] Furthermore, temperature distribution feature vectors are extracted from the infrared thermal imaging sequence, including: In each frame of the infrared thermal imaging sequence, the forehead region, nose tip region, and cheek region are located, with the cheek region including the left cheek region and the right cheek region. Calculate the slope of the curve showing the change in average temperature of the forehead region over time within a time window, and use the slope value as the rate of change in forehead temperature. Calculate the difference between the maximum and minimum average temperatures of the nasal tip region within the time window, and use this difference as the nasal tip temperature fluctuation range. Calculate the average temperature difference between the left and right cheek regions within a time window, and use the difference as the cheek temperature asymmetry index.

[0013] Furthermore, skin conductance response feature vectors are extracted from the skin conductance activity signals, including: During the resting period when there is no response to the skin electrical activity signal, the average value of the skin conductance level is calculated as the baseline value of the skin conductance level; The number of skin conductance responses is counted within a time window. Skin conductance response is defined as the process by which the skin conductance level rises from the baseline value of the skin conductance level to the baseline value of the skin conductance level after exceeding a preset amplitude threshold. The ratio of the number of skin conductance responses to the duration of the time window is used as the skin conductance response frequency. For each skin conductance response, the time difference from the start of the response to the recovery to the baseline level of skin conductance is calculated, and the average recovery time of all skin conductance responses is taken as the skin conductance recovery time.

[0014] Furthermore, the population stratification validation network matches initial candidate emotional states from multiple pre-defined population physiological feature templates based on the target object's skin color type and age range, including: The target population subtype is determined based on the combination of the target population’s skin color type and the target population’s age range. The skin color type includes light skin color type, medium skin color type and dark skin color type, and the age range includes first age range, second age range and third age range. Extract the population physiological feature template corresponding to the population subtype to which the target object belongs from the preset template library. The population physiological feature template contains the statistical distribution parameters of the fused physiological feature vectors corresponding to each emotional state under the population subtype. The statistical distribution parameters include the mean vector and covariance matrix corresponding to each emotional state. The Mahalanobis distance between the fused physiological feature vector and the statistical distribution parameters corresponding to each emotional state is calculated. Based on the Mahalanobis distance, the probability value of the fused physiological feature vector belonging to each emotional state is determined as the initial confidence level, and the initial candidate emotional states are obtained.

[0015] Furthermore, the population-stratified validation network modifies the initial candidate emotional states based on the individual baseline physiological characteristics of the target subjects, including: When the target object uses the system for the first time or when the system triggers the baseline update condition, it is detected whether the target object is in a resting state. The resting state is determined by detecting that the target object's facial movement amplitude is less than the movement threshold and there is no response of skin conductance signal within a preset time. Baseline multimodal physiological signals were collected for a preset duration while the target was in a resting state. The baseline multimodal physiological signals included facial video stream, infrared thermal imaging sequence, and skin electrical activity signals in a resting state. Extracting individual baseline fused physiological feature vectors from baseline multimodal physiological signals; Calculate the basic offset difference between the individual baseline fused physiological feature vector and the resting state standard vector in the population physiological feature template of the target object's subtype; An offset correction factor is generated based on the baseline offset difference. The initial confidence of each emotional state in the initial candidate emotional states is multiplied by the offset correction factor to obtain the corrected confidence. The emotional state with the highest confidence level after correction is taken as the final emotional state.

[0016] The beneficial effects of this invention compared to existing technologies are as follows: Existing technologies have three fundamental defects: First, the signal-to-noise ratio of remote photoplethysmography (TP) signals is significantly lower in people with dark skin than in people with light skin due to the strong absorption of light by melanin, resulting in serious skin color bias; second, single TP signals are easily affected by environmental interference such as changes in lighting, head movements, and facial occlusion, and are prone to failure in complex scenarios; third, single physiological indicators such as heart rate variability are difficult to distinguish between different emotional states corresponding to similar physiological reactions such as tension, excitement, and fatigue, and there are significant differences in the autonomic nervous system response patterns of people of different ages, making it impossible for general models to achieve personalized and accurate identification. This invention utilizes spectral adaptive enhancement technology to construct a near-infrared simulated channel based on skin color type, effectively improving signal extraction quality for dark-skinned individuals and fundamentally eliminating skin color bias to achieve fair recognition across different skin color groups. Through a multimodal signal pattern recognition network, it integrates remote photoplethysmography waveforms, temperature distribution characteristics, and skin conductance response characteristics. These three signals reflect different branches of the autonomic nervous system and possess physiological complementarity; when any signal is interfered with, compensation can be made by other signals, achieving modal complementarity and fault tolerance, significantly improving robustness in complex scenarios. Furthermore, a population-stratified verification network performs adaptive matching based on skin color and age, combined with individual baseline correction, enabling the system to accurately adapt to the differences in autonomic nervous system responses among different age groups, achieving personalized and accurate recognition. This invention achieves robust, fair, and accurate emotion recognition for diverse populations, providing a reliable non-contact emotion state monitoring solution for scenarios such as smart cockpits, security monitoring, and telemedicine.

[0017] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in this application.

[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is an overall flowchart of a method for recognizing the physiological and emotional states of diverse populations according to an embodiment of the present invention. Figure 2 This is a flowchart of multimodal signal processing and fusion in an embodiment of the present invention. Detailed Implementation

[0020] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention. References Figure 1 , Figure 2 This invention provides an embodiment of a method for recognizing the physiological and emotional states of diverse populations, comprising: Acquire facial video streams, infrared thermal imaging sequences, and skin electrical activity signals of the target object; The facial video stream is subjected to spectral adaptive enhancement based on the skin color type of the target object to obtain a spectral enhanced video stream. Spectral adaptive enhancement includes adjusting the weight of each color channel according to the skin color type. When the skin color type is dark, a near-infrared simulation channel is also constructed. The spectral-enhanced video stream is input into a spatially modulated attention network to obtain a remote photoplethysmography waveform; Temperature distribution feature vectors are extracted from infrared thermal imaging sequences. These feature vectors include the forehead temperature change rate, the nose tip temperature fluctuation amplitude, and the cheek temperature asymmetry index. The skin conductance response feature vector is extracted from the skin conductance activity signal. The skin conductance response feature vector includes the baseline value of skin conductance level, skin conductance response frequency, and skin conductance recovery time. Spatiotemporal alignment of remote photoplethysmography waveforms, temperature distribution feature vectors, and skin conductance response feature vectors; The validity of a signal is determined by the real-time signal-to-noise ratio of the remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector. For failed signals with a real-time signal-to-noise ratio lower than a preset threshold, alternative features of the failed signal are constructed based on the physiological correlation characteristics of other valid signals. The remote photoplethysmography waveform, temperature distribution feature vector, skin conductance response feature vector, and alternative features are input into a multimodal signal pattern recognition network. The multimodal signal pattern recognition network dynamically determines the fusion weights based on the effective signal-to-noise ratio of the remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector, and performs weighted fusion to output a fused physiological feature vector. The fused physiological feature vector is input into the population stratification validation network. The population stratification validation network matches the initial candidate emotional state from multiple preset population physiological feature templates based on the target object's skin color type and age range, and corrects the initial candidate emotional state based on the target object's individual baseline physiological characteristics, and outputs the final emotional state.

[0021] In this embodiment, facial video stream, infrared thermal imaging sequence, and skin conductance activity (SCEA) signal of the target object are acquired. The facial video stream is acquired using an RGB camera facing the target object, with a resolution of at least 640 pixels by 480 pixels and a frame rate of at least 20 frames per second. The infrared thermal imaging sequence is acquired using an infrared thermal imager to capture temperature distribution changes on the target object's face. The SCEA signal is acquired using a skin conductance sensor attached to the target object's hand or wrist, which measures changes in skin conductivity. It should be noted that the contact-based acquisition of the SCEA signal serves as an auxiliary mode, together with the non-contact acquisition of the facial video stream and infrared thermal imaging sequence, constituting a multimodal signal input. The SCEA signal is only used in scenarios where the target object actively cooperates. In completely non-contact applications, skin conductivity characteristic analysis based on facial video can be used as an alternative. This alternative infers the relative change in skin conductivity by analyzing changes in reflected light intensity in the facial skin area, thus maintaining the overall non-contact nature of the solution.

[0022] In this embodiment, the skin tone type of the target object is determined by analyzing the average pixel values ​​of multiple skin regions in the facial video stream. Specifically, the forehead, cheek, and chin regions are selected as skin regions from the facial video stream, and the average pixel values ​​of these regions in the RGB color space are calculated. Based on the distribution characteristics of the red, green, and blue channel components in the average pixel values, the skin tone type is classified into light, medium, and dark skin tones.

[0023] In this embodiment, the facial video stream is subjected to spectral adaptive enhancement based on the skin tone type of the target object to obtain a spectrally enhanced video stream. Spectral adaptive enhancement includes adjusting the weights of each color channel according to the skin tone type. When the skin tone type is dark, a near-infrared simulated channel is also constructed. When the skin tone type is light, the weight of the green channel is increased to enhance the ability to capture subtle changes in skin color, while the weights of the red and blue channels are decreased to suppress noise. When the skin tone type is dark, a near-infrared simulated channel is constructed. This channel is obtained by calculating the weighted difference between the pixel values ​​of the red and blue channels. The near-infrared simulated channel can penetrate deeper skin layers, enhancing the ability to capture subcutaneous blood flow signals in dark-skinned individuals. Simultaneously, the weight of the near-infrared simulated channel is increased, and the weight of the green channel is correspondingly decreased. The pixel values ​​of each channel are multiplied by their corresponding channel weights and then summed to generate the spectrally enhanced video stream.

[0024] In this embodiment, the Spatial Modulation Attention Network (SMOAttention Network) is a deep neural network based on the Transformer architecture, used to extract high-fidelity telephoto plot waveforms from a spectrally enhanced video stream. The SMOAttention Network takes the spectrally enhanced video stream as input and outputs a one-dimensional telephoto plot waveform. The SMOAttention Network includes a backbone network, a pipeline tokenization module, an encoder formed by stacking multiple SMOAttention modules, and a telephoto plot prediction head. The backbone network extracts the spatiotemporal features of the spectrally enhanced video stream through 3D convolution and temporal center-difference convolution. The pipeline tokenization module segments the feature map output by the backbone network into multiple spatiotemporal pipelines and projects each spatiotemporal pipeline as a token of fixed dimension. The SMOAttention module generates a spatial importance map and modulates the attention score of the multi-head self-attention based on the spatial importance map. The telephoto plot prediction head decodes the token sequence output by the encoder into a one-dimensional telephoto plot waveform.

[0025] In this embodiment, the remote photoplethysmography waveform refers to the photoplethysmography signal derived by analyzing the periodic, subtle color changes in the facial skin area caused by heartbeats. This waveform reflects the volume changes of peripheral blood vessels with the cardiac cycle, its frequency components correspond to the heart rate, and its waveform morphology can be used to calculate autonomic nervous system function indicators such as heart rate variability.

[0026] In this embodiment, the remote photoplethysmography (RPG) waveform, temperature distribution feature vector, and skin conductance response feature vector are spatiotemporally aligned. Spatiotemporal alignment includes two steps: timescale unification and event alignment. Timescale unification uses cubic spline interpolation to resample the three signals to the same sampling frequency, ensuring that all three signals have the same number of sampling points on the time axis. Event alignment involves detecting heartbeat events in the RPG waveform, temperature change initiation events in the temperature distribution feature vector, and skin conductance response initiation events in the skin conductance response feature vector. Based on the detected events, the three signals are time-shifted and corrected to align events reflecting the same physiological activity on the time axis.

[0027] In this embodiment, the signal validity is determined based on the real-time signal-to-noise ratio (SNR) of the remote photoplethysmography (TPM) waveform, temperature distribution feature vector, and skin conductance response feature vector. The real-time SNR is calculated as follows: For the TPM waveform, within a preset time window during which the target object's face remains still, a signal amplitude sequence is acquired. The average value of the preset proportion interval with the lowest amplitude value in the signal amplitude sequence is taken as the baseline noise amplitude. The ratio of the peak signal amplitude to the baseline noise amplitude within the current time window is used as the real-time SNR. For the temperature distribution feature vector, multiple sets of facial temperature distribution data are acquired while the target object is at rest. Temperature fluctuation components unrelated to emotional responses are extracted from each set of facial temperature distribution data, and a preset noise template is constructed. The matching degree between the temperature change curve within the current time window and the preset noise template is used as the real-time SNR. For the skin conductance response feature vector, during the resting period when the skin conductance activity signal is unresponsive, the low-frequency fluctuation amplitude of the skin conductance level signal is calculated as the baseline drift amplitude. The ratio of the skin conductance response amplitude to the baseline drift amplitude within the current time window is used as the real-time SNR. The calculated real-time signal-to-noise ratio is compared with a preset threshold. When the real-time signal-to-noise ratio is higher than or equal to the preset threshold, the signal is deemed valid; when the real-time signal-to-noise ratio is lower than the preset threshold, the signal is deemed invalid.

[0028] In this embodiment, the preset threshold is a judgment standard pre-set according to the signal type and application scenario, used to distinguish between valid and invalid signal states. For remote photoplethysmography waveforms, the first preset threshold is determined based on the typical signal-to-noise ratio distribution of the signal in the resting state, with a value ranging from 3 dB to 5 dB. For temperature distribution feature vectors, the second preset threshold is determined based on the distribution of the matching degree between the temperature change curve and the noise template under normal conditions, with a value ranging from 2.5 dB to 4 dB. For skin conductance response feature vectors, the third preset threshold is determined based on the distribution of the ratio of skin conductance response amplitude to baseline drift amplitude under normal conditions, with a value ranging from 2 dB to 3 dB. When the real-time signal-to-noise ratio of the signal is lower than the corresponding preset threshold, the signal is determined to be in an invalid state; when the real-time signal-to-noise ratio is higher than or equal to the preset threshold, the signal is determined to be valid.

[0029] In this embodiment, for failed signals with a real-time signal-to-noise ratio below a preset threshold, alternative features for the failed signals are constructed based on the physiological correlation characteristics of other valid signals. When the remote photoplethysmography (TPM) waveform fails, temperature fluctuation components related to the heart rate cycle are extracted from the temperature distribution feature vector, and time-domain features related to autonomic nervous activity are extracted from the skin conductance response feature vector. These feature components are input into a pre-trained long short-term memory (LSTM) network, which outputs an estimated value of the TPM waveform as an alternative feature. When the temperature distribution feature vector fails, heart rate variability features are extracted from the TPM waveform, and sympathetic nervous activity features are extracted from the skin conductance response feature vector. These feature components are input into a pre-trained LSM network, which outputs an estimated value of the temperature distribution feature vector as an alternative feature. When the skin conductance response feature vector fails, heart rate variability features are extracted from the TPM waveform, and vasomotor activity features are extracted from the temperature distribution feature vector. These feature components are input into a pre-trained LSM network, which outputs an estimated value of the skin conductance response feature vector as an alternative feature.

[0030] In this embodiment, the multimodal signal pattern recognition network is a deep neural network used to fuse multiple physiological signals for emotion recognition. The input to the multimodal signal pattern recognition network includes remote photoplethysmography waveforms, temperature distribution feature vectors, skin conductance response feature vectors, and surrogate features; the output is a fused physiological feature vector. The multimodal signal pattern recognition network includes a signal-to-noise ratio (SNR) estimation module, a spatiotemporal alignment module, and a complementary fusion module. The SNR estimation module calculates the real-time SNR of each of the three signals. The spatiotemporal alignment module unifies the three signals to a standardized time coordinate system. The complementary fusion module dynamically determines the fusion weights based on the real-time SNR and performs weighted fusion of the three signals.

[0031] In this embodiment, the remote photoplethysmography (RPG) waveform, temperature distribution feature vector, skin conductance response (SCR) feature vector, and substitute features are input into a multimodal signal pattern recognition network (MMC). The MMC dynamically determines the fusion weights based on the effective signal-to-noise ratio (SNR) of the RPG waveform, temperature distribution feature vector, and SCR feature vector, and performs weighted fusion to output a fused physiological feature vector. For signals determined to be valid, their real-time SNR is converted into fusion weights using a monotonically increasing mapping function, giving signals with higher SNRs higher fusion weights. For signals determined to be invalid, a preset substitute weight is used as the fusion weight, which is greater than the fusion weight obtained by converting the real-time SNR of the invalid signal using the mapping function. The RPG waveform, temperature distribution feature vector, and SCR feature vector are weighted and concatenated according to their corresponding fusion weights to generate a fused physiological feature vector. When a signal has a substitute feature, the substitute feature replaces the original signal in the weighted fusion.

[0032] In this embodiment, the crowd stratification verification network is a deep neural network used to achieve adaptive emotion recognition among crowds. The input to the crowd stratification verification network is a fused physiological feature vector, and the output is the final emotion state. The crowd stratification verification network includes a crowd candidate set generation layer and an individual baseline correction layer. The crowd candidate set generation layer is used to match initial candidate emotion states from multiple preset crowd physiological feature templates based on the target object's skin color type and age range. The individual baseline correction layer is used to correct the initial candidate emotion states based on the target object's individual baseline physiological characteristics.

[0033] In this embodiment, the preset population physiological feature template is a statistical model constructed based on diverse populations with different skin colors and age ranges. Each template corresponds to a population subtype, which is divided according to a combination of skin color type and age range. Each population physiological feature template contains statistical distribution parameters of the fused physiological feature vectors corresponding to each emotional state under that population subtype. The statistical distribution parameters include the mean vector and covariance matrix corresponding to each emotional state. The template is constructed as follows: a large amount of multimodal physiological data of target objects under that population subtype in resting state and under various emotional induction states are collected; fused physiological feature vectors are extracted from the multimodal physiological data; statistical analysis is performed on the fused physiological feature vectors under each emotional state; and the mean vector and covariance matrix are calculated as statistical distribution parameters.

[0034] In this embodiment, the fused physiological feature vector is input into a population-stratified validation network. The network matches initial candidate emotional states from multiple pre-defined population physiological feature templates based on the target object's skin color type and age range. Specifically, firstly, the target object's population subtype is determined based on the combination of its skin color type and age range. Then, population physiological feature templates corresponding to the target object's population subtype are extracted from a pre-defined template library. The Mahalanobis distance between the fused physiological feature vector and the statistical distribution parameters corresponding to each emotional state in the template is calculated. The probability value of the fused physiological feature vector belonging to each emotional state is determined based on the Mahalanobis distance as the initial confidence level. The initial confidence levels are sorted from high to low to obtain the initial candidate emotional states.

[0035] In this embodiment, the individual baseline physiological characteristics of the target object refer to the multimodal physiological signal characteristics of the target object in a resting state. The resting state is defined as the state in which the facial movement amplitude of the target object is less than a movement threshold and there is no response in skin electrical activity within a preset duration. The extraction method for the individual baseline physiological characteristics is as follows: when the target object is in a resting state, baseline multimodal physiological signals are collected for a preset duration. The baseline multimodal physiological signals include facial video streams, infrared thermal imaging sequences, and skin electrical activity signals in the resting state. A fused physiological feature vector is extracted from the baseline multimodal physiological signals as the individual baseline fused physiological feature vector.

[0036] In this embodiment, the initial candidate emotional states are corrected based on the individual baseline physiological characteristics of the target object, and the final emotional state is output. Specifically, the basic offset difference between the individual baseline fused physiological feature vector and the resting state standard vector in the physiological feature template of the target object's subtype is calculated. An offset correction factor is generated based on the basic offset difference, and the initial confidence of each emotional state in the initial candidate emotional states is multiplied by the offset correction factor to obtain the corrected confidence. The emotional state with the highest corrected confidence is output as the final emotional state. Further, the near-infrared simulation channel is constructed as follows: The average pixel value of multiple skin regions in the facial video stream of the target object is identified, and the skin color type of the target object is determined based on the average pixel value. The skin color types include light skin color, medium skin color and dark skin color. When the skin tone type is light skin tone, the green channel weight of the facial video stream is set as the first weight value, and the red and blue channel weights are set as the second weight values, with the first weight value being greater than the second weight value; When the skin tone type is dark skin tone, the weighted difference between the pixel values ​​of the red channel and the pixel values ​​of the blue channel in the facial video stream is used as the pixel value of the near-infrared analog channel. The weight of the near-infrared analog channel in the facial video stream is set as the third weight value, and the weight of the green channel in the facial video stream is set as the fourth weight value. The third weight value is greater than the fourth weight value. The pixel values ​​of each channel of the facial video stream are multiplied by their corresponding channel weights and then summed to generate a spectrally enhanced video stream.

[0037] In this embodiment, multiple skin regions in the facial video stream include the forehead, cheeks, and chin. The forehead region is located between the eyebrows and the hairline; the skin in this area is relatively flat and less affected by facial expressions. The cheek region is located between the eyes and the jawbone; this area has a large skin area and is rich in blood vessels. The chin region is located between the lower lip and the lower edge of the jawbone; the skin in this area is relatively stable and less affected by occlusion. Selecting these three regions as skin regions for skin color analysis can comprehensively reflect the overall skin color characteristics of the target subject's face.

[0038] In this embodiment, the average pixel values ​​of multiple skin regions in the facial video stream of the target object are identified. Specifically, firstly, the boundary coordinates of the forehead, cheek, and chin regions are located in each frame of the facial video stream using a facial landmark detection algorithm. Then, the red, green, and blue channel values ​​of all pixels within each skin region are extracted, and the average value of each channel is calculated to obtain the average pixel value of the red, green, and blue channels for each skin region. The average pixel values ​​of the three skin regions are then weighted and averaged to obtain a comprehensive average pixel value representing the overall skin tone of the target object's face.

[0039] In this embodiment, the skin tone type of the target object is determined based on the average pixel value. Specifically, the red, green, and blue channel components in the comprehensive average pixel value are normalized, and the ratios of the red channel component to the green channel component and the red channel component to the blue channel component are calculated. When the ratio of the red channel component to the green channel component is less than a first ratio threshold and the ratio of the red channel component to the blue channel component is less than a second ratio threshold, the skin tone type is determined to be light. When the ratio of the red channel component to the green channel component is greater than a third ratio threshold and the ratio of the red channel component to the blue channel component is greater than a fourth ratio threshold, the skin tone type is determined to be dark. When the ratio is in the range between light and dark skin tone types, the skin tone type is determined to be medium skin tone.

[0040] In this embodiment, when the skin tone type is light, the green channel weight of the facial video stream is set to a first weight value, and the red and blue channel weights are set to second weight values, with the first weight value being greater than the second weight value. For light skin tones, since the skin surface is most sensitive to changes in the absorption and reflection of green light, it can most effectively capture color changes caused by subcutaneous blood flow. Therefore, the green channel weight is set to a relatively high first weight value, ranging from 0.6 to 0.8. Meanwhile, the signal components in the red and blue channels are relatively weak and easily affected by ambient light; therefore, the red and blue channel weights are set to relatively low second weight values, ranging from 0.1 to 0.2.

[0041] In this embodiment, when the skin tone type is dark, the weighted difference between the pixel values ​​of the red channel and the blue channel in the facial video stream is used as the pixel value of the near-infrared simulated channel. The weight of the near-infrared simulated channel in the facial video stream is set as the third weight value, and the weight of the green channel in the facial video stream is set as the fourth weight value, with the third weight value being greater than the fourth weight value. For dark skin tones, the strong absorption of green light by melanin leads to severe attenuation of the green channel signal, making it difficult for the traditional green channel to effectively capture subcutaneous blood flow signals. The near-infrared simulated channel, by calculating the weighted difference between the red and blue channels, can simulate the penetration characteristics of near-infrared light, effectively penetrating deeper skin layers and enhancing the ability to capture subcutaneous blood flow signals in dark-skinned individuals. Therefore, the weight of the near-infrared simulated channel is set to a higher third weight value, ranging from 0.7 to 0.9, and the weight of the green channel is set to a lower fourth weight value, ranging from 0.1 to 0.3.

[0042] In this embodiment, the pixel values ​​of each channel of the facial video stream are multiplied by their corresponding channel weights and then summed to generate a spectrally enhanced video stream. Specifically, for each frame of the facial video stream, the pixel values ​​of the red, green, blue, and near-infrared analog channels are extracted. Each channel pixel value is multiplied by its corresponding channel weight, and then the products of all channels are summed to obtain the spectrally enhanced pixel values ​​of that frame. This process is repeated for each frame to generate a complete spectrally enhanced video stream. In this way, the spectrally enhanced video stream can adaptively highlight the spectral channels most effective for physiological signal extraction and suppress noise channels based on the skin color type of the target object, thereby improving the signal-to-noise ratio of subsequent remote photoplethysmography (RPG) waveform extraction. Further, a spatially modulated attention network obtains the RPG waveform, including: The spectral-enhanced video stream is divided into multiple time windows, each containing a preset number of consecutive video frames; Spatiotemporal feature extraction is performed on video frames within each waveform extraction time window. Spatiotemporal feature extraction includes extracting spatial features through three-dimensional convolution and extracting color change difference features between adjacent frames through temporal center difference convolution. Temporal center difference convolution enhances the ability to capture periodic changes in skin color by calculating the feature difference between the current frame and adjacent frames. The extracted spatiotemporal features are segmented into multiple spatiotemporal channels, and each spatiotemporal channel corresponds to the feature sequence of a local region of the face within the waveform extraction time window; A spatial importance map is generated based on the feature sequences of all spatiotemporal pipelines. Regions with high weight values ​​in the spatial importance map correspond to skin areas in the spectral enhanced video stream that are fully exposed, have small motion amplitude, and are unobstructed. Attention weighting is applied to the feature sequences of the spatiotemporal pipeline based on the spatial importance map, so that regions with high weight values ​​in the spatial importance map receive higher attention weights. The attention-weighted spatiotemporal pipeline feature sequence is decoded into a one-dimensional remote photoplethysmography waveform.

[0043] In this embodiment, the preset frame number refers to the number of consecutive video frames contained within each time window. This number is predetermined based on the frequency characteristics of the telephoto piezometry waveform and the video acquisition frame rate. The preset frame number is set to 160 frames. When the video acquisition frame rate is 30 frames per second, each time window corresponds to approximately 5.3 seconds of continuous video content. This duration can cover multiple complete cardiac cycles, ensuring that the extracted telephoto piezometry waveform contains sufficient heartbeat information for subsequent analysis.

[0044] In this embodiment, the spectral enhancement video stream is divided into multiple waveform extraction time windows, each containing a preset number of consecutive video frames. A sliding window approach is used to divide the spectral enhancement video stream, with a fixed window length of the preset number of frames and a window step size set to half of the preset number of frames. This sliding window approach ensures overlapping areas between adjacent waveform extraction time windows, guaranteeing information continuity in the temporal dimension and preventing information loss due to window boundaries.

[0045] In this embodiment, spatiotemporal feature extraction is performed on video frames within each waveform extraction time window. This extraction includes extracting spatial features through 3D convolution and extracting color change difference features between adjacent frames through temporal center-difference convolution. Temporal center-difference convolution enhances the ability to capture periodic changes in skin color by calculating the feature difference between the current frame and adjacent frames. The 3D convolution uses a 5x5x3 kernel layer to extract texture and structural features of the facial skin region in the spatial dimension and temporal change features between video frames in the temporal dimension. The temporal center-difference convolution uses a 3x1x1 kernel layer to calculate the average difference between the features of the current frame and the features of the previous and next frames, enhancing the ability to capture subtle periodic changes in skin color while suppressing interference from static backgrounds and slowly changing ambient lighting.

[0046] In this embodiment, the extracted spatiotemporal features are segmented into multiple spatiotemporal pipelines. Each spatiotemporal pipeline corresponds to the feature sequence of a local region of the face within the waveform extraction time window. The size of the spatiotemporal pipeline is set to 4 frames in the time dimension and 4 pixels by 4 pixels in the spatial dimension. By segmenting the spatiotemporal features into multiple spatiotemporal pipelines, the global facial features are decomposed into independent feature sequences of multiple local regions. Each spatiotemporal pipeline independently represents the spatiotemporal change pattern of the corresponding local facial region within the time window. This segmentation method enables the subsequent attention mechanism to perform differentiated processing on different facial regions.

[0047] In this embodiment, a spatial importance map is generated based on the feature sequences of all spatiotemporal pipelines. Regions with high weight values ​​in the spatial importance map correspond to skin areas in the spectrally enhanced video stream that are well-exposed, have minimal motion, and are unobstructed. The spatial importance map is generated using a learnable spatial modulator. The spatial modulator first performs average pooling on the feature sequences of all spatiotemporal pipelines to obtain a global feature vector. Then, it passes this global feature vector through a first fully connected layer, an activation function layer, and a second fully connected layer to generate a weight vector that matches the spatial dimension of the spatiotemporal pipelines. Finally, the weight vector is reshaped into a two-dimensional spatial importance map corresponding to the facial spatial layout. During training, the spatial modulator learns that regions with high weight values ​​correspond to skin areas with good signal quality, such as the forehead and cheeks—areas with well-exposed skin and rich blood vessels—while regions with low weight values ​​correspond to areas with poor signal quality, such as areas where glasses are worn, areas obscured by hair, or areas with large motion.

[0048] In this embodiment, attention weighting is applied to the feature sequences of the spatiotemporal pipeline based on the spatial importance map, giving higher attention weights to regions with high weight values ​​in the spatial importance map. The weight value corresponding to each spatial location in the spatial importance map is multiplied by the feature sequence of the spatiotemporal pipeline at that location, thus enhancing the feature sequences of regions with high weight values ​​and suppressing the feature sequences of regions with low weight values. In this way, the model can adaptively focus on facial regions with high signal quality while suppressing the negative impact of regions with poor signal quality on feature extraction.

[0049] In this embodiment, the attention-weighted spatiotemporal pipeline feature sequence is decoded into a one-dimensional remote photoplethysmography (RP-PGA) waveform. First, the attention-weighted spatiotemporal pipeline feature sequence is reconstructed back into a three-dimensional feature map. Then, the temporal resolution is progressively increased from 40 frames to 80 frames and then to 160 frames through two cascaded upsampling modules. The upsampling modules employ trilinear interpolation and 3x1x1 three-dimensional convolutions. Spatially, spatial pooling is performed using a learnable weighted averaging scheme to convert the three-dimensional feature map into a one-dimensional temporal feature sequence. Finally, the one-dimensional temporal feature sequence is projected onto a single channel through a 1x1 one-dimensional convolutional layer to obtain a 160-frame one-dimensional R-PGA waveform. Further, the real-time signal-to-noise ratio (SNR) of the R-PGA waveform, temperature distribution feature vector, and skin conductance response feature vector is calculated, including: The spatiotemporally aligned remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector are used as inputs; Within a preset time window in which the target object's face remains still, the signal amplitude sequence of the remote photoplethysmography waveform is acquired. The average value of the preset ratio range with the lowest amplitude value in the signal amplitude sequence is taken as the baseline noise amplitude. The ratio of the peak signal amplitude to the baseline noise amplitude within the current time window is taken as the first real-time signal-to-noise ratio of the remote photoplethysmography waveform. Multiple sets of facial temperature distribution data were collected when the target object was in a resting state. Temperature fluctuation components unrelated to emotional response were extracted from each set of facial temperature distribution data. A preset noise template was constructed, and the matching degree between the temperature change curve in the current time window and the preset noise template was used as the second real-time signal-to-noise ratio of the temperature distribution feature vector. During the resting period when there is no response to the skin conductance signal, the low-frequency fluctuation amplitude of the skin conductance level signal is calculated as the baseline drift amplitude, and the ratio of the skin conductance response amplitude to the baseline drift amplitude within the current time window is used as the third real-time signal-to-noise ratio of the skin conductance response feature vector.

[0050] In this embodiment, the spatiotemporally aligned remote photoplethysmography (TPM) waveform, temperature distribution feature vector, and skin conductance response feature vector are used as inputs. The spatiotemporally aligned TPM waveform, temperature distribution feature vector, and skin conductance response feature vector have the same number of sampling points on the time axis, and events reflecting the same physiological activity are aligned on the time axis. Using these three sets of signals as input data for subsequent signal-to-noise ratio (SNR) calculations ensures the temporal consistency of the SNR assessment.

[0051] In this embodiment, the signal-to-noise ratio (SNR) calculation time window for keeping the target object's face still refers to a fixed duration selected within a time interval during which the target object exhibits no significant head movement or facial expression changes. The SNR calculation time window is set to 10 seconds, which is sufficient to acquire remote photoplethysmography (PPG) waveform signals for multiple cardiac cycles. Within this time window, the target object's face remains still, and the amplitude of head movement and facial expression changes are both less than a preset motion threshold, thereby eliminating the interference of motion artifacts on signal amplitude analysis.

[0052] In this embodiment, acquiring the signal amplitude sequence of the remote photoplethysmography (TPM) waveform refers to recording the amplitude value of each sampling point of the TPM waveform in chronological order within a preset time window in which the target object's face remains still, forming an amplitude sequence that changes over time. The signal amplitude sequence reflects the complete waveform morphology of the TPM waveform within the time window, including the periodic peaks and troughs caused by the heartbeat.

[0053] In this embodiment, the average value of the lowest percentage interval in the signal amplitude sequence is used as the baseline noise amplitude. The preset percentage interval is set to 10% of the lowest amplitude values ​​in the signal amplitude sequence. All sampling points in the signal amplitude sequence are sorted in ascending order of amplitude value, and the first 10% of sampling points after sorting are selected. The average amplitude of these sampling points is calculated, and this average value is used as the baseline noise amplitude. The baseline noise amplitude represents the noise level during cardiac diastole or when the signal is weakest, and is an important benchmark for calculating the signal-to-noise ratio.

[0054] In this embodiment, the peak signal amplitude within the current time window refers to the maximum amplitude value extracted from the signal amplitude sequence of the remote photoplethysmography waveform within the same time window as the baseline noise amplitude. The peak signal amplitude typically corresponds to the peak systolic blood pressure during the cardiac cycle, reflecting the maximum change in peripheral blood volume during cardiac contraction.

[0055] In this embodiment, multiple sets of facial temperature distribution data are collected while the target object is in a resting state. A resting state refers to a stable state where the target object is free from significant emotional fluctuations, strenuous physical activity, and sudden changes in ambient temperature. After the target object enters a resting state, multiple sets of facial temperature distribution data are continuously collected using an infrared thermal imager. Each set of data corresponds to a facial temperature distribution image at a specific time point, and the acquisition duration for each set of data is consistent with the length of the time window.

[0056] In this embodiment, the target object is determined to be in a resting state by detecting that the facial movement amplitude is less than a movement threshold and there is no response in the electrodermal activity (EDA) signal within a preset duration. The preset duration is set to 30 seconds. The facial movement amplitude is calculated by analyzing the displacement changes of key facial points in the facial video stream. When the displacement changes of all key points are less than the movement threshold, the facial movement amplitude is determined to meet the resting condition. No response in the EDA signal means that the skin conductance level does not show a complete response waveform that rises above the preset amplitude threshold from the baseline value and then returns to the baseline value within the preset duration. When both the facial movement amplitude and the EDA signal meet the resting condition, the target object is determined to be in a resting state.

[0057] In this embodiment, temperature fluctuation components unrelated to emotional responses are extracted from facial temperature distribution data in each group to construct a preset noise template. Temperature fluctuation components unrelated to emotional responses include temperature changes caused by factors such as ambient temperature variations, device noise, and respiratory rhythm. By statistically analyzing facial temperature distribution data from multiple groups under resting conditions, common components of temperature changes in each group are extracted, local temperature change patterns related to specific emotions are removed, and the remaining temperature fluctuation components are used as the preset noise template. The preset noise template represents the normal temperature fluctuation range when there is no emotional response.

[0058] In this embodiment, the temperature change curve within the current time window refers to the curve showing the change of temperature values ​​extracted from the temperature distribution feature vector over time within a time window of the same length as the preset noise template. This curve reflects the dynamic changes in the forehead temperature, nose tip temperature, and cheek temperature of the target object within the current time window.

[0059] In this embodiment, the matching degree between the temperature change curve within the current time window and the preset noise template is calculated as follows: The temperature change curve and the preset noise template are compared point-by-point to obtain a difference sequence. The root mean square error (RMSE) of the difference sequence is calculated, and the reciprocal of the RMS error is used as the matching degree. A higher matching degree indicates that the current temperature change curve is closer to the normal temperature fluctuation pattern without emotional reaction; that is, the current temperature change contains fewer emotionally relevant components, and the signal-to-noise ratio (SNR) is lower. Conversely, a lower matching degree indicates that the current temperature change contains more emotionally relevant components, and the signal-to-noise ratio (SNR) is higher.

[0060] In this embodiment, the resting period of no response to the skin conductance signal refers to the time period during which the skin conductance level signal does not exhibit a complete response waveform that rises above a preset amplitude threshold from the baseline value and then returns to the baseline value. During this period, the skin conductance level remains relatively stable, with no obvious skin conductance response caused by sympathetic nerve activation. The duration of the resting period is set to at least 10 seconds to ensure sufficient data for baseline drift amplitude calculation.

[0061] In this embodiment, the low-frequency fluctuation amplitude of the skin conductance level signal is calculated as the baseline drift amplitude. During the resting period when the skin conductance signal is unresponsive, the skin conductance level signal is low-pass filtered to retain components with frequencies below 0.1 Hz. The difference between the maximum and minimum values ​​of the filtered signal is used as the baseline drift amplitude. The baseline drift amplitude reflects the degree of slow fluctuation of the skin conductance level signal in the unresponsive state and is an important benchmark for calculating the signal-to-noise ratio.

[0062] In this embodiment, the skin conductance response amplitude within the current time window refers to the amplitude value of the skin conductance response extracted from the skin conductance response feature vector within a time window of the same length as the baseline drift amplitude calculation. The skin conductance response amplitude is defined as the difference between the maximum value reached by the skin conductance level from the baseline value and the baseline value. This amplitude reflects the intensity of sympathetic nerve activation and is an important indicator for assessing emotional arousal. Further, the multimodal signal pattern recognition network dynamically determines the fusion weights and performs weighted fusion based on the effective signal-to-noise ratio of the remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector, including: The first real-time signal-to-noise ratio of the remote photoplethysmography waveform is converted into a first fusion weight through a first mapping function. The second real-time signal-to-noise ratio of the temperature distribution feature vector is converted into a second fusion weight through a second mapping function. The third real-time signal-to-noise ratio of the skin conductance response feature vector is converted into a third fusion weight through a third mapping function. The first, second, and third mapping functions are monotonically increasing functions, so that signals with high signal-to-noise ratios receive higher fusion weights. The remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector are weighted and concatenated according to the first fusion weight, the second fusion weight, and the third fusion weight to generate a fused physiological feature vector.

[0063] In this embodiment, the first mapping function is used to convert the first real-time signal-to-noise ratio (SNR) of the remote photoplethysmography (RPN) waveform into a first fusion weight. The first mapping function is a monotonically increasing function, specifically an sigmoid function. When the first real-time SNR is lower than a first lower SNR threshold, the output of the first mapping function approaches the first lower fusion weight value. When the first real-time SNR is higher than a first upper SNR threshold, the output of the first mapping function approaches the first upper fusion weight value. When the first real-time SNR is between the first lower and upper SNR thresholds, the output of the first mapping function monotonically increases with the increase of the first real-time SNR. Through the first mapping function, remote photoplethysmography waveforms with low SNR receive lower fusion weights, while remote photoplethysmography waveforms with high SNR receive higher fusion weights.

[0064] In this embodiment, the first real-time signal-to-noise ratio (SNR) of the remote photoplethysmography (RPN) waveform is converted into a first fusion weight using a first mapping function. After calculating the first real-time SNR, this value is input into the first mapping function. The first mapping function outputs the corresponding first fusion weight based on preset first SNR lower threshold, first SNR upper threshold, first fusion weight lower limit, and first fusion weight upper limit. The first fusion weight is used to control the contribution of the remote PPN waveform in subsequent weighted fusion.

[0065] In this embodiment, the second mapping function is used to convert the second real-time signal-to-noise ratio (SNR) of the temperature distribution feature vector into a second fusion weight. The second mapping function is a monotonically increasing function, specifically an sigmoid function. When the second real-time SNR is lower than the second SNR lower threshold, the output of the second mapping function approaches the second fusion weight lower threshold. When the second real-time SNR is higher than the second SNR upper threshold, the output of the second mapping function approaches the second fusion weight upper threshold. When the second real-time SNR is between the second SNR lower threshold and the second SNR upper threshold, the output of the second mapping function monotonically increases with the increase of the second real-time SNR. Through the second mapping function, temperature distribution feature vectors with low SNR receive lower fusion weights, and temperature distribution feature vectors with high SNR receive higher fusion weights.

[0066] In this embodiment, the second real-time signal-to-noise ratio (SNR) of the temperature distribution feature vector is converted into a second fusion weight using a second mapping function. After calculating the second real-time SNR, this value is input into the second mapping function. The second mapping function outputs the corresponding second fusion weight based on preset lower and upper threshold values ​​for the second SNR, as well as the lower and upper limits of the second fusion weight. The second fusion weight is used to control the contribution of the temperature distribution feature vector in subsequent weighted fusion.

[0067] In this embodiment, the third mapping function is used to convert the third real-time signal-to-noise ratio (SNR) of the skin electrical response feature vector into a third fusion weight. The third mapping function is a monotonically increasing function, specifically an sigmoid function. When the third real-time SNR is lower than the third SNR lower threshold, the output of the third mapping function approaches the third fusion weight lower threshold. When the third real-time SNR is higher than the third SNR upper threshold, the output of the third mapping function approaches the third fusion weight upper threshold. When the third real-time SNR is between the third SNR lower threshold and the third SNR upper threshold, the output of the third mapping function monotonically increases with the increase of the third real-time SNR. Through the third mapping function, skin electrical response feature vectors with low SNR receive lower fusion weights, and skin electrical response feature vectors with high SNR receive higher fusion weights.

[0068] In this embodiment, the third real-time signal-to-noise ratio (SNR) of the skin conductance response feature vector is converted into a third fusion weight using a third mapping function. After calculating the third real-time SNR, this value is input into the third mapping function. The third mapping function outputs the corresponding third fusion weight based on preset third SNR lower threshold, third SNR upper threshold, third fusion weight lower limit, and third fusion weight upper limit. The third fusion weight is used to control the contribution of the skin conductance response feature vector in subsequent weighted fusion.

[0069] In this embodiment, the remote photoplethysmography (TPM) waveform, temperature distribution feature vector, and skin conductance response feature vector are weighted and concatenated according to a first fusion weight, a second fusion weight, and a third fusion weight to generate a fused physiological feature vector. The specific weighted concatenation method is as follows: the TPM waveform is multiplied by the first fusion weight, the temperature distribution feature vector is multiplied by the second fusion weight, and the skin conductance response feature vector is multiplied by the third fusion weight. Then, the three weighted vectors are concatenated end-to-end along their feature dimensions to form a new vector with a dimension equal to the sum of the dimensions of the three original vectors. This new vector is the fused physiological feature vector. Through weighted concatenation, signals with high signal-to-noise ratio (SNR) occupy a larger proportion in the fused physiological feature vector, while signals with low SNR occupy a smaller proportion, thereby achieving SNR-adaptive multimodal signal fusion. Furthermore, it also includes determining that the remote photoplethysmography waveform is in a failed state when the first real-time signal-to-noise ratio of the remote photoplethysmography waveform is lower than the first preset threshold, constructing alternative features of the remote photoplethysmography waveform based on the temperature distribution feature vector and the skin conductance response feature vector, and using a preset alternative weight as the fusion weight of the remote photoplethysmography waveform. The alternative weight is greater than the fusion weight transformed by the mapping function of the first real-time signal-to-noise ratio, and does not exceed the maximum fusion weight of the effective signal. When the second real-time signal-to-noise ratio of the temperature distribution feature vector is lower than the second preset threshold, the temperature distribution feature vector is determined to be in a failed state, and alternative features of the temperature distribution feature vector are constructed based on the remote photoplethysmography waveform and the skin conductance response feature vector. When the third real-time signal-to-noise ratio of the skin electric response feature vector is lower than the third preset threshold, the skin electric response feature vector is determined to be in a failed state. Alternative features of the skin electric response feature vector are constructed based on the remote photoplethysmography waveform and temperature distribution feature vector. The alternative features are constructed by extracting feature components that are physiologically related to the failure signal from the effective signal, and mapping the feature components to the estimated value of the failure signal through a pre-trained regression model.

[0070] In this embodiment, the first preset threshold is a critical value used to determine whether the remote photoplethysmography (PPG) waveform is in a failed state. The first preset threshold is determined based on the signal-to-noise ratio (SNR) distribution of the PPG waveform under normal acquisition conditions. When the first real-time SNR of the PPG waveform is lower than the first preset threshold, it indicates that the signal is severely affected by noise and cannot reliably extract heart rate-related features. At this time, the PPG waveform is determined to be in a failed state.

[0071] In this embodiment, alternative features for remote photoplethysmography (TPM) waveforms are constructed based on temperature distribution feature vectors and skin conductance response feature vectors. When the TPM waveform fails, a temperature fluctuation component related to the heart rate cycle is extracted from the temperature distribution feature vector. This component reflects the temperature changes caused by the contraction and relaxation of subcutaneous blood vessels during the heart cycle. Simultaneously, temporal features related to autonomic nervous activity are extracted from the skin conductance response feature vector. These features reflect the activation state of the sympathetic nervous system. The temperature fluctuation component and the temporal features of autonomic nervous activity are concatenated to form a feature vector. This feature vector is input into a pre-trained Long Short-Term Memory (LSTM) network, which outputs an estimated value of the TPM waveform as the alternative feature.

[0072] In this embodiment, the preset substitution weight is a preset value used to replace the fusion weight of the failed signal. The preset substitution weight is greater than the fusion weight obtained by converting the real-time signal-to-noise ratio of the failed signal through a mapping function. When the remote photoplethysmography waveform fails, the preset substitution weight is used as the fusion weight of the signal to ensure that the substitution features corresponding to the failed signal can make a sufficient contribution in the weighted fusion, avoiding the signal being completely ignored due to an excessively low fusion weight. The preset substitution weight ranges from 0.3 to 0.5.

[0073] In this embodiment, the second preset threshold is a critical value used to determine whether the temperature distribution feature vector is in a failed state. The second preset threshold is determined based on the signal-to-noise ratio distribution of the temperature distribution feature vector under normal acquisition conditions. When the second real-time signal-to-noise ratio of the temperature distribution feature vector is lower than the second preset threshold, it indicates that the signal is severely affected by changes in ambient temperature or equipment noise, and cannot reliably reflect emotion-related temperature changes. At this time, the temperature distribution feature vector is determined to be in a failed state.

[0074] In this embodiment, a substitute feature for the temperature distribution feature vector is constructed based on the remote photoplethysmography (TPM) waveform and the skin conductance response feature vector. When the temperature distribution feature vector fails, heart rate variability (HRV) features, which reflect the overall activity state of the autonomic nervous system, are extracted from the TPM waveform. Sympathetic activation features, which reflect changes in emotional arousal, are extracted from the skin conductance response feature vector. The HRV features and sympathetic activation features are concatenated to form a feature vector, which is then input into a pre-trained long short-term memory (LSTM) network. The LSM network outputs an estimate of the temperature distribution feature vector as the substitute feature.

[0075] In this embodiment, the third preset threshold is a critical value used to determine whether the skin conductance response feature vector is in a failed state. The third preset threshold is determined based on the signal-to-noise ratio distribution of the skin conductance response feature vector under normal acquisition conditions. When the third real-time signal-to-noise ratio of the skin conductance response feature vector is lower than the third preset threshold, it indicates that the signal is severely interfered with by hand movement or poor electrode contact and cannot reliably reflect sympathetic nerve activity. At this time, the skin conductance response feature vector is determined to be in a failed state.

[0076] In this embodiment, alternative features for the skin conductance response feature vector are constructed based on the remote photoplethysmography (TP) waveform and the temperature distribution feature vector. When the skin conductance response feature vector fails, heart rate variability features, which reflect the overall activity state of the autonomic nervous system, are extracted from the TP waveform. Vasomotor activity features, which reflect changes in subcutaneous blood flow, are extracted from the temperature distribution feature vector. The heart rate variability features and vasomotor activity features are concatenated to form a feature vector, which is then input into a pre-trained long short-term memory (LSTM) network. The LSM network outputs an estimated value of the skin conductance response feature vector as the alternative feature.

[0077] In this embodiment, the alternative features are constructed as follows: feature components that are physiologically correlated with the failed signal are extracted from the valid signal, and the feature components are mapped to the estimated value of the failed signal through a pre-trained regression model. Specifically, firstly, it is identified which signals are currently in a valid state, and feature components that are physiologically correlated with the failed signal are extracted from the valid signal. Physiological correlation means that the valid signal and the failed signal are related in terms of physiological mechanisms. For example, both remote photoplethysmography waveforms and temperature distribution feature vectors are affected by cardiovascular activity, and both skin conductance response feature vectors and remote photoplethysmography waveforms are regulated by the autonomic nervous system. Then, the extracted feature components are input into a pre-trained long short-term memory network, and the long short-term memory network outputs the estimated value of the failed signal. This estimated value is used as an alternative feature in the subsequent weighted fusion.

[0078] In this embodiment, the feature components that are physiologically correlated with the failure signal refer to the feature components extracted from the valid signal that are intrinsically related to the failure signal in terms of physiological mechanisms. For example, the physiological correlation between the remote photoplethysmography waveform and the temperature distribution feature vector is reflected in the fact that both are affected by subcutaneous vasomotor activity, and changes in blood volume caused by heartbeats simultaneously lead to periodic changes in skin temperature and photoplethysmography signals. The physiological correlation between the remote photoplethysmography waveform and the skin conductance response feature vector is reflected in the fact that both are regulated by the autonomic nervous system, and sympathetic nerve activation simultaneously affects heart rate and skin conductance levels. The physiological correlation between the temperature distribution feature vector and the skin conductance response feature vector is reflected in the fact that sympathetic nerve activation during emotional arousal simultaneously leads to changes in skin temperature and skin conductance levels.

[0079] In this embodiment, the pre-trained regression model refers to a long short-term memory network pre-trained based on a large amount of training data. The input to this regression model is the physiologically relevant feature component extracted from the valid signal, and the output is an estimate of the invalid signal. The training method for the regression model is as follows: Collect all valid training samples containing remote photoplethysmography waveforms, temperature distribution feature vectors, and skin conductance response feature vectors; extract the feature components of each signal from the training samples; use one signal as the prediction target and the feature components of the other two signals as input; train the long short-term memory network to learn the mapping relationship from the input to the prediction target. In this way, a first regression model for predicting remote photoplethysmography waveforms, a second regression model for predicting temperature distribution feature vectors, and a third regression model for predicting skin conductance response feature vectors are trained respectively. Further, the temperature distribution feature vector is extracted from the infrared thermal imaging sequence, including: In each frame of the infrared thermal imaging sequence, the forehead region, nose tip region, and cheek region are located, with the cheek region including the left cheek region and the right cheek region. Calculate the slope of the curve showing the change in average temperature of the forehead region over time within a time window, and use the slope value as the rate of change in forehead temperature. Calculate the difference between the maximum and minimum average temperatures of the nasal tip region within the time window, and use this difference as the nasal tip temperature fluctuation range. Calculate the average temperature difference between the left and right cheek regions within a time window, and use the difference as the cheek temperature asymmetry index.

[0080] In this embodiment, the forehead region, nose tip region, and cheek region are located in each frame of the infrared thermal imaging sequence. The cheek region includes the left cheek region and the right cheek region. Specifically, facial key points are first located in each frame of the infrared thermal imaging sequence using a facial key point detection algorithm, including the center point of the forehead above the brow, the nose tip point, the center point of the left cheek, and the center point of the right cheek. The forehead region is defined by extending a preset radius outward from the center point of the forehead, covering the central area of ​​the forehead. The nose tip region is defined by defining the center point of the nose, covering the tip of the nose. The left cheek region is defined by defining the center point of the left cheek, covering the area from below the left cheekbone to above the left jaw angle. The right cheek region is defined by defining the center point of the right cheek, covering the area from below the right cheekbone to above the right jaw angle. This positioning method ensures consistency in the positioning of each region within each time window.

[0081] In this embodiment, the slope of the curve showing the average temperature of the forehead region changing over time within a time window is calculated, and the slope value is used as the forehead temperature change rate. Specifically, within the time window, the temperature values ​​of all pixels in the forehead region are extracted for each frame of the infrared thermal imaging sequence, and the average temperature of the forehead region is calculated to obtain a sequence showing the average forehead temperature changing over time. A linear fit is then performed on this sequence, and the slope of the fitted line is the forehead temperature change rate. A positive forehead temperature change rate indicates an increase in forehead temperature, while a negative rate indicates a decrease. The forehead temperature change rate reflects the rate of temperature change caused by the vasoconstriction or vasodilation of skin blood vessels due to sympathetic nerve activation during emotional arousal.

[0082] In this embodiment, the difference between the maximum and minimum average temperatures of the nasal tip region within a time window is calculated, and this difference is used as the nasal tip temperature fluctuation amplitude. Specifically, within the time window, the temperature values ​​of all pixels within the nasal tip region are extracted for each frame of the infrared thermal imaging sequence, and the average temperature of the nasal tip region is calculated, resulting in a sequence showing the change of the average nasal tip temperature over time. The maximum and minimum values ​​are then identified from this sequence, and the difference between them is calculated; this difference represents the nasal tip temperature fluctuation amplitude. The nasal tip temperature fluctuation amplitude reflects the intensity of fluctuations in nasal tip skin temperature during emotional changes. Because the nasal tip region is rich in blood vessels and less affected by ambient temperature, its temperature changes are closely related to emotional arousal.

[0083] In this embodiment, the average temperature difference between the left and right cheek regions within a time window is calculated, and this difference is used as the cheek temperature asymmetry index. Specifically, within the time window, the temperature values ​​of all pixels in the left and right cheek regions are extracted for each frame of the infrared thermal imaging sequence. The average temperature of the left and right cheek regions is calculated separately, resulting in sequences showing the average temperature of the left cheek changing over time and sequences showing the average temperature of the right cheek changing over time. The difference between the average temperature of the left and right cheeks at the same time point is calculated, resulting in a difference sequence. The average value of this difference sequence within the time window is calculated, and this average value is the cheek temperature asymmetry index. The cheek temperature asymmetry index reflects the symmetrical change in the temperature distribution on both sides of the face during emotional expression. A positive index indicates that the temperature of the left cheek is higher than that of the right cheek, and a negative index indicates that the temperature of the right cheek is higher than that of the left cheek. Further, a skin conductance response feature vector is extracted from the skin conductance activity signal, including: During the resting period when there is no response to the skin electrical activity signal, the average value of the skin conductance level is calculated as the baseline value of the skin conductance level; The number of skin conductance responses is counted within a time window. Skin conductance response is defined as the process by which the skin conductance level rises from the baseline value of the skin conductance level to the baseline value of the skin conductance level after exceeding a preset amplitude threshold. The ratio of the number of skin conductance responses to the duration of the time window is used as the skin conductance response frequency. For each skin conductance response, the time difference from the start of the response to the recovery to the baseline level of skin conductance is calculated, and the average recovery time of all skin conductance responses is taken as the skin conductance recovery time.

[0084] In this embodiment, the average skin conductivity level is calculated as the baseline value of skin conductivity level. Specifically, during the resting period when there is no response to the skin conductance signal, multiple sampling points of skin conductivity level signal are collected, and the arithmetic mean of all sampling points is calculated. This average value is used as the baseline value of skin conductivity level. The baseline value of skin conductivity level reflects the basic level of skin conductivity of the target object in the current state and is an important reference benchmark for determining whether a skin conductivity response has occurred.

[0085] In this embodiment, the number of skin conductance responses is counted within a time window. Specifically, within a preset time window, all sampling points of the skin conductance level signal are traversed, and all events that meet the definition of skin conductance response are identified and counted. The count result is used as the number of skin conductance responses within the time window. This number reflects the frequency of sympathetic nerve activation in the target object within the time window.

[0086] In this embodiment, skin conductance response is defined as the process by which skin conductance level rises from a baseline value exceeding a preset amplitude threshold and then returns to the baseline value. Specifically, when skin conductance level rises from the baseline value, exceeds the preset amplitude threshold, reaches a peak value, begins to decline, and eventually returns to near the baseline value, the complete process from the starting point of the rise to the return to the baseline value is considered as one skin conductance response. Skin conductance response reflects the transient activation event of the sympathetic nervous system and is closely related to emotional arousal.

[0087] In this embodiment, the baseline value of skin conductivity level refers to the stable value of skin conductivity level in the absence of skin conductivity response. The baseline value of skin conductivity level is calculated by averaging the skin conductivity level signal over a resting period. This value changes slowly over time and is affected by ambient temperature, humidity, and the target subject's basic physiological state. When judging skin conductivity response, the baseline value near the time of response occurrence should be used as the judgment benchmark.

[0088] In this embodiment, the preset amplitude threshold is a critical value used to determine whether a change in skin conductivity level constitutes a skin conductivity response. The preset amplitude threshold is determined based on the fluctuation range of the skin conductivity level signal in the resting state and is set to three times the standard deviation of the skin conductivity level in the resting state. When the increase in skin conductivity level from the baseline value exceeds the preset amplitude threshold, it is determined as the start of a skin conductivity response; when the increase falls back to below the preset amplitude threshold, it is determined as the end of a skin conductivity response.

[0089] In this embodiment, the ratio of the number of skin conductance responses to the duration of the time window is used as the skin conductance response frequency. Specifically, the total number of skin conductance responses occurring within the time window is counted, and this number is divided by the duration of the time window to obtain the average number of skin conductance responses per unit time. This ratio is the skin conductance response frequency. The skin conductance response frequency reflects the frequency of sympathetic nerve activation in the target subject within the time window and is an important indicator for assessing emotional arousal.

[0090] In this embodiment, for each skin conductance response, the time difference from the response start point to recovery to the baseline skin conductance level is calculated, and the average recovery time of all skin conductance responses is taken as the skin conductance recovery time. Specifically, for each skin conductance response, the moment when the skin conductance level rises above the baseline value beyond a preset amplitude threshold is recorded as the response start point, and the moment when the skin conductance level first recovers to the baseline value after falling is recorded as the response end point. The time difference between the response end point and the response start point is calculated as the recovery time of that response. The average recovery time is obtained by summing the recovery times of all skin conductance responses within the time window and dividing by the number of skin conductance responses. This average value is the skin conductance recovery time. Skin conductance recovery time reflects the speed at which the sympathetic nervous system recovers to a resting state after activation. A longer recovery time is usually associated with higher emotional arousal or greater psychological stress. Further, the population stratification validation network matches initial candidate emotional states from multiple preset population physiological characteristic templates based on the target object's skin color type and age range, including: The target population subtype is determined based on the combination of the target population’s skin color type and the target population’s age range. The skin color type includes light skin color type, medium skin color type and dark skin color type, and the age range includes first age range, second age range and third age range. Extract the population physiological feature template corresponding to the population subtype to which the target object belongs from the preset template library. The population physiological feature template contains the statistical distribution parameters of the fused physiological feature vectors corresponding to each emotional state under the population subtype. The statistical distribution parameters include the mean vector and covariance matrix corresponding to each emotional state. The Mahalanobis distance between the fused physiological feature vector and the statistical distribution parameters corresponding to each emotional state is calculated. Based on the Mahalanobis distance, the probability value of the fused physiological feature vector belonging to each emotional state is determined as the initial confidence level, and the initial candidate emotional states are obtained.

[0091] In this embodiment, the target object's population subtype is determined based on a combination of its skin tone type and age range. Specifically, the skin tone type and age range are combined using a Cartesian product to form multiple mutually exclusive population subtypes. When the skin tone type is light and the age range is the first age range, it corresponds to the first population subtype; when the skin tone type is light and the age range is the second age range, it corresponds to the second population subtype; and so on, forming nine mutually exclusive population subtypes. When the skin tone type is light and the age range is the third age range, the target object is determined to belong to the third population subtype. When the skin tone type is medium and the age range is the first age range, the target object is determined to belong to the fourth population subtype. When the skin tone type is medium and the age range is the second age range, the target object is determined to belong to the fifth population subtype. When the skin tone type is medium and the age range is the third age range, the target object is determined to belong to the sixth population subtype. When the skin tone type is dark and the age range is the first age range, the target object is determined to belong to the seventh population subtype. When the skin tone type is dark and the age range is the second age range, the target is determined to belong to the eighth subtype. When the skin tone type is dark and the age range is the third age range, the target is determined to belong to the ninth subtype.

[0092] In this embodiment, the age range includes a first age range, a second age range, and a third age range. The first age range corresponds to young adults aged 18 to 35, whose autonomic nervous system is highly responsive, and whose heart rate variability and skin conductance responses exhibit significant variations. The second age range corresponds to middle-aged adults aged 36 to 55, whose autonomic nervous system responses tend to be more stable, and whose heart rate variability and skin conductance responses exhibit moderate variations. The third age range corresponds to the elderly population aged 56 and above, whose heart rate variability is generally reduced, skin conductance response recovery time is prolonged, and autonomic nervous system response speed is slowed. By dividing the age into three ranges, the differences in physiological response patterns among different age groups can be effectively distinguished.

[0093] In this embodiment, the preset template library is a database storing multiple preset population physiological characteristic templates. Each template in the template library corresponds to a specific population subtype. The template library is pre-built and stored locally or in the cloud during the system training phase. The template library is constructed based on a large amount of multimodal physiological data of target objects of different skin colors and age ranges under various emotional states, and standardized reference templates are formed through statistical modeling.

[0094] In this embodiment, the population physiological feature template includes statistical distribution parameters of the fused physiological feature vectors corresponding to each emotional state under each population subtype. These statistical distribution parameters include the mean vector and covariance matrix corresponding to each emotional state. For each emotional state, the mean vector and covariance matrix are calculated from the fused physiological feature vectors collected from a large number of target objects belonging to that population subtype when they are in that emotional state. The mean vector reflects the typical fused physiological feature centers of that population subtype in that emotional state, and the covariance matrix reflects the dispersion of the fused physiological feature vectors of that population subtype in that emotional state and the correlation between features.

[0095] In this embodiment, the Mahalanobis distance between the fused physiological feature vector and the statistical distribution parameters corresponding to each emotional state is calculated. Specifically, for each emotional state, the mean vector and covariance matrix corresponding to that emotional state are obtained, and the Mahalanobis distance between the fused physiological feature vector and the mean vector of that emotional state is calculated. The Mahalanobis distance is calculated by multiplying the transpose of the mean vector minus the fused physiological feature vector by the inverse of the covariance matrix, and then multiplying by the sum of the fused physiological feature vector minus the mean vector. The Mahalanobis distance considers the variance differences of each feature dimension and the correlation between feature dimensions, and can more accurately measure the degree of deviation between the fused physiological feature vector and the typical distribution of each emotional state. The smaller the Mahalanobis distance, the closer the fused physiological feature vector is to the typical features of that emotional state.

[0096] In this embodiment, the probability value of the fused physiological feature vector belonging to each emotional state is determined based on Mahalanobis distance as the initial confidence level to obtain the initial candidate emotional states. Specifically, the Mahalanobis distance between the fused physiological feature vector and each emotional state is converted into a probability value using a probability density function. This probability value represents the likelihood that the fused physiological feature vector belongs to the corresponding emotional state. The probability values ​​corresponding to each emotional state are sorted from high to low, and the emotional state with the highest probability value or the top few emotional states are selected as the initial candidate emotional states, with the corresponding probability value serving as the initial confidence level for that candidate emotional state. The initial candidate emotional states reflect a preliminary judgment of the target object's current emotional state based on population statistical patterns without considering individual differences. Furthermore, the population stratified validation network corrects the initial candidate emotional states based on the individual baseline physiological characteristics of the target object, including: When the target object uses the system for the first time or when the system triggers the baseline update condition, it is detected whether the target object is in a resting state. The resting state is determined by detecting that the target object's facial movement amplitude is less than the movement threshold and there is no response of skin conductance signal within a preset time. Baseline multimodal physiological signals were collected for a preset duration while the target was in a resting state. The baseline multimodal physiological signals included facial video stream, infrared thermal imaging sequence, and skin electrical activity signals in a resting state. Extracting individual baseline fused physiological feature vectors from baseline multimodal physiological signals; Calculate the basic offset difference between the individual baseline fused physiological feature vector and the resting state standard vector in the population physiological feature template of the target object's subtype; An offset correction factor is generated based on the baseline offset difference. The initial confidence of each emotional state in the initial candidate emotional states is multiplied by the offset correction factor to obtain the corrected confidence. The emotional state with the highest confidence level after correction is taken as the final emotional state.

[0097] In this embodiment, the system triggers baseline updates under at least one of the following conditions: more than 30 days have passed since the last baseline acquisition, indicating that the target subject's physiological state may have changed over time; the target subject's emotional state identification results are non-resting state for three consecutive times, indicating that the target subject may be in a non-resting state for a long time and a new baseline needs to be established; the detection equipment has been replaced, indicating that the sensor characteristics of the new equipment may differ from those of the old equipment; the acquisition environment has changed significantly, including changes in ambient temperature exceeding 5 degrees Celsius or changes in ambient humidity exceeding 20%, indicating that environmental factors may affect the measurement results of skin conductance activity signals and infrared thermal imaging signals. When any of these conditions are met, the system triggers the baseline update process to re-acquire the target subject's individual baseline physiological characteristics.

[0098] In this embodiment, the resting state is determined by detecting that the target object's facial movement amplitude is less than a movement threshold and there is no response to the electrodermal activity signal within a preset time period. When the target object simultaneously meets the facial movement amplitude condition and the electrodermal activity signal condition, the target object is determined to be in a resting state. This determination method ensures that the target object is in a stable state without emotional fluctuations or physical activity during baseline acquisition.

[0099] In this embodiment, the preset duration refers to the length of the time window required to determine the resting state. The preset duration is set to 30 seconds, which is sufficient to eliminate transient motion interference and occasional skin conductance responses, ensuring the reliability of the resting state determination.

[0100] In this embodiment, facial motion amplitude refers to the degree of positional change of key facial points of the target object within a preset time period. Facial key points, including the center of the eyebrows, the tip of the nose, and the left and right corners of the mouth, are located in each frame of the facial video stream using a facial key point detection algorithm. The displacement of each key point between consecutive frames is calculated, and the maximum displacement of all key points is taken as the facial motion amplitude. Facial motion amplitude reflects the intensity of head movement and facial expression changes of the target object.

[0101] In this embodiment, the motion threshold is a critical value used to determine whether the target object is in a static facial state. The motion threshold is set to 2 millimeters. When the facial movement amplitude is less than 2 millimeters, the target object's face is determined to be in a static state, eliminating interference from head shaking or facial expression changes on signal acquisition.

[0102] In this embodiment, no response to the skin conductance signal means that the skin conductance level does not exhibit a complete response waveform that rises above a preset amplitude threshold from the baseline value and then returns to the baseline value within a preset duration. Specifically, the skin conductance level signal is iterated within the preset duration, and if no complete skin conductance response is detected, the skin conductance signal is determined to be unresponsive. This determination ensures that there are no sympathetic nerve activation events in the target subject during baseline acquisition.

[0103] In this embodiment, the preset duration refers to the length of the time window required to determine that there is no response to the skin conductance signal, which is consistent with the preset duration for determining the resting state and is set to 30 seconds.

[0104] In this embodiment, an individual baseline fused physiological feature vector is extracted from baseline multimodal physiological signals. Specifically, a facial video stream of a preset duration, an infrared thermal imaging sequence, and a skin conductance signal (SCES) of the target object in a resting state are used as baseline multimodal physiological signals. Following the same processing flow as real-time recognition, the facial video stream is spectrally adaptively enhanced and then input into a spatial modulation attention network to obtain a remote photoplethysmography (TPS) waveform. A temperature distribution feature vector is extracted from the infrared thermal imaging sequence, and a skin conductance response feature vector is extracted from the SCES signal. The three signals are spatiotemporally aligned and then input into a multimodal signal pattern recognition network for weighted fusion, outputting a fused physiological feature vector. This fused physiological feature vector is the individual baseline fused physiological feature vector. The individual baseline fused physiological feature vector reflects the personalized physiological characteristics of the target object in a resting state.

[0105] In this embodiment, the basic offset difference between the individual baseline fused physiological feature vector and the resting state standard vector in the physiological feature template of the target population subtype is calculated. Specifically, the mean vector corresponding to the resting state is extracted from the physiological feature template of the target population subtype as the resting state standard vector. The resting state standard vector is subtracted from the individual baseline fused physiological feature vector to obtain the difference vector. The magnitude or weighted sum of each dimension of the difference vector is calculated as the basic offset difference. The basic offset difference reflects the degree of deviation between the individual physiological characteristics of the target object and the average level of its population.

[0106] In this embodiment, the resting state standard vector in the population physiological characteristic template refers to the mean vector of the fused physiological characteristic vectors collected from a large number of target objects belonging to this population subtype in the resting state. This standard vector represents the typical physiological characteristic pattern of this population subtype in the resting state and serves as a reference benchmark for individual baseline correction.

[0107] In this embodiment, an offset correction factor is generated based on the baseline offset difference. Specifically, the baseline offset difference is converted into an offset correction factor using a monotonically decreasing function. When the baseline offset difference is small, it indicates that the individual characteristics of the target object are close to the average level of the population, the offset correction factor is close to 1, and the correction magnitude for the initial confidence level is small. When the baseline offset difference is large, it indicates that the individual characteristics of the target object differ significantly from the average level of the population, the offset correction factor deviates from 1, and the correction magnitude for the initial confidence level is large. The offset correction factor is used to adjust the initial confidence level so that the emotional state judgment result is more consistent with the individual characteristics of the target object.

[0108] In this embodiment, the initial confidence level of each emotional state in the initial candidate emotional states is multiplied by a shift correction factor to obtain the corrected confidence level. Specifically, for each emotional state in the initial candidate emotional states, the initial confidence level of that emotional state is multiplied by the shift correction factor, and the product is used as the corrected confidence level of that emotional state. The initial confidence level is scaled by the shift correction factor, adjusting the confidence level distribution in a direction that conforms to the individual characteristics of the target object.

[0109] In this embodiment, the emotional state with the highest corrected confidence level is taken as the final emotional state. Specifically, among all corrected confidence levels of emotional states, the emotional state corresponding to the maximum value is identified, and this emotional state is output as the final emotional state recognition result. Through individual baseline correction, the final output emotional state considers both population statistical patterns and the personalized physiological characteristics of the target object, thereby improving the accuracy of emotion recognition. Obviously, those skilled in the art can make various modifications and variations to this invention without departing from the spirit and scope of this invention. Thus, if these modifications and variations of this invention fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for recognizing physiological and emotional states of a diverse population, comprising: include: Acquire facial video streams, infrared thermal imaging sequences, and skin electrical activity signals of the target object; The facial video stream is subjected to spectral adaptive enhancement based on the skin color type of the target object to obtain a spectral enhanced video stream. Spectral adaptive enhancement includes adjusting the weight of each color channel according to the skin color type. When the skin color type is dark, a near-infrared simulation channel is also constructed. The spectral-enhanced video stream is input into a spatially modulated attention network to obtain a remote photoplethysmography waveform; Temperature distribution feature vectors are extracted from infrared thermal imaging sequences. These feature vectors include the forehead temperature change rate, the nose tip temperature fluctuation amplitude, and the cheek temperature asymmetry index. The skin conductance response feature vector is extracted from the skin conductance activity signal. The skin conductance response feature vector includes the baseline value of skin conductance level, skin conductance response frequency, and skin conductance recovery time. Spatiotemporal alignment of remote photoplethysmography waveforms, temperature distribution feature vectors, and skin conductance response feature vectors; The validity of a signal is determined by the real-time signal-to-noise ratio of the remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector. For failed signals with a real-time signal-to-noise ratio lower than a preset threshold, alternative features of the failed signal are constructed based on the physiological correlation characteristics of other valid signals. The remote photoplethysmography waveform, temperature distribution feature vector, skin conductance response feature vector, and alternative features are input into a multimodal signal pattern recognition network. The multimodal signal pattern recognition network dynamically determines the fusion weights based on the effective signal-to-noise ratio of the remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector, and performs weighted fusion to output a fused physiological feature vector. The fused physiological feature vector is input into the population stratification validation network. The population stratification validation network matches the initial candidate emotional state from multiple preset population physiological feature templates based on the target object's skin color type and age range, and corrects the initial candidate emotional state based on the target object's individual baseline physiological characteristics, and outputs the final emotional state.

2. The method of claim 1, wherein the method is a method of recognizing physiological and emotional states of a diverse population. Constructing a near-infrared simulation channel includes: The average pixel value of multiple skin regions in the facial video stream of the target object is identified, and the skin color type of the target object is determined based on the average pixel value. The skin color types include light skin color, medium skin color and dark skin color. When the skin tone type is light skin tone, the green channel weight of the facial video stream is set as the first weight value, and the red and blue channel weights are set as the second weight values, with the first weight value being greater than the second weight value; When the skin tone type is dark skin tone, the weighted difference between the pixel values ​​of the red channel and the pixel values ​​of the blue channel in the facial video stream is used as the pixel value of the near-infrared analog channel. The weight of the near-infrared analog channel in the facial video stream is set as the third weight value, and the weight of the green channel in the facial video stream is set as the fourth weight value. The third weight value is greater than the fourth weight value. The pixel values ​​of each channel of the facial video stream are multiplied by their corresponding channel weights and then summed to generate a spectrally enhanced video stream.

3. The method of claim 1, wherein the method is a method of recognizing physiological and emotional states of a diverse population. Spatial modulation attention network obtains long-range photoplethysmography waveforms, including: The spectral-enhanced video stream is divided into multiple time windows, each containing a preset number of consecutive video frames; Spatiotemporal feature extraction is performed on video frames within each waveform extraction time window. Spatiotemporal feature extraction includes extracting spatial features through three-dimensional convolution and extracting color change difference features between adjacent frames through temporal center difference convolution. Temporal center difference convolution enhances the ability to capture periodic changes in skin color by calculating the feature difference between the current frame and adjacent frames. The extracted spatiotemporal features are segmented into multiple spatiotemporal channels, and each spatiotemporal channel corresponds to the feature sequence of a local region of the face within the waveform extraction time window; A spatial importance map is generated based on the feature sequences of all spatiotemporal pipelines. Regions with high weight values ​​in the spatial importance map correspond to skin areas in the spectral enhanced video stream that are fully exposed, have small motion amplitude, and are unobstructed. Attention weighting is applied to the feature sequences of the spatiotemporal pipeline based on the spatial importance map, so that regions with high weight values ​​in the spatial importance map receive higher attention weights. The attention-weighted spatiotemporal pipeline feature sequence is decoded into a one-dimensional remote photoplethysmography waveform.

4. The method for recognizing the physiological and emotional states of diverse populations according to claim 1, characterized in that, Calculate the real-time signal-to-noise ratio of remote photoplethysmography waveforms, temperature distribution feature vectors, and skin conductance response feature vectors, including: The spatiotemporally aligned remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector are used as inputs; Within a preset time window in which the target object's face remains still, the signal amplitude sequence of the remote photoplethysmography waveform is acquired. The average value of the preset ratio range with the lowest amplitude value in the signal amplitude sequence is taken as the baseline noise amplitude. The ratio of the peak signal amplitude to the baseline noise amplitude within the current time window is taken as the first real-time signal-to-noise ratio of the remote photoplethysmography waveform. Multiple sets of facial temperature distribution data were collected when the target object was in a resting state. Temperature fluctuation components unrelated to emotional response were extracted from each set of facial temperature distribution data. A preset noise template was constructed, and the matching degree between the temperature change curve in the current time window and the preset noise template was used as the second real-time signal-to-noise ratio of the temperature distribution feature vector. During the resting period when there is no response to the skin conductance signal, the low-frequency fluctuation amplitude of the skin conductance level signal is calculated as the baseline drift amplitude, and the ratio of the skin conductance response amplitude to the baseline drift amplitude within the current time window is used as the third real-time signal-to-noise ratio of the skin conductance response feature vector.

5. The method for recognizing the physiological and emotional states of diverse populations according to claim 4, characterized in that, The multimodal signal pattern recognition network dynamically determines the fusion weights and performs weighted fusion based on the effective signal-to-noise ratio of remote photoplethysmography waveforms, temperature distribution feature vectors, and skin conductance response feature vectors, including: The first real-time signal-to-noise ratio of the remote photoplethysmography waveform is converted into a first fusion weight through a first mapping function. The second real-time signal-to-noise ratio of the temperature distribution feature vector is converted into a second fusion weight through a second mapping function. The third real-time signal-to-noise ratio of the skin conductance response feature vector is converted into a third fusion weight through a third mapping function. The first, second, and third mapping functions are monotonically increasing functions, so that signals with high signal-to-noise ratios receive higher fusion weights. The remote photoplethysmography waveform, temperature distribution feature vector, and skin conductance response feature vector are weighted and concatenated according to the first fusion weight, the second fusion weight, and the third fusion weight to generate a fused physiological feature vector.

6. The method for recognizing the physiological and emotional states of diverse populations according to claim 4, characterized in that, It also includes determining that the remote photoplethysmography waveform is in a failed state when the first real-time signal-to-noise ratio of the remote photoplethysmography waveform is lower than the first preset threshold, constructing alternative features of the remote photoplethysmography waveform based on the temperature distribution feature vector and the skin conductance response feature vector, and using the preset alternative weight as the fusion weight of the remote photoplethysmography waveform. The alternative weight is greater than the fusion weight transformed by the mapping function of the first real-time signal-to-noise ratio, and does not exceed the maximum fusion weight of the effective signal. When the second real-time signal-to-noise ratio of the temperature distribution feature vector is lower than the second preset threshold, the temperature distribution feature vector is determined to be in a failed state, and alternative features of the temperature distribution feature vector are constructed based on the remote photoplethysmography waveform and the skin conductance response feature vector. When the third real-time signal-to-noise ratio of the skin electric response feature vector is lower than the third preset threshold, the skin electric response feature vector is determined to be in a failed state. Alternative features of the skin electric response feature vector are constructed based on the remote photoplethysmography waveform and temperature distribution feature vector. The alternative features are constructed by extracting feature components that are physiologically related to the failure signal from the effective signal, and mapping the feature components to the estimated value of the failure signal through a pre-trained regression model.

7. The method for recognizing the physiological and emotional states of diverse populations according to claim 1, characterized in that, Extracting temperature distribution feature vectors from infrared thermal imaging sequences includes: In each frame of the infrared thermal imaging sequence, the forehead region, nose tip region, and cheek region are located, with the cheek region including the left cheek region and the right cheek region. Calculate the slope of the curve showing the change in average temperature of the forehead region over time within a time window, and use the slope value as the rate of change in forehead temperature. Calculate the difference between the maximum and minimum average temperatures of the nasal tip region within the time window, and use this difference as the nasal tip temperature fluctuation range. Calculate the average temperature difference between the left and right cheek regions within a time window, and use the difference as the cheek temperature asymmetry index.

8. The method for recognizing the physiological and emotional states of diverse populations according to claim 1, characterized in that, Extracting skin conductance response feature vectors from skin conductance activity signals includes: During the resting period when there is no response to the skin electrical activity signal, the average value of the skin conductance level is calculated as the baseline value of the skin conductance level; The number of skin conductance responses is counted within a time window. Skin conductance response is defined as the process by which the skin conductance level rises from the baseline value of the skin conductance level to the baseline value of the skin conductance level after exceeding a preset amplitude threshold. The ratio of the number of skin conductance responses to the duration of the time window is used as the skin conductance response frequency. For each skin conductance response, the time difference from the start of the response to the recovery to the baseline level of skin conductance is calculated, and the average recovery time of all skin conductance responses is taken as the skin conductance recovery time.

9. The method for recognizing the physiological and emotional states of diverse populations according to claim 1, characterized in that, The stratified demographic validation network matches initial candidate emotional states from multiple pre-defined demographic physiological feature templates based on the target audience's skin color type and age range, including: The target population subtype is determined based on the combination of the target population’s skin color type and the target population’s age range. The skin color type includes light skin color type, medium skin color type and dark skin color type, and the age range includes first age range, second age range and third age range. Extract the population physiological feature template corresponding to the population subtype to which the target object belongs from the preset template library. The population physiological feature template contains the statistical distribution parameters of the fused physiological feature vectors corresponding to each emotional state under the population subtype. The statistical distribution parameters include the mean vector and covariance matrix corresponding to each emotional state. The Mahalanobis distance between the fused physiological feature vector and the statistical distribution parameters corresponding to each emotional state is calculated. Based on the Mahalanobis distance, the probability value of the fused physiological feature vector belonging to each emotional state is determined as the initial confidence level, and the initial candidate emotional states are obtained.

10. The method for recognizing physiological and emotional states of diverse populations according to claim 9, characterized in that, The stratified population validation network modifies the initial candidate emotional states based on the individual baseline physiological characteristics of the target population, including: When the target object uses the system for the first time or when the system triggers the baseline update condition, it is detected whether the target object is in a resting state. The resting state is determined by detecting that the target object's facial movement amplitude is less than the movement threshold and there is no response of skin conductance signal within a preset time. Baseline multimodal physiological signals were collected for a preset duration while the target was in a resting state. The baseline multimodal physiological signals included facial video stream, infrared thermal imaging sequence, and skin electrical activity signals in a resting state. Extracting individual baseline fused physiological feature vectors from baseline multimodal physiological signals; Calculate the basic offset difference between the individual baseline fused physiological feature vector and the resting state standard vector in the population physiological feature template of the target object's subtype; An offset correction factor is generated based on the baseline offset difference. The initial confidence of each emotional state in the initial candidate emotional states is multiplied by the offset correction factor to obtain the corrected confidence. The emotional state with the highest confidence level after correction is taken as the final emotional state.