Method for assessing emotional retardation based on facial expression and skin electrical reaction deviation degree

CN122786014APending Publication Date: 2026-09-22SHANGHAI DONGHAI VOCATIONAL & TECH COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610938677.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0006]针对现有技术的不足,本发明提供了一种基于面部表情与皮肤电反应背离度的情感迟钝评估方法,解决了传统多模态情绪识别方法依赖面部表情与生理信号的一致性作为主要判断依据,忽略模态之间可能存在的背离关系对情感迟钝评估的重要性的问题

Benefits of technology

(1)本发明通过将人脸区域划分为面部关键特征局部区域,不直接对整张人脸进行粗粒度分类,而是对面部关键特征局部区域进行多尺度空间建模,提高了对弱表情和微小表情变化的捕捉能力;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122786014A_ABST
    Figure CN122786014A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of based on facial expression and skin galvanic reaction divergence degree emotional blunting evaluation method, including synchronous acquisition facial video sequence, skin galvanic signal, clinical reference label, facial video sequence and skin galvanic signal are divided into several analysis segments according to uniform time axis and are filtered respectively, construct facial expression response sequence and skin galvanic response sequence, calculate the divergence degree of facial expression response sequence and skin galvanic response sequence, calculate emotional blunting score and judge emotional blunting type, construct emotional blunting evaluation output, construct target loss function and train emotional blunting evaluation model, and then the emotional blunting evaluation result of patient is output by the emotional blunting evaluation model of training completion.The present application can retain and quantify the inconsistent relationship between facial overt response and peripheral physiological response, reduce the misjudgment risk caused by single visual judgment, improve the objectivity and explainability of depressive emotional blunting auxiliary evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of medical artificial intelligence and multimodal signal processing technology, and to a method for assessing emotional blunting based on the divergence between facial expression and electrodermal response. Background Technology

[0002] Emotional blunting is a common finding in mental health assessments, typically manifested as reduced facial expressions, diminished intonation, insufficient eye contact, reduced postural movements, and decreased external responses to emotional stimuli. Reduced emotional expression can be observed to varying degrees in depression and some chronic anxiety states. It's important to note that emotional blunting does not mean an individual lacks emotional experience. Some patients experience physical pain, tension, or other negative emotions, but these are not adequately expressed through facial expressions. In current clinical observation and intelligent assessment methods, facial expressions are crucial for determining emotional state. When individuals are in a state of pain, anxiety, or negative emotions, they typically exhibit outward behaviors such as brow retraction, eye tension, changes in the corners of the mouth, and increased facial muscle tension. Existing facial emotion recognition methods usually identify emotion categories through facial images and facial movement units. However, for individuals with emotional blunting, low facial expression does not equate to low distress or low arousal. Judging solely based on visual information can easily lead to misjudgments of patients as emotionally stable, without significant pain, or in a low-risk state.

[0003] Besides facial visual signals, electrodermal activity (EDA) is a commonly used peripheral physiological signal in emotion and stress research. EDA reflects changes in sweat gland activity, which is primarily regulated by the sympathetic nervous system; therefore, EDA can be used to characterize sympathetic-related peripheral arousal changes. Compared to facial expressions, EDA does not depend on the patient's voluntary expression, providing another perspective on an individual's physiological response to stimuli. EDA is significantly influenced by skin condition, ambient temperature, and individual baseline levels. In individuals with mood disorders such as depression, EDA responses are not always elevated; they may also manifest as low response, delayed response, or task-dependent abnormalities. Therefore, elevated EDA cannot be simply equated with anxiety, nor can decreased EDA be directly equated with emotional blunting.

[0004] Existing multimodal emotion recognition methods typically attempt to fuse facial videos and physiological signals to improve emotion classification accuracy. Common methods include feature concatenation, decision fusion, cross-modal attention, and multimodal Transformers. Most of these methods assume that information relevant to the target task in different modalities should be complementary or consistent, extracting common discriminative features between facial expressions and physiological signals. However, for emotion attenuation assessment, the truly clinically significant information is not the consistency between facial expressions and physiological signals, but rather the divergence between them. For example, some individuals may show weak facial expression changes but still exhibit significant EDA responses, indicating a state of limited external expression but strong autonomic nervous system responses. Other individuals may simultaneously exhibit low facial and EDA responses. Traditional multimodal fusion methods, which merely compress visual and EDA features into a unified classification vector, may treat this inconsistency as noise, thus losing crucial information in emotion attenuation assessment.

[0005] Therefore, the emotional lethargy assessment task observes facial expressions and EDA signals, but does not simply judge the strength of facial expressions or the level of EDA. Instead, it establishes quantitative indicators based on the relationship between the two responses to the same stimulus event. By calculating the divergence between facial expression responses and skin conductance responses, it objectively characterizes different response patterns that may occur in a state of emotional lethargy, such as low external expression but preserved physiological response, or low external expression and low physiological response. This improves the accuracy and interpretability of the auxiliary assessment of emotional lethargy. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a method for assessing emotional blunting based on the divergence between facial expression and electrodermal response (EDS). This method solves the problem of traditional multimodal emotion recognition methods relying primarily on the consistency between facial expressions and physiological signals, neglecting the importance of potential divergences between modalities in assessing emotional blunting. This invention does not use the absolute increase or decrease in EDS signals as a sole criterion. Instead, it obtains the divergence between facial expression and EDS based on differences in response amplitude, dynamic correlation, peak relationships, and inconsistencies in back attention segments. This preserves and quantifies the inconsistency between the two, using this inconsistency as the basis for assessing emotional blunting.

[0007] To achieve the above-mentioned technical objectives, the present invention provides the following technical solution: a method for assessing emotional blunting based on the divergence between facial expression and skin conductance response, comprising the following steps: S1. Simultaneously collect facial video sequences, electrodermal signals, and clinical reference labels from patients to form an emotional dullness assessment dataset; the clinical reference labels include the patient's emotional dullness score, emotional dullness type, and the true value of the deviation degree. S2. Divide the facial video sequence and the electrodermal signal into several analysis segments according to a unified time axis and filter them separately to obtain several effective facial video segments and the same number of effective electrodermal segments as the effective facial video segments. S3. Calculate the facial expression response intensity of each valid facial video segment and construct a facial expression response sequence; S4. Calculate the skin electrical response intensity of each effective skin electrical segment and construct the skin electrical response sequence; S5. Calculate the deviation between the facial expression response sequence and the skin conductance response sequence; the deviation is obtained by weighted fusion of amplitude deviation, dynamic correlation deviation, peak relationship deviation, and inverse attention deviation. S6. Calculate the degree of low facial reactivity and the degree of electrical skin reactivity based on the facial expression response sequence and the skin conductance response sequence, and then calculate the emotional dullness score by combining the deviation degree and determine the type of emotional dullness, and construct the emotional dullness assessment output; the emotional dullness type includes at least expression inhibition type emotional dullness and low reactivity type emotional dullness. S7. Construct a target loss function and train the emotional sluggishness assessment model using the emotional sluggishness assessment dataset; the emotional sluggishness assessment model takes facial video sequences and electrodermal signals as inputs and predicts the emotional sluggishness assessment output through steps S2-S6. S8. Real-time acquisition of facial video sequences and skin electrical signals of the patient, and prediction of emotional dullness assessment output by the trained emotional dullness assessment model, which serves as the patient's emotional dullness assessment result.

[0008] Optionally, the step of dividing the facial video sequence and the electrodermal signal into several analysis segments according to a unified time axis and filtering them separately to obtain several valid facial video segments and the same number of valid electrodermal segments as the valid facial video segments includes: The facial video sequence is divided into several analysis segments to obtain facial video analysis segments; The skin electrical signals were divided in the same way as the facial video analysis segments to obtain the same number of skin electrical analysis segments as the facial video analysis segments; The facial video analysis segments are filtered to remove segments that fail to detect faces, have severe occlusion, have large head rotation, or have abnormal lighting, in order to obtain the initial valid facial video segments. The skin conductance analysis segments are filtered to remove segments with poor contact, sudden spikes, obvious motion artifacts, or signal saturation, thus obtaining the initial effective skin conductance segments; The intersection of the initial valid facial video clips and the initial valid EEG clips on the time axis is taken, and the initial valid facial video clips and the initial valid EEG clips in each time interval of the intersection are retained to obtain the same number of valid facial video clips and valid EEG clips that are time-aligned.

[0009] Optionally, the step of calculating the facial expression response intensity of each valid facial video segment and constructing a facial expression response sequence includes: Face localization, key point detection and face alignment are performed on each frame of the effective facial video segment to obtain the face region and several facial key points. The face is divided into key facial feature regions based on key facial points; the key facial feature regions include one or more of the following: eyebrow region, periorbital region, nasal alar region, nasolabial fold region, corner of mouth region, and mandibular region. Each frame of the effective facial video segment is input into a convolutional neural network for feature extraction to obtain an initial full-face feature map. Based on the geometric coordinates of the local regions of the key facial features, several local feature blocks are extracted from the initial full-face feature map through region of interest alignment. Then, the local feature blocks are cascaded to obtain a spatial feature map. A multi-scale spatial enhancement module is constructed to perform convolution enhancement, attention enhancement, and multi-scale spatial enhancement on the spatial feature maps of each frame in the effective facial video segment, and to extract spatial representation vectors. A weakly supervised temporal attention mechanism is introduced to perform weighted fusion of the spatial representation vectors of each frame in the effective facial video clips, generating the clip-level facial expression response vectors of each effective facial video clip. Facial baseline representations are extracted from each valid facial video segment. The facial expression response intensity of each valid facial video segment relative to the facial baseline representation is calculated using segment-level facial expression response vectors, and a facial expression response sequence is constructed.

[0010] Optionally, the multi-scale spatial enhancement module performs convolution enhancement, attention enhancement, and multi-scale spatial enhancement on the spatial feature maps of each frame in the effective facial video segment, extracting spatial representation vectors, including: Deep convolution is performed on each frame of the valid facial video clip to extract texture features channel by channel and construct feature maps. Pointwise convolution is performed on the feature map obtained by depthwise convolution to achieve information fusion between channels; Two-dimensional self-attention weighting is performed on the feature map obtained by pointwise convolution to generate a self-attention response feature map, mathematically represented as follows: ; ; in, Indicates the self-attention weights; This indicates the Softmax normalization operation; , , These represent the query matrix, key matrix, and value matrix, respectively, obtained by linearly mapping the feature map obtained through pointwise convolution; The key matrix represents a feature dimension; the query matrix, key matrix, and value matrix all share the same feature dimension. Indicates the transpose operation; Represents the self-attention response feature map; Channel attention weights and spatial attention weights are calculated separately for the self-attention response feature maps, and then fused to obtain spatially enhanced features, mathematically represented as follows: ; ; ; in, , These represent channel attention weights and spatial attention weights, respectively. Indicates the Sigmoid activation operation; Indicates a multi-layer perceptron; This indicates an average pooling operation; This represents the max pooling operation; Indicates the convolution operation; This represents the channel average pooling operation performed along the channel dimension; This represents the channel max pooling operation performed along the channel dimension; Indicates a splicing operation; Represents spatially enhanced features; This represents element-wise multiplication; The spatial augmentation features are subjected to spatial dimension reduction and flattening to generate spatial representation vectors. This process is mathematically represented as follows: ; in, This represents the global average pooling operation, used to augment spatial features. Compression of two-dimensional spatial features into one-dimensional vectors; This indicates a fully connected layer, used for flattening. Indicates the first valid facial video segment Spatial representation vector of a frame image.

[0011] Optionally, the introduction of a weakly supervised temporal attention mechanism, which weights and fuses the spatial representation vectors of each frame in a valid facial video segment to generate a segment-level facial expression response vector for each valid facial video segment, includes: The process involves calculating temporal attention weights for the spatial representation vector, setting a temporal attention threshold, and filtering out vectors with low temporal attention weights. This process can be mathematically represented as follows: ; ; in, Indicates the first valid facial video segment Temporal attention weights for frame images; This indicates the Softmax normalization operation; This is the attention scalar mapping vector; This is the time-mapped weight matrix; This is the time-mapped bias vector; Indicates the transpose operation; Represents the hyperbolic tangent function; Indicates the first valid facial video segment The spatial representation vector of a frame image; Indicates the first valid facial video segment Temporal attention weights after filtering of frame images; The time attention threshold; The filtered temporal attention weights are normalized to obtain normalized frame-level weights. This process is mathematically represented as follows: ; in, Indicates the first valid facial video segment Normalized frame-level weights for frame images; The spatial representation vectors of each frame in a valid facial video segment are weighted and fused using normalized frame-level weights to generate a segment-level facial expression response vector for that valid facial video segment. This process is mathematically represented as follows: ; in, Indicates the first A segment-level facial expression response vector for each valid facial video clip.

[0012] Optionally, the facial reference characterization is a stationary segment or a sliding mean among each valid facial video segment. The stationary segment is determined by calculating the instantaneous variance of the motion displacement of facial key points within each valid facial video segment. Valid facial video segments with instantaneous variance lower than a preset stationary threshold are considered stationary segments. The sliding mean is obtained by taking the arithmetic mean of the segment-level facial expression response vectors of the current valid facial video segment and its preceding adjacent valid facial video segments through a preset span time sliding window. The calculation of the facial expression response intensity of each effective facial video segment relative to the facial reference using segment-level facial expression response vectors is mathematically represented as follows: ; in, Represents facial reference characteristics; Indicates the first Fragment-level facial expression response vectors for each valid facial video segment. Indicates the index of valid facial video clips. Total number of valid facial video clips; Represents the L2 norm computation operator; Indicates the first A valid facial video segment relative to a facial reference representation The intensity of facial expression responses; The facial expression response sequence is mathematically represented as follows: ; in, This represents a sequence of facial expression responses.

[0013] Optionally, calculating the skin electrical response intensity of each effective skin electrical segment and constructing a skin electrical response sequence includes: Each effective electrodermal cell segment undergoes preprocessing including noise reduction, outlier removal, smoothing, and standardization. Each effective skin electrodermal fragment after preprocessing was subjected to low-pass filtering to extract its tension component; The phase component is obtained by subtracting the effective skin electrical segment from its tension component; For each effective electrodermal segment, the local mean, fluctuation, peak amplitude, and energy of its phase components are calculated. The mathematical representation of the calculation method is as follows: ; ; ; ; in, , , , They represent the first Local mean, variability, peak amplitude, and energy of an effective skin electrodermal segment; Indicates the index of valid skin electrical fragments. The total number of effective skin electrodermal fragments is equal to the total number of effective facial video fragments; Indicates the first Number of sampling points within an effective skin electrical activity segment Indicates the first Within the first effective skin conductance fragment The signal amplitude at each sampling point Indicates the sampling point index; Indicates in Values ​​range from 1 to Take within range The maximum value; Perform a Fourier transform on the phase components to obtain the time-frequency representation of each effective electrodermal segment: ; in, Indicates the time offset of the phase component; is the rotation factor of the Fourier transform. For frequency variables, The imaginary unit; Represents a continuous physiological time series variable; Indicates a timing sliding window function; For the first Time shift of an effective skin electrodermal fragment and frequency variables Time-frequency representation; Calculate the energy within the target frequency band for each effective skin electrical activity segment based on time-frequency representation: ; in, No. Energy within the target frequency band of an effective skin electrodermal segment; , Representing frequency variables respectively The upper and lower bounds of the frequency band define the frequency range that the two bounds define; Indicates the calculation of the square of the modulus; The local mean, fluctuation, peak amplitude, and energy of the effective skin electric field fragments are constructed into a time-domain feature vector, and the energy within the target frequency band of the effective skin electric field fragments is constructed into a frequency-domain feature vector. ; ; in, , They represent the first The temporal feature vector of the first effective skin conductance segment, the first Frequency domain feature vectors of effective skin electrical activity segments; Indicates the transpose operation; By projecting the time-domain feature vector and the frequency-domain feature vector onto the same dimension through a linear mapping, we obtain the time-domain features and the frequency-domain features. Cross-domain attention calculation is performed between time-domain features and frequency-domain features to obtain the attention relationship between them. This process is mathematically represented as follows: ; ; in, This indicates the relationship of interest between time-domain features and frequency-domain features; This indicates the relationship of interest between frequency domain features and time domain features; This indicates the Softmax normalization operation; , , These represent the time-domain query matrix, time-domain key matrix, and time-domain value matrix, respectively, obtained through linear mapping of time-domain features; , , These represent the frequency domain query matrix, frequency domain key matrix, and frequency domain value matrix, respectively, obtained through linear mapping of frequency domain features; The characteristic dimension represents the time-domain key matrix. The time-domain query matrix, time-domain key matrix, time-domain value matrix, frequency-domain query matrix, frequency-domain key matrix, and frequency-domain value matrix all have the same characteristic dimension. Gated fusion of the interest relationships between time-domain and frequency-domain features is performed to generate a skin conductance response vector. The mathematical representation of this process is as follows: ; ; in, Indicates the first The gating coefficient of an effective skin electrodermal segment; Indicates the Sigmoid activation operation; , These represent the learnable gating weights and the gating bias term, respectively. Indicates a splicing operation; Indicates the first The skin electrical response vector of an effective skin electrical segment; This represents element-wise multiplication; Extract the baseline skin electrical activity (TEA) characterization and calculate the TEA response intensity of each effective TEA segment relative to this baseline characterization. The mathematical representation is as follows: ; in, Indicates the baseline characterization of skin electrical activity; Indicates the first Characterization of effective skin conductance segments relative to skin conductance baseline The intensity of skin conductance response; The L2 norm calculation operator is used; the skin electrophysiological benchmark is characterized as a stationary segment or a moving average in each effective skin electrophysiological segment. The stationary segment is determined by calculating the instantaneous variance of the signal amplitude at the sampling point in each effective skin electrophysiological segment and comparing it with a preset stationarity discrimination threshold. The moving average is obtained by taking the arithmetic mean of the skin electrophysiological response vectors of consecutive adjacent effective skin electrophysiological segments within a preset span of a time-series sliding window. The skin conductance response sequence is constructed based on the intensity of the skin conductance response, and its mathematical representation is as follows: ; in, This represents the skin conductance response sequence.

[0014] Optionally, calculating the deviation between the facial expression response sequence and the skin conductance response sequence includes: S51. Calculate the amplitude divergence between the intensity of each facial expression response and its corresponding skin conductance response intensity, mathematically represented as follows: ; in, Indicates the first Facial expression intensity of a valid facial video clip The corresponding first Skin electroreactivity intensity of each effective skin electroreactivity segment The amplitudes diverged between them; This indicates a normalization operation; This indicates the calculation of the modulus, used here to obtain the absolute value; S52. Define sliding windows for extracting facial expression response sequences and skin conductance response sequences, respectively. , Facial expression response subsequence was obtained. With skin electroreactivity sequence Then, the dynamic correlation divergence between the intensity of each facial expression response and its corresponding skin conductance response intensity is calculated, and the mathematical representation is as follows: ; in, Indicates the first The intensity of facial expression response in a valid facial video clip and its corresponding... The dynamic correlation between the intensity of skin electrical response of each effective skin electrical segment diverged; This indicates the time delay between visual response and physiological sympathetic nerve response. This indicates the preset allowable time delay range; Indicates in Evaluate expression within range The maximum value; This indicates the operator for calculating the Pearson correlation coefficient; Represents a sliding window Superimposed time delay The extracted skin electroreactivity sequence; S53, Obtain the sliding window respectively , Peak times of the inner facial expression response subsequence and the electrodermal response subsequence , Furthermore, the peak deviation between the intensity of each facial expression response and its corresponding skin conductance response intensity is calculated, and the mathematical representation is as follows: ; in, Indicates the first The intensity of facial expression response in a valid facial video clip and its corresponding... The peak relationship between the skin electrical response intensities of each effective skin electrical segment deviates; S54. Calculate the inverse attentional divergence between facial expression response sequences and skin conductance response sequences, including: S541. Using one-dimensional temporal convolutional layers, the facial expression response sequence and the skin conductance response sequence are respectively subjected to dimensionality increase and temporal local feature extraction to construct a facial feature sequence. With skin electrophysiological feature sequence ; S542. Calculate facial feature sequences With skin electrophysiological feature sequence The mathematical representation of the standard cross-modal attention matrix is ​​as follows: ; in, This represents a standard cross-modal attention matrix; This indicates the Softmax normalization operation; This represents the facial feature query matrix obtained by linear mapping of facial features in a facial feature sequence. This represents the skin electrophysiological feature key matrix obtained by linear mapping of skin electrophysiological features in a skin electrophysiological feature sequence; Indicates the transpose operation; The feature dimension of the column features after the skin electrodermal feature key matrix is ​​linearly mapped in the feature space. The facial feature query matrix and the skin electrodermal feature key matrix have the same feature dimension. S543. Calculate facial feature sequences With skin electrophysiological feature sequence The inverse cross-modal attention matrix between them is mathematically represented as follows: ; in, Representation compared to a conventional cross-modal attention matrix A matrix with the same dimension and all elements being 1; This represents the inverse cross-modal attention matrix; S544. Calculate the divergence feature representation based on the inverse cross-modal attention matrix, as follows: ; in, Indicates the divergence feature representation; This represents a time-based max-pooling operation; Represents the skin electrophysiological feature vector, obtained by analyzing the skin electrophysiological feature sequence. Obtain by performing a linear mapping; This represents element-wise multiplication; S545. The divergence feature is represented as a reverse attention divergence after being mapped through multiple perceptron layers, and its mathematical representation is as follows: ; in, This indicates a divergence in attention; Indicates a multi-layer perceptron; S55. Weighted fusion of amplitude divergence, dynamic correlation divergence, peak relationship divergence, and inverse attention divergence yields the divergence degree, mathematically represented as follows: ; in, Indicates the degree of divergence; , , These represent the time-series arithmetic mean of the amplitude divergence, dynamic correlation divergence, and peak relationship divergence between the intensity of all facial expression responses and their corresponding skin conductance responses, respectively. , , , These represent the weight parameters for amplitude divergence, dynamic correlation divergence, peak relationship divergence, and reverse attention divergence, respectively.

[0015] Optionally, the degree of facial low reactivity is calculated as follows: ; in, Indicates a low level of facial reactivity; This represents the mean intensity of facial expression responses within a facial expression response sequence. For facial reaction stability parameters; The degree of skin conductance response is calculated as follows: ; in, Indicates the degree of skin conductance response; This represents the mean intensity of the skin electrical response within the skin electrical response sequence; For the stability parameters of skin electrical response; The emotional blunting score is calculated as follows: ; in, Indicates emotional sluggishness; Indicates the degree of divergence; , , These are the weighting parameters for the facial low reactivity level, deviation degree, and reaction type adjustment items, respectively. As a reaction type regulating term, it is defined as follows: ; in, The preset threshold for low facial reactivity; The preset threshold for the degree of skin conductance response; The criteria for determining the expression-inhibited emotional sluggishness are as follows: ; in, The preset deviation threshold; The criteria for determining low-responsive emotional sluggishness are as follows: ; The output of the emotional blunting assessment is mathematically represented as follows: ; in, The output indicates an assessment of emotional blunting; This indicates a type of emotionally sluggishness.

[0016] Optionally, a rating regression loss can be constructed, mathematically represented as follows: ; in, This represents the score regression loss; This represents the emotional blunting score predicted by the emotional blunting assessment model; This represents the true value of the affective dullness score in the clinical reference label; Indicates the calculation of the modulus; The type classification loss is constructed and mathematically represented as follows: ; in, Represents the type classification loss; Indicates the first in the clinical reference label The true value for the type of emotional detachment; Represents logarithmic calculation; This indicates the degree of facial hyporesponsiveness predicted by the emotional dullness assessment model. With skin conductance The value mapped after inputting the Sigmoid soft threshold function with a temperature coefficient belongs to the first... The probability of being emotionally sluggish is mathematically represented as follows: ; ; in, Indicates the temperature coefficient; The divergence hold loss is constructed and mathematically represented as follows: ; in, This indicates a deviation from the expected loss; Indicates the degree of deviation from the predictions made by the emotional blunting assessment model; This represents the true value of the deviation from the clinical reference label; The target loss function is obtained by weighted summing of the rating regression loss, type classification loss, and deviation preservation loss, and its mathematical expression is as follows: ; in, Represent the target loss function; , These represent the weight parameters for the type classification loss and the deviation preservation loss, respectively.

[0017] By employing the above technical solution, the present invention provides a method for assessing emotional blunting based on the divergence between facial expression and skin conductance response, which has at least the following beneficial effects: (1) This invention divides the face region into local regions of key facial features. Instead of directly classifying the entire face in a coarse-grained manner, it performs multi-scale spatial modeling of local regions of key facial features, thereby improving the ability to capture weak expressions and subtle changes in expressions. (2) This invention extracts spatial representation vectors by constructing a multi-scale spatial enhancement module. Combined with convolution and two-dimensional self-attention weighting mechanism, the emotional dullness assessment model can adaptively enhance the spatial position related to emotional expression based on the relationship between different local regions of key facial features in the current frame image, and suppress the interference caused by background, lighting changes or non-key facial feature local regions. At the same time, the construction of a hierarchical attention structure enables the emotional dullness assessment model to highlight local regions related to emotional expression such as eyebrows, eye area, and corners of the mouth in frame-level feature extraction, and reduce the influence of irrelevant factors such as hair, background, clothing edges and lighting changes on subsequent judgments. As a result, the extracted spatial representation vectors not only contain local texture changes, but also the structural relationship between different local regions of key facial features, which is more suitable for describing the weak facial expression response, slow changes and indistinct local movements of patients with depression. (3) By introducing a weakly supervised temporal attention mechanism, this invention enables spatial features to provide more discriminative input for temporal attention, better adapting to the actual problems of weak facial reactions, unclear key frames, and difficulty in obtaining frame-level annotations in patients with depression, so that the emotional bluntness assessment model can automatically determine which time points can better reflect the patient's facial reaction status. (4) This invention extracts skin electroreactivity features from both the time and frequency domains simultaneously, and uses frequency domain energy to supplement the description of the periodic changes of skin electroreactivity signals at different time scales. This enables the emotional dullness assessment model to not only rely on instantaneous peak values, but also to characterize the slower autonomic nervous response trend. At the same time, the constructed bidirectional cross-domain attention mechanism enables the emotional dullness assessment model to utilize both transient fluctuations and spectral distribution in skin electroreactivity signals, avoiding information bias caused by relying solely on time domain statistics or solely on frequency domain energy. In addition, by constructing a gating mechanism to fuse the attention relationship between time domain features and frequency domain features, the contribution ratio of time domain features and frequency domain features can be adaptively adjusted according to the performance of skin electroreactivity signals in different patients and different segments, thereby obtaining a more robust skin electroreactivity characterization. (5) This invention constructs a deviation degree by integrating amplitude deviation, dynamic correlation deviation, peak relationship deviation and reverse attention deviation, and describes the degree of abnormality in the relationship between the two types of responses. It avoids misjudgment caused by amplitude fluctuation of a single segment and is more suitable for describing the persistent mismatch between facial expression and physiological response in patients with depression. At the same time, it further strengthens low consistency segments in the deep representation space, so that the emotional bluntness assessment model can learn more complex mismatch patterns between facial overt response and skin conductance response. (6) This invention calculates the emotional blunting score to help quantify the degree of insufficient facial outward reaction and abnormal relationship between it and the skin conductance response in patients with depression. It then constructs an emotional blunting assessment output that includes deviation degree, calculation of emotional blunting score and emotional blunting type. It can not only give a quantitative score, but also explain whether the score comes from low facial reaction and the presence of skin conductance or low facial and skin conductance response, thereby enhancing the interpretability of the assessment results. (7) By constructing a target loss function that includes rating regression loss, type classification loss and deviation preservation loss, this invention avoids weakening the inconsistency information between facial expression and skin conductance during the cross-modal fusion process of the emotional blunting assessment model, so that the emotional blunting assessment model can simultaneously learn the low facial expression features, skin conductance response features and the inconsistency relationship between the two of patients with depression. Compared with the traditional multimodal fusion model, this invention does not aim to eliminate modal differences, but retains the modal differences related to emotional blunting and transforms them into assessment indicators, thus making it more suitable for the specific task of emotional blunting in depression. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the process of the emotional blunting assessment method based on the divergence between facial expression and electrodermal response of the present invention. Figure 2 This is a schematic diagram of the facial expression response sequence construction process of the present invention; Figure 3 This is a schematic diagram of the skin electroreaction sequence construction process of the present invention; Figure 4 This is a schematic diagram illustrating the process of calculating the degree of divergence and the emotional bluntness score, and determining the type of emotional bluntness in this invention. Detailed Implementation

[0019] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. This will allow for a full understanding of how the present application uses technical means to solve technical problems and achieve technical effects, and to facilitate its implementation.

[0020] Those skilled in the art will understand that all or part of the steps in the implementation of the methods of the embodiments can be implemented by a program instructing related hardware. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0021] Please refer to Figures 1-4 This illustration demonstrates a specific implementation of the present embodiment. This embodiment uses facial video sequences and electrodermal signals from patients with depression as processing objects to construct facial expression response sequences and electrodermal response sequences, respectively. The degree of discrepancy between the two sequences is calculated at a unified time scale, thereby generating an emotional blunting score and an emotional blunting type. This approach preserves and quantifies the inconsistency between overt facial responses and peripheral physiological responses, reducing the risk of misjudgment caused by relying solely on visual assessment, and improving the objectivity and interpretability of the auxiliary assessment of emotional blunting in depression.

[0022] Please refer to Figure 1 This embodiment proposes a method for assessing emotional blunting based on the divergence between facial expression and skin conductance response. The method includes the following steps: S1. Simultaneously collect facial video sequences, electrodermal signals, and clinical reference labels from patients with depression to form an emotional blunting assessment dataset.

[0023] The input data for this invention consists of a pre-acquired and synchronized video sequence of a patient with depression and electrodermal signals. As a preferred embodiment of step S1, it specifically includes: The dataset for assessing emotional blunting is defined as follows: ; in, This represents a dataset for assessing emotional blunting. Represents a facial video sequence; This represents the skin electrical signals synchronized with the facial video sequence; This indicates a clinical reference label.

[0024] The clinical reference labels, including affective retardation scores, affective retardation types, and true values ​​of deviation, are labeled by professionals. The method of this invention does not rely on a specific scale, but can utilize the scale and manual labeling to train model parameters.

[0025] S2. Perform effective segmentation and quality control. Divide the facial video sequence and electrodermal signals into several analysis segments according to a unified time axis and filter them separately to obtain several effective facial video segments and the same number of effective electrodermal segments as the effective facial video segments.

[0026] As a preferred embodiment of step S2, it specifically includes: S21. For facial video sequences First, the video is divided into several analysis segments according to a unified timeline, resulting in facial video analysis segments: ; in, This represents the total number of facial video analysis segments; Indicates the first A facial video analysis segment.

[0027] S22, corresponding to step S21, the skin electrical signal Following the same method as the facial video analysis segments, the same number of electrodermal analysis segments were obtained as the facial video analysis segments: ; in, Indicates the first There are 100 skin conductance analysis segments, and the total number of skin conductance analysis segments is the same as the total number of facial video analysis segments. .

[0028] S23. To ensure the input quality of the emotional blunting assessment model, the facial video analysis segments are filtered to remove segments with failed face detection, severe occlusion, significant head deviation, or abnormal lighting, thus obtaining the initial effective facial video segments.

[0029] S24. Filter the skin conductance analysis segments to remove segments with poor contact, sudden spikes, obvious motion artifacts, or signal saturation to obtain the initial effective skin conductance segments.

[0030] S25. Take the intersection of the initial valid facial video clips and the initial valid EKS clips on the time axis, and retain the initial valid facial video clips and initial valid EKS clips in each time interval of the intersection to obtain the same number of valid facial video clips and valid EKS clips that are time-aligned.

[0031] S3. Calculate the facial expression response intensity of each valid facial video segment and construct a facial expression response sequence.

[0032] This step extracts temporal features characterizing the intensity of facial responses from acquired facial videos of patients with depression. Unlike typical facial expression recognition tasks, this invention does not primarily target discrete expression categories such as happiness, sadness, anger, and fear. Instead, it focuses on features such as diminished amplitude of facial responses, insufficient local muscle activity, missing key expression frames, and a lack of overall dynamic variation in continuous videos. For assessing emotional blunting, the key value of facial videos lies not in determining what standard expression the patient displays, but in quantifying whether their external facial response channels are in a low-expression state. Therefore, the facial video processing is designed as a continuous modeling process involving spatial fine-grained enhancement, temporal keyframe aggregation, and response intensity curve generation.

[0033] As a preferred embodiment of step S3, it specifically includes: S31. Record the valid facial video segments as: ; in, Indicates the current patient's number One valid facial video clip, For indexing valid facial video clips; This indicates the number of currently valid facial video clips. Frame image, This represents the total number of frames in the currently valid facial video segment.

[0034] Face localization, key point detection, and face alignment are performed on each frame of the valid facial video clip to obtain the face region and several facial key points.

[0035] S32. Divide the face region into key facial feature local regions based on key facial points; the key facial feature local regions include one or more of the following: eyebrow region, periorbital region, nasal wing region, nasolabial fold region, corner of mouth region, and mandibular region.

[0036] Specifically, by detecting and obtaining the coordinates of facial key points, a geometric boundary extraction algorithm is used to determine the minimum bounding rectangle or polygon region of each key feature area. For example, a preset set of key point coordinates around the left and right eyes is selected, and the eye periorbital region is defined by calculating the extreme values ​​of the coordinates and expanding them outward by a fixed pixel width; similarly, the eyebrow region and the corner of the mouth region are defined.

[0037] The aforementioned regions are closely related to low facial expression characteristics in patients with depression, such as reduced eyebrow movement, insufficient periorbital tension changes, weakened mouth corner movements, and overall decreased facial muscle dynamics. Therefore, this invention does not directly perform coarse-grained classification of the entire face, but instead performs multi-scale spatial modeling of local areas of key facial features to improve the ability to capture weak expressions and subtle changes in facial expressions.

[0038] S33. In the spatial feature extraction stage, the present invention constructs a multi-scale spatial enhancement module to perform convolution enhancement, attention enhancement and multi-scale spatial enhancement on the spatial feature maps of each frame image in the effective facial video segment, and extracts spatial representation vectors.

[0039] The multi-scale spatial enhancement module, based on depthwise separable convolution, splits the standard convolution into two parts: channel-wise depthwise convolution and pointwise convolution. This reduces the number of parameters while preserving the local differences in different facial key feature regions. As a preferred implementation of step S33, it specifically includes: S331. Input each frame of the effective facial video segment into a convolutional neural network for feature extraction to obtain an initial full-face feature map.

[0040] In this embodiment, a shallow convolutional layer of the ResNet-18 architecture, which has been pre-trained on a face dataset, is used as the backbone network of the emotional dullness assessment model of this invention. The size-normalized images of each frame are input into ResNet-18 for hierarchical feature extraction, thereby realizing shallow semantic representation of a single frame of full-face image.

[0041] S332. Based on the geometric coordinates of the local regions of the key facial features in their respective frames, feature extraction and scale alignment are performed on the initial full-face feature map through a Region of Interest (ROI) alignment operation to obtain local feature blocks. Subsequently, multiple extracted local feature blocks are concatenated along the channel dimension to generate a spatial feature map containing fine-grained local linkage features. .

[0042] S333. Perform depthwise convolution on the spatial feature maps of each frame image, extract texture features channel by channel, and construct feature maps: ; in, This represents the feature map obtained from depthwise convolution; This indicates a depthwise convolution operation.

[0043] S334. Perform pointwise convolution on the feature map obtained by depthwise convolution to fuse information between channels: ; in, This represents the feature map obtained by pointwise convolution; Pointwise convolution operation.

[0044] To broaden the range of facial structural changes perceived by the emotional blunting assessment model, dilated convolutions are introduced into the convolutional layers of the depth convolution. This allows the model to capture the spatial dependencies between the eyebrows, eyes, nasolabial folds, and corners of the mouth without significantly increasing computational cost. It also avoids the problem that traditional local convolutions only focus on small-scale textures and ignore overall facial structural changes.

[0045] S335. Relying solely on convolutional operations is insufficient to fully characterize the relationships between distant local regions of key facial features. For example, the low expression state of patients with depression may not only manifest as weakened movement in a single region, but rather as insufficient overall coordinated change in areas such as the eyebrows, periorbital area, and corners of the mouth. Therefore, this invention further introduces a two-dimensional self-attention mechanism into the multi-scale spatial enhancement module.

[0046] Two-dimensional self-attention weighting is performed on the feature map obtained by pointwise convolution to generate a self-attention response feature map, mathematically represented as follows: ; ; in, Indicates the self-attention weights; This indicates the Softmax normalization operation; , , These represent the query matrix, key matrix, and value matrix, respectively, and are feature maps obtained through pointwise convolution. Obtained by performing a linear mapping; The key matrix represents a feature dimension; the query matrix, key matrix, and value matrix all share the same feature dimension. Indicates the transpose operation; This represents the self-attention response feature map. The two-dimensional self-attention weighting mechanism enables the emotion blunting assessment model to adaptively strengthen spatial locations related to emotion expression based on the interrelationships between different local regions of key facial features in the current frame image, while suppressing interference from background, lighting changes, or local regions of non-key facial features.

[0047] S336. In order to further improve the filtering capability of single-frame features, a hierarchical attention structure combining channel attention and spatial attention is adopted. Channel attention is used to determine the contribution of different feature channels to the representation of facial expression intensity, while spatial attention is used to locate the region where expression changes are more concentrated.

[0048] Channel attention weights and spatial attention weights are calculated separately for the self-attention response feature maps, and then fused to obtain spatially enhanced features, mathematically represented as follows: ; ; ; in, , These represent channel attention weights and spatial attention weights, respectively. Indicates the Sigmoid activation operation; Indicates a multi-layer perceptron; This indicates an average pooling operation; This represents the max pooling operation; Indicates the convolution operation; This represents the channel average pooling operation performed along the channel dimension; This represents the channel max pooling operation performed along the channel dimension; Indicates a splicing operation; Represents spatially enhanced features; This indicates element-wise multiplication.

[0049] In the above operations, and By operating on the spatial dimension of the feature map, the spatial resolution is compressed within each feature channel to extract the global channel descriptor, which is then used to calculate the channel attention weights. and This operation, applied to the channel dimension of the feature map, calculates the average and maximum values ​​across channels at each spatial pixel location in the full feature map, compressing multi-dimensional channel information into a two-dimensional representation of a single spatial plane. This is used to accurately locate and capture the spatial distribution features of subtle facial movements, and then calculates the spatial attention weights.

[0050] Hierarchical attention structures enable emotion dullness assessment models to highlight local areas related to emotion expression, such as eyebrows, eye area, and corners of the mouth, during frame-level feature extraction, while reducing the influence of irrelevant factors such as hair, background, clothing edges, and lighting changes on subsequent judgments.

[0051] S337. Perform spatial dimension reduction and flattening on the spatial enhancement features to generate spatial representation vectors for each frame image. The mathematical representation of this process is as follows: ; in, This represents a global average pooling operation, used to compress across the height and width dimensions, eliminating spatially redundant noise and transforming it into channel-level feature vectors, thus enhancing spatial features. Compression of two-dimensional spatial features into one-dimensional vectors; This indicates a fully connected layer, used for flattening. Indicates the first Spatial representation vector of a frame image.

[0052] The spatial representation vector extracted by this invention not only includes local texture changes, but also the structural relationships between different local regions of key facial features, which is more suitable for describing the weak amplitude, slow changes and indistinct local movements of facial expressions in patients with depression.

[0053] S34. In the temporal feature modeling stage, a weakly supervised temporal attention mechanism is introduced to perform weighted fusion of the spatial representation vectors of each frame image in the effective facial video segments based on temporal attention weights, thereby generating segment-level facial expression response vectors for each effective facial video segment.

[0054] The weakly supervised temporal attention mechanism introduced in this invention is used to automatically identify key frames with high facial expression value from effective facial video segments with only fragment-level annotation. Facial reactions of patients with depression are usually not continuous, intense or stable. If all frames are simply averaged, key subtle changes are easily buried by a large number of calm frames. Therefore, this invention assigns different weights to different frames through temporal attention, enabling the emotional blunting assessment model to automatically determine which time points better reflect the patient's facial reaction state.

[0055] As a preferred embodiment of step S34, it specifically includes: S341. Calculate the temporal attention weights for the spatial representation vectors of each frame image. Simultaneously, to prevent the emotional blunting assessment model from being affected by low-quality or invalid still frames, a temporal attention threshold is further set to filter out low temporal attention weights. This process is mathematically represented as follows: ; ; in, Indicates the first valid facial video segment Temporal attention weights for frame images; This indicates the Softmax normalization operation; This is the attention scalar mapping vector; This is the time-mapped weight matrix; This is the time-mapped bias vector; Indicates the transpose operation; Represents the hyperbolic tangent function; Indicates the first valid facial video segment The spatial representation vector of a frame image; Indicates the first valid facial video segment Temporal attention weights after filtering of frame images; Temporal attention threshold

[0056] S342. Normalize the filtered temporal attention weights to obtain normalized frame-level weights. The mathematical representation of this process is as follows: ; in, Indicates the first valid facial video segment Normalized frame-level weights for frame images;

[0057] S343. The spatial representation vectors of each frame in the effective facial video segment are weighted and fused using normalized frame-level weights to generate the segment-level facial expression response vector of the effective facial video segment. This process is mathematically represented as follows: ; in, Indicates the first A segment-level facial expression response vector for each valid facial video clip.

[0058] Fragment-level facial expression response vectors represent the overall facial outward responses in effective facial video segments after spatial enhancement and temporal aggregation. Unlike traditional 3D convolution that directly couples to extract spatiotemporal features, this step first performs fine-grained spatial enhancement and then weakly supervised temporal aggregation, making spatial features provide more discriminative input for temporal attention. This approach is better suited to the practical problems of weak facial responses, unclear keyframes, and difficulty in obtaining frame-level annotations in patients with depression.

[0059] S35. Extract facial reference representations based on each effective facial video segment, calculate the facial expression response intensity of each effective facial video segment relative to the facial reference representation using segment-level facial expression response vectors, and construct a facial expression response sequence.

[0060] As a preferred embodiment of step S35, it specifically includes:

[0061] The facial benchmark is defined as a stable segment or sliding average among each valid facial video segment.

[0062] The stable segments are determined by calculating the instantaneous variance of the motion displacement of facial key points within each valid facial video segment. Valid facial video segments with an instantaneous variance lower than a preset stability threshold are considered expressionless resting state segments and are thus identified as stable segments. The moving average is obtained by calculating the arithmetic mean of the segment-level facial expression response vectors of the current valid facial video segment and its preceding adjacent valid facial video segments through a preset time sliding window.

[0063] The calculation of the facial expression response intensity of each effective facial video segment relative to the facial reference using segment-level facial expression response vectors is mathematically represented as follows: ; in, Represents facial reference characteristics; Indicates the first Fragment-level facial expression response vectors for each valid facial video segment. Indicates the index of valid facial video clips. Total number of valid facial video clips; This represents the L2 norm computation operator, used to measure the Euclidean distance between two vectors; Indicates the first A valid facial video segment relative to a facial reference representation The intensity of facial expression responses. If If the level remains consistently low, it indicates that the patient exhibits insufficient facial movement variation across multiple valid facial video clips, exhibiting characteristics of low facial responsiveness.

[0064] Facial benchmark representation This serves as a visual reference for patients whose emotions are not aroused or who are in a basal state, and includes two specific implementation methods: The first type is the relatively stable segment: A pre-defined, unstimulated resting state phase is introduced when acquiring facial video sequences. Within this phase, the inter-frame motion or geometric displacement of key facial points (such as eyebrows, periorbital area, and corners of the mouth) in each facial video analysis segment is extracted, and their instantaneous variance within the current segment time is calculated. When the instantaneous variance of the geometric displacement within a valid facial video segment is lower than a pre-defined stability threshold, it indicates that the patient's face is not actively expressing anything and muscle tension is stable. Therefore, the segment-level facial expression response vector of this segment is used as the facial baseline representation of the relatively stable segment. ; The second method is moving average: To eliminate the gradual interference from slow illumination drift and minor head pose adjustments on global features, a time-sliding window is set. The size of this window is preset to include the current valid facial video segment. A number of consecutive adjacent segments (of which (This refers to the span of neighboring segments before and after the current segment). The facial expression response vectors at the segment level within this time sliding window are summed temporally and their arithmetic mean is taken. This mean vector is then used as the dynamic facial reference representation for the current segment. Its mathematical expression is: , For the summation index, Indicates the first A segment-level facial expression response vector for each valid facial video clip.

[0065] After obtaining the facial reference representation, the facial expression response intensity of each effective facial video segment relative to the facial reference representation is calculated using segment-level facial expression response vectors.

[0066] The facial expression response sequence is mathematically represented as follows: ; in, This represents a sequence of facial expression responses.

[0067] Facial expression response sequences do not represent specific emotion categories, but rather the degree of intensity of facial overt responses over time. They form the visual basis for subsequent calculations of the divergence between facial expressions and electrodermal responses.

[0068] The facial expression response sequence construction process of this invention can be referred to Figure 2 .

[0069] S4. Calculate the skin electrical response intensity of each effective skin electrical segment and construct the skin electrical response sequence.

[0070] This step is used to extract skin conductance response sequences from the obtained skin conductance signals that can characterize changes in peripheral autonomic nervous system responses. Skin conductance activity can reflect changes in sweat gland activity, and sweat gland activity has a clear connection with sympathetic nervous system activity, making it suitable for characterizing the peripheral physiological responses of patients with depression in video recordings. This invention does not presuppose that the skin conductance response of patients with depression is necessarily elevated, nor does it directly use the strength of skin conductance as a diagnostic criterion. Instead, it uses skin conductance response as a physiological reference for modeling the relationship between skin conductance response and facial external responses. As a preferred embodiment of step S4, it specifically includes:

[0071] S41. The effective skin electrical activity fragment is denoted as: ; in, Indicates the current patient's number One effective skin electrical activity segment, For effective skin electrodermal fragment indexing; Indicates the first Number of sampling points within an effective skin electrodermal segment.

[0072] First, each effective electrodermal segment is preprocessed by denoising, outlier removal, smoothing, and standardization to reduce the impact of sensor contact instability, instantaneous motion artifacts, and differences in signal amplitude scale.

[0073] Subsequently, the preprocessed effective skin conduction fragments were subjected to low-pass filtering to extract their tension components: ; in, Indicates the first Tension components of an effective skin electrical segment; This indicates a low-pass filter operation.

[0074] S42. Difference between each effective skin electrical segment and its tension component to obtain its phase component: ; in, Indicates the first Phase components of an effective skin electrical segment.

[0075] The tension component reflects slower changes in skin conductance levels, while the phase component reflects relatively rapid event-related fluctuations in skin conductance. In this invention, the tension component is used to describe the slow trend of changes in a patient's skin conductance activity, while the phase component is used to describe faster changes in autonomic nervous system responses.

[0076] S43. Because patients with depression may experience weakened, delayed, or asynchronous autonomic nervous system responses with facial expressions, simply using the average value of skin conductance cannot fully characterize their physiological state. This invention extracts skin conductance response characteristics from both the time and frequency domains.

[0077] In time-domain modeling, this invention extracts local dynamic features and global trend features for phase components. For each effective electrodermal cell segment, the local mean, fluctuation degree, peak amplitude, and energy of its phase components are calculated. The mathematical representation of this calculation method is as follows: ; ; ; ; in, , , , They represent the first Local mean, variability, peak amplitude, and energy of an effective skin electrodermal segment; Indicates the index of valid skin electrical fragments. The total number of effective skin electrodermal fragments is equal to the total number of effective facial video fragments; Indicates the first Number of sampling points within an effective skin electrical segment; Indicates the first Within the first effective skin conductance fragment The signal amplitude at each sampling point is the phase component sequence obtained after decoupling the electrodermal signal. Indicates the sampling point index. Indicates in Values ​​range from 1 to Take within range The maximum value;

[0078] Furthermore, the rise slope, recovery time, and number of peaks can be calculated to describe whether the skin conductance response has a significant transient activation process.

[0079] S44. The above-mentioned time-domain features can reflect the amplitude, duration and local fluctuation of the skin conductance response, but they are still insufficient to express the frequency structure of the skin conductance signal at different time scales.

[0080] In frequency domain modeling, this invention performs Fourier transform on the phase components to extract energy distribution features within different frequency ranges, obtaining the time-frequency representation of each effective electrodermal segment: ; in, Indicates the time offset of the phase component; is the rotation factor of the Fourier transform. For frequency variables, The imaginary unit; Represents a continuous physiological time series variable; Indicates a timing sliding window function; For the first Time shift of an effective skin electrodermal fragment and frequency variables Time-frequency representation; Indicates in Value Within the range of expressions Seeking information about The points.

[0081] S45. Calculate the energy within the target frequency band of each effective skin electrodermal segment based on time-frequency representation: ; in, No. Energy within the target frequency band of an effective skin electrodermal segment; , Representing frequency variables respectively The upper and lower bounds of the frequency band define the frequency range that the two bounds define; Indicates the calculation of the square of the modulus; Indicates in Value Within the range of expressions Seeking information about The integral of frequency domain energy can supplement the description of the periodic changes of electrodermal signals at different time scales, enabling the emotional blunting assessment model to not only rely on instantaneous peak values ​​but also to characterize slower autonomic neural response trends.

[0082] S46. In order to fuse time-domain and frequency-domain information, the present invention constructs a time-frequency fusion encoder for electrodermatology.

[0083] Let the first The temporal feature vector of each effective skin conductance segment is: , No. The frequency domain feature vector of each effective skin conductance segment is First, the two are projected onto the same dimension using a linear mapping to obtain the time-domain features and frequency-domain features: ; ; in, , They represent the first The first time-domain feature and the second Individual frequency domain features; , These represent the learnable weights and biases of the projection layer that projects the temporal feature vectors, respectively. , These represent the learnable weights and biases of the projection layer that project the frequency domain feature vectors, respectively.

[0084] S47. Perform cross-domain attention calculation between time-domain features and frequency-domain features to obtain the attention relationship between them. The mathematical representation of this process is as follows: ; ; in, This indicates the relationship of interest between time-domain features and frequency-domain features; This indicates the relationship of interest between frequency domain features and time domain features; This indicates the Softmax normalization operation; , , These represent the time-domain query matrix, time-domain key matrix, and time-domain value matrix, respectively, obtained through linear mapping of time-domain features; , , These represent the frequency domain query matrix, frequency domain key matrix, and frequency domain value matrix, respectively, obtained through linear mapping of frequency domain features; Indicates the transpose operation; Represents the characteristic dimension of the time-domain key matrix. The scaling factor used for dot product attention is that the time-domain query matrix, time-domain key matrix, time-domain value matrix, frequency-domain query matrix, frequency-domain key matrix, and frequency-domain value matrix all have the same feature dimension.

[0085] This bidirectional cross-domain attention mechanism enables the emotional blunting assessment model to simultaneously utilize transient fluctuations and spectral distribution in electrodermal signals, avoiding information bias caused by relying solely on time-domain statistics or frequency-domain energy.

[0086] S48. Gated fusion of the interest relationships between time-domain features and frequency-domain features is performed to generate a skin conductance response vector. The mathematical representation of this process is as follows: ; ; in, Indicates the first The gating coefficient of an effective skin electrodermal segment; Indicates the Sigmoid activation operation; , These represent the learnable gating weights and the gating bias term, respectively. Indicates a splicing operation; Indicates the first The skin electrical response vector of an effective skin electrical segment; This represents element-wise multiplication;

[0087] The aforementioned gating mechanism can adaptively adjust the contribution ratio of time-domain and frequency-domain features based on the performance of electrodermal signals in different patients and different segments, thereby obtaining a more robust characterization of electrodermal response.

[0088] S49. In order to convert the electrodermal response into a time series that can correspond to the facial expression response sequence, the present invention further calculates the response intensity of each effective electrodermal segment relative to the patient's electrodermal baseline characterization.

[0089] Extract the baseline skin electrical activity (TEA) characterization and calculate the TEA response intensity of each effective TEA segment relative to this baseline characterization. The mathematical representation is as follows: ; in, Indicates the baseline characterization of skin electrical activity; Indicates the first Characterization of effective skin conductance segments relative to skin conductance baseline The intensity of skin conductance response; The L2 norm calculation operator is used. The skin electrophysiological (EEM) baseline is characterized as a stationary segment or a moving average among each effective EEM segment. A stationary segment is determined by calculating the instantaneous variance of the signal amplitude at sampling points within each effective EEM segment and comparing it with a preset stationarity threshold. The moving average is obtained by calculating the arithmetic mean of the EEM response vectors of consecutive adjacent effective EEM segments within a preset time-series sliding window. Thus, the EEM response sequence is obtained, mathematically represented as follows: ; in, This represents the skin conductance response sequence.

[0090] The skin conductance response sequence describes the strength of the patient's skin conductance response as a function of video time, rather than simply the absolute value of skin conductance. This approach can reduce the impact of individual differences in baseline skin conductance levels on the results and provide a physiological reference at a unified time scale for subsequent calculations of the divergence between facial expressions and skin conductance responses.

[0091] The skin electroreactivity sequence construction process of this invention can be referred to Figure 3 .

[0092] S5. Construct a divergence calculation module to calculate the divergence between the facial expression response sequence and the skin conductance response sequence.

[0093] This step is used to calculate the degree of inconsistency between facial expression response sequences and skin conductance response sequences at a unified time scale. Since the two sequences have been mapped at the segment level along the same video timeline, the divergence can be calculated at three levels: segment level, sliding window level, and deep feature level. As a preferred embodiment of step S5, it specifically includes: S51. Calculate the amplitude divergence between the intensity of each facial expression response and its corresponding skin conductance response intensity, mathematically represented as follows: ; in, Indicates the first Facial expression intensity of a valid facial video clip The corresponding first Skin electroreactivity intensity of each effective skin electroreactivity segment The amplitude divergence between them ; This indicates a normalization operation; in this embodiment, max-min normalization is used.

[0094] Amplitude divergence describes the difference between the intensity of facial expression responses and the intensity of skin conductance responses within the same segment. If a patient's facial response remains consistently weak while their skin conductance response is relatively strong, the amplitude divergence is elevated. If both facial and skin conductance responses are low or both are high, further interpretation based on subsequent correlation and type determination is required. Therefore, amplitude divergence is not the sole criterion for judgment, but rather a fundamental component of the degree of divergence.

[0095] S52. Define sliding windows for extracting facial expression response sequences and skin conductance response sequences, respectively. , Facial expression response subsequence was obtained. With skin electroreactivity sequence Considering that there may be a certain physiological delay in the skin conductance response relative to visible facial changes, this invention operates within an acceptable delay range. Internal calculation of maximum correlation: ; in, Indicates the first The intensity of facial expression response in a valid facial video clip and its corresponding... The dynamic correlation between the intensity of skin electrical response of each effective skin electrical segment diverged; This indicates the time delay between visual response and physiological sympathetic nerve response. This indicates the preset allowable time delay range; Indicates in Evaluate expression within range The maximum value; This indicates the operator for calculating the Pearson correlation coefficient; Represents a sliding window Superimposed time delay The extracted skin electroreactivity sequence.

[0096] Dynamic correlation divergence is used to describe whether facial expressions and electrodermal responses exhibit synchronous changes over a continuous period. When facial expression sequences and electrodermal response sequences lack synchronous changes over a period of time, the intensity of facial expression responses... Increase. Facial expression response intensity. It can avoid misjudgments caused by fluctuations in amplitude of a single segment and is more suitable for describing the persistent mismatch between facial expressions and physiological reactions in patients with depression.

[0097] S53, Obtain the sliding window respectively , Inner facial expression response subsequence With skin electroreactivity sequence peak time , Furthermore, the peak deviation between the intensity of each facial expression response and its corresponding skin conductance response intensity is calculated, and the mathematical representation is as follows: ; in, Indicates the first The intensity of facial expression response in a valid facial video clip and its corresponding... The peak relationship between the skin electrical response intensities of each effective skin electrical segment deviates; This indicates the calculation of the modulus.

[0098] Peak relationship divergence is used to describe the correspondence between facial expression response peaks and skin conductance response peaks. In this embodiment, if the peak of the facial expression response subsequence is missing while the peak of the skin conductance response subsequence is present, it is recorded as a facial peak missing divergence; if both peaks are missing, the current sliding window is more likely to be a low-response type; if the two peaks correspond to each other within the allowable delay range, the peak relationship divergence is low. This processing method can distinguish between two different states: no facial response but skin conductance response, and no obvious facial or skin conductance response.

[0099] S54. Calculate the inverse attentional divergence between facial expression response sequences and skin conductance response sequences, including: S541, Original Facial Expression Response Sequence With skin electrical response sequence Since the facial expression response sequence and the electrodermal response sequence are scalar time series with one-dimensional intensity, direct high-dimensional matrix dot products and deep feature interactions cannot be performed. Therefore, two parallel one-dimensional temporal convolutional layers are constructed. The facial expression response sequence and the electrodermal response sequence are fed into the one-dimensional temporal convolutional layer for local temporal sliding window convolution and channel dimensionality upscaling, mapping them from one-dimensional scalar sequences to feature channels with a preset number of channels. A high-dimensional temporal tensor is used to construct facial feature sequences in a high-dimensional feature space. With skin electrophysiological feature sequence This lays the foundation for feature representation in the subsequent establishment of deep spatial cross-modal association.

[0100] S542. Calculate facial feature sequences With skin electrophysiological feature sequence The mathematical representation of the standard cross-modal attention matrix is ​​as follows: ; in, This represents a standard cross-modal attention matrix, reflecting the consistent correspondence between facial features and electrodermal features; This indicates the Softmax normalization operation; This represents the facial feature query matrix obtained by linear mapping of facial features in a facial feature sequence. This represents the skin electrophysiological feature key matrix obtained by linear mapping of skin electrophysiological features in a skin electrophysiological feature sequence; Indicates the transpose operation; The feature dimension represents the column features of the electrodermal feature key matrix after linear mapping in the feature space. The facial feature query matrix and the electrodermal feature key matrix have the same feature dimension. This is then used as the corresponding dot product scaling factor.

[0101] S543, Unlike traditional multimodal fusion methods that directly enhance conventional cross-modal attention matrices. The present invention further constructs a reverse attention matrix for the consistent segments represented.

[0102] Calculate facial feature sequences With skin electrophysiological feature sequence The inverse cross-modal attention matrix between them is mathematically represented as follows: ; in, Representation compared to a conventional cross-modal attention matrix A matrix with the same dimension and all elements being 1; This represents the inverse cross-modal attention matrix, used to highlight locations where facial responses and skin conductance responses have low consistency, enabling the emotional blunting assessment model to focus on time segments where the two do not match.

[0103] S544. Calculate the divergence feature representation based on the inverse cross-modal attention matrix, as follows: ; in, Indicates the divergence feature representation; This represents a time-dimension max pooling operation, used to extract the deepest features with the strongest conflict. Represents the skin electrophysiological feature vector, obtained by analyzing the skin electrophysiological feature sequence. Obtain by performing a linear mapping; This indicates element-wise multiplication.

[0104] S545. The divergence feature is represented as a reverse attention divergence after being mapped through multiple perceptron layers, and its mathematical representation is as follows: ; in, This indicates a divergence in attention; This represents a multi-layer perceptron.

[0105] Inverse attention divergence is used to extract inconsistent fragments that are difficult to capture by traditional statistical indicators. Its role is not to replace amplitude, correlation and peak indicators, but to further amplify low-consistency fragments in the deep representation space, enabling the emotional blunting assessment model to learn more complex mismatch patterns between facial overt responses and skin conductance responses.

[0106] S55. Weighted fusion of amplitude divergence, dynamic correlation divergence, peak relationship divergence, and inverse attention divergence yields the divergence degree, mathematically represented as follows: ; in, Indicates the degree of divergence; , , These represent the time-series arithmetic mean of the amplitude divergence, dynamic correlation divergence, and peak relationship divergence between the intensity of all facial expression responses and their corresponding skin conductance responses, respectively. , , , These represent the weight parameters for amplitude divergence, dynamic correlation divergence, peak relationship divergence, and reverse attention divergence, respectively.

[0107] Deviation The higher the value, the greater the inconsistency between the patient's facial overt response and the skin electrophysiological response. This indicator is neither the intensity of facial expression nor the intensity of skin conductance, but rather a quantitative result of the degree of abnormality in the relationship between the two types of responses.

[0108] The calculation process for the deviation degree and emotional bluntness score, and the process for judging the type of emotional bluntness in this invention can be found in the following: Figure 4 .

[0109] S6. Calculate the degree of low facial response and the degree of electrical skin response based on the facial expression response sequence and the electrical skin response sequence. Then, combine the deviation degree to calculate the emotional dullness score and determine the type of emotional dullness, and construct the emotional dullness assessment output.

[0110] After obtaining the deviation between the facial expression response sequence and the skin conductance response sequence, this invention further generates an emotional blunting score. This score does not directly replace clinical diagnosis, but is used to assist in quantifying the degree of insufficient facial overt responses in patients with depression and the abnormal relationship between them and skin conductance responses. As a preferred embodiment of step S6, it specifically includes:

[0111] S61. Calculate the degree of low facial reactivity based on the facial expression response sequence, as follows: ; in, Indicates a low level of facial reactivity. The higher the value, the weaker the patient's facial expression changes and the less expressive their reactions. This represents the mean intensity of facial expression responses within a facial expression response sequence. For facial response stability parameters.

[0112] S62. Calculate the degree of skin conductance response based on the skin conductance response sequence. The calculation method is as follows: ; in, Indicates the degree of skin conductance response. The higher the value, the more pronounced the skin conductance response; The lower the value, the weaker the skin conductance response; This represents the mean intensity of the skin electrical response within the skin electrical response sequence; This is a stable parameter for skin conductance response.

[0113] S63. Calculate the emotional retardation score as follows: ; in, Indicates emotional sluggishness; Indicates the degree of divergence; The response type modifier is used to distinguish different types of emotional sluggishness, ensuring that the emotional sluggishness score reflects not only the degree of facial hyporesponsiveness but also the relationship between facial response and skin conductance. It is defined as the following piecewise function: ; in, The preset threshold for low facial reactivity; The preset threshold for the degree of skin conductance response; , , These are the weighting parameters for the facial low reactivity level, deviation degree, and reaction type adjustment terms, respectively.

[0114] The specific construction logic and clinical significance of piecewise functions in the emotional blunting assessment task are as follows: When the first stage condition is met (i.e., low facial reactivity) Greater than the preset threshold And the degree of skin conductance response Greater than the preset threshold When this occurs, the sensitive regulatory mechanism of expression-inhibited emotional sluggishness is activated. These patients exhibit a clear conflict between limited external behavioral expression and strong internal sympathetic arousal. The model utilizes this conflicting characteristic... and Nonlinear multiplication and weighting are applied as an adjustment term to amplify the divergence coupling gain in the representation space, thereby increasing the final emotional blunting score. This approach can effectively capture high-risk patients who are experiencing inner anguish but remain expressionless, allowing for subsequent assessment using a global deviation threshold. This provides a robust scoring basis for the final diagnosis of expression-inhibited emotional retardation.

[0115] When the second stage condition is met (i.e., low facial reactivity) Greater than the preset threshold And the degree of skin conductance response Less than or equal to the preset threshold When the patient was found to be in a state of low activation both physiologically and behaviorally, the compensatory mechanism of low-responsive emotional sluggishness was activated. Because both their facial outward behavior and peripheral sympathetic nervous system were in a state of widespread low activation and low response, the calculated global deviation was significant. The values ​​will experience a pathological natural decline, introducing... As a moderating output, the robustness and fairness of the low-response emotional sluggishness score are ensured by the inverse incentive characteristic that the lower the autonomous physiological response, the greater the compensation gain.

[0116] If facial low reactivity The preset threshold was not reached. This indicates that the facial emotion expression channel is in a healthy, normally activated response state. In this case, the moderating term output is directly set to 0, and no subtyping intervention is initiated. Through this nonlinear conditional piecewise moderating term design guided by medical prior knowledge, the model is endowed with a high discriminative quantification ability for multiple heterogeneous emotional blunting patterns, making the output results clinically valuable and interpretable.

[0117] S64. Based on this, the present invention classifies emotional blunting-related reactions into at least two categories.

[0118] S641. The first category is emotional sluggishness due to expression inhibition, and the criteria for determination are as follows: ; in, This is the preset deviation threshold. Expression inhibition-type emotional blunting indicates that the patient's facial outward response is weak, but the skin electrophysiological response is still present, and there is a significant deviation between facial expression and peripheral physiological response.

[0119] S642. The second category is low-responsive emotional sluggishness, and the criteria for determination are as follows: ; Low-responsive emotional blunting indicates that patients not only have weaker facial expression changes but also weaker skin conductance, suggesting that they may be exhibiting a more generalized state of low-activity response.

[0120] In actual implementation, if the facial reactivity level is low If the threshold is not reached, or if the relationship between facial reaction and skin conductance response is stable and consistent, a high emotional dullness result will not be output; for samples with insufficient video quality, severe skin conductance signal artifacts, or insufficient number of effective segments, an uncertain result will be output to avoid misjudging low-quality data as emotional dullness.

[0121] S65. Construct the output of the emotional blunting assessment, mathematically represented as: ; in, The output indicates an assessment of emotional blunting; This indicates an emotionally sluggish type. Emotionally sluggish type The degree of facial low reactivity obtained directly through calculation skin conductance response level and the degree of divergence Logical condition determination is performed against the corresponding preset threshold to obtain the result. When both conditions are met... At that time, determine the type of emotional sluggishness This is characterized by expression-inhibited emotional sluggishness; when simultaneously satisfying... At that time, determine the type of emotional sluggishness This is a low-response emotional sluggishness. The way the emotional sluggishness assessment output is constructed allows this invention not only to provide a quantitative score, but also to explain whether the score comes from low facial reactivity and the presence of skin conductance, or from low facial and skin conductance reactivity, thereby enhancing the interpretability of the assessment results.

[0122] S7. Construct a target loss function and train the emotional sluggishness assessment model using the emotional sluggishness assessment dataset; the emotional sluggishness assessment model takes facial video sequences and electrodermal signals as inputs and predicts the emotional sluggishness assessment output through steps S2-S6.

[0123] The training process of this invention does not employ a single fusion classification objective, but instead simultaneously constrains facial low-response representation, skin conductance representation, and the divergence between the two. As a preferred embodiment of step S7, constructing the target loss function specifically includes: S71. Construct the rating regression loss, mathematically represented as follows: ; in, This represents the score regression loss; This represents the emotional blunting score predicted by the emotional blunting assessment model; This represents the true value of the affective dullness score in the clinical reference label; This indicates the calculation of the modulus.

[0124] S72. Construct the type classification loss, mathematically represented as follows: ; in, Represents the type classification loss; Indicates the first in the clinical reference label The true value for the type of emotional detachment; Represents logarithmic calculation; This indicates the degree of facial hyporesponsiveness predicted by the emotional dullness assessment model. With skin conductance The value mapped after inputting the Sigmoid soft threshold function with a temperature coefficient belongs to the first... The probability of being emotionally sluggish is mathematically represented as follows: ; ; in, This represents the temperature coefficient.

[0125] S73. To ensure that the emotional blunting assessment model does not weaken the inconsistency information between facial and skin conductance during cross-modal fusion, a divergence preservation loss is further constructed, mathematically represented as follows: ; in, This indicates a deviation from the expected loss; Indicates the degree of deviation from the predictions made by the emotional blunting assessment model; This represents the true value of the deviation from the clinical reference label.

[0126] S74. The target loss function is obtained by weighted summing of the rating regression loss, the type classification loss, and the deviation preservation loss, as follows: ; in, Represent the target loss function; , These represent the weight parameters for the type classification loss and the deviation preservation loss, respectively.

[0127] Through this joint optimization approach, the affective blunting assessment model can simultaneously learn facial low-expression features, skin conductance characteristics, and the inconsistency between the two in patients with depression. Compared to traditional multimodal fusion models, this invention does not aim to eliminate modal differences, but rather preserves modal differences related to affective blunting and transforms them into assessment indicators, thus making it more suitable for the specific task of assessing affective blunting in depression.

[0128] S8. Real-time acquisition of facial video sequences and skin electrical signals of the patient, and prediction of emotional dullness assessment output by the trained emotional dullness assessment model, which serves as the patient's emotional dullness assessment result.

[0129] This invention constructs facial expression response sequences and skin conductance response sequences separately, and calculates the degree of divergence between the two. This allows the assessment results to reflect not only the strength of the facial overt response, but also the correspondence between facial expression and physiological response. This reduces the risk of underestimation caused by relying solely on visual signals, and avoids directly equating an absolute increase or decrease in skin conductance response with emotional dullness. Furthermore, different response types can be output based on the degree of low facial response, the degree of skin conductance response, and the degree of divergence, thereby improving the interpretability and distinguishability of emotional dullness assessment results.

[0130] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0131] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0132] The above embodiments provide a detailed description of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for assessing emotional blunting based on the divergence between facial expression and electrodermal response, characterized in that, include: S1. Simultaneously collect facial video sequences, electrodermal signals, and clinical reference labels from patients to form an emotional dullness assessment dataset; the clinical reference labels include the patient's emotional dullness score, emotional dullness type, and true value of deviation degree. S2. Divide the facial video sequence and the electrodermal signal into several analysis segments according to a unified time axis and filter them separately to obtain several effective facial video segments and the same number of effective electrodermal segments as the effective facial video segments. S3. Calculate the facial expression response intensity of each valid facial video segment and construct a facial expression response sequence; S4. Calculate the skin electrical response intensity of each effective skin electrical segment and construct the skin electrical response sequence; S5. Calculate the deviation between the facial expression response sequence and the skin conductance response sequence; the deviation is obtained by weighted fusion of amplitude deviation, dynamic correlation deviation, peak relationship deviation, and inverse attention deviation. S6. Calculate the degree of low facial reactivity and the degree of electrical skin reactivity based on the facial expression response sequence and the skin conductance response sequence, and then calculate the emotional dullness score by combining the deviation degree and determine the type of emotional dullness, and construct the emotional dullness assessment output; the emotional dullness type includes at least expression inhibition type emotional dullness and low reactivity type emotional dullness. S7. Construct a target loss function and train the emotional sluggishness assessment model using the emotional sluggishness assessment dataset; the emotional sluggishness assessment model takes facial video sequences and electrodermal signals as inputs and predicts the emotional sluggishness assessment output through steps S2-S6. S8. Real-time acquisition of facial video sequences and skin electrical signals of the patient, and prediction of emotional dullness assessment output by the trained emotional dullness assessment model, which serves as the patient's emotional dullness assessment result.

2. The method for assessing emotional blunting based on the divergence between facial expression and electrodermal response as described in claim 1, characterized in that: The process of dividing the facial video sequence and electrodermal signals into several analysis segments along a unified time axis and filtering them separately yields several valid facial video segments and the same number of valid electrodermal signals as the valid facial video segments, including: The facial video sequence is divided into several analysis segments to obtain facial video analysis segments; The skin electrical signals were divided in the same way as the facial video analysis segments to obtain the same number of skin electrical analysis segments as the facial video analysis segments; The facial video analysis segments are filtered to remove segments that fail to detect faces, have severe occlusion, have large head rotation, or have abnormal lighting, in order to obtain the initial valid facial video segments. The skin conductance analysis segments are filtered to remove segments with poor contact, sudden spikes, obvious motion artifacts, or signal saturation, thus obtaining the initial effective skin conductance segments; The intersection of the initial valid facial video clips and the initial valid EEG clips on the time axis is taken, and the initial valid facial video clips and the initial valid EEG clips in each time interval of the intersection are retained to obtain the same number of valid facial video clips and valid EEG clips that are time-aligned.

3. The method for assessing emotional blunting based on the divergence between facial expression and electrodermal response as described in claim 1, characterized in that: The calculation of facial expression response intensity for each valid facial video segment and the construction of a facial expression response sequence include: Face localization, key point detection and face alignment are performed on each frame of the effective facial video segment to obtain the face region and several facial key points. The face is divided into key facial feature regions based on key facial points; the key facial feature regions include one or more of the following: eyebrow region, periorbital region, nasal alar region, nasolabial fold region, corner of mouth region, and mandibular region. Each frame of the effective facial video segment is input into a convolutional neural network for feature extraction to obtain an initial full-face feature map. Based on the geometric coordinates of the local regions of the key facial features, several local feature blocks are extracted from the initial full-face feature map through region of interest alignment. Then, the local feature blocks are cascaded to obtain a spatial feature map. A multi-scale spatial enhancement module is constructed to perform convolution enhancement, attention enhancement, and multi-scale spatial enhancement on the spatial feature maps of each frame in the effective facial video segment, and to extract spatial representation vectors. A weakly supervised temporal attention mechanism is introduced to perform weighted fusion of the spatial representation vectors of each frame in the effective facial video clips, generating the clip-level facial expression response vectors of each effective facial video clip. Facial baseline representations are extracted from each valid facial video segment. The facial expression response intensity of each valid facial video segment relative to the facial baseline representation is calculated using segment-level facial expression response vectors, and a facial expression response sequence is constructed.

4. The method for assessing emotional blunting based on the divergence between facial expression and electrodermal response as described in claim 3, characterized in that: The multi-scale spatial enhancement module performs convolution enhancement, attention enhancement, and multi-scale spatial enhancement on the spatial feature maps of each frame in the effective facial video segment, extracting spatial representation vectors, including: Deep convolution is performed on each frame of the valid facial video clip to extract texture features channel by channel and construct feature maps. Pointwise convolution is performed on the feature map obtained by depthwise convolution to achieve information fusion between channels; Two-dimensional self-attention weighting is performed on the feature map obtained by pointwise convolution to generate a self-attention response feature map, mathematically represented as follows: ; ; in, Indicates the self-attention weights; This indicates the Softmax normalization operation; , , These represent the query matrix, key matrix, and value matrix, respectively, obtained by linearly mapping the feature map obtained through pointwise convolution; The key matrix represents a feature dimension; the query matrix, key matrix, and value matrix all share the same feature dimension. Indicates the transpose operation; Represents the self-attention response feature map; Channel attention weights and spatial attention weights are calculated separately for the self-attention response feature maps, and then fused to obtain spatially enhanced features, mathematically represented as follows: ; ; ; in, , These represent channel attention weights and spatial attention weights, respectively. Indicates the Sigmoid activation operation; Indicates a multi-layer perceptron; This indicates an average pooling operation; This represents the max pooling operation; Indicates the convolution operation; This represents the channel average pooling operation performed along the channel dimension; This represents the channel max pooling operation performed along the channel dimension; Indicates a splicing operation; Represents spatially enhanced features; This represents element-wise multiplication; The spatial augmentation features are subjected to spatial dimension reduction and flattening to generate spatial representation vectors. This process is mathematically represented as follows: ; in, This represents the global average pooling operation, used to augment spatial features. Compression of two-dimensional spatial features into one-dimensional vectors; This indicates a fully connected layer, used for flattening. Indicates the first valid facial video segment Spatial representation vector of a frame image.

5. The method for assessing emotional blunting based on the divergence between facial expression and electrodermal response as described in claim 3, characterized in that: The introduced weakly supervised temporal attention mechanism weighted and fused the spatial representation vectors of each frame in the effective facial video segment to generate segment-level facial expression response vectors for each effective facial video segment, including: The process involves calculating temporal attention weights for the spatial representation vector, setting a temporal attention threshold, and filtering out vectors with low temporal attention weights. This process can be mathematically represented as follows: ; ; in, Indicates the first valid facial video segment Temporal attention weights for frame images; This indicates the Softmax normalization operation; This is the attention scalar mapping vector; This is the time-mapped weight matrix; This is the time-mapped bias vector; Indicates the transpose operation; Represents the hyperbolic tangent function; Indicates the first valid facial video segment The spatial representation vector of a frame image; Indicates the first valid facial video segment Temporal attention weights after filtering of frame images; The time attention threshold; The filtered temporal attention weights are normalized to obtain normalized frame-level weights. This process is mathematically represented as follows: ; in, Indicates the first valid facial video segment Normalized frame-level weights for frame images; The spatial representation vectors of each frame in a valid facial video segment are weighted and fused using normalized frame-level weights to generate a segment-level facial expression response vector for that valid facial video segment. This process is mathematically represented as follows: ; in, Indicates the first A segment-level facial expression response vector for each valid facial video clip.

6. The method for assessing emotional blunting based on the divergence between facial expression and electrodermal response as described in claim 3, characterized in that: The facial reference characterization is either a stable segment or a sliding mean among each effective facial video segment. The stable segment is determined by calculating the instantaneous variance of the motion displacement of facial key points within each effective facial video segment. Effective facial video segments with instantaneous variance lower than a preset stability threshold are considered stable segments. The sliding mean is obtained by taking the arithmetic mean of the segment-level facial expression response vectors of the current effective facial video segment and its preceding adjacent effective facial video segments through a preset time sliding window. The calculation of the facial expression response intensity of each effective facial video segment relative to the facial reference using segment-level facial expression response vectors is mathematically represented as follows: ; in, Represents facial reference characteristics; Indicates the first Fragment-level facial expression response vectors for each valid facial video segment. Indicates the index of valid facial video clips. Total number of valid facial video clips; Represents the L2 norm computation operator; Indicates the first A valid facial video segment relative to a facial reference representation The intensity of facial expression responses; The facial expression response sequence is mathematically represented as follows: ; in, This represents a sequence of facial expression responses.

7. The method for assessing emotional blunting based on the divergence between facial expression and electrodermal response as described in claim 1, characterized in that: The calculation of the skin electrical response intensity of each effective skin electrical segment and the construction of the skin electrical response sequence include: Each effective electrodermal cell segment undergoes preprocessing including noise reduction, outlier removal, smoothing, and standardization. Each effective skin electrodermal fragment after preprocessing was subjected to low-pass filtering to extract its tension component; The phase component is obtained by subtracting the effective skin electrical segment from its tension component; For each effective electrodermal segment, the local mean, fluctuation, peak amplitude, and energy of its phase components are calculated. The mathematical representation of the calculation method is as follows: ; ; ; ; in, , , , They represent the first Local mean, variability, peak amplitude, and energy of an effective skin electrodermal segment; Indicates the index of valid skin electrical fragments. The total number of effective skin electrodermal fragments is equal to the total number of effective facial video fragments; Indicates the first Number of sampling points within an effective skin electrical activity segment Indicates the first Within the first effective skin conductance fragment The signal amplitude at each sampling point Indicates the sampling point index; Indicates in Values ​​range from 1 to Take within range The maximum value; Perform a Fourier transform on the phase components to obtain the time-frequency representation of each effective electrodermal segment: ; in, Indicates the time offset of the phase component; is the rotation factor of the Fourier transform. For frequency variables, The imaginary unit; Represents a continuous physiological time series variable; Indicates a timing sliding window function; For the first Time shift of an effective skin electrophysiological segment and frequency variables Time-frequency representation; Calculate the energy within the target frequency band for each effective skin electrical activity segment based on time-frequency representation: ; in, No. Energy within the target frequency band of an effective skin electrodermal segment; , Representing frequency variables respectively The upper and lower bounds of the frequency band define the frequency range that the two bounds define; Indicates the calculation of the square of the modulus; The local mean, fluctuation, peak amplitude, and energy of the effective skin electric field fragments are constructed into a time-domain feature vector, and the energy within the target frequency band of the effective skin electric field fragments is constructed into a frequency-domain feature vector. ; ; in, , They represent the first The temporal feature vector of the first effective skin conductance segment, the first Frequency domain feature vectors of effective skin electrical activity segments; Indicates the transpose operation; By projecting the time-domain feature vector and the frequency-domain feature vector onto the same dimension through a linear mapping, we obtain the time-domain features and the frequency-domain features. Cross-domain attention calculation is performed between time-domain features and frequency-domain features to obtain the attention relationship between them. This process is mathematically represented as follows: ; ; in, This indicates the relationship of interest between time-domain features and frequency-domain features; This indicates the relationship of interest between frequency domain features and time domain features; This indicates the Softmax normalization operation; , , These represent the time-domain query matrix, time-domain key matrix, and time-domain value matrix, respectively, obtained through linear mapping of time-domain features; , , These represent the frequency domain query matrix, frequency domain key matrix, and frequency domain value matrix, respectively, obtained through linear mapping of frequency domain features; The characteristic dimension represents the time-domain key matrix. The time-domain query matrix, time-domain key matrix, time-domain value matrix, frequency-domain query matrix, frequency-domain key matrix, and frequency-domain value matrix all have the same characteristic dimension. Gated fusion of the interest relationships between time-domain and frequency-domain features is performed to generate a skin conductance response vector. The mathematical representation of this process is as follows: ; ; in, Indicates the first The gating coefficient of an effective skin electrodermal segment; Indicates the Sigmoid activation operation; , These represent the learnable gating weights and the gating bias term, respectively. Indicates a splicing operation; Indicates the first The skin electrical response vector of an effective skin electrical segment; This represents element-wise multiplication; Extract the baseline skin electrical activity (TEA) characterization and calculate the TEA response intensity of each effective TEA segment relative to this baseline characterization. The mathematical representation is as follows: ; in, Indicates the baseline characterization of skin electrical activity; Indicates the first Characterization of effective skin conductance segments relative to skin conductance baseline The intensity of skin conductance response; The L2 norm calculation operator is used; the skin electrophysiological benchmark is characterized as a stationary segment or a moving average in each effective skin electrophysiological segment. The stationary segment is determined by calculating the instantaneous variance of the signal amplitude at the sampling point in each effective skin electrophysiological segment and comparing it with a preset stationarity discrimination threshold. The moving average is obtained by taking the arithmetic mean of the skin electrophysiological response vectors of consecutive adjacent effective skin electrophysiological segments within a preset span of a time-series sliding window. The skin conductance response sequence is constructed based on the intensity of the skin conductance response, and its mathematical representation is as follows: ; in, This represents the skin conductance response sequence.

8. The method for assessing emotional blunting based on the divergence between facial expression and electrodermal response as described in claim 1, characterized in that: The calculation of the deviation between the facial expression response sequence and the skin conductance response sequence includes: S51. Calculate the amplitude divergence between the intensity of each facial expression response and its corresponding skin conductance response intensity, mathematically represented as follows: ; in, Indicates the first Facial expression intensity of a valid facial video clip The corresponding first Skin electroreactivity intensity of each effective skin electroreactivity segment The amplitudes diverged between them; This indicates a normalization operation; This indicates the calculation of the modulus, used here to obtain the absolute value; S52. Define sliding windows for extracting facial expression response sequences and skin conductance response sequences, respectively. , Facial expression response subsequence was obtained. With skin electroreactivity sequence Then, the dynamic correlation divergence between the intensity of each facial expression response and its corresponding skin conductance response intensity is calculated, and the mathematical representation is as follows: ; in, Indicates the first The intensity of facial expression response in a valid facial video clip and its corresponding... The dynamic correlation between the intensity of skin electrical response of each effective skin electrical segment diverged; This indicates the time delay between visual response and physiological sympathetic nerve response. This indicates the preset allowable time delay range; Indicates in Evaluate expression within range The maximum value; This indicates the operator for calculating the Pearson correlation coefficient; Represents a sliding window Superimposed time delay The extracted skin electroreactivity sequence; S53, Obtain the sliding window respectively , Peak times of the inner facial expression response subsequence and the electrodermal response subsequence , Furthermore, the peak deviation between the intensity of each facial expression response and its corresponding skin conductance response intensity is calculated, and the mathematical representation is as follows: ; in, Indicates the first The intensity of facial expression response in a valid facial video clip and its corresponding... The peak relationship between the skin electrical response intensities of each effective skin electrical segment deviates; S54. Calculate the inverse attentional divergence between facial expression response sequences and skin conductance response sequences, including: S541. Using one-dimensional temporal convolutional layers, the facial expression response sequence and skin conductance response sequence are respectively subjected to dimensionality increase and temporal local feature extraction to construct a facial feature sequence. With skin electrophysiological feature sequence ; S542. Calculate facial feature sequences With skin electrophysiological feature sequence The mathematical representation of the standard cross-modal attention matrix is ​​as follows: ; in, This represents a standard cross-modal attention matrix; This indicates the Softmax normalization operation; This represents the facial feature query matrix obtained by linear mapping of facial features in a facial feature sequence. This represents the skin electrophysiological feature key matrix obtained by linear mapping of skin electrophysiological features in a skin electrophysiological feature sequence; Indicates the transpose operation; The feature dimension of the column features after the skin electrodermal feature key matrix is ​​linearly mapped in the feature space. The facial feature query matrix and the skin electrodermal feature key matrix have the same feature dimension. S543. Calculate facial feature sequences With skin electrophysiological feature sequence The inverse cross-modal attention matrix between them is mathematically represented as follows: ; in, Representation of the conventional cross-modal attention matrix A matrix with the same dimension and all elements being 1; This represents the inverse cross-modal attention matrix; S544. Calculate the divergence feature representation based on the inverse cross-modal attention matrix, as follows: ; in, Indicates the divergence feature representation; This represents a time-based max-pooling operation; Represents the skin electrophysiological feature vector, obtained by analyzing the skin electrophysiological feature sequence. Obtain by performing a linear mapping; This represents element-wise multiplication; S545. The divergence feature is represented as a reverse attention divergence after being mapped through multiple perceptron layers, and its mathematical representation is as follows: ; in, This indicates a divergence in attention; Indicates a multi-layer perceptron; S55. Weighted fusion of amplitude divergence, dynamic correlation divergence, peak relationship divergence, and inverse attention divergence yields the divergence degree, mathematically represented as follows: ; in, Indicates the degree of divergence; , , These represent the time-series arithmetic mean of the amplitude divergence, dynamic correlation divergence, and peak relationship divergence between the intensity of all facial expression responses and their corresponding skin conductance responses, respectively. , , , These represent the weight parameters for amplitude divergence, dynamic correlation divergence, peak relationship divergence, and reverse attention divergence, respectively.

9. The method for assessing emotional blunting based on the divergence between facial expression and electrodermal response as described in claim 1, characterized in that: The degree of facial low reactivity is calculated as follows: ; in, Indicates a low level of facial reactivity; This represents the mean intensity of facial expression responses within a facial expression response sequence. For facial reaction stability parameters; The degree of skin conductance response is calculated as follows: ; in, Indicates the degree of skin conductance response; This represents the mean intensity of the skin electrical response within the skin electrical response sequence; For the stability parameters of skin electrical response; The emotional blunting score is calculated as follows: ; in, Indicates emotional sluggishness; Indicates the degree of divergence; , , These are the weighting parameters for the facial low reactivity level, deviation degree, and reaction type adjustment items, respectively. As a reaction type regulating term, it is defined as follows: ; in, The preset threshold for low facial reactivity; The preset threshold for the degree of skin conductance response; The criteria for determining the expression-inhibited emotional sluggishness are as follows: ; in, This is the preset deviation threshold; The criteria for determining low-responsive emotional sluggishness are as follows: ; The output of the emotional blunting assessment is mathematically represented as follows: ; in, The output indicates an assessment of emotional blunting; This indicates a type of emotionally sluggishness.

10. The method for assessing emotional blunting based on the divergence between facial expression and electrodermal response as described in claim 1, characterized in that: The construction of the target loss function includes: The rating regression loss is constructed and mathematically represented as follows: ; in, This represents the score regression loss; This represents the emotional blunting score predicted by the emotional blunting assessment model; This represents the true value of the affective dullness score in the clinical reference label; Indicates the calculation of the modulus; The type classification loss is constructed and mathematically represented as follows: ; in, Represents the type classification loss; Indicates the first in the clinical reference label The true value for the type of emotional detachment; Represents logarithmic calculation; This indicates the degree of facial hyporesponsiveness predicted by the emotional dullness assessment model. With skin conductance The value mapped after inputting the Sigmoid soft threshold function with a temperature coefficient belongs to the first... The probability of being emotionally sluggish is mathematically represented as follows: ; ; in, Indicates the temperature coefficient; The divergence hold loss is constructed and mathematically represented as follows: ; in, This indicates a deviation from the expected loss; Indicates the degree of deviation from the predictions made by the emotional blunting assessment model; This represents the true value of the deviation from the clinical reference label; The target loss function is obtained by weighted summing of the rating regression loss, type classification loss, and deviation preservation loss, and its mathematical expression is as follows: ; in, Represent the target loss function; , These represent the weight parameters for the type classification loss and the deviation preservation loss, respectively.