Facial dimension emotion-based disturbance of consciousness auxiliary diagnosis method and model

By using a facial dimension emotion estimation model, combining global and local neural networks with name calling and pain stimuli, the changes in facial expressions of DOC patients are quantified, solving the problem of high misdiagnosis rate in existing technologies and achieving more accurate diagnosis of disorders of consciousness.

CN121885152APending Publication Date: 2026-04-17SOUTH CHINA NORMAL UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTH CHINA NORMAL UNIV
Filing Date
2025-12-30
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies for diagnosing disorders of consciousness suffer from high misdiagnosis rates, expensive equipment, and complex operation, making it difficult to accurately assess the state of consciousness in patients with DOC.

Method used

This study employs a facial-based emotional-assisted diagnostic method for disorders of consciousness (DOC). By evoking emotional responses in DOC patients through name calling and pain stimuli, and combining a global and local bi-branch neural network model, it quantifies facial expression changes and provides objective evidence for consciousness assessment.

Benefits of technology

It significantly improves the accuracy of MCS patient identification, reduces the risk of misdiagnosis, provides more refined quantitative evidence, and assists in clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121885152A_ABST
    Figure CN121885152A_ABST
Patent Text Reader

Abstract

The invention relates to a disturbance of consciousness auxiliary diagnosis method and model based on facial dimension emotion estimation. The method belongs to the cross technical field of deep learning and clinical medicine. According to the method, facial expression videos of a patient under resting, name calling stimulation and pain stimulation are collected, and global features and local five-sense-organ patch features are extracted by using RetinaFace, MaxViT and HRNet models. A global branch captures overall time sequence dependence through a time sequence convolutional network and a Transform encoder, a local branch and multi-head attention modeling local subtle change through a residual network, and a wake-up degree sequence is output through dynamic weighted fusion. And calculating three indexes, namely, a calling name / resting wake-up degree difference, a pain / resting wake-up degree difference and a wake-up degree standard difference under pain, based on the sequence. And comparing each index with a preset threshold, and judging the consciousness state category of the patient. The method is simple and convenient to operate, low in cost and capable of providing objective quantitative basis for clinical diagnosis and reducing the misdiagnosis rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of deep learning, computer vision, and clinical medicine, and in particular to a method and system for assisting in the diagnosis of consciousness disorders based on facial dimension emotion. Background Technology

[0002] Assessing the level of consciousness in patients with disorders of consciousness is crucial for diagnosis and treatment, but it remains a significant challenge in clinical practice. Disorders of Consciousness (DOC) encompass two main diagnostic categories: Vegetative State (VS) / Unresponsive Wakefulness Syndrome (UWS) and Minimally Consciousness State (MCS). VS / UWS presents with marked arousal without consciousness, while MCS exhibits inconsistent but reproducible signs of consciousness. In recent years, with improved survival rates for traumatic brain injury, hypoxic-ischemic encephalopathy after cardiac arrest, and stroke, the incidence of DOC has also increased. Statistics show that approximately 300,000 patients in the United States are in varying degrees of chronic disorder of consciousness, including about 120,000 children; my country has over 500,000 DOC patients, with approximately 100,000 new cases added annually, resulting in annual treatment costs of 30-50 billion yuan. This population not only places a heavy burden on the healthcare system and society but also imposes ongoing psychological and economic stress on patients' families.

[0003] Accurate diagnosis is crucial for providing effective treatment and prognosis for patients with Disorders of Consciousness (DOC). Because VS / UWS and MCS patients exhibit different behavioral responses, current clinical methods for assessing consciousness in DOC patients primarily rely on behavioral scales, such as the JFK Coma Recovery Scale-corrected (CRS-R) and the Glasgow Coma Scale (GCS). However, these diagnostic methods, depending solely on the patient's behavioral responses, are prone to misdiagnosis. Statistics show that approximately 40% of MCS patients may be misdiagnosed as VS patients. With technological advancements, many researchers have proposed using neuroimaging techniques such as positron emission tomography (PET), functional magnetic resonance imaging (fMRI), and diffusion tensor imaging (DTI), as well as electrophysiological techniques like electroencephalography (EEG), to aid in the diagnosis of consciousness disorders, thereby improving diagnostic accuracy. However, neuroimaging methods are expensive, the equipment is bulky and scarce, and there are radiation risks. On the other hand, head trauma in DOC patients can make electrode placement difficult, posing a challenge to EEG assessment methods. Therefore, it is crucial to explore simple and cost-effective methods for the auxiliary diagnosis of disorders of consciousness.

[0004] Studies have shown that emotions and states of consciousness are closely related in patients with DOC (Disease of Consciousness). The valence-arousal dimensional affective model is one of the most widely used theoretical frameworks for quantifying and representing emotions. This model treats emotional states as points in a two-dimensional continuous space, where valence describes the positive or negative degree of emotion, and arousal reflects the activation level of emotion. Therefore, the dimensional affective model can model subtle, complex, and continuous emotional behaviors.

[0005] Facial expressions are a direct manifestation of emotions. Compared to techniques like fMRI and EEG, which require complex equipment and are cumbersome to operate, facial expressions, as a natural and continuous behavioral signal, have significant advantages such as ease of acquisition, non-invasiveness, and low cost. Name-calling stimuli and pain stimuli, as emotional stimuli, can elicit changes in facial expressions. Based on this, we propose a simple and practical deep neural network model for facial emotion recognition and apply this model to patients with DOC (Disease of Constipation) to quantify their facial expressions, thereby reflecting their level of consciousness and assisting existing clinical diagnostic methods in accurately assessing patients' level of consciousness.

[0006] Existing publicly disclosed patents involving cognitive state recognition systems and methods based on facial expressions include CNCN112201343A, "Cognitive State Recognition System and Method Based on Facial Micro-expressions"; CN114565957A, "Consciousness Assessment Method and System Based on Micro-expression Recognition"; CN117643453A, "Auxiliary Diagnostic Method for Consciousness Disorders Based on Facial Expressions"; and CN119069110A, "An Auxiliary Diagnostic System for Consciousness Disorders Based on Pain Intensity Estimation". The core content of these methods all involves using video data to capture changes in the subject's facial expressions and classifying them through feature extraction and machine learning models to determine the subject's level of consciousness or state. All four methods have shortcomings. The first and third methods use facial action units (AUs) to quantify facial movements, but the number of publicly available AU datasets is limited, and the number of pain-related AU types is also small, affecting the accuracy and reliability of these methods in practical applications. The second method only uses auditory stimulation, which is difficult to apply to patients with hearing impairment. The fourth method only uses pain as a stimulus paradigm, resulting in a single source of data stimulation, which affects the accuracy of the assessment. Existing published patents related to auxiliary diagnostic systems and methods for DOC patients include CN116491901A, "Auxiliary Diagnostic System for Patients with Disorders of Consciousness Based on P300 EEG and Transfer Learning," and CN113116306A, "An Auxiliary Diagnostic System for Disorders of Consciousness Based on Auditory Evoked EEG Signal Analysis." Both patents attempt to assess the level of consciousness in DOC patients using EEG data. However, DOC patients often cannot actively cooperate with the EEG collection process, and DOC patients with severe head injuries cannot wear EEG caps, making it difficult to collect EEG signals from DOC patients. In addition, high-quality EEG acquisition equipment is expensive and requires professional maintenance. Summary of the Invention

[0007] To address the problems in existing technologies, this paper provides a diagnostic method and model for consciousness disorders based on facial dimension emotion estimation. It uses standardized multimodal emotional stimuli and a customized deep neural network to quantify subtle changes in the patient's facial arousal, providing objective and quantitative auxiliary evidence for clinical diagnosis.

[0008] To achieve the above objectives, the present invention provides a method and model for assisting in the diagnosis of consciousness disorders based on facial dimension emotion estimation, as follows:

[0009] A diagnostic method for consciousness disorders based on facial dimension emotion estimation includes:

[0010] To achieve the above objectives, the present invention adopts the following technical solution:

[0011] S1. Collect facial videos of patients with impaired consciousness in a resting state, when receiving name-calling stimuli, and when receiving pain stimuli;

[0012] S2. Preprocess the facial video to extract global spatial feature sequences and local patch sequences;

[0013] S3. Input the global spatial feature sequence and local patch sequence into the facial dimension emotion estimation model to obtain the temporal prediction value of the patient's facial emotion in each state; the model is a spatiotemporal feature fusion network based on global and local dual branches, which outputs the prediction value by dynamically weighting the global spatiotemporal features and local spatiotemporal features.

[0014] S4. Based on the predicted arousal time series value, calculate the first difference value, the second difference value, and the fluctuation value; the first difference value is the difference between the average arousal level in the calling state and the resting state; the second difference value is the difference between the average arousal level in the painful stimulus state and the resting state; the fluctuation value is the standard deviation of arousal level in the painful stimulus state.

[0015] S5. Based on the behavioral scale methods commonly used in clinical practice, combined with the first difference value, the second difference value, and the fluctuation value, determine the patient's state of consciousness category.

[0016] Furthermore, step S5 specifically includes:

[0017] The first difference value, the second difference value, and the fluctuation value are compared with the first preset threshold, the second preset threshold, and the third preset threshold, respectively.

[0018] If at least one of the three conditions is greater than its corresponding preset threshold, the patient is determined to be in a state of minimum consciousness; otherwise, the patient is determined to be in a vegetative state / unresponsive arousal syndrome.

[0019] Furthermore, the first preset threshold, the second preset threshold, and the third preset threshold are obtained through subject operating characteristic curve analysis on a sample dataset containing samples diagnosed by behavioral scales.

[0020] Further, the preprocessing in step S2 includes:

[0021] The RetinaFace model is used to detect and align faces in video frames to obtain a standardized facial image sequence.

[0022] The standardized facial image sequence is input into a pre-trained MaxViT feature extractor to generate the global spatial feature sequence;

[0023] Based on facial key points, multiple image blocks of preset facial feature regions are cropped from the standardized facial image sequence to form the local patch sequence;

[0024] Furthermore, the regions in the local patch sequence include the left eye, right eye, nose, left corner of the mouth, and right corner of the mouth; the preprocessing also includes generating a facial mask using HRNet and superimposing the mask information into the local patch sequence.

[0025] Furthermore, the global branch of the facial dimension emotion estimation model includes a temporal convolutional network module and a Transformer encoder module connected in sequence;

[0026] The temporal convolutional network module is configured to use a causal dilated convolutional structure to extract multi-scale local temporal patterns;

[0027] The Transformer encoder module is configured to perform global context dependency modeling.

[0028] The local branches of the facial dimension emotion estimation model include:

[0029] A spatial feature encoder, consisting of at least one residual network block, is used to extract spatial features from image blocks of each key region.

[0030] The temporal modeling module employs a multi-head attention mechanism to perform temporal dimension modeling of the spatial feature sequences of each key region.

[0031] The region aggregation module uses an attention mechanism to calculate the weights of features in each key region and performs weighted fusion.

[0032] Furthermore, the dynamic weight fusion is achieved through the following steps:

[0033] Calculate the initial weighted average of the global spatiotemporal features and the local spatiotemporal features to obtain the initial fused features;

[0034] Calculate the similarity between the global spatiotemporal features and the local spatiotemporal features and the initial fused features, respectively;

[0035] The similarity is normalized to generate a first dynamic weight for the global spatiotemporal features and a second dynamic weight for the local spatiotemporal features;

[0036] The global and local spatiotemporal features are weighted and summed using the first and second dynamic weights.

[0037] Furthermore, the facial dimension emotion estimation model was trained using the PyTorch framework and an Nvidia GeForce GTX 3090 24G GPU. During training, the AdamW optimizer was used, with an initial learning rate of 1e-4, a batch size of 16, a patch size of 16, and a global spatial feature dimension of 512. The model was trained on the publicly available large-scale natural scene facial expression dataset Aff-wild2. The Aff-wild2 dataset consists of 558 videos downloaded from YouTube, totaling 2,786,201 video frames and a total duration exceeding 43 hours. Four experts labeled each video frame with valence and arousal values ​​(range [-1, +1]). The videos in the dataset were divided into segments with a window size w of 120 and a stride s of 80 as input.

[0038] On the other hand, the present invention also provides a neural network model for facial dimension emotion estimation, comprising:

[0039] The global spatiotemporal feature extraction branch takes the global spatial feature sequence of the facial image as input and outputs global spatiotemporal features.

[0040] The local spatiotemporal feature extraction branch takes as input a sequence of local image patches of multiple key facial regions and outputs local spatiotemporal features.

[0041] The feature fusion module is used to adaptively calculate the contribution weights of the global spatiotemporal features and the local spatiotemporal features based on the input content, and perform dynamic weighted fusion to output the predicted value of the sentiment dimension.

[0042] The global spatiotemporal feature extraction branch includes a temporal convolutional network and a Transformer encoder connected in sequence; the local spatiotemporal feature extraction branch includes a residual network spatial encoder, a multi-head attention temporal modeler, and a keypoint attention aggregation module.

[0043] Compared with the prior art, the beneficial effects of the present invention are:

[0044] This invention employs a dual-stimulation paradigm of "name calling + pain," synergistically inducing emotional responses in DOC patients through auditory and somatosensory pathways, effectively improving the reliability of emotion induction. Building upon this, the model precisely captures subtle facial arousal changes via a global and local dual-branch network, and for the first time applies facial arousal statistics to the assessment of consciousness level. This method, combining multimodal stimulation with dimensional affective computation, significantly improves the accuracy of identifying MCS patients, providing more objective and refined quantitative evidence for clinical diagnosis, assisting physicians in diagnosis, and effectively reducing the risk of misdiagnosis.

[0045] This invention proposes a two-branch dimensional emotion estimation network model. This model is designed to address the characteristics of DOC patients, whose facial expressions are weak and whose subtle emotional changes are difficult to capture. The global branch captures overall facial contextual information, while the local branch focuses on the subtle dynamics of key areas such as the eyes and mouth. An efficient fusion mechanism integrates the two sets of information. This model has a sophisticated structure and an efficient fusion mechanism, providing an effective solution for weak emotion estimation. Attached Figure Description

[0046] Figure 1 Flowchart of a method for assisting in the diagnosis of consciousness disorders based on facial expression-based emotion;

[0047] Figure 2 A paradigm diagram for collecting facial expression data from DOC patients;

[0048] Figure 3 Schematic diagram of a model for an auxiliary diagnostic method for consciousness disorders based on facial emotion;

[0049] Figure 4 Schematic diagram of the local feature extraction module;

[0050] Figure 5 The ratio distribution curves of VS / UWS patients and MCS patients on the Diff (name-rest) index;

[0051] Figure 6 The ratio distribution curves of VS / UWS patients and MCS patients on the Diff (pain-rest) index;

[0052] Figure 7 The distribution curves of the ratio of VS / UWS patients and MCS patients on the Std (pain) index. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described below with reference to the accompanying drawings and preferred embodiments. However, the scope of protection of this invention is not limited to these embodiments. Any modifications and variations made within the scope of the claims should be covered within the scope of protection of this invention.

[0054] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0055] Example 1

[0056] like Figure 1-4 This embodiment discloses an auxiliary diagnostic method for disorders of consciousness based on facial dimension emotion estimation.

[0057] 1. Data Collection

[0058] Facial expression videos of patients should be recorded using a 1080p, 60fps camera or smartphone. The patient's face should occupy the main part of the frame, and the area below the shoulders should not be visible. A medical assistant should guide the subject and their family, explaining the experimental purpose and precautions. Those present should be reminded not to make any disruptive noises during the test and not to enter restricted areas. Patient information should be recorded, including subject number, name, gender, and medical condition. A clear start signal is required before recording (e.g., if there is a painful stimulus, medical staff should inform the recorder before starting finger pressure). If any abnormalities occur during the data collection process, the subject's video should be re-recorded. This data collection paradigm strictly adheres to the World Medical Association's Ethical Guidelines (Declaration of Helsinki).

[0059] Video data collection for DOC patients includes the following three states, such as Figure 2 As shown:

[0060] Resting state: The patient is in a resting state, which means that the patient is in a calm state without any stimulation. At this time, the patient needs to remain awake and record their facial expression video without any external stimulation for at least 10 seconds.

[0061] Name calling stimulation: This process involves medical staff or a family member calling the patient's name from the left side for 10 seconds. After a 1-minute rest, the name calling stimulation process is repeated from the right side. Facial expressions are recorded during the name calling stimulation.

[0062] Painful stimulation: Healthcare professionals apply a short (approximately 5 seconds) noxious stimulus to the patient's finger using standardized instruments (such as a pressure analgesic device) or standardized techniques (such as pressing the nail bed until it turns white). Facial expressions are recorded for 10 seconds after the onset of the painful stimulus. After a 1-minute rest, the painful stimulation process is repeated on the right side, continuously recording at least 10 seconds of video.

[0063] Because DOC patients may have varying degrees of unilateral brain injury, unilateral data acquisition may not be accurate. Therefore, name-calling stimulation and pain stimulation were collected twice, once from each side. The above acquisition process consisted of 5 video segments (resting-state video V). Rest Left-side name-calling stimulation video V LName The right-side calling-name stimulation video V RName Left-side pain stimulation video V LPain Right-side pain stimulation video V RPain ).

[0064] 2. Data Preprocessing and Feature Extraction

[0065] Each acquired video segment underwent standardized preprocessing: The acquired video segments were 10 seconds long, and 300 frames were extracted. The Retinaface model was used to perform face alignment, face detection, and keypoint detection on the input video frames, obtaining 112×112 pixel facial images. The global spatial feature sequence was generated from the face-aligned image sequence through spatial feature extraction, with a feature dimension of 512. A MaxViT model pre-trained on the static facial expression dataset AffectNet was used as the facial spatial feature extractor.

[0066] Local patches are cropped regions based on facial keypoint areas detected by RetinaFace. Five core expression regions are cropped from each aligned image frame: left eye, right eye, nose, left corner of mouth, and right corner of mouth. The cropped size is 16×16 pixels. Specifically, to reduce noise, we use HRNet to identify 98 facial coordinate points to generate a facial mask, and add additional facial mask information to the local branch data to improve facial feature recognition. The local patch sequence is a sequence of image patches cropped from facial keypoint areas after stacking the aligned image sequence and the facial mask.

[0067] The input video sequence is set to a length of T. After data preprocessing, the input video finally yields a global spatial feature sequence I with a dimension of T×512. G And a local patch sequence I with dimensions of 5×T×16×16×4 L This is used as input to the network model later.

[0068] 3. Facial Expression Dimension Emotion Estimation Model

[0069] The global spatial feature sequence and the local patch sequence are input into the core of this invention—the facial dimension emotion estimation model. For example... Figure 3 and Figure 4 As shown, the model consists of three main parts:

[0070] Global and local branch feature extraction is mainly used for feature extraction and modeling of video data, extracting temporal and spatial information from the video. The feature fusion part fuses the extracted global and local spatiotemporal features to generate the final prediction result.

[0071] Global branch feature extraction part: Global spatial feature sequence I obtained from data preprocessing G The data input for the global branch is used, where a three-layer temporal convolutional network and a Transformer encoder are cascaded together to extract features, generating a global spatiotemporal feature F with dimension T×512. G :

[0072]

[0073] TCN, through its hierarchical structure and causal dilated convolutions with varying dilation rates, excels at capturing local temporal patterns and multi-scale features. It efficiently extracts temporal patterns and dependencies of features over time in short, medium, and long-term sequences, providing highly semantic temporal feature representations for subsequent Transformers. The Transformer architecture performs exceptionally well in various natural language processing tasks. Unlike recurrent neural networks, Transformers achieve stronger feature representation capabilities by explicitly constructing contextual dependencies.

[0074] Local branch feature extraction section: Local patch sequence I generated by data preprocessing L This serves as the data input for the local branch. Subtle changes in facial features have a crucial impact on facial expression prediction. In the local branch, to effectively capture fine-grained spatiotemporal dynamic features, we designed a local feature extraction module. Two residual blocks are used in the local branch to extract spatial features from the local patch sequence. The number of feature channels in these two layers is set to 64 and 128 respectively. Each residual block contains two convolutional layers (Conv) with a kernel size of 3. A batch normalization (BN) layer is used after each convolutional layer to improve training efficiency and stability. In addition, a convolutional layer with a kernel size of 1 is used to adjust the number of input channels for residual connections. To reduce computation and highlight important information in the features, we perform adaptive max pooling after extracting spatial features from the residual blocks, generating pooled features with dimensions of 5×T×128. Subsequently, multi-head attention is used for time series modeling, while maintaining the same feature dimension.

[0075] The above steps generated a spatiotemporal feature sequence {S} of five key points. i ∈R T x 128 If i=1,2,3,4,5}, then after passing through the keypoint attention aggregation module, we will process the spatiotemporal feature sequence {S}. i Attention is calculated along the keypoint dimension, and the spatiotemporal feature sequence of keypoints is explicitly weighted and aggregated using attention weights. After passing through a fully connected layer, a local spatiotemporal feature F with dimension T×512 is finally generated. L Attention aggregation based on key points can be represented by the following formula:

[0076]

[0077] Where i represents the key point number, i∈{1,2,3,4,5}, A i represents the attention weight of keypoint sequence i, FC(·) represents a fully connected layer, and ⊗ represents element-wise multiplication.

[0078] Feature fusion: Different facial expressions are influenced by different facial regions, and the same expression can manifest differently in different individuals. Due to these expressive differences, fixed-weight feature fusion strategies cannot adapt to this dynamic nature. The proposed facial expression-based sentiment estimation model introduces a feature fusion module. This module dynamically evaluates and quantifies the contribution weights of global and local features based on cosine similarity, adaptively optimizing the feature fusion process. The input to the feature fusion module is the global spatiotemporal feature F. G Local spatiotemporal features F L The global spatiotemporal features F G and local spatiotemporal features F L The initial fusion feature F is calculated using a weighted average of 0.5:0.5. C F G F L F C The data are fed into fully connected layers for high-level semantic encoding and spatial alignment to generate features F. ’ G F ’ L F ’ C :

[0079]

[0080] Where k represents the element index, and FC(·) represents a fully connected layer.

[0081] Subsequently, we calculated F respectively. ’ C With F ’ G and F ’ L Cosine similarity:

[0082]

[0083] Where COS(·) is a cosine similarity function.

[0084] Cosine similarity value κ G and κ L This reflects the importance of both global and local features. These values ​​are normalized and calculated to generate the global and local contribution weights α. G and α L Each value is assigned to the corresponding feature vector for dynamic feature fusion calculation. α G and α L The definition is as follows:

[0085]

[0086] Finally, after feature fusion, we feed the features into a fully connected layer and Tanh to generate the final dimensionality sentiment prediction result. pred This includes arousal sequences and efficacy value sequences:

[0087]

[0088] 4. Calculation and Judgment of Consciousness Level Assessment Indicators

[0089] Existing research indicates that the emotions and consciousness states of DOC patients are closely related. Based on this finding, we induce changes in facial expressions in DOC patients through emotional stimulation, quantify these changes using a facial expression dimension emotion estimation model, and perform statistical analysis on the quantified data to determine the level of consciousness in DOC patients.

[0090] In the previous step, the video clip was fed into the facial expression dimension sentiment estimation model to predict the facial expression arousal sequence value corresponding to that video clip. For each arousal sequence, we calculated statistical metrics for arousal, including the mean and standard deviation. Rest (Arousal), Mean LName (Arousal), Mean RName (Arousal), Mean LPain (Arousal), Mean RPain (Arousal), Std LPain (Arousal), Std RPain (Arousal)). Subsequently, for the average arousal levels of both sides under the same stimulus condition, we selected the larger value as the average value for that stimulus condition, and the standard deviation was similarly calculated:

[0091]

[0092] For quantified patient facial expression data, we calculated the difference between the mean arousal level for name-calling stimuli and the resting state (Diff(name-rest)), the difference between the mean arousal level for pain stimuli and the resting state (Diff(pain-rest)), and the standard deviation of arousal level for pain stimuli (Std(pain)):

[0093]

[0094] For the three metrics Diff (name-rest), Diff (pain-rest), and Std (pain), a classification threshold T is set for each. N T P T SThis study analyzed data from 37 patients with DOC (Disease of Occurrence) to explore the optimal discrimination threshold for differentiating between VS / UWS and MCS. Figure 5-7 The analysis curve shown indicates that T was ultimately selected. N =0.05, T P =0.15 and T S =0.03 is used as the threshold. Our classification criteria are as follows: If a patient's score in any of the three indicators, Diff (name-rest), Diff (pain-rest), and Std (pain), is higher than the threshold of that indicator, the patient is judged to have minimum consciousness state MCS; otherwise, the patient is judged to be in vegetative state / unresponsive arousal syndrome VS / UWS.

[0095] 5. Auxiliary diagnosis

[0096] The classification results obtained from the above automatic analysis are presented side by side with the assessment results of clinical doctors' behavioral scales (such as the CRS-R scale), providing doctors with an objective quantitative reference to assist them in making the final diagnosis of the state of consciousness, thereby effectively reducing the misdiagnosis that may be caused by relying solely on behavioral assessment.

[0097] Example 2

[0098] On the other hand, this embodiment also discloses a neural network model for facial dimension emotion estimation, including:

[0099] The global spatiotemporal feature extraction branch takes the global spatial feature sequence of the facial image as input and outputs global spatiotemporal features.

[0100] The local spatiotemporal feature extraction branch takes as input a sequence of local image patches of multiple key facial regions and outputs local spatiotemporal features.

[0101] The dynamic feature fusion module is used to adaptively calculate the contribution weights of the global spatiotemporal features and the local spatiotemporal features based on the input content, and perform dynamic weighted fusion to output the predicted value of the sentiment dimension.

[0102] The global spatiotemporal feature extraction branch includes a three-layer temporal convolutional network and a Transformer encoder connected in sequence; the local spatiotemporal feature extraction branch includes a residual network spatial encoder, a multi-head attention temporal modeler, and a keypoint attention aggregation module.

[0103] The facial dimension emotion estimation model is implemented based on the PyTorch framework and trained on an Nvidia GeForce GTX3090 24G GPU. During training, the AdamW optimizer was used, with an initial learning rate of 1e-4, a batch size of 16, a patch size of 16, and a global spatial feature dimension of 512.

[0104] The model was trained on the publicly available large-scale natural scene facial expression dataset Aff-wild2: this dataset contains 558 YouTube videos (total duration over 43 hours, totaling 2,786,201 frames), each frame annotated with attentional valence and arousal by at least two professional annotators within the interval [-1, +1]. The videos in the dataset were divided into segments with a window size w of 120 and a stride s of 80 as input.

[0105] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for assisting in the diagnosis of consciousness disorders based on facial expression-based emotion, characterized in that, Includes the following steps: S1. Collect facial videos of patients with impaired consciousness in a resting state, when receiving name-calling stimuli, and when receiving pain stimuli; S2. Preprocess the facial video to extract global spatial feature sequences and local patch sequences; S3. Input the global spatial feature sequence and local patch sequence into the facial dimension emotion estimation model to obtain the temporal prediction value of the patient's facial emotion in each state; the model is a spatiotemporal feature fusion network based on global and local dual branches, which outputs the prediction value by dynamically weighting the global spatiotemporal features and local spatiotemporal features. S4. Based on the predicted arousal time series value, calculate the first difference value, the second difference value, and the fluctuation value; the first difference value is the difference between the average arousal level in the calling state and the resting state; the second difference value is the difference between the average arousal level in the painful stimulus state and the resting state; the fluctuation value is the standard deviation of arousal level in the painful stimulus state. S5. Based on the behavioral scale methods commonly used in clinical practice, combined with the first difference value, the second difference value, and the fluctuation value, determine the patient's state of consciousness category.

2. The method according to claim 1, characterized in that, Step S5 specifically also includes: The first difference value, the second difference value, and the fluctuation value are compared with the first preset threshold, the second preset threshold, and the third preset threshold, respectively. If at least one of the three conditions is greater than its corresponding preset threshold, the patient is determined to be in a state of minimum consciousness; otherwise, the patient is determined to be in a vegetative state / unresponsive arousal syndrome.

3. The method according to claim 2, characterized in that, The first, second, and third preset thresholds were obtained through subject operating characteristic curve analysis on a sample dataset containing samples diagnosed by behavioral scales.

4. The method according to claim 1, characterized in that, The preprocessing in step S2 includes: The RetinaFace model is used to detect and align faces in video frames to obtain a standardized facial image sequence. The standardized facial image sequence is input into a pre-trained MaxViT feature extractor to generate the global spatial feature sequence; Based on facial key points, multiple image blocks of preset facial feature regions are cropped from the standardized facial image sequence to form the local patch sequence.

5. The method according to claim 4, characterized in that, The regions in the local patch sequence include the left eye, right eye, nose, left corner of the mouth, and right corner of the mouth; the preprocessing also includes generating a facial mask using HRNet and overlaying the mask information onto the local patch sequence.

6. The method according to claim 1, characterized in that, The global branch of the facial dimension emotion estimation model includes a temporal convolutional network module and a Transformer encoder module connected in sequence; The temporal convolutional network module is configured to use a causal dilated convolutional structure to extract multi-scale local temporal patterns; The Transformer encoder module is configured to perform global context dependency modeling; The local branches of the facial dimension emotion estimation model include: A spatial feature encoder, consisting of at least one residual network block, is used to extract spatial features from image blocks of each key region. The temporal modeling module employs a multi-head attention mechanism to perform temporal dimension modeling of the spatial feature sequences of each key region. The region aggregation module uses an attention mechanism to calculate the weights of features in each key region and performs weighted fusion.

7. The method according to claim 1, characterized in that, The dynamic weight fusion is achieved through the following steps: Calculate the initial weighted average of the global spatiotemporal features and the local spatiotemporal features to obtain the initial fused features; Calculate the similarity between the global spatiotemporal features and the local spatiotemporal features and the initial fused features, respectively; The similarity is normalized to generate a first dynamic weight for the global spatiotemporal features and a second dynamic weight for the local spatiotemporal features; The global and local spatiotemporal features are weighted and summed using the first and second dynamic weights.

8. The method according to claim 1, characterized in that, The facial dimension emotion estimation model was trained as follows: it was trained on the PyTorch framework and an Nvidia GeForce GTX 3090 24G GPU; during training, the AdamW optimizer was used, with an initial learning rate of 1e-4, a batch size of 16, a patch size of 16, and a global spatial feature dimension of 512; the model was trained on the publicly available large-scale natural scene facial expression dataset Aff-wild2.

9. A neural network model for facial dimension emotion estimation, characterized in that, include: The global spatiotemporal feature extraction branch takes the global spatial feature sequence of the facial image as input and outputs global spatiotemporal features. The local spatiotemporal feature extraction branch takes as input a sequence of local image patches of multiple key facial regions and outputs local spatiotemporal features. The feature fusion module is used to adaptively calculate the contribution weights of the global spatiotemporal features and the local spatiotemporal features based on the input content, and perform dynamic weighted fusion to output the predicted value of the sentiment dimension.

10. The neural network model according to claim 9, characterized in that, The global spatiotemporal feature extraction branch includes a temporal convolutional network and a Transformer encoder connected in sequence; the local spatiotemporal feature extraction branch includes a residual network spatial encoder, a multi-head attention temporal modeler, and a keypoint attention aggregation module.

Citation Information

Patent Citations

  • Cognitive state recognition system and method based on facial micro-expressions

    CN112201343A

  • Auxiliary diagnosis system for disturbance of consciousness based on auditory evoked electroencephalogram signal analysis

    CN113116306A

  • Micro-expression recognition-based consciousness assessment method and system

    CN114565957A

  • Auxiliary diagnosis system for patients with disturbance of consciousness based on P300 electroencephalogram and transfer learning

    CN116491901A

  • Facial expression-based disturbance of consciousness auxiliary diagnosis method

    CN117643453A