A method for dynamically evaluating the degree of sub-conscious disorder under multi-modal
By using a multimodal physiological signal dynamic assessment method combined with a deep learning model, the problem of high misdiagnosis rate of chronic consciousness disorders was solved, and accurate assessment and dynamic feedback of consciousness status were achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-21
- Publication Date
- 2026-03-24
AI Technical Summary
Existing clinical diagnostic criteria, such as CRS-R, have a high rate of misdiagnosis when assessing chronic disorders of consciousness, and it is difficult to accurately distinguish between minimally conscious states and vegetative states.
A multimodal dynamic assessment method for the degree of consciousness impairment was adopted. By acquiring multiple physiological signals (EEG, skin conductance, electrooculography, electrocardiography, and blood oxygenation signals), preprocessing and feature extraction were performed, and a deep learning model was used for feature fusion to calculate the consciousness impairment assessment score.
It enables dynamic and accurate assessment of the consciousness status of patients with chronic consciousness disorders, reducing the misdiagnosis rate and improving the accuracy and efficiency of assessment.
Smart Images

Figure CN120661086B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of dynamic assessment of consciousness disorders, and in particular to a method for dynamic assessment of the degree of consciousness disorders under multimodal conditions. Background Technology
[0002] Chronic disorder of consciousness is a pathological process caused by severe brain injury or cerebrovascular disease, resulting in loss of consciousness in which the patient is unable to correctly perceive their own state or the objective environment and cannot respond appropriately to external stimuli. The main subtypes of chronic disorder of consciousness include vegetative state and minimally conscious state.
[0003] Accurate classification is crucial for subsequent treatment strategies in patients with chronic disorders of consciousness. The current clinical diagnostic standard is the Coma Recovery Scale-Revised (CRS-R), which is the only standardized neuropsychological assessment scale, providing detailed scoring instructions for 6 major items and 23 sub-items. However, due to challenges such as the diversity of clinical symptoms of chronic disorders of consciousness, the instability of subtle behavioral differences between minimally conscious and vegetative states, and the variability in physicians' clinical experience, behavioral assessment results have a high rate of misdiagnosis. Summary of the Invention
[0004] The purpose of this invention is to provide a method for dynamic assessment of the degree of consciousness impairment under multimodal conditions, aiming to solve the problem of dynamic assessment of the degree of consciousness impairment.
[0005] This invention provides a method for dynamic assessment of the degree of consciousness impairment in a multimodal context, including:
[0006] S1. Acquire physiological signals of the subject under multimodal conditions;
[0007] S2. Preprocess the physiological signals to obtain the preprocessed signals;
[0008] S3. Calculate the features of physiological signals based on preprocessed signals to obtain multidimensional features;
[0009] S4. Preprocess the multidimensional features to obtain preprocessed data;
[0010] S5. Merge the preprocessed data to obtain fusion features;
[0011] S6. The fusion feature input pattern recognition output module obtains the subject's consciousness impairment assessment score.
[0012] By employing the embodiments of the present invention, dynamic physiological signals under multimodal conditions are obtained, multidimensional features are acquired, and the consciousness impairment assessment of the test subject is judged based on the multidimensional features, making the multidimensional assessment more accurate.
[0013] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0014] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0015] Figure 1 This is a flowchart of a multimodal dynamic assessment method for the degree of consciousness impairment according to an embodiment of the present invention. Detailed Implementation
[0016] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Method Implementation Examples
[0018] According to embodiments of the present invention, a method for dynamic assessment of the degree of consciousness impairment in a multimodal context is provided. Figure 1 This is a flowchart of the multimodal dynamic assessment method for the degree of consciousness impairment according to an embodiment of the present invention, such as... Figure 1 As shown, it specifically includes:
[0019] S1. Acquire physiological signals of the subject under multimodal conditions;
[0020] S1 specifically includes: real-time acquisition of the subject's electroencephalogram (EEG), electrodermal conductance (TEA), electrooculogram (EOG), and blood oxygen saturation signals in both music and quiet states.
[0021] In this embodiment of the invention, five types of sensors—electroencephalogram (EEG), electrodermal conductance (EDA), electrocardiogram (ECG), electrooculogram (EOG), and blood oxygen saturation—are installed on the subject to achieve synchronous and dynamic assessment of the patient's neural activity, autonomic nervous system, cardiovascular system, vision, and metabolic status. In the ward environment, selected music is played in a loop for 30 minutes every 3 hours from 8:00 AM to 5:00 PM, for a total of 4 times per day. The patient's various physiological indicators are recorded in real time. The music chosen is either music the patient prefers or symphonic music with significant variations in melody and rhythm.
[0022] S2. Preprocess the physiological signals to obtain the preprocessed signals;
[0023] Under static conditions, the raw EEG signals underwent preprocessing operations such as downsampling, filtering, interpolation, artifact removal, rereference, and segmentation. The electrodermal signal underwent preprocessing operations such as low-pass filtering, baseline correction, and artifact removal. The blood oxygen and electrocardiogram signals underwent preprocessing operations mainly based on filtering. The electrooculogram signal underwent preprocessing operations such as filtering, endpoint detection, and normalization. The raw EEG signals were acquired by 64 electrodes, representing different brain regions.
[0024] S3. Calculate the features of physiological signals based on preprocessed signals to obtain multidimensional features;
[0025] S3 specifically includes:
[0026] In a quiet state, the power spectral density characteristics and Hilbert yellow entropy characteristics of EEG signals were calculated.
[0027] Eigenvalue 1: Power spectral density
[0028] The short-time Fourier transform method involves performing a Fourier transform on the electroencephalogram (EEG) signal, namely:
[0029] (1)
[0030] in, It is the angular frequency, where t is the time variable. This is the integration variable, used to perform integration operations on the EEG signal along the time axis. For EEG signals at time The value of . For window functions, It is a complex exponential function and a core component of the Fourier transform.
[0031] The formula for calculating power spectral density is as follows:
[0032] (2)
[0033] The power spectral density of brainwave power. This is the amplitude obtained after the short-time Fourier transform.
[0034] Eigenvalue 2: Hilbert-Huang entropy
[0035] Signal entropy is a very good measure of signal complexity. The Hilbert transform formula for an EEG signal x(t) is as follows:
[0036] (3)
[0037] Where t is the time variable, For integration variables, This is an electroencephalogram (EEG) signal.
[0038] In the music state, phase lock value characteristics and event-related desynchronization or synchronization index characteristics are calculated based on EEG signals. Different values reflect the synchronization or desynchronization of neuronal firing.
[0039] In a musical context, the frequency-following response is extracted through time-frequency analysis. Its phase-locking value with the music and the event-related desynchronization or synchronization index are calculated, denoted as characteristic values P1 and P2, respectively. The frequency-following response is the steady-state response wave of the EEG to the periodic frequency components in the music, used to assess the brain's synchronization and coupling characteristics to audio stimuli. The core of the analysis lies in the extraction of two types of time-frequency features:
[0040] ① Phase Lock Value (PLV): The EEG signal and the music signal are resampled to obtain the sampled EEG signal and the sampled music signal. The time-frequency features of the EEG signal and the music signal with the same sampling rate are extracted based on the short-time Fourier transform. The phase value of the time-frequency features of the two is extracted using Hilbert transform. The difference between the two time-frequency feature phase values is calculated to obtain the phase difference value. The correlation between music and EEG is analyzed for the phase difference value of each frequency to obtain the phase lock value feature.
[0041] Feature 1: Phase Locking Value
[0042] Method: The Short-Time Fourier Transform (STFT) formula is as follows:
[0043] (4)
[0044] Where t is the time variable, For integration variables, To sample EEG signals or sample music signals at time The value of h(τ−t) is the window function (such as the Hamming window). It is suitable for analyzing steady-state frequency components (such as the fundamental frequency of music), but its resolution is limited by the window length. = This refers to the time-frequency characteristics of EEG signals or music signals, i.e., amplitude.
[0045] The amplitude extracted from the short-time Fourier transform is converted into a complex number form using the Hilbert transform, as shown in the following formula:
[0046] (5)
[0047] t is a time variable. For frequency;
[0048] Calculated It can be represented as:
[0049] (6)
[0050] Where A and B These represent the real and imaginary parts of the transformed sampled EEG signal or sampled music signal, respectively.
[0051] The calculated phase is:
[0052] (7)
[0053] The phases of the sampled EEG signal and the sampled music signal at the same time t and the same frequency f are respectively denoted as... and The formula for calculating the phase difference is:
[0054] (8)
[0055] The phase lock value (PLV), which reflects the temporal phase synchronization between the EEG signal and the musical stimulus, is calculated using the following formula:
[0056] (9)
[0057] in, Let N be the phase difference, and N be the number of sampling points within the selected time window.
[0058] PLV is suitable for steady-state frequency tracking (such as music fundamental frequency) and directly characterizes neural synchronization, but its time-frequency resolution is limited by the window length and it has low sensitivity to transient changes.
[0059] ② Event-related desynchronization or synchronization (ERD / ERS) index: After quantifying the dynamic changes of frequency band energy through a secondary time-frequency analysis method (ZAM distribution), the event-related desynchronization / synchronization index is calculated for the frequency band of interest within the time window.
[0060] Feature 2: Event-related desynchronization or synchronization index (based on quadratic time-frequency analysis)
[0061] Method: Time-frequency energy is extracted based on ZAM distribution, and event-related desynchronization or synchronization index features are calculated based on the time-frequency energy.
[0062] The ZhaoAtlasMarks (ZAM) distribution is given by the following formula:
[0063] (10)
[0065] Where t is the time variable, f is the frequency variable, and s is the integral variable for calculating local autocorrelation. τ is the delay variable, which determines the upper and lower limits of integration.
[0066] Points limit Let X(t) be a window function, and X(t) be the EEG signal. This represents the complex conjugate of the signal.
[0067] By reducing cross-term interference through ZAM distribution, the time-frequency resolution of transient features (such as musical dynamics and consonant initiation) can be enhanced.
[0068] The ERD / ERS exponent is calculated based on the time-frequency energy extracted from the ZAM distribution, using the following formula:
[0069] (11)
[0070] Among them, time-frequency energy It is calculated by averaging the amplitude of the time-frequency distribution over frequency and time, that is:
[0071] (12)
[0072] in, and Each time window represents a different time window. and each frequency band The number of sampling points within the time window was set to 3 seconds, the sliding overlap ratio was set to 50%, and the frequency bands selected were the beta band (13-30Hz) and the gamma band (31-45Hz). The time-frequency distribution is calculated using the ZAM method. Similarly, numerical... The calculation is as follows:
[0073] (13)
[0074] It is the time-frequency distribution estimated from EEG segments recorded within a reference time interval before the stimulus begins. This represents the number of samples in the next interval.
[0075] Suitable for non-steady-state signal analysis, it quantifies cortical activation modes through dynamic energy changes.
[0076] Phase-locked values capture the temporal synchronicity of the brainstem cortex, while the ERD / ERS index reflects the frequency band energy modulation of the cortex. The two complement each other to construct a model of auditory perception and emotional processing pathways.
[0077] In both quiet and music-related states, the power spectral density characteristics of the skin are calculated based on the skin conductance signal, the frequency characteristics of the eye are calculated based on the electrooculogram (EOG) signal, the power spectral density characteristics of the heart are extracted from the heart function based on the heart function signal, and the blood oxygen saturation is calculated based on the blood oxygen signal.
[0078] Skin conductance feature extraction:
[0079] The power spectral density of the skin charge in the range of 0.03 to 0.5 Hz was calculated using short-time Fourier transform and denoted as P3 as a characteristic of this parameter.
[0080] Blood oxygen feature extraction:
[0081] Blood oxygen saturation SpO2 is calculated using detection light of multiple wavelengths and denoted as P4.
[0082] SpO2 = 110-25R; (14)
[0083] R=(AC660 / DC660) / (AC940 / DC940); (15)
[0084] Wherein, AC660 / DC660 is the AC to DC ratio of the received red light portion; AC940 / DC940 is the AC to DC ratio of the received infrared light portion.
[0085] ECG feature extraction:
[0086] Heart rate variability was extracted, and the RR interval sequence was converted into a frequency domain signal through fast Fourier transform. The power spectral density of low frequency (0.04-0.15Hz) and high frequency (0.15-0.40Hz) of LF was calculated. The sympathetic-vagal balance coefficient was constructed by quantizing the LF / HF ratio and denoted as P5.
[0087] (16)
[0088] P LE It is the power spectral density of low-frequency LEDs, P HE It is the power spectral density of high-frequency HE.
[0089] Electrooculogram (EOG) feature extraction;
[0090] Vertical and horizontal electrooculogram (EOG) signals were extracted, and then continuous wavelet transform was used to extract blinking and saccadic movements from the eye movement data. The average frequency of these two signals was used as the EOG feature, denoted as P6 and P7, respectively.
[0091] A spike in electrical potential signal appears during blinking or saccades. Wavelet transform can detect this spike in potential signal, such as the Mexican hat wavelet transform, whose function is as follows:
[0092] (17)
[0093] Where t is the time variable. After performing the Mexican hat wavelet transform, the trough signals are inverted, turning all troughs into peaks. Then, the findpeaks peak detection function is used to detect all peaks and troughs, obtaining the electrooculographic features detected. During saccade detection, an appropriate threshold range is set according to the characteristics of human eye saccades, generally between 20 and 200 ms, which needs to be adjusted based on the actual signal sampling rate. All peaks within the set threshold range are marked, and then the peak and trough values are encoded, with peak values encoded as "1" and trough values encoded as "0". The "01" or "10" sequences within the time window threshold of 20-200 ms are used as candidate segments for saccades. Then, the slope between the start and end points of the candidate segment is calculated, and the correlation coefficient between the slope and the feature vector is calculated. A corresponding threshold is set, and when the calculated coefficient value is within the threshold range, the candidate segment is considered as the desired saccade segment. During blink detection, the signal polarity is verified using the same encoding method. If a positive peak corresponds to a blink (closed eye) and a negative peak corresponds to an open eye, then the "010" sequence is selected as a blink candidate, and a corresponding maximum segment length threshold is set, typically between 100 and 400 ms. Then, the correlation coefficient between the slope and the candidate blink segments can be calculated. When it reaches the pre-set threshold, the candidate blink segment is identified as a blink segment.
[0094] Then the mean frequency of saccades and blinks was calculated for each type of chronic consciousness disorder patient.
[0095] S4. After preprocessing the multidimensional features, input them into the neural network fusion model to obtain the fused features;
[0096] In this embodiment of the invention, seven feature variables were extracted: power spectral density of EEG (P1 under static conditions), Hilbert-Huang entropy of EEG (P2 under static conditions), phase lock value of frequency following response (P1 under music mode), event-related desynchronization / synchronization index (P2 under music mode), power spectral density of skin conductance in 0.03~0.5Hz (P3), HbO concentration (P4), sympathetic-vagal balance coefficient in heart rate variability (P5), blink frequency (P6), and scan rate (P7). Among them, features P1 and P2 are selected differently according to different modes of quiet and music.
[0097] The preprocessing of multidimensional features specifically includes:
[0098] The multidimensional features are normalized, and the corresponding two-dimensional spatial coordinates of the phase-lock value feature, event-related desynchronization or synchronization index feature, EEG power spectral density feature, and EEG Hilbert yellow entropy feature of each electrode after normalization are encoded to obtain single-electrode encoded features. The single-electrode encoded features include: phase-lock value encoded features, event-related desynchronization or synchronization index encoded features, EEG power spectral density encoded features, and EEG Hilbert yellow entropy encoded features.
[0099] In this embodiment of the invention, a min-max normalization method is employed to uniformly map feature values to the [0, 1] interval, and location encoding is performed using the two-dimensional spatial coordinates of each electrode. In location embedding, this invention uses the two-dimensional spatial coordinates of each electrode because the electrode location helps estimate the correlation between brain regions. The location of each electrode is defined by the well-established 10⁻⁵ system.
[0100] The 10-5 system, published in 2001, describes the placement of EEG acquisition electrodes. It is called the 5% system or the 10-5 system because it uses a 5% proportion of the total length along the contour lines between skull landmarks, while the 10-20 and 10-10 systems use 20% and 10% of the distance, respectively.
[0101] The neural network fusion model specifically includes:
[0102] The multidimensional feature fusion module is used to fuse preprocessed multidimensional features to obtain multidimensional fused features;
[0103] The multi-dimensional feature fusion module specifically includes:
[0104] The single-electrode encoded features are input into the convolutional layer for preliminary fusion to obtain preliminary fused features;
[0105] The initial fused features and the remaining normalized multidimensional features are input into a depthwise separable convolution to perform feature fusion and obtain deep fused features. The deep fused features are then input into a batch normalization and GeLU activation layer to obtain multidimensional single-electrode features.
[0106] The self-attention channel spatial fusion features are input into the feedforward network layer to obtain enhanced nonlinear features. The feedforward network layer includes a fully connected layer and a Gaussian error linear unit activation function. The enhanced nonlinear features are weighted and averaged to obtain fused features.
[0107] The channel space feature fusion module is used to perform linear transformation on multi-dimensional single-electrode features to generate queries, keys, and values. The queries, keys, and values are split into multiple heads, and the attention of each head is calculated and then concatenated. The attention score is calculated to obtain the channel weights, which are then assigned to the values. Next, a dropout operation is performed to randomly discard some attention weights to obtain the weighted representation of each head. The weighted sum of multiple heads is then linearly transformed to obtain the transformed features. The transformed features are then subjected to dropout operation, residual connection, and layer normalization to obtain the self-attention channel space fusion features.
[0108] The self-attention channel spatial fusion features are input into the feedforward network layer to obtain enhanced nonlinear features. The feedforward network layer includes a fully connected layer and a Gaussian error linear unit activation function. The enhanced nonlinear features are weighted and averaged to obtain fused features.
[0109] The channel space feature fusion module is specifically used to: split the single-layer self-attention of each head into three levels of attention, the three levels of attention including: local attention, mid-range attention and global attention, and calculate the attention of each head based on the three levels of attention.
[0110] In this embodiment of the invention, the structural local attention of the third-level attention is to perform a 3×3 depthwise convolution on Q and K of each head, forcing the attention to only calculate the relationship between adjacent electrodes.
[0111] Mid-range attention involves applying 7×7 dilated convolutions (dilation introduces gaps in the convolution block to expand the range) to Q and K, and (dilation=2) expands the receptive field to the mid-range range.
[0112] Global attention: Traditional self-attention mechanism, without convolution operation, directly calculates the interaction of all electrode pairs.
[0113] Finally, the three attentions are merged to obtain the attention of each head.
[0114] In this embodiment of the invention, based on the traditional multi-head self-attention mechanism, the fused features are transformed linearly to generate a query (Q), a key (K), and a value (V), which are d-dimensional. Q represents each channel used for comparison, K represents all comparisons with channels in Q, and V represents the representation in the high-level feature space. Q, K, and V are split into h heads, and attention is calculated for each head before concatenation, as shown in the following formula:
[0115] (18)
[0116] in, The weight matrices for different heads are multiplied by the concatenated multi-head attention matrix to obtain the final output. The single-layer self-attention of each head is decomposed into three levels of attention: local, mid-range, and global. The modeling of inter-electrode dependencies is expanded from the local to the global level. Local attention targets a 3×3 region, focusing only on electrodes adjacent to each channel; mid-range attention targets a 7×7 region, progressively expanding the receptive field; global attention performs fully connected interactions to obtain long-distance channel spatial dependencies across brain regions. An exponential decay of channel (electrode) distance is introduced into the attention score calculation, as shown in the following formula:
[0117] (19)
[0118] in, The transpose of K. Let λ be the distance between lead i and lead j, λ be the distance attenuation coefficient, and d represent... It is d-dimensional.
[0119] For each head, the product result is divided by the square root of d as data normalization to ensure the effectiveness of the Softmax function. This invention obtains the weight of each channel and assigns it to V. Then, a dropout operation is performed to randomly discard some attention weights to prevent overfitting, obtaining a weighted representation of each head. Finally, multiple heads are summed with weights and subjected to a linear transformation to maintain consistency between the output dimension and the module input dimension.
[0120] Next, dropout, residual connections, and layer normalization are performed to help the network learn and converge adaptive weights faster, alleviating the gradient vanishing problem caused by large convolutional kernels.
[0121] The self-attention channel spatial fusion features are input into the feedforward network layer to obtain enhanced nonlinear features. The feedforward network layer includes a fully connected layer and a Gaussian error linear unit activation function. The enhanced nonlinear features are weighted and averaged to obtain fused features.
[0122] In this embodiment of the invention, the nonlinearity of the feedforward network layer is enhanced by inputting features obtained through self-attention. An FFN is constructed using two fully connected layers and a Gaussian error linear unit (GeLU) activation function, as shown in the following formula:
[0123] (20)
[0124] Where x represents the input feature, (x) refers to the Sigmoid function. , Here are the weight matrix and bias terms for the first fully connected layer. , This represents the weight matrix and bias terms of the second fully connected layer.
[0125] Then, the present invention adjusts the dropout rate to 0.1 and performs the same operation again to obtain better training results. After repeating the above operation multiple times, the present invention performs a weighted average on the spatial fusion output to obtain a global high-level representation that considers all channels and features.
[0126] S5. The fusion feature input pattern recognition output module obtains the subject's consciousness impairment assessment score.
[0127] S5 specifically includes:
[0128] The input pattern recognition module with fused features is processed to obtain the probability of the subject's minimum conscious state and vegetative state, and the level of consciousness score is calculated based on the probability.
[0129] The process of inputting the fused features into the pattern recognition output module for processing to obtain the probability of the subject's minimum state of consciousness and vegetative state, and calculating the consciousness level score based on the probability, specifically includes: performing global average pooling on the fused features to obtain key features; inputting the key features into a fully connected layer with a Softmax function to obtain the probability of the minimum state of consciousness and the probability of the vegetative state for detecting the consciousness level; and multiplying the probability of the minimum state of consciousness and the probability of the vegetative state by 100 to obtain the consciousness impairment assessment score.
[0130] In this embodiment of the invention, global average pooling is applied to the model output layer to compress the channel dimension while preserving key features. A fully connected layer with a Softmax function maps the output values to the [0, 1] interval, and two neurons are output to represent the probabilities of detecting the minimum conscious state (MCS) and vegetative state (VS) levels of consciousness. The SIOU loss function is used for optimization.
[0131] The output value is then multiplied by 100 to obtain the final consciousness level score, ranging from 0 to 100. The formula is as follows:
[0132] ; (twenty one)
[0133] in, Here, z is the Softmax function, and z is the weighted sum of the output layer.
[0134] In this embodiment of the invention, to ensure that the model results are independent of the dataset partitioning, a k-fold cross-validation method is used. The dataset mentioned in the model training is divided into k equal parts (k=5 in this embodiment), and training and validation are performed k times. In each validation, one part is used as the validation set, and the other k-1 parts are used as the training set. Different parts are used as the validation set each time. After k times, the entire dataset participates in both the training and validation sets. Finally, when calculating the model accuracy, the average value of the different validation data is taken to obtain the final model accuracy.
[0135] In this embodiment of the invention, multiple physiological signal sensors, a memory, a processor, and a computer program capable of executing the overall technical solution are employed. The system provides real-time assessment of consciousness levels in both quiet and music-based modes, assigns scores, stores dynamic physiological signals and real-time assessment scores over 72 hours, and stores the daily average assessment scores over the years.
[0136] Beneficial effects:
[0137] 1. This invention uses personalized musical stimulation (music the patient prefers) or symphonic music with large melodic variations to assess the patient's condition in real time based on music therapy and indicate the effectiveness of the treatment strategy;
[0138] 2. Existing technologies typically analyze cross-frequency coupling and brain region connectivity of EEG in patients with impaired consciousness under music stimulation based on frequency domain and connectivity analysis. This provides intermittent assessment and fails to provide real-time feedback on changes in EEG signals with musical melody. This technology employs time-frequency analysis to extract EEG features, extracting the frequency-following response of the patient's EEG within a smaller time window and calculating its phase-locking value with the musical melody and its event correlation / discorrelation index. This enables dynamic calculation of music-EEG coupling, capturing dynamic changes in the patient's state of consciousness. Simultaneously, it integrates other physiological signals to overcome the limitations of relying on a single signal.
[0139] 3. For training a bimodal model in both quiet and musical scenarios, a novel approach is employed: large-kernel deep separable convolutions are used to fuse single-channel EEG features with other multi-physiological signals. Hierarchical expansion of attention optimizes spatial feature fusion, enabling accurate and efficient assessment and classification of consciousness disorders, covering subtle changes in patients' levels of consciousness throughout the day. The system outputs a consciousness level score for the corresponding modality in that scenario, using quantifiable numerical values to reduce the comprehension burden on doctors and patients' families, thereby improving medical efficiency.
[0140] 4. Continuously collect multiple physiological signals (EEG, GSR, etc.) under quiet and musical stimulation conditions.
[0141] Seven features were extracted from SpO2, EOG, and ECG respectively. After fusing the features through a large kernel depthwise separable convolution, a bimodal model based on hierarchical expanded attention was trained. This invention enables effective dynamic assessment of the degree of consciousness impairment in patients with consciousness disorders in a dual-modal environment of music and quiet.
[0142] 5. Under music-driven conditions, calculate the phase-locked value of the frequency-following response and the event-related synchronization / desynchronization index. The frequency-following response, i.e., the steady-state response wave of the EEG to the periodic frequency components in music, is used to assess the brain's synchronization and coupling characteristics to audio stimuli. To achieve dynamic matching with the music, two time-frequency analysis methods—short-time Fourier transform and ZAM distribution—are used to extract two types of frequency-following response features. The frequency characteristics of music and EEG are extracted within a short time window to reflect the patient's following response to the frequency characteristics in the musical melody from the perspectives of phase-phase coupling and energy distribution, dynamically assessing the patient's state of consciousness.
[0143] 6. A dual-modal deep learning model trained for quiet and musical stimuli based on large-kernel convolution and hierarchical expanded attention: Two different modal models are trained using different input and output data for different scenarios, namely music and quiet. By employing large-kernel deep separable convolution, it can cover more local relationships between feature dimensions when applied to multimodal physiological signals. At the same time, hierarchical expanded attention is used to achieve cross-channel spatial fusion, gradually expanding the receptive field with three levels of attention: "local-mid-range-global". A spatial decay factor is introduced to focus on key brain regions and avoid global computational redundancy. The scene is determined by the music playback device and the evaluation model is switched accordingly, which is suitable for daily assessment of patients with impaired consciousness and for evaluating the effect of music therapy.
[0144] 7. This method integrates multiple physiological signals to extract the frequency-following response features of the coupling between melody and EEG under music (phase-locking value and event correlation / decorrelation index are extracted after time-frequency analysis). Based on the adaptive attention mechanism, it can assess the level of consciousness of patients with consciousness disorders under quiet and music modalities, which helps to objectively reflect the treatment effect.
[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions to the technical solutions of the embodiments of the present invention do not cause the essence of the corresponding technical solutions to deviate from the scope of the present solution.
Claims
1. A method for dynamic assessment of the degree of consciousness impairment in a multimodal context, characterized in that, include, S1. Acquire physiological signals of the subject under multimodal conditions; S2. Preprocess the physiological signals to obtain the preprocessed signals; S3. Calculate the features of physiological signals based on preprocessed signals to obtain multidimensional features; S4. Preprocess the multidimensional features to obtain preprocessed data; S5. Merge the preprocessed data to obtain fusion features; S6. Input the fused feature into the pattern recognition output module to obtain the subject's consciousness impairment assessment score; Specifically, S1 includes: real-time acquisition of multi-electrode electroencephalogram (EEG), electrodermal conductance (TEA), electrooculogram (EOG), electrocardiogram (ECG), and blood oxygen saturation signals of the subject in both music and quiet states. S3 specifically includes: in a quiet state, calculating the EEG power spectral density feature and the EEG Hilbert yellow entropy feature based on the preprocessed EEG signal of each electrode; in a music state, calculating the phase lock value feature and the event-related desynchronization or synchronization index feature based on the preprocessed EEG signal of each electrode; in both quiet and music states, calculating the skin spectral density feature based on the preprocessed skin spectral signal, calculating the eye frequency feature based on the preprocessed eye spectral signal, extracting the heart spectral density feature from the heart spectral signal based on the preprocessed heart spectral signal, and calculating the blood oxygen saturation feature based on the preprocessed blood oxygen signal; S4 specifically includes: normalizing the multidimensional features, and encoding the corresponding two-dimensional spatial coordinates of the phase lock value feature, event-related desynchronization or synchronization index feature, EEG power spectral density feature and EEG Hilbert yellow entropy feature of each electrode after normalization to obtain single-electrode encoding features. The single-electrode encoding features include: phase lock value encoding features, event-related desynchronization or synchronization index encoding features, EEG power spectral density encoding features and EEG Hilbert yellow entropy encoding features. S6 specifically includes: processing the fusion feature input pattern recognition output module to obtain the probability of the test subject's minimum conscious state and vegetative state, and calculating the consciousness level score based on the probability.
2. The method according to claim 1, characterized in that, The calculation of phase-locked value features and event-related desynchronization or synchronization index features based on preprocessed EEG signals for each electrode specifically includes: The EEG signal and music signal are resampled to obtain sampled EEG signal and sampled music signal. The time-frequency features of the EEG signal and the music signal with the same sampling rate are extracted based on the short-time Fourier transform. The phase value of the time-frequency feature is extracted by Hilbert transform. The difference between the two time-frequency feature phase values is calculated to obtain the phase difference value. The correlation value between music and EEG is analyzed for the phase difference value of each frequency to obtain the phase lock value feature. Time-frequency energy is extracted based on ZAM distribution, and event-related desynchronization or synchronization index features are calculated based on the time-frequency energy.
3. The method according to claim 1, characterized in that, S5 specifically includes: The single-electrode encoded features are input into the convolutional layer for preliminary fusion to obtain preliminary fused features; The initial fused features and the remaining normalized multidimensional features are input into a depthwise separable convolution to perform feature fusion and obtain deep fused features. The deep fused features are then input into a batch normalization and GeLU activation layer to obtain multidimensional single-electrode features. A linear transformation is performed on the multidimensional single-electrode features to generate queries, keys, and values. The queries, keys, and values are then split into multiple heads. The attention of each head is calculated separately and then concatenated. The attention score is calculated to obtain the channel weights. The weights are assigned to the values. Then, a dropout operation is performed to randomly discard some attention weights to obtain the weighted representation of each head. The weighted sum of the multiple heads is then linearly transformed to obtain the transformed features. The transformed features are then subjected to dropout operation, residual connection, and layer normalization to obtain the self-attention channel space fusion features. The self-attention channel spatial fusion features are input into the feedforward network layer to obtain enhanced nonlinear features. The feedforward network layer includes a fully connected layer and a Gaussian error linear unit activation function. The enhanced nonlinear features are weighted and averaged to obtain fused features.
4. The method according to claim 3, characterized in that, The specific steps of calculating the attention of multiple heads include: splitting the single-layer self-attention of each head into three levels of attention, namely: local attention, mid-range attention and global attention, and calculating the attention of each head based on the three levels of attention.
5. The method according to any one of claims 1 to 4, characterized in that: The process of inputting the fused features into the pattern recognition output module for processing to obtain the probability of the subject's minimum state of consciousness and vegetative state, and calculating the consciousness level score based on the probability, specifically includes: performing global average pooling on the fused features to obtain key features, inputting the key features into a fully connected layer with a Softmax function to obtain the probability of the minimum state of consciousness and the probability of the vegetative state for detecting the consciousness level, and multiplying the probability of the minimum state of consciousness and the probability of the vegetative state by 100 to obtain the consciousness impairment assessment score.
Citation Information
Patent Citations
Multi-modal fusion alertness assessment method and system based on cross attention
CN120124011A
Multimodal analysis combining monitoring modalities to elicit cognitive states and perform screening for mental disorders
US20210319897A1