Device for evaluating consciousness level and storage medium
By collecting clinical information and task-oriented EEG signals, and using a multimodal large language model to perform cross-modal attention calculation, the subjectivity and insufficient information fusion of existing consciousness assessment methods are solved, and the automated and accurate assessment of consciousness level is achieved.
Patent Information
- Application Number
- CN202511777891.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-02-27
AI Technical Summary
Existing methods for assessing consciousness are highly subjective and lack sufficient integration of multi-source information, making it difficult to achieve accurate and automated assessments of consciousness levels.
By collecting clinical information and task-oriented EEG signals, extracting frequency domain features and spatiotemporal features, using a multimodal large language model to perform cross-modal attention calculation, and fusing EEG features and text features, consciousness assessment can be achieved.
It improves the objectivity and accuracy of consciousness assessment, realizes automated assessment of consciousness level, and adapts to the non-standard responses of patients with different brain injuries.
Smart Images

Figure CN121570132A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application generally relates to the technical field of medical artificial intelligence. More specifically, the present application relates to an apparatus for consciousness level assessment and a computer readable storage medium. BACKGROUND
[0002] Objective assessment of consciousness level is a core challenge in neuroscience and clinical medicine, and its results directly affect the diagnosis and treatment of patients with consciousness disorders, prognosis and rehabilitation intervention effect. Whether it is consciousness impairment caused by cerebrovascular diseases such as stroke and hypoxic encephalopathy, or post-traumatic consciousness disorder, accurate judgment of the consciousness state of patients (such as coma, micro-consciousness state, etc.) is a key prerequisite for clinical diagnosis and treatment.
[0003] Currently, the mainstream consciousness assessment method in the clinic relies on behavior scales (such as the Coma Recovery Scale-Revised, CRS-R), but such methods are easily affected by the subjectivity of the evaluator, the limited motor function of the patient (such as limb paralysis), and other factors, and the objectivity and accuracy of the assessment results are limited; at the same time, existing assessment models based on electroencephalogram signals treat electroencephalogram signals as pure numerical sequences, and fail to deeply associate them with the semantic connotation of the stimulus, the corresponding neural pathways, and the patient's clinical pathophysiological knowledge, and lack effective fusion mechanisms for electroencephalogram features and clinical information, making it difficult to fully exploit the synergistic value of multi-source information and unable to meet the clinical demand for accurate and automated consciousness assessment.
[0004] Therefore, there is an urgent need to provide a scheme for consciousness level assessment in order to solve the problems of strong subjectivity and insufficient multi-source information fusion of existing assessment methods, and to realize automatic and accurate assessment of the consciousness level of patients. SUMMARY
[0005] In order to at least solve one or more of the above-mentioned technical problems, the present application proposes a scheme for consciousness level assessment in multiple aspects.
[0006] In a first aspect, this application provides an apparatus for assessing level of consciousness, comprising: a processor; and a memory storing computer instructions for assessing level of consciousness, wherein when the computer instructions are executed by the processor, the following operations are performed: acquiring clinical information of the subject to be assessed and task-state EEG signals under a target stimulus paradigm; extracting frequency domain features of a specific frequency band and spatiotemporal features of target event-related potentials based on the task-state EEG signals, and combining them into a corresponding EEG topographic map; inputting the corresponding EEG topographic map into a multimodal large language model, and using an image encoder in the multimodal large language model to extract image features to obtain EEG features; inputting the clinical information into the multimodal large language model, and using a text encoder in the multimodal large language model to extract text features to obtain text features; and using a cross-modal fusion module in the multimodal large language model to perform cross-modal attention calculation fusion of the EEG features and the text features to achieve consciousness assessment, thereby outputting a consciousness assessment result.
[0007] In some embodiments, the clinical information includes at least one or more of the following: brain injury location, degree of injury, disease etiology, and medical history; the target stimulus paradigm includes one or more of the following: auditory stimulus paradigm, visual stimulus paradigm, or pain stimulus paradigm.
[0008] In some embodiments, the device further performs the following operations to extract frequency domain features of a specific frequency band: calculating energy values in the specific frequency band based on the task-state EEG signal to extract the frequency domain features of the specific frequency band.
[0009] In some embodiments, the device further performs the following operations to extract the spatiotemporal features of the target event-related potentials: extracting the effective signal corresponding to the target event-related potentials based on the task-state EEG signal; calculating the average voltage of the effective signal at key time points to obtain a voltage distribution map; and calculating the amplitude and latency based on the voltage distribution map to extract the spatiotemporal features of the target event-related potentials, wherein the target event-related potentials include P300 and N1.
[0010] In some embodiments, the apparatus further performs the following operations: filtering the voltage distribution map using a target kernel function.
[0011] In some embodiments, the device further performs the following operations: calculating variance information of brainwave topography at different time periods; and fusing the variance information in cross-modal attention calculation.
[0012] In some embodiments, the apparatus further performs operations of: constructing a mapping relationship, wherein the mapping relationship comprises a corresponding mapping between a stimulation mode, a neural pathway, a key brain region, and a feature type; and embedding the mapping relationship into the text feature to guide cross-modal attention calculation.
[0013] In some embodiments, the apparatus further performs operations of: adaptively adjusting a weight of a target spatiotemporal feature based on a text feature embedded with the mapping relationship, to optimize a spatiotemporal feature of the target event-related potential.
[0014] In some embodiments, the apparatus further performs operations of: calculating a correlation feature between the spatiotemporal feature of the target event-related potential and the variance information; and fusing the correlation feature with the electroencephalogram feature.
[0015] In a second aspect, the present application provides a computer-readable storage medium having stored thereon computer program instructions for consciousness level assessment, which, when executed by one or more processors, cause the operations performed by the apparatus of one or more embodiments of the preceding first aspect to be implemented.
[0016] Through the scheme for consciousness level assessment as provided above, the embodiments of the present application solve the problems of strong subjectivity, difficulty in cross-pathway fusion, and insufficient adaptation for brain injury patients of traditional assessment methods, and improve the objectivity, accuracy, and automation level of consciousness assessment, by collecting clinical information and multi-paradigm task-state electroencephalogram signals, extracting frequency domain features and ERP spatiotemporal features and constructing an electroencephalogram topographic map, using a double-encoder of a multi-modal large language model to extract features, and achieving precise consciousness assessment through cross-modal attention fusion. BRIEF DESCRIPTION OF DRAWINGS
[0017] The above and other objects, features and advantages of the example embodiments of the present application will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which several embodiments of the present application are shown by way of example, and wherein like or corresponding elements show like or corresponding parts, by referring to which; drawings: Figure 1 is an exemplary structural block diagram illustrating an apparatus 100 for consciousness level assessment according to an embodiment of the present application; Figure 2 is an exemplary flow block diagram illustrating operations 200 implemented by an apparatus for consciousness level assessment according to an embodiment of the present application; Figure 3 is an exemplary schematic diagram illustrating an electroencephalogram topographic map according to an embodiment of the present application; Figure 4 is an exemplary flow block diagram illustrating the overall consciousness level assessment according to an embodiment of the present application; Figure 5 An exemplary structural block diagram of an electronic device 500 according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort fall within the scope of the present application.
[0019] It should be understood that the terms “comprise” and “include” used in the specification and claims of the present application indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0020] It should also be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification and claims of the present application, the singular forms “a”, “an” and “the” are intended to include the plural forms, unless the context clearly indicates otherwise. It should be further understood that the term “and / or” used in the specification and claims of the present application means any combination of one or more of the associated listed items and all possible combinations thereof, and includes these combinations.
[0021] As used in the specification and claims of the present application, the term “if’ can be interpreted as “when” or “upon” or “in response to a determination” or “in response to detecting” depending on the context. Similarly, the phrase “if determined” or “if detected [the described condition or event]” can be interpreted as meaning “upon determining” or “in response to determining” or “upon detecting [the described condition or event]” or “in response to detecting [the described condition or event]” depending on the context.
[0022] The specific embodiments of the present application will be described in detail below in conjunction with the accompanying drawings.
[0023] Figure 1 is an exemplary structural block diagram showing an apparatus 100 for consciousness level assessment according to an embodiment of the present application. As shown in Figure 1As shown in FIG. 1, the apparatus 100 can include a processor 110 and a memory 120. The processor 110 can include, for example, a general-purpose processor (“CPU”) or a special-purpose graphics processor (“GPU”), and the memory 120 can store program instructions executable on the processor. In some embodiments, the memory 120 can include, but is not limited to, a resistive random access memory (“RRAM”), a dynamic random access memory (“DRAM”), a static random access memory (“SRAM”), and an enhanced dynamic random access memory (“EDRAM”).
[0024] Further, the memory 120 can store program instructions for consciousness level assessment, and when the program instructions are executed by the processor 110, the apparatus 100 can implement the following operations: collecting clinical information of a subject to be assessed and a task-state electroencephalogram signal under a target stimulus paradigm; extracting frequency-domain features of a specific frequency band and spatiotemporal features of a target event-related potential based on the task-state electroencephalogram signal, and combining the features into a corresponding electroencephalogram topography; inputting the corresponding electroencephalogram topography into a multi-modal large language model, using an image encoder in the multi-modal large language model to extract image features, and obtaining electroencephalogram features; inputting the clinical information into the multi-modal large language model, using a text encoder in the multi-modal large language model to extract text features, and obtaining text features; using a cross-modal fusion module in the multi-modal large language model to perform cross-modal attention calculation and fusion on the electroencephalogram features and the text features to implement consciousness assessment, and outputting a consciousness assessment result. Details of the operations implemented by the apparatus 100 according to embodiments of the present application will be described below. Figure 2 The operations implemented by the apparatus 100 according to embodiments of the present application will be described in detail.
[0025] Figure 2 is an exemplary flow chart illustrating operations 200 implemented by an apparatus for consciousness level assessment according to embodiments of the present application. As shown in Figure 2 As shown in FIG. 1, at step S201, the clinical information of a subject to be assessed and a task-state electroencephalogram signal under a target stimulus paradigm are collected.
[0026] In some embodiments, the clinical information includes at least one or more of a brain injury site, an injury degree, a disease etiology, and a diagnosis and treatment history. The target stimulus paradigm includes one or more of an auditory stimulus paradigm, a visual stimulus paradigm, or a pain stimulus paradigm.
[0027] It can be understood that the clinical information refers to the structured data reflecting the basic state of brain function of the patient and the disease background, and is a core bridge for associating the electroencephalogram features with the pathophysiological mechanism. In some implementation scenarios, the clinical information can be obtained through, for example, automatic reading through an electronic medical record system (EMR) interface, manual input by medical staff through a special input terminal, and the like. For example, the specific brain injury site is determined, such as the left temporal lobe, the right occipital lobe, the bilateral frontal lobe, and the like; the injury degree is recorded according to the clinical classification standard, such as mild / moderate / severe; the pathogenic cause is collected, such as cerebral apoplexy, anoxic encephalopathy, traumatic brain injury, or other cerebrovascular diseases; and the onset time, surgical intervention, or rehabilitation treatment measures, and the like are collected.
[0028] The target stimulation paradigm refers to a standardized stimulation presentation mode capable of stably inducing event-related potential (ERP) components, such as auditory, visual, or pain, and the like. Preferably, at least two paradigms (for example, auditory and visual stimulation) can be used to cover multiple neural pathways, obtain multi-dimensional electroencephalogram responses, and avoid evaluation blind spots caused by single pathway damage. The task-state electroencephalogram signal is a continuous signal of the neural electrical activity of the brain recorded by a multi-channel electroencephalogram device when the subject receives standardized stimulation, and includes spontaneous electroencephalogram signals and ERPs directly related to the stimulation, and is a core data source for extracting objective features related to consciousness. In implementation scenarios, a multi-channel electroencephalogram recorder (for example, 32 or 64 channels) can be used to cover the core brain regions with a standard electrode distribution. For example, the temporal lobe electrodes T3, T4, T5, T6 related to the auditory pathway, or the occipital lobe electrodes O1, O2, Oz related to the visual pathway, and the like.
[0029] Then, at step S202, the frequency domain features of a specific frequency band and the spatiotemporal features of a target event-related potential are extracted based on the task-state electroencephalogram signal, and are combined into a corresponding electroencephalogram topographic map. In some implementation scenarios, before the features are extracted, the original task-state electroencephalogram signal can be preprocessed through denoising, artifact removal, and the like. For example, the preprocessed electroencephalogram signal has a significantly reduced noise and a clear and prominent ERP waveform, which provides a guarantee for the accuracy of subsequent feature extraction.
[0030] In some embodiments, the apparatus described above can further perform the following operation to extract the frequency domain features of a specific frequency band: calculating an energy value in the specific frequency band based on the task-state electroencephalogram signal to extract the frequency domain features of the specific frequency band. Through the energy distribution of the electroencephalogram signal in a specific frequency range, the intensity of the brain electrical activity in the frequency band can be reflected. In some implementation scenarios, the aforementioned specific frequency band may, for example, be an Alpha wave (8~13 Hz) or a Beta wave (13~30 Hz), which can cover the main frequency domain range related to consciousness evaluation.
[0031] Specifically, first, the pre-processed electroencephalogram signal can be based on the stimulation application time point as 0 time point, and a fixed length data segment, i.e., Epoch, is intercepted from the continuous electroencephalogram signal. For example, as an example, during the experiment, when a certain specific stimulus (such as a sound, a visual image) appears, a mark is made in the continuous electroencephalogram signal through software to lock the precise time of the stimulus occurrence. Taking the event marker as the time zero point, data from 200 ms before the event to 600 ms after the event is intercepted from the continuous electroencephalogram signal to form an Epoch with a total duration of 800 ms. Usually one stimulus corresponds to one Epoch, and the Epoch signals under the same paradigm are superimposed and averaged for subsequent calculation.
[0032] Next, the Epoch signal is subjected to fast Fourier transform, and the power spectral density under a specific frequency band is calculated by the periodogram method. The power spectral density in the specific frequency band is integrated to obtain the frequency band energy value, so as to obtain the frequency domain feature of the specific frequency band.
[0033] In some embodiments, the above device further performs the following operations to extract the spatiotemporal features of the target event-related potential: extracting the effective signal corresponding to the target event-related potential based on the task-state electroencephalogram signal; calculating the voltage mean value of the effective signal at the key time point to obtain a voltage distribution map; and calculating the amplitude and latency based on the voltage distribution map to extract the spatiotemporal features of the target event-related potential, wherein the target event-related potential includes P300 and N1.
[0034] It can be understood that P300 is a positive waveform (potential value higher than baseline) appearing 250-350 ms after stimulation, which is related to the cognitive processing of brain attention, memory, etc. to the stimulus. N1 is a negative waveform (potential value lower than baseline) appearing 0-100 ms after stimulation, which is related to the early perception processing of the brain to the stimulus. In an implementation scenario, based on the feature period of the ERP component, the effective signal segment is intercepted from the superimposed and averaged Epoch signal, for example, the P300 effective signal segment is 250-350 ms, covering the rising to peak period of P300; the N1 effective signal segment is 0-100 ms, covering the appearance to valley period of N1.
[0035] In one exemplary scenario, for P300, 5 key time points can be selected at an interval of 50 ms, i.e., 250 ms, 275 ms, 300 ms, 325 ms, and 350 ms, and the voltage mean value at each time point is the average value of the 25 ms time window before and after the corresponding time point. For example, the mean value at 250 ms = the average of the voltage values at 225-275 ms, the mean value at 300 ms = the average of the voltage values at 275-325 ms, and so on, to ensure that the mean value can reflect the stable potential level at the time point.
[0036] For N1, 4 key time points are selected at 50ms intervals, i.e. 50ms, 75ms, 100ms, 125ms and 150ms. The voltage mean value at each time point is the average value of the 25ms time window before and after the corresponding time point. For example, the mean value at 50ms time point = the average of the voltage values of 25~75ms.
[0037] Further, by adopting a two-dimensional canvas of, for example, 512x512 pixels, the positions of the brain regions are defined according to the real topological relationship of the brain, ensuring that the mapping of the electrode positions and the brain regions is accurate and unbiased. The electrode voltage mean value at each key time point is mapped to the corresponding pixel of the canvas according to the voltage value and the color gradient, and a voltage distribution map is obtained.
[0038] Then, the amplitude and latency are calculated based on the voltage distribution map. The amplitude of P300 is the maximum value in the voltage mean value at the key time point, for example, represented as p300_amplitude = np.max (data, axis = 0). The latency of P300 is the key time point corresponding to the P300 amplitude, for example, represented as p300_latency = np.argmax (data, axis = 0). Similarly, the amplitude of N1 is the minimum value in the voltage mean value at the key time point, for example, represented as N1_amplitude = np.min (data, axis = 0), and the latency of N1 is the key time point corresponding to the N1 amplitude, for example, represented as N1_latency = np.argmin (data, axis = 0). Wherein, data is the voltage matrix (i.e. voltage distribution), axis = 0 is the maximum (minimum) value along the time dimension or the index corresponding to the maximum (minimum) value along the time dimension.
[0039] In the implementation scenario, the P300 amplitude and latency and the N1 amplitude and latency are calculated based on different paradigms respectively. For example, for P300 amplitude under auditory paradigm, p300_h_amplitude = F_p300_amplitude (EEG_h), P300 latency p300_h_latency = F_p300_latency (EEG_h); N1 amplitude N1_h_amplitude = F_N1_amplitude (EEG_h), N1 latency N1_h_latency = F_N1_latency (EEG_h).
[0040] The P300 amplitude in the visual paradigm is p300_v_amplitude = F_p300_amplitude (EEG_v), the P300 latency is p300_v_latency = F_p300_latency (EEG_v), the N1 amplitude is N1_v_amplitude = F_N1_amplitude (EEG_v), and the N1 latency is N1_v_latency = F_N1_latency (EEG_v).
[0041] The voltage distribution map intuitively presents the differences in the brain region distribution of the ERP. For example, the P300 of a person with clear consciousness has high voltage in the temporal lobe / occipital lobe, and the coma has no obvious high voltage area, which provides spatial characteristics for the model. The amplitude and latency can objectively reflect the cognitive processing ability related to consciousness. The standardization of feature extraction is ensured by operatorization and formulaization calculation to avoid human error, and the feature extraction logic under different stimulation paradigms is unified to lay the foundation for cross-channel fusion.
[0042] In some implementation scenarios, by adopting, for example, a dual-color channel to carry the frequency domain features and the space-time features, or by adopting, for example, a multi-frame sequence diagram, the frequency domain energy values corresponding to each electrode are mapped to the same two-dimensional canvas as the voltage distribution map according to the brain topological position, to form a frequency domain energy distribution map, so as to ensure the spatial position alignment with the time domain features, and to form an electroencephalogram topographic map in combination. In order to provide complete spatial, temporal, and frequency three-dimensional feature inputs for the multi-modal model. The pixel value range of the combined electroencephalogram topographic map can be normalized to 0-255 to adapt to the input requirements of the image encoder.
[0043] In some embodiments, the apparatus described above can further perform the following operation: filtering the voltage distribution map using a target kernel function. The target kernel function refers to a filter window used to smooth the voltage distribution map, which realizes the sharing of information of adjacent electrodes by reducing noise. Specifically, the voltage distribution map can be filtered by adopting a 10x10 size mean filter kernel or a Gaussian kernel. The mean filter can effectively eliminate the abnormal fluctuations of a single electrode (such as “isolated bright spots” caused by random noise), so that the real spatial distribution of the ERP features is smoother. The Gaussian smoothing can better preserve the detailed features of the voltage gradient while reducing noise, so that the filtered voltage distribution map has low noise and coherent information, and avoids the influence of abnormal data of a single pixel on the feature extraction accuracy of the subsequent image encoder.
[0044] Based on the electroencephalogram topography under different paradigms, at step S203, the corresponding electroencephalogram topography is input to the multi-modal large language model, the image encoder in the multi-modal large language model is used for image feature extraction, and the electroencephalogram feature is obtained. In some implementation scenarios, the image encoder may be, for example, a Transformer or a convolutional neural network (CNN), etc., which converts multi-dimensional information such as spatial distribution, numerical size and time sequence change into a structured vector, captures the spatial correlation between brain regions, and obtains accurate electroencephalogram features.
[0045] Further, at step S204, the clinical information is input to the multi-modal large language model, and the text feature is obtained by using the text encoder in the multi-modal large language model for text feature extraction. The text encoder is used to analyze the text semantic features, can divide the clinical information according to the model dictionary, convert it into a structured word embedding vector, and obtain the text feature.
[0046] Finally, at step S205, the cross-modal attention calculation fusion of the electroencephalogram feature and the text feature is realized by using the cross-modal fusion module in the multi-modal large language model to realize the consciousness evaluation, and the consciousness evaluation result is output.
[0047] Specifically, assuming that the electroencephalogram feature is Feature_I and the text feature is Feature_L, when performing cross-modal attention calculation, first, the similarity matrix of Feature_I and Feature_L is calculated. The similarity matrix reflects the correlation degree of the text semantic and the electroencephalogram feature, for example, the similarity of the left temporal lobe injury and the temporal lobe electroencephalogram feature, p300_h_amplitude is high. Then, after normalization by Softmax, Feature_I is weighted and summed to obtain the semantic-guided electroencephalogram feature vector Feature_I'. By concatenating Feature_I' and Feature_L, a cross-modal fusion feature vector is formed, and the cross-modal fusion feature vector is subjected to feature fusion and strengthening through a fully connected layer and an activation function, and an ultimate reasoning feature vector is output.
[0048] In an implementation scenario, the inference feature vector is classified by, for example, a decoder, to obtain a final consciousness evaluation result. The consciousness evaluation result includes, for example, {consciousness_level: coma / unresponsive wakefulness syndrome / minimally conscious state / clear consciousness; confidence_level: high / medium / low; key_evidence: supporting evidence, for example, p300_v_amplitude is 2.8 μV (normal range 2~5 μV), p300_v_amplitude_various is 0.8 μV², and occipital visual pathway voltage distribution is obvious; clinical_correlation: associated with clinical information, for example, the evaluation result is consistent with the expectation that the left hemisphere appears compensatory after the right temporal parietal lobe of the patient is damaged}.
[0049] The cross-modal attention calculation realizes accurate alignment of text semantics and electroencephalogram features, so that the model preferentially pays attention to electroencephalogram features related to clinical background and mapping rules, and solves the technical problem that the fusion of electroencephalogram features and clinical information is only on the surface. The fused feature vector contains both electroencephalogram objective features and text semantic knowledge, providing comprehensive and accurate basis for consciousness level evaluation.
[0050] In some embodiments, the above apparatus further performs the following operations: calculating variance information of the electroencephalogram topography in different time periods, and fusing the variance information in the cross-modal attention calculation. It can be understood that the variance information refers to the variance of the amplitude and latency of each electrode between dozens of acquisitions, reflecting the time stability of the ERP feature. The feature with high stability (small variance) has more evaluation reference value, that is, the time variability feature is fused in the cross-modal attention calculation.
[0051] In some implementation scenarios, the variance of the voltage amplitude of all valid Epochs can be calculated for the same electrode and the same key time point, and the variance of the latency of all valid Epochs can be calculated for the same electrode. Similarly, the amplitude variance p300_h_amplitude_various and latency variance p300_h_latency_various of the respective p300 under the auditory and visual paradigms, and the amplitude variance N1_v_amplitude_various and latency variance N1_v_latency_various of the respective N1 can be calculated. Before cross-modal attention calculation, the normalized variance set and the electroencephalogram feature vector can be spliced, and then the cross-modal attention calculation operation is performed by the cross-modal fusion module.
[0052] The variance information (temporal variability feature) can distinguish the feature difference caused by the consciousness level from the feature fluctuation caused by the noise, and integrating such features into the cross-modal fusion can make the model preferentially rely on stable features for reasoning, suppress the interference of unstable features, and improve the reliability of the evaluation results.
[0053] In some embodiments, the above device can further perform the following operations: constructing a mapping relationship, wherein the mapping relationship contains the corresponding mapping between the stimulation mode, the neural pathway, the key brain area, and the feature type; and embedding the mapping relationship into the text feature to guide the cross-modal attention calculation. That is, by providing prior knowledge for cross-modal fusion, the model is guided to focus on core brain area features, solving the problem of cross-pathway fusion difficulty.
[0054] Specifically, based on the neuroscientific knowledge graph embedded in the multi-modal large language model, three types of core association rules can be extracted to form a structured mapping description. For example, the auditory stimulation paradigm corresponds to the auditory pathway and the bilateral temporal lobe (T3 / T4) brain area, and the P300 / N1 voltage distribution and Alpha wave energy features (corresponding to p300_h_amplitude, N1_h_amplitude, etc. under the auditory paradigm) need to be focused on. The visual stimulation paradigm corresponds to the visual pathway and the bilateral occipital lobe (O1 / O2) brain area, and the P300 / N1 voltage distribution and Beta wave energy features (corresponding to p300_v_amplitude, N1_v_amplitude, etc. under the visual paradigm) need to be focused on. The pain stimulation paradigm corresponds to the pain pathway and the bilateral parietal lobe (P3 / P4) brain area, and the N1 voltage distribution and Beta wave energy features need to be focused on.
[0055] Then, the mapping description is converted into natural language and embedded into the comprehensive text input (Prompt) to form a comprehensive text input (Prompt). The mapping rules are used as the guiding part of the Prompt, and the clinical information is input into the text encoder together. Based on this, the model can accurately focus on the key brain area during cross-modal fusion, avoiding indiscriminate processing of whole brain features, and improving the efficiency and relevance of fusion.
[0056] In some embodiments, the above device can further perform the following operations: based on the text feature embedded with the mapping relationship, adaptively adjusting the weight of the target spatiotemporal feature to optimize the spatiotemporal feature of the target event-related potential. That is, dynamically adjusting the importance weight of different spatiotemporal features to adapt to the non-standard response of brain injury patients.
[0057] As an example, based on the brain injury location in the clinical information, combined with the mapping relationship, the damaged neural pathway is determined, for example, left temporal lobe injury, the auditory pathway is determined to be damaged; right occipital lobe injury, the visual pathway is determined to be damaged. It is determined by the overlap degree of the brain injury location and the key brain area in the mapping relationship and the preset value (for example, 50%). For example, the left temporal lobe injury covers 60% of the key brain area of the auditory pathway, and it is determined that the auditory pathway is damaged. Then, the corresponding spatio-temporal feature weight of the damaged pathway is multiplied by the down-regulation coefficient, and the heavier the damage, the smaller the coefficient. For example, when the auditory pathway is damaged, the weights of p300_h_amplitude and N1_h_amplitude are down-regulated to 0.6.
[0058] In some implementation scenarios, the corresponding spatio-temporal feature weight of the pathway complementary to the damaged pathway function can also be multiplied by an up-regulation coefficient, and the heavier the damage, the larger the coefficient. For example, when the auditory pathway is moderately damaged, the weights of p300_v_amplitude and N1_v_amplitude in the visual pathway are up-regulated to 1.3. By down-regulating the weight of the damaged pathway and up-regulating the weight of the compensatory pathway, it is avoided to misjudge the feature anomaly caused by damage as low consciousness level, and at the same time, it is preferred to rely on high reliability features, which improves the accuracy and robustness of the evaluation result.
[0059] In some embodiments, the above device can further perform the following operations: calculating the correlation feature between the spatio-temporal feature and the variance information of the target event-related potential, and fusing the correlation feature with the electroencephalogram feature. The correlation feature refers to the internal correlation index between different dimension features (spatio-temporal feature, variance information), which strengthens the logical relationship between the features and solves the problem of insufficient discrimination of a single feature.
[0060] In some implementation scenarios, the correlation feature can include, for example, intra-component correlation, such as the correlation between the amplitude and latency of the same ERP component; cross-feature correlation, such as the correlation between the ERP amplitude and the corresponding frequency band energy value under the same pathway; cross-paradigm correlation, such as the amplitude consistency coefficient of the same ERP component under different stimulation paradigms; and variance correlation, such as the correlation between the ERP amplitude and the corresponding variance information. The foregoing correlations can be obtained based on correlation analysis calculation, and can be mapped to the range of 0~1 to construct a correlation feature set.
[0061] The correlation feature set and the electroencephalogram feature vector are spliced to form an extended feature vector, and the extended feature vector is input into a cross-modal fusion module to perform collaborative reasoning with the text feature vector to obtain the final consciousness evaluation result. Based on this, the correlation feature strengthens the internal logic between different dimension features, improves the discrimination and robustness of the features, and enables the model to capture more complex consciousness-related neural activity rules.
[0062] As described above, this embodiment of the application collects clinical information and task-oriented EEG signals under the target stimulus paradigm from the subject being evaluated, extracts frequency domain features of specific frequency bands and spatiotemporal features of P300 and N1 components, and constructs an EEG topography map. By using the image encoder and text encoder of a multimodal large language model to extract EEG features and text features respectively, and achieving consciousness assessment through cross-modal attention calculation fusion, it effectively solves the problems of strong subjectivity and insufficient fusion of EEG signals and clinical information in traditional behavioral scale assessments, and achieves automated and objective assessment of consciousness level.
[0063] Furthermore, the introduction of variance information enhances the judgment of feature stability, and the construction of mapping relationships between stimuli, neural pathways, brain regions, and features solves the problem of cross-pathway feature fusion. The adaptive adjustment of spatiotemporal feature weights adapts to the non-standard responses of brain injury patients, further improving the accuracy, robustness, and clinical suitability of the assessment results.
[0064] Figure 3 This is an exemplary schematic diagram illustrating an electroencephalogram (EEG) topography according to an embodiment of this application. For example... Figure 3 The image shows a brain topographic map corresponding to task-state EEG signals. It presents the voltage distribution characteristics of electrodes (e.g., labeled with E-series numbers) in different brain regions in a circular topological format, using color gradients to reflect the spatial differences in voltage distribution across different brain regions. In practical applications, energy distribution can also be fused and used as input to an image editor to extract EEG features.
[0065] Figure 4 This is an exemplary overall flowchart illustrating an assessment of level of consciousness according to an embodiment of this application. Figure 4 As shown, in steps S401 and S402, clinical information and task-oriented EEG signals under the target stimulus paradigm are collected from the subject of evaluation, respectively. Clinical information may include, for example, the location, degree, and cause of brain injury. The target stimulus paradigm may include, for example, auditory and visual stimulus paradigms. For the clinical information, in step S403, it is input into the text encoder of the multimodal large language model for text feature extraction, and in step S404, the text features are obtained.
[0066] For task-oriented EEG signals, in steps S405 and S406, frequency domain features of specific frequency bands and spatiotemporal features of target event-related potentials are extracted. For example, the energy distribution under alpha waves can be extracted. The amplitude, latency, and voltage distribution of P300 and N1 corresponding to visual and auditory signals are extracted respectively. In step S407, these are combined to form the corresponding EEG topography map. Next, in step S408, the data is input to the image encoder in the multimodal large language model for image feature extraction, and in step S409, the EEG features are obtained.
[0067] At step S411, the consciousness evaluation result is obtained by inputting the text features and the electroencephalogram features into the cross-modal fusion module in the multi-modal large language model for cross-modal attention calculation fusion at step S410.
[0068] According to the foregoing, variance information can also be calculated and fused into the cross-modal attention calculation, for example, at step S412. Alternatively, at step S413, the correlation features between the spatiotemporal features and the variance information are calculated to fuse the correlation features with the electroencephalogram features before cross-modal attention calculation fusion. In addition, at step S414, a mapping relationship can also be constructed and embedded in the text features to adaptively adjust the weights of the target spatiotemporal features in the cross-modal attention calculation fusion, thereby improving the accuracy of the consciousness level evaluation. For more details, reference can be made to the description made in the foregoing Figure 2
[0069] Figure 5 An exemplary structural block diagram of the electronic device 500 of the embodiments of the present application is shown. It can be understood that the device implementing the scheme of the present application can be a single device (e.g., a computing device) or a multifunctional device including various peripheral devices.
[0070] As shown in Figure 5 The electronic device of the present application can include a central processor or central processing unit (“CPU”) 511, which can be a general-purpose CPU, a special-purpose CPU, or other execution units for information processing and program running. Further, the electronic device 500 can also include a mass storage 512 and a read-only memory (“ROM”) 513, wherein the mass storage 512 can be configured to store various types of data, including various clinical information, task-state electroencephalogram signals, electroencephalogram features, text features, consciousness evaluation results, algorithm data, intermediate results, and various programs required for running the electronic device 500. The ROM 513 can be configured to store data and instructions required for power-on self-test, initialization of various functional modules in the system, driver programs for basic input / output of the system, and booting of the operating system of the electronic device 500.
[0071] Optionally, the electronic device 500 can further include other hardware platforms or components, such as a tensor processing unit (“TPU”) 514, a graphics processing unit (“GPU”) 515, a field programmable gate array (“FPGA”) 516, and a machine learning unit (“MLU”) 517 as shown. It can be understood that although a plurality of hardware platforms or components are shown in the electronic device 500, they are merely exemplary and not limiting, and a person skilled in the art can add or remove corresponding hardware according to actual needs. For example, the electronic device 500 can only include a CPU, related storage devices, and interface devices to implement the operations performed by the device for consciousness level assessment of the present application.
[0072] In some embodiments, in order to facilitate the transmission and interaction of data with external networks, the electronic device 500 of the present application further includes a communication interface 518, so that it can be connected to a local area network / wireless local area network (“LAN / WLAN”) 505 through the communication interface 518, and then connected to a local server 506 or connected to the Internet 507 through the LAN / WLAN. Alternatively or additionally, the electronic device 500 of the present application can also be directly connected to the Internet or a cellular network based on wireless communication technology through the communication interface 518, such as based on 3rd generation (“3G”), 4th generation (“4G”), or 5th generation (“5G”) wireless communication technology. In some application scenarios, the electronic device 500 of the present application can also access the servers 508 and databases 509 of external networks as needed in order to obtain various known algorithms, data, and modules, and can remotely store various data, such as various types of data or instructions for presenting example clinical information, task-state electroencephalogram signals, electroencephalogram features, text features, consciousness assessment results, etc.
[0073] The peripherals of the electronic device 500 can include a display device 502, an input device 503, and a data transmission interface 504. In one embodiment, the display device 502 can include, for example, one or more speakers and / or one or more visual displays configured to audibly and / or visually display operations performed by the device for consciousness level assessment. The input device 503 can include, for example, a keyboard, a mouse, a microphone, a gesture capture camera, and other input buttons or controls configured to receive input of audio data and / or user instructions. The data transmission interface 504 can include, for example, a serial interface, a parallel interface, or a universal serial bus interface (“USB”), a small computer system interface (“SCSI”), Serial ATA, FireWire, PCI Express, and a high-definition multimedia interface (“HDMI”), and the like, configured for data transmission and interaction with other devices or systems. According to the scheme of the present application, the data transmission interface 504 can receive the collected task-state electroencephalogram signals from the electroencephalogram device and transmit the task-state electroencephalogram signals or various other types of data or results to the electronic device 500.
[0074] The above CPU 511, mass storage 512, ROM 513, TPU 514, GPU 515, FPGA 516, MLU 517, and communication interface 518 of the electronic device 500 of the present application can be connected to each other through a bus 519, and data interaction is realized with the peripherals through the bus. In one embodiment, through the bus 519, the CPU 511 can control other hardware components and their peripherals in the electronic device 500.
[0075] The above description in combination with the accompanying drawings Figure 5 An electronic device that can be used to implement the present application is described. It needs to be understood that the device structure or architecture herein is only exemplary, and the implementation manner and implementation entity of the present application are not limited thereto, but changes can be made without departing from the spirit of the present application.
[0076] According to the above description in combination with the accompanying drawings, those skilled in the art can also understand that the embodiments of the present application can also be implemented by a software program. Therefore, the present application also provides a computer readable storage medium having computer readable instructions stored thereon for consciousness level assessment, which can be used to implement the present application in combination with the accompanying drawings Figure 1 、 Figure 2 The operations performed by the device for consciousness level assessment.
[0077] It should be noted that, although the operations of the methods of the present application are described in a particular order in the drawings, this is not meant to be limiting or insinuate that the operations must be performed in that particular order, or that all of the illustrated operations must be performed to achieve the desired result. On the contrary, the steps depicted in the flowcharts can be changed, eliminated, combined, and / or rearranged in other ways. Additionally or alternatively, some steps can be performed in parallel with one another.
[0078] It should be understood that the terms "first," "second," "third," and "fourth" and the like in the description and in the claims, if any, are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. The terms "including," "containing," "comprising," and the like are used herein to mean that the specified features, integers, steps, operations, elements, and / or components are included. It is further to be understood that the use of "including," "containing," "comprising," "having," and the like, does not exclude or remove the presence of any additional features, integers, steps, operations, elements, components, and / or groups thereof.
[0079] It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in this specification and the claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. It should be further understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It is to be understood that the terms "comprises" and "comprising" and the like are not intended to exclude the presence of one or more additional features, integers, steps, operations, elements, and / or groups described herein.
[0080] While the present application has been illustrated and described in detail in the drawings and foregoing description, such illustration and description is to be considered illustrative or exemplary and not restrictive; the present application is not limited to the disclosed embodiments. Numerous alternative modifications of the methods and apparatuses described herein will be apparent to those skilled in the art and it is intended to fall within the scope of the present application. Further, it should be understood that where the descriptions of the present application include or recite "a" or "an" element of the present application the descriptions of the present application also include and recite "one," "at least one," and "one or more" although the phrase "one or more" can not be expressly recited in the description of the present application. The appended claims are intended to cover all such alterations and modifications as are within the true spirit and scope of the present application.
Claims
1. A device for assessing level of consciousness, comprising: processor; as well as A memory storing computer instructions for assessing the level of consciousness, which, when executed by a processor, cause the following operations to be performed: Collect clinical information and task-oriented EEG signals under the target stimulus paradigm from the subjects to be evaluated; Based on the task-state EEG signals, frequency domain features of specific frequency bands and spatiotemporal features of target event-related potentials are extracted and combined to form a corresponding EEG topography map; The corresponding EEG topographic map is input into a multimodal large language model, and the image encoder in the multimodal large language model is used to extract image features to obtain EEG features. The clinical information is input into the multimodal large language model, and the text encoder in the multimodal large language model is used to extract text features to obtain text features; The cross-modal fusion module in the multimodal large language model is used to perform cross-modal attention calculation fusion on the EEG features and the text features to realize consciousness assessment and output consciousness assessment results.
2. The device according to claim 1, wherein the clinical information includes at least one or more of the following: brain injury location, degree of injury, disease etiology, and medical history; and the target stimulation paradigm includes one or more of the following: auditory stimulation paradigm, visual stimulation paradigm, or pain stimulation paradigm.
3. The apparatus of claim 1, wherein the apparatus further performs the following operations to extract frequency domain features of a specific frequency band: The energy value at a specific frequency band is calculated based on the task-state EEG signal to extract the frequency domain features of the specific frequency band.
4. The apparatus of claim 1, wherein the apparatus further performs the following operations to extract spatiotemporal features of the target event-related potential: Based on the task-state EEG signals, extract the effective signals corresponding to the target event-related potentials; Calculate the average voltage of the effective signal at key time points to obtain a voltage distribution map; The amplitude and latency are calculated based on the voltage distribution map to extract the spatiotemporal characteristics of the target event-related potentials, wherein the target event-related potentials include P300 and N1.
5. The apparatus of claim 4, wherein the apparatus further performs the following operations: The voltage distribution map is filtered using a target kernel function.
6. The apparatus of claim 1, wherein the apparatus further performs the following operations: Calculate the variance information of EEG topography at different time periods; The variance information is fused during cross-modal attention computation.
7. The apparatus according to claim 1 or 6, wherein the apparatus further performs the following operations: Construct a mapping relationship, wherein the mapping relationship includes a corresponding mapping between stimulation mode, neural pathway, key brain region and feature type; The mapping relationship is embedded in the text features to guide cross-modal attention computation.
8. The apparatus of claim 7, wherein the apparatus further performs the following operations: The weights of the target spatiotemporal features are adaptively adjusted based on the text features embedded in the mapping relationship to optimize the spatiotemporal features of the target event-related potential.
9. The apparatus of claim 6, wherein the apparatus further performs the following operations: Calculate the spatiotemporal characteristics of the target event-related potentials and the correlation characteristics between the variance information; The correlation features are fused with the EEG features.
10. A computer-readable storage medium having stored thereon computer program instructions for assessing level of consciousness, which, when executed by one or more processors, cause to perform the operations performed by the apparatus according to any one of claims 1-9.
Citation Information
Patent Citations
Construction method of multi-mode consciousness level recognition model based on EEG (electroencephalogram)
CN118335312A
Brain function evaluation system for patients with disturbance of consciousness based on neural multi-mode monitoring technology
CN119274789A
Cognitive state evaluation method in closed environment
CN119745319A
Multi-modal cognitive impairment evaluation system based on multi-dimensional cognitive function hypergraph
CN119818025A
Intelligent cognitive assessment method and system based on virtual reality environment
CN120130925A
Cited By
Slow wave detection method and device in electroencephalogram, program product and electronic equipment
CN122045708A
Methods, devices, software products and electronic equipment for slow wave detection in electroencephalography (EEG)
CN122045708B