Equipment and system for nervous system disease risk assessment

By using a multimodal fusion assessment model to extract features from eye-tracking and electroencephalogram (EEG) data and fuse cross-modal features, the problem of insufficient sensitivity and universality in the early diagnosis of neurological diseases is solved. This enables parallel diagnosis and risk assessment of multiple diseases, improving the accuracy and efficiency of diagnosis.

CN120977589APending Publication Date: 2025-11-18HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511486435.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing diagnostic methods for neurological diseases lack sensitivity and universality in the early stages. Traditional methods are costly and invasive, and single-modal diagnostic methods are not effective.

Method used

A multimodal fusion assessment model is adopted, which extracts features from eye-tracking data and EEG data and fuses cross-modal features through a dual-branch deep neural network. It generates cross-fusion features using an interactive attention mechanism and combines self-supervised learning and supervised fine-tuning to train the model and output risk assessment parameters for neurological diseases.

Benefits of technology

It improves the diagnostic sensitivity and applicability of neurological diseases, especially in the early stages such as mild cognitive impairment, with more significant accuracy. It can diagnose multiple diseases in parallel and generate objective risk assessment reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977589A_ABST
    Figure CN120977589A_ABST
Patent Text Reader

Abstract

The invention relates to a device for risk assessment of nervous system diseases, and a processor of the device is configured to receive eye movement data and electroencephalogram data which are associated in a time sequence, and fuse a double-branch deep neural network in a model through multiple modes to assess the risk of the nervous system diseases, performing feature extraction and attention mechanism-based cross-modal feature fusion on the eye movement data and the electroencephalogram data to obtain cross fusion features; and outputting risk assessment results and risk assessment parameters of at least one type of nervous system diseases through an output layer of the multi-modal fusion assessment model. According to the method and the device, the problem that diagnosis sensitivity and universality are insufficient due to the fact that diagnosis only depends on a single information source in related technologies is solved, disease features can be more comprehensively constructed from two levels of neural activity and behavior response by fusing data of two modes of eye movement and electroencephalogram, parallel diagnosis is further performed for multiple types of diseases, and the diagnosis efficiency is improved. The sensitivity and applicability of neuropsychiatric diagnosis are improved, and particularly, the method has more remarkable accuracy in early stages such as mild cognitive impairment and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and in particular to a device and system for risk assessment of nervous system diseases. BACKGROUND

[0002] Early diagnosis of nervous system diseases has always been a challenging problem in clinical practice. Common neurodegenerative or nervous system diseases include Parkinson's Disease (PD), Alzheimer's Disease (AD), and Tourette Syndrome / Tic Disorder.

[0003] In the related art, traditional diagnostic methods mainly rely on cognitive function evaluation scales, brain imaging examinations (such as MRI, PET), and cerebrospinal fluid biomarker detection. These methods have limited recognition ability in the early stages of the disease, and have problems such as high cost, some examinations are invasive, and equipment popularization is not high, which makes it difficult to meet the needs of large-scale early screening.

[0004] In addition, the method of diagnosing through machine learning algorithm in the related art mostly focuses on the solution for a single disease, or only relies on a single type of sensor for diagnosis. These methods often result in insufficient sensitivity and universality of diagnosis.

[0005] Therefore, in view of the problems of poor sensitivity and insufficient universality of the nervous system disease diagnosis method in the related art, an effective solution has not been proposed. SUMMARY

[0006] The embodiments of the present application provide a device and system for risk assessment of nervous system diseases to at least solve the problem of poor sensitivity of nervous system disease risk assessment method in the related art.

[0007] In a first aspect, the embodiments of the present application provide a device for risk assessment of nervous system diseases, which comprises: a memory configured to store a multi-modal fusion assessment model, and at least one processor in communication connection with the memory, the processor being configured to: receive timing-associated eye movement data and electroencephalogram data, extract features from the eye movement data and the electroencephalogram data and perform cross-modal feature fusion based on an attention mechanism through a double-branch deep neural network in the multi-modal fusion assessment model to obtain cross-fusion features; Based on the cross-fusion features, a risk assessment parameter of at least one type of nervous system disease is output through an output layer of the multi-modal fusion evaluation model.

[0008] In some embodiments, the eye movement data and the electroencephalogram data after preprocessing are subjected to feature extraction and cross-modal feature fusion based on an attention mechanism through a double-branch deep neural network in the multi-modal fusion evaluation model, including: The eye movement data and the electroencephalogram data after preprocessing are subjected to time windowing processing and time alignment to obtain eye movement time sequence features and electroencephalogram time sequence and frequency domain features; The eye movement time sequence features and the electroencephalogram time sequence and frequency domain features are subjected to intra-modal encoding to obtain eye movement embedding features and electroencephalogram embedding features; Through an interactive attention mechanism, the eye movement embedding features and the electroencephalogram embedding features are learned and guided to each other to generate enhanced features reflecting the correlation between eye movement and electroencephalogram, and the enhanced features are subjected to feature optimization, and through a fusion network, the cross-fusion features are obtained based on the results of the feature optimization.

[0009] In some embodiments, the enhanced features include eye movement enhanced features and electroencephalogram enhanced features, and through an interactive attention mechanism, the eye movement embedding features and the electroencephalogram embedding features are learned and guided to each other to generate enhanced features reflecting the correlation between eye movement and electroencephalogram, including: Through an interactive attention mechanism, based on the eye movement time sequence embedding features and the electroencephalogram time sequence embedding features, bidirectional guided enhancement and residual fusion are performed to obtain eye movement intermediate features and electroencephalogram intermediate features, respectively, The eye movement intermediate features and the electroencephalogram intermediate features are respectively subjected to self-attention refinement through a Transformer encoder to obtain eye movement enhanced features and electroencephalogram enhanced features, respectively.

[0010] In some embodiments, the enhanced features are subjected to feature optimization, and through a fusion network, the cross-fusion features are obtained based on the results of the feature optimization, including: The eye movement enhanced features and the electroencephalogram enhanced features are spliced to obtain combined features; Through a multi-head Fusion-Transformer structure, feature extraction and global pooling are performed based on the combined features to obtain the cross-fusion features.

[0011] In some embodiments, the training process of the multi-modal fusion evaluation model includes: In the pre-training stage, unlabelled eye movement data and electroencephalogram data are used to perform time augmentation on the eye movement data and the electroencephalogram data through self-supervised time sequence contrast learning, and to make eye movement time sequence embedding features and electroencephalogram time sequence embedding features in the same time window close to each other through cross-modal contrast learning, to obtain an initial model. In the supervised fine-tuning stage, the initial model is fine-tuned by using labelled eye movement data and electroencephalogram data, adopting a hierarchical unfreezing training strategy, and through class resampling and time sequence data augmentation, to obtain a trained multi-modal fusion evaluation model.

[0012] In some embodiments, the training process of the multi-modal fusion evaluation model further includes: In the case of freezing the main network parameters of the multi-modal fusion evaluation model, the parameters of a trainable adapter are trained and updated according to the eye movement data and the electroencephalogram data of a user to be evaluated, wherein the trainable adapter is embedded in the fusion layer of the double-branch deep learning network.

[0013] In some embodiments, the processor is further configured to: obtain baseline data of a user to be evaluated in a resting state, and eye movement data and electroencephalogram data in a preset task state; based on the baseline data, perform individual baseline normalization and Z-score standardization processing on the eye movement data and the electroencephalogram data in the preset task state, to obtain personalized eye movement data and personalized electroencephalogram data, and perform nervous system disease risk assessment based on the personalized eye movement data and the personalized electroencephalogram data through the multi-modal fusion evaluation model; determine whether statistical drift occurs according to the ratio of alpha waves to beta waves of the electroencephalogram data and the pupil baseline of the eye movement data, and if so, recalibrate the multi-modal fusion evaluation model.

[0014] In some embodiments, the processor is further configured to: perform denoising and image enhancement on the eye movement data to obtain an enhanced image, extract multiple eye movement indicators based on the enhanced image, compose the pre-processed eye movement data according to the eye movement indicators, and perform feature extraction and cross-modal feature fusion based on the pre-processed eye movement data; perform de-aliasing on the electroencephalogram data, obtain power values and event-related potentials in each preset frequency band based on the de-aliased electroencephalogram data, and compose the pre-processed electroencephalogram data according to the power values and the event-related potentials, and perform feature extraction and cross-modal feature fusion based on the pre-processed electroencephalogram data.

[0015] In some embodiments, outputting, by an output layer of the multi-modal fusion evaluation model, the risk evaluation parameter of the multiple types of nervous system diseases comprises: performing linear transformation on the cross-fusion feature through a full connection layer to obtain a linear transformation result, wherein the linear transformation result is used to reflect a risk prediction parameter of the cross-fusion feature with respect to each type of nervous system disease; outputting, by an activation function, the risk evaluation parameter of at least one type of nervous system disease based on the linear transformation result.

[0016] In a second aspect, the embodiments of the present application provide a nervous system disease risk evaluation system, which comprises: a head-mounted device configured to collect eye movement data of a user to be evaluated by using an inbuilt camera to collect eye images of the user to be evaluated; an EEG acquisition device arranged at different cerebral cortices and configured to collect multi-channel electroencephalogram data of the user while the head-mounted device collects the eye images; a device for nervous system disease risk evaluation as described in the first aspect, configured to receive the eye movement data collected by the head-mounted device and the electroencephalogram data collected by the EEG acquisition device, and output the risk evaluation parameter of at least one type of nervous system disease.

[0017] Compared with the related art, the device for nervous system disease risk evaluation provided by the embodiments of the present application comprises a memory configured to store a multi-modal fusion evaluation model, and at least one processor configured to: receive time-series correlated eye movement data and electroencephalogram data, perform feature extraction and cross-modal feature fusion based on an attention mechanism on the eye movement data and the electroencephalogram data through a double-branch deep neural network in the multi-modal fusion evaluation model to obtain cross-fusion features, and output the risk evaluation parameter of at least one type of nervous system disease through an output layer of the multi-modal fusion evaluation model based on the cross-fusion features. The present application solves the problem in the related art that only a single information source is relied on for diagnosis, resulting in insufficient sensitivity and universality of diagnosis. By fusing data of two modalities of eye movement and electroencephalogram, disease features can be more comprehensively constructed from two aspects of neural activity and behavioral response, and further parallel diagnosis can be performed for multiple types of diseases, thereby improving the sensitivity and applicability of neuropsychiatric diagnosis, and in particular, more significant accuracy is achieved in the early stage of mild cognitive impairment. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings, which are included to provide a further understanding of the present application, form a part of the present application and illustrate the illustrative embodiments of the present application and its description, which are used to explain the present application and do not constitute improper limitations on the present application. In the drawings: Figure 1is an application architecture diagram of a device for nervous system disease risk assessment according to an embodiment of the present application; Figure 2 is a running flowchart of a device for nervous system disease risk assessment according to an embodiment of the present application; Figure 3 is an application scenario diagram according to an embodiment of the present application; Figure 4 is a diagram of a multi-modal fusion evaluation model according to an embodiment of the present application; Figure 5 is a diagram of a nervous system disease risk assessment system according to an embodiment of the present application; Figure 6 is a diagram of a head-mounted device according to an embodiment of the present application. DETAILED DESCRIPTION

[0019] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application is described and explained below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of the present application.

[0020] Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application, and for those of ordinary skill in the art, the present application can be applied to other similar scenarios without creative efforts based on these drawings. In addition, it can be understood that although the efforts made in this development process can be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some designs, manufacturing or production changes based on the technical content disclosed in the present application are only routine technical means and should not be understood as insufficient disclosure of the present application.

[0021] In the present application, "embodiment" means that the specific features, structures or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present application. The phrase appears at various places in the specification does not necessarily all refer to the same embodiment, nor is it mutually exclusive or alternative to other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in the present application can be combined with other embodiments without conflict.

[0022] Unless otherwise defined, technical terms and scientific terms used in the present application shall have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terms "a", "an", "one", "this", and similar referents in the context of describing the application are to be construed to be open-ended, referring to one or more than one, unless otherwise noted. The terms "comprising", "comprises", "including", "includes" and "containing", "contains" shall be construed as containing the stated steps, elements or features but do not preclude the presence or addition of one or more other steps, elements or features. The term "connected" and / or "coupled" to be construed as possibly including a direct connection between two elements and / or an indirect connection between two elements through one or more intermediate elements. The term "plurality" means two or more. The term "and / or" includes all combinations of one or more of the associated listed items. The term "first", "second", "third", etc. are used to distinguish similar objects, not to denote a specific order.

[0023] In this document, it is to be understood that the terms involved can be technical means or other summary technical terms for implementing part of the present application, for example, the terms can include: AR glasses: refers to wearable augmented reality glasses, with built-in camera and display, used for augmented reality interaction. In the present application, it is used for shooting eye image; EEG (Electroencephalogram): electroencephalogram, a non-invasive technique for recording brain electrical signals through scalp electrodes, reflecting the activity of brain neuron groups; Eye tracking (ET): a technology for recording eye movement trajectory and pupil changes, used for analyzing attention and visual behavior; P300: an event-related potential (ERP) component, representing cognitive information processing, often showing reduced amplitude and prolonged latency in diseases such as depression; Self-attention mechanism: a kind of attention mechanism in deep learning, used to learn the association between different time points in time series data, used in the present application to fuse eye movement and EEG sequences; AUC (Area Under the Curve): Area under the curve, a common evaluation index in machine learning and medical statistics. The larger the AUC, the stronger the overall discriminant ability of the model at various thresholds. AUC can be understood as: the probability that a positive sample is ranked ahead of a negative sample by the model. For example, AUC = 0.85 means that in 85% of cases, the model will judge the positive sample to have a higher score than the negative sample. For example, judging whether EEG / eye movement signals indicate early Parkinson's disease → the higher the AUC, the better the diagnostic ability.

[0024] Early diagnosis of nervous system diseases is quite challenging. For example, in Alzheimer's disease (AD), traditional diagnosis relies on cognitive assessment, brain imaging, and biomarkers, which have problems such as difficulty in early identification, high cost, and low accessibility. With the aging of the population, the incidence of AD is rising, and there is an urgent need for convenient and non-invasive early screening methods. Recent studies have shown that electroencephalogram (EEG) and eye tracking technology have significant potential due to their non-invasive, low-cost, and high temporal resolution. AD patients have abnormal EEG rhythms, such as decreased alpha wave rhythm and increased theta / delta power, and abnormal eye movements, such as delayed saccadic response. Combining these two can improve the accuracy of early diagnosis. Parkinson's disease (PD) patients also have typical eye movement and EEG abnormalities. In terms of eye movement, blink frequency decreases, duration extends, and reflexive saccadic movement slows down, which can be used as early indicators before motor symptoms appear. In terms of EEG, patients with cognitive impairment in PD have increased theta wave relative power, decreased beta wave power, and decreased alpha / theta power ratio, which can be used to assess cognitive status. Depression diagnosis lacks objective markers and currently relies mainly on self-report scales and subjective diagnosis. Studies have found that patients with major depressive disorder have prolonged latency and reduced amplitude of ERPP300 components, and attention bias in emotional processing tasks. The dual-mode model of simultaneous collection of EEG and eye movement signals has significantly higher recognition accuracy for depression than single-mode models, such as 83.4% for mild depressive patients.

[0025] Tic disorders (such as Tourette syndrome) often start with simple motor tics in the head and face, and there is no recognized bioelectric indicator, so diagnosis mainly relies on clinical observation. In summary, most existing related technology research focuses on solutions for a single disease or relies solely on a single type of sensor (such as using only the camera of a mobile device to capture eye movement) for diagnosis. These methods have insufficient sensitivity and universality in diagnosis.

[0026] Therefore, the embodiments of the present application provide a device for nervous system disease risk assessment, Figure 1 is an application architecture diagram of a device for nervous system disease risk assessment according to the embodiments of the present application, like Figure 1As shown, the collection device 10 and the risk assessment device 11 communicate through a network. The collection device 10 can be, but is not limited to, a head-mounted device (such as augmented reality (AR) glasses) integrated with a camera and a multi-channel electroencephalogram (EEG) collection device. The user collects and uploads his / her eye movement data and EEG data to the risk assessment device 11 through the collection device 10. After processing by the risk assessment device 11, a comprehensive diagnostic report of the risk assessment parameters corresponding to at least one type of nervous system disease can be generated and visualized. The collection device 10 can be AR glasses, VR glasses, or other head-mounted devices and EEG information collection devices. The risk assessment device 11 can be any kind of operation processing device with storage and running of a lightweight model.

[0027] The risk assessment device 11 includes a memory configured to store a multi-modal fusion assessment model, and at least one processor in communication connection with the memory. The processor is configured to execute, for example, Figure 2 As shown in the flow steps, Figure 2 is a running flow diagram of a device for nervous system disease risk assessment according to an embodiment of the present application. As shown in the flow steps, Figure 2 The flow includes the following steps: S201, receiving time-correlated eye movement data and EEG data; The eye movement data and EEG data are collected by the user in a preset task state. Specifically, the preset task state refers to a test task state with specific requirements and controls set for the user when collecting eye movement and EEG data. The task state is designed to induce and record specific physiological responses related to nervous system diseases. In this embodiment, the design ensures the standardization and comparability of the data, thereby improving the accuracy of the diagnosis.

[0028] The head-mounted device preferably has a built-in black-and-white GlobalShutter camera with a field of view that can completely cover the subject's eyes, which is used to collect images of the wearer's eyes in real time and at a high frame rate to extract eye movement data such as pupil diameter, gaze point position, and blinking. The EEG collection device, for example, a portable non-invasive dry electrode cap, has electrodes covering the visual cortex (such as the occipital lobe), frontal lobe, parietal lobe, and other key brain regions related to visual processing, cognitive function, and emotional regulation, which are used to synchronously obtain electroencephalogram signals. The eye movement data and EEG data are synchronized by precise time stamps to ensure the accuracy of subsequent fusion analysis.

[0029] Further, after collecting the raw eye movement data and EEG data, the risk device performs a preprocessing step on the raw eye movement data and EEG data to purify and preliminarily extract features.

[0030] Specifically, the eye movement data is denoised and image enhanced, a plurality of eye movement indicators are extracted based on the enhanced image, and the preprocessed eye movement data is obtained according to the enhanced image and the eye movement indicators. For example, the pupil boundary is accurately extracted by the corneal reflection method or the convolutional neural network algorithm, the pupil diameter change curve, the gaze point sequence, the saccade speed, the number and frequency of blinking are calculated, and the preprocessed eye movement data is obtained. Exemplarily, for the eye movement modality data ET, the input original eye movement data (such as the eye video frame shot by a black and white global shutter sensor, the frame rate is 120-240 fps) is processed and extracted into a series of derived signals as output, which together constitute the preprocessed eye movement data, for example: pupil diameter sequence p(t) (obtained by elliptical fitting), gaze point trajectory g(t)=(x(t),y(t)) (mapped to the field of view coordinates), blinking binary state b(t)∈{0,1}, blinking duration, saccade speed / acceleration, fixation duration, and pupil corresponding curve.

[0031] At the same time, the original electroencephalogram data is filtered and de-noised, the power value and event-related potential in each preset frequency band are obtained based on the de-noised electroencephalogram data, and the preprocessed electroencephalogram data is obtained according to the power value and the event-related potential, for example: environmental and electromyographic artifacts are filtered out by a band-pass filter (0.5-40 Hz) and a power frequency trap, and eye movement, electromyographic interference is further removed by independent component analysis (ICA) or artifact subspace reconstruction (ASR), and then the power value of each preset frequency band (such as alpha, beta, theta wave) and the event-related potential (ERP, such as P300) are extracted, to obtain the preprocessed electroencephalogram data.

[0032] Exemplarily, for the electroencephalogram modality EEG, the preprocessed electroencephalogram data can be but is not limited to: time domain: band-passed original waveform segment, frequency domain: short-time power spectral density (Welch), theta, alpha, beta band energy and ratio (such as alpha / theta); event-related potential (ERP): P300 amplitude / latency is extracted under task event alignment.

[0033] In this embodiment, the pre-processing step is to improve the signal-to-noise ratio and extract basic features. By using digital signal processing and computer vision technology, the original electroencephalogram and eye movement signals are denoised, de-noised and characterized, and the complex original data stream is converted into a structured, higher information density feature sequence, providing clean, standardized input data for effective learning and accurate analysis of subsequent deep learning models.

[0034] S202, feature extraction and cross-modal feature fusion based on attention mechanism are performed on the eye movement data and the electroencephalogram data by a double-branch deep neural network in the multi-modal fusion evaluation model to obtain cross-fusion features; specifically, this step includes the following sub-steps: S2021, the eye movement data and the electroencephalogram data after preprocessing are subjected to time windowing processing and time alignment. For example, both signals are segmented into time windows with a length of 2.0 seconds and a step of 0.5 seconds, and strict alignment of the windows is ensured (with an error of less than 5 milliseconds) through system timestamps to obtain eye movement time sequence features and electroencephalogram time sequence and frequency domain features.

[0035] In this embodiment, time windowing and synchronous alignment ensure that the two heterogeneous signals of eye movement and electroencephalogram are matched in time. The original data are converted into uniform time sequence features that can be processed by the model.

[0036] S2022, the aligned eye movement time sequence features and the electroencephalogram time sequence and frequency domain features are subjected to intra-modal encoding respectively to obtain eye movement embedding features and electroencephalogram embedding features. Specifically, this embodiment is processed by two independent deep learning network branches for eye movement time sequence features and electroencephalogram time sequence features respectively.

[0037] For the eye movement (ET) encoder, the input is a multi-dimensional time sequence composed of pupil diameter, gaze point coordinates, blink state, etc., and the network structure adopts the design of a one-dimensional convolutional neural network (1D-CNN) followed by a Transformer encoder, which is used to capture the time sequence dependence within the eye movement sequence, and outputs eye movement embedding features.

[0038] For the electroencephalogram (EEG) encoder, the design in this embodiment is a time-frequency joint structure, one branch processes time domain waveforms, and the other branch processes frequency domain features (such as treating power spectrum as a pseudo-image and using a lightweight 2D-CNN to extract features), and the outputs of the two branches are fused to obtain electroencephalogram embedding features.

[0039] Specifically, the exemplary process of intra-modal encoding in this embodiment is as follows: 1. ET encoder (time sequence + numerical hybrid) Input: composed of [pupil diameter , gaze point trajectory , blink binary state , saccade velocity , saccade acceleration ], shape .

[0040] Structure: 1D-CNN (pyramid dilated convolution) → positional encoding → TransformerEncoder (2–4 layers, 4–8 heads) → layer normalization.

[0041] Specifically, convolutional neural networks (1D-CNNs) efficiently identify local, short-term behavioral patterns, such as a quick saccade or a long fixation, by scanning the entire eye movement time series through a sliding window.

[0042] Subsequently, the data enters the Transformer encoder. The core of the Transformer is the self-attention mechanism, which can analyze the correlation between any two time points in the entire sequence. This allows the model to not only see local actions, but also understand the long-range dependencies and importance of these actions in the entire task process. For example, in a complex search task, which gaze point is the most critical.

[0043] Output: Temporal embedding features With global aggregation vector ,in, Temporal embedded features are obtained through a pooling operation. All time-point information is aggregated (e.g., by averaging, taking the maximum value, or other more complex aggregation methods) to obtain a vector that represents the global information of the entire eye-tracking sequence, thus creating temporal embedding features. It is a local, temporal representation that retains detailed information at each time point, and is a globally aggregated vector. It is a global, aggregated representation that compresses the information of the entire sequence to capture macroscopic features.

[0044] 2. EEG encoder (Time-frequency combination) Branch A (Temporal Domain): Depthwise-Separable1D-CNN + Squeeze-and-Excitation → TransformerEncoder; The temporal branch is used to process the raw EEG waveforms. It uses an efficient deep separable 1D convolutional network (Depthwise-Separable1D-CNN) and an attention mechanism (Squeeze-and-Excitation) to process temporal signals, and uses a Transformer to capture the dynamic change patterns of EEG signals over time.

[0045] Branch B (frequency domain): Stack the power spectrum / band energy into pseudo-image, extract the spectral texture with lightweight 2D-CNN (or ConvNeXt-Tiny); the frequency domain branch is used to process the power spectrum of electroencephalogram or the energy features of different frequency bands (such as α, β, θ waves). The frequency domain data is stacked into a “pseudo-image”, and the texture features are extracted using a lightweight two-dimensional convolutional neural network (2D-CNN).

[0046] Fusion: ; global vector , shape is unified to dimension d (such as 128 / 256); finally, the features extracted by the time domain branch and the frequency domain branch are spliced and fused to generate time sequence embedding features representing the comprehensive information of electroencephalogram and global aggregation vector . It ensures that the analysis of brain activity neither ignores the instantaneous neural response nor loses the continuous state rhythm.

[0047] This step S2022 converts the complex time sequence information of the two signals of eye movement and electroencephalogram into higher-level and more abstract embedding features through two independent deep learning network branches. This intra-modal encoding method can capture unique patterns within each signal, providing high-quality and structured input for subsequent cross-modal fusion.

[0048] S2023, through the interaction attention mechanism, the eye movement embedding feature and the electroencephalogram embedding feature are learned and guided each other to generate enhanced features reflecting the correlation between eye movement and electroencephalogram, and finally obtain cross-fusion features. This step can be divided into a middle-layer interaction sub-step and a cascaded fusion sub-step, wherein: The middle-layer interaction sub-step uses the interaction attention mechanism (Co-Attention) for bidirectional guidance and enhancement. Specifically, the electroencephalogram embedding feature is taken as the query (Query), and the eye movement embedding feature is taken as the key (Key) and value (Value), to calculate the electroencephalogram enhanced feature guided by the eye movement feature, and perform the reverse to calculate the eye movement enhanced feature guided by the electroencephalogram feature. Then, through residual connection for fusion, and through Transformer encoder for self-attention refinement, the final eye movement enhanced feature and electroencephalogram enhanced feature are obtained.

[0049] The cascaded fusion sub-step splices the above two enhanced features in the time dimension or the channel dimension to obtain a combined feature, and inputs it into a multi-head fusion Transformer structure for deep feature extraction and global pooling, and finally obtains the cross-fusion feature.

[0050] Specifically, the exemplary process of cross-modal fusion in this embodiment is as follows: 1. Middle-layer interaction A. Co-Attention:

[0051] Let , Get EEG-guided ET representation; Do it again in reverse, ET-guided EEG representation.

[0052] B. Residual Fusion: ,

[0053] C. Self-Attention Refinement: Refine the cross-modal temporal representation through 1-2 layers of Transformer Encoder respectively.

[0054] It should be noted that the cascaded fusion sub-step allows the two modal information to be deeply and mutually confirmed through the attention mechanism, and finally generates a fused feature that is more discriminative than single information; among them, the middle layer interaction is used to guide and enhance each other, compared with simple data splicing, this step allows eye movement and EEG features to learn from each other and discover the potential relationship between the two, specifically, First, calculate the EEG-guided ET representation ( ): For each time point in the eye movement sequence, the model queries the entire EEG sequence and weights the most relevant EEG information into the current eye movement feature according to the correlation size. Similarly, do it once in reverse to calculate the ET-guided EEG representation ( ). Further, through a residual connection, the new features generated by the other party are added back to the original features respectively ( , ), to obtain the eye movement intermediate feature and the EEG intermediate feature respectively. It can be understood that in this embodiment, through this residual connection method, the original information is retained, and the supplementary information from the other modality is added on the basis of the original information, achieving information enhancement.

[0055] Subsequently, refine and each through a self-attention layer to further optimize the internal relationship of the sequence that has fused new information, and output the enhanced eye movement feature and the enhanced EEG feature.

[0056] 2. Cascaded fusion Concatenate and in the time dimension or the channel dimension, and then extract the final fused sequence through the Fusion-Transformer (multi-head 4-8, hidden 256-512) , and the cross-fusion feature zF is obtained through global pooling .

[0057] It should be noted that after the middle layer interaction, the two enhanced feature sequences ( and ) are spliced (Concat) in the time or channel dimension to form a wider and more comprehensive feature set; and the spliced feature is sent to the final fusion Transformer network for deep refining, extracting the most core cross-modal correlation information, and obtaining the cross-fusion feature zF for final classification and diagnosis through global pooling operation.

[0058] In step S2023, the embedded features of eye movement and electroencephalogram are bidirectionally learned and enhanced by using the interaction attention mechanism. Not only the correlation between the two modalities is deeply mined, but also the cross-fusion feature containing rich information is finally integrated through cascaded fusion, providing a comprehensive basis for the final disease diagnosis.

[0059] It should be noted that step S202 is the core step of the process, which solves the problem of how to effectively fuse heterogeneous time series data through a double-branch network structure and a cross-modal attention mechanism. This step not only extracts deep features from eye movement and electroencephalogram signals independently, but more importantly, it realizes deep interaction and complementarity of the two modalities at the feature level through interaction attention and cascaded fusion, thereby generating cross-fusion features that are richer and more discriminative than single-modal information.

[0060] S203, based on the cross-fusion feature, outputs at least one risk assessment parameter of a nervous system disease through the output layer of the multi-modal fusion evaluation model.

[0061] Among them, the cross-fusion feature output by step S202 is input into a classification head composed of one or more fully connected layers.

[0062] In an exemplary embodiment, the output layer is represented as: , for example, K=4 (PD / AD / depression / tic), where, is the cross-fusion feature vector containing all information of eye movement and electroencephalogram; the cross-fusion feature (zF) is input into a fully connected layer, which performs linear transformation on the cross-fusion feature vector, where, is a linear classifier used to convert a complex feature vector into a risk prediction parameter matching the number of diseases K. The linear classifier learns the mapping relationship between physiological signal features and different nervous system diseases (such as Parkinson's disease, Alzheimer's disease, etc.), and the result after linear transformation is an untreated raw risk score given by the model for each disease.

[0063] Further, the original risk scores do not directly show the risk assessment information of the disease, and need to be converted into standard probability values; sigmoid activation function, which independently compresses any original risk score into the interval [0, 1], outputs a risk assessment parameter, which represents the probability or confidence of suffering from each disease, and K represents the total number of diseases, for example, K = 4.

[0064] After outputting the risk assessment parameters for multiple nervous system diseases (such as Parkinson's disease, Alzheimer's disease, depression, etc.), these quantitative probability values are compared with preset disease diagnosis thresholds.

[0065] When the risk probability of any detected disease exceeds its corresponding dynamic threshold, an early warning mechanism is automatically triggered. In addition, whether it triggers a warning or as part of regular health monitoring, the device for nervous system disease risk assessment provided in the present application generates a comprehensive nervous system health diagnosis report reflecting the risk assessment results, which includes but is not limited to providing the following key information: Multimodal feature analysis results: list the key biomarker features that contribute to diagnosis, such as changes in alpha rhythm in electroencephalogram, abnormalities in P300 event-related potentials, or specific indicators such as saccade delay and reduced blink frequency in eye movement data; Disease risk assessment: clearly show the risk assessment probability for multiple diseases, and may include a comprehensive neurological health score; Historical data and trend monitoring: retain historical detection data to monitor changes in the user's neurological health status and provide a basis for long-term tracking.

[0066] The report will eventually be presented to the user or doctor, with the purpose of providing an objective, quantitative reference. The recommendations in the report will clearly indicate that the user should refer to them, or when the risk is high, explicitly prompt the need for further clinical examination for final diagnosis.

[0067] In this embodiment, step S203 corresponds to the decision output link, which converts highly abstract cross-fusion features into specific and interpretable clinical diagnosis recommendations. This step uses a multi-label classification strategy to achieve parallel and synchronous screening of multiple diseases, improving diagnosis efficiency and universality. In addition, by setting up an early warning mechanism and generating a comprehensive report, not only the final risk assessment is given, but also multi-dimensional feature analysis is provided, providing objective data support for users and doctors.

[0068] Through the above steps S201 to S203, the eye movement and electroencephalogram data of the patient in a specific state are first acquired. These raw data are carefully preprocessed, such as denoising and feature extraction, to lay the foundation for subsequent analysis. Further, a double-branch deep learning network is used to first independently encode the eye movement and electroencephalogram data, and then use cross-modal fusion technology based on attention mechanism to deeply fuse the two different types of features to obtain a unified feature vector containing the correlation information of both. Finally, based on the fused feature vector, through a multi-task output layer, multiple nervous system diseases are diagnosed in parallel, and the risk assessment parameters of each disease are output. The disease characteristics can be more comprehensively constructed from the two aspects of neural activity and behavioral response, and further parallel diagnosis can be performed for multiple diseases, thereby improving the sensitivity and applicability of neuro-psychiatric diagnosis.

[0069] Figure 3 is a schematic diagram of an application scenario according to an embodiment of the present application, as shown in Figure 3 , the system can be applied in an application environment as shown in Figure 3 , as shown in Figure 3 , the terminal and the processor communicate through the network, the user wears the terminal, collects eye movement data and electroencephalogram data, and uploads the user's own eye movement data and electroencephalogram data to the processor. After being processed by the processor, a comprehensive risk assessment report containing disease risk assessment suggestions can be generated, and can be visualized and output. Among them, the terminal can be but is not limited to various AR, VR and XR head-mounted devices, and the server can be implemented by an independent server or a server cluster composed of multiple servers.

[0070] Figure 4 is a schematic diagram of a multi-modal fusion evaluation model according to an embodiment of the present application, as shown in Figure 4 , the eye movement time sequence feature and the electroencephalogram time sequence feature are respectively subjected to eye movement feature encoding and electroencephalogram feature encoding. Further, the encoded features enter the middle layer interaction module, and through the interactive attention bidirectional guidance enhancement, cross-modal representations are obtained. Then these cross-modal representations are respectively refined by the Transformer encoder, and then subjected to time dimension or channel dimension splicing operation. Finally, the spliced features are input into the multi-head Fusion-Transformer structure to generate cross-fusion features. The cross-fusion features can more comprehensively constitute the diagnostic features from the two aspects of neural activity and behavioral response, thereby improving the accuracy and reliability of neuro-psychiatric disease diagnosis.

[0071] In some embodiments, the process of training the multi-modal fusion evaluation model is divided into two stages of pre-training and supervised fine-tuning.

[0072] In the pre-training stage, a large amount of unlabelled eye movement and electroencephalogram data are used to initialize the model parameters through self-supervised learning. Specifically, self-supervised temporal contrastive learning can be used to perform time augmentation (such as time masking and jittering) on the eye movement and electroencephalogram data, and cross-modal contrastive learning is used to make the eye movement embedding features and the electroencephalogram embedding features in the same time window close to each other in the feature space, so as to learn the internal correlation between the two modalities and obtain an initial model.

[0073] In the supervised fine-tuning stage, the pre-trained initial model is fine-tuned on a data set with clear disease labels. A hierarchical unfreezing training strategy is used to first train the classification head, then gradually unfreeze and fine-tune the fusion layer and the modality encoder, and finally use class resampling or temporal data augmentation techniques (Mixup / CutMix) to obtain a trained multi-modal fusion evaluation model.

[0074] In this embodiment, the pre-training + fine-tuning paradigm is used to learn general physiological signal representations using a large amount of unlabeled data, solving the problem of the scarcity of labeled data in the medical field. In addition, self-supervised and cross-modal contrastive learning enable the model to pre-master the deep correlation between eye movement and electroencephalogram, and the hierarchical fine-tuning strategy ensures stable and efficient convergence of the model on specific tasks, ultimately obtaining a high-precision and high-generalization diagnostic model.

[0075] In some embodiments, considering the problem of inaccurate prediction results caused by individual physiological feature differences, the device for neurological disease risk assessment is further configured with an individual adaptation strategy to improve the generalization ability and robustness of the model on different individuals. When the user uses it for the first time, the baseline data in the resting state is obtained, and the subsequent collected data is processed by individual baseline normalization and Z-score standardization based on the baseline data to eliminate the influence of individual physiological differences.

[0076] It should be noted that the physiological baseline values of different users are different (for example, the diameters of some people's pupils are naturally larger than those of other people), and if a unified standard is used for measurement, it is easy to produce misjudgment. The device for neurological disease risk assessment in this embodiment collects the eye movement and electroencephalogram data of the user in the resting state as the "baseline data" of the user. Subsequently, when performing a pre-set cognitive task (such as visual search, memory pairing, etc.), the eye movement and electroencephalogram data are collected again as "task state data".

[0077] Further, using baseline data as a reference, the data in the task state is individually baseline normalized and Z-score standardized. This process converts the original task data (such as the specific millimeter number of pupil diameter) into a relative value, which represents the degree of change of the task state relative to the user's own resting state. The processed data is called personalized eye movement data and personalized electroencephalogram data.

[0078] For example, the pupil of user A expands from 3mm in the resting state to 5mm in the task state, and user B expands from 5mm to 7mm. Although the absolute values are different, both of them expand by 2mm. Through normalization processing, the data analyzed by the model is converted from the absolute values of "5mm" and "7mm" to the relative value of "expanding by 2mm compared with the baseline". This kind of more comparable personalized feature enables the model to eliminate the physiological basis differences of different users and focus on finding the relative change pattern related to the disease, thereby greatly improving the accuracy and universality of diagnosis.

[0079] After the personalized processing of the eye movement and electroencephalogram data, the input to the pre-trained multi-modal fusion deep learning model, since the input data has eliminated most of the "noise" caused by individual differences, the model can more effectively learn and identify the biomarker features that are truly related to a specific disease (such as Alzheimer's disease, Parkinson's disease, etc.). Directly avoid the problem of "inaccurate diagnosis or performance degradation" caused by different physiological bases of users.

[0080] Secondly, a trainable adapter (Adapter) is embedded in the fusion layer of the double-branch deep learning network. When adapting for a new user, the parameters of the backbone network can be frozen, and only the parameters of the adapter are updated and trained according to a small amount of data of the user, realizing light and efficient individualization adaptation.

[0081] In addition, since the physiological state of the same user will fluctuate over time (for example, due to fatigue, emotion or environmental light changes), this may cause data "drift", affecting the long-term accuracy of the model. The device for risk assessment of nervous system diseases is also configured to continuously monitor the ratio of alpha wave and beta wave of electroencephalogram data and the pupil baseline of eye movement data, and when statistical drift is detected in the indicators, the recalibration of the model is automatically triggered.

[0082] It should be noted that the parameters learned by the model during training are based on the average level of a large number of people. Over time, the physiological state of the user may "drift" due to fatigue, emotion, medication or disease progression, which may cause the previously established diagnostic criteria of the model to be no longer accurate and applicable. Recalibration means fine-tuning the parameters of the model online according to the new changes, so that it can adapt to the current actual state of the user, thereby ensuring that the risk assessment result remains accurate and reliable over a long period of use by the user.

[0083] In this embodiment, the device for risk assessment of nervous system diseases ensures robustness and reliability in practical application through individual adaptation strategy, solves the problem of model performance decline caused by individual differences and physiological state changes through baseline calibration, lightweight parameter adaptation and statistical drift monitoring. This step enables the general model to quickly and cost-effectively adapt to new users, significantly improving the personalized accuracy and long-term stability of the system, and enhancing the practical application value of the product.

[0084] The following is described through a specific application process embodiment.

[0085] Step one: task and data collection. The subject wears AR glasses integrated with a camera and a synchronous EEG acquisition device. The system presents a series of cognitive tasks through the display interface of the AR glasses, such as visual search tasks. In the task, a target stimulus (such as the letter "T") and multiple interference items (such as the letter "L") are presented on the interface, and the subject is required to find the target within a specified time. During this period, the subject's eye movement trajectory, gaze duration, pupil diameter, blink frequency, and other eye movement data are continuously received, and their electroencephalogram signals (such as P300 waveforms) are simultaneously recorded.

[0086] Step two: data processing and fusion. The collected raw data is denoised and feature extracted. Subsequently, the aligned eye movement and electroencephalogram feature sequences are extracted by the feature encoder to extract high-dimensional embedding features of the two modalities, and then through the interaction attention and cascading fusion mechanism, cross-fusion features that can comprehensively reflect the user's neuro-behavioral state are generated.

[0087] Step three: diagnosis and output. The cross-fusion features are received and inferred through the trained multi-modal fusion evaluation model. The model outputs the risk probability of the subject suffering from various nervous system diseases (such as AD, PD, depression, etc.) in parallel. If a probability value exceeds a preset alarm threshold, a comprehensive diagnostic report will be generated, which may include specific abnormal feature indicators (such as "low alpha / theta power ratio", "P300 latency extension", etc.) and corresponding risk assessment, providing an objective basis for clinical diagnosis.

[0088] The device for risk assessment of nervous system diseases based on multi-modal information fusion provided in this embodiment can be applied in the following exemplary scenarios, but is not limited to: 1. Health screening for the elderly population: In community health or nursing homes, provide convenient cognitive and neurological function screening for the elderly. Wear AR glasses in a friendly and high-compliance manner to regularly detect early signs of Alzheimer's disease or mild cognitive impairment (MCI), enabling early warning and early intervention.

[0089] 2. Movement disorder diagnosis and treatment assistance: In the early detection of Parkinson's disease, it can be used in neurology clinics or home care. The abnormality of eye movement and brain electrical indicators monitored can be used as a reference for early diagnosis of PD to assist clinicians in making judgments.

[0090] 3. Mental health monitoring: In psychological counseling, psychiatry or school psychological assessment, the system can be used for screening and efficacy evaluation of depressive symptoms. Combined with traditional questionnaires, it can improve the identification rate of depressive mood disorders and provide individualized feedback.

[0091] 4. Daily remote monitoring: Taking advantage of the mobility and portability of AR glasses, the system can support continuous monitoring in home or mobile environments. Especially for individuals with ADHD / tic disorders or affected by work stress, background monitoring during daily work or study can achieve long-term non-invasive health tracking.

[0092] 5. Intelligent elderly care and rehabilitation scenarios: Integrating the system into smart elderly care glasses, brain-computer interfaces or VR / AR rehabilitation devices can be extended for user state monitoring. Task testing in virtual reality environments can improve interactivity and accuracy, providing technical support for smart healthcare and smart elderly care.

[0093] In a second aspect, the embodiments of the present application also provide a nervous system disease risk assessment system, Figure 5 is a schematic diagram of the nervous system disease risk assessment system according to the embodiments of the present application, as Figure 5 shown, the system comprises a head-mounted device 510, an EEG acquisition device 520 and a risk assessment device 530, wherein: The head-mounted device 510 acquires eye movement data of the user through the built-in camera; Wherein, the head-mounted device 510 can be an augmented reality (AR) glasses, in this embodiment, it acquires the eyeball image of the user through the built-in camera, and obtains real-time eye movement data.

[0094] In a preferred embodiment, in order to ensure the accuracy and speed of eye tracking, the built-in camera is a black and white GlobalShutter (GlobalShutter) camera. It can be understood that this camera can effectively avoid image smearing and distortion caused by rapid eye movement, and provide high-quality raw image data for subsequent pupil positioning and gaze tracking algorithms.

[0095] The EEG acquisition device 520 is used to synchronously acquire multi-channel electroencephalogram data of the user; Wherein, the EEG acquisition device 520 is used to synchronously acquire multi-channel electroencephalogram data of the user while the head-mounted device 510 acquires the eyeball image. It can be set to cover multiple electrodes of different brain cortexes, such as focusing on the occipital lobe, frontal lobe and parietal lobe regions related to vision, cognition and emotion.

[0096] In a preferred embodiment, in order to improve the convenience of use and the comfort of the user, the EEG acquisition device is a wearable non-invasive dry electrode cap. The dry electrode cap does not require conductive paste, is convenient to wear, is suitable for long-term continuous monitoring in a daily environment or a non-clinical environment, and realizes non-invasive and non-invasive diagnosis.

[0097] The risk assessment device 530 can be an external computer or a cloud server connected by wireless or wired means. The function of the device is to run the trained multi-modal fusion assessment model, receive the eye movement data from the head-mounted device 510 and the electroencephalogram data from the EEG acquisition device 520, execute the calculation logic described in the foregoing embodiments, finally perform parallel risk assessment on multiple nervous system diseases, and output the risk assessment result, such as risk score or comprehensive risk assessment report.

[0098] In a third aspect, the embodiments of the present application also provide a head-mounted device, Figure 6 is a schematic diagram of the head-mounted device according to the embodiments of the present application, as Figure 6 shown, the head-mounted device not only contains a built-in camera 610 for collecting user eye images to obtain eye movement data, but also integrates EEG electrodes 620 for collecting multi-channel electroencephalogram data at the corresponding positions of the head-wearing structure (such as the temple, headband or inner lining) of the head-mounted device. These electrodes are designed to closely fit different brain cortex regions of the user's scalp, so as to realize the synchronous collection of electroencephalogram signals while collecting eye images.

[0099] In addition, the head-mounted device includes a central processing unit (CPU) preloaded with a trained multi-modal fusion assessment model. It directly processes the collected eye movement data and electroencephalogram data on the device side, completes the whole process from data fusion to disease diagnosis, and finally outputs the diagnosis result.

[0100] The highly integrated design in this embodiment integrates data acquisition, processing and diagnosis functions into one, greatly improving the portability and ease of use of the device, making it an independent and complete neurological health monitoring terminal, especially suitable for long-term and continuous health tracking in personal home or mobile scenarios.

[0101] The technical features of the above embodiments can be combined in any way. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0102] The above embodiments only express several implementation ways of the present application, and the description is specific and detailed, but it should not be understood as a limitation to the patent scope of the application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, several modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A device for risk assessment of neurological diseases, characterized in that, The device used for risk assessment of neurological disorders includes: A memory configured to store a multimodal fusion evaluation model, and at least one processor communicatively connected to the memory, the processor being configured to: Receives time-series correlated eye-tracking and electroencephalogram (EEG) data; The eye-tracking data and EEG data are subjected to feature extraction and attention-based cross-modal feature fusion through the dual-branch deep neural network in the multimodal fusion evaluation model to obtain cross-fusion features. Based on the aforementioned cross-fusion features, the output layer of the multimodal fusion assessment model outputs risk assessment parameters for at least one type of neurological disease.

2. The device for risk assessment of neurological diseases according to claim 1, characterized in that, The multimodal fusion evaluation model utilizes a dual-branch deep neural network to perform feature extraction and attention-based cross-modal feature fusion on preprocessed eye-tracking and electroencephalogram (EEG) data, including: The eye-tracking data and the electroencephalogram (EEG) data are processed by time windowing and time alignment to obtain eye-tracking temporal features and EEG temporal and frequency domain features. Intramodal encoding is performed on the eye movement temporal features and the EEG temporal and frequency domain features to obtain eye movement embedding features and EEG embedding features; Through an interactive attention mechanism, the eye-tracking embedding features and the EEG embedding features learn from and guide each other to generate enhanced features that reflect the eye-tracking-EEG correlation. The enhanced features are then optimized, and the cross-fusion features are extracted and integrated based on the results of the feature optimization through a fusion network.

3. The device for risk assessment of neurological diseases according to claim 2, characterized in that, The enhancement features include eye-tracking enhancement features and electroencephalogram (EEG) enhancement features. Through an interactive attention mechanism, the eye-tracking embedding features and the EEG embedding features learn from and guide each other to generate enhancement features reflecting the eye-tracking-EEG correlation, including: Through an interactive attention mechanism, based on the eye-tracking temporal embedding features and the EEG temporal embedding features, bidirectional guided enhancement and residual fusion are performed to obtain intermediate eye-tracking features and intermediate EEG features, respectively. The eye-tracking intermediate features and the EEG intermediate features are refined by self-attention using a Transformer encoder to obtain eye-tracking enhancement features and EEG enhancement features, respectively.

4. The device for risk assessment of neurological diseases according to claim 2, characterized in that, The enhanced features are optimized, and the cross-fusion features are extracted and integrated based on the optimization results through a fusion network, including: The eye-tracking enhancement features and the electroencephalogram (EEG) enhancement features are concatenated to obtain combined features; The cross-fusion features are obtained by performing feature extraction and global pooling based on the combined features using a multi-head Fusion-Transformer structure.

5. The device for risk assessment of neurological diseases according to claim 1, characterized in that, The training process of the multimodal fusion evaluation model includes: In the pre-training phase, using unlabeled eye-tracking data and EEG data, the eye-tracking data and EEG data are temporally augmented through self-supervised temporal contrastive learning. Furthermore, through cross-modal contrastive learning, the eye-tracking temporal embedding features and EEG temporal embedding features within the same time window are made to be similar to each other, thus obtaining an initial model. During the supervised fine-tuning phase, labeled eye-tracking and EEG data are used, and a hierarchical unfreezing training strategy is employed. The initial model is then fine-tuned through category resampling and temporal data augmentation to obtain a trained multimodal fusion evaluation model.

6. The device for risk assessment of neurological diseases according to claim 5, characterized in that, The training process of the multimodal fusion evaluation model also includes: With the parameters of the main network of the multimodal fusion evaluation model frozen, the parameters of the trainable adapter are trained and updated based on the eye-tracking data and EEG data of the user to be evaluated, wherein the trainable adapter is embedded in the fusion layer of the dual-branch deep learning network.

7. The device for risk assessment of neurological diseases according to claim 1, characterized in that, The processor is also configured to: Acquire baseline data of the user to be evaluated in a resting state, as well as eye movement data and electroencephalogram data in a preset task state; Based on the baseline data, individual baseline normalization and Z-score standardization are performed on the eye movement data and EEG data under the preset task state to obtain personalized eye movement data and personalized EEG data. Then, based on the personalized eye movement data and personalized EEG data, a risk assessment of nervous system diseases is performed through the multimodal fusion assessment model. Based on the ratio of alpha and beta waves in the EEG data and the pupillary baseline in the eye movement data, it is determined whether statistical drift occurs. If so, the multimodal fusion evaluation model is recalibrated.

8. The device for risk assessment of neurological diseases according to claim 1, characterized in that, The processor is also configured to: The eye-tracking data is denoised and image-enhanced to obtain an enhanced image. Multiple eye-tracking metrics are extracted based on the enhanced image. The preprocessed eye-tracking data is then constructed based on the eye-tracking metrics. Feature extraction and cross-modal feature fusion are performed based on the preprocessed eye-tracking data. The EEG data is despoofed. Based on the despoofed EEG data, the power values ​​and event-related potentials in each preset frequency band are obtained. The preprocessed EEG data is then constructed based on the power values ​​and event-related potentials. Feature extraction and cross-modal feature fusion are then performed based on the preprocessed EEG data.

9. The device for risk assessment of neurological diseases according to claim 1, characterized in that, The output layer of the multimodal fusion assessment model outputs risk assessment parameters for various neurological diseases, including: A fully connected layer is used to perform a linear transformation on the cross-fusion features to obtain a linear transformation result, wherein the linear transformation result is used to reflect the risk prediction parameters of the cross-fusion features relative to various neurological diseases; Based on the linear transformation result, the activation function outputs risk assessment parameters for at least one type of neurological disease.

10. A system for risk assessment of neurological diseases, characterized in that, The system includes: Head-mounted display devices are used to capture images of the user's eyes through a built-in camera to obtain eye movement data; EEG acquisition devices are installed in different cerebral cortex areas to acquire multi-channel electroencephalogram (EEG) data of the user while the head-mounted display acquires eye images. The device for risk assessment of neurological diseases as described in any one of claims 1 to 9 is configured to receive eye-tracking data acquired by the head-mounted display device and electroencephalogram (EEG) data acquired by the EEG acquisition device, and output risk assessment parameters for at least one type of neurological disease.

Citation Information

Patent Citations

  • Multi-modal feature fusion emotion recognition method based on gating cross-attention mechanism

    CN117370828A

  • Disease risk prediction method and device, equipment and storage medium

    CN119207783A

  • Multi-modal fusion alertness assessment method and system based on cross attention

    CN120124011A

  • Intelligent system and method for early screening of depression based on electroencephalogram-eye movement multi-modal data fusion

    CN120227030A

Cited By

  • Method, medium and equipment for early warning risk of severity of illness state of enteritis patient

    CN121393907A

  • Intestinal inflammation patient severity risk early warning method, medium and device

    CN121393907B