Intelligent system and method for early screening of depression based on electroencephalogram-eyemovement multimodal data fusion
The intelligent system for early screening of depression, which integrates EEG and eye-tracking multimodal data, combines EEG and eye-tracking signals to design specific cognitive task paradigms. By utilizing virtual reality and brain-computer interface technologies, it solves the problems of insufficient accuracy and stability in early screening of mild depression, and achieves highly sensitive and efficient depression screening.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies for early screening of mild depression suffer from insufficient specificity of single-modal physiological signals, imperfect multimodal fusion methods, insensitivity to cognitive task stimulus paradigms, and a lack of integration with advanced technologies, resulting in insufficient accuracy and generalization ability in identifying mild depression.
An intelligent system for early screening of depression employs EEG-eye movement multimodal data fusion. By combining EEG and eye movement signals, a specific cognitive task paradigm is designed, and virtual reality and brain-computer interface technologies are utilized. Through multimodal data fusion algorithms and graph neural network models, a highly sensitive detection method for patients with mild depression is achieved.
It significantly improves the accuracy and stability of mild depression screening, provides a portable, real-time early screening platform for depression, supports big data analysis and clinical intervention for mental health, and promotes the application of intelligent mental health monitoring and intervention technologies.
Smart Images

Figure CN120227030B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent mental health screening, and particularly relates to an electroencephalogram-eyemovement multi-modal data fusion early screening intelligent system and method for depression. BACKGROUND
[0002] Depression is a common and severe mental illness that seriously affects the social function of patients, and its core symptoms include low mood and decreased interest. Depression has a high prevalence rate and high harm, and has become a global public health problem that needs to be addressed. Mild depression is in the early stage of depression, and if it can be detected and intervened early, it is of great significance to prevent the disease from worsening. However, the early symptoms of mild depression are often not obvious, and the rate of seeking medical treatment is low. Traditional diagnosis mainly relies on patient self-reports and clinical interviews, which can easily lead to missed diagnosis or misdiagnosis. Existing scale assessment methods (such as the PHQ-9 questionnaire) have subjectivity, and auxiliary means such as brain function imaging and voice analysis have limited specificity in detecting mild symptoms, resulting in insufficient timely attention and identification of mild depression.
[0003] Electroencephalogram (EEG) is a non-invasive neurophysiological signal that is very sensitive to small changes in brain activity and has a high temporal resolution of milliseconds, which can objectively reflect the brain function state. In the study of depression, some literature has reported that EEG signals such as alpha wave asymmetry and some event-related potentials (ERPs) are associated with depressive symptoms. However, relying solely on EEG single modality for depression detection faces the problem of large individual differences and insufficient specificity, and the differences in brain electrical characteristics of different subjects can lead to unstable classification results and poor generalization performance.
[0004] Eye tracking technology can record the gaze points, gaze duration, pupil diameter changes and other behavioral data of the subjects. These eye movement characteristics can intuitively reflect the individual's attention allocation and emotional response. For example, depressed individuals may have negative attention bias when viewing emotional stimuli, which is manifested as longer gaze on negative stimuli or more difficulty in moving their eyes away from negative information. Using eye movement data alone for depression recognition has made some progress, but it still has the problem of incomplete features and limited accuracy.
[0005] Depression recognition based on multi-modal data fusion has become a research hotspot in recent years. Electroencephalogram (EEG) and eye movement are two complementary physiological signal sources, which have great potential in depression detection: EEG provides internal neural activity information, and eye movement reflects external behavioral characteristics. The combination of the two is expected to more comprehensively characterize depression-related states. However, current research on EEG-eye movement fusion for depression recognition is still in its infancy. Existing methods mostly simply concatenate the features of the two signals and input them into a classifier, lacking in-depth modeling of the correlation and complementary relationship between multi-modal data. In addition, the cognitive task paradigms used in the past are relatively single, such as free browsing of emotional faces, static stimulus response, etc., which may not be sensitive enough to the subtle symptoms of mild depression. In terms of technical implementation, there is currently no solution to introduce advanced technologies such as virtual reality, brain-computer interface, remote sensing signal processing, etc. into this field to improve screening performance.
[0006] In summary, the existing technology has the following shortcomings in the early objective screening of mild depression: (1) limited single-modal physiological signal means, making it difficult to accurately identify mild depression; (2) multi-modal fusion methods are not perfect, and the correlation between different signals has not been fully explored, and the fusion strategy is simple, resulting in the need to improve the recognition accuracy and generalization ability; (3) there is a lack of carefully designed cognitive task stimulation paradigms for mild depression characteristics, and the existing test sensitivity is insufficient; (4) advanced technologies from other fields have not been organically integrated to achieve a new architecture and performance improvement for depression screening. To solve the above problems, it is necessary to provide a new intelligent screening system and method for depression in the early stage combining EEG and eye movement to improve the detection sensitivity and accuracy of mild depression. SUMMARY
[0007] Technical purpose: In view of the deficiencies of the prior art in the early objective screening of depression, the present application discloses an intelligent system and method for early screening of depression based on EEG-eye movement multi-modal data fusion, which can sensitively capture the subtle cognitive abnormalities of mild depression patients by fusing EEG and eye movement physiological signals and designing specific cognitive task paradigms, and achieve objective and efficient screening of early risk of depression.
[0008] Technical solution: To achieve the above technical purpose, the present application adopts the following technical solution:
[0009] An intelligent system for early screening of depression based on EEG-eye movement multi-modal data fusion, comprising:
[0010] An EEG signal acquisition module for acquiring EEG signals of a subject during the execution of a cognitive task and outputting multi-channel EEG data;
[0011] An eye movement tracking module for synchronously capturing eye movement information of the subject and generating eye movement data, including gaze point position, gaze duration and pupil diameter change;
[0012] a cognitive task presentation module configured to provide pre-designed cognitive tasks and stimulus scenarios to the subject to induce cognitive processes and emotional responses related to depression, the cognitive tasks including emotional stimulus browsing and cognitive function testing sessions;
[0013] a data processing and fusion analysis module configured to receive the EEG data and eye movement data, pre-process and extract features therefrom, and use a multi-modal data fusion algorithm to correlate and analyze the EEG features and eye movement features, and output a depression risk assessment result for the subject;
[0014] a result output module configured to present or store the assessment result in a human-machine readable form, including displaying the subject's depression tendency score, a discrimination conclusion and corresponding explanation information.
[0015] Preferably, the cognitive task presentation module includes a virtual reality display device, and the cognitive tasks are performed in a virtual reality scenario to simulate real-life scenarios and enhance the immersion of the subject in the cognitive tasks.
[0016] The cognitive task paradigm includes simultaneously presenting positive and negative emotional stimuli to detect the attention bias of the subject, measuring the behavioral and physiological responses of the subject in high cognitive load tasks, and observing the emotional responses of the subject in reward decision-making scenarios, thereby stimulating specific response patterns of the mild depression subject from multiple angles.
[0017] Preferably, the data processing and fusion analysis module includes:
[0018] a signal synchronization unit configured to time-align the EEG data and eye movement data according to a unified timestamp or a trigger signal;
[0019] a pre-processing unit configured to filter and remove artifacts from the EEG data, and smooth and calibrate the eye movement data;
[0020] a feature extraction unit configured to extract frequency domain features, event-related potential features, and brain network connection features from the EEG data, and extract gaze distribution, gaze sequence, and pupil change features from the eye movement data;
[0021] a multi-modal fusion unit configured to input the extracted EEG features and eye movement features into a multi-modal data fusion algorithm model, establish a correlation between the two and output a fusion feature representation;
[0022] a decision classification unit configured to perform pattern recognition or classification on the fusion feature representation to generate a depression risk assessment result.
[0023] Preferably, the multi-modal fusion unit constructs an association graph containing EEG feature nodes and eye movement feature nodes, and runs a graph neural network model on the graph to extract the association pattern between EEG and eye movement data, and the multi-modal fusion unit combines the EEG feature vector HE and the eye movement feature vector HO and their element-wise product using the following weighted fusion formula to construct the fusion feature:
[0024]
[0025] wherein, represents the feature vector extracted from the EEG signal, containing n feature components, represents the feature vector extracted from the eye movement data, containing m feature components, represents the element-wise product of the vectors, i.e. multiplying the corresponding elements in HE and HO to generate a vector of length min(n,m), representing the interaction between EEG and eye movement features, and a, b and g are weight coefficients, taking values in the range (0, 1), used to adjust the contribution proportion of EEG single mode, eye movement single mode and their interaction term to the final fusion feature.
[0026] Preferably, the EEG signal acquisition module is a wearable wireless device containing multiple dry electrode sensors and signal amplifiers, capable of transmitting EEG data in real time to the data processing and fusion analysis module through Bluetooth;
[0027] The eye movement tracking module includes a high-frame-rate infrared eye tracker or an eye movement sensor integrated into a head-mounted display, and a synchronous triggering mechanism is used to ensure millisecond-level alignment of the eye movement data and the EEG data.
[0028] An intelligent method for early screening of depression based on EEG-EO multi-modal data fusion, applied to an intelligent system for early screening of depression based on EEG-EO multi-modal data fusion as described above, comprising the following steps:
[0029] Through the EEG signal acquisition module and the eye movement tracking module, an eye movement calibration program is executed and baseline EEG signals and eye movement signals in a resting state are recorded;
[0030] A predetermined cognitive task sequence is presented to the subject through the cognitive task presentation module, which includes multiple links of emotional stimulus browsing and cognitive function testing, while key events are marked in real time;
[0031] EEG signals and eye movement signals are collected simultaneously while the subject is performing the cognitive task, and the two types of data are time-synchronized according to the event markers;
[0032] The collected electroencephalogram (EEG) signals are filtered, denoised and artifact processed, and the eye movement signals are smoothed and gazing event detected to obtain high-quality EEG data and eye movement data sequences;
[0033] A plurality of EEG feature indexes representing the neural activity of the subject are extracted from the pre-processed EEG data, and a plurality of eye movement feature indexes representing the gaze behavior and pupil response of the subject are extracted from the eye movement data;
[0034] The EEG features and eye movement features are input into a multi-modal fusion analysis model to fuse the information of the two modalities to identify depression-related patterns and output the depression risk assessment result of the subject;
[0035] The discrimination result and related information are presented to the user through a result output module or stored in a database, and subsequent processing suggestions are provided when needed.
[0036] Preferably, the multi-modal fusion analysis includes: constructing a correlation graph between the EEG features and the eye movement features, calculating the edge weight according to the synchronous activity of the EEG channel and the line of sight interest region, and extracting the cross-modal feature pattern using a graph neural network; or combining the EEG features, the eye movement features and their interaction terms to form a fusion feature vector through a formula fusion method, and inputting the fusion feature into a pre-trained machine learning classifier for depression state discrimination.
[0037] Preferably, the cognitive task presentation module presents the cognitive task using virtual reality technology, each task scene is displayed in an immersive three-dimensional environment, and the subject's perspective and operation behavior in the virtual environment are recorded; the task-induced effect is improved through rich multi-sensory stimulation, thereby amplifying the reaction differences of the mild depression subjects in the EEG and eye movement.
[0038] Preferably, the extracted EEG features include at least one of the following: power spectrum intensity of different frequency bands, amplitude and latency of event-related potentials, and functional connectivity network indicators;
[0039] The extracted eye movement features include at least one of the following: gaze time proportion to negative stimuli, transfer frequency of gaze points, and maximum expansion amplitude of pupil to stimuli; the psychological state of the subject is more comprehensively represented through the combination of multi-dimensional features.
[0040] Beneficial effects: The depression early screening intelligent system and method provided by the present application have the following beneficial effects:
[0041] 1. The present application realizes high sensitivity detection of mild depressive symptoms by adopting deep fusion of EEG and eye movement modal signals and design of a new type of cognitive task stimulation, captures subtle electrical activity changes in each region of the brain using the high time resolution of EEG, and records fixation behavior, gaze time and pupil dynamic changes through eye tracking, so that the system can extract multi-dimensional features such as specific frequency band power changes, ERP components and abnormal fixation bias. The proposed fusion unit adopts a multi-modal weighted fusion formula, which introduces an element-by-element product interaction term, explicitly models the coupling effect between EEG and eye movement features, thereby helping to capture the feature linkage effect of "abnormal eye movement brain response", and significantly improves the accuracy and stability of the screening results.
[0042] 2. The present application introduces virtual reality and brain-computer interface technology in system design, and constructs a highly immersive and multi-sensory stimulation cognitive task paradigm. The paradigm can stimulate the emotional and cognitive responses of subjects under different cognitive loads by displaying emotional stimuli, performing working memory or attention tasks, and participating in decision feedback and other multi-link tasks, further amplifying the neural and behavioral abnormalities related to mild depression; the virtual reality environment not only improves the realism of the task and the user's participation, but also realizes real-time closed-loop monitoring of data acquisition and task feedback through brain-computer interface, thereby ensuring high precision of data acquisition and flexible regulation of system response, providing strong technical support for early diagnosis.
[0043] 3. The present application adopts wireless wearable devices and high-frame-rate eye tracking instruments to construct a portable, real-time online monitoring intelligent screening platform; the system can quickly and non-invasively screen for early depression in a wide range of people through multi-modal data synchronous acquisition, signal preprocessing, feature extraction and deep fusion analysis. At the same time, the fine-grained physiological and behavioral data obtained by the platform provide objective basis for subsequent psychological health big data analysis, clinical precise intervention and self-regulation and rehabilitation training, and promote the wide application of intelligent psychological health monitoring and intervention technology in the field of public health. BRIEF DESCRIPTION OF DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, brief introduction will be given below to the drawings needed to be used in the embodiments or prior art description.
[0045] Figure 1 is the overall block diagram of the system of the present application;
[0046] Figure 2 is the method flowchart of the present application;
[0047] Figure 3 is the multi-modal signal preprocessing and feature extraction flowchart of the present application. DETAILED DESCRIPTION
[0048] The application will be described in more detail below by way of a preferred embodiment, but without intending to limit the application thereto, with reference to the accompanying drawings.
[0049] As shown in the drawings, an electroencephalogram-eyemovement multimodal data fusion early screening intelligent system for depression includes: Figure 1
[0050] An electroencephalogram signal acquisition module is configured to acquire electroencephalogram (EEG) signals of a subject during the execution of a cognitive task and output multi-channel EEG data.
[0051] The EEG signal acquisition module is configured to acquire the EEG signals of the subject in real time through a head-mounted EEG device. The EEG device can adopt a wearable wireless headgear, and a plurality of electrode channels are placed in the relevant cortical areas of the brain to record the EEG activities generated during the cognitive task.
[0052] An eye movement tracking module is configured to synchronously capture eye movement information of the subject and generate eye movement data, including gaze point position, gaze duration and pupil diameter change.
[0053] The eye movement tracking module is configured to record eye movement data such as the line-of-sight position and pupil change of the subject through an eye tracker or a camera. The eye movement tracking can adopt a high-precision infrared eye tracker or a tracking device integrated in a VR headset / smart glasses to ensure the synchronous acquisition of the eye movement information during the cognitive task.
[0054] A cognitive task presentation module is configured to provide a pre-designed cognitive task and stimulus scenario to the subject to induce depression-related cognitive processes and emotional responses, and the cognitive task includes an emotional stimulus browsing and a cognitive function test link.
[0055] The cognitive task presentation module is configured to present a customized cognitive task paradigm to stimulate specific cognitive processes and emotional responses. The task content is presented through a computer device screen or a VR virtual reality scene, and the task paradigm includes various stimulus scenarios and interactions, such as emotional picture free browsing, scenario memory / attention task, decision feedback task, etc., to comprehensively induce depression-related cognitive behavior characteristics such as attention bias, memory processing and decision response. Preferably, a virtual reality technology is adopted to create an immersive environment to improve the task participation and ecological effectiveness.
[0056] A data processing and fusion analysis module is configured to receive the EEG data and the eye movement data, pre-process and extract features therefrom, and use a multimodal data fusion algorithm to correlate and model the EEG features and the eye movement features and comprehensively analyze them to output a depression risk assessment result of the subject.
[0057] The data processing and fusion analysis module is used for synchronous processing, feature extraction and multi-modal fusion modeling of the collected electroencephalogram and eye movement data, and realizes depression risk assessment. The module includes a signal preprocessing unit (such as filtering, artifact removal, signal synchronization), a feature extraction unit (extracting frequency domain, time domain and brain network features from electroencephalogram, and extracting gaze duration, fixation sequence, pupil diameter change and other features from eye movement), and a fusion decision unit. The fusion decision unit is based on the new multi-modal data fusion algorithm architecture proposed in the application, which deeply fuses and models EEG and eye movement features, excavates the correlation pattern between the two, and finally gives the depression risk assessment result of the subject through the classifier or prediction model.
[0058] Specifically, the signal synchronization unit is used for time alignment of the electroencephalogram data and the eye movement data according to a unified timestamp or a trigger signal.
[0059] The preprocessing unit is used for filtering and artifact removal of the electroencephalogram data, and smoothing and calibration of the eye movement data.
[0060] The feature extraction unit is used for extracting frequency domain features, event-related potential features, brain network connection features from the electroencephalogram data, and extracting fixation distribution, gaze sequence, pupil change features from the eye movement data.
[0061] The multi-modal fusion unit is used for inputting the extracted electroencephalogram features and eye movement features into a multi-modal data fusion algorithm model, establishing the correlation between the two and outputting a fusion feature representation.
[0062] The decision classification unit is used for pattern recognition or classification of the fusion feature representation, and generates a depression risk assessment result.
[0063] The multi-modal fusion unit constructs an association graph containing electroencephalogram feature nodes and eye movement feature nodes, and runs a graph neural network model on the graph to extract the correlation pattern between the electroencephalogram and eye movement data. The multi-modal fusion unit adopts the following weighted fusion formula to combine the electroencephalogram feature vector HE and the eye movement feature vector HO and their element-wise product, thereby constructing a fusion feature:
[0064]
[0065] wherein, represents a feature vector extracted from the EEG signal, containing n feature components, represents a feature vector extracted from the eye movement data, containing m feature components, represents the element-wise product of vectors, that is, the multiplication of corresponding elements in HE and HO, generating a vector of length min(n, m), representing the interaction between electroencephalogram and eye movement features, and alpha, beta and gamma are weight coefficients, whose values are in the range of (0, 1), used to adjust the contribution proportion of electroencephalogram single mode, eye movement single mode and the interaction term to the final fusion features.
[0066] The result output module is configured to present or store the evaluation result in a human-computer readable form, including displaying the depression tendency score of the subject, the discrimination conclusion and the corresponding explanation information.
[0067] The result output and feedback module is configured to output the screening result and provide visual feedback or subsequent processing interface. The result can include a quantitative depression risk index, a judgment of whether there is mild depression, and a related feature prompt. The output form can be a mobile App interface display, a computer screen report, or uploading to a clinical system for reference by a doctor.
[0068] As shown in Figure 2 The present application also provides an electroencephalogram-eyeball movement multi-modal data fusion early screening intelligent method for depression, which is applied to the electroencephalogram-eyeball movement multi-modal data fusion early screening intelligent system for depression as described above, and includes the following steps:
[0069] S1. The electroencephalogram signal acquisition module and the eyeball movement tracking module are used to perform an eyeball movement calibration program and record baseline electroencephalogram (EEG) signals and eyeball movement signals in a resting state;
[0070] S2. The cognitive task presentation module is used to show a predetermined cognitive task sequence to the subject, and the cognitive task includes multiple links of emotion stimulus browsing and cognitive function testing, while real-time marking of key events is performed.
[0071] The subject is presented with a pre-designed cognitive task sequence in a computing device or a VR environment. The task paradigm includes multiple stages, such as: emotion stimulus browsing (showing positive and negative pictures or faces to induce attention bias and emotional response), working memory or attention tasks (such as N-Back test, Stroop color word task to induce cognitive load and reaction time change), decision and feedback tasks (such as a simple risk decision game, observing decision behavior under the influence of emotion). Each stage is accompanied by standardized instructions, and event markers in the task scenario are continuously recorded.
[0072] S3. The electroencephalogram (EEG) signals and the eyeball movement signals are synchronously collected while the subject is performing the cognitive task, and the two kinds of data are time-synchronized according to the event markers.
[0073] The EEG module and the eye movement module synchronously collect data in the whole process of the subject performing the task. Time alignment of the EEG and eye movement data is achieved through unified time stamps or trigger signals, ensuring the correspondence of the two modal data on each stimulation event. In this process, the application uses brain-computer interface technology to achieve efficient data transmission and real-time monitoring, ensuring the stability and reliability of data acquisition.
[0074] S4, filtering and denoising and artifact processing are performed on the collected EEG signals, and smoothing filtering and gaze event detection are performed on the eye movement signals, to obtain high-quality EEG data and eye movement data sequences;
[0075] The original EEG and eye movement data are preprocessed respectively. The EEG preprocessing includes filtering (such as 0.5-45Hz band-pass filtering to remove power frequency and electromyographic noise), eye movement artifact correction (such as removing the interference of blinking artifacts on EEG), segmentation and baseline correction (dividing EEG segments according to task events and deducting the baseline). The eye movement data preprocessing includes visual line coordinate filtering and smoothing, identification and removal of blinking / gaze null segments, and calculation of basic eye movement indicators (such as the number of fixation points per second, average gaze duration, etc.).
[0076] S5, multiple EEG feature indicators representing the neural activity of the subject are extracted from the preprocessed EEG data, and multiple eye movement feature indicators representing the fixation behavior and pupil response of the subject are extracted from the eye movement data;
[0077] The EEG features of the subject under each task condition are calculated, such as conventional frequency domain features (power spectrum intensity of each frequency band, such as alpha wave, beta wave power, etc.), time domain features (such as P300, N200, etc. Event-related potential amplitude and latency), nonlinear dynamics features (such as approximate entropy, Lyapunov exponent), and brain network features (functionally connected or synchronicity indicators such as coherence, Pearson correlation, etc. are calculated according to different brain region channels). These features comprehensively reflect the state differences of brain cognitive processing and emotional response.
[0078] The eye movement behavior features under the corresponding task situation are calculated, such as fixation distribution features (gaze time proportion in different interest regions, especially the difference in attention to emotional stimulus regions, used to measure negative or positive attention bias), fixation sequence and transfer mode (use scanning path to measure attention flexibility, such as single and slow fixation path of depressed individuals), pupil diameter change (pupil response caused by emotional or cognitive load, used to measure emotional arousal and concentration), etc. If necessary, multi-scale feature extraction methods in remote sensing image analysis can be combined to analyze the eye movement trajectory at multiple time scales to capture information on both short instantaneous fixation and long-term trends.
[0079] As Figure 3The multi-modal signal preprocessing and feature extraction flowchart is shown, which shows the whole data preprocessing process from the original EEG and eye movement data collection to signal filtering, artifact removal, data segmentation, and finally feature extraction. In the figure, the EEG signal and the eye movement signal are processed respectively, and then fused into a feature extraction node to further extract frequency domain features, time domain features, nonlinear features and behavior indicators. This process provides clean and structured feature data for the subsequent multi-modal fusion analysis module, which is the basic link of the whole system.
[0080] S6, inputting the EEG features and eye movement features into a multi-modal fusion analysis model, fusing information of the two modalities to identify depression-related patterns, and outputting a depression risk assessment result of the subject;
[0081] The extracted EEG features and eye movement features are input into the multi-modal fusion analysis model of the application for correlation modeling and depression state discrimination. Unlike the simple feature splicing method, the application designs a new multi-modal fusion algorithm architecture, including the following innovative steps:
[0082] Feature normalization and dimensionality reduction: first, the EEG and eye movement features are normalized (for example, Z-score standardization), and dimensionality reduction techniques (such as principal component analysis PCA or autoencoder) are used to reduce redundancy and extract key information subspace representation of each modality, denoted as H E (electroencephalogram feature vector) and H O (eye movement feature vector).
[0083] Correlation graph construction: drawing on the methods of brain network analysis and remote sensing multi-source data fusion, a correlation graph model is established between the EEG features and the eye movement features. Specifically, the EEG channels / features and the eye movement features are taken as nodes, and the edge connection is established according to their correlation degree in the same task event window, and the edge weight can be calculated by the statistical correlation or mutual information between the two signals. For example, for the i-th EEG feature and the j-th eye movement feature, the initial correlation weight is defined as: where E i (t) represents the value of the i-th component of the EEG feature at time t, O j (t) represents the value of the j-th component of the eye movement feature at time t, T is the total number of time frames of the task process, and ω(t) is a time weight function (used to highlight the contribution of a specific event time, which can be taken as ω(t) = 1 to represent equal weight). The product of the two modal signals at the same time is accumulated by the above formula to obtain the original correlation degree between the i-th and j-th features. Then, w ij is normalized to obtain the correlation matrix W, which is used as the adjacency matrix of the multi-modal graph. This graph intuitively represents the coupling strength between the EEG and eye movement features during the entire task.
[0084] Multimodal graph neural network fusion: The above association graph is input into a graph neural network model for feature fusion and extraction. Graph neural networks can extract deep cross-modal association patterns by iteratively propagating and aggregating node neighborhood information. For example, at each layer of the graph neural network, the feature representations of the EEG nodes and eye movement nodes will be updated according to the adjacency relationship is the feature vector of node v in the k-th layer of the graph neural network, and the update formula is:
[0085]
[0086] where N(v) represents the set of neighbor nodes of node v (determined by the association graph), d v , d u is the node degree for normalization, W (k) and V (k) are trainable weight matrices, and σ(·) is a nonlinear activation function. Through multi-layer propagation, the information of EEG and eye movement features is cross-fused in the graph structure, gradually forming a comprehensive representation that contains both EEG features and eye movement features.
[0087] Fusion feature representation and discrimination: After several layers of graph neural networks, the high-level representation vector of each node is obtained. All node representations can be aggregated (such as taking the average or maximum value) to form a global fusion feature vector, or the graph network can be connected to a readout layer to obtain a direct task discrimination output. In the preferred implementation of the present application, for the depression recognition task, the EEG- eye movement comprehensive feature vector after fusion is further input into a fully connected neural network or a support vector machine (SVM) classifier, and the output prediction result is output. For example, the classifier outputs the discrimination category of whether the subject is mildly depressed, or a depression risk probability value.
[0088] In addition, to improve the robustness and explainability of the model, a multi-evidence decision fusion mechanism is introduced in the algorithm architecture of the present application: sub-models are trained for different types of stimulus tasks to obtain discrimination results (such as emotion preference sub-models, cognitive performance sub-models, etc.), and then these sub-results are fused through weighted voting or a secondary classifier to obtain the final comprehensive judgment. This multi-evidence fusion method ensures the use of specific diagnostic information provided by each cognitive task to improve the overall recognition accuracy.
[0089] S7, the discrimination result and related information are presented to the user through the result output module or stored in the database, and subsequent processing suggestions are provided when needed.
[0090] The data processing module sends the analysis results to the result output module. The system generates a screening report for the subject, including a quantitative depression tendency score, a corresponding conclusion (such as "suspected mild depression" or "no obvious abnormalities found in screening"), and abnormal behavior / physiological indicators observed during the task process. The report can be displayed on the application interface for the user or further viewed by professionals. If used in conjunction with clinical practice, the report can be uploaded to a medical database to assist doctors in making diagnostic decisions.
[0091] Based on the screening results, the system can provide follow-up recommendations or guide the user to the corresponding mental health services. For example, when a depression tendency is detected, it is recommended to seek professional consultation or further medical diagnosis. The system of the present application can also extend to implement simple feedback functions due to its brain-computer interface capability, such as adjusting the stimulus intensity in real time during the task process (increasing the task interest according to the decrease in concentration monitored by EEG), or providing a relaxation training module to relieve the stress of the subject after screening. However, these are not core steps and can be implemented according to application requirements.
[0092] Embodiment 1
[0093] The depression early screening system provided in this embodiment includes a wearable EEG acquisition headgear, a desktop eye tracking camera, a virtual reality task presentation terminal, a backend data processing and analysis server, and a client result display device. The subject wears the EEG acquisition headgear, which is uniformly provided with a plurality of electrodes (such as 32 leads placed according to the international 10-20 system) for acquiring EEG signals of various regions of the brain. The desktop eye tracking camera is fixed below the display screen and is used to capture the pupils and corneal reflections of the subject's eyes in real time to calculate the gaze point. The virtual reality task presentation terminal consists of a high-performance computer and a VR head-mounted display, which is used to present a three-dimensional virtual task environment and record user behavior. The backend data processing and analysis server is connected to the EEG acquisition headgear and the desktop eye tracking camera through a wireless network, receives the collected data stream, and executes the method process of the present application. The final analysis results are sent to the client result display device (which can be a tablet computer or a mobile phone, etc.) through the network, and are displayed to the user or doctor in the form of a report in a special application App.
[0094] In this embodiment, the EEG acquisition headgear uses dry electrode technology to improve wear comfort, and the built-in amplifier and Bluetooth module realize wireless transmission of signals; the desktop eye tracking camera is an infrared eye tracker with a sampling rate of 120Hz, which is placed about 60cm away from the user; the VR head-mounted display provides an immersive task scene, including sound and tactile feedback and other multi-channel stimuli. The overall architecture of the system ensures the synchronization of data acquisition and good balance of user experience, and can be applied to hospital clinics, school psychology laboratories, etc. to screen and evaluate subjects for mild depression.
[0095] Example 2
[0096] This embodiment details the procedure of the early screening method for depression according to the present application, and each step is as follows.
[0097] Step S1: The subject connects the device and calibrates. The subject first wears the EEG brain electrical acquisition headgear and sits in front of the table-mounted eye tracking camera (or wears a VR head-mounted display instead of the camera, depending on the implementation environment). Start the system software, check the connection of each channel lead of the brain electrical, adjust the electrode impedance to the appropriate range (such as <10kΩ). Start the eye tracking calibration program, ask the subject to fix his gaze on the calibration points appearing on the screen in turn, calculate the accurate eye movement mapping parameters to ensure the accuracy of the subsequent eye movement data. After completing the calibration, record the brain electrical in 1 minute of resting state and the eye movement data in the natural gaze state as the individual baseline.
[0098] Step S2: Task paradigm presentation. After calibration, the system enters the formal test, and according to the designed cognitive task paradigm, the system presents each task stimulus step by step:
[0099] First, perform the emotion picture free-viewing task: simultaneously display multiple face pictures containing different emotions (positive / neutral / negative) in the VR scene or on the screen. Instruct the subject to freely view for a certain time (for example, 30 seconds). In this process, the system records the fixation sequence and duration distribution. It is expected that depressed individuals may tend to fixate on negative expressions or pay less attention to positive expressions, thus reflecting negative attention bias.
[0100] Next, perform the Stroop color word test (or similar attention control task): continuously present several color words, and the word meaning and font color are consistent or inconsistent, and the subject is required to ignore the word meaning and say the font color. This stage examines the cognitive control ability and reaction time, and depressed individuals may show slower reaction and higher error rate, while the brain electrical may show reduced P300 amplitude in the frontal region, etc.
[0101] Then, perform the working memory task (such as N-Back): randomly present letters or numbers, and the subject needs to judge whether the current item is the same as the item N steps ago. This stage induces memory load, and observes the change of brain electrical theta wave power and the degree of eye movement concentration (such as whether to frequently get distracted and shift the gaze) of depressed individuals under high load.
[0102] Finally, perform the decision feedback task: for example, play a simple virtual money reward and punishment game, the subject selects from several options and gets immediate reward or punishment feedback. This stage aims to induce differences in reward processing of depressed individuals (such as lack of pleasure response to rewards, and reduced EEG reward-related positive waves, etc.).
[0103] The whole task flow is automatically controlled and promoted by the system, and a short break can be set between stages. Key events (stimulus appearance, subject response) send trigger signals to the EEG record to achieve precise alignment of EEG and task scenarios. In the VR environment, the subject's operation (such as triggering selection by gazing at an object for more than a threshold time) is also recorded to replace the traditional key response.
[0104] Step S3: Multimodal data synchronous acquisition. During step S2, the EEG headgear and the desktop eye tracking camera continue to work. Through the system software, the EEG data stream and the eye movement data stream are implemented on a unified time axis to ensure that the two kinds of data are aligned on the same event reference. For example, when a picture appears, the eye movement records the area where the subject first gazes and the pupil changes, and the EEG records the visual evoked potential, which are associated through a common event tag. In the case of using a VR headset, eye tracking is completed by infrared sensors inside the headset, and EEG is collected by the headset, and the two are synchronized through the built-in synchronization module of the VR system. During the whole acquisition process, BCI technology is used to monitor data quality: if electrode shedding or signal saturation is detected, the system will prompt adjustment in time; if the subject closes his eyes for too long or is absent-minded (judged according to the lack of eye movement focus and the significant increase of EEG alpha wave), the system can pause the task to ensure data validity.
[0105] Step S4: Data preprocessing and feature extraction. After the task is completed, the backend data processing and analysis server preprocesses and extracts features from the acquired raw data. In this embodiment, MATLAB and Python hybrid programming is used to realize signal processing, in which EEG preprocessing uses band-pass filtering (1-40Hz) and independent component analysis (ICA) to remove eye movement artifacts, and eye movement data is identified by a self-defined algorithm to identify gaze and saccade sequences. In terms of feature extraction, some task-specific features are additionally added: for example, in the emotion browsing task, the total gaze time of the subject on each type of emotional face and the first gaze latency are calculated as behavior indicators; in the decision-making task, the EEG reward positive component (such as a positive wave with a 200-300ms latency) amplitude after the appearance of the reward stimulus is extracted. These features, together with conventional frequency features, constitute the input for subsequent analysis.
[0106] Step S5: Multimodal fusion analysis and discrimination. The system performs fusion analysis on the extracted EEG and eye movement features. In this embodiment, we use a weighted interactive fusion method to construct the fusion feature H fusion . Specifically, the EEG feature vector H E is set as a vector of dimension n = 50 (obtained by automatic encoder compression to obtain a 50-dimensional representation), and the eye movement feature vector H OThe vector m = 20 is set, and then the shorter vector is extended to the same length (in this case, the eye movement vector is padded with zeros to 50 dimensions to correspond to element-wise multiplication with the EEG vector). Take the initial weight α = β = 0.5, γ = 0.5, and substitute it into the formula to calculate the fusion feature. The obtained Η fusion (50 dimensions) is then input into a two-layer fully connected neural network classifier, the hidden layer uses ReLU activation, and the output layer is Softmax two classification (depression / non-depression). The classifier model has been trained in advance with a large amount of existing labeled data and has good performance. In actual screening, the classifier outputs a prediction score for the test subject, and the classifier outputs a prediction score for the test subject. For example, if the output is [0.82, 0.18], it means that the model considers the probability of the "depression" category to be 82%, and accordingly judges that the test subject may belong to the high risk of mild depression. In addition to the neural network, this embodiment also tries a support vector machine (SVM) classifier, and comparison shows that the accuracy of the two is similar, but the neural network has better scalability. fusion The prediction score is given, for example, the output is [0.82, 0.18], which means that the model considers the probability of the "depression" category to be 82%, and accordingly judges that the test subject may belong to the high risk of mild depression. In addition to the neural network, this embodiment also tries a support vector machine (SVM) classifier, and comparison shows that the accuracy of the two is similar, but the neural network has better scalability.
[0107] It should be noted that the weights α, β and γ can be automatically adjusted to the best value in the model training process. In our offline training, the final learned α ≈ 0.4, β ≈ 0.3, γ ≈ 0.8, which means that the interaction term has a large weight in the decision, which verifies the importance of the EEG- eye movement feature interaction for recognizing depression. In addition, if the graph neural network fusion path is used, only the preprocessed features need to be imported into the constructed heterogeneous graph, and the GNN model as described above is used for training and prediction. Both schemes are within the protection scope of the present application.
[0108] Step S6: result output and explanation. Once the discrimination model obtains the result, the system generates a screening report. In this embodiment, the result report of a certain test subject shows that its depression risk score is 0.82 (threshold value 0.5 or more), combined with its task performance characteristics, the system judges that "there is a possibility of mild depression". The report lists some objective indicators supporting this conclusion in detail: for example, "the test subject's gaze time ratio on negative expressions in the emotional face browsing task is as high as 45% (the normal mean is about 30%), showing a significant negative attention bias; at the same time, its EEG alpha wave asymmetry index is 0.2 (which should be negative for normal), indicating that the left frontal lobe is relatively inactive, which is consistent with the depression pattern". These explainable feedbacks are helpful for users or clinicians to understand the screening results. For individuals with positive screening, the system will suggest them to seek professional diagnosis as soon as possible, and can provide the contact information of local psychological counseling agencies. In subsequent services, if the user agrees, the system can upload the multi-modal raw data and results of this screening to the cloud database anonymously, accumulating data for objective diagnosis of depression.
[0109] The above examples mainly focus on the test verification of the college student group. The results show that the system has a discrimination accuracy of more than 90% for students with mild depression determined by psychological questionnaire screening, and a specificity of more than 85% for non-depression healthy controls. The whole screening process takes about 20 minutes, and the user compliance is good.
[0110] The above only describes the preferred embodiments of the present application, and it should be pointed out that those skilled in the art can make some improvements and refinements without departing from the principles of the present application, and these improvements and refinements should also be considered as the protection scope of the present application.
Claims
1. An intelligent system for early screening of depression based on EEG-eye-tracking multimodal data fusion, characterized in that, include: The EEG signal acquisition module is used to acquire the subject's EEG signals during cognitive tasks and output multi-channel EEG data. The eye-tracking module is used to synchronously capture the subject's eye movement information and generate eye movement data, including fixation point position, gaze duration, and pupil diameter changes. The cognitive task presentation module is used to provide subjects with pre-designed cognitive tasks and stimulus scenarios to induce cognitive processes and emotional responses related to depression. The cognitive tasks include browsing of emotional stimuli and cognitive function testing. The data processing and fusion analysis module is used to receive the EEG data and eye movement data, preprocess them, extract features, and use a multimodal data fusion algorithm to perform correlation modeling and comprehensive analysis of EEG features and eye movement features, and output the depression risk assessment results of the subjects. The data processing and fusion analysis module includes: The signal synchronization unit is used to time-align EEG data and eye-tracking data according to a unified timestamp or trigger signal. The preprocessing unit is used to filter and remove artifacts from EEG data, and to smooth and calibrate eye-tracking data. The feature extraction unit is used to extract frequency domain features, event-related potential features, and brain network connectivity features from EEG data, as well as gaze distribution, gaze sequence, and pupil change features from eye movement data. The multimodal fusion unit is used to input the extracted EEG features and eye-tracking features into the multimodal data fusion algorithm model, establish the correlation between the two, and output the fused feature representation. The multimodal fusion unit constructs an association graph containing EEG feature nodes and eye-tracking feature nodes, and runs a graph neural network model on this graph to extract the association patterns between EEG and eye-tracking data. The multimodal fusion unit uses the following weighted fusion formula to combine EEG feature vectors... With eye movement feature vector Combining this with its element-wise product, a fusion feature is constructed: ; in, This represents the feature vector extracted from the EEG signal, containing n feature components. This represents a feature vector extracted from eye-tracking data, containing m feature components. This represents the element-wise product of vectors, i.e., the product of vectors. and The corresponding elements are multiplied to generate a vector of length min(n,m), which represents the interaction between EEG and eye movement features. α, β and γ are weight coefficients with values in the range (0,1), which are used to adjust the contribution ratio of EEG monomodality, eye movement monomodality and their interaction to the final fused feature. A decision classification unit is used to perform pattern recognition or classification on the fused feature representation to generate a depression risk assessment result. The results output module is used to present or store the assessment results in a human-computer readable form, including displaying the subject's depression tendency score, judgment conclusion, and corresponding interpretation information.
2. The intelligent system for early screening of depression based on EEG-eye-tracking multimodal data fusion according to claim 1, characterized in that, The cognitive task presentation module includes a virtual reality display device, in which the cognitive task is performed in a virtual reality scene to simulate a real-world scenario and enhance the immersion of the cognitive task for the subject. The cognitive task paradigm includes: simultaneously presenting positive and negative emotional stimuli to detect the subject's attentional bias, measuring their behavioral and physiological responses in high cognitive load tasks, and observing their emotional responses in reward decision-making scenarios, thereby stimulating specific response patterns in mildly depressed subjects from multiple perspectives.
3. The intelligent system for early screening of depression based on EEG-eye-tracking multimodal data fusion according to claim 1, characterized in that, The EEG signal acquisition module is a wearable wireless device containing multiple dry electrode sensors and a signal amplifier, which can transmit EEG data to the data processing and fusion analysis module in real time via Bluetooth. The eye-tracking module includes a high-frame-rate infrared eye tracker or an eye-tracking sensor integrated into a head-mounted display, and ensures millisecond-level alignment of eye-tracking data with EEG data through a synchronous triggering mechanism.
Citation Information
Patent Citations
Depression identification method and system based on multi-modal data fusion model
CN116010901A