Mild cognitive impairment auxiliary recognition system and method based on eye movement characteristics

By designing a multimodal eye movement feature fusion network of emotion recognition and memory paradigm, combining static visual attention heatmap and dynamic eye movement sequences, the problem of MCI misdiagnosis in traditional methods is solved, and high-accuracy AD and MCI diagnosis is achieved, providing a more comprehensive cognitive function evaluation.

CN120241067APending Publication Date: 2025-07-04SHANDONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510367022.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Prior Art In the early diagnosis of Alzheimer's disease, traditional methods rely on expensive and invasive biomarker detection and neuropsychological testing, resulting in misdiagnosis or underestimation of MCI, and lack of multi-dimensional cognitive evaluation. Single-dimensional eye movement research fails to fully capture dynamic changes in the cognitive process.

Method used

The emotion recognition and memory paradigm is designed, and the multimodal fusion network combining static visual attention heatmaps and dynamic eye movement sequences is extracted, and the temporal and spatial characteristics of eye movement are enhanced. The MobileNet V2 and CBAM attention modules are used to enhance feature extraction, and the timing features are captured using 1D-CNN, and the multimodal fusion classification module is used to achieve accurate recognition of AD, MCI and health status.

Benefits of technology

It improves the diagnostic accuracy of AD and MCI, achieves 92.54% AD diagnosis and 90.64% MCI diagnosis, provides a more comprehensive cognitive function assessment, and enhances the ability to identify cognitive impairments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120241067A_ABST
    Figure CN120241067A_ABST
Patent Text Reader

Abstract

The invention discloses a mild cognitive impairment auxiliary recognition system and method based on eye movement characteristics, and relates to the technical field of computer-aided diagnosis, and the system comprises a data obtaining module which is used for obtaining eye movement data of a patient in the emotion classification task process; the data preprocessing module is used for performing data preprocessing on the acquired eye movement data; the auxiliary recognition module is used for inputting the preprocessed eye movement data into an auxiliary recognition model based on a double-flow fusion network, the auxiliary recognition model comprises a spatial feature extraction module, a time sequence feature extraction module and a multi-modal fusion classification module, key features of the eye movement hotspot map are extracted through the spatial feature extraction module, and the key features of the eye movement hotspot map are extracted through the time sequence feature extraction module; the time sequence features of the dynamic eye movement sequence are extracted through the time sequence feature extraction module, the features extracted by the two feature extraction modules are spliced and fused through the multi-modal fusion and classification module, and final accurate classification and recognition results of the Alzheimer's disease, the mild cognitive impairment and the health are output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer-aided diagnosis, and particularly to an auxiliary recognition system and method for mild cognitive impairment based on eye movement characteristics. Background Technique

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Alzheimer's disease (AD) is a progressive neurodegenerative disease. Brain neuron damage and synaptic dysfunction lead to impaired cognitive functions such as memory, executive function, and social interaction ability, ultimately impairing the ability to perform daily activities. Currently, there is no effective cure for AD. Mild cognitive impairment (MCI) is the early stage of Alzheimer's disease (AD). Effectively detecting the early stage of AD and intervening in a timely manner to slow down the progression of the disease is the key to solving the development of AD disease at present. That is, timely and accurate MCI diagnosis plays a crucial role in the treatment of AD and preventing the progression of the disease.

[0004] Traditional AD diagnosis relies on complex medical screenings, including neuropsychological tests, neuroimaging, and biomarker analysis. However, biomarker detection and neuroimaging examinations are expensive and invasive, making it difficult to meet the needs of large-scale clinical diagnosis. Neuropsychological tests such as the Mini-Mental State Examination (MMSE) and the Montreal Cognitive Assessment Scale (MoCA) are widely used in clinical practice to evaluate the severity of cognitive impairment. However, these evaluations rely heavily on the expertise of clinicians and may lead to differences in the assessment of the disease severity, especially in the early stage of AD when symptoms are mild or unclear. That is, the limitations of traditional diagnostic methods may lead to misdiagnosis or underestimation of MCI.

[0005] Currently, multiple studies have pointed out that eye movements can be used as potential biomarkers for AD diagnosis. There is a certain correlation between AD-related pathology and damage to the oculomotor nerve pathway. Therefore, eye movement experimental paradigms have been designed to explore abnormal eye movement behaviors in AD patients. For example, by designing short-term visual memory tasks such as the Visual Short-Term Memory task (VSTM) and the Visual Paired Comparison task (VPC), and then combining machine learning and deep learning models to detect visual short-term memory deficits in AD and MCI. However, the singularity of experimental paradigms and evaluation dimensions limits the development of existing research. Current research focuses on the evaluation of single dimensions such as visual memory, saccades, or free viewing, lacking a comprehensive evaluation of multi-dimensional cognitive abilities. The detection performance for MCI patients is average, while multi-dimensional cognitive evaluation can more effectively stimulate eye movement behaviors and related brain activities, thus capturing more complex cognitive characteristics. In addition, current research usually uses heatmaps and manually extracted eye movement metrics, failing to fully consider the dynamic changes in eye movements during the cognitive process. In fact, eye movement behaviors dynamically reflect the behavioral performance, memory, and cognitive processes during the execution of cognitive tasks. The temporal information of cognitive tasks is of great significance for the depth and accuracy of cognitive function evaluation. Summary of the Invention

[0006] To address the deficiencies of the above-mentioned prior art, the present invention provides a system and method for assisting in the identification of mild cognitive impairment based on eye movement characteristics. By designing an emotion recognition and memory paradigm to examine the differences in visual attention and eye movement behaviors among AD, MCI, and HC (Healthy Control) during the two processes of emotional memory and emotional selection, and designing a multi-modal fusion network that integrates eye movement temporal features and spatial features, the accuracy of final identification and classification is improved by combining the complementary information of static visual attention heatmaps and dynamic eye movement sequences.

[0007] In the first aspect, the present invention provides a system for assisting in the identification of mild cognitive impairment based on eye movement characteristics.

[0008] A system for assisting in the identification of mild cognitive impairment based on eye movement characteristics includes:

[0009] A data acquisition module, configured to acquire eye movement data of a patient during an emotion classification task; wherein, the emotion classification task includes emotional memory and emotional selection, and the eye movement data includes eye movement heatmap data and dynamic eye movement sequence data;

[0010] A data preprocessing module, configured to perform data preprocessing on the acquired eye movement data;

[0011] An auxiliary recognition module is used to input the preprocessed eye movement data into an auxiliary recognition model based on a two-stream fusion network. The auxiliary recognition model includes a spatial feature extraction module, a temporal feature extraction module, and a multi-modal fusion classification module. The spatial feature extraction module extracts the key features of the eye movement heat map weighted by spatial and channel attention, the temporal feature extraction module extracts the temporal features of the dynamic eye movement sequence, and the multi-modal fusion and classification module splices and fuses the features extracted by the two feature extraction modules to output the final classification and recognition results of Alzheimer's disease, mild cognitive impairment, and health.

[0012] A further technical solution, the emotion classification task includes:

[0013] Emotional memory: Display facial emotion target pictures for emotion recognition and memory within a set short time duration.

[0014] Emotional selection: After the target picture disappears and a blank screen is displayed for a set short time duration, display multiple facial emotion pictures including the target emotion for selection.

[0015] A further technical solution, the dynamic eye movement sequence data is the temporal data of 6 eye movement features, and the 6 eye movement features include the left and right eye fixation position coordinates of both eyes gazing at the screen and the pupil diameter sizes of the left and right eyes.

[0016] A further technical solution, the data preprocessing includes:

[0017] Perform data augmentation processing on the eye movement heat map data, including: using the central cropping method to process the collected eye movement heat map to generate a heat map with unified size and channels, and then using the random horizontal flip, random rotation, and random cropping methods to enhance the heat map data;

[0018] Perform interpolation filling and data augmentation on the dynamic eye movement sequence data, including:

[0019] Select the median value of the eye movement sequence length as the final eye movement sequence length. For data longer than the sequence length, perform downsampling. For sequences with a length shorter than the sequence length, use the linear interpolation method to linearly interpolate the sequence data; then perform data augmentation on the eye movement sequence data such as randomly adding noise, randomly cropping and filling, and speed transformation.

[0020] A further technical solution, the spatial feature extraction module uses the MobileNet V2 convolutional module as the backbone network, and introduces the CBAM attention module in the backbone network to enhance the channel and feature expression capabilities;

[0021] The input eye movement heat map passes through the standard convolutional layer and multiple Bottleneck layers of the MobileNet V2 convolutional module to extract initial features; the initial features are then input into the CBAM attention module, and the channel attention weight and spatial attention weight are extracted through the channel attention mechanism and spatial attention mechanism respectively. The initial features are weighted by the channel attention weight and spatial attention weight to obtain the weighted key features;

[0022] Among them, the channel attention weight for extracting the initial features through the attention mechanism is as follows: global average pooling and global max pooling are respectively used to perform pooling processing on the initial features, and the features after the two poolings are then input into a shared multi-layer perceptron to obtain the global average channel feature and global max channel feature. The two obtained features are added and activated through the Sigmoid function to obtain the final channel attention weight;

[0023] The spatial attention weight for extracting the initial features through the spatial attention mechanism is as follows: global average pooling and global max pooling are respectively used to perform pooling processing on the initial features in the channel direction, and the features after the two poolings are then input into a shared multi-layer perceptron to obtain the global average spatial feature and global max spatial feature. The two obtained features are concatenated in the channel dimension and then activated through the Sigmoid function to obtain the final spatial attention weight.

[0024] In a further technical solution, the temporal feature extraction module uses 1D-CNN as the backbone network for sequence feature extraction. Among them, 1D-CNN slides a convolutional kernel on the input dynamic eye movement sequence for feature extraction, captures the local temporal dependence relationship in the sequence through multiple one-dimensional convolutional layers and pooling layers, performs temporal feature extraction, and finally outputs the extracted temporal features through a fully connected layer.

[0025] In a second aspect, the present invention provides a method for assisting in the identification of mild cognitive impairment based on eye movement features.

[0026] A method for assisting in the identification of mild cognitive impairment based on eye movement features includes:

[0027] Obtain the eye movement data of a patient during an emotion classification task; among them, the emotion classification task includes emotion memory and emotion selection, and the eye movement data includes eye movement heat map data and dynamic eye movement sequence data;

[0028] Perform data preprocessing on the obtained eye movement data;

[0029] Input the preprocessed eye movement data into the auxiliary recognition model based on the dual-stream fusion network. The auxiliary recognition model includes a spatial feature extraction module, a temporal feature extraction module, and a multi-modal fusion classification module. The spatial feature extraction module extracts the key features of the eye movement heat map weighted by spatial and channel attention. The temporal feature extraction module extracts the temporal features of the dynamic eye movement sequence. The multi-modal fusion and classification module splices and fuses the features extracted by the two feature extraction modules, and outputs the final classification recognition results of Alzheimer's disease, mild cognitive impairment, and health.

[0030] In a third aspect, the present invention also provides an electronic device, including: a memory for storing executable instructions; a processor for implementing the above-mentioned auxiliary recognition system for mild cognitive impairment based on eye movement features, or performing the steps of the above-mentioned auxiliary recognition method for mild cognitive impairment based on eye movement features when executing the executable instructions stored in the memory.

[0031] In a fourth aspect, the present invention also provides a computer-readable storage medium storing executable instructions for causing a processor to implement the above-mentioned auxiliary recognition system for mild cognitive impairment based on eye movement features, or perform the steps of the above-mentioned auxiliary recognition method for mild cognitive impairment based on eye movement features when executing the executable instructions.

[0032] In a fifth aspect, the present invention also provides a computer program product, which includes executable instructions stored in a computer-readable storage medium; wherein, when the processor of the electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, it implements the above-mentioned auxiliary recognition system for mild cognitive impairment based on eye movement features, or performs the steps of the above-mentioned auxiliary recognition method for mild cognitive impairment based on eye movement features.

[0033] The above one or more technical solutions have the following beneficial effects:

[0034] 1. The present invention provides an auxiliary recognition system and method for mild cognitive impairment based on eye movement features. By designing an emotion recognition and memory paradigm to examine the visual attention and eye movement behavior differences of AD, MCI, and HC in the two processes of emotional memory and emotional selection, and designing a multi-modal fusion network that fuses eye movement time features and spatial features, by combining the complementary information of static visual attention heat maps and dynamic eye movement sequences, the accuracy of the final recognition and classification is improved. The effectiveness of this method is verified on the constructed multi-modal dataset, achieving a diagnosis of AD with an accuracy of 92.54% and a diagnosis of MCI with an accuracy of 90.64%. This further confirms the effectiveness of the multi-modal eye movement feature fusion method based on the emotion recognition and memory paradigm proposed by the present invention in the diagnosis of cognitive impairment patients.

[0035] 2. By designing a cognitive experiment paradigm and combining emotion recognition with visual memory, the present invention can provide a more multi-dimensional cognitive function assessment. That is, through the cognitive processes of emotion recognition and memory and emotion selection, the differences of an individual in aspects such as visual memory, emotion recognition, and information processing ability can be comprehensively reflected, so as to capture more complex cognitive characteristics and improve the accuracy of the final recognition. By proposing a multi-modal fusion framework based on the spatio-temporal feature fusion of eye movements in a two-stream network, combining eye movement heat map data and dynamic eye movement sequence data, the visual attention features and the temporal features of the changes in eye movement behaviors during the visual cognitive process can be effectively captured, and through fusing the complementary spatio-temporal features of eye movements, that is, fusing the visual attention features and the dynamic change features of eye movements, a more accurate and robust diagnosis of cognitive impairment can be realized.

[0036] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0038] Figure 1 It is an architecture diagram of the system for assisting in the identification of mild cognitive impairment based on eye movement features according to an embodiment of the present invention;

[0039] Figure 2 It is an overall flowchart of the method for assisting in the identification of mild cognitive impairment based on eye movement features according to an embodiment of the present invention;

[0040] Figure 3 It is a schematic diagram of the experimental paradigm of emotion recognition and memory tasks in an embodiment of the present invention;

[0041] Figure 4 It is a heat map of 2 AD patients, 2 MCI patients, and 2 healthy subjects during the memory process in an embodiment of the present invention;

[0042] Figure 5 It is a heat map of 2 AD patients, 2 MCI patients, and 2 healthy subjects during the selection process in an embodiment of the present invention;

[0043] Figure 6 It is a schematic diagram of the change of the left eye fixation position and pupil diameter over time in an embodiment of the present invention;

[0044] Figure 7 It is a schematic diagram of the change of the right eye fixation position and pupil diameter over time in an embodiment of the present invention;

[0045] Figure 8 This is a schematic diagram of the eye movement spatio-temporal feature fusion framework based on emotion recognition and memory in the embodiments of the present invention. Detailed implementation manners

[0046] It should be noted that the following detailed descriptions are all exemplary and are only for describing the specific implementation manners, aiming to provide further explanations for the present invention and are not intended to limit the exemplary embodiments according to the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those of ordinary skill in the technical field to which the present invention belongs. In addition, it should also be understood that when the terms "include" and / or "comprise" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.

[0047] Embodiment 1

[0048] This embodiment provides an auxiliary recognition system for mild cognitive impairment based on eye movement features, as Figure 1 shown, specifically including:

[0049] A data acquisition module, configured to acquire eye movement data of a patient during an emotion classification task; wherein, the emotion classification task includes emotion memory and emotion selection, and the eye movement data includes eye movement heat map data and dynamic eye movement sequence data;

[0050] A data preprocessing module, configured to perform data preprocessing on the acquired eye movement data;

[0051] An auxiliary recognition module, configured to input the preprocessed eye movement data into an auxiliary recognition model based on a two-stream fusion network. The auxiliary recognition model includes a spatial feature extraction module, a temporal feature extraction module, and a multi-modal fusion classification module. The spatial feature extraction module extracts key features of the eye movement heat map weighted by spatial and channel attention, the temporal feature extraction module extracts temporal features of the dynamic eye movement sequence, and the multi-modal fusion and classification module splices and fuses the features extracted by the two feature extraction modules, and outputs the final classification recognition results of Alzheimer's disease, mild cognitive impairment, and health.

[0052] The auxiliary recognition system for mild cognitive impairment based on eye movement features provided in this embodiment is introduced in more detail through the following content.

[0053] (1) Data acquisition module

[0054] First, considering that the emotion recognition ability of AD patients is impaired, and this defect is already obvious in the early stage of the disease. Experiments have shown that patients with mild cognitive impairment show moderate impairment in recognizing fear, sadness, and anger, especially in facial and vocal emotion recognition. Therefore, the decline in emotion recognition ability can be used as an important indicator for identifying early symptoms of mild cognitive impairment. For this reason, this embodiment proposes an eye movement tracking cognitive experiment paradigm, which evaluates cognitive functions from multiple dimensions such as visual memory, emotion recognition, and information processing through two cognitive processes: emotion memory and emotion selection, and captures more complex cognitive features through multi-dimensional cognitive function evaluation to improve the accuracy of final recognition.

[0055] Specifically, in the recognition and memory paradigm of facial expressions, seven facial expressions including happy, disgusted, afraid, angry, neutral, sad, and surprised are used to examine the emotion recognition ability and short-term memory ability of participants. In the task adopted in this embodiment, the facial emotion pictures used are selected from the existing RaFD (Radboud Faces Database) facial stimulus set, and include frontal facial pictures of the above seven facial expressions. Among them, the emotion classification task includes two stages: the memory stage and the selection stage. In the emotion memory stage, a facial emotion picture, that is, the target picture, is displayed in the center of the screen for a set duration of 5 seconds for emotion recognition and memory; in the emotion selection stage, after the target image disappears, a blank screen is displayed for a set duration of 5 seconds. After this delay, four emotion images including the target emotion are displayed as options. Preferably, to reduce the influence of the human face, the human faces in the target picture and all option images are different from the human face in the original target picture. The entire paradigm is as Figure 3 shown.

[0056] Secondly, the Tobii Pro Fusion 250 (250HZ) eye tracker is used to collect the eye movement data of patients during the emotion classification task. Specifically, the eye movement tracking task paradigm is displayed on a monitor with a resolution of 1920×1080. The eye tracker is placed under the monitor, and binocular tracking is used. The subject is 60 - 80 centimeters away from the monitor and is told to keep the head fixed during the task. In addition, a binocular calibration process is performed between each task to maintain the accuracy of the eye tracker throughout the task; if the accuracy is poor, recalibration is performed. The entire eye movement data collection process is carried out in a professional eye movement collection room to ensure the quietness of the collection room and the stability of the light source, so as to maintain consistent collection conditions for different subjects and ensure the accuracy of final recognition.

[0057] The eye movement data collected in the above manner includes eye movement heat map data and dynamic eye movement sequence data, which are used as the source input of the fusion model to extract discriminative representation features for classifying AD patients, MCI patients, and normal people.

[0058] Among them, the eye movement heat map (abbreviated as heat map) represents the visual attention distribution of the subject, as Figure 4 and Figure 5 shown in the heat maps of two randomly selected AD patients, two MCI patients, and two normal individuals during the memory process and the selection process respectively. There are certain differences in the visual attention heat maps of AD patients, MCI patients, and healthy control subjects during emotion recognition and recall selection. For example, the visual attention of healthy subjects is more concentrated, and they are more inclined to fixate on the eyes and mouth of the face and other parts expressing emotions, while cognitively impaired patients are more likely to pay attention to positions unrelated to emotions such as the forehead or hair. Moreover, compared with healthy controls, the visual attention of cognitively impaired patients is more dispersed. During the recall and selection process, healthy controls can find the target pictures more efficiently and accurately, and AD patients are more inclined to make wrong choices compared with MCI patients.

[0059] The dynamic eye movement sequence data is the original eye movement data of the subject during the execution of the emotion recognition and memory tasks, which reflects the dynamic changes of the subject's eye movement behavior over time t during the completion of the cognitive experimental paradigm. The original eye movement data is preprocessed through the filter provided by the TobiiPro Lab software to extract the dynamic eye movement sequence data, that is, the left and right eye fixation position coordinates (x left , y left , x right , y right ) of both eyes fixating on the screen and the pupil diameter sizes of the left and right eyes (pupil left , pupil right ) at each time stamp t after denoising, a total of six time series data of eye movement features. As Figure 6 and Figure 7 shown in the changes of the fixation positions and pupil diameter sizes of AD patients, MCI patients, and healthy individuals over time t during the memory and recognition phases (where the eye movement changes during the selection processes of angry and disgust are shown in the figure), AD patients and MCI patients show longer fixation times compared with healthy control subjects, with smaller pupil dilation and dilation delays.

[0060] (2) Data preprocessing module

[0061] For subsequent feature extraction and fusion, this embodiment preprocesses the heat map data and eye movement sequence data through this module, such as data enhancement and interpolation filling.

[0062] Among them, data augmentation processing is performed on the eye movement heatmap data, including: using the central cropping method to process the collected eye movement heatmap to generate a heatmap with unified size and channels, such as generating heatmap data with a size of (224, 224) and 3 channels; then using random horizontal flipping, random rotation (±15°), and random cropping methods to augment the heatmap data.

[0063] Interpolation filling and data augmentation are performed on the dynamic eye movement sequence data, including: to ensure the consistency of the eye movement dynamic sequence length, the median value of the eye movement sequence length is selected as the final eye movement sequence length. For data longer than the sequence length, downsampling is performed. For sequences with a length shorter than the sequence length, linear interpolation is used to linearly interpolate the sequence data. The advantage of this linear interpolation is that it can maintain the timeliness and trend characteristics of the data during the sequence length adjustment process, thereby minimizing the impact of data transformation on the dynamic sequence information; then, random noise addition, random cropping and filling, speed transformation, etc. are performed on the eye movement sequence data for data augmentation.

[0064] (3) Auxiliary recognition module

[0065] In this auxiliary recognition module, this embodiment proposes an auxiliary recognition model based on a two-stream fusion network to extract and fuse multi-modal features from the eye movement heatmap and eye movement dynamic sequence data to achieve the recognition of AD, MCI, and HC. The proposed network model framework is as Figure 8 shown, and it consists of three parts: a spatial feature extraction module, a temporal feature extraction module, and a multi-modal fusion classification module. Among them, the spatial and channel features of the eye movement heatmap are extracted by the spatial feature extraction module, the temporal features of the dynamic eye movement sequence are extracted by the temporal feature extraction module, and the features extracted by the two feature extraction modules are concatenated and fused by the multi-modal fusion and classification module to output the final classification recognition results of Alzheimer's disease AD, mild cognitive impairment MCI, and healthy HC.

[0066] (3.1) Spatial feature extraction module

[0067] The eye movement heatmap reflects the visual attention distribution of the subject when viewing different emotional faces, including the spatial distribution of fixation points and emotional correlation, etc. Considering that the MobileNet network model can optimize the performance and computational complexity of image feature extraction through depthwise separable convolution, inverted residual structure, and linear Bottleneck, based on the above characteristics, this embodiment uses the MobileNet V2 convolutional module as the backbone network of the spatial feature extraction module, and introduces a CBAM (Convolutional Block Attention Module) attention module into the backbone network to enhance the channel and feature expression capabilities and more accurately capture the key regions in the eye movement heatmap.

[0068] Specifically, first, input the eye movement hotspot map image F with a size of 224×224×3 img , and extract the initial features with a size of 7×7×1280 through the standard convolutional layer and multiple Bottleneck layers of the MobileNet V2 convolutional module as follows:

[0069] Secondly, since the key features in the hotspot map are relatively concentrated, in order to reduce the interference of background redundant information in the picture and focus on the key hotspot features, a CBAM module is introduced on the basis of the initial feature extraction. Combining the channel attention and spatial attention mechanisms, the initial features are further optimized.

[0070] In this embodiment, the extracted initial features are input into the CBAM attention module, and the CBAM attention module jointly models the spatial and channel features, that is, extracts the channel attention weight and spatial attention weight through the channel attention mechanism and spatial attention mechanism respectively. By calculating the attention weight, the network's attention to the key channels is enhanced, and the focusing ability on the key features is improved. The initial feature map is weighted by the channel attention weight and spatial attention weight to obtain the weighted key features, which are:

[0071]

[0072] where M c is the channel attention weight of CBAM, and M s is the spatial attention weight of CBAM. ⊙ represents the Hadamard product, that is, multiplying the corresponding position elements.

[0073] Among them, the channel attention weight of the initial features extracted by the attention mechanism is:

[0074] In the CBAM attention module, global average pooling (GAP) and global max pooling (GMP) are used to pool the initial features in two different ways respectively to extract the importance information of the channels and capture the important relationships of each channel. The feature vectors after the two poolings are then input into the shared multi-layer perceptron (MLP) to obtain the global average channel feature Avg c and the global max channel feature Max c . The two obtained features are added and activated through the Sigmoid function to obtain the final channel attention weight M c , which can be expressed as: M c = sigmoid(Avg c + Max c ).

[0075] The spatial attention weights for extracting initial features using the spatial attention mechanism are as follows:

[0076] In the CBAM attention module, global average pooling (GAP) and global max pooling (GMP) are used in the channel direction to perform pooling operations on the initial features respectively, so as to extract the important information in the feature space in the channel direction. The feature vectors after the two poolings are then input into a shared multi-layer perceptron (MLP) to obtain the global average spatial feature Avg s and the global max spatial feature Max s . The two obtained features are concatenated in the channel dimension and then activated through the Sigmoid function to obtain the final spatial attention weight M s , which can be expressed as: M s = sigmoid(conv(concat[Avg s , Max s )).

[0077] (3.2) Temporal Feature Extraction Module

[0078] In this embodiment, the temporal feature extraction module uses 1D-CNN as the backbone network for sequence feature extraction. The convolution kernel slides on the input dynamic eye movement sequence for feature extraction. Through multiple one-dimensional convolutional layers and pooling layers, the local temporal dependencies in the sequence are captured to extract representative temporal features. Specifically, the input eye movement sequence is F seq ∈R 1250×6 , the length of this eye movement sequence is 1250, and the dimension is 6. Feature extraction is performed through 3 one-dimensional convolutional layers (the convolution kernel size is 3, and the number of convolution kernels is 32, 64, and 128 respectively), and finally the extracted temporal features are output through the fully connected layer as:

[0079]

[0080] (3.3) Multimodal Fusion Classification Module

[0081] To fully combine the spatial features and temporal features, the feature maps of the spatial and temporal modules are fused before the fully connected layer to retain the interactivity of the features of each branch network. The feature vectors extracted by the above two modules are fused through the feature concatenation operation to form a joint feature representation, and classification is completed through the fully connected layer and the Softmax activation function, which can be expressed as:

[0082]

[0083] As another implementation, for the above network model, this embodiment uses Python (version 3.10) and Pytorch (version CUDA 12.2) as the backend of the deeplearning network framework. The training and testing processes use the binary cross-entropy loss function and are completed on an Nvidia GeForce RTX 4090 GPU with 24GB of memory. For each training, a batch size of 32 samples is randomly selected (batch size = 32), the number of training epochs is 200 (epoch = 200), the cross-entropy loss function and the Adam optimizer are selected, the initial learning rate is 1e-4, and a learning rate decay strategy is adopted. After every 40 epochs, the learning rate decays to 0.90 times the original value. To reduce overfitting, dropout is used, where p = 0.9.

[0084] In addition, eye movement data was collected from 15 AD patients, 23 MCI patients, and 32 healthy controls (HC). All participants underwent a detailed assessment of the MoCA neuropsychological test by an expert, and the assessment results were used as labels for the collected data to construct the training dataset for the above network model. Among them, the AD patients (10 females and 5 males) had an average age of 66.53 ± 9.19 years, the MCI patients (11 females and 12 males) had an average age of 67.57 ± 9.47 years, and the healthy controls (20 females and 12 males) had an average age of 60.56 ± 4.92 years. There were no significant differences in age and educational background among all experimental participants. In addition, criteria were set to exclude subjects diagnosed with any neurological diseases other than AD: those with uncorrected visual dysfunction, hearing loss, aphasia, or inability to complete clinical examinations or scale assessments; those with a history of mental disorders and illicit drug abuse. The statistical characteristics of the participants are shown in Table 1 below.

[0085] Table 1 Statistical characteristics of the subjects

[0086]

[0087] For the collection of eye movement data for each subject, during the recognition process, the subject needs to complete a total of 7 selection tasks. Before the task starts, the subject is explained and demonstrated to ensure that the subject understands the task process and is familiar with mouse operations. During the task, the participant will not receive any prompts and is free to observe the images and make selections. The entire experimental process lasts about 3 minutes.

[0088] To evaluate the above model, 10-fold cross-validation was adopted, that is, the participants were divided into 10 non-overlapping subsets. Each time, the data of all participants in one of the subsets was used as the test set, and the data of all participants in the remaining 9 subsets was used as the training set. When dividing the subsets, it was ensured that the proportions of AD, MCI, and HC participants in each subset were roughly balanced. This process was repeated 10 times, and finally, the average performance of all folds was taken as the evaluation metric of the model. On this basis, accuracy, precision, recall, and F1-Score were used as evaluation criteria to comprehensively evaluate the classification performance of the model. The formula calculations for the above criteria are as follows:

[0089]

[0090] Among them, TP (True Positive) represents the number of samples that are actually positive classes and are correctly classified as positive classes; TN (True Negative) represents the number of samples that are actually negative classes and are correctly classified as negative classes; FP (False Positive) represents the number of samples that are actually negative classes and are misclassified as positive classes; FN (False Negative) represents the number of samples that are actually positive classes and are misclassified as negative classes.

[0091] To further verify the effectiveness and superiority of the proposed solution in this embodiment, it was verified through the following experiments and experimental results.

[0092] (1) Ablation experiment

[0093] A. AD vs MCI vs HC

[0094] To evaluate the effectiveness of the above solution for the diagnosis of AD and MCI, an ablation experiment was conducted. The classification performance of the baseline model was compared with that of the proposed feature extraction module, and the classification performance of the single-modal model was also compared with that of the proposed fusion model. Experiments on patients with cognitive impairment and healthy controls were carried out on the constructed dataset. The following Tables 2, 3, and 4 respectively show the classification experiment results of patients with cognitive impairment and healthy controls.

[0095] Table 2 Results of the three-class ablation experiment for AD, MCI, and HC

[0096]

[0097] Table 2 shows the experimental results of different methods for classifying AD, MCI, and HC. It can be seen that the proposed model achieved the maximum classification accuracy of 78.31%. The classification accuracy using only heatmap data was 72.92%, and the classification accuracy using only eye movement sequence data was the lowest at 72.02%. After adding the CBAM module to the spatial feature extraction model, the classification accuracy increased by 1.79%, indicating that the attention mechanism enhances the key spatial and channel features of the heatmap, which can improve the classification performance. When fusing the eye movement sequence data and heatmap data, the accuracy, precision, and recall all increased significantly. Compared with the classification results of the direct concatenation model, the accuracy increased by 1.67%, indicating that there is complementary information between the heatmap data and the eye movement sequence data. Fusing the two modalities helps to more comprehensively represent the features of AD, MCI, and HC, improving the robustness and accuracy of classification, and demonstrating the superiority of the proposed scheme in this embodiment for the current three-classification task.

[0098] B. AD vs HC

[0099] Table 3 Ablation Experiment Results for AD and HC Classification

[0100]

[0101] Table 3 presents the experimental results of different methods in the AD and HC classification experiment. After fusing the heatmap data and the sequence data, the classification accuracy of the model was the highest, reaching 92.54%. The classification accuracies of the eye movement sequence data and the heatmap data were 85.56% and 89.62% respectively. Compared with the eye movement sequence, the heatmap data has better classification performance, indicating that the heatmap contains more information related to the classification task. The ability of the eye movement sequence data to distinguish between cognitively impaired patients and healthy controls is limited, which may be due to the weak ability of sequence features to represent complex patterns. Compared with the baseline model, the classification accuracy of this model increased by 6.98% and 2.92% respectively. Therefore, the fusion of heatmap features and eye movement sequence features significantly improves the classification performance.

[0102] C. MCI vs HC

[0103] Table 4 Design and Results of Ablation Experiment for MCI and HC Classification

[0104]

[0105] Table 4 shows the experimental results of different methods for classifying MCI and HC patients. It can be seen from Table 4 that the proposed multi-modal fusion model in this embodiment achieved an accuracy of 90.64%, a precision of 89.04%, a recall of 88.72%, and an F1-score of 88.91%, which is significantly better than the classification performance based on a single modality. This method performs well in the identification of MCI patients, demonstrating its effectiveness.

[0106] (2) Performance in the memory process and the selection process

[0107] To evaluate the influence of different cognitive stages on classification performance, experiments were conducted in the memory stage and the selection stage of the emotion recognition and memory paradigm. Tables 5, 6, and 7 below show the classification performance in the memory and selection stages. In the task of classifying AD, MCI, and HC, the accuracy in the memory stage was 68.72%, the precision was 67.63%, and the recall was 67.82%, indicating that this stage provided certain classification information and performed well in distinguishing AD, MCI, and HC. The classification performance in the selection stage reached an accuracy of 70.29%, and at the same time, the precision and recall were improved. In addition, similar trends were also observed in the classification of AD and HC, as well as MCI and HC, indicating the robustness of the selection stage. This shows that the selection stage provides additional discriminant information, possibly due to the inclusion of more complex features (such as decision-making), which enhances the model's ability to distinguish different groups. The combined data from the memory stage and the selection stage achieved the best classification performance. This indicates that the eye movement features from the memory and selection stages provide a more comprehensive cognitive feature representation and can better classify AD, MCI, and HC. These results verify the effectiveness and advancement of the proposed solution in this embodiment.

[0108] Table 5 Classification results in the memory stage and the selection stage (AD vs MCI vs HC)

[0109] Task Phase Accuracy Precision Recall F1-score Memory 0.6872 0.6763 0.6782 0.6774 Selection 0.7029 0.6933 0.7029 0.6971 Combination of Two Phases 0.7831 0.7919 0.7974 0.7935

[0110] Table 6 Classification results in the memory stage and the selection stage (AD vs HC)

[0111] Task Phase Accuracy Precision Recall F1-score Memory 0.8322 0.8299 0.8232 0.8252 Selection 0.8841 0.8860 0.8481 0.8791 Combination of Two Phases 0.9254 0.9378 0.9334 0.9351

[0112] Table 7 Classification results in the memory stage and the selection stage (MCI vs HC)

[0113] Task Phase Accuracy Precision Recall F1-score Memory 0.8150 0.8021 0.8154 0.8095 Selection 0.8778 0.8806 0.8778 0.8786 Combination of Two Phases 0.9064 0.8904 0.8872 0.8891

[0114] (3) Comparison with other methods

[0115] The classification results and performance comparison between the method proposed in this embodiment and the state-of-the-art method are shown in Table 8 below. Zuo et al. and Sun et al. used heatmap data and achieved classification accuracies of 84% and 85% for AD and HC, respectively. Yin et al. used the extracted eye movement metrics and achieved an 86% classification accuracy for AD and HC. Most of the eye movement-based AD diagnosis works rely on unimodal methods, such as heatmaps or manually extracted eye movement features (fixation and saccade metrics). In addition, most studies focus on the discrimination between AD and HC, while the classification of MCI and HC has been less explored. This embodiment has achieved good classification performance in all classification tasks. Specifically, the unimodal methods perform well in the classification tasks, with an accuracy of 90% using heatmap data and an accuracy of 86% using sequence data for classifying AD and HC. In the classification of MCI and HC, this embodiment has achieved the best performance in terms of accuracy, precision, and recall. In addition, the multimodal fusion method (heatmap + sequence) of this embodiment outperforms all unimodal methods in the classification tasks, achieving an accuracy of 78% in AD vs MCI vs HC, an accuracy of 93% in AD vs HC, and an accuracy of 91% in MCI vs HC. These results indicate that heatmap and sequence data provide complementary information, and their fusion enables the model to effectively utilize these two types of features.

[0116] Comparison Results with the Current State-of-the-Art Methods

[0117]

[0118] The deep learning-based method proposed in this embodiment diagnoses MCI based on the eye movement data collected from the visual tasks of emotion recognition and memory. The eye movement data exists in the form of heatmaps and eye movement sequences, reflecting the visual attention information and dynamic eye movement changes of the participants. From Figures 4 - 7It can be seen that there are significant differences in the visual attention distribution and eye movement changes among AD patients, MCI patients, and HC. This embodiment proposes a dual-stream fusion network based on deep learning to extract and integrate heatmap features and eye movement sequence features to achieve the classification of AD patients, MCI patients, and HC. In the fusion network, the MobilenetV2 network and 1DCNN are respectively used for the extraction and encoding of eye movement features of the heatmap and eye movement sequence, and the feature representation is enhanced through the attention mechanism. Finally, the spatial features and temporal features are fused to achieve the classification of AD, MCI, and HC. As can be seen from Table 8, the proposed method is compared with the state-of-the-art classification network on the collected eye movement tracking data. It has the best classification performance on AD and HC, with the highest accuracy, precision, and recall rate, and performs well in the three-class classification of AD, MCI, and HC. After fusing the heatmap data and eye movement sequence data, the best classification accuracy is obtained. Compared with the method using single-modal heatmap data or single-modal eye movement index data, the classification accuracy, precision, and recall rate are all significantly improved. The eye movement sequence and heatmap data have complementary characteristics in this task, and fusing the two data modalities can significantly improve the classification performance of the model. In addition, as shown in Tables 5, 6, and 7, the classification performance of AD, MCI, and HC is better in the selection stage. After combining the data in the memory stage and the selection stage, the best classification performance can be obtained, indicating that the proposed experimental paradigm can comprehensively capture the abnormal eye movements of AD and MCI.

[0119] In summary, the multi-modal fusion framework proposed in this embodiment can effectively improve the early detection of AD and MCI. By fusing the eye movement heatmap data and sequence data into the auxiliary diagnosis of AD and MCI, through the fused enhanced spatio-temporal feature information, the cognitive differences among AD, MCI, and HC participants can be more comprehensively characterized, and the classification and recognition performance can be improved.

[0120] Embodiment 2

[0121] This embodiment provides a method for assisting in the identification of mild cognitive impairment based on eye movement features, as Figure 2 shown, including:

[0122] Obtain the eye movement data of the patient during the emotion classification task; wherein, the emotion classification task includes emotion memory and emotion selection, and the eye movement data includes eye movement heatmap data and dynamic eye movement sequence data;

[0123] Perform data preprocessing on the obtained eye movement data;

[0124] Input the preprocessed eye movement data into the auxiliary recognition model based on the dual-stream fusion network. The auxiliary recognition model includes a spatial feature extraction module, a temporal feature extraction module, and a multimodal fusion classification module. The key features of the eye movement heat map weighted by spatial and channel attention are extracted by the spatial feature extraction module, the temporal features of the dynamic eye movement sequence are extracted by the temporal feature extraction module, and the features extracted by the two feature extraction modules are concatenated and fused by the multimodal fusion and classification module to output the final classification and recognition results of Alzheimer's disease, mild cognitive impairment, and health.

[0125] Embodiment III

[0126] This embodiment provides an electronic device, including: a memory for storing executable instructions; a processor for implementing the above method provided by this embodiment or implementing the above system provided by this embodiment when executing the executable instructions stored in the memory.

[0127] Embodiment IV

[0128] This embodiment also provides a computer-readable storage medium storing executable instructions. When the executable instructions are executed by a processor, the processor will be caused to execute the above method provided by this embodiment or implement the above system provided by this embodiment.

[0129] Embodiment V

[0130] This embodiment provides a computer program product. The computer program product includes executable instructions, which are a kind of computer instructions; the executable instructions are stored in a computer-readable storage medium. When the processor of an electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the electronic device is caused to execute the above method provided by this embodiment or implement the above system provided by this embodiment.

[0131] The steps involved in Embodiments II to V above correspond to those in Embodiment I. For specific implementation manners, reference may be made to the relevant description part of Embodiment I. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to execute any method in the present invention.

[0132] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device. Thus, they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0133] The above are only the preferred embodiments of the present invention. Although the specific implementation manners of the present invention have been described in conjunction with the accompanying drawings, they are not limitations on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solutions of the present invention are still within the protection scope of the present invention.

Claims

1. An auxiliary recognition system for mild cognitive impairment based on eye movement characteristics, characterized in that, Including: A data acquisition module for acquiring eye movement data of a patient during an emotion classification task; wherein, the emotion classification task includes emotion memory and emotion selection, and the eye movement data includes eye movement heat map data and dynamic eye movement sequence data; A data preprocessing module for preprocessing the acquired eye movement data; An auxiliary recognition module for inputting the preprocessed eye movement data into an auxiliary recognition model based on a two-stream fusion network. The auxiliary recognition model includes a spatial feature extraction module, a temporal feature extraction module, and a multi-modal fusion classification module. The key features of the eye movement heat map based on spatial and channel attention weighting are extracted by the spatial feature extraction module, the temporal features of the dynamic eye movement sequence are extracted by the temporal feature extraction module, and the features extracted by the two feature extraction modules are spliced and fused by the multi-modal fusion and classification module to output the final classification and recognition results of Alzheimer's disease, mild cognitive impairment, and health.

2. The mild cognitive impairment assisted recognition system based on eye movement features according to claim 1, wherein The emotion classification task includes: Emotion memory, which is: displaying facial emotion target pictures for emotion recognition and memory within a set short time duration; Emotion selection, which is: waiting for the target picture to disappear and displaying a blank screen for a set short time duration, and then displaying multiple facial emotion pictures including the target emotion for selection.

3. The mild cognitive impairment auxiliary recognition system based on eye movement characteristics according to claim 1, characterized in that, The dynamic eye movement sequence data is the temporal data of 6 eye movement features, and the 6 eye movement features include the left and right eye fixation position coordinates and the pupil diameter sizes of both eyes fixating on the screen.

4. The mild cognitive impairment assisted recognition system based on eye movement features according to claim 1, characterized in that, The data preprocessing includes: Performing data augmentation on the eye movement heat map data, including: processing the collected eye movement heat map by the central cropping method to generate a heat map with unified size and channels, and then performing augmentation on the heat map data by the random horizontal flipping, random rotation, and random cropping methods; Performing interpolation filling and data augmentation on the dynamic eye movement sequence data, including: Selecting the median value of the eye movement sequence length as the final eye movement sequence length, downsampling the data longer than the sequence length, and linearly interpolating the sequence data with a length shorter than the sequence length by the linear interpolation method; and then performing data augmentation on the eye movement sequence data by randomly adding noise, randomly cropping and filling, and speed transformation.

5. The mild cognitive impairment auxiliary recognition system based on eye movement characteristics according to claim 1, characterized in that, The spatial feature extraction module uses the MobileNet V2 convolutional module as the backbone network, and introduces the CBAM attention module in the backbone network to enhance the channel and feature expression capabilities; The input eye movement heat map passes through the standard convolutional layer and multiple Bottleneck layers of the MobileNet V2 convolutional module to extract initial features; The initial features are then input into the CBAM attention module, and the channel attention weight and spatial attention weight are extracted respectively through the channel attention mechanism and the spatial attention mechanism. The initial features are weighted by the channel attention weight and the spatial attention weight to obtain the weighted key features; Among them, the channel attention weight for extracting the initial features by using the attention mechanism is as follows: global average pooling and global max pooling are respectively used to perform pooling processing on the initial features, and the features after the two poolings are then input into a shared multi-layer perceptron to obtain the global average channel feature and the global max channel feature. The two obtained features are added and activated through the Sigmoid function to obtain the final channel attention weight; The spatial attention weight for extracting the initial features by using the spatial attention mechanism is as follows: global average pooling and global max pooling are respectively used to perform pooling processing on the initial features in the channel direction, and the features after the two poolings are then input into a shared multi-layer perceptron to obtain the global average spatial feature and the global max spatial feature. The two obtained features are concatenated in the channel dimension and then activated through the Sigmoid function to obtain the final spatial attention weight.

6. The mild cognitive impairment assisted recognition system based on eye movement features according to claim 1, wherein The temporal feature extraction module uses 1D-CNN as the backbone network for sequence feature extraction. Among them, 1D-CNN uses a convolutional kernel to slide on the input dynamic eye movement sequence for feature extraction, captures the local temporal dependence relationship in the sequence through multiple one-dimensional convolutional layers and pooling layers, performs temporal feature extraction, and finally outputs the extracted temporal features through a fully connected layer.

7. An auxiliary recognition method for mild cognitive impairment based on eye movement characteristics, characterized in that, Including: Obtain the eye movement data of the patient during the emotion classification task; among them, the emotion classification task includes emotion memory and emotion selection, and the eye movement data includes eye movement heat map data and dynamic eye movement sequence data; Perform data preprocessing on the obtained eye movement data; Input the preprocessed eye movement data into an auxiliary recognition model based on a two-stream fusion network. The auxiliary recognition model includes a spatial feature extraction module, a temporal feature extraction module, and a multi-modal fusion classification module. The key features based on spatial and channel attention weighting of the eye movement heat map are extracted by the spatial feature extraction module, the temporal features of the dynamic eye movement sequence are extracted by the temporal feature extraction module, and the features extracted by the two feature extraction modules are concatenated and fused by the multi-modal fusion and classification module, and the final classification and recognition results of Alzheimer's disease, mild cognitive impairment, and health are output.

8. An electronic device, characterized in that, Including: A memory for storing executable instructions; A processor, when executing the executable instructions stored in the memory, implements the mild cognitive impairment auxiliary recognition system based on eye movement features according to any one of claims 1-6, or executes the steps of the mild cognitive impairment auxiliary recognition method based on eye movement features according to claim 7.

9. A computer-readable storage medium, characterized in that, Stored with executable instructions for causing the processor to implement the mild cognitive impairment auxiliary recognition system based on eye movement features according to any one of claims 1-6, or execute the steps of the mild cognitive impairment auxiliary recognition method based on eye movement features according to claim 7 when executing the executable instructions.

10. A computer program product, characterized in that, The computer program product includes executable instructions, and the executable instructions are stored in a computer-readable storage medium; When the processor of the electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the system for assisting in identifying mild cognitive impairment based on eye movement features according to any one of claims 1-6 is implemented, or the steps of the method for assisting in identifying mild cognitive impairment based on eye movement features according to claim 7 are executed.

Citation Information

Cited By

  • Cognitive disorder assessment method and system

    CN121890951A

  • A method and system for assessing cognitive impairment

    CN121890951B