A system for assisting in the screening of mild cognitive impairment

By combining a multimodal fusion method of cognitive scales and near-infrared time-series data, and utilizing XGBoost and ResNet models, the accuracy and efficiency issues in screening for mild cognitive impairment were addressed, achieving efficient identification of mild cognitive impairment.

CN116999027BActive Publication Date: 2026-05-15INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310969606.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-03
Publication Date
2026-05-15
Estimated Expiration
2043-08-03

AI Technical Summary

Technical Problem

Existing technologies have low accuracy and efficiency in screening for mild cognitive impairment, making it difficult to effectively distinguish between normal individuals and patients with mild cognitive impairment, and the diagnostic results are greatly affected by the subjectivity of doctors.

Method used

A multimodal fusion method is adopted, which combines cognitive scales and near-infrared time series data. The confidence scores are obtained by using XGBoost and ResNet models respectively, and the fusion module integrates multiple confidence scores to improve accuracy.

Benefits of technology

It improved the accuracy and efficiency of screening for mild cognitive impairment, with a classification accuracy rate of 93.9%, and reduced the influence of subjectivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116999027B_ABST
    Figure CN116999027B_ABST
Patent Text Reader

Abstract

The application provides a system for assisting screening of mild cognitive impairment, comprising: a first modality verification module for obtaining a cognitive score vector of a subject, determining a first and a second confidence according to the cognitive score vector, the cognitive score vector consisting of scores of the subject tested on at least two cognitive scales; a second modality verification module for obtaining multi-channel near-infrared time series data collected from the subject by a plurality of brain electrodes, converting the multi-channel near-infrared time series data into a Grimm angular field image, determining a third and a fourth confidence according to fusion features extracted from the Grimm angular field image, wherein the first confidence table and the third confidence indicate a probability of normal cognitive ability, and the second confidence table and the fourth confidence indicate a probability of existing mild cognitive impairment; and a fusion module for fusing the first confidence and the third confidence to obtain a fifth confidence, and fusing the second confidence and the fourth confidence to obtain a sixth confidence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of neural network technology, specifically to the fields of cognitive impairment diagnosis, intelligent medicine and multimodal fusion, and more specifically, to a system for assisting in the screening of mild cognitive impairment. Background Technology

[0002] Cognitive ability is the capacity for clear thinking, learning, and memory.

[0003] Based on different cognitive abilities, cognitive impairment can be divided into normal cognition (NC), mild cognitive impairment (MCI), and Alzheimer's disease (AD, also known as dementia in some contexts).

[0004] Normal cognitive function is fundamental to daily life, affecting all aspects of learning and work. Today, with the continuous changes in the social environment, such as an aging population, cognitive health has gradually become a social issue. Statistics from the World Health Organization (WHO) in 2019 indicate that approximately 50 million older adults worldwide experience cognitive decline. Among those aged 60 and over, the incidence of Alzheimer's disease is as high as 5% to 8%.

[0005] There is currently no effective treatment for the cognitive decline symptoms in Alzheimer's patients. The only way to slow the progression of the disease is through early screening and intervention.

[0006] In early screening, it is relatively easy to distinguish Alzheimer's patients from people with normal cognition because there are obvious differences in the outward manifestations of their cognitive abilities.

[0007] However, mild cognitive impairment, as a precursor stage of Alzheimer's disease, is an intermediate state between normal cognition and dementia. In reality, because patients with mild cognitive impairment exhibit mild symptoms and are not significantly different from normal individuals in daily life, distinguishing between a normal person and someone with mild cognitive impairment presents considerable difficulty, posing a significant challenge to screening efforts. Furthermore, since the assessment of cognitive ability is highly subjective, the opinions of clinicians can greatly influence the accuracy of the diagnosis.

[0008] Therefore, to reduce the influence of subjectivity, doctors often use various tools for screening. There is already considerable research on the early screening of mild cognitive impairment, which can be conducted through a variety of methods, including those based on electrophysiological signals (electromyography, electroencephalography, etc.), cognitive scales, brain imaging techniques (MRI, etc.), and biomarkers (protein deposits in cerebrospinal fluid, Aβ). 40 / 42 Screening methods include ratios, patient behavior (eye movements, gait, gestures, etc.).

[0009] To improve the accuracy and efficiency of screening, existing technologies also utilize models to predict cognitive abilities. However, both accuracy and efficiency need improvement. Summary of the Invention

[0010] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a system for assisting in the screening of mild cognitive impairment.

[0011] The objective of this invention is achieved through the following technical solution:

[0012] According to a first aspect of the present invention, a system for assisting in the screening of mild cognitive impairment is provided, comprising: a first modality testing module, configured to acquire a cognitive score vector of a subject, and determine a first confidence level and a second confidence level based on the cognitive score vector, wherein the cognitive score vector consists of scores of the subject on at least two cognitive scales, the first confidence level representing the probability of normal cognitive ability, and the second confidence level representing the probability of the presence of mild cognitive impairment; and a second modality testing module, configured to acquire multi-channel data collected by multiple brain electrodes from the subject during a delayed matching test task. The system uses near-infrared time-series data, converts it into a Grimm field image, extracts fusion features of the time and spatial domains related to cognitive ability from the Grimm field image, and determines a third confidence level and a fourth confidence level based on the fusion features. The third confidence level indicates the probability of normal cognitive ability, and the fourth confidence level indicates the probability of mild cognitive impairment. A fusion module is used to fuse the first confidence level and the third confidence level to obtain a fifth confidence level, and to fuse the second confidence level and the fourth confidence level to obtain a sixth confidence level.

[0013] Optionally, at least two cognitive scales are selected in the following manner: based on the scores of multiple individuals on a variety of preset cognitive scales and the cognitive status of each individual determined by experts, a chi-square test is used to perform a correlation analysis of mild cognitive impairment on the multiple cognitive scales to obtain a chi-square test score; from the multiple cognitive scales, the cognitive scales with chi-square test scores greater than or equal to a predetermined threshold are selected.

[0014] Optionally, the at least two cognitive scales include: MMSET, BNT, STROOP2, AVLT1, MOCAT, STROOP3, SDMT, AVLT4, AVLT6, and AVLT5, or combinations thereof.

[0015] Optionally, the first modality testing module includes a trained XGBoost model configured to determine a first confidence level and a second confidence level for the test based on the cognitive score vector.

[0016] Optionally, the trained XGBoost model is trained as follows: a first training set is obtained, which includes first samples from multiple individuals and a first label corresponding to each first sample. The first sample is a cognitive score vector composed of the scores of the corresponding individual on at least two cognitive scales. The first label indicates whether the corresponding individual is cognitively normal or has mild cognitive impairment. The XGBoost model is trained once or multiple times using the first training set under binary classification supervision to obtain the trained XGBoost model.

[0017] Optionally, the second modality testing module includes a trained ResNet model configured to extract cognitive ability-related temporal and spatial fusion features from the Grimm field image and determine a third confidence level and a fourth confidence level based on the fusion features, wherein the ResNet model includes a ResNet18 model.

[0018] Optionally, the trained ResNet model is trained in the following manner: a second training set is obtained, which includes second samples from multiple people and a second label corresponding to each second sample. The second sample is a Grimm field image of the corresponding person, and the second label indicates whether the corresponding person is cognitively normal or has mild cognitive impairment. The ResNet model is trained once or multiple times using the second training set to obtain the trained ResNet model.

[0019] Optionally, the multi-channel near-infrared time-series data consists of 39 channels of data collected from multiple infrared electrodes distributed in a dispersed manner in the brain.

[0020] Optionally, the multi-channel near-infrared time-series data is processed by noise reduction, conversion to oxyhemoglobin concentration, and conversion to Grimm angle field data to obtain a Grimm angle field image.

[0021] Optionally, the number of samples in the first training set belonging to patients with normal cognition and mild cognitive impairment is balanced, and the number of samples in the second training set belonging to patients with normal cognition and mild cognitive impairment is balanced. Attached Figure Description

[0022] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:

[0023] Figure 1 This is a schematic diagram of a system for assisting in the screening of mild cognitive impairment according to an embodiment of the present invention;

[0024] Figure 2 This is a schematic diagram illustrating the principle of the chi-square test according to an embodiment of the present invention;

[0025] Figure 3 This is a schematic diagram illustrating the principle of data balancing according to an embodiment of the present invention;

[0026] Figure 4 This is a schematic diagram showing the location of brain electrodes according to an embodiment of the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the invention.

[0028] As mentioned in the background section, existing technologies also utilize models to predict cognitive abilities in order to improve the accuracy and efficiency of screening. However, accuracy and efficiency need further improvement. The inventors, through research, believe that, firstly, existing technologies often aim for models to have more classification capabilities, able to distinguish between normal individuals, MCI patients, and AD patients. However, since AD ​​patients are generally easy to identify, this invention sacrifices this classification function, allowing the model to focus more on the binary classification problem of distinguishing between normal individuals and MCI patients, thereby improving classification accuracy. Second, near-infrared time-series data is a time-series signal. Since conventional methods only extract time-domain features (or temporal features), the accuracy of cognitive ability classification based on these time-domain features is low. Therefore, this invention converts near-infrared time-series data into Grimm field images and extracts fusion features of the time and spatial domains related to cognitive abilities from these Grimm field images. Cognitive ability classification is then performed based on these fusion features, improving classification accuracy. Third, relying solely on one modality from cognitive scales or electrophysiological signals for auxiliary judgment in cognitive ability screening suffers from low single-modality accuracy and poor model prediction performance, resulting in low accuracy in distinguishing between normal individuals and MCI patients. Therefore, this invention integrates binary classification confidence scores based on cognitive scales and binary classification confidence scores from near-infrared time-series data to improve the accuracy of cognitive ability classification. Through these improvements, the model's performance in identifying mild cognitive impairment is enhanced.

[0029] To better understand the system of this invention, we will first briefly introduce its overall architecture. The system of this invention aims to determine whether a subject has normal cognition or mild cognitive impairment, i.e., a binary classification problem. Wherein:

[0030] The system of this invention comprises three modules, namely:

[0031] The first modality test module is used to determine two confidence levels for binary classification (i.e., the first confidence level representing normal cognitive ability and the second confidence level representing mild cognitive impairment, abbreviated as a1 and b1, respectively) based on the cognitive score vector composed of the scores of the subjects on at least two cognitive scales.

[0032] The second modality testing module is used to determine two confidence levels for binary classification based on multi-channel near-infrared time-series data collected from subjects by multiple brain electrodes during a delayed matching test task (i.e., a third confidence level representing normal cognitive ability and a fourth confidence level representing mild cognitive impairment, abbreviated as a2 and b2, respectively).

[0033] The fusion module is used to obtain confidence level A (i.e., the fifth confidence level) by fusing a1 and a2, and to obtain confidence level B (i.e., the sixth confidence level) by fusing b1 and b2.

[0034] As can be seen from the definitions of a1 and a2, b1 and b2 above, confidence level A can reflect the probability that the subject is a person with normal cognitive ability (1-A is B), and confidence level B can reflect the probability that the subject has mild cognitive impairment (1-B is A).

[0035] Therefore, either confidence level A or confidence level B can assist medical personnel in assessing the cognitive status of the subject. However, it should be noted that the confidence level A or confidence level B determined by the system of this invention is merely a machine-predicted confidence level, serving only as an auxiliary diagnostic tool and cannot be used as the final diagnostic result. For example, doctors may also need to consider the subject's performance in other aspects to determine whether a subject has mild cognitive impairment.

[0036] For a better understanding of the system of the present invention, see [link to relevant documentation]. Figure 1 The following sections will introduce the first modality test module 100, the second modality test module 200, and the fusion module 300 respectively.

[0037] I. First Modal Testing Module 100

[0038] According to one embodiment of the present invention, a first modality testing module is used to obtain the cognitive score vector of the subject, and to determine a first confidence level that the subject belongs to a cognitively normal person and a second confidence level that belongs to a person with mild cognitive impairment based on the cognitive score vector. The cognitive score vector is composed of the subject's scores on at least two cognitive scales. For example, the cognitive score vector is composed of scores from tests on the Mini-Mental State Examination (MMSET), the Montreal Cognitive Assessment Scale (MoCAT), the Auditory Word Learning Test (AVLT), the Stroop Color Word Test, the Symbolic Number Transformation Test (SDMT), the Boston Naming Test (BNT), and the Clock Drawing Test (CDT), or combinations thereof. This embodiment achieves at least the following beneficial technical effects: the present invention utilizes a cognitive score vector composed of the subject's scores on at least two cognitive scales for the first test, combining the advantages of multiple cognitive scales to improve the accuracy of the final result.

[0039] Cognitive scales are the most common method for assessing cognitive abilities. Many cognitive scales exist, but their ability to detect different types of cognitive impairment varies. In early screening for Mild Cognitive Impairment (MCI), cognitive scales with low sensitivity can affect the accuracy of classification results. To improve the accuracy of the final results, screening can be performed first. According to one embodiment of the present invention, the at least two cognitive scales are selected as follows: based on the scores of multiple individuals on a set of preset cognitive scales and the cognitive status of each individual determined by experts, a chi-square test is used to perform a correlation analysis of mild cognitive impairment on the multiple cognitive scales to obtain a chi-square test score; from the multiple cognitive scales, cognitive measures with chi-square test scores greater than or equal to a predetermined threshold are selected. Since not all cognitive scales are suitable for predicting MCI, feature correlation tests are used to improve the effectiveness of cognitive scales in predicting MCI classification results. The chi-square test is generally used to study the difference between categorical and non-categorical data, and can effectively determine the difference and correlation between two sets of variables. By analyzing the scores of multiple individuals on multiple cognitive scales and performing a chi-square test, the correlation between the cognitive scales and their corresponding label categories is calculated. This process eliminates some cognitive scales with low correlation, preventing them from affecting the representation results. The technical solution of this embodiment achieves at least the following beneficial effects: In the cognitive scale representation part, this embodiment of the invention screens effective scales by performing feature correlation tests on each cognitive scale, further improving the accuracy of the final results.

[0040] For example, referring to Table 1, the inventors conducted a chi-square test on the multiple cognitive scales shown in the first column of Table 1, and the corresponding scores are shown in the second column of Table 1. Implementers can set a predetermined threshold (e.g., 2 or 4) as needed, leaving only the at least two cognitive scales. Preferably, the at least two cognitive scales include: MMSET, BNT, STROOP2, AVLT1, MOCAT, STROOP3, SDMT, AVLT4, AVLT6, and AVLT5, or combinations thereof.

[0041] Table 1

[0042]

[0043]

[0044] Note:

[0045] STROOP1 is a cognitive scale for recognizing multiple words that do not confuse colors; for example, words that all express colors semantically, such as "red," "yellow," "blue," and "purple," but all of which are black.

[0046] STROOP2 is a cognitive scale about reading the colors of multiple circles (or squares);

[0047] STROOP3 is a cognitive scale for reading multiple words with confusing colors. In this scale, the color that each word represents semantically is different from the color that the word appears to have (illustrated, for example, the color of the word "red" is set to blue, the color of the word "blue" is set to green, the color of the word "green" is set to purple, etc., i.e. confusing colors).

[0048] AVLT1 is a cognitive scale for the ability to recall words immediately after hearing them;

[0049] AVLT4 is a cognitive scale that assesses the ability to recall words after a 5-minute delay following hearing them.

[0050] AVLT5 is a cognitive scale that assesses the ability to recall words after a 20-minute delay following hearing them.

[0051] AVLT6 is a cognitive scale that tests the ability to recall several words after being given a prompt, following the completion of the AVLT5 test.

[0052] Based on the chi-square test scores of the cognitive scale, a first predictive model can be constructed to determine the first and second confidence levels based on the cognitive score vector. According to one embodiment of the invention, the first predictive model can be a random forest, Naive Bayes, support vector machine, or XGBoost model. The inventors conducted comparative experiments on these four machine learning models. The comparative experiment based on the cognitive scale was a single-modal training: a total of 82 subjects were included, comprising 50 MCIs and 32 NCs. In the single-modal training, the dataset was randomly shuffled and subjected to five-fold cross-validation, with the average of the five-fold cross-validation results used as the final model result. Each subject corresponds to one sample, and one sample includes gender, education level, age (these three are considered as cognitive scales), and 12 cognitive scales such as STROOP and CDTT, i.e., one sample includes 15 dimensions of features. The Sklearn third-party package was imported, and comparisons were made based on the four models, with the input dimensions and training methods remaining consistent across the four models. Simultaneously, based on the chi-square test results, the accuracy of models with 3, 5, and 7 features removed was compared. As shown in Table 2, after removing some cognitive scales with low relevance from Table 1, the results of the comparative experiment were as follows. The comparative experiment revealed that the XGBoost model outperformed the other three models. Therefore, the XGBoost model will be the preferred model in a more detailed description below.

[0053] Table 2

[0054]

[0055] According to one embodiment of the present invention, a first modality testing module includes a trained XGBoost model configured to determine a first confidence level and a second confidence level for a test based on a cognitive score vector. The trained XGBoost model is trained as follows: obtaining a first training set, the first training set including first samples from multiple individuals and a first label corresponding to each first sample, wherein the first sample is a cognitive score vector composed of the scores of the corresponding individual on at least two cognitive scales, and the first label indicates whether the corresponding individual is cognitively normal or has mild cognitive impairment; and performing one or more supervised binary classification training operations on the XGBoost model using the first training set to obtain the trained XGBoost model. See again. Figure 1The XGBoost model consists of an XGBoost decision tree layer and a softmax layer. The XGBoost decision tree layer is used to obtain the corresponding decision value based on the first input sample, and the softmax layer obtains the first confidence and second confidence scores for the first sample based on the decision value. During training, the classification loss is calculated based on the first and second confidence scores and the first label for each first sample, and the XGBoost decision tree layer is adjusted according to the classification loss. Due to the relatively small sample size, this training can use leave-one-out validation, employ the softmax layer to obtain binary classification probabilities, remove the first 5 cognitive scales, and keep the rest consistent with the unimodal training.

[0056] According to one embodiment of the invention, the number of samples in the first training set belonging to patients with normal cognition and those with mild cognitive impairment is balanced. Preferably, see... Figure 3 In the first training set, where the number of the first samples belonging to patients with normal cognition and mild cognitive impairment is unbalanced, a data balancing operation is performed to achieve numerical balance by adding new samples by selecting two or more samples from the smaller side and calculating the average.

[0057] II. Second Modal Testing Module

[0058] The second modality testing module is used to acquire multi-channel near-infrared time-series data collected from subjects during a delayed matching test task by multiple brain electrodes. The multi-channel near-infrared time-series data is converted into a Grimm field image. Fusion features of the temporal and spatial domains related to cognitive ability are extracted from the Grimm field image. A third confidence level and a fourth confidence level are determined based on the fusion features. The third confidence level indicates the probability of normal cognitive ability, and the fourth confidence level indicates the probability of mild cognitive impairment. Preferably, the multi-channel near-infrared time-series data consists of 39 channels of data collected from multiple distributed infrared electrodes on the brain. According to one embodiment of the present invention, see... Figure 4 Multi-channel near-infrared time-series data is composed of near-infrared time-series data collected from multiple channels in the brain of a person using multiple infrared emitting electrodes and multiple infrared receiving electrodes in multiple brain electrodes. These channels include the left frontal cortex (LFC), right frontal cortex (RFC), left parietal cortex (LPC), right parietal cortex (RPC), left occipital cortex (LOC), and right occipital cortex (ROC). Channels numbered 1-39 are as follows: Figure 4 As shown. It should be understood that, if needed, implementers can add or remove some electrodes using similar methods to achieve similar effects.

[0059] In their research on near-infrared characterization methods, the inventors considered that while near-infrared is a high-temporal-resolution detection method, its spatial resolution is slightly insufficient compared to other physiological signals. Therefore, to leverage the advantages of computer vision and incorporate spatial attention information into the data, the inventors converted the data from each channel of the multi-channel near-infrared time-series data (a one-dimensional near-infrared time-series signal) into images for processing. Different grimm angle field images were generated using different channels used in near-infrared detection. Thus, converting the one-dimensional near-infrared time-series signal into an image allows for the integration of temporal and spatial attention, improving the feature learning effect of near-infrared signals.

[0060] Subsequently, a ResNet model can be used to extract fusion features from the Grimm field image and perform classification prediction. According to one embodiment of the present invention, the second modality testing module includes a trained ResNet model configured to extract cognitive ability-related temporal and spatial fusion features from the Grimm field image and determine a third and fourth confidence level based on the fusion features. The ResNet model includes a ResNet18 model. This embodiment achieves at least the following beneficial technical effects: using a ResNet18 model to extract fusion features allows for better extraction of fusion features related to mild cognitive recognition after near-infrared data is converted into a Grimm field image, resulting in better classification performance. This will be demonstrated experimentally later.

[0061] Indicatively, during near-infrared data acquisition, the electrode positions on the brain form 39 channels, and the data obtained is actually from these 39 channels. During data acquisition, 82 subjects (or relevant personnel) completed a delayed matching test task (or delayed matching task, which involves showing a target image, then a period of black screen, followed by the display of four images (containing the target image and other images used for confusion), from which the subject selects the target image. This experiment was repeated ten times to obtain a time-series data. One acquisition consisted of 10 cycles of the delayed matching task. After being segmented according to task time, the data format was 82*10*39*539, where 539 is the length of the data points for one task. The illustrative task flow for each delayed matching task was: 15 seconds of rest, followed by 12 seconds of image observation. The black screen memory lasts 12 seconds, and the image option test lasts 10 seconds; the total time is 49 seconds, with 11 near-infrared data acquisitions per second, resulting in 539 data points per delayed matching task. Due to the large number of data points, windowing transformation can be performed, reducing each set of four points to a single data point through windowing (e.g., averaging the values ​​of four points). GAF image conversion is then performed to obtain a 135*135 image, resulting in a final format of 82*10*39*135*135. A single sample is 39*135*135, with the number of channels used as the third dimension, resulting in an image with a third dimension of 39. Therefore, it can be considered that spatial attention is integrated without requiring additional computation to emphasize it.

[0062] Following the above process, 820 images were obtained. The dataset was randomly shuffled and then divided into a training set:validation set:test set ratio of 8:1:1. This training set was used as the second training set and input into ResNet18 for training. Alternatively, due to the smaller number of images, the dataset could be divided without the proportional ratio, instead using a leave-one-out-of-one validation set, with the remaining samples used as the second training set.

[0063] According to one embodiment of the present invention, the trained ResNet model is trained as follows: A second training set is obtained, the second training set including second samples from multiple individuals and a second label corresponding to each second sample, wherein the second sample is a Grimm's angle field image of the corresponding individual, and the second label indicates whether the corresponding individual is cognitively normal or has mild cognitive impairment; the ResNet model is then subjected to one or more supervised binary classification training sessions using the second training set to obtain the trained ResNet model. See again. Figure 1The ResNet model contains ResNet network layers and softmax layers. The ResNet network layers extract fusion features based on the second sample, and the softmax layers output third and fourth confidence scores based on these fusion features. Binary cross-entropy loss is calculated based on the output third and fourth confidence scores and the corresponding second label. The gradient is calculated based on the binary cross-entropy loss, and backpropagation is used to update the trainable parameters of the ResNet network layers. According to one embodiment of the present invention, the pre-training parameters of the ResNet18 model are removed; that is, the initial trainable parameters of the ResNet18 model are obtained through random initialization, the training epochs are set to 50, and the output dimension is modified to binary classification.

[0064] According to one embodiment of the present invention, the number of samples belonging to cognitively normal and mild cognitive impairment patients in the second training set is balanced. Preferably, if the number of samples belonging to cognitively normal and mild cognitive impairment patients in the second training set is unbalanced, data balancing is performed by adding new samples by selecting two or more samples from the smaller side and calculating the average to achieve balance. Illustratively, in near-infrared modality studies, because the number of MCI patients and normal subjects is not the same, there are differences in weights during model training, resulting in lower classification accuracy for normal subjects. Therefore, the inventors generated new data by assigning the same weights to normal subjects, thus balancing the number of subjects in both categories.

[0065] To achieve better results, the multi-channel near-infrared time-series data can be preprocessed before being converted into a Grimm angle field image. According to one embodiment of the present invention, the multi-channel near-infrared time-series data undergoes noise reduction processing, conversion to oxyhemoglobin concentration processing, and conversion to Grimm angle field data processing to obtain a Grimm angle field image.

[0066] The following explanation uses specific formulas as examples:

[0067] Methods for detecting brain activity mainly include electroencephalogram (EEG), magnetic resonance imaging (MRI), and functional near-infrared spectroscopy (fNIRS). FNIRS falls between the two, offering a more convenient and user-friendly approach. FNIRS can detect changes in optical density caused by variations in blood oxygen concentration in the brain. These indicators can be used to calculate absorbance using Beer-Lambert's law (Equation 1).

[0068]

[0069] Where Abso represents absorbance or light attenuation, I t Let I0 represent the intensity of the incident light, I0 represent the intensity of the transmitted light, K represent the absorption coefficient or molar absorption coefficient, l represent the thickness of the absorbing medium, and c represent the concentration of the absorbing substance. During detection, light attenuates due to absorption or scattering. Beer-Lambert's law combines these parameters to calculate the magnitude of the attenuation. Since photons do not travel in a straight line from the source to the detector, there is a certain error. Therefore, a modified Beer-Lambert's law is introduced, resulting in the following formula 2:

[0070]

[0071] Where C represents concentration, ε represents extinction coefficient, L represents actual optical path length, G represents the sum of light intensity attenuation caused by factors other than oxyhemoglobin and deoxyhemoglobin, HbO2 represents oxyhemoglobin, and Hb represents hemoglobin. Correspondingly, ε represents the extinction coefficient of HbO2. Hb This represents the extinction coefficient of Hb. C represents the concentration of HbO2. Hb This represents the concentration of Hb. Once the type of substance and the wavelength of the incident light are determined, the extinction coefficient ε of that substance can also be determined. By calculating the light intensity attenuation per unit time, the influence of the G factor can be eliminated. Combining the two wavelengths used in near-infrared detection, the concentrations of deoxyhemoglobin (Hb) and oxyhemoglobin (HbO2) can be calculated using the following formula 3.

[0072]

[0073] When neurons are excited, the concentration of HbO2 in the surrounding tissue increases, while the concentration of Hb decreases. Therefore, HbO2 was chosen for subsequent data analysis because its results are more suitable for near-infrared data processing. In the concentration change curve, there are common instances of sudden changes due to errors. We use mean filtering and motion artifact removal to smooth the curve and reduce the impact of noise. If resting-state data is available, baseline removal can also be performed to obtain more obvious task-driven concentration changes. Near-infrared data is generally processed using one-dimensional time series of concentration changes. These series can be used for classification, prediction, and other tasks using powerful machine learning methods. However, from a data dimension perspective, one-dimensional time series considers more features in the time domain and ignores features in the spatial domain. Adding spatial attention weights to the processing of one-dimensional time series can better integrate the feature information from the time and spatial domains, resulting in more accurate classification results. Converting one-dimensional time series into images for processing is effective; therefore, this invention attempts to convert near-infrared one-dimensional time series into Gramian Angular Field (GAF) images. Essentially, it involves magnifying the data from each time point in a one-dimensional time series into various positions within a square matrix, using Equation 4. This calculation method, while utilizing the high temporal resolution of near-infrared technology, also leverages the advantages of computer vision, combined with spatial attention, for processing.

[0074]

[0075] Where, φ n This represents the nth number in the input one-dimensional time series.

[0076] III. Integration Module

[0077] The fusion module is used to fuse the first confidence level and the third confidence level to obtain a fifth confidence level, and to fuse the second confidence level and the fourth confidence level to obtain a sixth confidence level.

[0078] According to one embodiment of the present invention, when the data has been balanced or the data for the two classes (NC and MCI) are already balanced, during the training of the XGBoost and ResNet models, the fusion module averages the first and third confidence scores to obtain a fifth confidence score; and averages the second and fourth confidence scores to obtain a sixth confidence score. To improve modal accuracy, the inventors fused the binary classification confidence scores based on cognitive scale and near-infrared data in the fusion module, and the results showed that the accuracy was significantly better than the single-modal classification results. If the data for the two classes is imbalanced during the training of the XGBoost and ResNet models, the weights of the binary classification confidence scores output by the two models can be adjusted multiple times to fuse the first and third confidence scores using a weighted summation; or, the second and fourth confidence scores can be fused. This leads to other implementation methods.

[0079] To verify the effectiveness of the system of the present invention, the inventors also conducted the following comparative experiments.

[0080] (1) The model used in the comparison experiment of near-infrared modes

[0081] In this invention, after reviewing the literature, several models with relatively good performance were selected and used for prediction on our data. Among them:

[0082] The system in this invention is compared with existing technology models:

[0083] 1) CNN: The CNN model has 17 layers, including one input layer, four convolutional layers, a batch normalization layer, two dropout layers, one max pooling layer, four fully connected layers, and one output layer;

[0084] 2) DCNN: The DCNN model consists of six different convolutional blocks and uses ELU as the activation function. DCNN also uses the original near-infrared one-dimensional time series signal as input;

[0085] 3) GAF-CNN: After converting the near-infrared one-dimensional time series into GAF images, the classification results are obtained through CNN. This CNN consists of 18 layers, including an input layer, four pairs of convolutional layers, a max pooling layer, a batch normalization layer, two fully connected layers, a dropout layer, and a softmax layer.

[0086] 4) GAF ResNet18: Compared to GAF-CNN, GAF-ResNet18 modifies the model from CNN to ResNet18 and adjusts the input and output dimensions to conform to the format of GAF images;

[0087] 5) Proposed system (unbalanced data): The accuracy of the system proposed in this invention without employing a data balancing strategy;

[0088] 6) The proposed system (balanced data);

[0089] The results of the comparative experiments on the near-infrared modes are shown in Table 2:

[0090] Table 2

[0091]

[0092] As shown in Table 2, after data balancing, the classification accuracy of near-infrared single-mode reached 88.02%, which is an improvement compared to existing technologies.

[0093] (2) Multimodal fusion of cognitive scales and near-infrared spectroscopy

[0094] Multimodal analysis is currently a highly effective method for improving model accuracy, with multimodal fusion being the most widely used. The ability to establish connections between fuzzy, heterogeneous data can significantly enhance machine learning. Multimodal analysis undoubtedly demonstrates powerful learning and analytical capabilities; therefore, this paper uses this method to fuse near-infrared spectroscopy and cognitive scales.

[0095] During data collection, considering the randomness of the population, the number of participants in the MCI and NC categories was not balanced. This imbalance could cause changes in model weights, leading to inaccurate results. Therefore, a data balancing method was proposed to make the learning ability of each type in the model approach the same. The main method is to achieve balance by generating fewer data points for the smaller category. To avoid inter-patient differences, each data point in the smaller category is assigned the same weight. New data is generated by averaging two adjacent data points, and so on, until all data have been averaged. This is to preserve the specificity of the data and prevent the generated new data from being too similar, which could lead to model overfitting. Afterwards, leave-one-out validation and the Softmax function are still used to obtain the binary classification probability values.

[0096] This invention chooses to perform fusion at the fusion module (equivalent to the decision layer) primarily because in other machine learning processes, such as the feature layer, it is difficult to find suitable fusion methods for the signal data of the cognitive scale and the near-infrared GAF ​​image data. Simultaneously, fusion at the decision layer can better distinguish the results of different subjects, which is more beneficial for clinical diagnosis. During the fusion process, leave-one-out validation was used to obtain the binary classification probabilities of the cognitive scale from the XGBoost model and the near-infrared binary classification probabilities from the ResNet18 model. Each subject's probability was assigned the same weight, and the average was calculated to obtain the final classification result for the current subject. Table 3 shows a comparison of the classification accuracy after combining and complementing the data, using both the cognitive scale and near-infrared single-modal and multi-modal fusion methods. It can be seen that the accuracy is significantly improved in the case of multi-modal fusion.

[0097] Table 3 Comparison of Near-Infrared Data Accuracy Before and After Balancing

[0098]

[0099] In the multimodal fusion experiment, the classification accuracy of cognitive quantitative data monomodal, near-infrared monomodal, and multimodal fusion was compared. The experimental results are shown in Table 3. The results show that after data balancing, the classification accuracy of near-infrared data on both aMCI and NC data increased, and the imbalance was alleviated. The accuracy of multimodal fusion reached 93.9%, providing important evidence for early screening of MCI patients.

[0100] In summary, when researching classification methods for early screening of dementia, the inventors first classified subjects using cognitive scales and near-infrared monomodal methods. Subject labeling primarily utilized the Mini-Mental State Examination (MMSE) and expert clinical diagnoses for initial classification. Secondly, in monomodal classification experiments using CNNs (Convolutional Neural Networks), the highest accuracy reached 87.38% for the cognitive scale and 88.02% for the near-infrared model. Therefore, the inventors believe that the shortcomings of existing technologies stem from a failure to consider the subjectivity of diagnosis, as well as low monomodal accuracy and poor model prediction. The inventors believe these shortcomings can be addressed through multimodal fusion. Multimodal methods, through various combinations, crossovers, and complementarities of two or more modalities, acquire the ability to process and interpret multimodal messages and demonstrate the correlation between multiple messages, yielding richer feature information than monomodal methods. Utilizing multimodal fusion can solve the problem of low monomodal accuracy in MCI screening methods, leading to accurate diagnoses. Feature transfer between different channels is also a characteristic of multimodal learning, which can learn to assist channels lacking modeling resources. Through multimodal learning, various features can be combined through representation, interpretation, and fusion processes to ultimately describe an accurate result. Compared to unimodal learning, multimodal learning is more reliable, uses more features and parameters, and is better suited to the high-precision requirements of the future.

[0101] In summary, this invention improves the accuracy of modal representation of cognitive scales by screening effective cognitive scales; enhances the learning effect of near-infrared signal features by converting one-dimensional time-series near-infrared signals into images, integrating temporal and spatial attention; solves the problem of imbalanced data among different categories of samples through data balancing processing; and improves the accuracy of cognitive assessment through multimodal fusion of near-infrared scales. Experimental validation using real clinical data shows results significantly higher than existing screening methods, with a classification accuracy of 93.9%.

[0102] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.

[0103] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.

[0104] Computer-readable storage media can be tangible devices that hold and store instructions for use by an instruction execution device. Computer-readable storage media can be, for example, including but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.

[0105] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A system for assisting in the screening of mild cognitive impairment, characterized in that, include: The first modality testing module is used to obtain the cognitive score vector of the subject and determine a first confidence level and a second confidence level based on the cognitive score vector. The cognitive score vector consists of the subject's scores on at least two cognitive scales. The first confidence level represents the probability of normal cognitive ability, and the second confidence level represents the probability of mild cognitive impairment. The second modality testing module is used to acquire multi-channel near-infrared time-series data collected from subjects by multiple brain electrodes during a delayed matching test task, convert the multi-channel near-infrared time-series data into a Grimm field image, extract fusion features of the time and spatial domains related to cognitive ability from the Grimm field image, and determine a third confidence level and a fourth confidence level based on the fusion features, wherein the third confidence level indicates the probability of normal cognitive ability and the fourth confidence level indicates the probability of mild cognitive impairment. The fusion module is used to fuse the first confidence level and the third confidence level to obtain a fifth confidence level, and to fuse the second confidence level and the fourth confidence level to obtain a sixth confidence level.

2. The system according to claim 1, characterized in that, The at least two cognitive scales were selected in the following manner: Based on the scores of multiple individuals on various pre-set cognitive scales and the cognitive status of each individual determined by experts, the chi-square test was used to conduct a correlation analysis of mild cognitive impairment on the various cognitive scales to obtain the chi-square test score. From the various cognitive scales, select the cognitive scales whose chi-square test scores are greater than or equal to a predetermined threshold.

3. The system according to claim 2, characterized in that, The at least two cognitive scales include: MMSET, BNT, STROOP2, AVLT1, MOCAT, STROOP3, SDMT, AVLT4, AVLT6, and AVLT5, or combinations thereof.

4. The system according to claim 3, characterized in that, The first modality testing module includes a trained XGBoost model configured to determine a first confidence level and a second confidence level for the test based on a cognitive score vector.

5. The system according to claim 4, characterized in that, The trained XGBoost model was trained in the following manner: Obtain a first training set, which includes first samples from multiple people and a first label corresponding to each first sample. The first sample is a cognitive score vector composed of the scores of the corresponding person on at least two cognitive scales. The first label indicates whether the corresponding person is cognitively normal or has mild cognitive impairment. The XGBoost model is trained by performing one or more binary classification supervised trainings using the first training set to obtain the trained XGBoost model.

6. The system according to claim 4, characterized in that, The second modality testing module includes a trained ResNet model configured to extract cognitive ability-related temporal and spatial fusion features from the Grimm field image and determine a third confidence level and a fourth confidence level based on the fusion features. The ResNet model includes a ResNet18 model.

7. The system according to claim 6, characterized in that, The trained ResNet model was obtained in the following manner: Obtain a second training set, which includes second samples from multiple people and a second label corresponding to each second sample. The second sample is a Grimm field image of the corresponding person, and the second label indicates whether the corresponding person is cognitively normal or has mild cognitive impairment. The ResNet model is trained once or multiple times using the second training set to obtain the trained ResNet model.

8. The system according to claim 7, characterized in that, Multichannel near-infrared time-series data consists of 39 channels of data collected from multiple infrared electrodes distributed in a dispersed manner in the brain.

9. The system according to claim 8, characterized in that, The multi-channel near-infrared time-series data were processed by noise reduction, conversion to oxyhemoglobin concentration, and conversion to Grimm angle field data to obtain the Grimm angle field image.

10. The system according to claim 5, characterized in that, The number of samples in the first training set belonging to patients with normal cognition and those with mild cognitive impairment is balanced.

11. The system according to claim 7, characterized in that, The sample size in the second training set is balanced between patients with normal cognition and those with mild cognitive impairment.