Alzheimer's disease computer-aided diagnosis system based on unified multi-modal mechanism visual language model

By adopting a visual language model with a unified multimodal mechanism in the Alzheimer's diagnosis system, the problems of multimodal data fusion and information interaction are solved, and higher diagnostic accuracy and model robustness are achieved.

CN120108688APending Publication Date: 2025-06-06CHONGQING UNIV

Patent Information

Application Number
CN202510017092.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing Alzheimer's diagnostic system has difficulties in multimodal data fusion and information interaction, resulting in insufficient diagnostic accuracy and adaptability.

Method used

The visual language model based on a unified multimodal mechanism is adopted, and efficient fusion and cross-attention processing of image and text data is achieved through the multimodal data input module, feature extraction and projection module, exception marking token integration module, unified multimodal attention mechanism module, exception detection module and diagnostic result generation module.

Benefits of technology

It improves the accuracy of early diagnosis of Alzheimer's disease, enhances the robustness and clinical practicality of the model, and can better adapt to the multimodal data characteristics of different patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108688A_ABST
    Figure CN120108688A_ABST
Patent Text Reader

Abstract

The invention discloses an Alzheimer's disease computer-aided diagnosis system based on a unified multi-modal mechanism visual language model. The Alzheimer's disease computer-aided diagnosis system comprises a multi-modal data input module, a feature extraction and projection module, an abnormal mark token integration module, a unified multi-modal attention mechanism module, an anomaly detection module and a diagnosis result generation module. The method has the outstanding advantages of improving the early diagnosis precision of the Alzheimer's disease, enhancing the model robustness and improving the clinical practicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Alzheimer's disease diagnosis, and in particular to an Alzheimer's disease computer-aided diagnosis system based on a unified multimodal mechanism visual language model. Background Art

[0002] Alzheimer's Disease (AD) is a neurodegenerative disease characterized by progressive cognitive decline. It is the main type of senile dementia and imposes a huge burden on patients, their families and society. At present, the diagnosis of AD mainly relies on the combination of medical imaging data (such as MRI, CT, etc.) and neuropsychological cognitive assessment scales (such as MMSE, CDR, etc.). However, due to the unclear early symptoms of patients and complex data, the traditional diagnostic system has certain subjectivity and uncertainty. In order to improve the accuracy and efficiency of diagnosis, computer-assisted diagnosis (CAD) methods based on artificial intelligence (AI) have received widespread attention in recent years.

[0003] In the prior art, many studies mainly diagnose Alzheimer's disease based on single-modal data. For example:

[0004] 1. Single-modality method based on image data: Analyze medical imaging data through deep learning technology to extract structural and functional brain features for identifying early signs of AD. Typical studies of this type of method include extracting MRI image features based on convolutional neural networks (CNN) and using them for the classification and staging of AD. However, due to ignoring the patient's cognitive test results and other non-imaging information, this method often fails to fully reflect the multidimensional characteristics of the disease.

[0005] 2. Single-modal method based on text data: With the patient's cognitive assessment scale as the core, combined with machine learning technology to analyze text data and mine the patient's cognitive impairment pattern. Although this method is helpful in assessing the patient's cognitive state, it lacks a direct reflection of brain structure and functional abnormalities, and the comprehensiveness of the diagnosis is limited.

[0006] 3. Multimodal diagnosis attempts: In recent years, multimodal diagnosis systems have begun to emerge. Some studies have tried to combine medical imaging data with neuropsychological scale data to improve diagnostic performance through simple feature splicing or multimodal fusion networks. However, these methods generally have the following shortcomings:

[0007] (1) Modality alignment problem: Image and text data come from different sources and their feature expressions are inconsistent, which makes it difficult to align the modalities and affects the diagnostic effect.

[0008] (2) Insufficient information interaction: Existing methods lack deep interaction mechanisms when processing multimodal data and cannot fully utilize the potential correlation between image and text data.

[0009] (3) Limited adaptability and generalization capabilities: In real clinical scenarios, data from different patients may have missing modalities or uneven quality. Existing methods perform poorly in dealing with such problems, limiting their practical application.

[0010] The main reasons for the above problems are: on the one hand, the differences in the characteristics of multimodal data make information interaction and alignment complicated; on the other hand, existing models mostly focus on a single task or modality, and there is relatively little research on multimodal interaction mechanisms. At the same time, due to the high cost of obtaining and annotating medical data, the training and verification of multimodal models face the difficulty of insufficient data.

[0011] Therefore, developing a visual language model (VLM) based on a unified multimodal mechanism that can efficiently fuse image and text data, overcome the difficulty of modal alignment, and provide higher accuracy and adaptability in the auxiliary diagnosis of Alzheimer's disease is an important direction of current research. Summary of the invention

[0012] The purpose of the present invention is to provide a computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal mechanism visual language model, comprising a multimodal data input module, a feature extraction and projection module, an abnormal marking token integration module, a unified multimodal attention mechanism module, an abnormality detection module, and a diagnosis result generation module;

[0013] The multimodal data input module acquires the multimodal data and transmits it to the feature extraction and projection module;

[0014] The multimodal data includes textual modal data and visual modal data; the visual modal data includes structural image data and metabolic image data;

[0015] The feature extraction and projection module extracts features from the multimodal data to obtain text features and initial visual features;

[0016] The feature extraction and projection module transmits the text features to the unified multimodal attention mechanism module;

[0017] The feature extraction and projection module projects the initial visual features into the shared embedding space to obtain a visual feature vector;

[0018] The abnormal mark token integration module concatenates the abnormal mark token with the visual feature to obtain the enhanced visual feature, and transmits it to the unified multimodal attention mechanism module;

[0019] The unified multimodal attention mechanism module performs cross-attention processing on text features and enhanced visual features to obtain multimodal fusion features, and transmits them to the anomaly detection module and the diagnosis result generation module;

[0020] The anomaly detection module stores a linear classification model;

[0021] The anomaly detection module inputs the multimodal fusion features into the linear classification module, generates anomaly detection results of visual modality data, and transmits them to the diagnosis result generation module;

[0022] The diagnosis result generation module stores a large language model;

[0023] The diagnosis result generation module will be input into the large language model to obtain a diagnosis report, including cognitive status assessment, imaging abnormality description and diagnosis conclusion.

[0024] Further, the text modality data is clinical text data, including demographic information and neuropsychological assessment results of the patient;

[0025] Said demographic information includes age and gender;

[0026] The neuropsychological assessment results include the Mini-Mental State Examination results and the Clinical Dementia Score;

[0027] The structural image data is a brain structural image, obtained by structural magnetic resonance imaging;

[0028] The metabolic image data is a metabolic activity image, which is obtained by positron emission tomography.

[0029] Furthermore, the feature extraction and projection module uses the pre-trained CLIP visual encoder ViT-L / 14 to extract features from the visual modality data in the multimodal data to obtain a feature vector z i ,Right now:

[0030]

[0031] Among them, g v (·) represents the feature extraction function of the visual encoder, An indexed collection of visual modalities.

[0032] Furthermore, the feature extraction and projection module uses the text embedding function g t (·) Extract features from text modal data in multimodal data to obtain feature vector h i ,Right now:

[0033]

[0034] in, An index collection of text modal.

[0035] Further, the visual feature vector is as follows:

[0036] The feature vectors z of all visual modalities i Projection into a shared embedding space

[0037]

[0038] in, is the projection matrix;

[0039] Further, the enhanced visual features are as follows:

[0040]

[0041] in, Represents a splicing operation, and is the enhanced feature vector;

[0042] Used to capture structural abnormalities in structural magnetic resonance (sMRI) modalities, Used to capture metabolic abnormalities in the positron emission tomography (PET) modality; the structural magnetic resonance (sMRI) modality is used to reflect brain structure; the positron emission tomography (PET) modality is used to reflect brain metabolic activity.

[0043] Further, the multimodal fusion features include interactive information between structure and metabolism, interactive information between metabolism and clinical text, and interactive information between structure and clinical text;

[0044] The structural and metabolic interaction information is shown below:

[0045]

[0046] in, Represents the features after the fusion of structural and metabolic information

[0047] Metabolic and clinical text interaction information is as follows:

[0048]

[0049] in, Represents the features after the fusion of metabolic information and clinical text information.

[0050] The interaction between the structure and clinical text is shown below:

[0051]

[0052] in, It represents the features after the fusion of structural information and clinical text information.

[0053] Further, the linear classification model is as follows:

[0054]

[0055] Among them, σ(·) is the Sigmoid activation function, and is the weight matrix, and is the bias term.

[0056] Furthermore, the input of the large language model is abnormality detection results and multimodal feature information, and the output is a diagnosis report; the abnormality detection results include MRI abnormalities, PET abnormalities, and patient clinical information; the patient clinical information includes age, gender, and cognitive assessment scores; the diagnosis report includes cognitive status assessment, imaging abnormality description, and diagnosis conclusion.

[0057] Furthermore, the diagnosis conclusion is normal cognition, mild cognitive impairment or Alzheimer's disease.

[0058] The technical effect of the present invention is unquestionable. The present invention performs multimodal data fusion diagnosis for different disease states of Alzheimer's disease (including Alzheimer's disease AD, mild cognitive impairment MCI and cognitive normal CN). Compared with the traditional method that only relies on a single data modality (such as MRI images or PET images), the present invention uses MRI, PET images and clinical text information (including demographic data and neuropsychological assessment results) for organic integration to provide a more comprehensive and detailed information representation for the diagnostic model. The present invention has outstanding advantages in improving the accuracy of early diagnosis of Alzheimer's disease, enhancing model robustness, and improving clinical practicality. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a flow chart of the present invention;

[0060] Figure 2 It is a model structure diagram of the present invention;

[0061] Figure 3 It is a schematic diagram of the structure of the unified multi-modal mechanism of the present invention. DETAILED DESCRIPTION

[0062] The present invention is further described below in conjunction with the embodiments, but it should not be understood that the above subject matter of the present invention is limited to the following embodiments. Without departing from the above technical ideas of the present invention, various substitutions and changes are made according to the common technical knowledge and customary means in the art, which should all be included in the protection scope of the present invention.

[0063] Embodiment 1:

[0064] See also Figures 1 to 3 , a computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal mechanism visual language model, including a multimodal data input module, a feature extraction and projection module, an abnormal marking token integration module, a unified multimodal attention mechanism module, an abnormality detection module, and a diagnosis result generation module;

[0065] The multimodal data input module acquires the multimodal data and transmits it to the feature extraction and projection module;

[0066] The multimodal data includes textual modal data and visual modal data; the visual modal data includes structural image data and metabolic image data;

[0067] The feature extraction and projection module extracts features from the multimodal data to obtain text features and initial visual features;

[0068] The feature extraction and projection module transmits the text features to the unified multimodal attention mechanism module;

[0069] The feature extraction and projection module projects the initial visual features into the shared embedding space to obtain a visual feature vector;

[0070] The abnormal mark token integration module concatenates the abnormal mark token with the visual feature to obtain the enhanced visual feature, and transmits it to the unified multimodal attention mechanism module;

[0071] The unified multimodal attention mechanism module performs cross-attention processing on text features and enhanced visual features to obtain multimodal fusion features, and transmits them to the anomaly detection module and the diagnosis result generation module;

[0072] The anomaly detection module stores a linear classification model;

[0073] The anomaly detection module inputs the multimodal fusion features into the linear classification module, generates anomaly detection results of visual modality data, and transmits them to the diagnosis result generation module;

[0074] The diagnosis result generation module stores a large language model;

[0075] The diagnosis result generation module will be input into the large language model to obtain a diagnosis report, including cognitive status assessment, imaging abnormality description and diagnosis conclusion.

[0076] The text modality data is clinical text data, including demographic information and neuropsychological assessment results of the patient;

[0077] Said demographic information includes age and gender;

[0078] The neuropsychological assessment results include the Mini-Mental State Examination results and the Clinical Dementia Score;

[0079] The structural image data is a brain structural image, obtained by structural magnetic resonance imaging;

[0080] The metabolic image data is a metabolic activity image, which is obtained by positron emission tomography.

[0081] The feature extraction and projection module uses the pre-trained CLIP visual encoder ViT-L / 14 to extract features from the visual modality data in the multimodal data and obtains the feature vector z i ,Right now:

[0082]

[0083] Among them, g v (·) represents the feature extraction function of the visual encoder, An indexed collection of visual modalities.

[0084] The feature extraction and projection module uses the text embedding function g t (·) Extract features from text modal data in multimodal data to obtain feature vector h i ,Right now:

[0085]

[0086] in, An index collection of text modal.

[0087] The visual feature vector is as follows:

[0088] The feature vectors z of all visual modalities i Projection into a shared embedding space

[0089]

[0090] in, is the projection matrix;

[0091] The enhanced visual features are as follows:

[0092]

[0093] in, Represents a splicing operation, and is the enhanced feature vector;

[0094] Used to capture structural abnormalities in structural magnetic resonance (sMRI) modalities, Used to capture metabolic abnormalities in the positron emission tomography (PET) modality; the structural magnetic resonance (sMRI) modality is used to reflect brain structure; the positron emission tomography (PET) modality is used to reflect brain metabolic activity. Multimodal data includes structural magnetic resonance (sMRI) modality and structural magnetic resonance (sMRI) modality.

[0095] The multimodal fusion features include structure-metabolism interaction information, metabolism-clinical text interaction information, and structure-clinical text interaction information;

[0096] The structural and metabolic interaction information is shown below:

[0097]

[0098] in, Represents the features after the fusion of structural and metabolic information

[0099] Metabolic and clinical text interaction information is as follows:

[0100]

[0101] in, Represents the features after the fusion of metabolic information and clinical text information.

[0102] The interaction between the structure and clinical text is shown below:

[0103]

[0104] in, It represents the features after the fusion of structural information and clinical text information.

[0105] The linear classification model looks like this:

[0106]

[0107] Among them, σ(·) is the Sigmoid activation function, and is the weight matrix, and is the bias term.

[0108] The input of the large language model is abnormality detection results and multimodal feature information, and the output is a diagnosis report; the abnormality detection results include MRI abnormalities, PET abnormalities, and patient clinical information; the patient clinical information includes age, gender, and cognitive assessment score; the diagnosis report includes cognitive status assessment, imaging abnormality description and diagnosis conclusion.

[0109] The diagnosis conclusion is normal cognition, mild cognitive impairment or Alzheimer's disease.

[0110] Embodiment 2:

[0111] A computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal mechanism visual language model, comprising a multimodal data input module, a feature extraction and projection module, an abnormal marking token integration module, a unified multimodal attention mechanism module, an abnormality detection module, and a diagnosis result generation module;

[0112] The multimodal data input module acquires the multimodal data and transmits it to the feature extraction and projection module;

[0113] The multimodal data includes textual modal data and visual modal data; the visual modal data includes structural image data and metabolic image data;

[0114] The feature extraction and projection module extracts features from the multimodal data to obtain text features and initial visual features;

[0115] The feature extraction and projection module transmits the text features to the unified multimodal attention mechanism module;

[0116] The feature extraction and projection module projects the initial visual features into the shared embedding space to obtain a visual feature vector;

[0117] The abnormal mark token integration module concatenates the abnormal mark token with the visual feature to obtain the enhanced visual feature, and transmits it to the unified multimodal attention mechanism module;

[0118] The unified multimodal attention mechanism module performs cross-attention processing on text features and enhanced visual features to obtain multimodal fusion features, and transmits them to the anomaly detection module and the diagnosis result generation module;

[0119] The anomaly detection module stores a linear classification model;

[0120] The anomaly detection module inputs the multimodal fusion features into the linear classification module, generates anomaly detection results of visual modality data, and transmits them to the diagnosis result generation module;

[0121] The diagnosis result generation module stores a large language model;

[0122] The diagnosis result generation module will be input into the large language model to obtain a diagnosis report, including cognitive status assessment, imaging abnormality description and diagnosis conclusion.

[0123] Embodiment 3:

[0124] A computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model, the technical content of which is the same as that of Example 2, and further, the text modality data is clinical text data, including demographic information and neuropsychological assessment results of the patient;

[0125] Said demographic information includes age and gender;

[0126] The neuropsychological assessment results include the Mini-Mental State Examination results and the Clinical Dementia Score;

[0127] The structural image data is a brain structural image, obtained by structural magnetic resonance imaging;

[0128] The metabolic image data is a metabolic activity image, which is obtained by positron emission tomography.

[0129] Embodiment 4:

[0130] A computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal mechanism visual language model, the technical content of which is the same as any one of Embodiments 2-3, and further, the feature extraction and projection module uses the pre-trained CLIP visual encoder ViT-L / 14 to extract features from the visual modality data in the multimodal data to obtain a feature vector z i ,Right now:

[0131]

[0132] Among them, g v (·) represents the feature extraction function of the visual encoder, An indexed collection of visual modalities.

[0133] Embodiment 5:

[0134] A computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model, the technical content of which is the same as any one of Embodiments 2-4, further, the feature extraction and projection module utilizes a text embedding function g t (·) Extract features from text modal data in multimodal data to obtain feature vector h i ,Right now:

[0135]

[0136] in, An index collection of text modal.

[0137] Embodiment 6:

[0138] A computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal mechanism visual language model, the technical content is the same as any one of embodiments 2-5, and further, the visual feature vector is as follows:

[0139] The feature vectors z of all visual modalities i Projection into a shared embedding space

[0140]

[0141] in, is the projection matrix;

[0142] Embodiment 7:

[0143] A computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model, the technical content of which is the same as any one of Embodiments 2-6, and further, the enhanced visual features are as follows:

[0144]

[0145] in, Represents a splicing operation, and is the enhanced feature vector;

[0146] Used to capture structural abnormalities in structural magnetic resonance (sMRI) modalities, Used to capture metabolic abnormalities in the positron emission tomography (PET) modality; the structural magnetic resonance (sMRI) modality is used to reflect brain structure; the positron emission tomography (PET) modality is used to reflect brain metabolic activity.

[0147] Embodiment 8:

[0148] A computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal mechanism visual language model, the technical content of which is the same as any one of Embodiments 2-7, and further, the multimodal fusion features include structure and metabolism interaction information, metabolism and clinical text interaction information, and structure and clinical text interaction information;

[0149] The structural and metabolic interaction information is shown below:

[0150]

[0151] in, Represents the features after the fusion of structural and metabolic information

[0152] Metabolic and clinical text interaction information is as follows:

[0153]

[0154] in, Represents the features after the fusion of metabolic information and clinical text information.

[0155] The interaction between the structure and clinical text is shown below:

[0156]

[0157] in, It represents the features after the fusion of structural information and clinical text information.

[0158] Embodiment 9:

[0159] A computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model, the technical content of which is the same as any one of Embodiments 2-8, and further, the linear classification model is as follows:

[0160]

[0161] Among them, σ(·) is the Sigmoid activation function, and is the weight matrix, and is the bias term.

[0162] Embodiment 10:

[0163] A computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model, the technical content of which is the same as any one of Examples 2-9, and further, the input of the large language model is abnormality detection results and multimodal feature information, and the output is a diagnosis report; the abnormality detection results include MRI abnormalities, PET abnormalities, and patient clinical information; the patient clinical information includes age, gender, and cognitive assessment scores; the diagnosis report includes cognitive status assessment, imaging abnormality description, and diagnosis conclusion.

[0164] Embodiment 11:

[0165] A computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model, the technical content of which is the same as any one of Examples 2-10, and further, the diagnosis conclusion is normal cognition, mild cognitive impairment or Alzheimer's disease.

[0166] Embodiment 12:

[0167] A computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model, the technical content of which is the same as any one of Embodiments 2-11, and further includes a visualization interface;

[0168] The visualization interface is used to visualize the diagnosis report.

[0169] Embodiment 13:

[0170] A computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal mechanism visual language model includes a multimodal data input module, a feature extraction and projection module, an abnormal marking token integration module, a unified multimodal attention mechanism module, and a diagnosis result generation module.

[0171] The implementation steps of the Alzheimer's disease computer-aided diagnosis system are as follows:

[0172] Step 1: Multimodal data input

[0173] The model receives data in different modalities, including:

[0174] Clinical text data: patient demographic information and neuropsychological assessment results.

[0175] Image Data:

[0176] Structural data: Brain structure images obtained through structural magnetic resonance imaging (sMRI).

[0177] Metabolic data: Metabolic activity images obtained by positron emission tomography (PET).

[0178] The method of the present invention obtains multimodal data based on publicly available data sets, including the Alzheimer's Disease Neuroimaging Initiative (ADNI), the Open Imaging Study Series (OASIS), and the National Alzheimer's Disease Coordinating Center (NACC). The data is preprocessed as follows:

[0179] Structured clinical text data: Clinical text data includes demographic information and neuropsychological assessment results. Demographic information includes age and gender; neuropsychological assessment includes Mini-Mental State Examination (MMSE) and Clinical Dementia Rating (CDR). In order to make full use of these non-imaging data, this method converts them into text form. Specifically: the non-imaging data of each sample is converted into text. For example, if a patient's MMSE score is 29, the corresponding text expression is: "The mini-mental state examination score is 29 out of 30."

[0180] Select slice data: Since the image data is three-dimensional data, this method selects the middle slice of each patient's coronal position for analysis.

[0181] Symbolically represented as input data set x = (x MRI ,x PET ,x clinical ), where x MRI is the MRI image, x PET is the PET image, x clinical Contains clinical text information.

[0182] Step 2: Feature extraction projection

[0183] For visual modality data (MRI and PET images), the pre-trained CLIP visual encoder ViT-L / 14 is used for feature extraction to obtain the feature vector z i :

[0184]

[0185] Among them, g v (·) represents the feature extraction function of the visual encoder, An indexed collection of visual modalities.

[0186] For text modality data, use the text embedding function g t (·) Process the clinical text to obtain the feature vector h i :

[0187]

[0188] in, An index collection of text modal.

[0189] The feature vectors z of all visual modalities i Projection into a shared embedding space

[0190]

[0191] in, The projection matrix is ​​used in the projection process. Convolution extraction modules, including ResNet modules and adaptive average pooling layers, are used to achieve effective local context modeling.

[0192] Step 3: Abnormal Marking Token Integration

[0193] In order to enhance the model's attention to key areas, a learnable anomaly marker token is introduced. Define a set of anomaly marker tokens Among them I tokens is the token index set corresponding to different modalities. Specifically, Used to capture structural abnormalities in sMRI modalities, Used to capture metabolic abnormalities in PET modalities.

[0194] Concatenate the anomaly marker token with the feature vector of the corresponding modality:

[0195]

[0196] in, Represents a splicing operation, and It is the enhanced feature vector, which contains abnormal label information.

[0197] Step 4: Unify the multimodal attention mechanism

[0198] like Figure 3 As shown in the figure, a unified multimodal attention mechanism is used to perform cross-attention processing on the integrated features to achieve comprehensive fusion between different modalities. Specifically, three cross attentions (CA) are used to capture the interactive information between the two modalities:

[0199] Structural and metabolic interaction information:

[0200]

[0201] in, Represents the features after the fusion of structural and metabolic information

[0202] Metabolic information and clinical text interaction information:

[0203]

[0204] in, Represents the features after the fusion of metabolic information and clinical text information.

[0205] Interaction between structural information and clinical text:

[0206]

[0207] in, It represents the features after the fusion of structural information and clinical text information.

[0208] The above-mentioned cross-attention mechanism ensures the comprehensive integration of structural, metabolic, and clinical information, providing rich feature representation for subsequent diagnosis.

[0209] Step 5: Anomaly Detection and Classification

[0210] The output features of the cross-attention mechanism are input into the linear classification layer for anomaly detection:

[0211]

[0212] Among them, σ(·) is the Sigmoid activation function, and is the weight matrix, and is the bias term. Through the above classification, it is judged whether the sMRI and PET images are abnormal.

[0213] Step 6: Generate final diagnosis results

[0214] The results of anomaly detection and multimodal feature information are input into a large language model (LLM) to generate a comprehensive diagnosis conclusion. The specific steps include:

[0215] The abnormal detection results of each modality (such as abnormal MRI display, abnormal PET display, etc.) and the patient's clinical information (age, gender, cognitive assessment score, etc.) are used as input.

[0216] The large language model comprehensively considers all input information and generates a detailed diagnostic report, including cognitive status assessment, description of imaging abnormalities, and final diagnosis conclusion (normal cognition, mild cognitive impairment, or Alzheimer's disease).

[0217] Embodiment 14:

[0218] Validation of a computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model, the contents are as follows:

[0219] In the experiment, the widely used ADNI public dataset was selected as the main evaluation platform, and the generalization ability of the model was further tested on the OASIS and NACC datasets. To evaluate the performance of the method, the traditional indicators of accuracy, recall and F1-Score were used to measure the performance of the model.

[0220] Robust performance across datasets

[0221] In order to test the robustness and extensibility of the method of the present invention, after completing the model training on the ADNI dataset, the present invention tested it on three datasets from different sources: ADNI, OASIS and NACC. These three datasets differ in data distribution and subject population characteristics. Table 1 summarizes the test results of AD and CN, AD and MCI, MCI and CN, and AD vs. MCI vs. CN multi-classification tasks on each dataset.

[0222] Table 1: Comparison of diagnostic performance of the method of the present invention on the ADNI, OASIS, and NACC datasets (indicators: Accuracy, Recall, F1-Score)

[0223]

[0224]

[0225] It can be seen that the method of the present invention achieves a high accuracy of 0.9919 in the AD vs. CN task on the ADNI dataset, and also reaches 0.9450 and 0.9102 on the OASIS and NACC datasets, respectively, indicating that the method has extremely high accuracy in identifying Alzheimer's patients and normal individuals. At the same time, in the more challenging AD vs. MCI and MCI vs. CN tasks, the method of the present invention can still achieve a relatively high accuracy (such as the accuracy of 0.9122 for the AD vs. MCI task on the ADNI dataset), and the performance on OASIS and NACC is only slightly reduced. This shows that even in the face of different data distributions and population characteristics, the method of the present invention can still robustly classify the disease level accurately and has good cross-dataset generalization capabilities.

[0226] Effects of different modality combinations on diagnostic performance

[0227] To further explore the role and contribution of each data modality (MRI, PET, clinical assessment information), the present invention conducts comparative experiments on multiple modality configurations on the ADNI dataset. Table 2 shows the performance changes when using only MRI, only PET, and combining clinical assessment information with one or more of the imaging modalities.

[0228] As can be seen in Table 2, when only a single modality is used (such as MRI only or PET only), the diagnostic performance is relatively low. When two modalities are combined (such as MRI+PET or MRI+clinical assessment), the performance is improved, but it is still not as good as the improvement brought by integrating all MRI, PET and clinical assessment data. The comprehensive integration of multimodal data can enable the model to show stronger discrimination ability and robustness in complex AD grading and diagnosis tasks.

[0229] Table 2: Performance of different modality combinations on the ADNI test set (indicator: Accuracy)

[0230]

[0231] Note: Table 2 shows typical results of each modality combination. It can be seen that when MRI, PET and clinical evaluation data are integrated simultaneously, the accuracy of the method of the present invention is as high as 0.9919 in the AD vs. CN task, and it also achieves significant improvements in the MCI vs. CN and AD vs. MCI vs. CN tasks.

[0232] Based on the above experimental results and analysis, the present invention organically integrates MRI, PET and clinical evaluation data, and introduces abnormal labeling tokens and a unified multimodal attention mechanism. It not only achieves a high accuracy rate close to 100% in the classic AD vs. CN task, but also maintains a high level of performance in the more challenging AD vs. MCI and multi-classification diagnosis tasks. At the same time, the present invention has excellent generalization capabilities in different data sets (ADNI, OASIS, NACC), and can maintain stable performance under diverse populations and different data characteristics. The realization of these technical effects fully demonstrates the outstanding advantages of the method of the present invention in improving the accuracy of early diagnosis of Alzheimer's disease, enhancing model robustness, and improving clinical practicality.

Claims

1. A computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model, characterized by: It includes multimodal data input module, feature extraction and projection module, abnormal labeling token integration module, unified multimodal attention mechanism module, abnormal detection module, and diagnosis result generation module; The multimodal data input module acquires the multimodal data and transmits it to the feature extraction and projection module; The multimodal data includes textual modal data and visual modal data; the visual modal data includes structural image data and metabolic image data; The feature extraction and projection module extracts features from the multimodal data to obtain text features and initial visual features; The feature extraction and projection module transmits the text features to the unified multimodal attention mechanism module; The feature extraction and projection module projects the initial visual features into the shared embedding space to obtain a visual feature vector; The abnormal mark token integration module concatenates the abnormal mark token with the visual feature to obtain the enhanced visual feature, and transmits it to the unified multimodal attention mechanism module; The unified multimodal attention mechanism module performs cross-attention processing on text features and enhanced visual features to obtain multimodal fusion features, and transmits them to the anomaly detection module and the diagnosis result generation module; The anomaly detection module stores a linear classification model; The anomaly detection module inputs the multimodal fusion features into the linear classification module, generates anomaly detection results of visual modality data, and transmits them to the diagnosis result generation module; The diagnosis result generation module stores a large language model; The diagnosis result generation module will be input into the large language model to obtain a diagnosis report, including cognitive status assessment, imaging abnormality description and diagnosis conclusion.

2. The computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model according to claim 1, characterized in that: The text modality data is clinical text data, including demographic information and neuropsychological assessment results of the patient; Said demographic information includes age and gender; The neuropsychological assessment results include the Mini-Mental State Examination results and the Clinical Dementia Scale; The structural image data is a brain structural image, obtained by structural magnetic resonance imaging; The metabolic image data is a metabolic activity image, which is obtained by positron emission tomography.

3. The computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model according to claim 1, characterized in that: The feature extraction and projection module uses the pre-trained CLIP visual encoder ViT-L / 14 to extract features from the visual modality data in the multimodal data and obtains the feature vector z i ,Right now: Among them, g v (·) represents the feature extraction function of the visual encoder, is the index set of visual modalities; 4. The computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model according to claim 1, characterized in that: The feature extraction and projection module uses the text embedding function g t (·) Extract features from text modal data in multimodal data to obtain feature vector h i ,Right now: in, An index collection of text modal.

5. The computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model according to claim 1, characterized in that: The visual feature vector is shown below: The feature vectors z of all visual modalities i Projection into a shared embedding space in, is the projection matrix; 6. The computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model according to claim 1, characterized in that: The enhanced visual features are as follows: in, Represents a splicing operation, and is the enhanced feature vector; Used to capture structural abnormalities in structural magnetic resonance (sMRI) modalities, Used to capture metabolic abnormalities in the positron emission tomography (PET) modality; the structural magnetic resonance (sMRI) modality is used to reflect brain structure; the positron emission tomography (PET) modality is used to reflect brain metabolic activity.

7. The computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model according to claim 1, characterized in that: The multimodal fusion features include structure-metabolism interaction information, metabolism-clinical text interaction information, and structure-clinical text interaction information; The structural and metabolic interaction information is shown below: in, Represents the features after the fusion of structural and metabolic information Metabolic and clinical text interaction information is as follows: in, Represents the features after the fusion of metabolic information and clinical text information. The interaction between the structure and clinical text is shown below: in, It represents the features after the fusion of structural information and clinical text information.

8. The Alzheimer's disease computer-aided diagnosis system based on a unified multimodal visual language model according to claim 1, characterized in that: The linear classification model looks like this: Among them, σ(·) is the Sigmoid activation function, and is the weight matrix, and is the bias term.

9. The computer-aided diagnosis system for Alzheimer's disease based on a unified multimodal visual language model according to claim 1, characterized in that: The input of the large language model is abnormality detection results and multimodal feature information, and the output is a diagnosis report; the abnormality detection results include MRI abnormalities, PET abnormalities, and patient clinical information; the patient clinical information includes age, gender, and cognitive assessment score; the diagnosis report includes cognitive status assessment, imaging abnormality description and diagnosis conclusion.

10. The Alzheimer's disease computer-aided diagnosis system based on a unified multimodal visual language model according to claim 1, characterized in that: The diagnosis conclusion is normal cognition, mild cognitive impairment or Alzheimer's disease.

Citation Information

Patent Citations

  • Alzheimer disease diagnosis method based on multi-modal cross attention

    CN118116573A

  • Brain disease early-stage intelligent grading screening system based on multi-modal calculation

    CN118248318A

  • GFE-Mama neural network-based interpretable Alzheimer's progress classification method

    CN118918363A

  • Alzheimer disease risk early warning method, device and system based on handwriting recognition

    CN119230103A

  • Method for Allocating Power in SWIPT-Enabled NOMA System in Distributed Antenna System and Apparatus thereof

    KR1020250127438A

Cited By

  • Alzheimer's disease preclinical risk quantitative evaluation method and system

    CN121460166A