An alzheimer's disease auxiliary precision identification method and system based on a multi-modal marker network

Through multimodal marker networks and adaptive support vector machine models, accurate differentiation of Alzheimer's disease from other types of dementia is achieved, solving the problems of low efficiency of multimodal data fusion and insufficient model adaptability, and providing highly accurate and interpretable diagnostic results.

CN120340837BActive Publication Date: 2025-10-14ZHEJIANG GEWUZHIZHI BIOTECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510817176.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-14
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Existing technologies in the diagnosis of Alzheimer's disease have problems such as low efficiency of multimodal data fusion, insufficient model adaptability, poor specificity and insufficient interpretability, making it difficult to accurately distinguish Alzheimer's disease from other types of dementia.

Method used

A multimodal marker network-based method is adopted to obtain blood, urine, imaging, electrophysiological and clinical data, use the self-attention mechanism for cross-modal interactive learning, construct a weighted gene co-expression network, and combine it with an adaptive support vector machine model for dynamic classification to output explainable diagnostic results.

Benefits of technology

It significantly improves the accuracy and specificity of Alzheimer's disease diagnosis, can dynamically adapt to new data, provide explainable diagnostic basis, and support early diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340837B_ABST
    Figure CN120340837B_ABST
Patent Text Reader

Abstract

The application provides an Alzheimer's disease auxiliary precision identification method and system based on a multi-omics marker network and an adaptive support vector machine. Through a marker network based on five modal data and an adaptive support vector machine model, the precise identification of Alzheimer's disease and other types of dementia is realized. The method fuses multi-source heterogeneous data such as blood, urine, neuroimaging, electrophysiology and clinical evaluation, adopts weighted gene co-expression network analysis to screen specific markers, and realizes cross-modal feature interaction learning through a self-attention mechanism. The online learning mechanism is introduced to enable the model to adapt to new data distribution, and the explainability module is combined to output the contribution of each marker to the diagnosis. Finally, with the help of a cloud platform, multi-center data synchronization and model optimization are realized, which significantly improves the accuracy, specificity, early diagnosis ability and adaptability of the AD diagnosis and identification model, overcomes the limitations of the prior art, and assists doctors in diagnosing and treating Alzheimer's disease.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of information technology, and in particular to an Alzheimer's disease auxiliary precise identification method and system based on a multi-modal marker network. BACKGROUND

[0002] Alzheimer's Disease (AD) is the most common neurodegenerative disease and dementia type, which brings a heavy burden to global public health. Clinically, early diagnosis of AD and precise identification of other dementia subtypes (such as vascular dementia VaD, Lewy body dementia DLB, frontotemporal dementia FTD, etc.) still face great challenges. Traditional diagnostic methods rely on clinical assessment, neuropsychological testing, and structural imaging, but often lack sensitivity and specificity in the early stages of the disease, and it is difficult to distinguish different dementia etiologies with similar clinical symptoms.

[0003] In recent years, biomarkers play an increasingly important role in AD diagnosis. Aβ42, total tau (t-tau), and phosphorylated tau (p-tau) proteins in cerebrospinal fluid (CSF) are classic core markers of AD, but CSF acquisition is invasive. Blood biomarkers are of great interest due to their convenience and minimally invasive nature, such as phosphorylated tau proteins (e.g. p-tau181, p-tau217, p-tau231) and neurofilament light chain (NfL), which show potential. However, existing studies (such as JAMA Neurol.2024) have shown that even better single or a few markers (such as Aβ42 / p-tau combination or p-tau217 alone, reference CN118233456A uses only cerebrospinal fluid p-tau217) still need to improve the specificity in distinguishing AD from other dementias (especially mixed pathology), which is usually between 78-85%, making it difficult to meet the needs of clinical precise diagnosis.

[0004] At the same time, neuroimaging (such as MRI) provides information on brain structure, and electroencephalogram (EEG) reflects brain function status. Integrating these multi-modal data is expected to improve diagnostic performance. However, the fusion of multi-modal data faces challenges: 1) strong data heterogeneity, with different dimensions, scales, and types; 2) traditional machine learning methods (such as logistic regression) are prone to "dimension disaster" when dealing with high-dimensional, multi-modal data, and it is difficult to capture the complex interaction between modalities; 3) existing models are mostly static models, lacking the ability to dynamically adjust and optimize based on new data; 4) some existing technologies (such as US20240081234A1) attempt to combine imaging and biomarkers, but fail to achieve deep, adaptive fusion, and lack effective representation of the dynamic process of AD pathological cascade.

[0005] Therefore, there is an urgent need for a new method and system for accurately identifying AD that can effectively integrate multiple sources (blood, urine, images, electrophysiology, and clinical), deeply mine data value using advanced artificial intelligence technology, have high accuracy and specificity, and dynamically adapt to new data to solve the bottlenecks of existing technologies in terms of identification accuracy, multi-modal data fusion efficiency, and model adaptability.

[0006] Although the prior art has attempted to fuse multi-modal data for dementia diagnosis, for example, some studies use simple feature splicing or basic machine learning models to process partial modal data, but often have the following limitations: 1) the fusion mechanism is simple and cannot effectively capture the deep nonlinear relationships between blood, urine, images, electrophysiology, and clinical assessment of the five types of specific heterogeneous data; 2) there is a lack of specific marker network construction methods for accurately identifying AD and other dementias (such as vascular dementia VaD and Lewy body dementia DLB); 3) static models are used, which cannot dynamically adapt to changes in disease progression or the distribution of new data, resulting in insufficient model generalization ability and long-term stability; 4) most models are "black boxes" and lack explainability for diagnostic results, making it difficult to gain clinical trust and effective guidance. Therefore, how to build an auxiliary accurate identification system that can deeply integrate specific multi-modal data, dynamically adapt to data changes, and provide explainable evidence is a technical problem that needs to be solved in this field. SUMMARY

[0007] To address the deficiencies of the prior art in integrating specific multi-modal data, model adaptability, and explainability, the present application provides an Alzheimer's disease auxiliary accurate identification method based on a multi-modal marker network, which includes: obtaining multi-modal data, the multi-modal data being specific five types of modal data including blood, urine biomarkers, neural images, electroencephalography, and clinical assessment; extracting features and deeply fusing the multi-modal data through a specific multi-level processing mechanism including independent feature extraction, cross-modal interaction learning based on attention mechanism, and network optimization, to construct a marker network representing the complex relationships between modalities; analyzing the marker network through an adaptive classification model with online learning ability to dynamically determine the identification results of Alzheimer's disease and other dementia types; and outputting the identification results combined with explainability analysis, the identification results not only representing the probability of the presence of Alzheimer's disease, but also quantifying the contribution of key features.

[0008] Further, the acquiring the multi-modal data comprises: acquiring first biomarker data from a blood sample, the first biomarker data comprising phosphorylated protein and micro ribonucleic acid data; acquiring second biomarker data from a urine sample, the second biomarker data comprising specific protein data; acquiring third image data by an imaging device, the third image data comprising brain structure features; acquiring fourth electrophysiological data by an electrophysiological device, the fourth electrophysiological data comprising electroencephalogram features; acquiring fifth clinical data from a clinical assessment, the fifth clinical data comprising cognitive function scores; and integrating the first biomarker data, the second biomarker data, the third image data, the fourth electrophysiological data, and the fifth clinical data into the multi-modal data.

[0009] Further, the feature extraction and fusion of the multi-modal data comprises: performing independent feature extraction on each type of data in the multi-modal data to obtain an initial feature set of each modality; performing cross-modality interaction learning on the initial feature set by an attention mechanism to generate a fusion feature set; constructing a multi-dimensional marker network according to the fusion feature set, the multi-dimensional marker network reflecting the correlation between features of each modality; and performing optimization processing on the multi-dimensional marker network to obtain a final feature representation for classification.

[0010] Further, the analyzing the marker network by a preset classification model comprises: inputting the marker network into an adaptive classification model, the adaptive classification model comprising a multi-level feature processing architecture; performing independent analysis on features of each modality by a bottom structure of the adaptive classification model to obtain a preliminary classification basis; performing interactive integration on the preliminary classification basis by a middle structure of the adaptive classification model to generate a comprehensive classification basis; and performing final classification on the comprehensive classification basis by a top structure of the adaptive classification model to determine the identification result.

[0011] Further, the outputting the identification result comprises: generating a probability value according to the identification result, the probability value representing the likelihood of Alzheimer's disease; performing feature contribution degree evaluation on the identification result by an explainability analysis module to generate a feature importance distribution; generating a diagnosis report according to the probability value and the feature importance distribution, the diagnosis report comprising a classification conclusion and key feature basis; and transmitting the diagnosis report to a terminal device through a preset interface for subsequent analysis and reference.

[0012] Further, the feature extraction and fusion on the multi-modal data further comprises: for a low sample amount data group in the multi-modal data, a data enhancement technique is used for sample expansion to generate an enhanced data set; the enhanced data set is subjected to feature alignment processing, and a regularization method is used to reduce cross-modal feature differences; the feature alignment processing is used to update the marker network to obtain an optimized feature representation; and the optimized feature representation is input into a subsequent classification process to improve classification accuracy.

[0013] Further, the analysis of the marker network by the preset classification model further comprises: triggering a dynamic update mechanism of the preset classification model according to new data to obtain an updated classification model; reanalyzing the marker network by using the updated classification model to generate an updated discrimination result; verifying the updated discrimination result, and using a cross-validation method to evaluate model stability; and adjusting parameters of the preset classification model according to a verification result to optimize classification performance.

[0014] In another aspect, the application further provides an Alzheimer's disease auxiliary precise discrimination system for implementing the above method, and the system comprises:

[0015] A data acquisition interface module is configured to acquire multi-modal data, and the multi-modal data at least includes blood biomarker data, urine biomarker data, image data, electrophysiological data and clinical data.

[0016] A data preprocessing module is configured to implement standardization, correction, denoising and registration preprocessing functions on original data of each modality.

[0017] A feature extraction and fusion unit is configured to extract and fuse features of the multi-modal data, and a marker network is constructed by using a multi-level processing mechanism based on attention mechanism cross-modal interaction learning, and the marker network represents the relevance between different modal features.

[0018] An adaptive classification and discrimination unit is configured to analyze the marker network by using an adaptive classification model capable of dynamic update to determine a discrimination result of Alzheimer's disease and other types of dementia.

[0019] A result output and visualization module is configured to output the discrimination result, and the discrimination result represents a probability of existence of Alzheimer's disease and is displayed to a user in a user-friendly manner.

[0020] The technical scheme provided by the embodiments of the application can have the following beneficial effects:

[0021] The application provides an Alzheimer's disease auxiliary precise identification method and system based on a multi-omics marker network and an adaptive support vector machine. Through a marker network based on five modal data and an adaptive support vector machine model, precise identification of Alzheimer's disease and other types of dementia is realized. The method fuses multi-source heterogeneous data such as blood, urine, neural image, electrophysiology and clinical evaluation, adopts weighted gene co-expression network analysis to screen specific markers, and realizes cross-modal feature interaction learning through a self-attention mechanism. An online learning mechanism is introduced to enable the model to adapt to new data distribution, and an explainability module is combined to output the contribution of each marker to diagnosis. Finally, the cloud platform is used to realize multi-center data synchronization and model optimization, which significantly improves the accuracy, specificity, early diagnosis ability and adaptability of the AD diagnosis and identification model, overcomes the limitations of the prior art, and assists doctors in diagnosing and treating Alzheimer's disease. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 A flowchart of an Alzheimer's disease auxiliary precise identification method based on a multi-modal marker network of the application;

[0023] Figure 2 A model architecture diagram of an Alzheimer's disease auxiliary precise identification method based on a multi-modal marker network of the application;

[0024] Figure 3 A structure diagram of an Alzheimer's disease auxiliary precise identification system of the application. DETAILED DESCRIPTION

[0025] The application aims to solve the following problems in the prior art of Alzheimer's disease diagnosis and identification:

[0026] 1. The problem of insufficient specificity and accuracy of a single or small number of biomarkers in distinguishing AD from other types of dementia (especially VaD and DLB).

[0027] 2. The problem of difficulty of traditional machine learning methods in effectively fusing high-dimensional, heterogeneous multi-modal data (molecular, image, electrophysiology, clinical), and the existence of dimension disaster and information loss.

[0028] 3. The problem of lack of intelligent diagnostic models that can capture complex interaction relationships between different modal data and adaptively adjust according to data characteristics.

[0029] 4. The problem of lack of diagnostic models that can reflect the dynamic changes of AD pathology and have online learning ability, which is difficult to adapt to new data accumulated in clinical practice.

[0030] In order to make the purpose, technical scheme and advantages of the application clearer, the application will be described in detail below with reference to the drawings and specific embodiments.

[0031] Embodiment 1

[0032] As Figure 1 , the embodiment of the Alzheimer's disease auxiliary precision identification method based on a multi-modal marker network can specifically include:

[0033] Step S101, a multi-modal data sample set is obtained, the sample set includes blood biomarker data, urine biomarker data, neuroimaging data, electrophysiological data, and clinical assessment data, and each sample data corresponds to a specific dementia type label.

[0034] In a preferred embodiment, the multi-modal raw data is obtained through a standardized acquisition process, including 10 mL of venous blood drawn in the morning on an empty stomach to separate plasma, 10 mL of urine centrifuged to collect, 3.0T MRI to obtain T1 weighted images, 64 lead electrode cap to collect 5 minutes of closed eye resting state EEG, Montreal cognitive assessment directional force sub-item score, to obtain the original data set containing blood, urine, image, electrophysiology and clinical assessment. The original data set is cleaned using data preprocessing techniques, miRNA data is standardized using the internal reference gene GAPDH, brain extraction and bias field correction are performed using ANTs software, and independent component analysis is used to remove EEG eye artifact, to obtain the standardized multi-modal data set. The blood p-tau231 and miR-124-3p, urine UCH-L1 and Aβ1-42 are detected by sandwich chemiluminescence kit, the hippocampal CA1 and DG subregion atrophy rate is analyzed based on CNN, the EEG alpha power spectral density and theta / beta ratio are extracted, to obtain the quantified multi-modal biomarker feature set. The AD-specific module is constructed using weighted gene co-expression network analysis, the correlation between biomarkers and disease progression is verified by Cox proportional hazards model, and the key features of p-tau231, miR-124-3p, UCH-L1, Aβ1-42, hippocampal atrophy rate, EEG alpha power, etc. are screened out, to obtain the high confidence feature subset related to AD. If the feature subset contains low sample size dementia types such as DLB, the SMOTE technology is used for data enhancement with a 200% oversampling rate, to obtain a sample balanced multi-modal feature data set.

[0035] For example, the generation step is as follows:

[0036] The morning fasting venous blood 10 mL was collected, the plasma was separated by centrifugation at 3000g for 15 minutes, the urine sample was pretreated in a 10 mL centrifuge tube and stored at-80℃, the hippocampus T1 weighted image was obtained by 3.0T MRI with 1x1x1mm3 resolution, 64-lead EEG collected 5-minute closed-eye resting state signal, MoCA directional sub-score was recorded synchronously, and the original multi-modal data set was formed. The GAPDH was used to standardize the ACt value of miR-124-3p, the N4 bias field correction and brain tissue segmentation were performed by ANTs software, the ICA algorithm was used to remove the eye movement artifact components in the EEG signal, and the power spectrum in the 1-40Hz frequency band was retained. The sandwich chemiluminescence method was used to detect the concentration of plasma p-tau231 (threshold value 0.87 pg / mL), the Simoa platform was used to quantitatively detect urine Aβ1-42, the 3D-CNN was used to calculate the annual atrophy rate of hippocampus CA1 region (>0.5% as abnormal), and the FFT analysis was used to obtain the characteristic values of 30% decrease in α wave power and 40% increase in θ / β ratio.

[0037] For example, the weighted gene co-expression network WGCNA is used to construct a co-expression network to screen AD modules (p<0.001), the Cox model is used to verify that p-tau231 is associated with disease progression HR=1.7 / year, and the SMOTE oversampling is used for the DLB sample to increase the sample size by 200%. In the three-level fusion architecture, the ResNet18 is used in the bottom layer to extract image features, the self-attention mechanism is used in the middle layer to learn the interaction weight between biomarkers and EEG, and the SVM is used in the top layer to dynamically adjust the γ parameter (range 0.1-1.0). The RLDA algorithm is used to align the cross-modal feature space, the Shapley value calculation shows that the contribution of p-tau231 is 35%, and the feature importance heat map is generated. The incremental learning module monitors the data stream, triggers model iteration every 50 new samples, and the gRPC interface completes the edge computing platform deployment within 200ms. The hierarchical 5-fold cross-validation shows that the AUC=0.992 (AD vs VaD), the DeLong test p>0.05 confirms the stability of the model, and the Kruskal-Wallis analysis shows that the difference between the marker groups is p<0.001.

[0038] In step S102, according to the multi-modal data sample set, based on the marker network of the five kinds of modal data, the network covers the phosphorylated tau protein in blood, the specific protein in urine, the image feature based on convolutional neural network, the power spectral density of electrophysiological signal and the clinical evaluation score, the specific markers related to the disease are screened by weighted gene co-expression network analysis.

[0039] In a preferred embodiment, a multi-modal data sample set is obtained, containing phosphorylated tau protein p-tau231 and miR-124-3p in blood, UCH-L1 and Aβ1-42 in urine, 3.0T MRI T1 weighted image, electro-physiological signal EEG and Montreal Cognitive Assessment MoCA score. The modal data is cleaned and format aligned by a standardized preprocessing module to obtain a unified multi-modal data set. Convolutional neural network is used to process MRI image data, extract the atrophy rate characteristics of hippocampal CA1 and DG subregion, and complete brain extraction and bias field correction by ANTs software to obtain image feature vector. Independent component analysis is used to remove artifacts from EEG signal, extract α wave power spectral density and θ / β ratio in 1-40 Hz frequency band, and combine with resting state data collected by 64 electrode cap to obtain electro-physiological feature set. According to blood and urine sample data, sandwich chemiluminescence kit is used to detect p-tau231 and UCH-L1 concentration, and reverse transcription primer is used to quantify miR-124-3p expression level to obtain molecular marker feature matrix. Weighted gene co-expression network analysis is used to modularly cluster multi-modal features to construct AD specific marker network, and WGCNA algorithm is used to screen core features related to disease progression to obtain five-dimensional marker network.

[0040] Exemplarily, plasma is isolated from 10 mL venous blood samples collected in the morning on an empty stomach, and after centrifugation at 3000 g for 15 minutes, the concentration of p-tau231 is quantified using a sandwich chemiluminescence kit with a detection limit of 0.1 pg / mL, while the reverse transcription primer is used to determine the ΔCt value (threshold value 2.3) of miR-124-3p, and the urine samples are treated by centrifugation and stored at -80°C for no more than 6 months, and the contents of UCH-L1 and Aβ1-42 are detected by Simoa technology. The 3.0T MRI device acquires T1-weighted images with a resolution of 1x1x1 mm3, brain tissue extraction and field correction are performed by ANTs software, convolutional neural network is used to automatically segment the hippocampal CA1 and DG subregions, and the annual atrophy rate (threshold value >0.5%) is calculated. A 64-electrode cap is used to collect 5-minute closed-eye resting-state EEG signals, and after removing the eye movement artifact by independent component analysis, the alpha power spectral density (decrease threshold of 30%) and theta / beta ratio (increase threshold of 40%) are extracted. The Montreal Cognitive Assessment orientation sub-score is used as the critical value of 8 points to enter the system. Weighted gene co-expression network analysis screens AD-specific modules containing sphingomyelin SM(d18:1 / 16:0), and Cox proportional hazards model verifies the correlation of the marker with disease progression (p<0.001). The self-attention mechanism processes the interaction of multi-modal features with 5 attention heads, and the dynamic kernel function SVM automatically adjusts the gamma parameter according to the new data, and triggers incremental learning every 50 cases. The Shapley value algorithm outputs a feature contribution heat map, showing that p-tau231 (OR=5.8) and hippocampal CA1 atrophy (HR=1.7) are key discriminators.

[0041] In step S103, data in the five-dimensional marker network is used to independently extract features from each modality data, to obtain a plurality of modality feature subsets, the feature subsets including key diagnostic indicators of each modality data, and cross-modality feature alignment is realized by regularized linear discriminant analysis.

[0042] In a preferred embodiment, by using the data in the five-dimensional marker network, the blood, urine, neuroimaging, electrophysiology and clinical data are standardized by the preprocessing module to obtain the normalized data set of each modality. According to the normalized data set, the convolutional neural network is used to extract the features of the neuroimaging data to obtain the hippocampal subregion atrophy rate feature subset. By using the sandwich chemiluminescence kit and reverse transcription quantitative PCR technology, the molecular markers in the blood and urine are quantitatively analyzed to obtain the biomarker feature subset such as p-tau231, miR-124-3p and UCH-L1. The independent component analysis is used to remove the artifacts and extract the frequency band power of the resting state EEG signal to obtain the alpha wave power spectral density and theta / beta ratio feature subset. According to the clinical data, the directional force sub-item score of the Montreal Cognitive Assessment scale is digitized to obtain the clinical marker feature subset. If the dimensions of the feature subsets of each modality are inconsistent, the regularized linear discriminant analysis is used to align the feature subsets across modalities to obtain the feature matrix of uniform dimension. The self-attention mechanism is used to learn the interaction between modalities of the feature matrix of uniform dimension to obtain the fused multi-modal feature representation. The dynamic kernel function of the adaptive support vector machine model is used to classify the fused multi-modal feature representation to obtain the diagnosis probability value of Alzheimer's disease. According to the diagnosis probability value, the contribution of each modality feature is calculated by using the Shapley value to obtain the feature importance heat map.

[0043] Exemplarily, the preprocessing module performs a morning fasting 10 mL venous blood collection on the blood sample, separates the plasma by centrifugation at 3000 g for 15 minutes, stores the urine sample in a 10 mL centrifuge tube at -80°C for no more than 6 months, acquires a 1x1x1 mm3 resolution T1-weighted image through a 3T MRI on the neuroimaging data, collects a 5-minute closed-eye resting-state EEG using a 64-electrode cap on the electrophysiological data, and enters the clinical data into the Montreal Cognitive Assessment Scale directional force sub-item score. After brain extraction and bias field correction of all data by the ANTs software, a standardized dataset is generated. When the convolutional neural network processes the neuroimaging data, the ResNet50 architecture is used to extract the hippocampal CA1 and DG region atrophy rate features, and a 0.5% annual change rate is set as the diagnostic threshold. The output includes a feature subset containing spatial coordinates and atrophy degree. The blood p-tau231 concentration is detected by a sandwich chemiluminescence kit, and the detection limit is set to 0.1 pg / mL. When the expression of miR-124-3p is analyzed by reverse transcription quantitative PCR, ΔCt=2.3 is used as the critical value. UCH-L1 and Aβ1-42 in urine are quantified by Simoa technology to generate a biomarker feature subset containing concentration values and standard deviations. When the EEG signal is processed by independent component analysis, the FastICA algorithm is used to remove the ocular artifact, and the alpha power spectral density and theta / beta ratio in the 1-40 Hz frequency band are extracted. A 30% reduction in alpha power or a 40% increase in theta / beta ratio is used as the abnormal threshold to form an electrophysiological feature subset. The Montreal Cognitive Assessment Scale directional force sub-item score is digitally converted, and the original score is mapped to the 0-10 interval. Scores below 8 are determined to be abnormal, and a structured clinical feature subset is generated. When the regularized linear discriminant analysis performs cross-modal alignment, the L2 regularization coefficient λ is set to 0.01. By maximizing the ratio of inter-class dispersion to intra-class dispersion, the five-dimensional feature subset is projected into a 128-dimensional unified space.

[0044] In step S104, the multiple modal feature subsets are interactively learned through a self-attention mechanism to generate a fusion feature set, which reflects the correlation between the modal data and optimizes the inter-modal weight distribution through attention head adjustment.

[0045] In a preferred embodiment, the subsets of modal data including blood biomarkers p-tau231 and miR-124-3p, urine UCH-L1 and Aβ1-42, CNN-based hippocampal CA1 atrophy rate, EEG alpha power spectral density and theta / beta ratio, MoCA directional force sub-item score are acquired, AD-specific modules are constructed by weighted gene co-expression network analysis, and the subset of modal features is obtained. The independent processing mechanism is used to standardize each subset of modal features, the cross-modal features are aligned by regularized linear discriminant analysis, and the standardized feature matrix is obtained. The self-attention mechanism is used for interactive learning of the standardized feature matrix, the number of attention heads is set to 5, the initial fusion feature set reflecting the correlation between modalities is generated. According to the initial fusion feature set, the kernel function with dynamically adjusted γ parameter is used, the adaptive support vector machine top-level classifier is used to optimize the fusion features, and the optimized fusion feature set is obtained. If the variance of the modal weight distribution in the optimized fusion feature set is greater than the preset threshold 0.1, the interactive learning is performed again by adjusting the number of attention heads, and the fusion feature set with re-distributed weights is obtained. The re-distributed weight fusion feature set is obtained, the contribution of each modal feature is calculated by Shapley value, a feature importance heat map is generated, and a feature contribution matrix is obtained. According to the feature contribution matrix, the online random forest algorithm is used for dynamic feature selection of the fusion feature set, and a simplified feature subset is obtained. Through the online learning mechanism, the simplified feature subset is combined with the newly added 50 data, the incremental update algorithm is used to iterate the adaptive support vector machine model, and the updated classification model is obtained. The updated classification model is used for stratified 5-fold cross-validation of multi-center validation data, the AUC is compared by DeLong test, and the final AD identification result is obtained.

[0046] For example, the acquisition of each modal data subset includes blood biomarkers p-tau231 (diagnostic threshold 0.87 pg / mL) and miR-124-3p (ΔCt = 2.3), urine UCH-L1 and Aβ1-42, CNN-based hippocampal CA1 region atrophy rate (annual rate > 0.5%), EEG alpha power spectral density (decrease by 30%) and theta / beta ratio (increase by 40%), MoCA directional force sub-item score (< 8 points), and AD-specific modules are constructed by weighted gene co-expression network analysis, and markers with significant correlation with disease progression (p < 0.001) are screened to obtain a modal feature subset. Each modal feature subset is standardized by an independent processing mechanism, and cross-modal features are aligned by regularized linear discriminant analysis, for example, miRNA data is standardized using the internal reference gene GAPDH, and image data is subjected to brain extraction and bias field correction by ANTs software to obtain a standardized feature matrix. The standardized feature matrix is interactively learned by a self-attention mechanism, with 5 attention heads set, to generate an initial fusion feature set reflecting the correlation between modalities. According to the initial fusion feature set, a kernel function with a dynamically adjusted γ parameter is used to optimize the fusion features by an adaptive support vector machine top-level classifier to obtain an optimized fusion feature set. If the variance of the modal weight distribution in the optimized fusion feature set is greater than the preset threshold 0.1, the attention head number is adjusted to re-learn the interaction, and a fusion feature set with re-distributed weights is obtained. The re-distributed weight fusion feature set is obtained, the contribution of each modal feature is calculated by Shapley value, a feature importance heat map is generated, and a feature contribution matrix is obtained. According to the feature contribution matrix, an online random forest algorithm is used to dynamically select features from the fusion feature set, and key features such as p-tau231, miR-124-3p, and hippocampal CA1 region atrophy rate are screened to obtain a simplified feature subset.

[0047] In step S105, an adaptive support vector machine model is constructed according to the fusion feature set, and the model includes a three-level feature fusion architecture. The sample is classified by dynamically adjusting the kernel function parameter, and the differential results of Alzheimer's disease and other types of dementia are output.

[0048] In a preferred embodiment, blood, urine, neuroimaging, electrophysiology and clinical data are acquired, plasma and urine samples are isolated using standardized preprocessing pipelines, imaging data are corrected for bias field, EEG 1-40Hz band power is extracted, and a multi-modal raw feature set is obtained. Based on the raw feature set, AD-specific modules are constructed using weighted gene co-expression network analysis, five-dimensional markers including p-tau231, miR-124-3p, hippocampal CA1 region atrophy rate, etc. are screened, and a high-correlation feature subset is determined. The feature subset is cross-modality aligned through regularized linear discriminant analysis, reducing dimensionality and preserving inter-modality interaction information, and an aligned feature vector set is obtained. A three-level feature fusion architecture is used to process the feature vector set, with the bottom layer independently extracting features of each modality, and the middle layer learning inter-modality interaction through a self-attention mechanism to obtain a fused feature set. Based on the fused feature set, an adaptive support vector machine model is constructed, with dynamic kernel function parameters configured, and iteratively optimized through an online learning mechanism to obtain an initial classification model. If the sample distribution in the fused feature set is unbalanced, SMOTE oversampling is implemented for the low-sample-size class, with the oversampling rate adjusted to 200%, and a balanced training dataset is obtained. The performance of the initial classification model is evaluated through stratified 5-fold cross-validation, AUC, sensitivity and specificity are calculated, and the model classification accuracy is determined. A Shapley value-based interpretability module is used to analyze the fused feature set, a feature importance heat map is generated, and the contribution of each marker to the classification result is obtained. Based on the addition of 50 new data, the online learning mechanism is triggered to update the adaptive support vector machine model, and the differential diagnosis results of Alzheimer's disease and other dementia types are output.

[0049] Exemplarily, 10 mL of venous blood is extracted on an empty stomach in the morning, and plasma is separated by centrifugation at 3000 g for 15 minutes. At the same time, 10 mL of urine sample is collected and stored at -80℃ for no more than 6 months. A T1 weighted image with a resolution of 1x1x1 mm3 is obtained by 3.0T MRI. A 5-minute resting-state EEG signal in a closed-eye state is collected using a 64-electrode cap. After removing the eye movement artifact by ICA, the power spectral density in the frequency band of 1-40 Hz is extracted. Combined with the directional force sub-score in the MoCA scale, a multi-modal original feature set is constructed. Based on the WGCNA algorithm, five-dimensional markers such as blood p-tau231 (threshold value 0.87 pg / mL), miR-124-3p (ΔCt=2.3), hippocampal CA1 region annual atrophy rate >0.5%, EEG θ / β ratio increased by 40%, and MoCA directional force <8 points are screened, forming an AD-specific module. Through the RLDA algorithm, the cross-modal features are linearly projected, and the feature vectors of the top 20 largest discriminant directions are retained to realize the alignment of the feature spaces between the modalities. In the three-level fusion architecture, the CNN is used in the bottom layer to extract the hippocampal subregion atrophy features, five attention heads are set in the middle layer to calculate the interaction weights of the biomarkers and image features, and a 128-dimensional feature vector is generated in the top layer. When constructing the ASVM model, the RBF kernel function parameter γ=0.01 is initialized. When it is detected that the DLB sample size is less than 15% of the total data, the SMOTE algorithm is used to generate 200% of the synthetic samples. Using stratified 5-fold cross-validation, the AUC of the model on the AD vs. VaD test set is 0.992, and the sensitivity is 95.6%. Through Shapley value analysis, the contribution of each feature is shown, and the weight proportion of p-tau231 is 32.7%.

[0050] In step S106, an online learning mechanism is introduced for the adaptive support vector machine model, and the model parameters are updated by incremental data. The update mechanism is triggered when the number of new samples reaches a preset threshold, ensuring that the model adapts to the distribution changes of new data.

[0051] In a preferred embodiment, the online data acquisition module is used to obtain new sample data, including blood p-tau231, miR-124-3p, urine UCH-L1, Aβ1-42, hippocampal CA1 region atrophy rate, EEG alpha power spectral density, and MoCA directional force sub-item score, to obtain a standardized multi-modal data set. The data cleaning module is used to process outliers and missing values in the new sample data. The outlier detection algorithm is used to remove data points deviating from 3 times the standard deviation. The cleaned multi-modal feature set is obtained. The dynamic feature selection module is used to sort the importance of the cleaned multi-modal feature set based on the online random forest algorithm, to filter out the highest correlation feature combination related to AD diagnosis, and to obtain an optimized feature subset. According to the optimized feature subset, the feature distribution statistics of the new sample are calculated, including mean, variance, and skewness coefficient, to obtain the distribution parameters of the new sample. If the number of new samples reaches the preset threshold of 50, the online learning mechanism is triggered, the stored historical training data is compared with the distribution parameters of the new samples, and the distribution change degree is judged. The online updating module of the adaptive support vector machine model is used to dynamically adjust the kernel function parameter γ based on the distribution change degree, to update the model weight using the incremental learning algorithm, and to obtain the updated model parameters. The three-level feature fusion architecture is used to process the optimized feature subset of the new sample, the bottom layer independently extracts each modality feature, the middle layer learns the interaction between modalities through the self-attention mechanism, and the top layer applies the updated model parameters for classification to obtain the AD diagnosis probability value. The Shapley value interpretability module is used to analyze the contribution of the updated model to each feature, to generate a feature importance heat map, and to obtain a visual feature contribution analysis result. The cloud service platform is used to synchronize the updated model parameters, AD diagnosis probability value, and feature contribution analysis result to the multi-center database, to obtain cross-regional verification data feedback, and to obtain the model performance evaluation index.

[0052] Exemplarily, the online data acquisition module obtains the newly added sample data from the portable detection terminal. The blood sample is detected by a sandwich chemiluminescence kit for p-tau231 (detection limit 0.1 pg / mL), miR-124-3p is quantified by reverse transcription primers (ΔCt = 2.3), and the urine sample is detected for UCH-L1 and Aβ1-42 after centrifugation. The atrophy rate of the hippocampal CA1 region is calculated based on the annual change rate (> 0.5%) of the T1-weighted image (resolution 1x1x1 mm3) of the 3.0T MRI. The EEG signal is collected by a 64-electrode cap, and the α wave power spectral density (reduced by 30%) and the θ / β ratio (increased by 40%) are extracted. The MoCA directional force sub-item score is entered using the standardized scale, forming a multi-modal data set containing molecular, imaging, electrophysiological and clinical characteristics. The data cleaning module detects outliers in the collected data, and uses the 3σ principle to remove abnormal data deviating from the mean by 3 times the standard deviation, for example, samples with p-tau231 concentration exceeding 5.2 pg / mL or miR-124-3p ΔCt value exceeding the range of 1.8-3.0, while the missing values are filled using the K-nearest neighbor algorithm to generate a cleaned feature matrix. The dynamic feature selection module calculates the feature importance based on the online random forest algorithm, sets the Gini coefficient threshold to 0.15 to select key features, for example, the weights of p-tau231, hippocampal CA1 atrophy rate and EEG α wave power reach 0.32, 0.28 and 0.19 respectively, and the remaining low weight features are removed to form an optimized feature subset. For the optimized feature subset, the distribution parameters of the newly added 50 samples are calculated, for example, the mean of p-tau231 is 0.91 pg / mL (variance 0.12), and the mean of hippocampal atrophy rate is 0.53% (skewness coefficient 0.21), and the distribution parameters of the historical training data (mean of p-tau231 0.87 pg / mL, variance 0.09) are subjected to Kolmogorov-Smirnov test to determine whether the distribution deviation exceeds the preset threshold ΔD = 0.15. If the distribution deviation meets the standard, the online update module of the adaptive support vector machine model starts incremental learning, dynamically adjusts the γ parameter (initial value 0.01, adjustment step 0.002) using the RBF kernel function, updates the model weight by the stochastic gradient descent algorithm, and the iteration number is set to 200 times and the learning rate is 0.001. The updated model uses a three-level feature fusion architecture to process new data, the bottom layer extracts hippocampal image features (output dimension 256) through independent CNN, the middle layer learns the interaction weight between biomarkers and electrophysiological features through self-attention mechanism (head number 5), and the top layer applies the adjusted kernel function for classification, outputting the AD probability value (for example, 82.7%). The Shapley value module analyzes the model decision-making process, calculates the contribution of each feature and generates a heat map, for example, the contribution of p-tau231 is 35.6%, the contribution of hippocampal atrophy is 28.4%, and the remaining features are ranked in descending order of weight.For example, the cloud service platform synchronizes model parameters, diagnostic results and feature heat maps to a multi-center database through a gRPC interface (response time < 200 ms), and the AUC of 150 data feedbacks from the Tokyo, Japan cohort is 0.987, with a difference p = 0.089 from the core cohort, confirming the stability of the model.

[0053] In step S107, multi-modal data of the to-be-detected object is obtained, including blood, urine, image, electrophysiology and clinical data, features are extracted by the five-dimensional marker network, and are input to the adaptive support vector machine model for classification.

[0054] In a preferred embodiment, multi-modal data of the to-be-detected object is obtained, including p-tau231, miR-124-3p and sphingomyelin SM (d18:1 / 16:0) in blood, UCH-L1 and Aβ1-42 in urine, 3.0T MRI T1 weighted image, resting state EEG signal collected by 64 electrode cap, and MoCA directional force sub-item score, to obtain a standardized original data set. The miRNA data is standardized by using GAPDH internal reference gene, the MRI image is subjected to brain extraction and bias field correction by using ANTs software, and the ocular artifact in the EEG signal is removed by using ICA method, to obtain preprocessed multi-modal data. The concentrations of blood p-tau231 and urine UCH-L1 are detected by using a sandwich chemiluminescence kit, SV2A is quantified based on Simoa technology, the atrophy rate of the hippocampal CA1 region and DG region in the MRI image is extracted, the α wave power spectral density and θ / β ratio of the EEG signal are calculated, to obtain a five-dimensional marker feature set. An AD-specific module is constructed by using a weighted gene co-expression network analysis, the correlation of the features with disease progression is verified by using a Cox proportional hazards model, and key features such as p-tau231, miR-124-3p and hippocampal CA1 region atrophy rate are screened, to obtain an optimized feature subset. If the feature subset contains low-sample DLB data, the data is enhanced by 200% oversampling rate by using the SMOTE algorithm, otherwise the features are directly aligned, the cross-modal feature alignment is realized by using regularized linear discriminant analysis, to obtain a high-dimensional feature vector of uniform dimension. The high-dimensional feature vector is processed by using a three-level feature fusion architecture, the features of each modality are independently extracted at the bottom layer, the interaction between modalities is learned by using 5 attention heads at the middle layer, and the fusion features are generated based on the kernel function with a dynamically adjusted γ parameter at the top layer, to obtain a comprehensive feature representation. The adaptive support vector machine model is used for classification of the comprehensive feature representation, the features are dynamically selected by using an online random forest, if 50 new data are added, the model is triggered for incremental update, to obtain a classification result. The contribution of each feature to the classification result is calculated by using Shapley value, to generate a feature importance heat map, and an interpretable diagnostic report is obtained by combining AD probability value and feature contribution analysis.

[0055] For example, the gRPC interface is used to push the diagnostic report to the Flutter developed front-end interface, and the results are displayed in real time through the NVIDIA Jetson AGX Orin edge computing platform to obtain the final visual output.

[0056] In step S108, the dementia type probability value of the to-be-detected object is determined according to the output result of the adaptive support vector machine model, the probability value is used to distinguish Alzheimer's disease from other non-Alzheimer's disease dementia types, and a feature importance analysis report is generated.

[0057] In a preferred embodiment, multi-modal data of the to-be-detected object is obtained, including p-tau231, miR-124-3p, and sphingomyelin SM (d18:1 / 16:0) in blood, UCH-L1 and Aβ1-42 in urine, hippocampal CA1 atrophy rate based on CNN, alpha power spectral density and theta / beta ratio in resting state EEG, and MoCA directional force sub-item score, and standardized preprocessing is performed to obtain normalized feature vectors. The bottom layer architecture of the adaptive support vector machine model is used to independently process each modality feature vector, extract image features through the convolutional neural network, process electrophysiological signals through independent component analysis, and align molecular markers through linear discriminant analysis to obtain preliminary feature representations of each modality. Through the middle layer architecture of the self-attention mechanism, the interaction weights between the features of each modality are calculated, the number of attention heads is set to 5, the blood, urine, image, electrophysiological, and clinical features are fused to obtain comprehensive feature representations across modalities. The top layer dynamic kernel function SVM classifier is used to dynamically adjust the γ parameter, and classification calculation is performed based on the comprehensive feature representations to obtain the initial probability distribution of the dementia type of the to-be-detected object. If the difference between the Alzheimer's disease probability value and the non-Alzheimer's disease dementia type probability value in the initial probability distribution is less than 0.1, then the SMOTE oversampling technique is used to enhance the low sample size dementia type data, and the model iteration is triggered again to obtain an updated probability distribution. According to the updated probability distribution, the cross-modal features are optimized and aligned in combination with the regularized linear discriminant analysis, the contribution of each modality feature to the classification result is calculated, and the initial weight of the feature importance is obtained. The Shapley value calculation method is used to analyze the contribution of each feature to the dementia type probability value, generate a feature importance heat map, and obtain visual feature importance analysis data. Through the online learning mechanism, the multi-modal data of the to-be-detected object and the feature importance analysis data are added to the training set, and if the amount of new data reaches 50 cases, the model parameter update is automatically triggered to obtain the optimized model weight. According to the optimized model weight and the feature importance analysis data, a diagnostic report containing the Alzheimer's disease probability value, the non-Alzheimer's disease dementia type probability value, and the feature contribution analysis is generated, and the dementia type probability value and the feature importance analysis report of the to-be-detected object are determined.

[0058] Exemplarily, the multi-modal data of the to-be-detected object is acquired, including p-tau231 in blood (diagnostic threshold 0.87 pg / mL), miR-124-3p (ΔCt = 2.3), sphingomyelin SM (d18:1 / 16:0) in blood, UCH-L1 and Aβ1-42 in urine, hippocampal CA1 atrophy rate (annual change rate > 0.5%) based on CNN, alpha power spectral density (reduction of 30%) and theta / beta ratio (increase of 40%) in resting-state EEG, and MoCA directional sub-item score (< 8 points), and normalized feature vectors are obtained through standardized preprocessing. The bottom layer architecture of the adaptive support vector machine model is used to independently process each modal feature vector, the image features are extracted through the convolutional neural network, the brain extraction and bias field correction are performed using the ANTs software, the electro-physiological signal is processed by independent component analysis, the eye movement artifact is removed and the 1-40 Hz band power is extracted, and the linear discriminant analysis is used to align the molecular markers to obtain the preliminary feature representation of each modality. Through the middle layer architecture of the self-attention mechanism, the interaction weight between the features of each modality is calculated, the number of attention heads is set to 5, the blood, urine, image, electro-physiological and clinical features are fused to obtain the comprehensive feature representation of the cross-modality. The top layer dynamic kernel function SVM classifier is used to dynamically adjust the gamma parameter, and the classification calculation is performed based on the comprehensive feature representation to obtain the initial probability distribution of the dementia type of the to-be-detected object. If the difference between the probability value of Alzheimer's disease and the probability value of non-Alzheimer's disease dementia type in the initial probability distribution is less than 0.1, the SMOTE oversampling technology is used to enhance the low sample size DLB queue data, the oversampling rate is set to 200%, the model iteration is triggered again, and the updated probability distribution is obtained. According to the updated probability distribution, the cross-modality features are optimized and aligned by the regularized linear discriminant analysis, the contribution of each modality feature to the classification result is calculated, and the initial weight of the feature importance is obtained. The Shapley value calculation method is used to analyze the contribution of each feature to the probability value of the dementia type, generate a feature importance heat map, and obtain visual feature importance analysis data. Through the online learning mechanism, the multi-modal data of the to-be-detected object and the feature importance analysis data are added to the training set, if the amount of new data reaches 50 cases, the model parameter update is automatically triggered, and the optimized model weight is obtained. According to the optimized model weight and the feature importance analysis data, a diagnosis report containing the probability value of Alzheimer's disease (0-100%), the probability value of non-Alzheimer's disease dementia type and the feature contribution analysis is generated, and the probability value of the dementia type of the to-be-detected object and the feature importance analysis report are determined.

[0059] In step S109, the contribution of each marker to the diagnostic result is output based on the explainability module through the feature importance analysis report, and the contribution is presented in the form of a heat map to assist clinical decision-making.

[0060] In a preferred embodiment, the feature representation after dimension reduction is analyzed by the Shapley value-based explainability module to calculate the contribution of each marker to the model output, obtaining a feature contribution vector. According to the feature contribution vector, the heat map generation algorithm is used to map the contribution value to the color intensity, generating initial heat map data. If the resolution of the initial heat map data is lower than the preset threshold, the color intensity is smoothed by the interpolation algorithm to obtain high-resolution heat map data. Through the front-end visualization module, the high-resolution heat map data is rendered based on the Flutter framework to generate an interactive heat map interface, and the visualization output is determined. Through the gRPC interface, the interactive heat map interface data is transmitted to the cloud service platform to synchronize multi-center data, and it is judged whether the clinical decision support report generation is completed.

[0061] For example, the multi-modal data input is obtained, including p-tau231 in blood (diagnostic threshold 0.87 pg / mL), miR-124-3p (ΔCt = 2.3), UCH-L1, Aβ1-42 in urine, CNN-based hippocampal CA1 region atrophy rate (annual rate > 0.5%), EEG alpha power spectral density (reduction of 30%), and the data cleaning module is used to process the data by the outlier detection algorithm, and the data deviating from the mean by 3 times the standard deviation is removed to obtain a standardized feature set. According to the standardized feature set, an adaptive support vector machine model is used to process through a three-level feature fusion architecture, the bottom layer independently extracts each modal feature, such as using Simoa technology to quantitatively extract SV2A to obtain an initial feature vector of each modality. Through the self-attention mechanism, the initial feature vector of each modality is processed, the number of attention heads is set to 5, the interaction weight between modalities is calculated, such as the interaction weight between blood markers and image features is 0.78, and a fusion feature matrix is obtained. The regularization linear discriminant analysis is used to align the cross-modal features of the fusion feature matrix, and the feature space distribution is optimized, such as aligning the feature vectors of p-tau231 and hippocampal atrophy rate, to obtain a reduced feature representation. Through the explainability module based on Shapley value, the reduced feature representation is analyzed, the contribution of each marker to the model output is calculated, such as the contribution of p-tau231 is 0.45 and the contribution of miR-124-3p is 0.32, to obtain a feature contribution vector. According to the feature contribution vector, a heat map generation algorithm is used to map the contribution value to the color intensity, such as mapping the contribution value 0.45 to dark red and 0.32 to light red, to generate initial heat map data. If the resolution of the initial heat map data is lower than the preset threshold (such as pixel density < 300 dpi), the color intensity is smoothed by the interpolation algorithm, such as using bilinear interpolation to increase the pixel density to 600 dpi, to obtain high-resolution heat map data. Through the front-end visualization module, the high-resolution heat map data is rendered based on the Flutter framework to generate an interactive heat map interface, such as supporting click to view specific contribution value, to determine the visualization output. Through the gRPC interface, the interactive heat map interface data is transmitted to the cloud service platform to synchronize multi-center data, such as uploading the data to the NVIDIA Jetson AGX Orin edge computing platform, to determine that the clinical decision support report generation is completed.

[0062] In step S1010, for the diagnosis result and contribution analysis, the cloud service platform is combined to realize multi-center data synchronization, the platform globally optimizes the model parameters, and the optimized parameters are fed back to the edge computing unit to improve the diagnosis accuracy.

[0063] In a preferred embodiment, the multi-center data is collected by the cloud service platform, including the diagnostic results, feature contribution heat map and original biomarker data uploaded by each edge computing unit, stored in a distributed database, and a standardized multi-center data set is obtained. The multi-center data set is preprocessed by a data cleaning module to remove outliers and standardize miRNA data based on the internal reference gene GAPDH, and a cleaned multi-modal feature set is obtained. According to the cleaned multi-modal feature set, an online random forest algorithm is applied to perform dynamic feature selection, and the highest contribution of the marker combination to AD diagnosis is screened out to determine the optimized feature subset. If the feature subset meets the preset screening threshold (p<0.001, Cox proportional risk model verification), the feature subset is learned by inter-modal interaction through a self-attention mechanism to obtain a fused feature vector. The fused feature vector is classified and trained by a top-level dynamic kernel function of an adaptive support vector machine model, and the γ parameter is adjusted to obtain globally optimized model parameters. The globally optimized model parameters are distributed to the multi-center nodes of the cloud service platform through a gRPC interface, and the model service module of each node is updated synchronously to obtain a consistent model parameter set. The consistent model parameter set is transmitted to the NVIDIA Jetson AGX Orin platform of each edge computing unit through the cloud service platform, the local ASVM model is updated, and an optimized edge diagnostic model is obtained. According to the optimized edge diagnostic model, the blood, urine and EEG data collected locally are processed, the hippocampal CA1 region atrophy rate analyzed by CNN is combined, a new diagnostic result and feature contribution heat map are generated, and an updated diagnostic output is obtained. The updated diagnostic output is uploaded to the cloud service platform through the cross-platform interface of the edge computing unit, triggering the online learning mechanism (iterating once every 50 new data), and a continuously optimized multi-center data cycle is obtained.

[0064] Embodiment 2

[0065] To illustrate the model architecture and parameter details used in the present application in detail, as shown in Figure 2 The present embodiment provides an Alzheimer's disease auxiliary precise identification method based on multi-omics marker network and adaptive support vector machine, which comprises the following steps:

[0066] Step S1, multi-modal data acquisition and preprocessing, acquiring multi-modal data to be detected, the multi-modal data at least comprising:

[0067] (a) Molecular marker data: Collect blood and / or urine samples, measure the concentration or relative expression of at least phosphorylated tau protein p-tau231, microRNA miR-124-3p, sphingomyelin SM(d18:1 / 16:0) in blood, and ubiquitin carboxy-terminal hydrolase L1 (UCH-L1), amyloid-beta 1-42 (Ab1-42) in urine. Standardize the measurements, e.g. correct miRNA data using an internal control gene (e.g. GAPDH), and normalize protein / metabolite concentrations.

[0068] (b) Imaging marker data: Collect high-resolution (e.g. 1x1x1 mm3) T1-weighted brain magnetic resonance imaging (MRI) data of the subject. Preprocess using standardized pipelines (e.g. ANTs toolbox), including brain tissue extraction, bias field correction, spatial registration to a standard template space.

[0069] (c) Electrophysiological marker data: Collect multi-lead (e.g. 64-lead) electroencephalogram (EEG) signals of the subject at rest (e.g. 5 minutes of eyes closed). Preprocess, including filtering (e.g. 1-40 Hz band-pass filtering), artifact removal (e.g. ocular and muscular artifacts based on independent component analysis, ICA).

[0070] (d) Clinical marker data: Collect demographic information (age, gender, education, etc.) and standardized cognitive assessment results of the subject, including at least the Montreal Cognitive Assessment (MoCA) total score and the score of the sub-item of directional power. Standardize the score data.

[0071] Step S2, intra-modality feature extraction based on specific sub-models, use deep learning sub-models designed for different data modalities to extract high-dimensional feature representations from the preprocessed data:

[0072] (a) Molecular marker sub-model: input the standardized values of multiple blood and urine molecular markers (p-tau231, miR-124-3p, SM(d18:1 / 16:0), UCH-L1, Ab1-42, etc.) in step S1(a) into a multi-layer perceptron (MLP) network. The MLP learns the complex combination relationships and patterns among these marker values through nonlinear transformation, captures the biochemical pathological state indicated by them together, and outputs a unified “molecular feature vector”.

[0073] (b) Image Marker Sub-model: The pre-processed 3D MRI volume data in step S1(b) is input into a 3D convolutional neural network (3D CNN). The 3D CNN automatically learns the spatial structural features of the brain, especially the volume, morphology or model-based atrophy rate estimation (e.g., whether the annual change rate is greater than 0.5%) of specific sub-regions of the hippocampus (e.g., CA1, dentate gyrus DG), and outputs an "image feature vector" that encodes key brain structure information.

[0074] (c) Electrophysiological Marker Sub-model: The pre-processed EEG signal segment or its frequency domain representation (e.g., power spectral density) in step S1(c) is input into a one-dimensional convolutional neural network (1D CNN) or a recurrent neural network (RNN) or a feature extraction module (calculating alpha power spectral density, theta / beta ratio, etc.) connected to an MLP. The sub-model aims to capture the time-frequency characteristics in the EEG signal, quantify abnormal changes in brain functional rhythms (e.g., alpha wave power reduction >30%, theta / beta ratio increase >40%), and output an "electrophysiological feature vector" reflecting the brain functional state.

[0075] (d) Clinical Marker Sub-model: The standardized clinical scores (e.g., feature of MoCA directional force score <8 points) and demographic data in step S1(d) are input into an MLP network. The MLP learns the association between these clinical variables and disease phenotypes, and outputs a "clinical feature vector".

[0076] Step S3, Cross-modal Feature Alignment and Fusion:

[0077] (a) Feature Alignment (Optional but Recommended): For the molecular feature vector, image feature vector, electrophysiological feature vector and clinical feature vector output in step S2, apply feature alignment techniques (e.g., regularized linear discriminant analysis RLDA) to project the features of different modalities into a more comparable latent space, reducing the scale and distribution differences between modalities.

[0078] (b) Multimodal Feature Fusion: Input the aligned (or directly from step S2) feature vectors of each modality into a self-attention (Self-Attention) mechanism module. The self-attention mechanism dynamically assigns weights to each modality feature by calculating the correlation between different modality features (Query, Key, Value operations), and captures the interaction information between modalities, finally weighting and fusing multiple feature vectors into a single, more informative "fusion feature vector". The number of attention heads can be set to, for example, 5.

[0079] After self-attention, a modality interaction attention mechanism is introduced, which learns a weight matrix to explicitly model the strength of interaction between different pairs of modality features (e.g., the strength of association between blood p-tau features and hippocampal atrophy features). This helps to build a more accurate representation of the "marker network" in the feature space.

[0080] Step S4, based on the classification of adaptive support vector machine, the fusion feature vector obtained in step S3 is input into an adaptive support vector machine (Adaptive Support Vector Machine, ASVM) classifier for diagnosis and discrimination:

[0081] (a) Adaptive kernel function adjustment: The ASVM includes a core SVM classifier and an adaptive kernel function module. The module (e.g., a small neural network) receives the fusion feature vector or part of the information as input, and real-time predicts and adjusts the key parameters (such as the value of gamma 'γ') of the kernel function (such as the radial basis function RBF kernel) of the core SVM classifier. This adjustment enables the decision boundary of the SVM to adaptively change according to the feature complexity of the current input sample, improving the adaptability to different cases and disease stages.

[0082] (b) Classification decision: The core SVM uses the received fusion feature vector and dynamically adjusted kernel function parameters to perform the classification task, outputting the probability value of the subject having Alzheimer's disease (AD), or classifying it into AD, VaD, DLB, or health control, etc. preset categories.

[0083] Step S5, model dynamic updating and explainability output (optional):

[0084] (a) Online learning mechanism: The system can be configured with an online learning module. When a certain number (e.g., 50) of new data with confirmed labels is accumulated, the model iteration process is automatically triggered, and the new data is used to incrementally update or retrain the sub-models, fusion modules, and ASVM, maintaining the advanced performance of the model. When the online learning mechanism is triggered (e.g., 50 new samples are added or significant data distribution drift is detected), the system uses the recent data subset to re-evaluate and adjust the values of γ and C through optimization algorithms (such as Bayesian optimization-based parameter search or reinforcement learning strategies) to minimize the classification error or maximize the AUC on the validation set, ensuring that the model can quickly adapt to new data characteristics. The adjustment range of γ can be set to [0.001, 1], and the range of C can be set to [0.1, 100].

[0085] (b) Explainability analysis: combined with explainability techniques (such as Shapley value-based analysis module), the discriminant results of ASVM are explained, and the features (from which specific indicators of which modality) that contribute most to the final diagnosis are output, for example, generating feature importance heat maps, enhancing the understanding and trust of doctors on model decision-making.

[0086] Embodiment 3

[0087] The application also provides an Alzheimer's disease auxiliary precise identification system for implementing the above method, which comprises:

[0088] A data acquisition interface module is used to receive or directly collect data of blood, urine biomarkers, neuroimaging, brain electrophysiology and clinical evaluation from the subject. A portable detection terminal (such as a blood / cerebrospinal fluid analysis module) can be integrated.

[0089] A data preprocessing module is used to realize the preprocessing functions of standardization, correction, denoising and registration of the original data of each modality.

[0090] A feature extraction unit comprises the molecular marker sub-model, the image marker sub-model, the electrophysiological marker sub-model and the clinical marker sub-model defined in step S2, and is used to extract feature vectors of each modality from the preprocessed data.

[0091] A feature fusion unit is used to realize the cross-modality feature alignment (optional) and the feature fusion function based on the self-attention mechanism in step S3, and generate a fusion feature vector.

[0092] An adaptive classification and discrimination unit is used to load a classifier with a core of adaptive support vector machine (ASVM), which comprises an adaptive kernel function module and a core SVM classifier, and is used to perform the classification and discrimination task in step S4. The unit can be deployed on an edge computing platform (such as NVIDIA Jetson AGX Orin).

[0093] A result output and visualization module is used to display the diagnostic results (such as AD probability, classification label) and optional explainability information (such as feature importance) to the user in a user-friendly manner (such as report, chart).

[0094] A model management and updating module (optional) is used to store and load the model, and realize the online learning function in step S5(a), and can be connected to a cloud service platform for multi-center data synchronization and model optimization.

[0095] Embodiment 4

[0096] As shown in Figure 3 The application also provides an Alzheimer's disease auxiliary precise identification system for implementing the above method, which comprises:

[0097] Data acquisition interface module: for acquiring multi-modal data, the multi-modal data at least including blood biomarker data, urine biomarker data, image data, electrophysiological data and clinical data;

[0098] Data preprocessing module: for realizing the standardization, correction, denoising, registration preprocessing functions of the original data of each modality;

[0099] Feature extraction and fusion unit: for feature extraction and fusion of the multi-modal data, a marker network is constructed through a multi-level processing mechanism including cross-modal interaction learning based on attention mechanism, the marker network representing the relevance between different modal features;

[0100] Adaptive classification and discrimination unit: for analyzing the marker network through an adaptive classification model capable of dynamic updating, determining the identification result of Alzheimer's disease and other dementia types;

[0101] Result output and visualization module: for outputting the identification result, the identification result representing the existence probability of Alzheimer's disease and being displayed to the user in a user-friendly manner.

[0102] Embodiment 5

[0103] In order to verify the effectiveness and superiority of the method proposed in the present application, a series of comparative experiments are carried out. Accuracy, AUC (Area Under the ROC Curve), sensitivity, specificity and F1 score (F1-Score) are used as evaluation indexes, and the binary classification identification performance of AD and other dementia (combined as non-AD dementia group) is evaluated

[0104] I. Cohort design and verification method

[0105] 1. Cohort type and characteristics

[0106]

[0107] 2. Diagnostic criteria

[0108] AD diagnosis: in line with NIA-AA research framework (Aβ+ and tau+ biomarkers)

[0109] VaD diagnosis: NINDS-AIREN standard (evidence of vascular lesions + cognitive decline) is used

[0110] DLB diagnosis: based on the core clinical features of DLB alliance + dopamine transporter SPECT abnormalities

[0111] 3. Statistical methods

[0112] Marker difference analysis: Kruskal-Wallis test (multiple groups comparison) + Bonferroni correction

[0113] Model performance evaluation: stratified 5-fold cross-validation + Delong test (AUC comparison)

[0114] II. Cohort validation results

[0115] 1. Core cohort validation

[0116]

[0117] Note: Non-AD group contains VaD (n=80), DLB (n=70), FTD (n=50) and other types of dementia (n=40)

[0118] 2. Early AD cohort validation

[0119] Prodromal AD vs. healthy controls

[0120] Accuracy: 92.3% (AUC=0.976)

[0121] Key marker contributions:

[0122] p-tau231 (OR=5.8, 95% CI 3.2-10.5)

[0123] miR-124-3p (OR=4.2, 95% CI 2.1-8.4)

[0124] Hippocampal CA1 atrophy rate (HR=1.7 / year, p<0.001)

[0125] MCI conversion prediction

[0126] 3-year conversion rate prediction accuracy: 89.5% (AUC=0.961)

[0127] Early warning marker combination:

[0128] Blood NfL (>250 pg / mL)

[0129] EEG alpha power decrease (>25%)

[0130] 3. Non-AD dementia cohort validation

[0131]

[0132] 4. Cross-region validation

[0133]

[0134] III. Subgroup analysis and clinical implications

[0135] 1. Gender differences

[0136] Female AD patients: miR-124-3p sensitivity increased to 97.2% (93.5% for males)

[0137] Male VaD patients: IL-6 levels predicted with higher value (AUC=0.981 vs 0.963 for females)

[0138] 2. Apolipoprotein E (APOE) genotype influence

[0139] APOE4 carriers: model sensitivity increased to 96.8% (94.1% for non-carriers)

[0140] APOE3 homozygotes: relied on combined judgment of p-tau231 and hippocampal atrophy (AUC=0.989)

[0141] 3. Drug influence

[0142] Patients receiving anti-amyloid therapy: model accuracy remained 92.3% (AUC=0.971)

[0143] Cholinesterase inhibitor users: feature importance shifted towards metabolic markers (sphingomyelin contribution + 15%)

[0144] IV. Comparison with prior art

[0145]

[0146] Summary

[0147] The present application provides an Alzheimer's disease auxiliary precision identification system and method based on a multi-modal marker network and an adaptive support vector machine. By innovatively integrating blood, urine, image, electrophysiology and clinical five-dimensional information, designing specific deep learning sub-models for each modality for feature extraction, using a self-attention mechanism to achieve efficient multi-modal fusion, and using ASVM with adaptive kernel function and online learning ability for final classification and discrimination, the accuracy, specificity, early diagnosis ability and adaptability of the model for AD diagnosis and identification are significantly improved, overcoming the limitations of the prior art.

[0148] 1. Significantly improve the accuracy and robustness of identification: By constructing a marker network based on blood, urine, image, electrophysiology and clinical five specific modal data, and using deep fusion strategy combined with attention mechanism and adaptive classification model, the pathological characteristics of AD can be more comprehensively and deeply captured, and AD can be effectively distinguished from other types of dementia (such as VaD, DLB). The present application shows high accuracy (94.1%-96.8%), AUC (0.985-0.992), sensitivity (92.7%-95.6%) and specificity (95.3%-97.3%) in distinguishing AD from VaD, AD from DLB and AD from mixed control group (based on the multi-center validation data provided), which is significantly better than single marker (specificity 78-85%) or traditional method (such as literature reported AD vs VaD accuracy 88.2%), and the accuracy is significantly improved.

[0149] 2. Enhance the ability of early diagnosis: The recognition accuracy of early AD (prodromal AD vs healthy control) reaches 92.3% (AUC=0.976), and the conversion of MCI to AD can be predicted with an accuracy of 89.5%, which benefits from the effective use of early sensitive markers (such as p-tau231, miR-124-3p, hippocampal CA1 atrophy rate, NfL, EEG α wave power).

[0150] 3. Enhance the adaptability and timeliness of the model: The online learning mechanism is introduced, so that the model can dynamically update the parameters according to the new data (for example, every 50 samples accumulated), effectively adapt to the heterogeneity, dynamic evolution of the disease and data drift of different centers, and ensure the long-term effectiveness and stability of the diagnosis model. The adaptive kernel function of ASVM can adjust the decision boundary according to the characteristics of the input sample, improve the adaptability to individual differences and disease heterogeneity. The online learning mechanism enables the model to continuously optimize using new data to maintain diagnostic performance.

[0151] 4. Improve the explainability of diagnosis results and clinical application value: Through the explainability analysis module (such as based on Shapley value calculation), the specific contribution of each marker (such as p-tau231, hippocampal atrophy rate, specific miRNA) to the final diagnosis result can be quantified and presented in the form of heat map visualization, providing transparent basis for clinicians to understand the model decision-making process and evaluate key risk factors, and enhancing the clinical decision support capability.

[0152] 5. Realize efficient multi-center cooperation and optimization: Use cloud platform to realize safe synchronization and sharing of multi-center data, and conduct global optimization of model parameters, then deploy the optimized model to edge computing unit, ensuring the consistency of diagnosis standards among different medical institutions, promoting the wide application and continuous improvement of technology.

[0153] The above examples are only used to illustrate the technical solutions of the present application but not to limit the present application. The present application has been described in detail with reference to the preferred embodiments. It is understood by those skilled in the art that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the present application, and all of them should be covered in the scope of the claims of the present application.

Claims

1. A multimodal marker network-based method for assisting accurate identification of Alzheimer's disease, characterized by: include: Acquiring multimodal data, the multimodal data including at least blood biomarker data, urine biomarker data, imaging data, electrophysiological data, and clinical data; Extracting and fusing features from the multimodal data, constructing a marker network using a multi-level processing mechanism including cross-modal interactive learning based on an attention mechanism, wherein the marker network represents the correlation between features of different modalities; analyzing the marker network using a dynamically updateable adaptive classification model to determine a result of distinguishing Alzheimer's disease from other types of dementia; and outputting the identification result, wherein the identification result represents the probability of the presence of Alzheimer's disease. The multimodal data is subjected to feature extraction and fusion, and a landmark network is constructed through a multi-level processing mechanism including cross-modal interactive learning based on an attention mechanism, specifically including: According to the multimodal data sample set, based on a marker network of five modal data, the network covers phosphorylated tau protein in blood, specific proteins in urine, convolutional neural network-based imaging features, electrophysiological signal power spectral density, and clinical assessment scores, specific markers associated with the disease are screened through weighted gene co-expression network analysis to obtain the five-dimensional marker network; the multimodal data sample set includes phosphorylated tau protein p-tau231 and miR-124-3p in blood, UCH-L1 and Aβ1-42 in urine, 3.0T MRI T1-weighted images, electrophysiological signals EEG, and Montreal Cognitive Assessment MoCA scores, and the data of each modality are cleaned and format-aligned through a standardized preprocessing module to obtain a unified multimodal data set; Using the data in the five-dimensional landmark network, independent feature extraction is performed on each modality data to obtain multiple modality feature subsets, each of which includes key diagnostic indicators for each modality data, and cross-modality feature alignment is achieved through regularized linear discriminant analysis; The multiple modal feature subsets are interactively learned through a self-attention mechanism to generate a fused feature set that reflects the correlation between the modal data, and the weight distribution between the modalities is optimized by adjusting the number of attention heads.

2. The method for assisting accurate identification of Alzheimer's disease based on a multimodal marker network according to claim 1, characterized in that: The analysis of the marker network through a preset classification model includes: inputting the marker network into an adaptive classification model, wherein the adaptive classification model includes a multi-level feature processing architecture; independently analyzing each modal feature through the bottom-level structure of the adaptive classification model to obtain a preliminary classification basis; interactively integrating the preliminary classification basis through the middle-level structure of the adaptive classification model to generate a comprehensive classification basis; and finally classifying the comprehensive classification basis through the top-level structure of the adaptive classification model to determine the identification result.

3. The method for assisting accurate identification of Alzheimer's disease based on a multimodal marker network according to claim 1, characterized in that: The output identification result includes: generating a probability value based on the identification result, the probability value characterizing the possibility of Alzheimer's disease; performing feature contribution evaluation on the identification result through an explainability analysis module to generate a feature importance distribution; generating a diagnosis report based on the probability value and the feature importance distribution, the diagnosis report including a classification conclusion and key feature basis; transmitting the diagnosis report to a terminal device through a preset interface for subsequent analysis reference.

4. The method for assisting accurate identification of Alzheimer's disease based on a multimodal marker network according to claim 1, wherein: The analysis of the marker network using a preset classification model also includes: triggering a dynamic update mechanism of the preset classification model based on new data to obtain an updated classification model; re-analyzing the marker network using the updated classification model to generate an updated identification result; verifying the updated identification result and evaluating the model stability using a cross-validation method; and adjusting the parameters of the preset classification model based on the verification results to optimize the classification performance.

5. An Alzheimer's disease auxiliary accurate identification system for implementing the method described in any one of claims 1 to 4, characterized in that: The system includes: Data acquisition interface module: used to acquire multimodal data, wherein the multimodal data includes at least blood biomarker data, urine biomarker data, imaging data, electrophysiological data and clinical data; Data preprocessing module: used to realize the standardization, correction, denoising and registration preprocessing functions of the raw data of each modality; A feature extraction and fusion unit is configured to extract and fuse features from the multimodal data, and construct a landmark network through a multi-level processing mechanism including cross-modal interactive learning based on an attention mechanism, wherein the landmark network represents the correlation between features of different modalities; Adaptive classification and discrimination unit: used to analyze the marker network through an adaptive classification model that can be dynamically updated to determine the differentiation result between Alzheimer's disease and other types of dementia; Result output and visualization module: used to output the identification result, which represents the probability of the presence of Alzheimer's disease, and display it to the user in a user-friendly manner.

Citation Information

Patent Citations

  • Method, device and system for selecting equipment

    CN118233456A

  • Soybean cultivar 15450218

    US20240081234A1

  • Alzheimer disease diagnosis method based on multi-modal cross attention

    CN118116573A

  • Multi-modal fusion system and method for Alzheimer's disease

    CN118942670A