Method for establishing a clinically interpretable PD diagnostic model based on multi-sequence MRI brain features

By acquiring brain imaging features through multimodal MRI and combining them with machine learning methods, a PD diagnostic model was constructed, which solved the problem of the lack of objective imaging biomarkers in existing technologies and achieved improved PD diagnosis and generalization capabilities across disease sets.

CN119786012BActive Publication Date: 2025-10-03ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411760433.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-10-03
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

Existing technologies lack objective imaging biomarkers in the diagnosis of Parkinson's disease (PD), resulting in moderate diagnostic efficacy and a lack of independent data verification. In addition, the diagnostic efficacy of models with small samples and single-modality imaging data decreases with increasing sample size.

Method used

Brain imaging features were acquired through multimodal MRI (quantitative susceptibility imaging, high-resolution T1-weighted imaging, and arterial spin labeling imaging). A PD diagnostic model was constructed using machine learning methods, including cortical gray matter segmentation, quantitative susceptibility map segmentation, principal component analysis, and feature selection. A binary logistic regression model was constructed to improve diagnostic efficacy.

Benefits of technology

The constructed PD diagnostic model has good diagnostic ability in the training set and independent validation set, good stability and clinical interpretability, can be significantly correlated with clinical symptoms, and performs well in PSP and MSA patients, but has limited diagnostic efficacy in ET patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119786012B_ABST
    Figure CN119786012B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for establishing a clinically interpretable PD diagnostic model based on multi-sequence magnetic resonance brain features, which aims to promote a more accurate diagnosis of PD by extracting multi-level and multi-dimensional brain imaging information from multimodal brain magnetic resonance imaging data and constructing a diagnostic model with the help of machine learning. The method includes: brain imaging data preprocessing, feature standardization and screening, model construction, independent validation set and cross-disease set validation, and obtains a set of brain function and metabolic characteristics mainly based on PD abnormal perfusion patterns and substantia nigra iron distribution. The constructed model has high diagnostic efficacy in both the training set and the independent validation set. Secondly, the clinical interpretability of the model is promoted by calibration curve, DCA, CIC analysis, and correlation analysis between the predicted individual scores and clinical symptoms obtained based on the model. The diagnostic efficacy of the PD diagnostic model constructed by the present invention in different disease sets shows that the model has good diagnostic efficacy for PSP and MSA diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of brain imaging, and in particular relates to a method for establishing a clinically interpretable PD diagnostic model based on multi-sequence magnetic resonance brain features. Background Art

[0002] Parkinson's disease (PD), the second most common neurodegenerative disorder, is currently diagnosed based on subjective and empirical evidence, relying on clinical symptoms and response to conventional anti-Parkinson drugs. In recent years, the role of imaging in the revision of clinical diagnostic criteria for PD has gained increasing attention. However, the diagnosis of PD still lacks a valid clinical reference standard, and the medical care of PD patients remains a focus of international attention. Therefore, the development of objective imaging markers for PD diagnosis and disease assessment is crucial.

[0003] Parkinson's disease (PD) presents a complex pattern of brain damage, involving extensive changes in brain structure, function, and metabolism. Noninvasive magnetic resonance imaging (MRI) technology can obtain multidimensional information of brain tissue in a single scan, providing an effective means of assessing neurodegeneration in PD. Previous researchers have exploited the technical advantages of MRI to explore the degenerative characteristics of the brain in PD and have attempted to construct objective biomarkers based on these findings. A PD diagnostic model constructed using MRI structural data from an international multicenter cohort database reported the beneficial value of MRI in PD diagnosis, but the model's diagnostic efficacy was moderate and lacked independent data validation. More studies have used small sample sizes and single-modality imaging data to explore the diagnostic value of PD, and the model's diagnostic efficacy tends to decrease with increasing sample size, which has hindered the clinical translation of relevant models.

[0004] Therefore, based on PD brain imaging technology, capturing the neurodegenerative characteristics of PD by integrating large-scale sample sizes with brain structure and functional indicators may help us understand the disease heterogeneity of PD, thereby promoting the further development of disease diagnosis and clinical evaluation markers. Summary of the Invention

[0005] The present invention aims to address the current lack of objective imaging biomarkers for PD by providing a method for establishing a clinically interpretable PD diagnostic model based on multi-sequence MRI brain features. This method transforms multimodal MRI (quantitative susceptibility imaging, high-resolution T1-weighted imaging, and arterial spin labeling imaging) into brain imaging features reflecting PD degeneration. This method then utilizes machine learning methods to construct a diagnostic model, improving the diagnostic efficacy of PD and validating its generalizability across multiple disease sets.

[0006] The object of the present invention is achieved through the following technical solution: The present invention provides a method for establishing a clinically interpretable PD diagnostic model based on multi-sequence magnetic resonance imaging brain features, comprising the following steps:

[0007] (1) T1-weighted images, enhanced sensitivity-weighted angiography (ESWAN) images, and arterial spin labeling (ASL) images were obtained from patients with PD, normal controls (NC), progressive supranuclear palsy (PSP), multiple system atrophy (MSA), and essential tremor (ET). PD patients and normal controls were divided into training sets and independent validation sets in proportion, while PSP patients, MSA patients, and ET patients were used as cross-disease datasets.

[0008] (2) The cortical gray matter of the T1-weighted image obtained in step (1) was segmented to obtain 68 cortical regions and 1 TIV; the ESWAN image obtained in step (1) was post-processed using the STAR-QSM algorithm to obtain a quantitative susceptibility map (QSM), and the QSM map was segmented using the DeepQSMSeg algorithm to obtain 10 subcortical nuclei in the bilateral substantia nigra, red nucleus, putamen, caudate nucleus, and globus pallidus; the ASL image obtained in step (1) was subjected to principal component analysis-proportional subsection model analysis to obtain the abnormal PD perfusion pattern;

[0009] (3) extracting brain image features for each cortical region, subcortical nucleus, and PD abnormal perfusion pattern segmented in step (2); and standardizing the brain image features using a general linear model;

[0010] (4) The least absolute shrinkage and selection operator (LASSO) method was used to select multiple brain imaging features with the highest contribution as input, and a PD diagnosis model was constructed using binary logistic regression and trained based on the training set;

[0011] (5) According to step (4), the individual prediction scores obtained based on the PD diagnostic model were correlated with the corresponding clinical symptoms. SHAP (SHapley Additive exPlanations) analysis was used to screen the top six brain imaging features in terms of contribution, and inter-group differences were compared in the training set, independent validation set, PSP, MSA, and ET patient data sets, and FDR (False Discovery Rate) correction was performed.

[0012] Furthermore, in step (1), a fast spoiler gradient sequence is used to obtain a T1-weighted image; a gradient echo sequence is used to obtain an ESWAN image; and a pulsed continuous arterial spin labeling technique is used to obtain an ASL image.

[0013] Furthermore, in step (2), the cortical gray matter of the T1-weighted image obtained in step (1) is segmented to obtain 68 cortical regions, specifically: the subcortical nuclei of the T1-weighted image are segmented using FreeSurfer software, and the recon-al preprocessing process is used to perform automated head motion correction, skull stripping, spatial normalization, spatial registration, and cortical segmentation, and then based on the Desikan-Killiany atlas, 68 cortical regions and total intracranial volume (TIV) are segmented;

[0014] The ESWAN image obtained in step (1) is post-processed using the STAR-QSM algorithm to obtain a quantitative magnetic susceptibility map, specifically: using Laplace-based phase unwrapping to unwrap the original phase map and retain the spatial low-frequency components of all brain tissues; using spherical mean filtering to remove background phases, thereby specifically obtaining brain tissue phase data; estimating magnetic susceptibility boundaries and removing magnetization artifacts; obtaining a tissue magnetic susceptibility map, and calculating the QSM image based on this;

[0015] The DeepQSMSeg algorithm tool sets the true value to the historical data of manual and semi-automatic segmentation; after using the DeepQSMSeg algorithm to segment the QSM image to obtain subcortical nuclei, it also includes: performing manual correction to correct the misalignment of surrounding tissue voxels caused by alignment deviation.

[0016] Furthermore, in step (2), the ASL image obtained in step (1) is subjected to principal component analysis-proportional subsection model analysis to obtain the abnormal PD perfusion pattern, specifically:

[0017] (2.2.1) ASL images and ASL-based cerebral blood flow (CBF) maps were preprocessed using SPM 12, including motion correction, spatial registration and normalization of the ASL and CBF maps using DARTEL, CBF map normalization, and spatial smoothing using a Gaussian kernel with a full width at half maximum (FWHM) of 6 mm.

[0018] (2.2.2) The principal component analysis-scaled sub-profile model analysis (PCA-SSM) method was used to construct the ASL_PDRP of the CBF map preprocessed in step (2.2.1), including: constructing individual sub-profiles, principal component analysis (PCA), scaled sub-profile model analysis (SSM), and constructing PD-related cerebral perfusion patterns;

[0019] (2.2.3) Topological profile scoring (TPR) was used to calculate individual ASL_PDRP expression scores in an independent validation set and PSP, MSA, and ET patient datasets.

[0020] Furthermore, in step (2.2.2), based on step (2.2.1), the constructed individual sub-profiles perform voxel-level analysis on the CBF maps of all participants in the training set and represent them as high-dimensional vectors; the principal component analysis performs matrix centering and covariance matrix decomposition to output the eigenvector and score corresponding to each principal component; the principal component with the highest score is selected to form a proportional sub-profile model, and the expression value ASL_PDRP of each participant on the proportional sub-profile model is calculated;

[0021] In step (2.2.3), based on the proportional sub-profile model obtained based on the training set in step (2.2.2), the individual standardized CBF map based on the independent validation set and the cross-disease dataset obtained by the TPR algorithm is multiplied and accumulated with the weights of the corresponding voxels in the proportional sub-profile model to obtain the individual's overall score on the proportional sub-profile model, which is used as the ASL_PDRP obtained based on TPR.

[0022] Furthermore, in step (3), the brain image features extracted for each cortical region, subcortical nucleus, and PD abnormal perfusion pattern segmented in step (2) are specifically:

[0023] (3.1) For the QSM map, the average magnetic susceptibility values ​​of the 10 subcortical nuclei were calculated as 10 brain imaging features;

[0024] (3.2) For T1-weighted images, 68 average values ​​of cortical thickness and 1 TIV were calculated based on the Desikan-Killiany atlas as 69 brain imaging features;

[0025] (3.3) For ASL images, ASL_PDRP or ASL_PDRP based on TPR is used as a brain image feature;

[0026] The general linear model is used to standardize the brain image features. Specifically, a general linear model (GLM) is constructed based on the brain image features and corresponding age and gender of the normal control group in the training set:

[0027] Y~β+β age X age +β gender X gender

[0028] Among them, Y is the brain image feature, X age is age, X gender For gender, β, β age , β genderare the parameters of the general linear model; the constructed general linear model is applied to all participants in step (1) to obtain estimated brain image features; it is applied to the brain image features extracted in step (3) to obtain brain image features that eliminate the effects of age and gender; and the brain image features after eliminating the effects are z-valued to [-1,1].

[0029] Furthermore, the method for using the PD diagnostic model is as follows: the T1-weighted image, ESWAN image and ASL image to be diagnosed are processed in steps (2) to (3) in sequence, and the influence of age and gender are eliminated based on the GLM; according to step (4), multiple brain image features with the highest contribution are screened using LASSO regression analysis, and the trained PD diagnostic model is input and expanded to be applied to the diagnosis of different types of neurodegenerative diseases.

[0030] Furthermore, in step (4), based on the current brain image features and the corresponding classification labels, the features with higher contributions are selected. Specifically, LASSO analysis is performed and cross-validation is performed using the R language package glmnet to select the optimal regularization parameter, and the 20 most important brain image features are screened. The selected important brain image features are subjected to binary logistic regression analysis using the R language package glm to construct a PD diagnostic model.

[0031] Furthermore, in step (5), based on the trained PD diagnostic model, the predict function of the R language package was used to calculate the individual prediction score, and further Pearson correlation analysis was performed with the UPDRS II and UPDRS III scores; SHAP analysis was used to screen the top 6 brain imaging features with the greatest contribution, and ANOVA analysis of variance was used to compare the differences between groups in different data sets, and FDR correction was performed.

[0032] Furthermore, the method also includes verifying the diagnostic ability of the trained PD diagnostic model across disease sets, specifically: using the predict function of the R language package to apply the trained PD diagnostic model to the MSA, PSP, and ET patient data sets, using the Delong test to evaluate the diagnostic efficacy, and performing FDR correction.

[0033] The beneficial effects of the present invention are:

[0034] (1) A set of 20 important brain imaging features, mainly abnormal perfusion patterns of PD and iron distribution in the substantia nigra, were obtained. The PD diagnostic model constructed based on the above imaging features had an area under the diagnostic curve (AUC value) of 0.839 (95% CI: 0.794-0.879) in the training set and an AUC value of 0.797 (95% CI: 0.682-0.886) in the independent validation set. This shows that the constructed PD diagnostic model has good stability and diagnostic ability in different PD populations and has potential clinical significance.

[0035] (2) Analysis of the calibration curve, decision curve analysis (DCA), and clinical impact curve (CIC) of the PD diagnostic model further demonstrated that the model was relatively reproducible between the training set and the independent validation set. The predicted individual scores obtained based on the model were significantly correlated with the severity of clinical symptoms (UPDRS II: R = 0.12, P = 0.01; UPDRS III: R = 0.13, P = 0.026), indicating that the model has the potential to serve as a clinically interpretable imaging marker for PD and has certain clinical practical value.

[0036] (3) The diagnostic efficacy of the PD diagnostic model constructed based on the above brain imaging features in different disease sets showed that the model had good diagnostic efficacy for PSP and MSA patients, but poor diagnostic efficacy in ET patients (PSP cohort vs ET cohort: P < 0.001; MSA cohort vs ET cohort: P = 0.002). The diagnostic AUC values ​​of the model in the PSP and MSA cohorts were 0.927 (95% CI: 0.844-0.986) and 0.902 (95% CI: 0.800-0.978), respectively, and the diagnostic AUC value in the ET cohort was 0.651 (95% CI: 0.543-0.755). The association between the imaging diagnostic markers constructed in the PD population and the above diseases is still unknown. This result further indicates that although the diagnostic model was constructed for PD patients, PSP and MSA patients may also benefit from it, and its generalization ability deserves further exploration. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1The preprocessing and analysis process for QSM and ASL images. (A.) The preprocessing and analysis process for QSM images, where a is the individual-space QSM image, b is the QSM template in standard space, c is the mask of subcortical nuclei segmented using the DeepQSMSeg algorithm, and d is the mean magnetic susceptibility of subcortical nuclei extracted after manual modification. (B.) The preprocessing and analysis process for ASL images, where a is the ASL image and the CBF image calculated based on the ASL image, b is the gray matter and white matter templates generated using T1-weighted image segmentation, c is the spatially normalized ASL and CBF images, d is the smoothed image, e is the voxel vector matrix of the CBF images for all participants in the training set, f is the PC1 as the ASL_PDRP, g is the topographic profile of the ASL_PDRP, and h is the individual ASL_PDRP expression calculated using TPR in other disease sets.

[0038] Figure 2 This is a flowchart for constructing a PD diagnostic model based on brain imaging features. (A.) shows the enrollment and data collection of four clinical cohorts. (B.) shows the construction process of the PD diagnostic model. A total of 80 image features were extracted using T1-weighted, ASL, and QSM images. A general linear model was constructed based on 122 normal controls from the training cohort, controlling for the effects of age and gender on the extracted image features. Data-driven LASSO regression was used to select the optimal number and best image features. A matrix of all subject features (80 × 369) and a matrix of subject feature selection (20 × 369) were obtained from the training set. Binary logistic regression analysis was used to construct the diagnostic model. (C.) shows the application and validation of the diagnostic model in different disease cohorts. The diagnostic model was applied to the MSA, PSP, and ET cohorts. Receiver operating characteristic (ROC) curves were compared using the Delong test, and area under the curve (AUC) values ​​(95% CI) were obtained using 1000 bootstrapping cycles. (D.) Evaluation of the clinical utility of the diagnostic model. The diagnostic efficacy of the model was verified in an independent validation set. Calibration curve, DCA, and CIC analyses were performed on the training set and independent validation set. SHAP analysis was used to identify brain imaging features that contributed significantly to the diagnostic model, followed by single-feature difference analysis. Correlation analysis was performed between individual prediction scores and clinical scale scores.

[0039] Figure 3 The ROC curves of the LASSO regression feature selection process and diagnostic model in the training set and independent validation set. (A.) Log Lambda (λ) and the error model of LASSO regression were used. Ten-fold cross validation was used. Cross validation was performed using the R language package "glmnet". λ was selected. min=20 as the optimal number of image features; (B.) Log(λ) and LASSO regression coefficients, and the optimal λ or penalty coefficient is determined from LASSO regression. (C.) ROC curves of the diagnostic model on the training set and independent validation set, respectively.

[0040] Figure 4 Figure 2. Clinical evaluation of the PD diagnostic model. (A) Calibration curve analysis of the diagnostic model in the training set; (B) Calibration curve analysis of the diagnostic model in the independent validation set. (C) Clinical decision curve analysis of the diagnostic model in the training set; (D) Clinical decision curve analysis of the diagnostic model in the independent validation set. (E) Clinical impact curve analysis of the diagnostic model in the training set; (F) Clinical impact curve analysis of the diagnostic model in the independent validation set.

[0041] Figure 5 Exploring the clinical utility of the PD diagnostic model. (A.) Pearson correlation analysis between the individual prediction scores of the included PD patients and the clinical UPDRS I scores; (B.) Pearson correlation analysis between the individual prediction scores of the included PD patients and the clinical UPDRS II scores; (C.) Pearson correlation analysis between the individual prediction scores of the included PD patients and the clinical UPDRS III scores; (D.) Pearson correlation analysis between the individual prediction scores of the included PD patients and the Purdue Peg Board total score; (E.) Pearson correlation analysis between the individual prediction scores of the included PD patients and the assembly score; (F.) SHAP mean absolute value bar chart to help screen for highly contributing imaging features.

[0042] Figure 6Figure 3: Differences in the standardized values ​​of the first six brain imaging features from the SHAP analysis, after removing the effects of age and gender. (A) Differences in the standardized values ​​of the first six brain imaging features in the training set: ASL_PDRP: P < 0.001; R_SN: P < 0.001; L_SN: P < 0.001. (B) Differences in the standardized values ​​of the first six brain imaging features in the independent validation set: ASL_PDRP: P = 0.015; R_SN: P = 0.008; L_SN: P = 0.008. (C.) The differences in the standardized values ​​of the first six brain imaging features in the PSP-NC cohort, including: ASL_PDRP: P < 0.001; R_SN: P < 0.001; L_SN: P < 0.001; R_pallium: P = 0.001; R_caudate: P = 0.030; rh_supermargin: P = 0.030; (D.) The differences in the standardized values ​​of the first six brain imaging features in the PSP-PD cohort, including: R_SN: P < 0.001; L_SN: P < 0.001; R_pallium: P < 0.001; R_caudate: P = 0.018). (E) Differences in the standardized values ​​of the first six brain imaging features in the MSA-NC cohort, including: ASL_PDRP: P = 0.007; L_SN: P = 0.007; (F) Differences in the standardized values ​​of the first six brain imaging features in the MSA-PD cohort. (G) Differences in the standardized values ​​of the first six brain imaging features in the ET-NC cohort, and (H) Differences in the standardized values ​​of the first six brain imaging features in the ET-PD cohort, including: ASL_PDRP: P = 0.017. All comparisons above were corrected for FDR. *P < 0.05, **P < 0.01, ***P < 0.001. DETAILED DESCRIPTION

[0043] The present invention aims to extract multi-level and multi-dimensional brain imaging information from multi-sequence brain magnetic resonance imaging (MRI) brain features and construct a clinically interpretable PD diagnostic model using machine learning. Multimodal MRI images include QSM, ASL, and T1-weighted images, from which brain imaging features that can comprehensively reflect the characteristics of PD disease can be extracted. The research framework defined in the present invention includes brain imaging data preprocessing, feature standardization and screening, model construction, independent validation set, and cross-disease validation. In addition, the performance of the present invention's model in a cross-disease validation set showed that the model has good diagnostic efficacy for PSP and MSA, but limited diagnostic efficacy in ET.

[0044] The research implemented in the present invention was approved by the Medical Ethics Committee of the Second Affiliated Hospital of Zhejiang University School of Medicine, and all PD patients, PSP patients, MSA patients and NC subjects signed informed consent.

[0045] The clinical and imaging data for this example were collected from August 2014 to December 2023. The diagnosis of PD was made by a senior neurologist according to the British PD Association Brain Bank criteria and the Movement Disorder Society Clinical Diagnostic Criteria. Initially, a total of 356 PD patients and 183 NC patients were recruited and underwent multimodal MRI scans and clinical assessments. Of these, 77 subjects were excluded due to head motion, significant brain atrophy, multiple microbleeds, history of head trauma surgery, cerebrovascular disease, and other neurological / psychiatric diseases. The final cohort included 307 PD patients and 155 NC patients. For PD patients receiving drug treatment, MRI scans and clinical assessments were performed during the anti-Parkinson's drug washout period (at least 12 hours after discontinuation of the drug). The diagnosis of MSA and PSP was based on the consensus diagnostic criteria, and the diagnosis of ET was based on the Movement Disorder Society statement; the same exclusion criteria were applied to the cross-disease dataset, and finally 18 MSA, 15 PSP, and 80 ET subjects were included.

[0046] (1) Acquisition of MRI sequences. All subjects were scanned using a 3.0T MRI scanner (GE Discovery 750) equipped with an 8-channel head coil. During the MRI scan, the head was fixed with a foam pad and earplugs were provided to reduce noise; the subjects remained awake with their eyes closed. If the image quality was affected by motion artifacts or other factors, the affected sequence was rescanned to ensure the best quality images for preprocessing. The scanning parameters of each MRI sequence are recorded as follows:

[0047] (1.1) High-resolution 3D T1-weighted images were acquired using a fast spoiled gradient sequence: repetition time = 7.336 ms; echo time = 3.036 ms; inversion time = 450 ms; flip angle = 11°; field of view = 260 × 260 mm 2 ; Matrix = 256 × 256; Slice thickness = 1.2 mm; Number of layers = 196.

[0048] (1.2) ESWAN images were acquired using a gradient echo sequence: repetition time = 33.7 ms; first echo time / interval / eighth echo time = 4.556 ms / 3.648 ms / 30.092 ms; flip angle = 20°; field of view = 240 × 240 mm 2 ; Matrix = 416 × 384; Slice thickness = 2 mm; Slice spacing = 0 mm; Number of layers = 64.

[0049] (1.3) ASL images were obtained using pulsed continuous arterial spin labeling (ASL): repetition time = 4632 ms; echo time = 10.536 ms; flip angle = 111°; field of view = 240 × 240 mm 2 ; Matrix = 128 × 128; Slice thickness = 4 mm; Slice spacing = 0 mm; Number of layers = 36.

[0050] (2) Preprocess the obtained T1-weighted images, ESWAN images and ASL images.

[0051] (2.1) The present invention uses the STAR-QSM algorithm in the STI Suite V3.0 software package to reconstruct the ESWAN data obtained in step (1) to obtain a QSM map (using the STISuite V3.0 software package developed by the University of California, Berkeley). In order to obtain the original magnetic susceptibility of each nucleus in the individual space, the DeepQSMSeg algorithm based on deep learning is used to segment the 10 subcortical nuclei in the individual space QSM map. The steps are as follows: Figure 1 A, specifically:

[0052] (2.1.1) The steps for reconstructing the QSM image mainly include: ① unwrapping the original phase map using Laplace-based phase unwrapping while retaining the spatial low-frequency components of all brain tissue; ② removing the background phase using spherical mean filtering to specifically obtain brain tissue phase data; ③ estimating the magnetic susceptibility boundary and removing magnetic artifacts; ④ obtaining the tissue magnetic susceptibility map, which is used to calculate the QSM image.

[0053] (2.1.2) An end-to-end deep learning-based tool (DeepQSMSeg algorithm) was used to segment subcortical nuclei. The DeepQSMSeg algorithm was developed by setting ground truth values ​​to historical data from previous manual and semi-automatic segmentations. It is primarily used to segment subcortical nuclei such as the substantia nigra and red nucleus. After automated segmentation, all segmented nucleus images were manually corrected by an experienced neuroradiologist to correct for misalignment of surrounding tissue voxels due to misregistration.

[0054] (2.1.3) After the nuclear segmentation was completed, the average magnetic susceptibility of the 10 subcortical nuclei, including the bilateral substantia nigra, red nucleus, putamen, caudate nucleus, and globus pallidus, was extracted.

[0055] (2.2) The cortical gray matter of the original individual T1-weighted image obtained in step (1) was segmented using the FreeSurfer open source software. The “recon-al” preprocessing process was used to segment the cortex of the individual spatial T1-weighted image, including automatic head motion correction, skull stripping, spatial normalization, spatial registration, and cortical segmentation. Based on the Desikan-Killiany atlas, 68 cortical regions (34 cortical regions in each hemisphere) and TIV ( Figure 2 B).

[0056] (2.3) Use SPM 12 to preprocess the ASL image obtained in step (1) and construct ASL_PDRP. The steps are as follows: Figure 1 B, specifically:

[0057] (2.3.1) The ASL images and the CBF images generated based on ASL were preprocessed using SPM 12. The Segment standard segmentation algorithm of SPM 12 performed tissue segmentation on the T1-weighted images of each subject and co-registered them to the MNI standard space to generate DARTEL gray matter and white matter templates based on the study population.

[0058] (2.3.2) Use the DARTEL toolkit to perform spatial registration and normalization of the ASL image and CBF image in the MNI standard space.

[0059] (2.3.3) The normalized CBF map was smoothed using a Gaussian kernel with a FWHM of 6 mm.

[0060] (2.3.4) Constructing the ASL_PDRP of CBF maps using the PCA-SSM method: Constructing individual subprofiles requires voxel-level analysis of the CBF maps of all participants in the training set, representing them as high-dimensional vectors and constructing a voxel matrix. The PCA method performs matrix centering and covariance matrix decomposition on this voxel matrix. Specifically, it performs log transformation and double decentered averaging. The resulting participant residual profile (SRP) image represents the deviation between the participant and group mean. Finally, covariance matrix decomposition is performed to output the eigenvectors (i.e., spatial patterns, representing the unique spatial distribution network of functionally related brain regions expressed to varying degrees for each subject) and scores (representing the degree of expression of the eigenvectors in the image data of a specific individual) corresponding to each principal component (PC). The linear combination of the top scores or individual PCs forms a proportional subprofile model (i.e., composite pattern), and the expression value (ASL_PDRP) of each subject on the proportional subprofile model is calculated.

[0061] (2.3.5) The individual standardized CBF map obtained by the TPR algorithm based on the independent validation set and the cross-disease dataset is multiplied and accumulated with the weight of the voxel in the proportional sub-profile model (i.e., the inner vector product of the participant's SRP vector and the final pattern vector) to obtain the individual's overall score on the proportional sub-profile model, which is used as the ASL_PDRP obtained based on TPR.

[0062] Among them, the individual standardized CBF map is a standardized image obtained by applying the above-mentioned pattern mask in the independent validation set and cross-disease dataset using the TPR algorithm, and also performing log transformation and double decentered averaging.

[0063] (3) Brain image feature screening and extraction.

[0064] (3.1) For the QSM image, the present invention calculated the average magnetic susceptibility values ​​of 10 subcortical nuclei as 10 brain image features, and preliminarily constructed a 369×10 feature matrix.

[0065] (3.2) For T1-weighted image segmentation, 68 cortical region thicknesses and 1 TIV feature were obtained in the bilateral cerebral hemispheres, initially forming a 369×69 feature matrix.

[0066] (3.3) For ASL images, ASL_PDRP or ASL_PDRP obtained based on TPR is used as a brain image feature alone, initially forming a 369×1 feature matrix.

[0067] (4) Brain image feature standardization: general linear model ( Figure 2 B).

[0068] Since age and gender are common confounding factors, the present invention uses the data of 122 NCs in the training set to construct a GLM to estimate the effects of age and gender of the normal population on each brain imaging feature:

[0069] Y~β+β age X age +β gender X gender

[0070] Among them, Y is the brain image feature, X age is age (continuous variable), X gender is gender (categorical variable), β, β age , β gender are the parameters of GLM.

[0071] For each brain feature Y i (122×1)(i=1~80), a set of β, β age , β gender , and the corresponding general linear models (80) were obtained. The constructed GLM was applied to the brain imaging features of all participants (including PD, PSP, MSA, and ET), and the standard brain imaging feature matrices (80×369, 80×93, 80×51, 80×48, and 80×113) were obtained after eliminating the effects of age and gender. The brain imaging features after eliminating the effects were then z-valued [-1, 1] using SPSS 25.0.

[0072] (5) Feature selection and model construction: LASSO analysis was performed and ten-fold cross validation was performed using the R language package “glmnet” to select the optimal regularization parameter, perform brain imaging feature selection, and build a PD diagnostic model ( Figure 2 B), as follows:

[0073] (5.1) Using the standardized brain image features preprocessed in step (4), LASSO analysis was performed and 10-fold cross validation was used to analyze the difference between the model prediction results and the real data, and then the 80 brain image features processed in step (4) were sorted and screened by deviation (Deviance) Figure 3 A); and select the 20 most important (i.e., smallest deviation) brain image features as the training set features input into the subsequent classification model ( Figure 3 B).

[0074] (5.2) Use logistic regression analysis to establish a diagnostic model. Figure 2 D. The present invention divides all PD and NC involved in this project into a training set (for training the classification model) and an independent validation set (for internal validation) according to the inclusion time point (around November 2021), with a split ratio of 4:1. In this embodiment: 247 PD patients and 122 NC included before November 2021 are used as the training set; 60 PD patients and 33 NC included after November 2021 are used as the independent validation set. The training set is used for 1000 iterations of sampling to train the PD diagnostic model, so that the model can be stably applied to other independent disease data sets ( Figure 2 B). The input of the model is the 20 brain image features selected in step (5.1) of the corresponding dataset.

[0075] (6) Clinical efficacy analysis and cross-disease dataset validation

[0076] (6.1) To enhance the clinical interpretability of the PD diagnostic model, we performed calibration curve, DCA, and CIC analyses on the training set and independent validation set, respectively. We also used Pearson correlation analysis to explore the relationship between the predicted individual scores obtained based on the model and the severity of clinical symptoms (UPDRS I-III). Figure 2 D).

[0077] (6.2) In order to further verify the generalization ability of the obtained model of the present invention in independent validation sets and cross-disease datasets. The present invention first used 60 PD patients and 33 normal controls recruited between November 2021 and December 2023 to independently validate the constructed diagnostic model. In addition, the present invention also included 18 MSA, 15 PSP and 80 ET patients to perform cross-disease validation on the constructed PD diagnostic model ( Figure 2 C) All procedures were performed independently except for using the previously estimated GLM to eliminate the effects of age and sex.

[0078] (7) To further explain the model, the present invention conducted a single feature analysis and used SHAP analysis to screen the top 6 brain imaging features with the greatest contribution ( Figure 2 D), and ANOVA analysis of variance was used to compare the differences between groups in standardized brain imaging features in different data sets such as training set, independent validation set, MSA cohort, PSP cohort and ET cohort, and FDR correction was performed ( Figure 6 ).

[0079] The method for using the PD diagnostic model of the present invention is as follows: the T1-weighted image, ESWAN image and ASL image to be diagnosed are processed in sequence according to steps (2) to (3), and then the influence of age and gender are eliminated based on the GLM according to step (4). Then, according to step (5), brain image features with higher contribution are screened using LASSO regression analysis, and the features are input into the trained PD diagnostic model, which is then expanded to be applied to the diagnosis of different types of neurodegenerative diseases.

[0080] In summary, the present invention proposes a method for establishing a clinically interpretable PD diagnostic model based on multi-sequence magnetic resonance imaging (MRI) brain features, which provides an objective basis for the clinical diagnosis of PD and helps improve its diagnostic accuracy, with good clinical translational value. The present invention establishes a model construction process that integrates training, independent validation, and cross-disease validation, and calculates the model diagnostic efficacy index AUC value and its confidence interval through 1000 iterations. The model construction process includes brain imaging data preprocessing, feature standardization and screening, model construction, independent validation set, and cross-disease set validation. The research characteristics and innovation of the present invention lie in obtaining a set of brain function and metabolic features based on ASL_PDRP and substantia nigra iron distribution. This not only reaffirms the diagnostic effect of ASL_PDRP on PD in previous studies, but also clarifies the important role of substantia nigra iron metabolism abnormalities in PD disease, and verifies the effectiveness of ASL_PDRP combined with brain iron metabolism and brain structure-related features as imaging biomarkers for PD. Secondly, the clinical interpretability of the model is promoted by analysis such as calibration curves, DCA, CIC, and correlation analysis between the predicted individual scores obtained based on the model and clinical symptoms. Finally, the diagnostic efficacy of the PD diagnostic model constructed based on the above-mentioned brain imaging features in different disease sets showed that the model had different diagnostic efficacy in different diseases (PSP, MSA, ET), with good diagnostic efficacy for PSP and MSA, but limited diagnostic efficacy in ET, indicating that the model has potential cross-disease generalization capabilities.

Claims

1. A method for establishing a clinically interpretable PD diagnostic model based on multi-sequence magnetic resonance imaging (MRI) brain features, characterized in that: The following steps are involved: (1) Obtain T1-weighted images, ESWAN images, and ASL images of PD patients, normal controls, PSP patients, MSA patients, and ET patients. PD patients and normal controls were divided into training sets and independent validation sets in proportion, while PSP patients, MSA patients, and ET patients were used as cross-disease datasets. (2) The cortical gray matter of the T1-weighted image obtained in step (1) was segmented to obtain 68 cortical regions and 1 TIV; the ESWAN image obtained in step (1) was post-processed using the STAR-QSM algorithm to obtain a QSM map, and the QSM map was segmented using the DeepQSMSeg algorithm to obtain 10 subcortical nuclei in the bilateral substantia nigra, red nucleus, putamen, caudate nucleus, and globus pallidus; the ASL image obtained in step (1) was subjected to principal component analysis-proportional subsection model analysis to obtain the abnormal PD perfusion pattern, specifically: (2.2.1) The ASL images and the CBF images generated based on the ASL were preprocessed using SPM 12, including motion correction, spatial registration and normalization of the ASL and CBF images using DARTEL, CBF image normalization, and spatial smoothing using a Gaussian kernel with a full width at half maximum of 6 mm. (2.2.2) Construct the ASL_PDRP of the CBF map preprocessed in step (2.2.1) using principal component analysis-proportional sub-profile model analysis, including: constructing individual sub-profiles, principal component analysis, proportional sub-profile model analysis, and constructing PD-related cerebral perfusion patterns; The individual sub-profile construction performed voxel-level analysis on the CBF maps of all participants in the training set and expressed them as high-dimensional vectors. The principal component analysis performed matrix centering and covariance matrix decomposition to output the eigenvector and score corresponding to each principal component. The principal component with the highest score was selected to form a proportional sub-profile model, and the expression value ASL_PDRP of each participant on the proportional sub-profile model was calculated. (2.2.3) Calculate individual ASL_PDRP expression scores using topological profile scoring in the independent validation set and PSP, MSA, and ET patient datasets. Based on the scaled sub-profile model obtained from the training set in step (2.2.2), multiply the individual standardized CBF maps obtained from the independent validation set and the cross-disease dataset using the TPR algorithm with the weights of the corresponding voxels in the scaled sub-profile model and add them together to obtain the individual's overall score on the scaled sub-profile model, which is used as the TPR-derived ASL_PDRP. (3) Extract brain image features for each cortical region, subcortical nucleus, and PD abnormal perfusion pattern segmented in step (2); standardize the brain image features using a general linear model; (4) The LASSO method was used to select multiple brain imaging features with the highest contribution as input, and a binary logistic regression model was used to construct a PD diagnosis model, which was trained based on the training set; (5) According to step (4), the individual prediction scores obtained based on the PD diagnostic model were correlated with the corresponding clinical symptoms; SHAP analysis was used to screen the top 6 brain imaging features in terms of contribution, and inter-group differences were compared in the training set, independent validation set, PSP, MSA, and ET patient data sets, and FDR correction was performed.

2. The method for establishing a clinically interpretable PD diagnostic model based on multi-sequence MRI brain features as claimed in claim 1, characterized in that: In step (1), a rapid spoiler gradient sequence is used to obtain T1-weighted images; a gradient echo sequence is used to obtain ESWAN images; and a pulsed continuous arterial spin labeling technique is used to obtain ASL images.

3. The method for establishing a clinically interpretable PD diagnostic model based on multi-sequence MRI brain features as claimed in claim 1, characterized in that: In step (2), the cortical gray matter of the T1-weighted image obtained in step (1) is segmented to obtain 68 cortical regions. Specifically, the subcortical nuclei of the T1-weighted image are segmented using FreeSurfer software, and the recon-al preprocessing process is used to perform automated head motion correction, skull stripping, spatial normalization, spatial registration, and cortical segmentation. Then, based on the Desikan-Killiany atlas, 68 cortical regions and TIV are segmented. The ESWAN image obtained in step (1) is post-processed using the STAR-QSM algorithm to obtain a quantitative magnetic susceptibility map, specifically: the original phase map is unwrapped using Laplace-based phase unwrapping, and the spatial low-frequency components of all brain tissues are retained; the background phase is removed using spherical mean filtering, thereby specifically obtaining brain tissue phase data; the magnetic susceptibility boundary is estimated and magnetization artifacts are removed; Obtain tissue magnetic susceptibility maps to calculate QSM images; The DeepQSMSeg algorithm tool sets the true value to the historical data of manual and semi-automatic segmentation; After using the DeepQSMSeg algorithm to segment the QSM image to obtain the subcortical nuclei, it also includes: manual correction to correct the misalignment of surrounding tissue voxels caused by alignment deviation.

4. The method for establishing a clinically interpretable PD diagnostic model based on multi-sequence MRI brain features as claimed in claim 1, characterized in that: In step (3), the brain image features extracted for each cortical region, subcortical nucleus, and PD abnormal perfusion pattern segmented in step (2) are specifically: (3.1) For the QSM map, the average magnetic susceptibility values ​​of the 10 subcortical nuclei were calculated as 10 brain imaging features; (3.2) For T1-weighted images, 68 average values ​​of cortical thickness and 1 TIV were calculated based on the Desikan-Killiany atlas as 69 brain imaging features; (3.3) For ASL images, ASL_PDRP or ASL_PDRP based on TPR is used as a brain image feature; The general linear model is used to standardize the brain image features. Specifically, a general linear model is constructed based on the brain image features and corresponding age and gender of the normal control group in the training set: Y~β+β age X age +b gender X gender Among them, Y is the brain image feature, X age is age, X gender For gender, β, β age , β gender are the parameters of the general linear model; the constructed general linear model is applied to all participants in step (1) to obtain the estimated brain imaging features; The brain image features extracted in step (3) are applied to obtain brain image features that eliminate the influence of age and gender; the brain image features after eliminating the influence are then z-valued to [-1, 1].

5. The method for establishing a clinically interpretable PD diagnostic model based on multi-sequence MRI brain features as claimed in claim 1, characterized in that: The method of using the PD diagnostic model is as follows: the T1-weighted image, ESWAN image and ASL image to be diagnosed are processed in steps (2) to (3) in sequence, and the influence of age and gender are eliminated based on the general linear model; according to step (4), multiple brain image features with the highest contribution are screened using LASSO regression analysis, and the trained PD diagnostic model is input and expanded to be applied to the diagnosis of different types of neurodegenerative diseases.

6. The method for establishing a clinically interpretable PD diagnostic model based on multi-sequence MRI brain features according to claim 1, characterized in that: In step (4), based on the current brain image features and the corresponding classification labels, the features with higher contributions are selected. Specifically, LASSO analysis is performed, cross-validation is performed using the R language package glmnet, the optimal regularization parameter is selected, and the 20 most important brain image features are screened. The selected important brain image features are subjected to binary logistic regression analysis using the R language package glm to construct a PD diagnostic model.

7. The method for establishing a clinically interpretable PD diagnostic model based on multi-sequence MRI brain features according to claim 1, characterized in that: In step (5), based on the trained PD diagnostic model, the predict function of the R language package was used to calculate the individual prediction score, and further Pearson correlation analysis was performed with the UPDRS II and UPDRS III scores; SHAP analysis was used to screen the top 6 brain imaging features with the greatest contribution, and ANOVA analysis of variance was used to compare the differences between groups in different data sets, and FDR correction was performed.

8. The method for establishing a clinically interpretable PD diagnostic model based on multi-sequence MRI brain features as claimed in claim 1, characterized in that: The method also includes verifying the diagnostic ability of the trained PD diagnostic model across disease sets, specifically by applying the trained PD diagnostic model to the MSA, PSP, and ET patient datasets using the predict function of the R language package, evaluating the diagnostic efficacy using the Delong test, and performing FDR correction.