A multimodal VHD risk prediction system based on deep learning and CMR valve feature extraction

CN122575702APending Publication Date: 2026-08-14FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004](1)影像表型提取与利用不足:缺乏基于先进方法(如深度学习)的高通量、自动化表型提取策略,难以充分挖掘CMR数据中蕴含的结构及血流动力学信息

Benefits of technology

[0061]本发明相较于现有技术,在多模态数据融合方式、风险预测性能、自动化处理水平、模型可解释性以及临床应用价值等方面具有一定改进和优化效果,具体如下:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575702A_ABST
    Figure CN122575702A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of auxiliary diagnostic technology, specifically a multimodal VHD risk prediction system based on deep learning for CMR valve feature extraction. The invention first preprocesses the input multi-source data; then, it defines the cardiac valve structural phenotype based on cine MRI data, segments the target image region using deep learning methods, and extracts 25 cardiac valve structural and functional phenotypes through automated quantitative analysis; subsequently, it performs GWAS analysis on the image phenotypes, combined with multi-omics analysis methods, to screen for shared pleiotropic risk sites associated with VHD and mechanical stress; finally, it integrates CMR structural and functional phenotypes, individual genotypes of shared pleiotropic risk sites, and clinical non-imaging data to construct a multimodal fusion prediction model, predicting and assessing an individual's current and future risk of VHD. This invention not only achieves automated quantification of CMR structural phenotypes but also enables early risk assessment and stratification of VHD.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of auxiliary diagnostic technology, specifically a multimodal VHD risk prediction system. Background Technology

[0002] Valvular heart disease (VHD) is a significant cardiovascular disease worldwide, with a prevalence of approximately 2.5%, which increases significantly with age. In developed countries, VHD is primarily characterized by degenerative changes such as aortic stenosis and mitral regurgitation. Valvular dysfunction can lead to a prolonged state of pressure or volume overload on the heart, inducing a series of pathological changes such as myocardial hypertrophy, ventricular dilation, fibrosis, and microvascular ischemia. These structural and functional changes often occur before the onset of clinical symptoms, but conventional diagnostic methods struggle to identify them in a timely manner. Meanwhile, genome-wide association studies (GWAS) have revealed numerous genetic loci related to cardiac structure, function, and hemodynamics in the cardiovascular field, indicating that genetic factors play a crucial role in cardiac development and disease progression.

[0003] Currently, echocardiography is the primary clinical method for diagnosing VHD, but it still has limitations in early risk identification and quantitative assessment. Cardiac magnetic resonance imaging (CMR), due to its superior tissue resolution and blood flow quantification capabilities, is widely used for assessing cardiac structure and function. In recent years, CMR-based imaging phenotypes (such as valvular structural phenotypes and hemodynamic parameters) have been shown to reflect early cardiac remodeling processes and have potential risk prediction value. Although high-precision information on cardiac structure and function can be obtained from CMR images, a systematic methodological framework is still lacking in how to efficiently and stably extract biologically meaningful quantitative phenotypes from imaging data, and how to further integrate genetic background and clinical characteristics for disease risk assessment. This deficiency, to some extent, restricts the effective application of relevant information in VHD risk prediction. Therefore, current technologies still have the following limitations in VHD risk prediction:

[0004] (1) Insufficient extraction and utilization of image phenotypes: There is a lack of high-throughput, automated phenotype extraction strategies based on advanced methods (such as deep learning), making it difficult to fully explore the structural and hemodynamic information contained in CMR data.

[0005] (2) Insufficient cross-level integration analysis: There is a lack of a unified analytical framework between imaging phenotypes, genetic variations and clinical features. A systematic connection from "imaging phenotypes - genetic mechanisms - clinical manifestations" has not yet been established, and there is a lack of biological explanatory power.

[0006] (3) Insufficient construction of multimodal prediction models: Existing methods mostly rely on a single data source and lack a comprehensive modeling strategy that integrates imaging phenotypes, genetic information and clinical characteristics, which limits the accuracy and generalization ability of risk prediction.

[0007] (4) Limited ability to identify early risks: Most studies focus on people who already have the disease, making it difficult to achieve risk stratification and prediction before the onset of the disease based on early imaging phenotypes and genetic information.

[0008] Therefore, there is an urgent need to develop a multimodal prediction system that can integrate CMR imaging phenotypes, genetic information and clinical characteristics to achieve early identification, risk stratification and analysis of potential mechanisms of VHD. Summary of the Invention

[0009] To address the shortcomings of the existing technologies, the present invention aims to propose a multimodal VHD risk prediction system based on deep learning and CMR valve feature extraction, so as to achieve early identification, risk stratification, and analysis of potential mechanisms of VHD.

[0010] The multimodal VHD risk prediction system provided by this invention includes: a data input module, a data preprocessing module, an image phenotype extraction module, a genetic information feature construction module, a multimodal feature fusion module, and a VHD risk prediction output module. The modules collaborate and exchange information through a data interface.

[0011] The data input module is used to acquire multi-source data, including CMR imaging data, genetic data, and clinical non-imaging data. CMR imaging data includes cardiac cine MRI and phase-contrast MRI (PC-MRI), used to characterize dynamic changes in cardiac structure and hemodynamic information, respectively. Genetic data includes single nucleotide polymorphism genotype data, gene expression prediction information, and protein-related genetic variation information, used for GWAS, expression quantitative trait locus analysis (eQTL), transcriptome association analysis (TWAS), proteome association analysis (PWAS), Mendelian randomization analysis (MR), and the construction of a multi-gene risk score (PRS). Clinical non-imaging data includes clinical phenotypic information such as age, sex, body mass index, waist-to-hip ratio, systolic blood pressure, blood lipid levels, glycated hemoglobin, and inflammatory markers, used to assist in disease risk assessment and model training.

[0012] The data preprocessing module is used to standardize and perform quality control on multimodal data, specifically including:

[0013] For cinema MRI and PC-MRI image data, grayscale normalization is first performed. For PC-MRI data, velocity encoding correction and background phase error correction are further performed to improve the accuracy of hemodynamic parameter measurement.

[0014] For genetic data, standard genetic data quality control procedures were employed to remove data with low minor allele frequencies (MAF < 1%), poor genotype imputation quality (INFO < 0.3), high deletion rates, or significant deviations from Hardy-Weinberg equilibrium (P < 1×10⁻⁻⁴). 6 The variant sites were identified. Simultaneously, samples with excessively high genotype loss rates, inconsistent sex information, abnormal heterozygosity, or recessive kinship (pi-hat > ​​0.2) were excluded. For candidate SNP sites, genotypes were converted into numerical variables (0 / 1 / 2) based on the number of "minor alleles" carried by each individual.

[0015] For clinical non-imaging data, missing values ​​were imputed (median imputed for continuous variables, mode imputed for categorical variables, with a missing value rate of <2%), and Z-score standardization was performed to eliminate the influence of dimensions.

[0016] The image phenotype extraction module includes a cine MRI image phenotype extraction unit and a PC-MRI blood flow phenotype extraction unit, used to extract structural and functional phenotypes from CMR images, including cardiac structural phenotypes and hemodynamic phenotypes; wherein:

[0017] The cinematic MRI image phenotype extraction unit, based on the nnU-Net deep learning segmentation network, automatically segments cardiac structures and extracts structural phenotypes based on keyframes of the cardiac cycle. The process includes segmentation and structural phenotype calculation, wherein:

[0018] The segmentation process begins by dividing the cardiac cycle. The cine MRI data consists of 50 frames, with end-systolic (ES), end-diastolic (ED), mid-systolic (MS), and mid-diastolic (MD) defined by the clinician as frames 20, 1, 10, and 34, respectively. Fifteen structural phenotypes were extracted from different chamber views (see Table 1), specifically: 2-chamber mitral annulus end-diastolic diameter, 2-chamber mitral annulus end-systolic diameter, 3-chamber mitral annulus end-diastolic diameter, and 3-chamber... The end-systolic diameter of the mitral annulus, the mid-systolic diameter of the mitral annulus in the three chambers, the mitral valve elevation height, the mitral valve elevation area, the length of the anterior leaflet of the mitral valve, the angle α between the anterior leaflet and the annulus diameter, the angle β between the posterior leaflet and the annulus diameter, and the angle θ between the elevations of the anterior and posterior leaflets of the mitral valve, the end-systolic diameter of the tricuspid annulus in the four chambers, the end-diastolic diameter of the mitral annulus in the four chambers, and the end-diastolic diameter of the tricuspid annulus in the four chambers. Clinicians then manually annotated 200 cine MRI images, with a model training-to-test ratio of 150:50. Finally, a total of 8 models were trained to extract 15 structural phenotypes, and the Dice similarity coefficient (DSC), intersection-over-union ratio (IoU), and 95% Hausdorff distance (HD) were used.95 The model's segmentation performance is quantitatively evaluated. DSC and IoU are used to measure the overlap between the predicted segmentation and the ground truth annotation, with values ​​ranging from 0 to 1. Higher values ​​indicate higher segmentation accuracy. The formulas for DSC and IoU are:

[0019] , (1)

[0020] , (2)

[0021] in, Indicates the predicted segmentation region. Indicates the actual labeled area. Indicates the number of pixels within the area.

[0022] HD 95 Used to evaluate the spatial distance between the predicted boundary and the true boundary, it is defined as the 95th percentile of the Hausdorff distance. This effectively reduces the impact of outliers; a smaller value indicates more accurate segmentation. The formula is:

[0023] (3)

[0024] in, Represents the set of points for the predicted segmentation region. Represents the set of points in the actual labeled region. This represents the Euclidean distance between two points. HD 95 The 95th percentile of the two-way Hausdorff distance is used to reduce the impact of outliers on the results.

[0025] The structural phenotype calculation process categorizes phenotype extraction into the following four types:

[0026] (a) Height of the bulge: the distance from the junction of the ventricle and atrium.

[0027] (b) Angle of elevation: determined by the three vertices of the elevation area, where the vertices of ∠α and ∠β, where the anterior and posterior leaflets of the mitral valve are located, are determined according to the endpoints of the junction between the ventricle and the atrium, and the vertex of ∠θ is the point farthest from the vertices of ∠α and ∠β.

[0028] (c) Valve annulus diameter: the length of the boundary between the ventricle and the atrium.

[0029] (d) Area of ​​bulge and length of leaflet: calculated directly from the results of model segmentation.

[0030] The PC-MRI blood flow phenotype extraction unit is used for quantitative analysis of aortic hemodynamics. The DeepFlow blood flow analysis tool based on deep learning is used to extract PC-MRI blood flow functional and structural phenotypes, a total of 10 (see Table 1), specifically: maximum aortic diameter, minimum aortic diameter, aortic area, aortic circumference, maximum forward blood flow velocity, minimum forward blood flow velocity, aortic regurgitation volume, aortic regurgitation fraction, stroke volume, and net flow rate.

[0031] The genetic information feature construction module is used to construct genetic risk features related to imaging phenotypes. This includes: performing genetic analysis on the imaging phenotypes obtained by the imaging phenotype extraction module using the GWAS method to screen for genetic variation sites with genome-wide significance; identifying genetic sites or related genes that reach genome-wide significance and show significance in at least two omics association analyses as pleiotropic candidate genes related to VHD; further identifying key genetic sites related to heart valve structure, function, and the occurrence of VHD to obtain a set of candidate risk sites; and extracting genotypic information of individuals at the aforementioned key risk sites to obtain a set of genetic risk features.

[0032] Furthermore, the specific process is as follows:

[0033] First, a GWAS method was used to perform genetic analysis on the image phenotypes obtained from the image phenotype extraction module (cine MRI structural phenotype + PC-MRI blood flow functional and structural phenotypes). Specifically, under an additive genetic model, linear regression analysis was used to assess the association between single nucleotide variants and phenotypes. Covariate corrections were applied to age, sex, body mass index, body surface area, waist-to-hip ratio, and principal components of population structure to screen for genetic variants that reached genome-wide significance (P < 5 × 10⁻⁶). -8 Then, the linkage disequilibrium score regression (LDSC) method was used to estimate the heritability of the imaging phenotypes and their pairwise genetic correlations to assess the shared genetic basis of different phenotypes. Subsequently, functional annotation and gene mapping analysis were performed based on significant genetic variation sites, and candidate variation sites were mapped to potential functional genes through location mapping and gene set analysis methods.

[0034] Furthermore, based on the eQTL prediction model, gene expression levels mediated by genetic variations were inferred, and the association between gene or protein expression and imaging phenotypes was assessed using TWAS and PWAS. Genetic loci or related genes that reached genome-wide significance and showed significance in at least two omics association analyses were identified as pleiotropic candidate genes associated with VHD. In addition, pathway enrichment and tissue-specific expression analyses were performed on the screened pleiotropic candidate genes to elucidate their potential biological functions, and fine mapping analysis was used to prioritize the screening of possible causal variation sites. Finally, key genetic loci associated with heart valve structure, function, and the occurrence of VHD were identified, resulting in a set of candidate risk sites.

[0035] Furthermore, considering that CMR imaging phenotypes reflect the mechanical properties of heart valves and hemodynamics (such as shear stress and pressure load) to a certain extent, a subset of genetic loci related to mechanical stress was further screened from the above loci to enhance the biological consistency between genetic characteristics and imaging phenotypes and improve the explanatory power of the model.

[0036] Furthermore, genotypic information of individuals at the aforementioned key risk sites is extracted to construct a set of genetic characteristics.

[0037] Optionally, a weighted genetic risk score (GRS) can be constructed based on the effect size at each point, and its calculation method is as follows:

[0038] , (4)

[0039] Where βᵢ represents the effect size of the i-th genetic locus, and Gᵢ represents the genotype code of the corresponding locus (0, 1, or 2, representing the copy number of the minor allele, respectively). By constructing GRS, the dimensionality of features can be reduced while preserving genetic information, and the risk characterization ability of genetic features can be enhanced.

[0040] The multimodal feature fusion module is used to fuse image phenotypic features, genetic features, and clinical non-imaging features. It includes population segmentation, model selection, feature fusion, and class imbalance feature selection, wherein:

[0041] Population segmentation divides the population with VHD into two groups: those who already have VHD and those who will develop VHD in the future. The definitions are as follows:

[0042] (a) VHD group: VHD was present at the time of CMR scan.

[0043] (b) Future VHD group: those who were not diagnosed with VHD at the time of CMR scan but will be diagnosed with VHD in the future.

[0044] Model selection involved inputting the extracted 25 image phenotypes into various machine learning models for modeling. The machine learning models included Support Vector Machine (SVM), Multilayer Perceptron (MLP), Gradient Boosting Decision Tree (GBDT), Random Forest (RF), and XGBoost models. The model with the best VHD prediction performance was used for modeling multimodal features.

[0045] Feature fusion integrates extracted image phenotypic features, genetic features, and clinical non-imaging features to construct a unified feature input matrix. The feature input matrix is ​​represented as follows:

[0046] , (5)

[0047] Among them, X img Representing CMR image phenotypic characteristics, X geno X represents a hereditary trait. clin This represents clinical, non-imaging features. Continuous variables in the multimodal features are standardized, and categorical variables are numerically coded.

[0048] In implementation, feature fusion can also employ deep learning-based feature representation and fusion methods to embed and encode image phenotypic features, genetic features, and clinical non-image features respectively, and achieve multimodal fusion through feature splicing or attention mechanisms to further improve prediction performance.

[0049] Imbalanced class feature selection involves filtering majority class samples in the training set using a downsampling strategy based on K-Means clustering. Specifically, the K-Means clustering is achieved by minimizing the within-class squared error, and its objective function is:

[0050] , (6)

[0051] Where J is the sum of squared errors within the class, and C k Represented as the first One cluster, Indicates the corresponding cluster center. This is represented as a data point belonging to that cluster. Then, stratified sampling is performed according to the sample distribution ratio of each cluster, so that the majority class samples after downsampling maintain the original distribution characteristics in the feature space, thereby avoiding information loss.

[0052] The VHD risk prediction output module is used to output the individual's VHD risk probability and risk stratification results. It includes model training, model interpretation, VHD disease prediction, and risk stratification, wherein:

[0053] Model training involved dividing the dataset into training and test sets in a 7:3 ratio, optimizing model parameters using a 10-fold cross-validation strategy, and selecting the model based on the area under the receiver operating characteristic (AUC) curve as the primary evaluation metric. The optimal model was selected to output the risk probability of VHD in individuals with existing VHD and individuals with potential VHD. The performance of the VHD disease prediction model based on multimodal features and ablation experiments are shown in Appendix Table 2.

[0054] Model interpretation employs a Shapley value-based (SHAP) method to interpret the model output, quantifying the contribution of each input feature to the prediction result.

[0055] VHD Disease Prediction and Risk Stratification: Disease prediction involves inputting the feature data of individuals in the test set into the optimal prediction model to obtain the corresponding VHD risk prediction probability. The model output is... , representing the probability of an individual experiencing VHD given feature X. Here, X represents the input multimodal feature set. Based on the model's output probability, the optimal classification threshold is determined by iterating through different classification thresholds and calculating the corresponding performance metrics (including F1 score, precision, accuracy, and recall). The optimal classification threshold T* is determined by maximizing the F1 score, i.e.:

[0056] , (7)

[0057] In the formula, T represents the classification threshold, and F1(T) represents the F1 score calculated at threshold T. By iterating through different thresholds and calculating the corresponding F1 scores, the threshold that maximizes the F1 score is selected as the optimal classification threshold T*, thus achieving a balance between precision and recall. Based on the optimal threshold, individuals are divided into high-risk and low-risk groups.

[0058] Risk stratification involves constructing a Cox proportional hazards model based on follow-up data, performing survival analysis on different risk groups, calculating their risk ratio for future VHD events, thereby validating the model's effectiveness in predicting diseases using longitudinal data, and achieving dynamic assessment of individual risk.

[0059] The system also includes a processor and a memory; the processor is used to perform CMR image analysis, genetic modeling, and risk prediction calculations; the memory is used to store computer-executable instructions to enable the processor to execute the above method steps. This invention can be implemented in software or combined with hardware acceleration units such as GPUs to improve computational efficiency.

[0060] The main technical features and beneficial effects of this invention are as follows:

[0061] Compared with existing technologies, this invention has certain improvements and optimizations in terms of multimodal data fusion methods, risk prediction performance, automation level, model interpretability, and clinical application value, as detailed below:

[0062] (1) This invention utilizes deep learning methods to automate the analysis of CMR images, achieving high-throughput quantitative extraction of cardiac structural and hemodynamic phenotypes. Specifically, a CMR image phenotype system comprising 15 structural phenotypes and 10 functional phenotypes is constructed, covering multidimensional information such as valve morphology, cardiac chamber geometry, and aortic hemodynamic phenotypes. Compared to traditional methods relying on manual measurement or semi-automatic analysis, this invention can stably extract highly consistent quantitative indicators under multiple cardiac cycles and multiple chamber perspectives, significantly improving image analysis efficiency and repeatability. Simultaneously, this automated process effectively reduces subjective errors introduced by human operation, thereby enhancing the transferability and generalizability of the model in large-scale population studies and real-world clinical scenarios. This high-throughput image phenotype extraction capability provides a stable and high-quality data foundation for subsequent multimodal fusion modeling.

[0063] (2) This invention constructs a multimodal joint modeling framework that integrates CMR image phenotypes, genetic variation information, and clinical non-image data to achieve unified representation and collaborative learning of multi-source data. Image phenotypes reflect the structural and functional state of the heart; genetic information reflects an individual's innate susceptibility; and clinical features reflect acquired influencing factors such as environmental exposure and metabolic status. Through multi-branch feature encoding and fusion mechanisms, this invention can effectively mine complementary information and potential nonlinear correlations between different modalities. Compared to traditional methods based solely on a single image feature or clinical variable, this invention demonstrates superior discriminative ability and generalization performance in VHD risk prediction tasks. In the embodiments, the multimodal model AUC reaches approximately 0.77, indicating good performance in risk discrimination, especially with significant application value in the early stages of disease identification.

[0064] (3) Traditional VHD diagnostic methods mainly rely on imaging assessments after structural changes or clinical symptoms appear, making it difficult to identify potential high-risk individuals in a timely manner. This invention integrates CMR structural phenotypes (such as mitral valve annulus diameter, tricuspid valve annulus diameter, and bulge area), functional hemodynamic phenotypes (such as regurgitation fraction and blood flow velocity), genetic risk information, and clinical non-imaging characteristics to achieve a comprehensive characterization of the cardiac structure, function, and genetic susceptibility of VHD patients. The predictive model constructed based on the above multidimensional features can identify individuals who have not yet shown obvious structural abnormalities but already have functional abnormalities or high-risk genetic characteristics. In follow-up data validation, the model of this invention can effectively predict the risk of future VHD before a definitive clinical diagnosis and achieve risk stratification, thereby enabling prospective identification of the disease. This method helps to shift the clinical intervention window forward to before irreversible structural remodeling occurs, providing a basis for early prevention and delaying disease progression.

[0065] (4) This invention not only constructs a predictive model but also enhances overall interpretability through multi-level genetic analysis and model interpretation methods. Specifically, key genetic loci are screened through GWAS, TWAS, and eQTL co-location analysis to support the association between imaging phenotypes and VHD risk at the genetic level; the contribution of genetic information in risk prediction is verified to improve the biological rationality of the model; and the contribution of each imaging feature to the prediction results is quantified by combining the SHAP method to achieve transparent interpretation of the model. The analysis results show that structural imaging phenotypes contribute more in the already diseased population, while functional hemodynamic phenotypes play a stronger role in predicting future disease risk. This hierarchical interpretation mechanism not only improves the credibility of the model but also reveals the potential biological basis of VHD occurrence and development from both imaging and genetic perspectives, breaking through the limitations of traditional "black box models".

[0066] (5) The model of this invention can output individual risk probabilities and achieve refined risk stratification based on these probabilities. By setting an optimal classification threshold, individuals can be divided into high-risk and low-risk groups, and the long-term disease risk of different risk groups can be assessed using the Cox proportional hazards model. This method provides clinicians with a quantitative and interpretable risk assessment tool, which can be used for priority screening and follow-up management of high-risk populations; optimization of individualized intervention timing; and formulation of long-term disease management strategies. This promotes the transformation of clinical practice from the traditional passive diagnosis model to an active prevention and risk management model.

[0067] (6) This invention constructs an end-to-end automated system architecture covering data input, data preprocessing, image phenotype extraction, genetic feature construction, multimodal fusion, and risk prediction output. Each module has clear data interfaces and functional boundaries. Compared to existing data analysis workflows that rely on multiple tools for decentralized processing or manual integration, this invention has higher system integration and engineering feasibility, facilitating standardized deployment in clinical environments. Furthermore, this multimodal fusion framework adopts a modular design, possessing excellent scalability. It can flexibly introduce new data types (such as proteomics, metabolomics, or wearable device data) or replace and upgrade existing model structures, thereby continuously improving system performance. This framework has good versatility and can be further extended to risk prediction tasks for other cardiovascular diseases such as cardiomyopathy and heart failure, showing broad prospects for clinical promotion and application. Attached Figure Description

[0068] Figure 1 This is a block diagram of the multimodal VHD risk prediction system based on deep learning and CMR valve feature extraction according to the present invention. Detailed Implementation

[0069] The present invention will be further described below through examples.

[0070] This embodiment provides a multimodal VHD risk prediction system based on deep learning and CMR valve feature extraction. The specific workflow is as follows:

[0071] 1. Multi-source data acquisition and data standardization preprocessing

[0072] First, multi-source data from the UK Biobank was acquired through the data input module, including CMR imaging data, genetic data, and clinical non-imaging data. The CMR imaging data included cardiac cine MRI and PC-MRI, used to characterize cardiac structural changes and hemodynamic information, respectively. The input data underwent standardization processing through the data preprocessing module, including CMR image normalization and blood flow correction, genetic data quality control and genotype encoding, and missing value imputation and standardization of clinical variables, to improve the consistency of the multi-source data and the stability of model training.

[0073] 2. Definition and Extraction of CMR Image Phenotypes

[0074] After data preprocessing, the CMR image data is automatically analyzed using an image phenotype extraction module to extract image phenotypes reflecting valve structure and function. This embodiment extracts a total of 25 CMR structural and functional phenotypes, defined in the table below:

[0075] Table 1. Definitions of 25 CMR structural and functional phenotypes

[0076] .

[0077] 3. Screening and modeling of genetic risk traits

[0078] After obtaining 25 CMR structural and functional imaging phenotypes, genetic association analysis (GRA) was conducted using a genetic information feature construction module to elucidate their potential genetic basis and construct genetic characteristics. Specifically, GWAS was performed based on the 25 valve imaging phenotypes extracted from CMR to assess the association between genetic variations and imaging phenotypes. LDSC analysis was combined with this to analyze the heritability and genetic relevance of the imaging phenotypes. The results showed that valve imaging phenotypes exhibit significant polygenic inheritance characteristics and share some genetic basis with VHD. At the genome-wide significance level, genetic loci associated with at least two imaging phenotypes were further screened, identifying 70 aortic valve-related loci, 31 mitral valve-related loci, and 9 tricuspid valve-related loci. Subsequently, multi-omics integration analysis was performed using TWAS, PWAS, and eQTL data, ultimately identifying 34 key risk loci and candidate functional genes supported by multi-omics evidence. Based on this, 20 core genetic loci related to mechanical stress and hemodynamic regulation were further screened to construct GRS, enhancing the biological consistency and disease risk characterization ability between genetic characteristics and valve imaging phenotypes.

[0079] 4. Multimodal feature fusion and model optimization

[0080] The multimodal feature fusion module sequentially performs group segmentation, model selection, feature fusion, and class imbalance feature selection on the input data. Finally, the best-performing RF model is selected for training.

[0081] 5. VHD Risk Prediction

[0082] The VHD risk prediction output module inputs the individual's multimodal feature data into the optimal prediction model to obtain the corresponding VHD risk prediction probability P. Ablation experiments showed that the multimodal model integrating imaging phenotype, genotype, demographics, anthropometry, and metabolite characteristics achieved AUCs of 0.775 and 0.720 in the VHD-affected group and the future VHD-affected group, respectively, significantly outperforming the single-modal model. The performance of the VHD disease prediction model based on multimodal features and the ablation experiments are shown in Table 2. Furthermore, a Cox proportional hazards model was constructed based on follow-up data to calculate the hazard ratio for future VHD events (HR=3.21, 95% CI: 2.42–4.27, p = 8.73 × 10⁻⁶). -16 This enables dynamic assessment of individual risks.

[0083] SHAP analysis results indicate that, besides the most important age characteristic, the structural phenotype of CMR is more important in the pre-existing disease population, while the functional hemodynamic phenotype of CMR contributes more significantly to predicting future disease risk. Integrating genetic characteristics can further reveal the potential link between imaging phenotypes and genetic regulation.

[0084] Table 2. Performance of VHD disease prediction model and ablation experiment results

[0085] .

Claims

1. A multimodal VHD risk prediction system, characterized in that, include: The system comprises a data input module, a data preprocessing module, an image phenotype extraction module, a genetic information feature construction module, a multimodal feature fusion module, and a VHD risk prediction output module. These modules collaborate and exchange information via data interfaces. The data input module is used to acquire multi-source data, including CMR imaging data, genetic data, and clinical non-imaging data. The CMR imaging data includes cardiac cine MRI and phase-contrast MRI (PC-MRI), used to characterize dynamic changes in cardiac structure and hemodynamic information, respectively. The genetic data includes single nucleotide polymorphism genotype data, gene expression prediction information, and protein-related genetic variation information, used for GWAS, expression quantitative trait locus analysis (eQTL), transcriptome association analysis (TWAS), proteome association analysis (PWAS), Mendelian randomization analysis (MR), and the construction of a multi-gene risk score (PRS). The clinical non-imaging data includes clinical phenotypic information such as age, sex, body mass index, waist-to-hip ratio, systolic blood pressure, blood lipid levels, glycated hemoglobin, and inflammatory markers, used to assist in disease risk assessment and model training. The data preprocessing module is used to standardize and perform quality control on multimodal data, specifically including: For cinema MRI and PC-MRI image data, grayscale normalization is first performed. For PC-MRI data, velocity encoding correction and background phase error correction are further performed to improve the accuracy of hemodynamic parameter measurement. For genetic data, a genetic data quality control process was adopted, removing data with minor allele frequencies (MAF) < 1%, genotype imputation quality (INFO) < 0.3, high deletion rates, or significant deviations from Hardy-Weinberg equilibrium (P < 1×10⁻⁶). -6 The variant sites; at the same time, samples with excessively high genotype deletion rates, inconsistent sex information, abnormal heterozygosity, or recessive kinship pi-hat > ​​0.2 were removed; For clinical non-imaging data, missing values ​​were imputed and Z-score normalization was performed to eliminate the influence of dimensions. The image phenotype extraction module includes a cine MRI image phenotype extraction unit and a PC-MRI blood flow phenotype extraction unit, used to extract structural and functional phenotypes from CMR images, including cardiac structural phenotypes and hemodynamic phenotypes; wherein: The cinematic MRI image phenotype extraction unit, based on the nnU-Net deep learning segmentation network, automatically segments the heart structure and extracts the structural phenotype based on key frames of the cardiac cycle; it includes a segmentation process and a structural phenotype calculation process. The PC-MRI blood flow phenotype extraction unit is used to perform quantitative analysis of aortic hemodynamics, and uses DeepFlow, a blood flow analysis tool based on deep learning, to extract PC-MRI blood flow functional and structural phenotypes. The genetic information feature construction module is used to construct genetic risk features related to image phenotypes; including: performing genetic analysis on the image phenotypes obtained by the image phenotype extraction module using the GWAS method to screen for genetic variation sites with genome-wide significance; identifying genetic sites or related genes that reach genome-wide significance and show significance in at least two omics association analyses as pleiotropic candidate genes related to VHD; further identifying key genetic sites related to heart valve structure, function, and VHD occurrence to obtain a set of candidate risk sites; and extracting genotypic information of individuals at the above key risk sites to obtain a set of genetic risk features. The multimodal feature fusion module is used to fuse imaging phenotypic features, genetic risk features, and clinical non-imaging features; it includes population segmentation, model selection, feature fusion, and class imbalance feature selection, wherein: Population segmentation divides the population with VHD into two groups: those who already have VHD and those who will develop VHD in the future. The definitions are as follows: (a) VHD group: those who had VHD at the time of CMR scan; (b) Future VHD group: those who were not diagnosed with VHD at the time of CMR scan but will be diagnosed with VHD in the future; Model selection involves inputting the extracted image phenotypes into various machine learning models for modeling. The machine learning models include Support Vector Machine (SVM), Multilayer Perceptron (MLP), Gradient Boosting Decision Tree (GBDT), Random Forest (RF), and XGBoost models. The model with the best VHD prediction performance is used for modeling multimodal features. Feature fusion integrates extracted image phenotypic features, genetic features, and clinical non-imaging features to construct a unified feature input matrix; the feature input matrix is ​​represented as follows: , Among them, X img Representing CMR image phenotypic characteristics, X geno X represents a genetic trait. clin It represents clinical non-imaging features; it standardizes continuous variables in multimodal features and numerically encodes categorical variables; Imbalanced class feature selection involves filtering majority class samples in the training set using a downsampling strategy based on K-Means clustering. Specifically, the K-Means clustering is achieved by minimizing the within-class squared error, and its objective function is: , Where J is the sum of squared errors within the class, and C k Represented as the first One cluster, Indicates the corresponding cluster center. The data points are represented as belonging to the cluster; then, stratified sampling is performed according to the sample distribution ratio of each cluster, so that the majority class samples after downsampling maintain the original distribution characteristics in the feature space, thereby avoiding information loss. The VHD risk prediction output module is used to output the probability of individual VHD occurrence and the risk stratification results.

2. The multimodal VHD risk prediction system according to claim 1, characterized in that, The cinema MRI image phenotype extraction unit in the image phenotype extraction module includes a segmentation process and a structural phenotype calculation process; wherein: The segmentation process is as follows: The cardiac cycle is divided; the cine MRI data consists of 50 frames, with end-systolic (ES), end-diastolic (ED), mid-systolic (MS), and mid-diastolic (MD) frames defined by the clinician as frames 20, 1, 10, and 34, respectively; 15 structural phenotypes are extracted from different chamber views, specifically: 2-chamber mitral annulus end-diastolic diameter, 2-chamber mitral annulus end-systolic diameter, 3-chamber mitral annulus end-diastolic diameter, 3-chamber mitral annulus end-systolic diameter, 3-chamber mid-systolic mitral annulus diameter, mitral valve bulge height, and mitral valve bulge area. The following parameters were measured: anterior leaflet length, angle α between the anterior leaflet and the annulus diameter, angle β between the posterior leaflet and the annulus diameter, angle θ between the bulges of the anterior and posterior leaflets, end-systolic diameter of the four-chamber mitral annulus, end-systolic diameter of the four-chamber tricuspid annulus, end-diastolic diameter of the four-chamber mitral annulus, and end-diastolic diameter of the four-chamber tricuspid annulus. Clinicians then manually annotated 200 cine MRI images, with a model training-to-test ratio of 150:

50. Finally, a total of eight models were trained to extract 15 structural phenotypes, and the Dice similarity coefficient (DSC), intersection-over-union ratio (IoU), and 95% Hausdorff distance (HD) were used. 95 The model's segmentation performance is quantitatively evaluated; DSC and IoU are used to measure the overlap between the predicted segmentation and the ground truth annotation, with values ​​ranging from 0 to 1. Higher values ​​indicate higher segmentation accuracy. The formulas for DSC and IoU are: , , in, Indicates the predicted segmentation region. Indicates the actual labeled area. Indicates the number of pixels within the area; HD 95 Used to evaluate the spatial distance between the predicted boundary and the true boundary, it is defined as the 95th percentile of the Hausdorff distance, which can effectively reduce the influence of outliers. The smaller the value, the more accurate the segmentation result. The formula is: , in, Represents the set of points for the predicted segmentation region. Represents the set of points in the actual labeled region. HD represents the Euclidean distance between two points. 95 The 95th percentile of the two-way Hausdorff distance is used to reduce the impact of outliers on the results; The structural phenotype calculation process is as follows: phenotype extraction is divided into the following four categories: (a) Height of the bulge: the distance from the junction of the ventricle and atrium. (b) Angle of elevation: determined by the three vertices of the elevation area, wherein the vertices of ∠α and ∠β, where the anterior and posterior leaflets of the mitral valve are located, are determined according to the endpoints of the junction between the ventricle and the atrium, and the vertex of ∠θ is the point farthest from the vertices of ∠α and ∠β. (c) Valve annulus diameter: the length of the boundary between the ventricle and atrium; (d) Area of ​​bulge and length of leaflet: calculated directly from the results of model segmentation.

3. The multimodal VHD risk prediction system according to claim 2, characterized in that, The PC-MRI blood flow phenotype extraction unit in the image phenotype extraction module extracts a total of 10 PC-MRI blood flow functional and structural phenotypes, specifically: maximum aortic diameter, minimum aortic diameter, aortic area, aortic circumference, maximum forward blood flow velocity, minimum forward blood flow velocity, aortic regurgitation volume, aortic regurgitation fraction, stroke volume, and net flow rate.

4. The multimodal VHD risk prediction system according to claim 3, characterized in that, The genetic information feature construction module includes the following process: First, the image phenotypes obtained from the image phenotype extraction module were subjected to genetic analysis using the GWAS method. Specifically, under an additive genetic model, linear regression analysis was used to assess the association between single nucleotide variants and phenotypes, and covariate corrections were performed for age, sex, body mass index, body surface area, waist-to-hip ratio, and principal components of population structure to screen for genome-wide significant genetic variant sites. Then, the linkage disequilibrium score regression (LDSC) method was used to estimate the heritability of the image phenotypes and their pairwise genetic correlations to assess the shared genetic basis of different phenotypes. Subsequently, functional annotation and gene mapping analysis were performed based on significant genetic variant sites, and candidate variant sites were mapped to potential functional genes through position mapping and gene set analysis methods. Furthermore, based on the eQTL prediction model, the gene expression level mediated by genetic variation was inferred, and the association between gene expression or protein expression and imaging phenotype was assessed by combining TWAS and PWAS. Genetic loci or related genes that reached genome-wide significance and showed significance in at least two omics association analyses were identified as pleiotropic candidate genes associated with VHD. In addition, pathway enrichment analysis and tissue-specific expression analysis were performed on the screened pleiotropic candidate genes to elucidate their potential biological functions, and possible causal variation sites were preferentially screened through fine mapping analysis. Finally, key genetic loci related to the structure, function and occurrence of heart valves and VHD were identified, resulting in a set of candidate risk loci. Furthermore, considering that CMR imaging phenotypes reflect the mechanical properties of heart valves and hemodynamics (such as shear stress and pressure load) to a certain extent, a subset of genetic loci related to mechanical stress was screened from the above loci to enhance the biological consistency between genetic characteristics and imaging phenotypes and improve the explanatory power of the model. Genotypic information of individuals at the aforementioned key risk sites was extracted to construct a set of genetic characteristics.

5. The multimodal VHD risk prediction system according to claim 4, characterized in that, In the genetic information feature construction module, a weighted genetic risk score (GRS) is constructed based on the effect size of each point, and its calculation method is as follows: , Wherein, βᵢ represents the effect size of the i-th genetic locus, and Gᵢ represents the genotype code of the corresponding locus: 0, 1, or 2, representing the copy number of the minor allele, respectively; by constructing GRS, the feature dimension is reduced while preserving genetic information, thereby enhancing the risk representation ability of genetic features.

6. The multimodal VHD risk prediction system according to claim 4, characterized in that, The VHD risk prediction output module includes model training, model interpretation, VHD disease prediction and risk stratification, wherein: Model training includes dividing the dataset into training and test sets in a 7:3 ratio, optimizing model parameters using a 10-fold cross-validation strategy, and selecting the model using the area under the receiver operating characteristic curve (AUC) as the main evaluation metric. The optimal model is selected to output the risk probability of VHD in individuals who already have it and individuals who will develop VHD in the future. Model interpretation employs a Shapley value-based (SHAP) method to interpret the model output, quantifying the contribution of each input feature to the prediction result; VHD Disease Prediction and Risk Stratification: Disease prediction involves inputting the feature data of individuals in the test set into the optimal prediction model to obtain the corresponding VHD risk prediction probability; the model output is... , representing the probability of an individual experiencing VHD given feature X; where X represents the input multimodal feature set; based on the model output probability, the optimal classification threshold is determined by iterating through different classification thresholds and calculating the corresponding performance metrics, including F1 score, precision, accuracy, and recall; the optimal classification threshold T* is determined by maximizing the F1 score, i.e.: , In the formula, T represents the classification threshold, and F1(T) represents the F1 score calculated at the threshold T. By iterating through different thresholds and calculating the corresponding F1 scores, the threshold that makes the F1 score reach the maximum value is selected as the optimal classification threshold T*, thereby achieving a balance between precision and recall. Based on the optimal threshold, individuals are divided into high-risk groups and low-risk groups. Risk stratification involves constructing a Cox proportional hazards model based on follow-up data, performing survival analysis on different risk groups, calculating their risk ratio for future VHD events, thereby validating the model's effectiveness in predicting diseases using longitudinal data, and achieving dynamic assessment of individual risk.

7. The multimodal VHD risk prediction system according to claim 1, characterized in that, It also includes a processor and a memory; the processor is used to perform CMR image analysis, genetic modeling and risk prediction calculations; the memory is used to store computer-executable instructions to enable the processor to perform the above method steps.