High-risk MASH intelligent identification model construction method and system based on multi-parameter MRI
The binary logistic regression model EFT1 constructed through multi-parameter MRI solves the problems of high trauma of liver biopsy and insufficient sensitivity of non-invasive models in traditional methods, and achieves efficient and accurate identification of intermediate and high-risk MASH with excellent specificity and sensitivity, reducing the risk of missed diagnosis and reducing costs.
Patent Information
- Application Number
- CN202510505808.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-09-12
AI Technical Summary
In existing technologies, traditional liver biopsy is highly invasive and cannot be monitored repeatedly. Non-invasive models such as FIB-4 and APRI are not sensitive enough in early fibrosis. Single-modality MRI cannot meet the multi-dimensional diagnostic needs of MASH. Existing imaging genomics methods ignore the synergistic effects between parameters, resulting in decreased diagnostic specificity.
Multi-parameter MRI was used to construct a binary logistic regression model (EFT1) with liver stiffness, fat quantification, and T1 relaxation time as parameters through MRE, PDFF, and T1 mapping image data. Combined with feature extraction and 5-fold cross-validation, the classification boundary was optimized to achieve intelligent identification of high-risk MASH.
It achieves unique recognition capabilities for medium- and high-risk MASH, reduces the risk of missed diagnosis, achieves 100% specificity, achieves an excellent balance between sensitivity and specificity, has high model robustness, reduces costs by approximately 60%, and takes ≤5 minutes for a single examination.
Smart Images

Figure CN120636830A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical models, and in particular to a method and system for constructing a high-risk MASH intelligent recognition model based on multi-parameter MRI. Background Art
[0002] Metabolic dysfunction-associated steatohepatitis (MASH) is a severe subtype of metabolic dysfunction-associated fatty liver disease (MASLD). High-risk patients with stage 2 or higher fibrosis are at increased risk for progression to cirrhosis and liver-related death. Identifying high-risk MASH is crucial for preventing disease progression and is a key point in clinical trials and drug treatment.
[0003] Methods for identifying MASH can be categorized as invasive or non-invasive. Liver biopsy is considered the gold standard, but limitations and risks make it difficult to use as a screening procedure. Researchers have begun developing non-invasive methods. Non-invasive serum scores such as the FIB-4 (Fibrosis index based on the 4 factors), the APRI (Aspartate aminotransferase-to-Platelet Ratio Index), the GPR (Gamma-glutamyl transpeptidase to Platelet Ratio), and the ELF (Enhanced liver fibrosis) score are widely recommended due to their availability and affordability. The FIB-4 effectively excludes advanced fibrosis, but results are inconclusive in 30% of cases, requiring additional testing. Its low specificity in elderly patients may lead to diagnostic errors. Due to variations in platelet counts and liver enzyme levels, the APRI, GPR, and ELF models face challenges in accuracy across different populations. Vibration-controlled transient elastography (VCTE) and magnetic resonance elastography (MRE) are the most extensively studied imaging techniques for the assessment of high-risk MASH. VCTE is recommended when the FIB-4 score is equivocal, but its reliability depends on operator and patient characteristics.
[0004] Therefore, traditional liver biopsy is highly invasive, has large sampling errors, and cannot be monitored repeatedly. Existing non-invasive models (such as FIB-4 and APRI) rely on serological indicators and are insufficiently sensitive to early fibrosis (≥F2). Single-modality MRI (magnetic resonance imaging, such as MRE or PDFF) only reflects a single pathological feature and cannot meet the multidimensional diagnostic needs of MASH. There is a lack of multi-parameter quantitative models that integrate liver fibrosis, steatosis, and inflammatory activity. Existing imaging omics methods ignore the synergistic effects between parameters, resulting in reduced diagnostic specificity. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for constructing a high-risk MASH intelligent identification model based on multi-parameter MRI to solve the problems raised in the above background technology.
[0006] To achieve the above-mentioned object of the invention, one aspect of the present invention provides a method for constructing a high-risk MASH intelligent identification model based on multi-parameter MRI, comprising the following steps:
[0007] Step S1, acquiring multimodal imaging data of the patient, wherein the multimodal imaging data includes an MRE (liver magnetic resonance elastography) image, a PDFF (liver proton density fat fraction) image, and a T1 mapping image;
[0008] Step S2, performing feature extraction on the collected multimodal image data;
[0009] Step S3: construct a binary logistic regression model EFT1 with liver stiffness, fat quantification, and T1 relaxation time as parameters, and perform training and testing. The formula of the binary logistic regression model EFT1 (custom name, E for elastography (MRE), F for fat quantification (PDFF), and T1 for T1 mapping) is:
[0010] EFT1=39.421logMRE+7.858logPDFF+28.883logT1-111.920;
[0011] Among them, MRE is the liver stiffness value obtained from the MRE image, PDFF is the fat quantitative value obtained from the PDFF image, and T1 is the T1 relaxation time value obtained from the T1 mapping image.
[0012] Step S4: perform risk assessment and visualization on the output results of the logistic regression model.
[0013] Furthermore, in step S1, an EPI (echo planar imaging) sequence is used to obtain a four-layer liver elastogram from the MRE image; a six-echo Dixon technique is used to obtain a PDFF image, from which fat fraction is synchronously quantified; and the MOLLI (modified look-lock inversion recovery) technique is used to measure the T1 relaxation time from the T1 mapping image.
[0014] Furthermore, the feature extraction method in step S2 is to outline the ROI (Region of Interest) on the MRE elastogram, exclude the interference area of blood vessels and artifacts, and generate a quantitative value of liver stiffness (ie, elasticity value).
[0015] Furthermore, in step S2, several ROIs are selected on the PDFF map to calculate the mean value of liver fat content.
[0016] Furthermore, step S3 includes the following steps:
[0017] Step S301: pre-process the collected data samples and divide the samples into a training set and a test set;
[0018] Step S302, feature screening is performed through single factor logistic regression and variance inflation factor;
[0019] Step S303: On the training set, all combinations of MRE, PDFF, and T1 are analyzed through full-subset regression to construct a binary logistic regression model. The best model is selected using the AUC (area under the receiver operating characteristic curve), AIC (Akaike information criterion), and BIC (Bayesian information criterion) criteria.
[0020] Step S304: Determine the classification threshold of the model through 5-fold cross-validation and evaluate it on the test set. The evaluation content includes AUC, sensitivity, specificity, approximation index, accuracy, precision, and F1 score.
[0021] Step S305: Determine the best model EFT1 based on the performance of the model on the model test set.
[0022] Furthermore, in step S4, a heat map of risk levels and fibrosis stages is output, wherein the risk levels include low, medium and high risk.
[0023] Another aspect of the present invention provides a high-risk MASH intelligent identification model construction system based on multi-parameter MRI, including an acquisition module, a feature extraction module, a model construction module, and a model output module, wherein:
[0024] The acquisition module is used to acquire multimodal imaging data of the patient, wherein the multimodal imaging data includes MRE images, PDFF images, and T1 mapping images;
[0025] The feature extraction module is used to extract features from the collected multimodal image data;
[0026] The model building module is used to construct a binary logistic regression model EFT1 with liver stiffness, fat quantification and T1 relaxation time as parameters and perform training and testing. The formula of the binary logistic regression model EFT1 is:
[0027] EFT1=39.421logMRE+7.858logPDFF+28.883logT1-111.920;
[0028] Among them, MRE is the liver stiffness value obtained from the MRE image, PDFF is the fat quantitative value obtained from the PDFF image, and T1 is the T1 relaxation time value obtained from the T1 mapping image.
[0029] The model output module is used to perform risk assessment and visualization on the output results of the logistic regression model.
[0030] Compared with the prior art, the present system and method have the following advantages:
[0031] 1. This invention integrates three-dimensional biomarkers of liver stiffness (kPa), fat fraction (%), and inflammatory T1 value (ms), and for the first time combines MRE (fibrosis quantification), PDFF (fatty degeneration), and T1 mapping (inflammatory activity) to construct a multidimensional diagnostic model. This model breaks through the limitations of traditional single-parameter diagnosis, has unique identification capabilities for intermediate- and high-risk MASH, and reduces the risk of missed diagnosis.
[0032] 2. This paper dynamically optimizes the classification boundary (critical value -0.431) through logistic regression to achieve the optimal balance between sensitivity and specificity, and eliminates feature redundancy through VIF control and full subset regression to ensure model robustness.
[0033] 3. This invention constructs a full-process system architecture, covering standardized image acquisition protocols, automated feature extraction modules and visual risk assessment interfaces, forming a complete technical closed loop with clinical translation advantages.
[0034] 4. This method matches the gold standard NAS (Nonalcoholic steatohepatitis activity score) ≥4 and fibrosis ≥F2 with a strict pathological definition, with a specificity of 100%. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 Flowchart of the method for constructing a high-risk MASH intelligent identification model based on multi-parameter MRI.
[0036] Figure 2 Schematic diagram of the principle of a method for constructing a high-risk MASH intelligent identification model based on multi-parameter MRI.
[0037] Figure 3 ROC curves are shown to compare the diagnostic performance of EFT1, FIB-4, APRI, GPR, and MAST in the training and test cohorts. DETAILED DESCRIPTION
[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0039] like Figure 1 , Figure 2 The figures are respectively a flow chart and a principle diagram of the method of the present invention. The embodiment of the present invention provides a method, and the specific steps are as follows:
[0040] Step S1: Acquire multimodal imaging data of a patient, wherein the multimodal imaging data includes MRE images, PDFF images, and T1mapping images.
[0041] A standardized scanning protocol was performed using a 3.0T MRI system (United Imaging uMR 790). 2D MRI scans were performed using a 60 Hz controller and driver (Resoundant, USA). MRE, PDFF, and T1 mapping scans were performed with the patient in the supine position. MRE was acquired in four axial planes using an echo-planar imaging (EPI) sequence with the following parameters: TR 1000.2 milliseconds; TE 44.6 milliseconds; field of view (FOV) 420 × 420 mm; matrix 96 × 96; slice thickness 8 mm; acceleration factor 2; and acquisition time 10 seconds. The MRI scanner automatically processed wave images into elastograms. A 3D chemical shift-encoded fat analysis and computational technology (FACT) sequence was used to acquire PDFF images. Imaging parameters were as follows: TR 12.06 milliseconds; TE 1.71, 3.22, 4.73, 6.24, 7.75, and 9.26 milliseconds; flip angle 3°; FOV 420 × 420 mm; matrix 240 × 240; slice thickness 10 mm; bandwidth 900 Hz; perceived acceleration 2; and acquisition time 15 seconds. T1 mapping was obtained using the Look-Locker inversion recovery (MOLLI) technique, and T1 relaxometry was performed using four single-breath-hold axial slices. Sequence parameters were as follows: TR 2.79 milliseconds; TE 1.29 milliseconds; inversion times 145 and 225 milliseconds; flip angle 35°; FOV 420 × 420 mm; matrix 256 × 256; slice thickness 8 mm; and acquisition time 19 seconds.
[0042] Step S2: extracting features from the collected multimodal image data.
[0043] Raw MRI images were transferred to a dedicated workstation at Shanghai United Imaging Healthcare (uWS, Shanghai United Imaging Healthcare). For liver stiffness measurement, regions of interest (ROIs) were manually delineated across the largest measurable portion of the liver in four elasticity images. Global mean liver stiffness was calculated as the average of the four ROIs. Water images, fat images, PDFF images, relaxation images, in-phase images, and out-of-phase images obtained from the six-echo sequence were arranged separately. In the FACT sequence, the PDFF image was placed within a circular region of interest to measure relevant quantitative indices. A total of twelve ROIs were selected (six in the left lobe and six in the right lobe), and the mean PDFF value was calculated. T1 mapping measurements were performed identically to those for MRE.
[0044] Step S3: Construct and train a binary logistic regression model (EFT1) that uses liver stiffness, fat quantification, and T1 relaxation time as parameters to reflect liver fibrosis, steatosis, and inflammation to accurately identify high-risk MASH. Establishing the EFT1 model includes the following steps:
[0045] In step S301, the collected data samples are preprocessed and the data are randomly stratified and sampled into a training group and a test group using a ratio of 7:3.
[0046] Step S302: Perform a preliminary screening using univariate logistic regression (p < 0.01), and eliminate multicollinearity issues with the screened variables using a variance inflation factor (VIF) > 10. Optimize feature combinations for these variables using full subset regression.
[0047] Step S303: On the training set, full subset regression analysis of MRE, PDFF, and T1 is performed, all combinations are traversed, and a binary logistic regression model optimization model is constructed.
[0048] As shown in Table 1 , the model was selected based on the criteria of large AUC value and small AIC and BIC values, and the logMRE+logPDFF+logT1 model was selected.
[0049] Table 1:
[0050]
[0051] In step S304, the classification cutoffs for the two models are determined through 5-fold cross-validation. Key parameters, including AUC, sensitivity, specificity, approximation index, accuracy, precision, and F1 score, are calculated. In an independent test group, the classification cutoffs are used to further validate model performance and select the optimal model.
[0052] Step S305: Based on the model's performance on the model test set, the optimal model, EFT1, is determined. As shown in Table 2, in the test set, the model, logMRE+logPDFF+logT1, not only avoids overfitting but also performs well in terms of sensitivity, accuracy, and F1 score on the test set. This model is designated as the optimal model, EFT1.
[0053] EFT1=39.421 logMRE+7.858 logPDFF+28.883 logT1-111.920
[0054] Table 2
[0055]
[0056] The AUC, sensitivity, and specificity of the best model were compared and evaluated with traditional scoring systems (FIB-4, APRI, GPR) and MAST (MRI-aspartate aminotransferase score) to comprehensively measure the diagnostic effect. The classification critical value of -0.431 was determined by 5-fold cross-validation. A multi-dimensional validation system was developed, including: Internal validation: The AUC of the training set reached 0.995 (95%CI: 0.985-1.000). Independent testing: The accuracy of the test set was 97.06%, and the F1 score was 0.978. Clinical comparison: significantly better than traditional models such as FIB-4 (ΔAUC +0.32) and MAST (ΔAUC +0.28). As shown in Tables 2 and 3, Figure 3 As shown, the EFT1 model demonstrated excellent diagnostic performance in both the training and test cohorts, with AUC values of 0.995 (95% CI 0.985-1.000) and 0.995 (95% CI 0.979-1.000), respectively. Five-fold cross-validation results showed an average accuracy of 98.8%, an average precision of 98.5%, and an average F1 score of 0.992 in the training cohort. At the optimal cutoff value of -0.431, the sensitivity and specificity were 98.2% and 96.3%, respectively. In an independent test cohort, the model continued to perform well, with an accuracy of 97.1%, a precision of 100%, and an F1 score of 0.978. Using the same cutoff value as in the test cohort, the sensitivity and specificity reached 95.7% and 100%, respectively. The MAST score performed second best in the training set, with an AUC of 0.911, a sensitivity of 91.1%, and a specificity of 88.9%. The AUC for the test cohort was 0.879, with a sensitivity of 90.9% and a specificity of 77.8%. Stability was good, but sensitivity and specificity were reduced and unbalanced. The overall performance of traditional scoring systems (FIB-4, APRI, GPR) was limited. FIB-4 had high specificity (92.6% in the training cohort and 100% in the test cohort), but very low sensitivity (40.9% in the test cohort), with a significant risk of missed diagnosis. In the test cohort, APRI's sensitivity increased to 86.4%, but its specificity decreased to 55.6%, with an increased misclassification rate. GPR had the lowest AUC (0.617 in the training cohort and 0.626 in the test cohort) and insufficient specificity (44.4% in the training cohort), limiting its clinical application value.
[0057] Table 3:
[0058]
[0059] Step S4: Risk assessment and visualization of the logistic regression model output results. The output results are divided into three risk levels: low, medium, and high risk, and a fibrosis staging heat map is provided.
[0060] Furthermore, this invention utilizes a cloud-edge computing architecture to automate the entire process of image analysis, feature extraction, and risk assessment. This technology demonstrates excellent clinical value, enabling accurate identification of high-risk individuals for MASH (NAS ≥ 4 and fibrosis ≥ F2). Furthermore, a single examination takes ≤ 5 minutes, reducing costs by approximately 60%.
[0061] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a high-risk MASH intelligent identification model based on multi-parameter MRI, characterized in that: The following steps are involved: Step S1, collecting multimodal imaging data of the patient, wherein the multimodal imaging data includes MRE images, PDFF images, and T1mapping images; Step S2, performing feature extraction on the collected multimodal image data; Step S3, constructing a binary logistic regression model EFT1 with liver stiffness, fat quantification and T1 relaxation time as parameters and performing training and testing. The formula of the binary logistic regression model EFT1 is: EFT1=39.421logMRE+7.858logPDFF+28.883logT1-111.920; Among them, MRE is the liver stiffness value obtained from the MRE image, PDFF is the fat quantitative value obtained from the PDFF image, and T1 is the T1 relaxation time value obtained from the T1mapping image. Step S4: perform risk assessment and visualization on the output results of the logistic regression model.
2. The method for constructing a high-risk MASH intelligent identification model based on multi-parameter MRI according to claim 1, characterized in that: In step S1, an EPI sequence is used to obtain a four-layer liver elastogram from the MRE image; a six-echo Dixon technique is used to obtain a PDFF image, and fat fraction is synchronously quantified from the PDFF image; and the MOLLI technique is used to measure the T1 relaxation time from the T1 mapping image.
3. The method for constructing a high-risk MASH intelligent identification model based on multi-parameter MRI according to claim 1, characterized in that: The feature extraction method in step S2 is to outline the ROI on the MRE elastogram, exclude the interference areas of blood vessels and artifacts, and generate a quantitative liver stiffness map.
4. The method for constructing a high-risk MASH intelligent identification model based on multi-parameter MRI according to claim 1, characterized in that: In step S2, several ROIs are selected on the PDFF map and the mean value of liver fat content is calculated.
5. The method for constructing a high-risk MASH intelligent identification model based on multi-parameter MRI according to claim 1, characterized in that: Step S3 includes the following steps: Step S301: pre-process the collected data samples and divide the samples into a training set and a test set; Step S302, feature screening is performed through single factor logistic regression and variance inflation factor; Step S303: On the training set, all subset regression analyses are performed on all combinations of MRE, PDFF, and T1 to construct a binary logistic regression model, and the best model is selected using the AUC, AIC, and BIC criteria. Step S304: Determine the classification threshold of the model through 5-fold cross-validation and evaluate it on the test set. The evaluation content includes AUC, sensitivity, specificity, approximation index, accuracy, precision, and F1 score. Step S305: Determine the best model EFT1 based on the performance of the model on the model test set.
6. The method for constructing a high-risk MASH intelligent identification model based on multi-parameter MRI according to claim 1, characterized in that: In step S4, a heat map of risk levels and fibrosis stages is output, wherein the risk levels include low, medium and high risk.
7. A high-risk MASH intelligent identification model construction system based on multi-parameter MRI, characterized by: It includes acquisition module, feature extraction module, model building module and model output module, among which: The acquisition module is used to acquire multimodal imaging data of the patient, wherein the multimodal imaging data includes MRE images, PDFF images, and T1 mapping images; The feature extraction module is used to extract features from the collected multimodal image data; The model building module is used to construct a binary logistic regression model EFT1 with liver stiffness, fat quantification and T1 relaxation time as parameters and perform training and testing. The formula of the binary logistic regression model EFT1 is: EFT1=39.421logMRE+7.858logPDFF+28.883logT1-111.920; Among them, MRE is the liver stiffness value obtained from the MRE image, PDFF is the fat quantitative value obtained from the PDFF image, and T1 is the T1 relaxation time value obtained from the T1mapping image. The model output module is used to perform risk assessment and visualization on the output results of the logistic regression model.