Liver fat quantitative evaluation method and system based on plasma proteomics
By combining plasma proteomics and machine learning algorithms, specific protein markers were screened out and liver fat content prediction models were established, which solved the invasive and high-cost problems of liver fat content evaluation in the prior art, and achieved high-precision non-invasive assessment and early screening.
Patent Information
- Application Number
- CN202510607716.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems such as strong invasiveness, high cost, long detection time and limited equipment when evaluating liver fat content, and cannot achieve high-precision, non-invasive dynamic monitoring and early screening.
By combining plasma proteomics and machine learning algorithms, specific protein markers were screened out, liver fat content prediction model was established, and proton density fat fractions were combined with magnetic resonance imaging to achieve high-precision non-invasive quantitative evaluation.
It realizes high-precision and non-invasive liver fat content assessment, breaks through the limitations of traditional imaging examinations, provides time window information for early screening and efficacy evaluation, reduces detection costs and improves the convenience of detection.
Smart Images

Figure CN120472989A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical detection technology, and in particular to a method and system for quantitatively evaluating liver fat based on plasma proteomics. Background Art
[0002] Liver fat content is an important prognostic indicator for assessing metabolic health and is closely associated with cardiovascular disease and metabolic dysfunction-related fatty liver disease. In particular, within the diagnosis and treatment of metabolic dysfunction-related fatty liver disease, quantitative assessment of liver fat has been elevated to a core position of importance, on par with traditional vital signs. Therefore, accurate assessment of liver fat content is crucial for early diagnosis, efficacy monitoring, and prognostic assessment.
[0003] Currently, clinical assessment of liver fat content relies primarily on two techniques: 1) Liver biopsy, the gold standard, suffers from inherent limitations such as large sampling error (reflecting only approximately 1 / 50,000 of liver volume), high invasiveness (complication rate approximately 0.5-3%), and high procedural risk. 2) MRI-PDFF (magnetic resonance imaging proton density fat fraction), while capable of accurately quantifying whole-liver fat (with >95% accuracy), is limited by high equipment cost (approximately 3,000-5,000 yuan per test), lengthy testing times (≥30 minutes per case), and contraindications for metal implants. Therefore, developing a novel peripheral blood-based molecular marker assay and establishing an alternative method for quantitative liver fat assessment holds significant clinical value in overcoming the limitations of existing technologies.
[0004] Plasma proteomics has become a revolutionary tool for disease biomarker research due to its high sensitivity (capable of detecting pg / mL levels of protein), high throughput (>1000 proteins per assay), and dynamic monitoring capabilities. Notably, a specific proteomic profile for liver fat content has yet to be established globally. This study, integrating high-throughput plasma protein analysis with machine learning algorithms, systematically screened for characteristic protein clusters associated with liver fat deposition. Its innovations include: 1) overcoming the invasive and costly limitations of existing technologies, establishing a minimally invasive (requiring only 200 μL of plasma) and reproducible dynamic monitoring protocol; and 2) constructing the first multi-protein scoring model for liver fat content, providing a molecular basis for early screening of fatty liver disease (providing 3-5 years of early warning) and therapeutic efficacy assessment. This research will promote a paradigm shift in liver fat assessment from imaging-based to molecular diagnostics. Summary of the Invention
[0005] In response to the deficiencies in the prior art, the present invention provides a method and system for quantitatively evaluating liver fat based on plasma proteomics.
[0006] In the first aspect, the present invention provides a method for quantitative assessment of liver fat based on plasma proteomics, comprising the following steps: obtaining an individual's plasma sample; obtaining plasma protein levels based on the plasma sample; obtaining the influence weight of liver fat content based on the plasma protein level and in combination with the proton density fat fraction of magnetic resonance imaging; establishing a liver fat content prediction model using the influence weight; and achieving quantitative assessment of liver fat content through the liver fat content prediction model. The present invention achieves high-precision non-invasive quantitative assessment of liver fat content by integrating high-dimensional plasma proteomics data with machine learning algorithms, breaking through the limitations of expensive and complex operation of traditional imaging examination equipment; by screening a combination of protein biomarkers specifically related to liver fatty degeneration, a highly specific prediction model is established, which is significantly better than the existing scoring model based on clinical indicators; by dynamically monitoring the relationship between plasma proteome changes and liver fat content, a high-precision prediction of liver fat content is achieved, providing time window information for early intervention that is difficult to obtain with traditional technologies.
[0007] Optionally, obtaining the plasma protein level according to the plasma sample includes: obtaining the test results of the individual plasma protein level using a plasma protein biomarker analysis platform according to the plasma sample; and standardizing the test results to obtain standardized plasma protein levels. The present invention realizes high-throughput and high-precision detection of individual plasma protein levels by combining the plasma protein biomarker analysis platform with standardized processing technology, breaking through the limitations of traditional protein detection methods in throughput and accuracy, and providing an important tool for large-scale population proteomics research; by applying standardized processing technology, batch effects and experimental errors between samples are eliminated, and the comparability and repeatability of test results in different laboratories and at different time points are significantly improved; by establishing a dynamic monitoring system for individualized plasma protein levels, a paradigm shift from single detection to continuous health monitoring is achieved, breaking through the temporal and spatial limitations of early disease warning and efficacy evaluation.
[0008] Optionally, the method of obtaining the influence weight of liver fat content based on the plasma protein level in combination with the proton density fat fraction of magnetic resonance imaging includes: allocating a training set and a validation set based on the standardized plasma protein level; obtaining the influence weight of different plasma protein levels on liver fat content based on the training set in combination with the proton density fat fraction of magnetic resonance imaging. The present invention breaks through the limitations of traditional statistical methods in variable screening by combining machine learning with plasma proteomics, and realizes the accurate identification of key biomarkers of liver fat content from massive proteins; by establishing a relationship model between protein level and liver fat content, the sensitivity and specificity of early risk assessment of fatty liver are significantly improved; by using a dynamic weight optimization algorithm driven by a validation set, the limitation of traditional biomarker research that relies only on a single data set is broken through, and the model has the ability to continuously self-optimize, providing a breakthrough solution for the non-invasive and accurate prediction of liver fat content.
[0009] Optionally, obtaining the weights of the effects of different plasma protein levels on liver fat content based on the training set in combination with the proton density fat fraction from magnetic resonance imaging includes: establishing a lasso regression model using a lasso regression algorithm based on the training set; training the lasso regression model using the age, gender, and levels of different types of proteins of the individuals in the training set to obtain a trained lasso regression model; and obtaining the weights of the effects of different plasma protein levels on liver fat content based on the trained lasso regression model in combination with the proton density fat fraction from magnetic resonance imaging. The present invention achieves accurate prediction of liver fat content based on a dynamic protein regulatory network by integrating the lasso regression algorithm with multidimensional clinical data and proteomics modeling. Lasso regression's adaptive variable screening feature innovatively automatically identifies key biomarker combinations from high-dimensional plasma protein data, overcoming the subjectivity and inefficiency of traditional manual marker screening and ensuring that the model is both biologically interpretable and clinically practical. A dynamic weight model with strong generalization capabilities is constructed through a collaborative optimization mechanism of the training and validation sets, overcoming the performance degradation problem of existing prediction methods when applied across populations.
[0010] Optionally, the structural formula of the lasso regression model is as follows: in, represents the proton density fat fraction on magnetic resonance imaging, Indicates the The expression levels of the proteins is the intercept term, is the regression coefficient of age, is the regression coefficient of gender, A numeric variable representing age, A categorical variable representing gender, For the The regression coefficient of protein is the error term, 2911 represents the number of protein types; the objective function of the lasso regression model is as follows: in, Indicates the The target variable for each sample, is the intercept term, is the regression coefficient of age, is the regression coefficient of gender, Indicates the The age value of the samples, Indicates the The gender classification variable of the samples, For the In the sample The concentration of protein, is the total number of samples, is the regularization parameter, 2911 represents the number of protein types, is the regression coefficient, For the The regression coefficient of a protein. This invention introduces regularization constraints into high-dimensional proteomics modeling, realizes the simultaneous optimization of automatic feature selection and weight distribution of multiple plasma proteins, breaks through the bottleneck of overfitting of traditional statistical methods in high-dimensional small sample data, and constructs for the first time a quantitative prediction model of liver fat that is both sparse and explanatory. By jointly modeling demographic variables of age and gender with protein expression levels under a unified mathematical framework, it overcomes the limitations of traditional clinical research that relies on complex, subjective or time-consuming clinical indicators, and significantly improves the stability and clinical applicability of the model across populations. By dynamically adjusting the balance between protein weight sparsity and model accuracy through regularization parameters, it breaks through the technical barrier of existing non-invasive diagnostic technologies that cannot quantify specific contributions, and provides a breakthrough computational tool for accurate quantification and personalized intervention of fatty liver.
[0011] Optionally, the weights of the effects of different plasma protein levels on liver fat content are obtained based on the trained lasso regression model and combined with the proton density fat fraction of magnetic resonance imaging, including: based on the trained lasso regression model, the proton density fat fraction of magnetic resonance imaging is normalized to obtain a normalization result; based on the normalization result, the hyperparameters of the trained lasso regression model are optimized using 10-fold cross validation to obtain an optimization result; and through the optimization result, the weights of the effects of different plasma protein levels on liver fat content are obtained. The present invention solves the model bias problem caused by the heterogeneity of multi-center medical imaging data by combining the normalization of magnetic resonance proton density fat fraction with the optimization of lasso regression hyperparameters, and realizes the standardization of the liver fat content prediction model; through the dynamic parameter optimization mechanism of 10-fold cross validation, it breaks through the limitations of traditional single data division, so that the model can significantly reduce the prediction error while maintaining the stability of protein feature selection.
[0012] Optionally, the influence weight is used to establish a liver fat content prediction model based on plasma protein. The liver fat content prediction model satisfies the following expression: in, represents the liver fat content calculated based on plasma protein levels, Indicates the The weight of the effect of proteins on liver fat content, 62 represents the number of proteins significantly correlated with liver fat content screened by lasso regression, Indicates the The expression levels of the 10 proteins screened out included IGFBP2, FABP4, MET, CPM, CES1, IGFBP1, CDHR2, RBP5, ERBB2, and SSC5D. To assess clinical applicability, a liver fat content prediction model (simplified model) based on the top 10 proteins was calculated. The liver fat content prediction model for the top 10 proteins satisfies the following expression: in, represents the liver fat content calculated based on plasma protein levels, Indicates the The weight of the effect of each protein on liver fat content, 10 represents the number of the top ten proteins significantly correlated with liver fat content screened by lasso regression, Indicates the The present invention converts the weighted combination of 10 or 62 key plasma proteins into quantifiable liver fat content based on plasma protein levels, thus realizing the transformation from complex proteomic data to clinically applicable indicators, breaking through the bottleneck of traditional biomarker research that is difficult to implement, and establishing for the first time a non-invasive liver fat assessment system that can be directly used in clinical practice; by accurately quantifying the contribution of each protein to liver fat content, it overcomes the limitation of existing fat assessment methods that only provide qualitative or semi-quantitative results; by constructing a prediction model based on 10 or 62 proteins, it breaks through the practical barriers of high-dimensional omics data in clinical translation, and provides a breakthrough blood test solution for early screening, dynamic monitoring and efficacy evaluation of fatty liver. Its convenience and accessibility are expected to change the current clinical practice model of fatty liver diagnosis and treatment.
[0013] Optionally, the quantitative assessment of liver fat content using the liver fat content prediction model includes: first, establishing a correlation analysis between the prediction model's calculation results and the MRI proton density fat fraction; and second, using a receiver operating characteristic (ROC) curve to evaluate the model's ability to discriminate fatty liver disease (MRI proton density fat fraction >5%). This technological breakthrough is reflected in three aspects: 1) It enables the objective grading of liver fat content by blood tests for the first time, overcoming the limitation of traditional liver function tests that cannot distinguish fat content levels; 2) By objectively evaluating the model's discriminative capabilities, it overturns the current situation of imaging examinations relying on subjective judgment, providing a repeatable quantitative standard for liver fat assessment and significantly improving the accuracy of early screening for fatty liver disease; 3) Through clinical validation, it converts laboratory markers into diagnostic evidence, providing a reliable objective standard for early screening for fatty liver disease.
[0014] Optionally, the liver fat content prediction model's ability to discriminate fatty liver disease is tested using a magnetic resonance proton density fat fraction >5% (5%-32% for mild, 33%-65% for moderate, and >66% for severe fatty liver) as the gold standard, using the area under the receiver operating characteristic curve (AUC). The results include: In both the training and validation sets, the model's AUC significantly outperformed that of traditional clinical models, demonstrating excellent discriminatory capabilities. This model, for the first time, accurately identifies mild fatty liver disease (PDFF 5%-32%) using blood testing, breaking through the technical bottleneck of existing blood testing technologies that cannot identify early-stage steatosis. Through rigorous training and validation set validation, this model redefines the accuracy standard for non-invasive testing in the diagnosis of fatty liver disease, enabling blood testing to achieve imaging-level discriminatory performance for the first time, providing a revolutionary technical approach for early screening of fatty liver disease.
[0015] In a second aspect, the present invention provides a plasma proteomics-based liver fat quantitative assessment system, comprising an input device, a processor, an output device, and a memory, wherein the input device, the processor, the output device, and the memory are interconnected, wherein the memory is used to store a computer program, the computer program including program instructions, the processor is configured to call the program instructions, and the system uses the plasma proteomics-based liver fat quantitative assessment method. The system provided by the present invention has high integration and smooth information transmission between various components. By integrating high-throughput proteomic detection, machine learning modeling, and imaging gold standard verification, it breaks through the invasiveness of traditional liver biopsy and the high cost limitations of imaging examinations, and achieves a non-invasive diagnostic technology breakthrough that can accurately quantify liver fat content with only a single blood draw. Through in-depth screening of 2911 plasma proteins and dynamic weight modeling of 10 key markers, a clinically interpretable multidimensional protein risk assessment algorithm was developed, whose identification efficiency surpasses existing blood test methods. By organically combining standardized detection processes with intelligent analysis platforms, it breaks through the technical barriers to fatty liver screening caused by uneven distribution of medical resources, significantly reducing detection costs and improving screening efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flow chart of a method for quantitatively assessing liver fat based on plasma proteomics according to an embodiment of the present invention; Figure 2 Schematic diagram of the relationship between liver fat content calculated based on 62 plasma protein levels and proton density fat fraction on magnetic resonance imaging in the training set and validation set of an embodiment of the present invention; Figure 3 Schematic diagram of the relationship between liver fat content calculated based on 10 plasma protein levels and proton density fat fraction on magnetic resonance imaging in the training set and validation set of an embodiment of the present invention; Figure 4 Schematic diagram of the structure of a liver fat quantitative assessment system based on plasma proteomics according to an embodiment of the present invention. DETAILED DESCRIPTION
[0017] Specific embodiments of the present invention will be described in detail below. It should be noted that the embodiments described herein are for illustrative purposes only and are not intended to limit the present invention. In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. However, it will be apparent to one of ordinary skill in the art that these specific details are not necessarily required to practice the present invention. In other instances, well-known circuits, software, or methods are not specifically described to avoid obscuring the present invention.
[0018] Throughout this specification, references to "one embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the present invention. Therefore, appearances of the phrases "in one embodiment," "in an embodiment," "an example," or "an example" in various places throughout this specification are not necessarily all referring to the same embodiment or example. Furthermore, the particular features, structures, or characteristics may be combined in any suitable combinations and / or subcombinations in one or more embodiments or examples. Furthermore, those of ordinary skill in the art will appreciate that the figures provided herein are for illustrative purposes only and are not necessarily drawn to scale.
[0019] See Figure 1 The embodiment of the present invention provides a method for quantitatively evaluating liver fat based on plasma proteomics, the method comprising the following steps: S1. Obtain a plasma sample from an individual.
[0020] In one example, plasma proteomics was measured in 53,017 participants using the antibody-based Olink Explore 3072 proximity extension assay (PEA).
[0021] Furthermore, blood samples were collected using 9 ml vacuum blood collection tubes containing ethylenediaminetetraacetic acid (EDTA).
[0022] The collected EDTA blood samples were centrifuged and divided into 850µl EDTA plasma, buffy coat, and erythrocyte coat aliquots, which were then stored in an automated 80 Sample archive.
[0023] Furthermore, plasma samples were selected. EDTA plasma aliquots of the selected subjects were extracted from the automated sample archive in a quasi-random order and stored at -80 in the freezer.
[0024] Next, thaw and pipette. Thaw the rack containing 85 aliquots. Using a pipetting robot with full sample tracking, pipette 60µl of EDTA plasma onto a polymerase chain reaction (PCR) plate.
[0025] Further, blind replicates were added and sealed. The plates were sealed with adhesive strips and stored at -80 Beforehand, two blinded replicates were added to each plate.
[0026] The samples were then shipped in sealed PCR plates on dry ice for analysis by Olink Analysis Service, along with an electronic sample list containing sample-level information.
[0027] S2. Obtaining plasma protein levels based on the plasma sample.
[0028] In one embodiment, individual plasma protein levels are detected using an Olink-based plasma protein biomarker analysis platform. That is, protein levels in individual plasma are detected using Olink's plasma protein biomarker analysis platform. The Olink technology employs a proximity extension assay, in which a pair of matched antibodies are labeled with unique complementary oligonucleotides (proximity probes), which bind to the corresponding target proteins in the sample. As a result, the probes enter a close proximity state and hybridize with each other, enabling DNA amplification of the protein signal. Antibodies targeting 2,923 unique proteins are distributed across eight 384-plex assay plates, each consisting of four dilution blocks to accommodate the different dynamic ranges of target proteins in plasma and serum. These assay plates focus on inflammatory, oncological, cardiovascular metabolic, and neurological proteins.
[0029] Specifically, the Olink platform was used to detect individual plasma protein levels and obtain detection data.
[0030] First, a proximity extension assay (PEA) technique was employed, using 2,923 specific antibodies distributed across eight 384-well plates, with each plate containing four dilution blocks (1:1 to 1:10,000) to cover the protein's dynamic range.
[0031] Furthermore, antibody-labeled complementary oligonucleotides (proximity probes) are bound to and hybridized with the target protein, and extended by DNA polymerase to generate an amplifiable DNA sequence.
[0032] Furthermore, two-stage PCR amplification and library construction were performed. Specifically, the primary PCR (PCR1) amplified protein-specific DNA sequences and combined the amplicons of the four abundant regions. The secondary PCR (PCR2) added a 96-index plate to construct sequencing libraries. After magnetic bead purification and quality control, the libraries were combined for high-throughput sequencing.
[0033] Furthermore, the sequencing data were converted into standardized protein-level expression values through the proprietary sequencing software Olink MyData Cloud to achieve quantitative analysis.
[0034] It should be noted that the above process not only eliminates the differences in experimental conditions between different running batches and different test plates, making the test data comparable, but also eliminates the dimensional differences between different variables, facilitating subsequent modeling.
[0035] S3. Obtaining the influence weight of liver fat content based on the plasma protein level.
[0036] In one embodiment, 5,320 individuals who had completed magnetic resonance imaging were screened and the DIXON (Digital Infrared Image) technique, which captures signals at different echo time points, was used to extract multi-echo signals and construct a proton density fraction (PDFF) map to quantify liver fat content. The PDFF values were taken from multiple circular regions of interest in the liver parenchyma.
[0037] Furthermore, the PDFF values were normalized to a uniform scale to facilitate correlation analysis with the normalized plasma protein levels.
[0038] Furthermore, the least absolute shrinkage and selection operator (LSS) regression algorithm was used to calculate the correlation coefficient between each standardized plasma protein level and PDFF as a reference for the influence weight.
[0039] Specifically, 5,320 participants were randomly assigned in a 7:3 ratio to a training set (3,726 participants) and a validation set (1,594 participants) for model training. The baseline population characteristics of the validation and training sets are shown in Table 1.
[0040] Table 1 Description of baseline population characteristics according to training and validation sets Variables are expressed as median (25th percentile, 75th percentile) or number of cases (percentage) Furthermore, the LASSO regression algorithm was used to build and train the model on the training set. The target variable of the model was the magnetic resonance imaging proton density fat fraction (MRI-PDFF), expressed as liver fat content. The input variables of the model included age, gender, and levels of 2,911 proteins. The LASSO regression model formula includes the LASSO regression model structure formula (linear prediction part) and the LASSO objective function formula (optimization problem).
[0041] Specifically, the LASSO regression model structure formula is as follows: in, is the target variable of the model, i.e., MRI-PDFF, which represents the liver fat content (continuous variable, following Gaussian distribution); Indicates the The expression level corresponding to the protein is the concentration of the protein in the plasma sample; is the intercept term, which represents the baseline predicted value when all input variables (age, sex, and protein concentration) are 0; is the coefficient of age, which indicates the expected change in MRI-PDFF for every 1-year increase in age when other variables remain unchanged; is the coefficient of sex, indicating the average difference of MRI-PDFF between males and females; Age is the numerical variable of age, which is the actual age of the subject; Sex is the categorical variable of sex, coded as 0 / 1, when the subject is male, Sex=1, when the subject is female, Sex=0; For the The coefficient of a protein indicates the direction and strength of the linear effect on MRI-PDFF when the protein concentration increases by 1 unit; is the error term, which obeys the normal distribution .
[0042] Furthermore, the LASSO regression model structure formula is used to establish a LASSO objective function, which is as follows: in, Indicates the The target variable for each sample, is the intercept term, is the regression coefficient of age, is the regression coefficient of gender, Indicates the The age value of the samples, Indicates the The gender classification variable of the samples, For the In the sample The concentration of protein, is the total number of samples, 2911 represents the number of protein types, is the regression coefficient, For the The regression coefficient of the protein represents the weight of its influence on liver fat content. is a regularization parameter used to control the sparsity of the model, that is, the strength of variable selection; when The bigger, the stronger the punishment, the more Being compressed to 0 means retaining less protein; The optimal value is selected through 10-fold cross validation, which means minimizing the prediction error. 2911 represents the number of protein types. Indicates the The contribution weight of each protein to MRI-PDFF is also called the influence weight: when When , the increased level of this protein is positively correlated with liver fat content; when When , the increased level of this protein was negatively correlated with liver fat content; when , is eliminated and has no contribution.
[0043] It should be noted that the objective function is optimized by minimizing the loss function and the L1 penalty, aiming to: 1. Build a high-precision prediction model to accurately estimate liver fat content.
[0044] 2. Implement automated feature selection: Identify key subsets from multiple proteins.
[0045] 3. Balance clinical and statistical needs: protect known important variables while exploring new biomarkers.
[0046] The solution to this optimization problem ultimately provides a sparse, stable, and interpretable prediction model that, while maintaining predictive accuracy, clearly shows which proteins are significantly associated with liver fat deposition; quantifies the independent effects of age and sex; and provides a list of priority targets for subsequent mechanistic studies.
[0047] S4. Using the influence weights, establish a liver fat content prediction model.
[0048] Please refer to Table 2, which shows the screened plasma proteins and weight coefficients related to liver fat content. The weight coefficient here is the influence weight mentioned in step S3.
[0049] Table 2 Components and weights of the proteomic liver fat content score In one embodiment, the influence weights are combined with the levels of different types of plasma proteins to establish a liver fat content prediction model based on plasma proteins. The liver fat content prediction model satisfies the following expression: in, represents the liver fat content calculated based on plasma protein levels, Indicates the The weight of the effect of each protein on liver fat content, 62 represents the number of proteins significantly correlated with liver fat content screened by lasso regression, Indicates the The expression level of each protein.
[0050] It should be noted that the protein risk score is the linear part of the lasso regression model, and the calculation process is: each protein level is multiplied by its corresponding weight coefficient and then summed.
[0051] Furthermore, to increase clinical applicability, a liver fat content prediction model (simplified model) based on the top 10 proteins (sorted by the absolute values of the coefficients in Table 2) was calculated. The 10-protein liver fat content prediction model satisfies the following expression: in, represents the liver fat content calculated based on plasma protein levels, Indicates the The weight of the influence of proteins on liver fat content (see Table 2), 10 represents the number of the top ten proteins significantly correlated with liver fat content screened by lasso regression, Indicates the The expression level of each protein.
[0052] S5. Achieve quantitative assessment of liver fat content through the liver fat content prediction model.
[0053] In one embodiment, the area under the receiver operating characteristic curve (AUC) is first used to test the ability of the liver fat content prediction model to identify high liver fat content (fatty liver) and obtain a test result.
[0054] Specifically, the relationship between liver fat content calculated based on 62 plasma protein levels and proton density fat fraction on magnetic resonance imaging in the training and validation sets can be found in Figure 2 , protein risk scores on the training set and validation set All were highly correlated with MRI-PDFF (Spearman correlation coefficient = 0.70).
[0055] Furthermore, the relationship between liver fat content calculated based on the levels of 10 plasma proteins (including IGFBP2, FABP4, MET, CPM, CES1, IGFBP1, CDHR2, RBP5, ERBB2, and SSC5D) and proton density fat fraction on magnetic resonance imaging in the training and validation sets can be found in Figure 3, the protein risk score S was highly correlated with MRI-PDFF in both the training and validation sets (Spearman correlation coefficient = 0.66).
[0056] The discriminatory capabilities of the two protein risk scoring models for liver fat content proposed in the present invention and existing clinical models are shown in Table 3.
[0057] Table 3 Evaluation of the discriminative ability of the model in the training set and validation set The clinical model included: age, sex, race, body mass index, Townsend deprivation index, smoking status, alcohol consumption, systolic blood pressure, low-density lipoprotein cholesterol, and glycated hemoglobin concentration.
[0058] As shown in Table 3, in both the training and validation sets, the liver fat content prediction model based on 62 plasma proteins had the best discriminative ability for patients with MRI-PDFF > 5% (fatty liver) (AUC>0.85), and its AUC value was significantly higher than that of the clinical model (AUC = 0.77). The liver fat content prediction model based on 10 plasma proteins (simplified model) also had good discriminative ability for patients with MRI-PDFF > 5% (fatty liver) (AUC [训练集]=0.839,AUC[验证集] =0.850), and its AUC values were significantly higher than those of the clinical model (AUC =0.77).
[0059] The present invention established two liver fat content prediction models through proteomic analysis optimization: a complex model containing 62 proteins and a simplified model that only requires 10 key proteins. Clinically verified, the simplified model maintains excellent diagnostic performance (AUC [验证集] =0.850), significantly reducing computational complexity and facilitating rapid clinical detection and application. Based on its excellent operability and diagnostic efficacy, this paper recommends a simplified 10-protein model as a standard prediction tool. This model meets the clinical need for convenient testing while ensuring accurate diagnosis of fatty liver disease, and has broad clinical application value. It should be noted that the 95% confidence interval (CI) is an indicator used in statistics to measure the uncertainty of estimated values, including AUC values, means and proportions. It is usually written as: estimated value (95% CI: lower limit, upper limit), which means that if the experiment is repeated 100 times using the same method, approximately 95 of the calculated intervals will contain the true AUC value, that is, the true value has a 95% probability of falling within this range.
[0060] Specifically, for the liver fat content prediction model, since its AUC value is significantly higher than that of the clinical model, it is suitable for clinical screening and early intervention; P<0.001 indicates that the predictive value of the liver fat content prediction model is statistically significantly higher than that of traditional clinical indicators. For the clinical model, the AUC value was 0.770, which was a moderate discriminatory ability, indicating that traditional clinical indicators had limited predictive ability for liver fat content.
[0061] The protein risk score achieved similar AUCs in the training and validation sets, indicating that the model was not overfitting and had good generalization ability. The clinical model achieved an AUC of 0.770 in both the training and validation sets, indicating that its predictive ability was stable but low.
[0062] Existing clinical models rely on conventional metabolic indices, such as body mass index, blood lipids, and blood glucose, which are correlated with liver fat content. However, these indices lack specificity. For example, obese individuals may not necessarily have fatty liver; lean individuals can also have fatty liver. They also fail to consider molecular biological mechanisms, such as proteomics and inflammatory markers. The present invention, however, leverages proteomics and specific biomarkers to directly reflect abnormal liver fat metabolism, significantly improving predictive accuracy. Furthermore, the simplified protein risk score of the present invention significantly outperforms traditional clinical models in distinguishing patients with MRI-PDFF >5% (fatty liver), with consistent performance across training and validation sets, demonstrating its robustness. Furthermore, the present prediction model has the potential to become a new standard for noninvasive diagnosis of liver fat content, addressing the shortcomings of existing clinical indices.
[0063] See Figure 4 , Figure 4 This is a schematic diagram of the structure of a plasma proteomics-based liver fat quantitative assessment system according to an embodiment of the present invention. The system includes an input device, a processor, an output device, and a memory, wherein the input device, the processor, the output device, and the memory are interconnected. The memory is used to store a computer program comprising program instructions, and the processor is configured to invoke the program instructions. The system utilizes the plasma proteomics-based liver fat quantitative assessment method described above.
[0064] In this embodiment, the input device includes a plasma sample collection module and a protein detection device, which is used to receive the plasma sample from the subject, quantitatively analyze the plasma protein expression level through a high-throughput proteomics detection platform, such as Olink, and transmit the original detection data, i.e., the concentrations of 10-62 proteins identified by the present invention, to the processor; the plasma sample collection module includes a centrifuge and a micropipette.
[0065] The processor includes a data preprocessing module, a machine learning modeling engine, and a model optimization module. Its functions include: normalizing the original protein data to eliminate batch effects; executing the liver fat content prediction model of the present invention and outputting the predicted value of liver protein content.
[0066] The output device includes a visualization terminal and a clinical decision support interface, and its functions include: outputting liver fat content prediction results, including protein risk score, fatty liver grade and key protein contribution analysis.
[0067] Provide ROC curve analysis results to assist in judging the model's identification ability and recommend personalized intervention plans, such as lifestyle adjustments or further imaging examinations.
[0068] The memory includes a local database and a cloud storage system, supporting structured data such as protein expression profiles, clinical data, and unstructured data such as model parameters and ROC curve data. Its functions include: archiving subjects' plasma protein data, clinical information, model training parameters, and historical prediction results; supporting incremental learning, dynamically optimizing the prediction model through new data, and improving generalization capabilities.
[0069] The core breakthrough of this system is to convert complex proteomic data into clinically applicable non-invasive diagnostic tools, significantly improving the popularity and accuracy of early screening for fatty liver.
[0070] In summary, the present invention aims to develop a non-invasive, high-precision, low-cost, easy-to-operate method for quantitative assessment of liver fat based on plasma proteomics, so as to solve the technical difficulties of existing technologies in terms of invasiveness, accessibility, dynamic monitoring, accuracy and cost-effectiveness. The technical problems are solved by the following technical solutions: First, non-invasive detection. Proteomic analysis is performed using plasma samples to avoid the invasive operation of traditional liver biopsy. The principle is to use large-scale proteomics methods to screen out specific protein marker groups related to liver fat content from plasma to achieve non-invasive detection; second, high-precision quantitative assessment. Based on machine learning algorithms, a liver fat content prediction model with high universality is constructed using only 10 proteins. The principle is that machine learning algorithms can learn complex relationships in data, integrate and analyze multi-dimensional data, optimize marker combinations, and improve the accuracy and specificity of detection; third, high specificity and reliability. Screening for a panel of protein markers specific for liver fat, optimizing the detection algorithm, and improving the specificity and reliability of the test. This is done by validating large-scale clinical samples to screen for protein markers highly correlated with liver fat content, and combining this with multivariate statistical analysis to ensure the stability and reliability of the test results. Fourth, low cost and high accessibility. Developing rapid detection kits based on enzyme-linked immunosorbent assays or mass spectrometry to reduce equipment dependency and testing costs. This is done by standardizing the kits and simplifying the operating procedures, making the test methods applicable to medical institutions at all levels, significantly reducing testing costs and equipment barriers.
[0071] The present invention effectively solves the following technical problems through the above technical solutions: 1. Invasiveness. Only non-invasive blood testing is required, eliminating the need for a liver biopsy. This effectively addresses the issues of poor patient compliance and surgical risks associated with the invasive procedure of traditional liver biopsy.
[0072] 2. Accuracy: By screening specific protein marker panels and combining them with machine learning algorithms, the accuracy and reliability of detection were significantly improved, solving the sampling error problem caused by spatial heterogeneity in traditional methods.
[0073] 3. Accessibility and cost-effectiveness: On the one hand, by developing a low-cost, easy-to-use test kit, we address the reliance of MRI-PDFF on expensive specialized imaging equipment, improving the accessibility of the technology. On the other hand, by optimizing the testing process and reducing equipment dependency, we significantly reduce testing costs, improve the cost-effectiveness of the technology, and address the high testing costs of existing technologies.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and description of the present invention.
Claims
1. A method for quantitative assessment of liver fat based on plasma proteomics, characterized in that: The method comprises the following steps: Obtaining plasma samples from individuals; obtaining a plasma protein level based on the plasma sample; Based on the plasma protein level and combined with the proton density fat fraction of magnetic resonance imaging, the influence weight of liver fat content is obtained; Using the influence weights, establishing a liver fat content prediction model; The liver fat content prediction model is used to achieve quantitative evaluation of the liver fat content.
2. The method for quantitative assessment of liver fat based on plasma proteomics according to claim 1, characterized in that: Obtaining the plasma protein level according to the plasma sample comprises: Obtaining individual plasma protein level test results using a plasma protein biomarker analysis platform based on the plasma sample; The detection results are normalized to obtain standardized plasma protein levels.
3. The method for quantitative assessment of liver fat based on plasma proteomics according to claim 1, characterized in that: The influence weight of liver fat content obtained based on the plasma protein level and the proton density fat fraction of magnetic resonance imaging includes: Based on the normalized plasma protein levels, training and validation sets were allocated; Based on the training set and combined with the proton density fat fraction of magnetic resonance imaging, the influence weights of different plasma protein levels on liver fat content are obtained.
4. The method for quantitative assessment of liver fat based on plasma proteomics according to claim 3, characterized in that: The method of obtaining the influence weights of different plasma protein levels on liver fat content based on the training set and combined with the proton density fat fraction of magnetic resonance imaging includes: Based on the training set, a lasso regression model is established using a lasso regression algorithm; Using the age, gender, and levels of different types of proteins of the individuals in the training set to train the lasso regression model, thereby obtaining a trained lasso regression model; Based on the trained lasso regression model and combined with the proton density fat fraction of magnetic resonance imaging, the influence weights of different plasma protein levels on liver fat content are obtained.
5. The method for quantitative assessment of liver fat based on plasma proteomics according to claim 4, characterized in that: The structural formula of the lasso regression model is as follows: in, represents the proton density fat fraction on magnetic resonance imaging, Indicates the The expression levels of the proteins is the intercept term, is the regression coefficient of age, is the regression coefficient of gender, A numeric variable representing age, A categorical variable representing gender, For the The regression coefficient of protein is the error term, 2911 represents the number of protein types; the objective function of the lasso regression model is as follows: in, Indicates the The target variable for each sample, is the intercept term, is the regression coefficient of age, is the regression coefficient of gender, Indicates the The age value of the samples, Indicates the The gender classification variable of the samples, For the In the sample The concentration of protein, is the total number of samples, is the regularization parameter, 2911 represents the number of protein types, is the regression coefficient, For the The regression coefficients of proteins.
6. The method for quantitative assessment of liver fat based on plasma proteomics according to claim 4, characterized in that: The influence weights of different plasma protein levels on liver fat content are obtained based on the trained lasso regression model and combined with the proton density fat fraction of magnetic resonance imaging, including: Based on the trained lasso regression model, normalizing the proton density fat fraction of the magnetic resonance imaging to obtain a normalized result; Based on the normalization processing result, 10-fold cross validation is used to optimize the hyperparameters of the trained lasso regression model and obtain an optimization result; The influence weights of different plasma protein levels on liver fat content are obtained through the optimization results.
7. The method for quantitative assessment of liver fat based on plasma proteomics according to claim 1, characterized in that: The use of the influence weights to establish a liver fat content prediction model includes: The influence weights are used in combination with plasma protein levels to establish a liver fat content prediction model, which satisfies the following expression: in, A simplified protein risk score indicating liver fat content, Indicates the The weight of the effect of each protein on liver fat content, 10 represents the number of the top ten proteins significantly correlated with liver fat content screened by lasso regression, Indicates the The expression levels of the proteins were compared; the 10 proteins screened out included IGFBP2, FABP4, MET, CPM, CES1, IGFBP1, CDHR2, RBP5, ERBB2 and SSC5D.
8. The method for quantitative assessment of liver fat based on plasma proteomics according to claim 1, characterized in that: The method of implementing a quantitative assessment of liver fat content by using the liver fat content prediction model includes: Using the area under the receiver operating characteristic curve (AOC) to test the ability of the liver fat content prediction model to identify high liver fat content and obtain test results; According to the test results, a quantitative assessment of the liver fat content is achieved.
9. The method for quantitative assessment of liver fat based on plasma proteomics according to claim 1, characterized in that: The use of the area under the receiver operating characteristic curve to test the ability of the liver fat content prediction model to identify high liver fat content and obtain the test results includes: In both the training and validation sets, the protein risk score for liver fat content had the best discriminatory ability for individuals with a proton density fat fraction >5% (fatty liver) on magnetic resonance imaging, with significantly higher area under the receiver operating characteristic curve values than the clinical model.
10. A plasma proteomics-based liver fat quantitative assessment system, the system using the plasma proteomics-based liver fat quantitative assessment method according to any one of claims 1 to 9, characterized in that: The system includes an input device, a processor, an output device and a memory, wherein the input device, the processor, the output device and the memory are connected to each other, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions.
Citation Information
Patent Citations
Non-alcoholic fatty liver disease diagnosis model construction method, system, equipment and medium
CN118398221A
Method and system for constructing liver fat noninvasive quantitative evaluation model based on PCD-CT
CN118782254A
Stomach cancer early-stage multi-molecular diagnosis model construction method based on plasma protein
CN119626519A
Screening diagnosing method for fatty liver degeneration accompanying abdominal obesity
RU2684201C1