Biomarker for predicting recurrence risk of liver cancer and application of biomarker
By screening and constructing a biomarker model based on proteomics, the problem of inaccurate prediction of liver cancer recurrence risk in the prior art is solved, efficient and accurate prediction of liver cancer recurrence risk is achieved, the sensitivity and specificity of diagnosis is improved, and personalized treatment is supported.
Patent Information
- Application Number
- CN202510725742.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The lack of high sensitivity proteomic biomarkers in the prior art is used to predict the risk of liver cancer recurrence, which leads to cumbersome and inaccurate diagnosis, making it difficult to effectively evaluate the risk of liver cancer recurrence after surgery.
Biomarkers such as IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A and B3GNT2 were screened through proteomics, and a risk prediction model for liver cancer recurrence was constructed, and analysed using high-performance liquid chromatography-tandem mass spectrometry technology was used to perform risk prediction with machine learning algorithms.
It realizes accurate, non-invasive and rapid prediction of the risk of recurrence of liver cancer, improves the sensitivity and specificity of diagnosis, reduces the risks of misdiagnosis and missed diagnosis, provides personalized treatment suggestions, and improves the treatment effect and quality of survival of liver cancer patients.
Smart Images

Figure CN120254288A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical technology, and particularly relates to a biomarker for predicting the recurrence risk of liver cancer and its application. Background Art
[0002] Liver cancer is the third leading cause of cancer-related deaths globally. Liver cancer is mainly divided into two major categories: hepatocellular carcinoma (HCC) and cholangiocarcinoma (ICC). Among them, hepatocellular carcinoma is the most common type, accounting for 90% of liver cancer cases. The incidence of liver cancer is related to multiple factors, including chronic hepatitis virus infection (such as hepatitis B virus HBV and hepatitis C virus HCV), chronic liver diseases (such as fatty liver or cirrhosis), alcoholism, and metabolic diseases (such as diabetes). In addition, autoimmune hepatitis, hemochromatosis, and exposure to certain environmental toxins are also potential risk factors for liver cancer.
[0003] The incidence of liver cancer is related to multiple factors, including chronic hepatitis virus infection, such as hepatitis B virus HBV and hepatitis C virus HCV; chronic liver diseases, such as fatty liver or cirrhosis; alcoholism, and metabolic diseases, such as diabetes. In addition, autoimmune hepatitis, hemochromatosis, and exposure to certain environmental toxins are also potential risk factors for liver cancer. Hepatectomy is the most important means for patients with hepatocellular carcinoma to achieve long-term survival benefits. However, even for early-stage liver cancer, about 30% of patients will experience recurrence within 2 to 5 years after surgery. Preventing the recurrence and metastasis of hepatocellular carcinoma and prolonging survival are the key goals pursued in the field of liver cancer treatment and one of the focuses of current clinical research work. If patients who do not benefit from surgical treatment or experience progression can be subjected to risk prediction and the treatment plan can be adjusted in a timely manner (such as adjuvant radiotherapy and chemotherapy, secondary surgical resection, targeted therapy, or immunotherapy, etc.), the clinical benefits of patients can be maximized. However, there is currently no consensus on the best tool for predicting the recurrence risk of hepatocellular carcinoma after surgery in the preoperative scenario.
[0004] Proteomics is the science that studies the protein composition, localization, changes, and their interaction laws in cells, tissues, or organisms, including the study of protein expression patterns and proteome function patterns. With the development of mass spectrometry technology, liquid chromatography-tandem mass spectrometry (LC-MS / MS) has become the most important tool in proteomics research. The development of proteomics is of great significance for finding disease diagnostic markers, screening drug targets, and toxicological research, and has therefore been widely applied in medical research.
[0005] Although there are currently some reports on biomarkers for evaluating the recurrence risk of liver cancer prognosis, they are basically evaluated and detected at the gene level, such as SNP sites, etc. Not only is the detection process more cumbersome and highly instrument-dependent, but the diagnostic sensitivity is still not high. There is a lack of biomarkers for diagnosing the recurrence risk of liver cancer in clinical practice, especially the discovery of high-sensitivity proteomic biomarkers for diagnosing the recurrence risk of liver cancer is of great significance. Summary of the Invention
[0006] In view of the problems existing in the prior art, the present invention provides a biomarker for predicting the recurrence risk of liver cancer and its application. By using proteomics methods, proteins with significantly different abundance levels in the blood of two groups of patients with recurrence and non-recurrence after liver cancer treatment are analyzed, and biomarkers for predicting the recurrence or metastasis risk of liver cancer are screened out. Furthermore, a prediction model for the recurrence or metastasis risk of liver cancer is constructed, which can accurately, non-invasively and efficiently predict the recurrence or metastasis risk of liver cancer and meet the clinical needs.
[0007] The first aspect of the present invention provides the use of a substance for detecting a biomarker in the preparation of a reagent for predicting the recurrence or metastasis risk of liver cancer, wherein the biomarker comprises one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A and B3GNT2.
[0008] The biomarker obtained by the present invention through proteomics can accurately predict the recurrence or metastasis risk of liver cancer, which is beneficial for doctors to judge the severity of the patient's condition and whether treatment needs to be adjusted, so as to provide more accurate treatment suggestions for patients and truly benefit liver cancer patients.
[0009] By using proteomics methods, the present invention collects plasma samples of patients who have recurrence within a short period (such as 1 to 5 years) after liver cancer treatment and patients who do not have recurrence within a short period. Different samples are analyzed by high-performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS). Based on the orthogonal partial least squares discriminant analysis and significance analysis methods, proteins with significant differences between liver cancer recurrence and non-recurrence are first screened, and 10 differential proteins with obvious relevance to the recurrence risk of liver cancer are obtained. These 10 proteins can be used to distinguish the risk of recurrence or metastasis after liver cancer treatment and have certain diagnostic efficacy.
[0010] Among them, the IGFBP4 is a protein or amino acid sequence with the UniProt database number P22692; ORM1 is a protein or amino acid sequence with the UniProt database number P02763; RNASE1 is a protein or amino acid sequence with the UniProt database number P07998; DEFA3 is a protein or amino acid sequence with the UniProt database number P59666; PGM5 is a protein or amino acid sequence with the UniProt database number Q15124; FBLN5 is a protein or amino acid sequence with the UniProt database number Q9UBX5; DCD is a protein or amino acid sequence with the UniProt database number P81605; SERPINB1 is a protein or amino acid sequence with the UniProt database number P25407; CORO1A is a protein or amino acid sequence with the UniProt database number P31146; B3GNT2 is a protein or amino acid sequence with the UniProt database number Q9NY97.
[0011] The present invention also surprisingly discovers that some of the protein markers obtained through proteomic screening are known markers that can be used for other cancers. For example, DEFA3 has been reported to be used for the prediction of respiratory tract infections, but this screening discovers that this marker can also be used for the prediction of the risk of recurrence or metastasis of liver cancer. It can be seen that for many different tumors, their protein markers are not completely separated or irrelevant. In fact, there are many cross - relationships or influences. Many protein markers can be used for the prediction of early - stage cancer, and also for the prognostic diagnosis of cancer. Even many protein markers can be used for the prediction and diagnosis of different stages of many different cancers. Therefore, there are still many brand - new functions in the field of proteomics waiting to be explored, and the market prospect is very broad.
[0012] In some embodiments, the biomarker can be 1 marker or a combination of several markers, such as a combination of 2 markers, a combination of 3 markers, a combination of 4 markers, a combination of 5 markers, a combination of 6 markers, a combination of 7 markers, a combination of 8 markers, a combination of 9 markers, a combination of 10 markers. In some specific embodiments, the biomarker comprises a combination of more than 2 markers, such as a combination of more than 3 markers. When a combination of markers constructs a prediction model for the risk of recurrence or metastasis of liver cancer, its AUC value is 0.801 - 0.949, the sensitivity is 82.9% - 95.4%, and the specificity is 83.3% - 95.5%.
[0013] Furthermore, the biomarker includes IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD.
[0014] Furthermore, to study the diagnostic efficacy of the research department for the risk of liver cancer recurrence, it is also necessary to combine different differential proteins to construct a diagnostic model, sort them according to the importance obtained by screening, and select different numbers of differential proteins with higher rankings for combination respectively to obtain 10 protein markers. After verification, it is found that constructing a model based on 7 protein markers has good risk prediction ability in the diagnosis of the risk of liver cancer recurrence. The data of detecting samples with the risk of liver cancer recurrence show that just using these 7 biomarkers to predict the risk of liver cancer recurrence, the AUC value can reach 0.949, and the diagnostic performance is good.
[0015] In some embodiments, the product is selected from one or more of reagents, reagent kits, chips, probes or membrane strips. The product is a product with the above-mentioned biomarkers as the detection target, and includes biological reagents and reagent kits suitable for the detection of the biomarkers, such as sample pretreatment reagents, antigens or antibodies; it can also be developed into standardized reagents or reagent kits, chips, probes or membrane strips suitable for the biomarkers.
[0016] Furthermore, the recurrence or metastasis risk refers to recurrence or metastasis within three years after liver cancer treatment; the recurrence or metastasis of liver cancer includes in-situ or adjacent area recurrence of liver cancer, and metastasis of liver cancer.
[0017] It can be understood that the short term in "recurrence within three years after liver cancer treatment" here is not an absolute and unchanging time node. It is only a time node summarized based on current clinical experience. With the passage of time or the improvement of other treatment methods, this time node or the length of time may change. For example, it may be within one year, within 360 days, or within half a year, within 180 days, or within two years, within two and a half years, within three and a half years, etc. It is possible to have recurrence of liver cancer after treatment.
[0018] In some embodiments, the recurrence or metastasis of liver cancer means that after surgical treatment of radical tumor resection, the patient has in-situ or adjacent area recurrence of liver cancer within three years, or has metastasis of liver cancer, such as lymph node metastasis, bone metastasis, etc.
[0019] In some embodiments, the product is used for predicting patients with the risk of liver cancer recurrence or metastasis.
[0020] In some embodiments, the product is used for detecting the detection amount of biomarkers in biological samples.
[0021] In some specific embodiments, the biological sample is selected from one or more of saliva, blood, urine, plasma, serum and spinal fluid.
[0022] In some specific embodiments, the detection amount includes the presence or absence, relative abundance or concentration of biomarkers.
[0023] The present invention screens for biomarkers for predicting the risk of hepatocellular carcinoma recurrence from blood. These biomarkers show significant differences in the blood of populations with and without hepatocellular carcinoma recurrence. By collecting a blood sample, it is possible to predict or assist in diagnosing whether an individual has hepatocellular carcinoma recurrence or not, or to predict or assist in diagnosing whether an individual has no recurrence, in-situ recurrence, or distant recurrence after hepatocellular carcinoma surgery by detecting these biomarkers in the individual's blood.
[0024] Further, the detection methods generally include radiological methods, immunological methods, fluorescence methods, flow-through fluorescence methods, latex turbidimetry methods, biochemical methods, enzymatic methods, hybridization methods, gas chromatography-mass spectrometry methods, liquid chromatography-mass spectrometry methods, chromatography methods, chemiluminescence methods, magnetoelectric methods, or photoelectric conversion methods.
[0025] The presence, absence, or high or low content of these biomarkers here is a relative concept. For example, when comparing the hepatocellular carcinoma recurrence group and the non-recurrence group, the content of these specific biomarkers is compared with the hepatocellular carcinoma recurrence group or the non-recurrence group as a reference. There may be some biomarkers with a relatively higher content in hepatocellular carcinoma recurrence compared to the non-recurrence group, and this increase is statistically significant, such as a significant or highly significant increase. Therefore, when judging these biomarkers, if it is a single biomarker, if the probability of a certain risk occurrence increases and the content of this biomarker changes, this change may be a relative increase or may also be a relative decrease, and the difference in this relative increase or decrease is significantly different, and of course, it can also be highly significantly different. Therefore, no matter what means are used for detection, a pre-specified value (cut-off value) can be used as a standard. If it is higher than this value, it is considered that the content has changed, and having such a result can all be used for prediction or diagnostic value.
[0026] Therefore, in some aspects, the biomarker of the present invention can be obtained by detecting the content of the biomarker in a sample by any known method, such as liquid chromatography, gas chromatography, mass spectrometry, LC-MS, gas chromatography-mass spectrometry (GC-MS), chromatography-mass spectrometry (CC-MS), liquid chromatography-tandem mass spectrometry (LC-MS-MS), nuclear magnetic resonance spectroscopy (NMR), immunochromatographic test strip, immunoreaction chip, capillary electrophoresis, infrared spectroscopy, etc. As long as it can be used to detect the content of protein biomarkers in a sample, it can be used for the diagnosis of recurrent and non-recurrent liver cancer. As long as the content of protein biomarkers in a sample can be detected, it can be used to predict or diagnose the probability of occurrence of a certain disease. It can be understood that the detection here is for the samples of individuals, and then compared with a preset standard, and the comparison results are used to judge or predict the occurrence status of the disease. For example, it can be used to predict the probability of recurrence risk of liver cancer. Such prediction or diagnosis is whether it occurs within a certain time. Of course, such detection can be continuous detection, and the progress of the disease can be inferred from the change of the content of certain substances.
[0027] In some ways, the relative abundance is the peak area of the biomarker in the detection spectrum obtained by high performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of a certain biomarker measured in a control sample is 300, and the average peak area measured in a recurrent liver cancer sample is 1800, then the abundance of the biomarker in the sample is considered to be 6 times that in the control sample.
[0028] The second aspect of the present invention provides the use of proteins as biomarkers in the preparation of products for predicting the risk of recurrence or metastasis of liver cancer. The proteins include one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2. The genes of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2 can be used for the auxiliary judgment of the risk of recurrence or metastasis of liver cancer, the evaluation of drug efficacy, etc. The inventors of the present invention found that IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2 are closely related to the risk of recurrence or metastasis of liver cancer.
[0029] The third aspect of the present invention provides a product for predicting the risk of recurrence or metastasis of liver cancer, including substances for detecting biomarkers, and the biomarkers include one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2.
[0030] The fourth aspect of the present invention provides a biomarker combination for predicting the risk of recurrence or metastasis of liver cancer, including IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD.
[0031] The fifth aspect of the present invention provides a method for constructing a prediction model for the risk of recurrence or metastasis of liver cancer for non-disease diagnosis purposes, including the following:
[0032] 1) Construct a sample data set based on the detected amounts of biomarkers in biological samples, where the biomarkers include one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2;
[0033] 2) Divide the data set into a test set and a training set, and construct and train the prediction model for the risk of recurrence or metastasis of liver cancer through machine learning methods.
[0034] In some embodiments, the biomarkers are IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD.
[0035] In some embodiments, the machine learning method is selected from at least one of gradient boosting algorithm, random forest algorithm, support vector machine algorithm, decision tree algorithm, K-nearest neighbor algorithm, logistic regression algorithm, and neural network algorithm. For example, it can be selected from the gradient boosting algorithm. Unlike generalized linear models that can output model formulas and cut-off values, all calculations in the gradient boosting algorithm are directly completed by machine learning. By directly inputting the detection values into the software system, the prediction results can be directly obtained.
[0036] In some embodiments, a prediction model for the risk of recurrence or metastasis of liver cancer is constructed using multiple machine learning methods, and it is preliminarily confirmed that for any one of the newly screened biomarkers alone, the change in its concentration can be used to distinguish between patients with recurrent liver cancer and those without recurrence, indicating that these biomarkers have extremely high diagnostic value.
[0037] The present invention discovers that by detecting the detected amounts of biomarkers in biological samples, then inputting the detected amounts into the formula of the prediction model for the risk of recurrence or metastasis of liver cancer to obtain the prediction score Y of the prediction model, and comparing this prediction score Y with the threshold (cut off) defined by the Youden index, if the prediction score Y > the threshold (cut off), it is judged as recurrence of liver cancer; if the prediction score y ≤ the threshold (cut off), it is judged as no recurrence of liver cancer.
[0038] In some embodiments, the equation of the prediction model for the risk of recurrence of liver cancer is:
[0039]
[0040] Among them, Y is the predicted score, i represents the i-th biomarker, m represents the number of biomarkers (m = 7), Xi represents the detected value (μg / mL) of the i-th biomarker, Ki represents the coefficient of the i-th biomarker, and b is a constant 2.5873933; the coefficients of the 7 biomarkers are as follows:
[0041]
[0042] When Y ≤ the cut-off value, the subject is a patient with non-recurrent liver cancer; when Y > the cut-off value, the subject is a patient with recurrent liver cancer. Specifically, the cut-off value is 0.5180802.
[0043] In some embodiments, it further includes step 3) using a test set to test the liver cancer recurrence or metastasis risk prediction model. After training is completed, the trained prediction model is verified using the test set. At the same time, the AUC value, specificity, and sensitivity are used as evaluation indicators to evaluate the effect of the prediction model.
[0044] The Receiver Operating Characteristic Curve (ROC curve) is a curve plotted with the true positive rate (sensitivity) as the ordinate and the false positive rate (1 - specificity) as the abscissa according to a series of different binary classification methods (cut-off values). The Area Under Curve of the subject curve is defined as the area under the ROC curve. The AUC value is often used to evaluate the diagnostic efficacy of the prediction model. The larger the AUC value, the better the diagnostic efficacy of the corresponding prediction model; conversely, the worse the diagnostic efficacy of the corresponding prediction model.
[0045] The prediction model constructed based on the combination of 7MP in the present invention can distinguish between recurrent and non-recurrent liver cancer after surgery. Its AUC in the training group is 0.949, the sensitivity is 0.954, and the specificity is 0.955; its AUC in the test group is 0.921, the sensitivity is 0.931, and the specificity is 0.932.
[0046] The sixth aspect of the present invention provides a liver cancer recurrence or metastasis risk prediction device, including a data acquisition unit and a calculation unit;
[0047] The data acquisition unit is used to acquire the detection amount data of the biomarker as described above in the biological sample of the subject, and the biomarker includes one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2;
[0048] The calculation unit is used to calculate and output the predicted score of the risk of liver cancer recurrence or metastasis of the subject based on the detection amount data, and make a judgment according to the threshold
[0049] In some embodiments, the biomarker includes IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD.
[0050] In some embodiments, the prediction device further includes a data storage unit and a data output unit; the data storage unit is used to store the detection amount (or detection value) of the biomarker; the data input interface is used to input the detection value of the biomarker, and the data output unit is used to output the prediction result.
[0051] Further, the detection value is the presence or absence, relative abundance, or concentration value of each biomarker.
[0052] The seventh aspect of the present invention provides a device, including a processor and a memory, the memory is used to store a computer program, and it is characterized in that the processor is used to execute the computer program stored in the memory so that the device executes the prediction method or the construction method as described above.
[0053] The eighth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program executes the prediction method or the construction method as described above when being processed.
[0054] In some embodiments, the computer-readable storage medium includes: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0055] In some embodiments, the being processed means being executed by one or more processors.
[0056] The ninth aspect of the present invention provides a method for predicting the risk of liver cancer recurrence or metastasis, including the following steps:
[0057] S1. Obtain the detection quantity data of biomarkers in the biological samples of the subjects, where the biomarkers include one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2;
[0058] S2. Process the detection quantity data by using the liver cancer recurrence risk prediction model obtained by the construction method described above to output the liver cancer recurrence or metastasis risk prediction result.
[0059] In some embodiments, the biomarkers include IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD.
[0060] On the other hand, the present invention provides a system for predicting the recurrence risk of liver cancer. The system includes a data analysis module, and the data analysis module is used to analyze the detection values of biomarkers, where the biomarkers include IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD.
[0061] Further, the data analysis module uses the detection values of biomarkers of known samples as a training set, and divides them into a liver cancer recurrence group and a non-liver cancer recurrence group according to whether liver cancer recurs, analyzes the relationship between the detection values of the liver cancer recurrence group and the non-liver cancer recurrence group, and constructs a model.
[0062] In some embodiments, a combined diagnosis model for predicting the recurrence risk of liver cancer is constructed by combining multiple machine learning methods, and it is initially confirmed that for any one of the selected novel biomarkers used alone, the change in its concentration can be used to distinguish high-risk and low-risk populations of liver cancer recurrence, indicating that these biomarkers have extremely high diagnostic value.
[0063] In some embodiments, the equation of the constructed model is:
[0064] Among them, Y is the predicted score, i represents the i-th biomarker, m represents the number of biomarkers (m = 7), Xi represents the detection value (μg / mL) of the i-th biomarker, Ki represents the coefficient of the i-th biomarker, and b is a constant 2.5873933; the coefficients of the 7 biomarkers are:
[0065]
[0066] When Y ≤ the cut-off value, the subject is non-recurrent liver cancer; when Y > the cut-off value, the subject is recurrent liver cancer. Specifically, the cut-off value is 0.5180802.
[0067] Further, the system further includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values of biomarkers; the data input interface is used to input the detection values of biomarkers, and the data output interface is used to output the prediction results.
[0068] In another aspect, the present invention provides a use of a biomarker for preparing a reagent for predicting non - recurrence, in - situ recurrence or distant metastasis after surgery in liver cancer patients, and the biomarker includes any one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2.
[0069] Further, the biomarker includes IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, and CORO1A.
[0070] In another aspect, the present invention provides a kit for predicting non - recurrence, in - situ recurrence or distant metastasis after surgery in liver cancer patients, and the kit includes a detection reagent for the biomarker as described above.
[0071] In another aspect, the present invention provides a biomarker combination for predicting non - recurrence, in - situ recurrence or distant metastasis after surgery in liver cancer patients, and the combination includes IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, and CORO1A.
[0072] In another aspect, the present invention provides a system for predicting non - recurrence, in - situ recurrence or distant metastasis after liver cancer surgery, and the system includes a data analysis module, and the data analysis module is used to analyze the detection values of biomarkers, and the biomarkers include IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, and CORO1A.
[0073] Further, the distant metastasis also includes simultaneous in - situ recurrence and distant metastasis, and as long as distant metastasis occurs, it is classified into the distant metastasis group.
[0074] Further, the data analysis module uses the detection values of biomarkers of known samples as a training set, and according to the situation after surgery of liver cancer patients, it is divided into a non - recurrence group, an in - situ recurrence group, and a distant metastasis group, analyzes the relationship between the detection values of the non - recurrence group, the in - situ recurrence group, and the distant metastasis group, and constructs a model.
[0075] Further, the model is constructed based on the gradient boosting algorithm.
[0076] Unlike generalized linear models, the gradient boosting algorithm does not output a model formula and a cut-off value. All calculations are directly performed by machine learning. By directly inputting the test values into the software system, the prediction results can be directly obtained.
[0077] Furthermore, the system further includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the test values of biomarkers; the data input interface is used to input the test values of biomarkers, and the data output interface is used to output the prediction results.
[0078] The beneficial effects of the present invention are as follows:
[0079] 1. The present invention has screened 10 novel biomarkers that can predict the risk of liver cancer recurrence or metastasis, and developed a new combination of protein biomarkers, which can effectively evaluate and diagnose the risk of liver cancer recurrence or metastasis, effectively distinguish between recurrent and non-recurrent liver cancer, and effectively distinguish between in-situ recurrence and distant metastasis of liver cancer. This not only improves the diagnostic accuracy of liver cancer recurrence and metastasis, but also provides an important biomarker basis for the personalized treatment and prognostic monitoring of liver cancer patients. In addition, the discovery and application of these biomarkers are expected to improve the treatment effect and quality of life of liver cancer patients, and provide a scientific basis for the early diagnosis and treatment strategy selection of liver cancer. Compared with traditional detection methods, the risk of misdiagnosis and missed diagnosis is reduced, providing strong support for the early detection and intervention of diseases.
[0080] 2. The combined differential diagnosis model of 7 biomarkers constructed by the present invention is convenient, fast, and the test results are highly consistent with the clinical gold standard test results. At the same time, the cost of predicting the risk of liver cancer recurrence is significantly reduced, and it has good application prospects.
[0081] 3. Based on the screened biomarkers for predicting the risk of liver cancer recurrence, a three-classification model that can simultaneously distinguish between the non-recurrence group, the in-situ recurrence group, and the distant metastasis group is further constructed, providing a more effective and accurate prediction and diagnosis mode. Detailed description
[0082] (1) Diagnosis or detection
[0083] The diagnosis or detection herein refers to the detection or assay of biomarkers in a sample, or the content of the target biomarker, such as the absolute content or relative content, and then it is used to illustrate whether the individual providing the sample may have or suffer from a certain disease, or the possibility of having a certain disease, based on the presence or quantity of the target biomarker. The meanings of diagnosis and detection herein can be interchanged. The result of this detection or the result of the diagnosis cannot be directly used as the direct result of having a disease, but is an intermediate result. If a direct result is to be obtained, other auxiliary means such as pathology or anatomy are required to confirm the presence of a certain disease. For example, the present invention provides a variety of new biomarkers related to the risk of liver cancer recurrence or metastasis, and the change in the content of these biomarkers has a direct correlation with whether the patient belongs to the population at risk of liver cancer recurrence or metastasis.
[0084] (2) The association between the biomarker or biological marker or differential protein and the risk of liver cancer recurrence or metastasis
[0085] The biomarker, biological marker and differential protein have the same meaning in the present invention. The association herein refers to the direct correlation between the appearance or change in the content of a certain biomarker in a sample and a specific disease. For example, a relative increase or decrease in content indicates that the possibility of having this disease is relatively higher than that of the healthy population.
[0086] If multiple different biomarkers in the sample appear simultaneously or there is a relative change in their content, it also indicates that the possibility of having this disease is relatively higher than that of the healthy population. That is to say, among the types of biomarkers, some biomarkers have a strong association with the disease, some have a weak association with the disease, or some even have no association with a certain specific disease. One or more of those biomarkers with a strong association can be used as biomarkers for diagnosing the disease, and those with a weak association can be combined with the strong ones to diagnose a certain disease, increasing the accuracy of the detection result.
[0087] Among the numerous biomarkers found in the serum of the present invention, these biomarkers can all be used to distinguish between recurrent and non-recurrent liver cancer, as well as non-recurrent, in-situ recurrence, and distant metastasis. These biomarkers can be directly detected or diagnosed individually as single biomarkers. Selecting such a biomarker indicates that the relative change in the content of this biomarker has a strong correlation with the risk of liver cancer recurrence or metastasis. Of course, it can be understood that simultaneous detection of one or more biomarkers with a strong correlation with the risk of liver cancer recurrence or metastasis can be selected. Normally understood, in some ways, selecting biomarkers with a strong correlation for detection or diagnosis can achieve a certain standard of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90%, or 95% accuracy. This indicates that these biomarkers can obtain an intermediate value for diagnosing a certain disease, but it does not mean that it can directly confirm the presence of a certain disease.
[0088] Of course, differential proteins with a larger ROC value can also be selected as diagnostic biomarkers. The so-called strength or weakness is generally calculated and confirmed through some algorithms, such as the contribution rate or weight analysis of the biomarker to the prediction of liver cancer recurrence or metastasis risk. Such calculation methods can include significance analysis (p-value or FDR value) and fold change. Multivariate statistical analysis mainly includes principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA), and orthogonal partial least squares discriminant analysis (OPLS-DA). Of course, other methods are also included, such as ROC analysis, etc. Of course, other model prediction methods are also possible. When specifically selecting biomarkers, the differential proteins disclosed in the present invention can be selected, or other existing and well-known biomarker combinations can be selected or combined for prediction through model methods.
[0089] (3) Definition of disease terms
[0090] Liver cancer is a malignant tumor that originates from liver cells and is usually manifested as abnormal cell growth and proliferation in liver tissue. These cells may invade the tissues surrounding the liver or spread to other parts of the body through the blood and lymphatic systems. The formation process of liver cancer may involve genetics, chronic hepatitis, cirrhosis, smoking, obesity, diabetes, and other lifestyle factors.
[0091] Early-stage liver cancer refers to liver cancer in the initial stage of development, without obvious vascular invasion, distant metastasis, or extensive invasion of surrounding tissues. This type of liver cancer is usually small in size (diameter ≤ 5 cm), has a low degree of malignancy, and can achieve a better prognosis through surgery or other local treatments.
[0092] Recurrence of liver cancer: It refers to the reappearance of liver cancer or its spread to other parts of the body after treatment. On the one hand, the recurrent liver cancer may occur in the primary site of the liver or its vicinity, which is called local recurrence (in-situ recurrence). Local recurrence means that the cancer has not spread widely but reappears in the liver or its directly adjacent tissues. On the other hand, liver cancer may also spread to other distant parts of the body, and this situation is called distant recurrence or metastasis. Even for early-stage liver cancer, about 30% of patients are prone to recurrence within 2-3 years after surgery, and the risk of recurrence is significantly reduced after 5 years. The recurrence of liver cancer is closely related to the tumor stage, the standardization of treatment, and postoperative management. The prognosis of liver cancer (prognosis refers to the expected survival situation after disease treatment) is not good because liver cancer is an aggressive malignant tumor with unclear early symptoms and is only discovered at an advanced stage, which leads to relatively poor treatment effects and prognosis.
[0093] Metastasis of liver cancer: Distant metastasis indicates that the cancer has spread to organs or tissues outside the liver through the blood or lymphatic system, such as the lungs, bones, etc. This type of recurrence usually indicates that the cancer has progressed to an advanced stage, and the complexity and difficulty of treatment increase accordingly.
[0094] Therefore, early detection and intervention of liver cancer recurrence or metastasis are crucial for improving treatment effects and increasing the survival rate. Biomarkers have accurate and specific predictive effects on liver cancer recurrence or metastasis, which is beneficial to improving the treatment effects of patients and improving their prognosis.
[0095] (4) The gold standard for liver cancer diagnosis: It is pathological diagnosis (i.e., pathological confirmation of puncture biopsy or surgical resection specimens), observing the presence of cancer cells under a microscope, and clarifying the nature of the tumor in combination with immunohistochemistry or molecular detection. Brief Description of the Drawings
[0096] Figure 1 It is a volcano plot for the differential analysis of protein markers in the high-risk group and low-risk group of liver cancer recurrence in Example 1.
[0097] Figure 2 It is a result diagram of ROC and OPLS-DA analysis in the high-risk group and low-risk group of liver cancer recurrence in Example 1.
[0098] Figure 3 It is a schematic diagram of the optimization results of the hyperparameters of the glmnet algorithm in Example 2;
[0099] Figure 4 It is an ROC curve diagram of the liver cancer recurrence risk prediction model in the training group in Example 2.
[0100] Figure 5 It is an ROC curve diagram of the liver cancer recurrence risk prediction model in the test group in Example 2. Detailed implementation mode
[0101] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention and do not limit it in any way. The reagents used in this embodiment are all known products and are obtained by purchasing commercially available products.
[0102] Example 1 Screening biomarkers for the risk of liver cancer recurrence using proteomics
[0103] By collecting plasma samples from liver cancer patients after radical surgery, enriching low-abundance proteins based on the method of removing high-abundance proteins by immunoaffinity chromatography, detecting the protein abundance in the samples by a high-performance liquid chromatography-tandem mass spectrometry device, and screening for differential proteins by analyzing the differences in their abundances between liver cancer recurrence and non-recurrence patients. The specific steps are as follows:
[0104] 1.1 Sample collection
[0105] A total of 100 blood samples were collected from early liver cancer patients (diameter ≤ 5 cm) 2 weeks after surgical resection. All liver cancer patients were confirmed by pathology of living tissues. Approximately 2 ml of peripheral blood samples of the subjects were collected, placed in a vacuum tube containing EDTA anticoagulant and mixed well, and centrifuged at 120 g at room temperature for 10 minutes. The supernatant was taken and repeated twice. Then, it was centrifuged at 360 g for 20 minutes. Then, the platelet samples were collected in a centrifuge tube and stored at -80 °C for later use. The 100 patients included 72 males and 28 females, with an average age of 40 years (25 - 69 years). All enrolled patients signed an informed consent form. Among them, all liver cancers were patients diagnosed by pathological histology. Inclusion criteria: (a) No history of other malignant tumors; (b) No patients with combined other malignant tumors or autoimmune diseases.
[0106] During the following three years, the patients were followed up every 6 months, with regular AFP and imaging examinations. Once recurrence or metastasis was found, it needed to be confirmed by pathological histology. A total of 32 early liver cancer patients had recurrence within three years after surgery, including 25 cases of in-situ local recurrence and 7 cases of distant metastasis (including 5 cases of lymph node metastasis, 1 case of lung metastasis, and 1 case of bone metastasis). Then, the samples stored at -80 °C were taken out, and the samples were divided into two groups: the samples of 68 patients without recurrence were classified into the low-risk group of liver cancer recurrence, and the 32 cases with recurrence or metastasis were classified into the high-risk group of liver cancer recurrence for the screening of proteomic biomarkers.
[0107] 1.2 Sample processing and enzymatic digestion
[0108] 1) First, the plasma samples were centrifuged on a centrifuge for 15 minutes (15,000 g), the supernatant was taken, filtered, and then 14 high-abundance proteins were removed by immunoaffinity chromatography.
[0109] 2) The low-abundance fractions were concentrated to 350 μL using a 3 kDa cut-off concentration tube in a centrifuge (4000 g, 1 hour).
[0110] 3) The concentrated solution was recovered and buffer exchanged using a desalting column with a cut-off molecular weight of 7 kDa on a centrifuge (1000 g, 2 minutes). The replacement solution was AEX-A (20 mM Tris, 4 M urea, 3% isopropanol, pH 8.0).
[0111] 4) Using AEX-A as a blank, the protein concentration in the sample was determined using the biuret method (BCA). According to the sample grouping, 25 μL TCEP was added to the sample and incubated at 37°C for 30 minutes for protein reduction. Then the corresponding TMT16-plex reagent was added and incubated at room temperature in the dark for 1 hour for TMT labeling reaction. After that, the sample was buffer exchanged using Zeba columns, and the replacement fluid was AEX-A. After mixing the TMT 16-plex labeled samples, 2 mL AEX-A was added to the mixed samples, and the final volume was 5.5 mL.
[0112] 5) Filter the sample using a 0.22 μm filter and separate the TMT 16-plex labeled sample using a 2D-HPLC system. Freeze-dry the collected fractions, and finally add Trypsin-Lysin C mixed enzyme, incubate the sample at 37°C for 5 hours, and add 5 μL 10% TFA to terminate the enzymatic reaction.
[0113] 6) A total of 60 2D-HPLC fractions after enzymatic digestion were used for nanoLC-MS / MS analysis.
[0114] 1.3 LC-MS / MS data acquisition
[0115] Each sample obtained in step 1.2 was separated using the nanoliter flow rate liquid chromatography system Easy nLC-1200 and detected online with the high-resolution mass spectrometer Q Exactive HF-X. The details are as follows:
[0116] 1) Separation: Mobile phase-A is 0.1% formic acid in water, and mobile phase-B is 0.1% formic acid in acetonitrile (80% acetonitrile, 20% water). The chromatographic column consists of an enrichment column and an analytical column and is balanced with 100% mobile phase-A. The sample is loaded onto the enrichment column (100μmID×4cmL, C18, 3μm, 100A) by an autosampler and separated by the analytical column (75μmID×25cmL, C18,3μm, 100A) at a flow rate of 300 nL / min.
[0117] 2) The sample after chromatographic separation was subjected to mass spectrometry analysis using a Q Exactive HF-X mass spectrometer. The detection mode was positive ion, the parent ion scan range was 350 - 1800 m / z, the resolution of the first-order mass spectrometry was 120,000 at 200 m / z, the AGC (Automatic gain control) target was 3e6, the Maximum IT was 50 ms, and the dynamic exclusion time was 40 s. The mass-to-charge ratios of polypeptides and polypeptide fragments were collected according to the data-dependent acquisition (DDA) method: 20 secondary spectra (MS / MS, MS2 scan) were collected after each full scan (first-order mass spectrometry). The MS2 ActivationType was HCD, the Isolation window was 0.7 m / z, the resolution of the second-order mass spectrometry was 30,000 at 200 m / z, the AGC target was 1e5, the Maximum IT was 65 ms, the Fixed first mass was 110.0 m / z, the Normalized Collision Energy was 32 ev, the Minimum AGC target was 2.00e4, the Charge exclusion was 1, 6 - 8, >8, the Multiple charge states was one charge state only, the Peptide match was preferred, and the Exclude isotopes was on.
[0118] 1.4 Data preprocessing
[0119] The MS / MS data was searched using Maxquant (v1.6.15.0). The data type was DIA proteomics data based on MS / MS reporter ion quantification. For the MS / MS spectra used for quantification, it was required that the proportion of precursor ions in the MS spectra was greater than 75%. The database source was the Homo_sapiens_9606_proteome_gene in the Uniprot database (release: 2021-10-14, sequence: 20,437), and a common contaminant library was added to the database. Contaminant proteins were removed during data analysis; the digestion method was set to Trypsin / P; the maximum number of missed cleavage sites was set to 2; the mass error tolerances for precursor ions in the first search and main search were set to 20 ppm and 5 ppm respectively, and the mass error tolerance for MS / MS fragment ions was 20 ppm. The fixed modification was cysteine alkylation, and the variable modifications were methionine oxidation and protein N-terminal acetylation. The false discovery rate (FDR) for protein identification and PSM identification was set to 1%.
[0120] 1.5 Differential analysis
[0121] A combination of univariate analysis and multivariate statistical analysis was used to screen for differential proteins and transcripts. Univariate analysis mainly included the significance analysis (p-value or FDR value) and fold change of characteristic molecules in different groups, and multivariate statistical analysis mainly included receiver operating characteristic curve (ROC) analysis and Boruta feature selection based on the random forest algorithm. All statistical analyses were performed using R, and the specific R-related information is shown in Table 1.
[0122] Table 1. R and its related information
[0123]
[0124] The variable importance for the projection (VIP) was calculated to measure the influence intensity and interpretability of the expression patterns of each protein on the classification and discrimination of each group of samples. Further, the Wilcoxon rank-sum test was performed to obtain the corrected p-value (FDR). According to the conditions of FDR < 0.01 and Fold change > 2, 77 downregulated proteins and 71 upregulated proteins were screened out (see details in Figure 1 ).
[0125] To evaluate the role of each protein biomarker in the diagnosis and prediction of the risk of liver cancer recurrence, the ROC and Boruta analysis methods were used to evaluate each protein biomarker, and the results are shown in Figure 2The abscissa is the AUC obtained from the ROC analysis, the ordinate is the -log10(FDR) calculated by the Wilcoxon test, and the size of the points represents the VIP value obtained from the Boruta analysis. Further screening was carried out according to VIP > 3 and AUC > 0.6, and a total of 10 more significant candidate protein markers were found, as shown in Table 2 for details.
[0126] Table 2. Differential markers for the risk of liver cancer recurrence
[0127]
[0128] Among them, the smaller the FDR value and / or the larger the VIP value, to a certain extent, it indicates that the difference in the protein between the high-risk group and the low-risk group of liver cancer recurrence is more significant, and at the same time, it also indicates that the protein may have higher diagnostic value.
[0129] Example 2. Construction and validation of a prediction model for the risk of liver cancer recurrence
[0130] In this example, a recurrence risk prediction model was constructed based on the combination of the 10 protein markers screened in Example 1 for research.
[0131] Although a single biomarker can also distinguish the recurrence risk of liver cancer patients after surgery, generally speaking, the combination of multiple biomarkers has higher accuracy in discrimination or prediction. However, for a single biomarker with higher accuracy in predicting the recurrence risk of liver cancer, its role in the combination with one or more other biomarkers may not necessarily be greater. At the same time, it is not that the more the number of biomarkers, the higher the prediction accuracy (AUC value) of their combination. Therefore, a large number of verification experiments are still needed.
[0132] 2.1 Obtaining data
[0133] A total of 320 blood samples were collected from early liver cancer patients 2 weeks after surgical treatment. All the enrolled patients signed informed consent forms. Among them, 68 patients had recurrence within three years after surgery (38 cases had local recurrence in situ and 30 cases had metastasis). Thus, they were divided into two groups: 252 samples from non-recurrent patients were classified into the low-risk group of liver cancer recurrence, and 68 samples from patients with recurrence or metastasis were classified into the high-risk group of liver cancer recurrence. They were randomly divided into a test group and a validation group. The test group included 126 samples from the low-risk group of liver cancer recurrence and 34 samples from the high-risk group of liver cancer recurrence. The validation group included 126 samples from the low-risk group of liver cancer recurrence and 34 samples from the high-risk group of liver cancer recurrence.
[0134] The relative abundances of 10 proteins in the serum of 320 samples were obtained using the same steps as in Steps 1.2 and 1.3 of Example 1.
[0135] 2.2 Data statistical analysis
[0136] In the training group, a combined diagnostic model of multiple liver cancer recurrence risk markers was constructed using a combination of multiple machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) was estimated using the predicted probability value with a 95% confidence interval (CI) to evaluate the discrimination ability of the multivariate diagnostic model.
[0137] Using the test group, the Youden index (YI) was calculated to determine the cut-off value for predicting the probability of distinguishing between the high-risk group and low-risk group of liver cancer recurrence. In addition, the ROCs of the models formed by combining different markers were constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R 3.6.1, and a p-value less than 0.05 was considered statistically significant.
[0138] 2.3 Construction of the prediction model
[0139] The steps for constructing the liver cancer recurrence risk prediction model are as follows:
[0140] S101. From the 10 protein markers of the samples in the training group, a concentration matrix of randomly selected 2 to 10 markers was used as the original training data set.
[0141] S102. The generalized linear model (glmnet) algorithm was selected for constructing the prediction model and the grid search range during the hyperparameter optimization process of the algorithm. In this step, the grid search range for hyperparameter optimization of each algorithm was set as shown in Table 3.
[0142] Table 3 Parameter grid of the glmnet algorithm
[0143]
[0144] S103. According to the algorithm and the hyperparameter setting range set in step S102, one of the hyperparameter combination methods was selected as the parameters for constructing the prediction model.
[0145] S104. The original data set was split into K subsets according to the K-fold cross-validation mechanism. To ensure that the proportion of majority-class samples and minority-class samples in each subset is the same as that in the original data set, the stratified K-fold cross-validation mechanism was used for data splitting.
[0146] S105. According to the K training data subsets obtained by splitting in step S104, one of the subsets was selected as the validation set Ddev.
[0147] S106. Combine the unselected subsets of the training data in step S105 to form the training data pool Dtrainl.
[0148] S107. Based on the training data set Dtrain obtained in step S106, construct a prediction model based on the selected supervised classification algorithm and hyperparameters.
[0149] S108. According to the prediction model obtained in step S107, evaluate it on the validation set Ddev to obtain the AUC value, and store the current prognosis prediction model and the corresponding AUC value in the prediction model pool Pool. Step S108 is to evaluate according to the prediction model obtained in step S107 on the validation set determined in the current iteration, and store both the model and the evaluation results in the prediction model pool for future use in predicting model selection. The evaluation mentioned in this step can be the AUC value or other reasonable metrics for evaluating the model performance.
[0150] S109. Determine whether each subset has been used as the validation set. Step S109 is to determine whether the K subsets obtained in step S104 have all been used as the validation set for model training. If all subsets have been used as the validation set and completed training, execute step S110; if there are subsets that have not been used as the validation set, execute step S105. This step ensures that every sample in the original data set has been used as the validation set, improves the model stability, and prevents the model from overfitting to a certain subset.
[0151] S110. Take the average value of the AUCs of all the models in the obtained prediction model pool Pool as the final performance evaluation value of the model for this combination method. And store the model parameters and the final performance evaluation AUC value in the optimal model pool Poolbest.
[0152] S111. Determine whether prediction models have been constructed for all combinations of hyperparameters. Step S111 is to determine whether prediction models have been constructed for all the algorithms and the corresponding combinations of hyperparameters obtained in step S102. If all combinations have completed model construction, execute step S112; if there are combinations that have not completed model construction, execute step S103.
[0153] S112. After completing step S111, it is necessary to check whether all combinations of hyperparameters have been used to construct prediction models. If it is found that there are still combinations of hyperparameters for which models have not been constructed, then it is necessary to return to step S103 to continue constructing the remaining models. If models have been constructed for all combinations of hyperparameters, then step S113 can be continued to select the best model from the model pool.
[0154] S113. Select the model with the largest AUC value from the model set Poolbest obtained in step S112 as the final prediction model for liver cancer recurrence risk.
[0155] S114. Repeat all the above steps until all combinations of markers are modeled.
[0156] By executing the above prediction model construction steps, the optimal model constructed by all combinations of markers was obtained. In order to compare the performance of the models under these different marker combinations, the ROC method was used to evaluate the AUC values of these prediction models in the test group, and the results are shown in Table 4.
[0157] Table 4. Comparison of the areas under the ROC curves of the models constructed with different marker combinations in the test group
[0158]
[0159] Table 4 shows the ranking of markers in Table 2 according to the larger VIP value and the smaller FDR value. Starting from the top-ranked markers, the markers are selected in turn for combination. From 2MP to 5MP, the AUC value, accuracy and sensitivity of the detection increase with the increase of markers. However, when B3GNT2 is added to 5MP to form 6MP, the diagnostic performance of the constructed model does not continue to increase, and the AUC value decreases. It is speculated that the addition of B3GNT2 may also introduce noise; while the diagnostic performance of the 6MP-1 model formed by removing B3GNT2 and further adding the FBLN5 marker is improved. Further adding FBLN5 to 6MP improves the model performance; adding DCD to the 6MP-1 model also improves the model performance. Compared with other marker combinations, the model mean of 7MP-2 is higher, and further adding can no longer improve the diagnostic efficiency. In addition, the number of markers is smaller, so it is the most preferred model.
[0160] 2.4 Optimization of model parameters
[0161] For the optimal marker combination of IGFBP4+ORM1+RNASE1+DEFA3+ PGM5+B3GNT2+DCD in step 2.3, the models constructed under 9 different combinations of glmnet algorithm hyperparameters were analyzed based on this marker combination, and the model performance was evaluated by AUC value (10-fold cross-validation method was used to calculate AUC during the modeling process). The results are shown in Tables 5 and Figure 3 shown.
[0162] Table 5. AUC of the constructed model under different combinations of glmnet algorithm hyperparameters
[0163]
[0164] As can be seen from Table 5, when the hyperparameter combination of the glmnet algorithm is alpha = 0.55 and lambda = 0.0055, the AUC reaches the maximum value of 0.949.
[0165] The equation for constructing the model based on the optimal hyperparameter combination is:
[0166]
[0167] Among them, Y is the predicted score, i represents the i-th biomarker, m represents the number of biomarkers (m = 7), Xi represents the measured value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is the constant 2.5873933; the coefficients of the 7 biomarkers are:
[0168] Table 6. The coefficients of the 7 biomarkers are
[0169]
[0170] Y = 2.509IGFBP4 + 4.869ORM1 + 1.160RNASE1 + 9.782DEFA3 + 6.506PGM5 + 5.632FBLN5 + 3.739DCD + 2.5873933
[0171] Determination of the diagnostic threshold for the liver cancer recurrence risk prediction model:
[0172] ## Setting levels: control = case, case = control
[0173] ## Setting direction: controls < cases
[0174] The ROC curve is plotted with the predicted scores in the training group, and the optimal prediction cut-off value of 0.5180802 is set according to the Youden index value. The results are shown in Figure 4 。
[0175] When the predicted score of the prediction model ≤ 0.5180802, it is considered that the liver cancer recurrence risk of the tested person is low.
[0176] When the predicted score of the prediction model > 0.5180802, it is considered that the tested person has a high liver cancer recurrence risk.
[0177] From Figure 4 it can be seen that the AUC of the prediction model in the training group is 0.949, the sensitivity is 0.954, and the specificity is 0.955.
[0178] 2.6 Validation of the Hepatocellular Carcinoma Recurrence Risk Prediction Model
[0179] The optimal model constructed was validated in the test group, and the ROC curve was plotted as Figure 5 shown.
[0180] From Figure 5 it can be seen that the AUC of the model in the test group was 0.921, the sensitivity was 0.931, and the specificity was 0.932.
[0181] In summary, the hepatocellular carcinoma recurrence risk prediction model constructed with 7 protein markers has good prediction performance and accuracy, and has the best diagnostic efficacy.
[0182] Example 3 Construction and Validation of a Three-Class Diagnostic Model
[0183] In this example, an attempt was made to construct a three-class combined diagnostic model for distinguishing between the non-recurrence group, in-situ recurrence group, and distant metastasis group of hepatocellular carcinoma prognosis, which specifically included the following processes: (1) construction and screening of the optimal diagnostic model; (2) verification of the effect of the optimal diagnostic model. The specific screening process and results are as follows (in the present invention, the AUC value was used as the evaluation index for the binary classification model in Example 2; when constructing a three-class model, since multiple categories are involved, the AUC value is usually not applicable, and in this example, indicators such as sensitivity, specificity, accuracy, and consistency were used to measure the diagnostic efficacy of the model):
[0184] 3.1 Construction and Screening of the Three-Class Diagnostic Model
[0185] For the test cohort of 190 patients with early hepatocellular carcinoma, all enrolled patients signed informed consent forms. Among them, 38 patients had recurrence within two years after surgery (20 had in-situ local recurrence and 18 had metastasis). Thus, they were divided into two groups: 152 patient samples without recurrence were classified into the low-risk group of hepatocellular carcinoma recurrence, and 38 patients with recurrence or metastasis were classified into the high-risk group of hepatocellular carcinoma recurrence, and were randomly divided into a test group and a validation group. The training group included 95 patients in the low-risk group of hepatocellular carcinoma recurrence and 19 patients in the high-risk group of hepatocellular carcinoma recurrence (including 10 cases of in-situ recurrence and 9 cases of metastasis); the test group included 95 patients in the low-risk group of hepatocellular carcinoma recurrence and 19 patients in the high-risk group of hepatocellular carcinoma recurrence (including 10 cases of in-situ recurrence and 9 cases of metastasis).
[0186] Based on the biomarker combination form IGFBP4 + ORM1 + RNASE1 + DEFA3 + PGM5 + FBLN5 + DCD screened in Example 2, this example further constructs a three-classification detection model that can effectively distinguish the non-recurrence group (low-risk group), in-situ recurrence group (in-situ recurrence group), and distant metastasis group (metastasis group) of liver cancer prognosis. All enrolled patients signed informed consent forms. Among them, all liver cancer patients were pathologically and histologically diagnosed, and the inclusion criteria were: (a) no history of other malignancies; (b) no patients with other malignancies or autoimmune diseases.
[0187] LC-MS / MS data collection and detection were performed on the collected serum samples to obtain the concentrations of seven protein biomarkers respectively.
[0188] The Shapiro Wilk test was used to evaluate the normal distribution, and the non-parametric Wilcoxon test was used to analyze the differences in blood biomarker concentrations between the non-recurrence group (low-risk group), in-situ recurrence group (in-situ recurrence group), and distant metastasis group (metastasis group) of liver cancer prognosis. A three-classification combined diagnostic model of 6 biomarkers was constructed using a combination of machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) was estimated using the predicted probability value with a 95% confidence interval (CI) to evaluate the discrimination ability of the multivariate diagnostic model. Using the test group, the Youden index (YI) was calculated to determine the predicted probability cut-off value for distinguishing the non-recurrence group (low-risk group), in-situ recurrence group (in-situ recurrence group), and distant metastasis group (metastasis group) of liver cancer prognosis. In addition, the ROCs of models constructed with different biomarker combinations were constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV), and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R3.6.1, and a p-value less than 0.05 was considered statistically significant.
[0189] In this example, in order to construct the optimal three-classification combined diagnostic model, after comparing the models constructed by 7 algorithms including gradient boosting, naive Bayes, support vector machine, neural network, generalized linear, and discriminant analysis, the gradient boosting method was selected as the best supervised classification algorithm for constructing the prediction model. The grid search range for hyperparameter optimization of the gradient boosting method is shown in Table 7 below.
[0190] Table 7. Parameter grid search range of the gradient boosting method
[0191]
[0192] Through optimization and screening in terms of accuracy, consistency, sensitivity, specificity, etc., the optimal parameter combination mode was determined as follows: interaction.depth 2, n.trees 150, shrinkage 0.1, n.minobsinnode 10.
[0193] Completely different two batches of samples were used for the training group and the test group. In this embodiment, the screening of biomarkers and the construction of the model were only carried out in the training group; the samples of the test group were only used to verify the diagnostic efficacy of the model. The specific results are shown in Table 8.
[0194] Table 8. Performance evaluation table for constructing a model to distinguish three classifications by gradient boosting method
[0195]
[0196] It can be seen from Table 8 that the gradient boosting model constructed based on seven protein biomarkers of IGFBP4 + ORM1 + RNASE1 + DEFA3 + PGM5 + FBLN5 + DCD can be used to predict whether liver cancer patients will have no recurrence, in-situ recurrence or distant metastasis after surgery. It also shows that the protein biomarkers screened in the present invention can be used to distinguish the recurrence risk of liver cancer patients after surgical treatment, and can also be used to distinguish whether it is in-situ recurrence or distant metastasis when the recurrence risk is high (when there is both in-situ recurrence and metastasis, it is also classified into the metastasis group). However, the diagnostic efficacy for distinguishing the in-situ recurrence group and the metastasis group is not ideal.
[0197] 3.2 Combined performance of the three-classification combined diagnostic model
[0198] In order to further improve the diagnostic value of the three-classification diagnostic model (gradient boosting) constructed by biomarker combinations of different proteins, in this embodiment, based on the 9 protein biomarkers screened in Example 1, the diagnostic models constructed by biomarker combinations of different proteins were compared in terms of performance in the test group. The specific combination forms of different models are shown in Table 9.
[0199] Table 9. Combination forms of different diagnostic models
[0200]
[0201] Comparison results of performance indicators of different diagnostic models constructed using the 10 biomarkers selected in Example 1. The calculation methods for the minimum, first quartile, median, mean, third quartile, and maximum of accuracy and consistency are as follows: (1) Sort the values of accuracy or consistency from smallest to largest; (2) Minimum: The first value after sorting; (3) First quartile (Q1): Multiply the number of data by 0.25. If the result is an integer, take the average of the values at this position and the next position; if not, round up to get the position, and the value at this position is Q1; (4) Median: If the number of data is odd, the median is the middle value; if even, it is the average of the two middle values; (5) Mean: The sum of all values divided by the number of data; (6) Third quartile (Q3): Multiply the number of data by 0.75, and the processing method is the same as Q1; (7) Maximum: The last value after sorting. Among them, the minimum and maximum can reflect the extreme situations of the data and show the worst and best performances that the model may exhibit; the quartiles can help understand the distribution range and dispersion degree of the data; below Q1 represents a lower performance level, and above Q3 represents a higher performance level; the median can reflect the performance at the middle level; the mean comprehensively reflects the overall average performance. Combining the above statistical values, the overall situation, distribution characteristics, and stability of the model performance can be comprehensively understood, providing a strong basis for model selection and optimization.
[0202] Table 10. Performance comparison of diagnostic models constructed based on different protein combination biomarkers
[0203]
[0204] As can be seen from Table 10, for the three-class diagnostic model, the nine-item combined detection model composed of 9 biomarkers has the best performance. This also clearly shows that on the basis of the seven biomarkers IGFBP4 + ORM1 + RNASE1 + DEFA3 + PGM5 + FBLN5 + DCD, adding the biomarkers SERPINB1 and CORO1A can further improve the diagnostic efficacy for distinguishing between non-recurrence, in-situ recurrence, or distant metastasis of the prognosis of liver cancer surgery. Therefore, it is preferred to use the three-class gradient boosting model constructed with these 9 protein biomarkers (IGFBP4 + ORM1 + RNASE1 + DEFA3 + PGM5 + FBLN5 + DCD + SERPINB1 + CORO1A) as the best combined diagnostic model.
[0205] 3.2 Determination and verification of the diagnostic performance of the three-class combined diagnostic model
[0206] 1) Determination of the diagnostic performance of the three-class combined diagnostic model
[0207] To more accurately determine the diagnostic performance and threshold of the model constructed in this embodiment for different disease classifications, a multi-classification model based on the gradient boosting (gbm) algorithm was used to perform predictive analysis in the training group, and the predicted probability values for the three-class classification (no recurrence, in-situ recurrence, or distant metastasis after liver cancer surgery) were calculated. The classification with the largest predicted probability value is the final prediction result of the system.
[0208] Among them, the meanings and calculation methods of each index are as follows:
[0209] Calculation results: The accuracy of the three-class combined diagnosis model in the training group is 0.88, and the consistency is 0.89. The diagnostic sensitivity for the group without recurrence after liver cancer surgery is 92.8%, and the specificity is 93.4%; the diagnostic sensitivity for the in-situ recurrence group after liver cancer surgery is 86.5%, and the specificity is 88.6%; the diagnostic sensitivity for the distant metastasis group after liver cancer surgery is 84.7%, and the specificity is 87.1%.
[0210] It should be noted that the three-class combined diagnosis model constructed by gradient boosting is a model constructed by machine learning and cannot fit a specific equation formula like a generalized linear model.
[0211] 2) Verification of the three-class combined diagnosis model
[0212] Based on the model constructed for the test group, the predictive performance was verified in the training group, and the specific results are as follows:
[0213] The accuracy is 0.87, and the consistency is 0.88. The diagnostic sensitivity for the group without recurrence after liver cancer surgery is 93.5%, and the specificity is 92.3%; the diagnostic sensitivity for the in-situ recurrence group after liver cancer surgery is 86.6%, and the specificity is 86.5%; the diagnostic sensitivity for the distant metastasis group after liver cancer surgery is 86.1%, and the specificity is 85.4%.
[0214] In summary, the three-class combined diagnosis model constructed in this embodiment, which includes 9 protein markers, has good diagnostic value for the three classifications of the group without recurrence after liver cancer surgery, the in-situ recurrence group after liver cancer surgery, and the distant metastasis group after liver cancer surgery.
[0215] All patents and publications mentioned in the specification of the present invention indicate that these are publicly available technologies in the art and can be used in the present invention. All patents and publications cited herein are equally listed in the references, just as each publication is specifically individually referenced. The present invention described herein can be implemented in the absence of any one or more elements, one or more limitations, where such limitations are not specifically stated. For example, in each instance herein, the terms "comprising", "consisting essentially of", and "consisting of" can be replaced by either of the remaining two terms. The so-called "a" herein merely means "one", and does not exclude including only one, nor does it exclude including more than two. The terms and expressions used herein are for the purpose of description and are not limiting, and there is no intention to indicate that the terms and interpretations described herein exclude any equivalent features, but it is understood that any suitable changes or modifications can be made within the scope of the present invention and the claims. It is understood that the embodiments described in the present invention are some preferred embodiments and features, and any person of ordinary skill in the art can make some changes and variations based on the essence described in the present invention, and these changes and variations are also considered to be within the scope of the present invention and the scope limited by the independent claims and the dependent claims.
Claims
1. Use of a biomarker in the preparation of a reagent for predicting the risk of recurrence or metastasis of liver cancer, wherein the biomarker comprises one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2.
2. The use according to claim 1, characterized in that, The biomarker comprises IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD.
3. The use according to claim 2, characterized in that, The risk of recurrence or metastasis refers to recurrence or metastasis within three years after liver cancer treatment.
4. The use according to claim 2, characterized in that, The biological sample is selected from one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid; and / or, the reagent is used to detect the presence, relative abundance, or concentration of the biomarker.
5. A biomarker combination for predicting the risk of recurrence or metastasis of liver cancer, characterized in that, Comprises IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD.
6. A method for constructing a risk prediction model for liver cancer recurrence or metastasis for non-diagnostic purposes, characterized in that, Includes the following: 1) Construct a data set based on the detected amount of the biomarker in the biological sample, wherein the biomarker comprises IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD; 2) Divide the data set into a test set and a training set, and construct and train the liver cancer recurrence or metastasis risk prediction model through machine learning methods.
7. A device for predicting the risk of recurrence or metastasis of liver cancer, characterized in that, Includes a data acquisition unit and a calculation unit; The data acquisition unit is used to obtain the detected amount data of the biomarker in the biological sample of the subject, wherein the biomarker comprises IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD; The calculation unit is used to calculate and output the predicted score of the subject's risk of liver cancer recurrence or metastasis based on the detected amount data, and make a judgment according to the cut-off value.
8. An apparatus, comprising a processor and a memory, the memory being configured to store a computer program, characterized in that The processor is used to execute the computer program stored in the memory, so that the device executes the construction method as described in claim 6.
9. A system for predicting the risk of liver cancer recurrence, characterized in that, The system includes a data analysis module, and the data analysis module is used to analyze the detection value of the biomarker, and the biomarker comprises IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD.
10. The system according to claim 9, wherein The data analysis module uses the detection values of the biomarkers of known samples as the training set, and divides them into a liver cancer recurrence group and a non-recurrence group according to whether liver cancer recurs, analyzes the relationship between the detection values of the liver cancer recurrence group and the non-recurrence group, and constructs a model. The equation of the model is: Among them, Y is the predicted score, i represents the i-th biomarker, m represents the number of biomarkers (m = 7), Xi represents the measured value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant 2.5873933; the coefficients of the biomarkers are as follows: , When Y ≤ cut-off value, the subject is non-recurrent liver cancer; when Y > cut-off value, the subject is recurrent liver cancer. Specifically, the cut-off value is 0.5180802.
Citation Information
Patent Citations
Saliva proteomics- based liver cancer spleen-deficiency syndrome biomarker detection method
CN108152357A
Application of marker
CN114317756A
Compositions and methods for cancer diagnosis and prognosis
US20220136064A1