A biomarker for predicting the risk of liver cancer recurrence and its application

By screening and constructing a risk prediction model for liver cancer recurrence based on biomarkers such as IGFBP4, the problem of inaccurate prediction of liver cancer recurrence in the prior art is solved, and high-sensitivity diagnosis and personalized treatment support are achieved.

CN120254288BActive Publication Date: 2025-08-22HANGZHOU GUANGKE ANDE BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510725742.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-08-22
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The lack of high sensitivity proteomic biomarkers in the prior art are used to predict the risk of liver cancer recurrence, resulting in inaccurate prediction of the risk of recurrence after liver cancer and difficult to achieve personalized treatment adjustment.

Method used

Biomarkers such as IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A and B3GNT2 were screened through proteomics to construct a risk prediction model for liver cancer recurrence, and blood samples were analyzed using high-performance liquid chromatography-tandem mass spectrometry technology, combined with orthogonal partial least squares discriminant analysis and significance analysis, a diagnostic model was constructed.

Benefits of technology

It achieves accurate, non-invasive and effective prediction of the risk of recurrence or metastasis of liver cancer, improves the sensitivity and specificity of diagnosis, supports the adjustment of personalized treatment plans, and reduces the risk of misdiagnosis and missed diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120254288B_ABST
    Figure CN120254288B_ABST
Patent Text Reader

Abstract

The present invention discloses a biomarker for predicting the risk of liver cancer recurrence and its application. The use of a substance for detecting biomarkers in the preparation of a product for predicting the risk of liver cancer recurrence, wherein the biomarker includes one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A and B3GNT2. The present invention uses a proteomics method to analyze the significant differences in proteins in the blood of patients with liver cancer recurrence and non-recurrence, screens out biomarkers that can be used to predict the risk of liver cancer recurrence, and further constructs a liver cancer recurrence risk prediction model. The biomarker and liver cancer recurrence risk prediction model of the present invention can achieve simple, accurate and rapid diagnosis of liver cancer recurrence risk, improve the diagnostic level, and are suitable for large-scale promotion and application. Since the biomarker is derived from serum, sampling has the advantages of being convenient, immediate and non-invasive, so the present invention has important clinical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedicine technology, and in particular to a biomarker for predicting the risk of liver cancer recurrence and its application. Background Art

[0002] Liver cancer is the third leading cause of cancer-related deaths worldwide. Liver cancer is divided into two main types: hepatocellular carcinoma (HCC) and cholangiocarcinoma (ICC). HCC is the most common type, accounting for 90% of liver cancer cases. The development of liver cancer is associated with multiple factors, including chronic hepatitis virus infection (such as hepatitis B virus (HBV) and hepatitis C virus (HCV)), chronic liver disease (such as fatty liver disease or cirrhosis), alcoholism, and metabolic diseases (such as diabetes). Autoimmune hepatitis, hemochromatosis, and exposure to certain environmental toxins are also potential risk factors for liver cancer.

[0003] The development of liver cancer is associated with multiple factors, including chronic hepatitis virus infection, such as hepatitis B virus (HBV) and hepatitis C virus (HCV); chronic liver disease, such as fatty liver disease or cirrhosis; alcoholism; and metabolic diseases, such as diabetes. Furthermore, autoimmune hepatitis, hemochromatosis, and exposure to certain environmental toxins are potential risk factors for liver cancer. Liver resection is the primary means of achieving long-term survival benefits for patients with hepatocellular carcinoma (HCC). However, even in early-stage HCC, approximately 30% of patients experience recurrence within 2 to 5 years after surgery. Preventing HCC recurrence and metastasis and prolonging survival are key goals in the treatment of HCC and are currently a focus of clinical research. If patients who do not benefit from surgical treatment or whose disease progresses can be risk-predicted and their treatment plans adjusted promptly (such as adjuvant chemoradiotherapy, secondary surgical resection, targeted therapy, or immunotherapy), the patient's clinical benefit can be maximized. However, there is currently no consensus on the optimal tool for predicting the risk of HCC recurrence in the preoperative setting.

[0004] Proteomics is the study of protein composition, localization, changes, and interactions within cells, tissues, or organisms. This includes the study of protein expression patterns and proteome functional patterns. With the advancement of mass spectrometry, liquid chromatography coupled to mass spectrometry (LC-MS / MS) has become the primary tool in proteomics research. The development of proteomics is crucial for identifying disease diagnostic markers, screening drug targets, and conducting toxicology studies, leading to its widespread application in medical research.

[0005] While there are reports of biomarkers for assessing liver cancer recurrence risk, these are primarily based on genetic assessments, such as single-nucleotide polymorphisms (SNPs). This makes the testing process more complex and instrument-dependent, while also resulting in low diagnostic sensitivity. Clinically, there is a lack of biomarkers for liver cancer recurrence risk diagnosis, and the discovery of highly sensitive proteomic biomarkers for this purpose is of great significance. Summary of the Invention

[0006] In response to the problems existing in the prior art, the present invention provides a biomarker for predicting the risk of liver cancer recurrence and its application. By using proteomics methods, by analyzing proteins with significantly different abundance levels in the blood of two groups of people with liver cancer who have relapsed and those who have not relapsed after treatment, biomarkers that can be used to predict the risk of liver cancer recurrence or metastasis are screened out, and a liver cancer recurrence or metastasis risk prediction model is further constructed. This can achieve accurate, non-invasive and efficient prediction of the risk of liver cancer recurrence or metastasis to meet clinical needs.

[0007] A first aspect of the present invention provides the use of a substance for detecting biomarkers in the preparation of a reagent for predicting the risk of recurrence or metastasis of liver cancer, wherein the biomarkers include one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A and B3GNT2.

[0008] The present invention provides biomarkers obtained by proteomics that can accurately predict the risk of liver cancer recurrence or metastasis, which helps doctors judge the severity of the patient's condition and whether the treatment plan needs to be adjusted, thereby providing patients with more accurate treatment recommendations and truly benefiting liver cancer patients.

[0009] The present invention utilizes proteomics methods to collect plasma samples from patients who experience short-term recurrence (e.g., within 1 to 5 years) of liver cancer treatment, as well as patients who do not experience short-term recurrence. The different samples are analyzed using high-performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS). Based on orthogonal partial least squares discriminant analysis and significance analysis, proteins with significant differences between patients with and without liver cancer recurrence are first screened. Ten differentially expressed proteins with a clear correlation with the risk of liver cancer recurrence are identified. These 10 proteins can be used to distinguish the risk of recurrence or metastasis of liver cancer after treatment, demonstrating a certain diagnostic efficacy.

[0010] Among them, the IGFBP4 is a protein or amino acid sequence with a UniProt database number of P22692; ORM1 is a protein or amino acid sequence with a UniProt database number of P02763; RNASE1 is a protein or amino acid sequence with a UniProt database number of P07998; DEFA3 is a protein or amino acid sequence with a UniProt database number of P59666; PGM5 is a protein or amino acid sequence with a UniProt database number of Q15124; FBLN5 is a protein or amino acid sequence with a UniProt database number of Q9UBX5; DCD is a protein or amino acid sequence with a UniProt database number of P81605; SERPINB1 is a protein or amino acid sequence with a UniProt database number of P25407; CORO1A is a protein or amino acid sequence with a UniProt database number of P31146; and B3GNT2 is a protein or amino acid sequence with a UniProt database number of Q9NY97.

[0011] The present inventors also surprisingly discovered that some of the protein markers obtained through proteomic screening are already known markers for other cancers. For example, DEFA3, previously reported for the prediction of respiratory infections, was found in this screening to be also useful for predicting the risk of liver cancer recurrence or metastasis. This demonstrates that the protein markers for many different tumors are not completely separate or unrelated; in fact, they exhibit numerous cross-relationships or influences. Many protein markers can be used for both early-stage cancer prediction and prognostic diagnosis, and even for the prediction and diagnosis of multiple cancers at different stages. Therefore, the field of proteomics holds many new capabilities yet to be explored, and the market prospects are enormous.

[0012] In some embodiments, the biomarker can be one marker or a combination of several markers, such as a combination of two markers, a combination of three markers, a combination of four markers, a combination of five markers, a combination of six markers, a combination of seven markers, a combination of eight markers, a combination of nine markers, or a combination of ten markers. In some specific embodiments, the biomarker comprises a combination of two or more markers, such as a combination of three or more markers. When the combination of markers is used to construct a prediction model for the risk of liver cancer recurrence or metastasis, its AUC value is 0.801-0.949, the sensitivity is 82.9%-95.4%, and the specificity is 83.3%-95.5%.

[0013] Furthermore, the biomarkers include IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5 and DCD.

[0014] Furthermore, to further investigate the diagnostic efficacy of the model for liver cancer recurrence risk, it was necessary to combine different differentially expressed proteins to construct a diagnostic model. These proteins were ranked according to their importance, and different numbers of differentially expressed proteins with the highest rankings were selected and combined to obtain 10 protein markers. After verification, it was found that the model based on seven protein markers had good risk prediction capabilities in liver cancer recurrence risk diagnosis. Data from liver cancer recurrence risk samples showed that using only these seven biomarkers to predict liver cancer recurrence risk achieved an AUC value of 0.949, demonstrating good diagnostic performance.

[0015] In some embodiments, the product is selected from one or more of a reagent, a kit, a chip, a probe, or a membrane strip. The product is a product that detects the biomarkers described above, and includes, for example, sample pretreatment reagents, antigens, or antibodies, and other biological reagents and kits suitable for detecting the biomarkers. The product can also be developed into standardized reagents or kits, chips, probes, or membrane strips suitable for detecting the biomarkers.

[0016] Furthermore, the risk of recurrence or metastasis refers to recurrence or metastasis within three years after treatment of liver cancer; the recurrence or metastasis of liver cancer includes recurrence of liver cancer in situ or adjacent areas, and metastasis of liver cancer.

[0017] It can be understood that the short term of "recurrence within three years after liver cancer treatment" here is not an absolute and unchanging time node, but a time node currently summarized based on clinical experience. With the passage of time or the improvement of other treatment methods, this time node or the length of time may change. For example, liver cancer may recur after treatment within one year, within 360 days, or within half a year, 180 days, or within two years, within two and a half years, within three and a half years, etc.

[0018] In some embodiments, the liver cancer recurrence or metastasis refers to the recurrence of liver cancer in situ or in adjacent areas within three years after radical tumor resection, or the occurrence of liver cancer metastasis, such as lymph node metastasis, bone metastasis, etc.

[0019] In some embodiments, the product is used to predict patients' risk of liver cancer recurrence or metastasis.

[0020] In some embodiments, the product is used to detect the amount of a biomarker in a biological sample.

[0021] In some specific embodiments, the biological sample is selected from one or more of saliva, blood, urine, plasma, serum and cerebrospinal fluid.

[0022] In some specific embodiments, the detected amount includes the presence or relative abundance or concentration of a biomarker.

[0023] The present invention uses blood screening to identify biomarkers that predict the risk of liver cancer recurrence. These biomarkers show significant differences in the blood of people with and without liver cancer recurrence. By collecting blood samples, these biomarkers can be detected in the individual's blood to predict or assist in diagnosing whether the individual has liver cancer recurrence or not, or to predict or assist in diagnosing whether the individual has not recurred after liver cancer surgery, has recurred in situ, or has distant recurrence.

[0024] Furthermore, the detection method generally includes a radiometric method, an immunological method, a fluorescence method, a flow cytometry, a latex turbidimetry method, a biochemical method, an enzymatic method, a hybridization method, a gas chromatography-mass spectrometry method, a liquid chromatography-mass spectrometry method, a chromatography method, a chemiluminescence method, a magnetoelectric method or a photoelectric conversion method.

[0025] The presence or absence, or level, of a biomarker is a relative concept. For example, when comparing a recurrent liver cancer group with a non-recurrent liver cancer group, the levels of these specific biomarkers are compared relative to the baseline of the recurrent or non-recurrent group. It's possible that certain biomarkers may be higher in recurrent liver cancer than in the non-recurrent group. This increase can be statistically significant, such as a significant or highly significant increase. Therefore, when assessing the presence of a single biomarker, if the probability of a particular risk is elevated, the biomarker's level may change. This change can be a relative increase or decrease. This relative increase or decrease is considered significant, and can even be highly significant. Therefore, regardless of the testing method, a predetermined cutoff value can be used as a standard. A value above this cutoff value is considered a change in the level, and such a result can be used for prognostic or diagnostic purposes.

[0026] Therefore, in some aspects, the biomarkers described herein can be obtained by detecting the marker content in a sample using any known method, such as liquid chromatography, gas chromatography, mass spectrometry, LC-MS, gas chromatography-mass spectrometry (GC-MS), chromatography-mass spectrometry (CC-MS), liquid chromatography-tandem mass spectrometry (LC-MS-MS), nuclear magnetic resonance spectroscopy (NMR), immunochromatographic test strips, immunoreaction chips, capillary electrophoresis, infrared spectroscopy, and the like. As long as the protein marker content in a sample can be detected, it can be used to predict or diagnose the probability of a disease. It is understood that the detection here involves testing an individual sample and then comparing it with a pre-set standard. The comparison results are used to determine or predict the disease state. For example, it can be used to predict the probability of liver cancer recurrence risk. This prediction or diagnosis is based on whether the disease will recur within a certain period of time. Of course, such detection can be continuous, and the progression of the disease can be inferred based on the changes in the content of certain substances.

[0027] In some embodiments, the relative abundance is the peak area of ​​the biomarker in the detection spectrum obtained by high-performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of ​​a biomarker measured in a control sample is 300 and the average peak area measured in a liver cancer recurrence sample is 1800, then the abundance of the biomarker in the sample is considered to be 6 times that in the control sample.

[0028] A second aspect of the present invention provides the use of proteins as biomarkers in the preparation of products for predicting the risk of liver cancer recurrence or metastasis. The proteins include one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2. The IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2 genes can be used to assist in determining the risk of liver cancer recurrence or metastasis, assess drug efficacy, and more. The inventors have discovered that IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2 are closely associated with the risk of liver cancer recurrence or metastasis.

[0029] A third aspect of the present invention provides a product for predicting the risk of recurrence or metastasis of liver cancer, comprising a substance for detecting biomarkers, wherein the biomarkers include one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A and B3GNT2.

[0030] A fourth aspect of the present invention provides a biomarker combination for predicting the risk of recurrence or metastasis of liver cancer, comprising IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5 and DCD.

[0031] A fifth aspect of the present invention provides a method for constructing a liver cancer recurrence or metastasis risk prediction model for purposes other than disease diagnosis, comprising the following steps:

[0032] 1) constructing a sample dataset based on the detected amount of biomarkers in the biological sample, wherein the biomarkers include one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2;

[0033] 2) dividing the data set into a test set and a training set, and constructing and training the liver cancer recurrence or metastasis risk prediction model through machine learning methods.

[0034] In some embodiments, the biomarkers are IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD.

[0035] In some embodiments, the machine learning method is selected from at least one of a gradient boosting algorithm, a random forest algorithm, a support vector machine algorithm, a decision tree algorithm, a K-nearest neighbor algorithm, a logistic regression algorithm, and a neural network algorithm. For example, the gradient boosting algorithm can be selected. Unlike generalized linear regression, which outputs a model formula and cutoff value, the gradient boosting algorithm performs all calculations directly through machine learning. Test values ​​can be directly input into the software system to directly obtain prediction results.

[0036] In some embodiments, multiple machine learning methods are used to construct a liver cancer recurrence or metastasis risk prediction model, and it is preliminarily confirmed that the concentration changes of any one of the screened new biomarkers alone can be used to distinguish between people with liver cancer recurrence and those without recurrence, indicating that these biomarkers have extremely high diagnostic value.

[0037] The present invention discovered that by detecting the detection amount of biomarkers in biological samples and then inputting the detection amount into the formula of a liver cancer recurrence or metastasis risk prediction model, a prediction score Y of the prediction model is obtained, and the prediction score Y is compared with the threshold (cut off) defined by the Youden Index. If the prediction score Y is greater than the threshold (cut off), it is judged that the liver cancer has recurred; if the prediction score y is less than or equal to the threshold (cut off), it is judged that the liver cancer has not recurred.

[0038] In some embodiments, the equation of the constructed liver cancer recurrence risk prediction model is:

[0039]

[0040] Where Y is the prediction score, i represents the i-th biomarker, m represents the number of biomarkers (m=7), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant of 2.5873933; the coefficients of the seven biomarkers are:

[0041]

[0042] When Y ≤ the cutoff value, the patient has no liver cancer recurrence; when Y > the cutoff value, the patient has liver cancer recurrence. Specifically, the cutoff value is 0.5180802.

[0043] In some embodiments, the method further includes step 3) testing the liver cancer recurrence or metastasis risk prediction model using a test set. After training, the trained prediction model is validated using the test set, and the effectiveness of the prediction model is evaluated using AUC, specificity, and sensitivity as evaluation indicators.

[0044] The Receiver Operating Characteristic Curve (ROC Curve) is a curve drawn based on a series of different binary classification methods (cut-off values), with the true positive rate (sensitivity) as the vertical axis and the false positive rate (1-specificity) as the horizontal axis. The area under the receiver operating characteristic curve (AUC) is defined as the area under the ROC curve. The AUC value is often used to evaluate the diagnostic efficacy of a prediction model. The larger the AUC value, the better the diagnostic efficacy of the corresponding prediction model; conversely, the lower the AUC value, the worse the diagnostic efficacy of the corresponding prediction model.

[0045] The prediction model constructed based on the combination of 7MP in the present invention can distinguish between recurrence and non-recurrence of liver cancer after surgery. In the training group, the AUC was 0.949, the sensitivity was 0.954, and the specificity was 0.955; in the test group, the AUC was 0.921, the sensitivity was 0.931, and the specificity was 0.932.

[0046] A sixth aspect of the present invention provides a device for predicting the risk of liver cancer recurrence or metastasis, comprising a data acquisition unit and a calculation unit;

[0047] The data acquisition unit is used to obtain the detection amount data of the biomarkers as described above in the biological sample of the subject, wherein the biomarkers include one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A and B3GNT2;

[0048] The calculation unit is used to calculate and output the predicted score of the subject's risk of liver cancer recurrence or metastasis based on the detection data, and make a judgment based on the threshold

[0049] In some embodiments, the biomarkers include IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD.

[0050] In some embodiments, the prediction device further includes a data storage unit and a data output unit; the data storage unit is used to store the detection amount (or detection value) of the biomarker; the data input interface is used to input the detection value of the biomarker, and the data output unit is used to output the prediction result.

[0051] Furthermore, the detection value is the presence or absence, relative abundance or concentration value of each biomarker.

[0052] The seventh aspect of the present invention provides a device comprising a processor and a memory, wherein the memory is used to store a computer program, and is characterized in that the processor is used to execute the computer program stored in the memory so that the device performs the prediction method as described above or the construction method as described above.

[0053] An eighth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is processed, the prediction method as described above or the construction method as described above is executed.

[0054] In some embodiments, the computer-readable storage medium includes various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0055] In some implementations, being processed refers to being executed by one or more processors.

[0056] A ninth aspect of the present invention provides a method for predicting the risk of liver cancer recurrence or metastasis, comprising the following steps:

[0057] S1. Obtaining detection amount data of biomarkers in a biological sample of a subject, wherein the biomarkers include one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2;

[0058] S2. Apply the liver cancer recurrence risk prediction model obtained by the construction method described above to process the detection data to output a liver cancer recurrence or metastasis risk prediction result.

[0059] In some embodiments, the biomarkers include IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD.

[0060] On the other hand, the present invention provides a system for predicting the risk of recurrence of liver cancer, the system comprising a data analysis module, the data analysis module being used to analyze the detection values ​​of markers, the markers comprising IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5 and DCD.

[0061] Furthermore, the data analysis module uses the detection values ​​of markers of known samples as a training set, divides the samples into a liver cancer recurrence group and a liver cancer non-recurrence group according to whether the liver cancer recurs, analyzes the relationship between the detection values ​​of the liver cancer recurrence group and the liver cancer non-recurrence group, and constructs a model.

[0062] In some embodiments, a combined diagnostic model for predicting the risk of liver cancer recurrence is constructed by combining multiple machine learning methods, and it is preliminarily confirmed that the concentration changes of any one of the screened new biomarkers alone can be used to distinguish between people at high risk and low risk of liver cancer recurrence, indicating that these biomarkers have extremely high diagnostic value.

[0063] In some embodiments, the equation of the constructed model is:

[0064] Where Y is the prediction score, i represents the i-th biomarker, m represents the number of biomarkers (m=7), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant of 2.5873933; the coefficients of the seven biomarkers are:

[0065]

[0066] When Y ≤ the cutoff value, the patient has no liver cancer recurrence; when Y > the cutoff value, the patient has liver cancer recurrence. Specifically, the cutoff value is 0.5180802.

[0067] Furthermore, the system also includes a data storage module, a data input interface and a data output interface; the data storage module is used to store the detection values ​​of biomarkers; the data input interface is used to input the detection values ​​of biomarkers, and the data output interface is used to output the prediction results.

[0068] In another aspect, the present invention provides a use of a marker for preparing a reagent for predicting whether a liver cancer patient will not relapse after surgery, relapse in situ, or have distant metastasis, wherein the marker comprises any one or more of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, CORO1A, and B3GNT2.

[0069] Furthermore, the markers include IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1 and CORO1A.

[0070] In another aspect, the present invention provides a kit for predicting whether liver cancer patients will not relapse after surgery, will relapse in situ, or will metastasize to distant sites. The kit comprises a detection reagent for the biomarker for the purpose described above.

[0071] In another aspect, the present invention provides a biomarker combination for predicting whether a liver cancer patient will not relapse after surgery, relapse in situ, or have distant metastasis, the combination comprising IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, and CORO1A.

[0072] In another aspect, the present invention provides a system for predicting whether liver cancer will not recur after surgery, will recur in situ, or will metastasize to distant sites. The system includes a data analysis module, which is used to analyze the detection values ​​of markers, including IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, and CORO1A.

[0073] Furthermore, the distant metastasis also includes simultaneous in situ recurrence and distant metastasis. As long as distant metastasis occurs, it will be classified into the distant metastasis group.

[0074] Furthermore, the data analysis module uses the detection values ​​of markers of known samples as a training set, and divides liver cancer patients into a non-recurrence group, an in situ recurrence group, and a distant metastasis group according to their post-operative conditions. The relationship between the detection values ​​of the non-recurrence group, the in situ recurrence group, and the distant metastasis group is analyzed to construct a model.

[0075] Furthermore, the model is constructed based on the gradient boosting algorithm.

[0076] Unlike generalized linear regression, the gradient boosting algorithm cannot output model formulas and cutoff values. All calculations are done directly by machine learning. The test values ​​can be directly input into the software system to obtain the prediction results.

[0077] Furthermore, the system also includes a data storage module, a data input interface and a data output interface; the data storage module is used to store the detection values ​​of biomarkers; the data input interface is used to input the detection values ​​of biomarkers, and the data output interface is used to output the prediction results.

[0078] The beneficial effects of the present invention are:

[0079] 1. The present invention has screened 10 new biomarkers that can predict the risk of liver cancer recurrence or metastasis, and developed a new protein marker combination that can effectively assess and diagnose the risk of liver cancer recurrence or metastasis, effectively distinguish between liver cancer recurrence and non-recurrence, and effectively distinguish between in situ recurrence and distant metastasis of liver cancer. This not only improves the diagnostic accuracy of liver cancer recurrence and metastasis, but also provides an important biomarker foundation for personalized treatment and prognosis monitoring of liver cancer patients. In addition, the discovery and application of these markers are expected to improve the treatment effect and quality of life of liver cancer patients, and provide a scientific basis for the early diagnosis and treatment strategy selection of liver cancer. Compared with traditional detection methods, it reduces the risk of misdiagnosis and missed diagnosis, providing strong support for early detection and intervention of the disease.

[0080] 2. The combined differential diagnosis model of seven biomarkers constructed in the present invention is convenient and fast, and the test results are highly consistent with the clinical gold standard test results. At the same time, it significantly reduces the cost of predicting the risk of liver cancer recurrence and has good application prospects.

[0081] 3. Based on the biomarkers screened for predicting the risk of liver cancer recurrence, a three-classification model was further constructed that can simultaneously distinguish between the non-recurrence group, the in situ recurrence group, and the distant metastasis group, providing a more effective and accurate predictive diagnostic model.

[0082] (1) Diagnosis or testing

[0083] The diagnosis or detection here refers to the detection or analysis of biomarkers in a sample, or the content of a target biomarker, such as the absolute content or relative content, and then the presence or amount of the target marker is used to indicate whether the individual providing the sample may have or suffer from a certain disease, or the possibility of having a certain disease. The meanings of diagnosis and detection here are interchangeable. The result of such a test or diagnosis cannot be directly used as a direct result of the disease, but is an intermediate result. If a direct result is obtained, other auxiliary means such as pathology or anatomy are required to confirm that the patient has a certain disease. For example, the present invention provides a variety of new biomarkers related to the risk of recurrence or metastasis of liver cancer. Changes in the content of these markers are directly correlated with whether the patient belongs to the group at risk of liver cancer recurrence or metastasis.

[0084] (2) Association between markers, biomarkers, or differentially expressed proteins and the risk of HCC recurrence or metastasis

[0085] The terms "marker," "biomarker," and "differential protein" have the same meaning in this invention. Association here refers to a direct correlation between the presence or change in the level of a biomarker in a sample and a specific disease. For example, a relative increase or decrease in the level indicates a higher likelihood of the individual having the disease compared to a healthy population.

[0086] The simultaneous presence of multiple markers in a sample, or the relative changes in their levels, indicate a higher likelihood of the individual having the disease compared to healthy individuals. This means that among marker types, some are strongly associated with disease, while others are weakly associated, or even unrelated to a particular disease. One or more markers with strong correlations can be used as diagnostic markers, while markers with weaker correlations can be combined with stronger markers to diagnose a disease, increasing the accuracy of test results.

[0087] The numerous serum biomarkers discovered by the present invention can be used to distinguish between recurrent and non-recurrent liver cancer, as well as between non-recurrent, in situ recurrence, and distant metastasis. These markers can be used alone as single markers for direct detection or diagnosis. The selection of such markers indicates that the relative change in their levels is strongly correlated with the risk of liver cancer recurrence or metastasis. Of course, it is understood that one or more markers with strong correlations with the risk of liver cancer recurrence or metastasis can be selected for simultaneous detection. It is generally understood that, in some approaches, selecting biomarkers with strong correlations for detection or diagnosis can achieve a certain standard of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90%, or 95%. This indicates that these markers can achieve intermediate values ​​for diagnosing a certain disease, but does not directly confirm the presence of a specific disease.

[0088] Of course, the differential protein with the larger ROC value can also be selected as a diagnostic marker. The so-called strength is generally calculated and confirmed by some algorithms, such as the contribution rate or weight analysis of the marker to the risk prediction of liver cancer recurrence or metastasis. Such calculation methods can be significance analysis (p value or FDR value) and fold change (Fold change). Multivariate statistical analysis mainly includes principal component analysis (PCA), partial least squares discriminant analysis (PLS-DA) and orthogonal partial least squares discriminant analysis (OPLS-DA), and of course other methods, such as ROC analysis, etc. Of course, other model prediction methods are also possible. When specifically selecting biomarkers, the differential proteins disclosed in the present invention can be selected, or other existing well-known marker combinations can be selected or combined to make predictions through model methods.

[0089] (3) Definition of disease terms

[0090] Liver cancer is a malignant tumor that originates in liver cells and typically manifests as abnormal cell growth and proliferation in liver tissue. These cells may invade surrounding tissues or spread to other parts of the body through the bloodstream and lymphatic system. The development of liver cancer may involve genetics, chronic hepatitis, cirrhosis, smoking, obesity, diabetes, and other lifestyle factors.

[0091] Early-stage liver cancer refers to the initial stages of tumor development, before significant vascular invasion, distant metastasis, or extensive invasion of surrounding tissues. This type of liver cancer is typically small (≤5 cm in diameter), less aggressive, and offers a better prognosis with surgery or other local treatments.

[0092] Liver cancer recurrence refers to the reappearance or spread of liver cancer to other parts of the body after treatment. Recurrence can occur in or near the original site of the liver, known as a local recurrence (in situ recurrence). Local recurrence means the cancer has not spread widely but has reappeared in the liver or its immediately adjacent tissues. Alternatively, liver cancer can spread to more distant sites in the body, known as distant recurrence or metastasis. Even for early-stage liver cancer, approximately 30% of patients experience recurrence within 2-3 years after surgery, with the risk of recurrence significantly decreasing after 5 years. Liver cancer recurrence is closely related to tumor stage, treatment compliance, and postoperative management. The prognosis (expected survival after treatment) for liver cancer is poor because it is a highly aggressive malignancy with few early symptoms and is often discovered in the late stages. This results in relatively poor treatment outcomes and a poor prognosis.

[0093] Metastasis of liver cancer: Distant metastasis means that the cancer has spread to organs or tissues outside the liver, such as the lungs and bones, through the blood or lymphatic system. This type of recurrence usually indicates that the cancer has progressed to an advanced stage, and the complexity and difficulty of treatment have increased accordingly.

[0094] Therefore, early detection and intervention of liver cancer recurrence or metastasis are crucial to improving treatment outcomes and increasing survival rates. Biomarkers have accurate and specific predictive efficacy for liver cancer recurrence or metastasis, which is beneficial for improving treatment outcomes and prognosis.

[0095] (4) The gold standard for liver cancer diagnosis is pathological diagnosis (i.e., pathological confirmation of a biopsy or surgical resection specimen), which involves observing the presence of cancer cells under a microscope and combining immunohistochemistry or molecular testing to clarify the nature of the tumor. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] Figure 1 This is a volcano plot of the differential analysis of protein markers in the high-risk group and low-risk group for liver cancer recurrence in Example 1.

[0097] Figure 2 Graphs showing the ROC and OPLS-DA analysis results for the high-risk and low-risk groups for liver cancer recurrence in Example 1.

[0098] Figure 3 This is a schematic diagram of the hyperparameter optimization results of the glmnet algorithm in Example 2;

[0099] Figure 4 2 is the ROC curve of the liver cancer recurrence risk prediction model in Example 2 in the training group.

[0100] Figure 5 This is the ROC curve of the liver cancer recurrence risk prediction model in Example 2 in the test group. DETAILED DESCRIPTION

[0101] The present invention will be described in further detail below in conjunction with the accompanying drawings and Examples. It should be noted that the following examples are intended to facilitate understanding of the present invention and do not serve to limit the present invention in any way. The reagents used in this example are all known products and were obtained by purchasing commercially available products.

[0102] Example 1 Screening biomarkers for liver cancer recurrence risk using proteomics

[0103] Plasma samples were collected from patients undergoing radical surgery for liver cancer. Low-abundance proteins were enriched using immunoaffinity chromatography to remove high-abundance proteins. Protein abundance in the samples was detected using a tandem high-performance liquid chromatography-mass spectrometry device. The differences in abundance between patients with and without liver cancer recurrence were analyzed to identify differentially expressed proteins. The specific steps are as follows:

[0104] 1.1 Sample Collection

[0105] Blood samples were collected from 100 patients with early-stage liver cancer (≤5 cm in diameter) 2 weeks after surgical resection. All liver cancer patients had biopsies confirmed by pathology. Approximately 2 ml of peripheral blood was collected from the subjects, placed in a vacuum tube containing EDTA anticoagulant, mixed, and centrifuged at 120 g for 10 minutes at room temperature. The supernatant was removed and repeated twice. The supernatant was then collected and centrifuged at 360 g for 20 minutes. Platelet samples were then collected in centrifuge tubes and stored at -80°C for later use. The 100 patients included 72 males and 28 females, with an average age of 40 years (range, 25-69 years). All enrolled patients signed informed consent. All patients with liver cancer were confirmed by pathological histology. Inclusion criteria included: (a) no history of other malignancies; (b) no patients with other malignancies or autoimmune diseases.

[0106] For the next three years, patients were followed up every six months, with regular AFP and imaging examinations. Any recurrence or metastasis was confirmed by pathological histology. A total of 32 patients with early-stage liver cancer relapsed within three years of surgery, including 25 with local in situ recurrence and 7 with distant metastasis (including 5 lymph node metastases, 1 lung metastasis, and 1 bone metastasis). Samples were then removed from the -80°C storage and divided into two groups: 68 patient samples without recurrence were classified as a low-risk group for colorectal cancer recurrence, and 32 samples with recurrence or metastasis were classified as a high-risk group for colorectal cancer recurrence. These samples were used for proteomic biomarker screening.

[0107] 1.2 Sample processing and enzymatic hydrolysis

[0108] 1) First, plasma samples were centrifuged for 15 minutes (15,000 g), and the supernatant was filtered and then subjected to immunoaffinity chromatography to remove 14 highly abundant proteins.

[0109] 2) The low-abundance fraction was concentrated to 350 μL using a 3 kDa cut-off concentration tube in a centrifuge (4000 g, 1 hour).

[0110] 3) The concentrate was recovered and buffer exchanged using a 7 kDa cutoff desalting column in a centrifuge (1000 g, 2 minutes) using AEX-A (20 mM Tris, 4 M urea, 3% isopropanol, pH 8.0).

[0111] 4) Determine the protein concentration in the sample using the biuret assay (BCA) with AEX-A as a blank. According to the sample grouping, add 25 μL of TCEP to the sample and incubate at 37°C for 30 minutes to reduce the protein. Then, add the corresponding TMT 16-plex reagent and incubate at room temperature in the dark for 1 hour to perform the TMT labeling reaction. Then, perform buffer exchange of the sample using Zeba columns with AEX-A. After mixing the TMT 16-plex-labeled sample, add 2 mL of AEX-A to the mixed sample to a final volume of 5.5 mL.

[0112] 5) Filter the sample through a 0.22 µm filter and separate the TMT 16-plex-labeled sample using a 2D-HPLC system. Freeze-dry the collected fractions and add Trypsin-Lysin C enzyme mix. Incubate the sample at 37°C for 5 hours to digest the sample. Terminate the digestion reaction by adding 5 µL of 10% TFA.

[0113] 6) A total of 60 2D-HPLC fractions after enzymatic digestion were used for nanoLC-MS / MS analysis.

[0114] 1.3 LC-MS / MS Data Acquisition

[0115] Each sample obtained in step 1.2 was separated using the nanoliter flow rate liquid chromatography system Easy nLC-1200 and detected online using a high-resolution mass spectrometer Q Exactive HF-X. The details are as follows:

[0116] 1) Separation: Mobile phase A consisted of 0.1% formic acid in water, and mobile phase B consisted of 0.1% formic acid in acetonitrile (80% acetonitrile, 20% water). The chromatographic column consisted of an enrichment column and an analytical column equilibrated with 100% mobile phase A. The sample was loaded via an autosampler onto the enrichment column (100 μm ID × 4 cmL, C18, 3 μm, 100 Å) and separated on the analytical column (75 μm ID × 25 cmL, C18, 3 μm, 100 Å) at a flow rate of 300 nL / min.

[0117] 2) After chromatographic separation, samples were analyzed by mass spectrometry using a Q Exactive HF-X mass spectrometer. Positive ion detection was used, with a parent ion scan range of 350–1800 m / z, a primary MS resolution of 120,000 at 200 m / z, an AGC (Automatic Gain Control) target of 3e6, a maximum IT of 50 ms, and a dynamic exclusion time of 40 s. The mass-to-charge ratios of peptides and peptide fragments were acquired using data-dependent acquisition (DDA): 20 secondary MS / MS spectra (MS2 scans) were collected after each full scan (MS2 scan). The MS2 Activation Type is HCD, the Isolation window is 0.7 m / z, the secondary mass spectrometry resolution is 30,000 at 200 m / z, the AGC target is 1e5, the Maximum IT is 65 ms, the Fixed first mass is 110.0 m / z, the Normalized Collision Energy is 32 eV, the Minimum AGC target is 2.00e4, the Charge exclusion is 1, 6-8, >8, the Multiple charge states is one charge state only, the Peptidematch is preferred, and the Exclude isotopes is on.

[0118] 1.4 Data Preprocessing

[0119] Secondary mass spectrometry data were retrieved using Maxquant (v1.6.15.0). The data type is DIA proteomics data based on secondary reporter ion quantification. The secondary spectrum used for quantification requires that the parent ion accounts for more than 75% in the primary spectrum. The database comes from the Homo_sapiens_9606_proteome_gene (release: 2021-10-14, sequence: 20,437) of the Uniprot database, and a common contamination library is added to the database. Contaminating proteins are deleted during data analysis; the enzyme cleavage method is set to Trypsin / P; the number of missed cleavage sites is set to 2; the parent ion mass error tolerance of the First search and Main search is set to 20 ppm and 5 ppm, respectively, and the mass error tolerance of the secondary fragment ion is 20 ppm. The fixed modification is cysteine ​​alkylation, and the variable modification is methionine oxidation and acetylation of the protein N-terminus. The FDR for protein identification and PSM identification is set to 1%.

[0120] 1.5 Gap Analysis

[0121] We screened for differentially expressed proteins and transcripts using a combination of univariate and multivariate statistical analyses. Univariate analysis primarily included significance analysis (p-value or FDR value) and fold change analysis of signature molecules across different groups. Multivariate statistical analysis primarily included receiver operating characteristic (ROC) curve analysis and Boruta signature screening based on the random forest algorithm. All statistical analyses were performed using R. Specific R information is provided in Table 1.

[0122] Table 1. R and related information

[0123]

[0124] The variable importance for the projection (VIP) was calculated to measure the influence and explanatory power of the expression pattern of each protein on the classification and discrimination of each group of samples. The Wilcoxon rank sum test was further performed to obtain the corrected p value (FDR). According to the conditions of FDR < 0.01 and Fold change > 2, 77 down-regulated proteins and 71 up-regulated proteins were screened (see Figure 1 ).

[0125] In order to evaluate the role of each protein marker in the diagnosis and prediction of liver cancer recurrence risk, ROC and Boruta analysis methods were used to evaluate each protein marker. The results are shown in Figure 2The horizontal axis represents the AUC obtained from ROC analysis, and the vertical axis represents the -log10 (FDR) calculated by the Wilcoxon test. The size of the dots represents the VIP value obtained from Boruta analysis. Further screening based on VIP > 3 and AUC > 0.6 identified 10 more significant candidate protein biomarkers, as detailed in Table 2.

[0126] Table 2. Differential markers of liver cancer recurrence risk

[0127]

[0128] Among them, the smaller the FDR value and / or the larger the VIP value, to a certain extent, it indicates that the difference between the protein at high risk of liver cancer recurrence and low risk of liver cancer recurrence is more significant, and it also indicates that the protein may have a higher diagnostic value.

[0129] Example 2: Construction and validation of a liver cancer recurrence risk prediction model

[0130] In this example, the 10 protein markers screened and obtained in Example 1 were combined to construct a recurrence risk prediction model for research.

[0131] While a single biomarker can differentiate the risk of recurrence in patients with liver cancer after surgery, combining multiple biomarkers generally offers greater accuracy in differentiation or prediction. However, a single biomarker that accurately predicts liver cancer recurrence risk may not necessarily be more effective when combined with one or more other biomarkers. Furthermore, a greater number of biomarkers does not necessarily equate to a higher predictive accuracy (AUC) for the combination. Therefore, extensive validation experiments are still needed.

[0132] 2.1 Get Data

[0133] A total of 320 blood samples were collected from patients with early-stage liver cancer 2 weeks after surgical treatment. All enrolled patients signed informed consent forms, of which 68 patients relapsed within three years after surgery (38 were local recurrences in situ and 30 were metastatic). They were divided into two groups: 252 patient samples without recurrence were classified as a low-risk group for liver cancer recurrence, and 68 samples with recurrence or metastasis were classified as a high-risk group for liver cancer recurrence. They were randomly divided into a test group and a validation group. The test group included 126 patients in the low-risk group for liver cancer recurrence and 34 patients in the high-risk group for liver cancer recurrence, and the validation group included 126 patients in the low-risk group for liver cancer recurrence and 34 patients in the high-risk group for liver cancer recurrence.

[0134] The same steps as steps 1.2 and 1.3 in Example 1 were used to obtain the relative abundance of 10 proteins in serum for 320 samples.

[0135] 2.2 Data Statistical Analysis

[0136] In the training group, a combined diagnostic model for multiple HCC recurrence risk markers was constructed using a combination of machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) was estimated using predicted probability values ​​with 95% confidence intervals (CI) to assess the discriminative ability of the multivariate diagnostic model.

[0137] Using the test group, the Youden index (YI) was calculated to determine the cut-off value for distinguishing the high-risk group for liver cancer recurrence from the low-risk group for liver cancer recurrence. In addition, the ROC of the model formed by the combination of different markers was constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive score (PPV), negative predictive score (NPV) and standard deviation (SD) were calculated to describe the experimental results of the study group. Statistical analysis was performed using R3.6.1, and p values ​​less than 0.05 were considered statistically significant.

[0138] 2.3 Construction of prediction model

[0139] The specific steps for constructing a liver cancer recurrence risk prediction model are as follows:

[0140] S101. Randomly select 2 to 10 marker concentration matrices from the 10 protein markers in the samples of the training group as the original training data set.

[0141] S102: Select the generalized linear model (glmnet) algorithm for building the prediction model and the grid search range during the algorithm's hyperparameter optimization process. In this step, the grid search range for the model's hyperparameter optimization is set for each algorithm as shown in Table 3.

[0142] Table 3 Parameter grid of the glmnet algorithm

[0143]

[0144] S103. According to the algorithm and hyperparameter setting range set in step S102, one of the hyperparameter combinations is selected as a parameter for constructing a prediction model.

[0145] S104: Split the original dataset into K subsets using a K-fold cross validation mechanism. To ensure that the ratio of majority class samples to minority class samples in each subset is the same as in the original dataset, a Stratified K-Folds cross validation mechanism is used for data segmentation.

[0146] S105 , selecting one of the K training data subsets obtained by segmentation in step S104 as a validation set Ddev.

[0147] S106: Combine the training data subsets not selected in step S105 to form a training data pool Dtrain1.

[0148] S107 . According to the training data set Dtrain obtained in step S106 , a prediction model is constructed based on the selected supervised classification algorithm and hyperparameters.

[0149] S108: Evaluate the prediction model obtained in step S107 on the validation set Ddev to obtain an AUC value, and store the current prognosis prediction model and the corresponding AUC value in the prediction model pool Pool. Step S108 involves evaluating the prediction model obtained in step S107 on the validation set determined in the current iteration, and storing both the model and the evaluation results in the prediction model pool for future selection and use by the base prediction model. The evaluation mentioned in this step can be an AUC value or other reasonable metric for evaluating model performance.

[0150] S109: Determine whether all subsets have been used as validation sets. Step S109 determines whether all K subsets obtained in step S104 have been used as validation sets and trained on the model. If all subsets have been used as validation sets and training has been completed, proceed to step S110; if any subsets have not been used as validation sets, proceed to step S105. This step ensures that every sample in the original dataset has been used as a validation set, improving model stability and preventing the model from overfitting to a particular subset.

[0151] S110: The average AUC value of all models in the prediction model pool Pool is used as the final performance evaluation value of the combined model. The model parameters and the final performance evaluation AUC value are stored in the optimal model pool Poolbest.

[0152] S111: Determine whether all hyperparameter combinations have been used to construct prediction models. Step S111 determines whether all algorithms and corresponding hyperparameter combinations obtained in step S102 have been used to construct prediction models. If all combinations have been used to construct models, step S112 is executed; if any combination has not been used to construct models, step S103 is executed.

[0153] S112. After completing step S111, it is necessary to check whether all hyperparameter combinations have been used to build prediction models. If it is found that there are still hyperparameter combinations for which models have not been built, it is necessary to return to step S103 and continue to build the remaining models. If all hyperparameter combinations have been used to build models, it is possible to continue to step S113 and select the best model from the model pool.

[0154] S113. From the model set Poolbest obtained in step S112, select the model with the largest AUC value as the final prediction model for liver cancer recurrence risk.

[0155] S114. Repeat all the above steps until all combinations of markers are modeled.

[0156] By executing the above prediction model construction steps, the optimal models for all combinations of markers were obtained. To compare the performance of the models under these different marker combinations, the ROC method was used to evaluate the AUC values ​​of these prediction models in the test group. The results are shown in Table 4.

[0157] Table 4. Comparison of the area under the ROC curve of the models constructed by different marker combinations in the test group

[0158]

[0159] Table 4 shows the ranking of the markers in Table 2 according to their highest VIP values ​​and lowest FDR values. Starting with the highest-ranked markers, marker combinations were selected sequentially, from 2MP to 5MP. The AUC, accuracy, and sensitivity of the model increased with the addition of more markers. However, when B3GNT2 was added to 5MP to form 6MP, the diagnostic performance of the constructed model did not improve further, and the AUC value actually decreased. This is speculated to be due to the introduction of noise by B3GNT2. However, the diagnostic performance of the 6MP-1 model, which removed B3GNT2 and added FBLN5, improved. Further addition of FBLN5 to 6MP further improved model performance, as did the addition of DCD to the 6MP-1 model. 7MP-2 had higher model means than the other marker combinations, and further additions failed to improve diagnostic performance. Furthermore, the model had fewer markers, making it the most preferred model.

[0160] 2.4 Optimization of model parameters

[0161] For the optimal marker combination of IGFBP4+ORM1+RNASE1+DEFA3+PGM5+B3GNT2+DCD in step 2.3, the models constructed under 9 different combinations of glmnet algorithm hyperparameters were analyzed based on this marker combination, and the model performance was evaluated by the AUC value (AUC was calculated using a 10-fold cross-validation method during the modeling process). The results are shown in Tables 5 and Figure 3 shown.

[0162] Table 5. AUC of the constructed model under different combinations of glmnet algorithm hyperparameters

[0163]

[0164] As can be seen from Table 5, when the hyperparameter combination of the glmnet algorithm is alpha = 0.55, lambda = 0.0055, the AUC reaches a maximum value of 0.949.

[0165] The equation for building a model based on the optimal hyperparameter combination is:

[0166]

[0167] Where Y is the prediction score, i represents the i-th biomarker, m represents the number of biomarkers (m=7), Xi represents the detection value of the i-th biomarker (μg / mL), Ki represents the coefficient of the i-th biomarker, and b is a constant of 2.5873933; the coefficients of the seven biomarkers are:

[0168] Table 6. The coefficients of the 7 biomarkers are

[0169]

[0170] Y=2.509IGFBP4+4.869ORM1+1.160RNASE1+9.782DEFA3+6.506PGM5+5.632FBLN5+3.739DCD+2.5873933

[0171] Determination of diagnostic threshold for liver cancer recurrence risk prediction model:

[0172] ## Setting levels: control = case, case = control

[0173] ## Setting direction: controls < cases

[0174] The ROC curve was drawn with the prediction scores in the training group, and the optimal prediction cutoff value was set to 0.5180802 according to the Youden index value. The results are shown in Figure 4 .

[0175] When the prediction score of the prediction model is ≤0.5180802, it is considered that the risk of liver cancer recurrence of the subject is low.

[0176] When the prediction score of the prediction model is greater than 0.5180802, the subject is considered to have a high risk of liver cancer recurrence.

[0177] from Figure 4 It can be seen that the prediction model has an AUC of 0.949, a sensitivity of 0.954, and a specificity of 0.955 in the training group.

[0178] 2.6 Validation of the HCC Recurrence Risk Prediction Model

[0179] The optimal model constructed was verified in the test group and the ROC curve was drawn as follows Figure 5 shown.

[0180] from Figure 5 It can be seen that the AUC of the model in the test group is 0.921, the sensitivity is 0.931, and the specificity is 0.932.

[0181] In summary, the liver cancer recurrence risk prediction model constructed using 7 protein markers has good predictive performance and accuracy and the best diagnostic efficacy.

[0182] Example 3 Construction and verification of a three-category diagnostic model

[0183] This example attempts to construct a three-category combined diagnostic model for distinguishing between the non-recurrence group, the in situ recurrence group, and the distant metastasis group. The model specifically includes the following steps: (1) construction and screening of the optimal diagnostic model; and (2) validation of the effectiveness of the optimal diagnostic model. The specific screening process and results are as follows (in the present invention, the two-category model in Example 2 uses the AUC value as the evaluation indicator; when a three-category model is constructed, since multiple categories are involved, the AUC value is generally not applicable. In this example, indicators such as sensitivity, specificity, accuracy, and consistency are used to measure the diagnostic efficacy of the model):

[0184] 3.1 Construction and screening of the three-category diagnostic model

[0185] For the test cohort of 190 patients with early-stage liver cancer, all enrolled patients signed informed consent. Among them, 38 patients relapsed within two years after surgery (20 of them had local recurrence in situ and 18 had metastasis). They were divided into two groups: 152 patient samples without recurrence were classified as the low-risk group for liver cancer recurrence, and 38 samples with recurrence or metastasis were classified as the high-risk group for liver cancer recurrence. They were randomly divided into a test group and a validation group. The training group included 95 patients in the low-risk group for liver cancer recurrence and 19 patients in the high-risk group for liver cancer recurrence (including 10 patients with in situ recurrence and 9 patients with metastasis); the test group included 95 patients in the low-risk group for liver cancer recurrence and 19 patients in the high-risk group for liver cancer recurrence (including 10 patients with in situ recurrence and 9 patients with metastasis).

[0186] This example further constructs a three-category detection model based on the marker combination IGFBP4+ORM1+RNASE1+DEFA3+PGM5+FBLN5+DCD screened in Example 2, capable of effectively distinguishing between a group with no HCC recurrence (low-risk group), a group with in situ HCC recurrence (in situ recurrence group), and a group with distant HCC metastasis (metastasis group). All enrolled patients provided informed consent. All HCC patients were confirmed by histopathology. Inclusion criteria included: (a) no history of other malignancies; and (b) no concurrent malignancies or autoimmune diseases.

[0187] The collected serum samples were subjected to LC-MS / MS data acquisition and detection to obtain the concentrations of seven protein markers.

[0188] The Shapiro-Wilk test was used to assess normal distribution, and the nonparametric Wilcoxon test was used to analyze differences in blood marker concentrations between the HCC prognosis non-recurrence group (low-risk group), HCC prognosis in situ recurrence group (in situ recurrence group), and HCC prognosis distant metastasis group (metastasis group). A three-class combined diagnostic model for the six markers was constructed using a combination of machine learning methods. The area under the receiver operator characteristic (ROC) curve (AUC) was estimated using the predicted probability values ​​with 95% confidence intervals (CI) to assess the discriminatory ability of the multivariate diagnostic model. Using the test set, the Youden index (YI) was calculated to determine the predicted probability cutoff value for distinguishing the HCC prognosis non-recurrence group (low-risk group), HCC prognosis in situ recurrence group (in situ recurrence group), and HCC prognosis distant metastasis group (metastasis group). In addition, the ROC values ​​of the models constructed with different marker combinations were constructed and compared. Standard descriptive statistics such as frequency, mean, median, positive predictive value (PPV), negative predictive value (NPV) and standard deviation (SD) were calculated to describe the experimental results of the study population. Statistical analysis was performed using R3.6.1, and a p value of less than 0.05 was considered statistically significant.

[0189] In this embodiment, in order to construct the optimal three-class joint diagnosis model, after comparing the model construction of seven algorithms including gradient boosting, naive Bayes, support vector machine, neural network, generalized linear, and discriminant analysis, the gradient boosting method was selected as the best supervised classification algorithm for constructing the prediction model. The grid search range for hyperparameter optimization of the gradient boosting method model is shown in Table 7 below.

[0190] Table 7. Parameter grid search range of gradient boosting method

[0191]

[0192] Through optimization screening in terms of accuracy, consistency, sensitivity, specificity, etc., the optimal parameter combination mode was determined to be: interaction.depth 2, n.trees 150, shrinkage 0.1, n.minobsinnode 10.

[0193] The training and test groups used two completely different batches of samples. This example only screened markers and constructed models from the training group; the test group samples were only used to verify the diagnostic efficacy of the model. The specific results are shown in Table 8.

[0194] Table 8. Performance evaluation table of the gradient boosting method to build a model to distinguish three categories

[0195]

[0196] As shown in Table 8, the gradient boosting model constructed based on the seven protein markers (IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD) can be used to predict whether HCC patients will experience no recurrence, in situ recurrence, or distant metastasis after surgery. This also demonstrates that the protein markers screened by this invention can be used to distinguish the risk of recurrence in HCC patients after surgical treatment and, when the risk of recurrence is high, to distinguish between in situ recurrence and distant metastasis (patients with both in situ recurrence and metastasis are also included in the metastasis group). However, the diagnostic efficacy for distinguishing between the in situ recurrence and metastasis groups is still not ideal.

[0197] 3.2 Combined performance of the three-class joint diagnosis model

[0198] To further enhance the diagnostic value of the three-category diagnostic model (gradient boosting) constructed using different protein combinations of biomarkers, this example compared the performance of diagnostic models constructed using different protein combinations of biomarkers in a test group based on the nine protein markers screened in Example 1. The specific combinations of the different models are shown in Table 9.

[0199] Table 9. Combinations of different diagnostic models

[0200]

[0201] The performance index comparison results of different diagnostic models constructed by the 10 biomarkers screened in Example 1 are shown. The calculation method of the minimum value, first quartile, median, mean, third quartile and maximum value of accuracy and consistency is as follows: (1) Sort the accuracy or consistency values ​​from small to large; (2) Minimum value: the first value after sorting; (3) First quartile (Q1): multiply the number of data by 0.25. If the result is an integer, take the average of the values ​​at this position and the next position; if it is not an integer, round up to get the position, and the value at this position is Q1; (4) Median: if the number of data is odd, the median is the middle value; if it is even, it is the average of the two middle values; (5) Mean: the sum of all values ​​divided by the number of data; (6) Third quartile (Q3): multiply the number of data by 0.75 and process it in the same way as Q1; (7) Maximum value: the last value after sorting. The minimum and maximum values ​​reflect data extremes, demonstrating the worst and best possible model performance. Quartiles help understand the data's distribution and dispersion. Q1 and below indicate lower performance, while Q3 and above indicate higher performance. The median reflects intermediate performance, and the mean comprehensively reflects the overall average performance. By combining these statistical values, we can gain a comprehensive understanding of the overall performance, distribution characteristics, and stability of the model, providing a strong basis for model selection and optimization.

[0202] Table 10. Performance comparison of diagnostic models based on different protein combination biomarkers

[0203]

[0204] Table 10 shows that for the three-category diagnostic model, the nine-item combined test model, consisting of nine markers, performed the best. This also clearly demonstrates that the addition of SERPINB1 and CORO1A to the seven-item combination of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, and DCD further improves the diagnostic efficacy of differentiating HCC surgical prognosis from non-recurrence, in situ recurrence, and distant metastasis. Therefore, the three-category gradient boosting model constructed using these nine protein markers (IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, and CORO1A) was selected as the optimal combined diagnostic model.

[0205] 3.2 Diagnostic Performance Measurement and Verification of the Three-Classification Joint Diagnosis Model

[0206] 1) Diagnostic performance measurement of the three-classification combined diagnosis model

[0207] In order to more accurately determine the diagnostic performance and threshold of the model constructed in this embodiment for different disease classifications, a multi-classification model of the gradient boosting (GBM) algorithm was used to perform predictive analysis in the training group, and the predicted results were calculated as the predicted probability values ​​of three categories (whether the prognosis of liver cancer surgery is no recurrence, in situ recurrence, or distant metastasis). The category with the largest predicted probability value is the final prediction result of the system.

[0208] The meaning and calculation method of each indicator are as follows:

[0209] Results: The three-class combined diagnostic model achieved an accuracy of 0.88 and a consistency of 0.89 in the training group. Its sensitivity for the group without HCC recurrence after surgery was 92.8% and its specificity was 93.4%. Its sensitivity for the group with in situ HCC recurrence after surgery was 86.5% and its specificity was 88.6%. Its sensitivity for the group with distant metastasis after surgery was 84.7% and its specificity was 87.1%.

[0210] It should be noted that the three-classification joint diagnosis model constructed by gradient boosting is a model constructed by machine learning and cannot fit a specific equation formula like a generalized linear model.

[0211] 2) Validation of the three-classification combined diagnosis model

[0212] The prediction performance of the model built based on the test group was verified in the training group. The specific results are as follows:

[0213] The accuracy was 0.87 and the consistency was 0.88. The sensitivity for diagnosing non-recurrence after HCC surgery was 93.5% and the specificity was 92.3%. The sensitivity for diagnosing in situ recurrence after HCC surgery was 86.6% and the specificity was 86.5%. The sensitivity for diagnosing distant metastasis after HCC surgery was 86.1% and the specificity was 85.4%.

[0214] In summary, the three-category combined diagnostic model containing 9 protein markers constructed in this example has good diagnostic value for the three categories of liver cancer non-recurrence group after surgery, liver cancer in situ recurrence group after surgery, and liver cancer distant metastasis group after surgery.

[0215] All patents and publications cited in this specification are intended to indicate that they are state of the art and that the present invention may be used. All patents and publications cited herein are incorporated by reference in their entirety, as if each publication were specifically incorporated by reference. The invention described herein may be practiced in the absence of any element or elements, limitation or limitations, unless otherwise specified. For example, in each instance, the terms "comprising," "consisting essentially of," and "consisting of" may be replaced with either of the other two terms. The term "a" or "an" herein simply means "one" and does not exclude the inclusion of only one or more. The terms and expressions used herein are intended to be descriptive, not limiting, and are not intended to exclude any equivalent features. However, it is understood that any suitable changes or modifications may be made within the scope of the present invention and the appended claims. It is understood that the embodiments described herein are preferred embodiments and features, and that modifications and variations can be made by persons of ordinary skill in the art based on the spirit of the present invention. Such modifications and variations are considered to be within the scope of the present invention and the scope of the independent and appended claims.

Claims

1. Use of a reagent for detecting a marker for preparing a reagent for predicting whether a liver cancer patient will not relapse, will have in situ recurrence, or will have distant metastasis within two years after surgery, characterized in that: The markers consist of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1 and CORO1A.

2. A kit for predicting whether a liver cancer patient will not relapse, relapse in situ, or metastasize within two years after surgery, characterized in that: The kit comprises a detection reagent for the biomarker for use as claimed in claim 1.

3. A biomarker combination for predicting whether a liver cancer patient will not relapse, relapse in situ, or metastasize within two years after surgery, characterized in that: The panel consists of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1, and CORO1A.

4. A system for predicting whether liver cancer will not recur, recur in situ, or metastasize within two years after surgery, characterized in that: The system includes a data analysis module for analyzing the detection values ​​of markers, wherein the markers consist of IGFBP4, ORM1, RNASE1, DEFA3, PGM5, FBLN5, DCD, SERPINB1 and CORO1A.

5. The system according to claim 4, wherein: The data analysis module uses the detection values ​​of markers of known samples as a training set, and divides liver cancer patients into a non-recurrence group, an in situ recurrence group, and a distant metastasis group according to their post-operative conditions. The relationship between the detection values ​​of the non-recurrence group, the in situ recurrence group, and the distant metastasis group is analyzed to construct a model; the model is constructed based on a gradient boosting algorithm; the system also includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values ​​of biomarkers; the data input interface is used to input the detection values ​​of biomarkers, and the data output interface is used to output prediction results.

Citation Information

Patent Citations

  • Saliva proteomics- based liver cancer spleen-deficiency syndrome biomarker detection method

    CN108152357A

  • Application of marker

    CN114317756A