A prediction model for identifying hepatocellular carcinoma and a construction method and application thereof

By constructing the GT2.0 predictive model for hepatocellular carcinoma based on 12 specific oligosaccharide chains and clinical indicators, the shortcomings of traditional serological markers in the diagnosis of early hepatocellular carcinoma have been addressed. This model achieves highly sensitive non-invasive diagnosis, especially in very early stages and in patients with negative markers, thereby improving the detection rate of early hepatocellular carcinoma.

CN122135792AActive Publication Date: 2026-06-02JIANGSU XIANSIDA BIOTECH CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU XIANSIDA BIOTECH CO LTD
Filing Date
2026-05-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In the current technology, traditional serological markers such as AFP and DCP have insufficient sensitivity in the early diagnosis of hepatocellular carcinoma, especially in very early-stage hepatocellular carcinoma and AFP/DCP negative patients, which limits their diagnostic efficacy and cannot meet clinical needs.

Method used

By acquiring data on the abundance, gender, and age of 12 specific oligosaccharide chains from the training sample set, a prediction model was constructed using the Bootstrap-integrated support vector machine algorithm. Combined with machine learning techniques, a high-precision hepatocellular carcinoma prediction model, GT2.0, was formed, enabling non-invasive and highly sensitive diagnosis of hepatocellular carcinoma.

Benefits of technology

The GT2.0 model achieved a diagnostic sensitivity of 74.77% in very early-stage hepatocellular carcinoma, significantly improving the detection rate of early-stage patients. In particular, its sensitivity exceeded 83% in AFP/DCP-negative patients, filling the diagnostic gap of existing biomarkers and providing a comprehensive serological diagnostic system with the potential for cross-center and cross-population applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135792A_ABST
    Figure CN122135792A_ABST
Patent Text Reader

Abstract

This invention discloses a predictive model for identifying hepatocellular carcinoma, its construction method, and its application. The construction method includes training the predictive model using a machine learning algorithm with gender, age, and 12 specific oligosaccharide chains as independent variables and a clinical diagnostic label as the dependent variable. The 12 specific oligosaccharide chains are NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb, NA4, NA4Fb, and NA4F2b. The predictive model achieves an accuracy of 89.43% on an independent validation set, a sensitivity of 74.77% for very early-stage hepatocellular carcinoma, and a sensitivity of 87.36% for patients with both AFP and DCP double-negative results. This invention also provides a predictive system and kit incorporating this model, offering an efficient and reliable solution for the early non-invasive diagnosis of hepatocellular carcinoma.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of molecular biomedicine and bioinformatics technology, specifically relating to a predictive model for identifying hepatocellular carcinoma, its construction method, and its application. Background Technology

[0002] Hepatocellular carcinoma (HCC) is the most common histological type of primary liver cancer, accounting for approximately 75%–85% of all primary liver cancers. HCC is characterized by its insidious onset, rapid progression, and poor prognosis, ranking among the leading causes of death among all malignant tumors. Currently, for early-stage HCC, radical treatments such as surgical resection, liver transplantation, or local ablation can significantly improve the 5-year survival rate (reaching over 60%–70%). However, because early-stage HCC often presents with no obvious clinical symptoms, over 60% of patients are already in the middle or late stages at the time of initial diagnosis, missing the optimal treatment window. Therefore, identifying highly sensitive and specific biomarkers is of significant clinical value for achieving early diagnosis of HCC and improving patient prognosis.

[0003] Currently, the widely used serological biomarker alpha-fetoprotein (AFP) has significant limitations: on the one hand, approximately 30%–40% of patients with early-stage hepatocellular carcinoma have normal serum AFP levels (false negatives); on the other hand, non-tumor conditions such as chronic hepatitis, cirrhosis, and pregnancy can also lead to elevated AFP levels (false positives). While other biomarkers such as abnormal prothrombin (DCP) and phosphatidylinositol proteoglycan-3 (GPC3) can partially compensate for the shortcomings of AFP, their diagnostic capabilities in the context of early-stage tumors and cirrhosis still fall short of clinical needs. Therefore, finding novel, highly specific biomarkers has become an urgent technical problem to be solved in the field of hepatocellular carcinoma diagnosis.

[0004] Protein glycosylation is one of the most common post-translational modifications of proteins, and most proteins in the blood undergo glycosylation. During the development and progression of hepatocellular carcinoma (HCC), the expression profiles of glycosyltransferases and glycosidases within tumor cells undergo significant changes, leading to specific alterations in the oligosaccharide chain structure of glycoproteins. These changes in oligosaccharide chain structure are not random but highly pathologically relevant. In the progression from hepatitis to cirrhosis to HCC, specific oligosaccharide chain structures (such as core fucosylation, bifurcated, and branched oligosaccharide chains) exhibit regular dynamic evolution. This makes oligosaccharide chain biomarkers theoretically more applicable and with higher pathological specificity than single protein biomarkers. Based on this, researchers in this field have attempted to use oligosaccharide chain biomarkers for the diagnosis, prognosis, and treatment response prediction of HCC.

[0005] For example, Chinese patent application CN120072335A discloses a G-GAAD model for early diagnosis of liver cancer. This model uses oligosaccharide chain detection value (GT), sex, age, AFP, and DCP as input variables and is constructed using logistic regression. Although this model improves the accuracy of liver cancer diagnosis to some extent, it still relies on the two traditional serological markers, AFP and DCP. As mentioned earlier, AFP and DCP have high false negative rates in early-stage liver cancer and some special types of liver cancer, which means that the diagnostic efficacy of the G-GAAD model in AFP / DCP-negative patients is still limited. In addition, this model uses a comprehensive oligosaccharide chain detection value (GT) rather than detailed information on multiple specific oligosaccharide chain structures, failing to fully utilize the differentiated changes of different oligosaccharide chains in the development and progression of hepatocellular carcinoma.

[0006] Early-stage hepatocellular carcinoma (HCC) refers to stage 0 or A HCC in the Barcelona Clinic Liver Cancer (BCLC) staging system. According to the "Guidelines for the Diagnosis and Treatment of Primary Liver Cancer," BCLC stage 0 is defined as a single tumor ≤2 cm in diameter, without vascular invasion, and with Child-Pugh A liver function; BCLC stage A is defined as a single tumor or up to three tumors with a maximum diameter ≤3 cm, without vascular invasion, and with Child-Pugh AB liver function. Early-stage HCC usually presents with no obvious clinical symptoms, representing the optimal time for radical treatment. However, the diagnostic sensitivity of existing serological markers at this stage is generally less than 40%, highlighting the urgent need for more sensitive detection methods in clinical practice.

[0007] In summary, there is an urgent need in the field for a novel, highly sensitive, and highly specific non-invasive diagnostic method for hepatocellular carcinoma that does not rely on traditional serological markers (such as AFP and DCP), particularly a solution that can improve the detection rate of very early-stage hepatocellular carcinoma and patients with double-negative AFP / DCP. This invention addresses this technological need. Summary of the Invention

[0008] To address the shortcomings of existing technologies, this invention provides a predictive model for identifying hepatocellular carcinoma, its construction method, and its application. By discovering clinical indicators combined with specific oligosaccharide chains and constructing a high-precision predictive model for the first time, it achieves rapid, non-invasive, and highly accurate identification of early-stage hepatocellular carcinoma, providing a crucial basis for early clinical treatment.

[0009] This invention is achieved through the following technical solution:

[0010] A method for constructing a predictive model for identifying hepatocellular carcinoma includes the following steps:

[0011] Step 1) Obtain the training sample set, which includes biological samples, clinical diagnostic labels, and the gender and age of multiple patients with known clinical diagnoses of hepatitis, cirrhosis, and hepatocellular carcinoma.

[0012] Step 2) Detect the abundance of 12 specific oligosaccharide chains in each sample of the training sample set to obtain oligosaccharide chain abundance data; the 12 specific oligosaccharide chains are: NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb, NA4, NA4Fb, and NA4F2b;

[0013] Step 3) Using the oligosaccharide chain abundance data, gender, and age as independent variables, and the clinical diagnostic label as the dependent variable, a machine learning algorithm is used to train the model to obtain the prediction model.

[0014] Preferably, the biological sample in step 1) is blood, serum or plasma from venous blood or peripheral blood of patients with hepatitis, cirrhosis or hepatocellular carcinoma.

[0015] Preferably, the machine learning algorithm in step 3) is the support vector machine algorithm integrated with Bootstrap.

[0016] A predictive model for identifying hepatocellular carcinoma, the predictive model being obtained by the construction method described above.

[0017] Preferably, the prediction model uses a predicted value ≥ 5.0 as the threshold for judging hepatocellular carcinoma.

[0018] Preferably, the hepatocellular carcinoma is a very early stage hepatocellular carcinoma, specifically hepatocellular carcinoma in Barcelona clinical stage 0 or A.

[0019] A predictive system for identifying hepatocellular carcinoma, comprising:

[0020] The data acquisition module is used to acquire the abundance information of 12 specific oligosaccharide chains in the sample to be tested, as well as the gender and age of the individual to be tested; the 12 specific oligosaccharide chains are composed of the following oligosaccharide chains: NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb, NA4, NA4Fb, and NA4F2b;

[0021] The prediction module stores the aforementioned prediction model.

[0022] The result output module is used to output the hepatocellular carcinoma prediction results calculated by the prediction model based on the oligosaccharide chain abundance information, gender, and age.

[0023] A computer-readable storage medium having the above-described prediction model stored thereon.

[0024] The application of reagents for detecting 12 specific oligosaccharide chains in the preparation of products for assisting in the identification of hepatocellular carcinoma, wherein the 12 specific oligosaccharide chains consist of the following oligosaccharide chains: NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb, NA4, NA4Fb, and NA4F2b.

[0025] Preferably, the product is a reagent kit.

[0026] A kit for assisting in the identification of hepatocellular carcinoma includes reagents for detecting the abundance of 12 specific oligosaccharide chains and a carrier for recording gender and age information; the 12 specific oligosaccharide chains consist of the following oligosaccharide chains: NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb, NA4, NA4Fb, and NA4F2b; the kit is used in conjunction with the aforementioned prediction model.

[0027] The beneficial effects of this invention are as follows:

[0028] (1) Through systematic glycomics analysis, this invention has for the first time discovered and verified a biomarker combination consisting of 12 specific oligosaccharide chains combined with gender and age. This invention not only reveals the specific expression characteristics of this biomarker combination in hepatocellular carcinoma patients, but more importantly, it elucidates for the first time the dynamic evolution of these 12 oligosaccharide chains throughout the entire process of "hepatitis → cirrhosis → hepatocellular carcinoma" and their synergistic association with gender and age. This discovery breaks through the limitations of existing technologies that only focus on a single disease stage or a single type of biomarker, providing a new perspective for a deeper understanding of the dynamic changes in glycosylation modifications during the development of hepatocellular carcinoma, and also providing a brand-new technical means for risk warning, early diagnosis and disease monitoring of hepatocellular carcinoma.

[0029] (2) This invention, through the cross-integration of systematic glycomics analysis and machine learning technology, for the first time uses 12 specific oligosaccharide chain biomarkers and clinical indicators such as gender and age as feature variables, and constructs a hepatocellular carcinoma prediction model (GlycanTest 2.0, GT2.0) using machine learning algorithms such as Bootstrap-integrated Support Vector Machine (SVM). This model achieves synergistic optimization of oligosaccharide chain biomarker combinations and artificial intelligence algorithms, and fully explores the nonlinear correlation information between oligosaccharide chain biomarkers and clinical indicators. Experimental results show that the GT2.0 model of this invention exhibits excellent predictive performance. In the training set, test set, and independent validation set, the accuracy of the model reaches 95.26%, 93.63%, and 89.43%, respectively, and the area under the receiver operating characteristic curve (AUC) reaches 0.984, 0.973, and 0.957, respectively, indicating that the GT2.0 model has extremely high diagnostic accuracy and good generalization ability, and has the potential for cross-center and cross-population applications.

[0030] (3) This invention shows significant advantages, particularly in the diagnosis of early-stage hepatocellular carcinoma. Experimental data show that for patients with very early-stage hepatocellular carcinoma at Barcelona Clinical Hepatocellular Carcinoma Stage 0 (BCLC), the diagnostic sensitivity of the GT2.0 model can reach 74.77%, while the sensitivity of the traditional clinical biomarker alpha-fetoprotein in early-stage hepatocellular carcinoma is less than 40%. This invention significantly improves the detection sensitivity of early-stage hepatocellular carcinoma, significantly increases the detection rate of early-stage patients, and provides patients with a valuable window for radical treatment, which is expected to significantly improve the overall prognosis of hepatocellular carcinoma patients.

[0031] (4) The GT2.0 model of this invention has a sensitivity of over 83% in hepatocellular carcinoma patients who are negative for both alpha-fetoprotein (AFP) and abnormal prothrombin (DCP). Specifically, the sensitivity reaches 87.53% in AFP-negative patients and 83.47% in DCP-negative patients. Even in patients who are negative for both, the sensitivity remains as high as 87.36%. This invention effectively fills the diagnostic gap of existing tumor markers in this population and provides a reliable auxiliary diagnostic method for hepatocellular carcinoma patients who are "marker-negative". Since GT2.0 is based on the principle of glycosylation modification, it has an essential technical complementarity with AFP and DCP. The combined application of the three can construct a comprehensive serological diagnostic system, significantly expanding the population coverage for hepatocellular carcinoma detection and reducing the overall missed diagnosis rate.

[0032] (5) This invention further verified the independent contribution of specific oligosaccharide chain components through ablation experiments. Experimental results showed that removing any one of the oligosaccharide chains NA4, NA4Fb, or NA4F2b from the model resulted in a statistically significant decrease in the AUC values ​​of the model in the training set, test set, and independent validation set (P < 0.05); the model performance decreased more significantly when all three were removed simultaneously (AUC dropped to 0.868). This result proves that NA4, NA4Fb, and NA4F2b each carry unique predictive information that cannot be compensated for by the other 11 oligosaccharide chains, as well as by gender and age, and the model performance is optimal when all three are present. Completely retaining NA4, NA4Fb, and NA4F2b in the GT2.0 model is a necessary and non-obvious technical means to obtain optimal predictive performance, and any simplification scheme that attempts to reduce or replace any of the oligosaccharide chains cannot achieve the technical effect described in this application.

[0033] (6) The test sample relied upon by this invention is blood (blood, serum, and plasma from venous or peripheral blood), which is easy to obtain, non-invasive, and has good patient compliance. The detection method provided by this invention has a standardized operating procedure, requires no complex equipment, and is easy to promote in primary hospitals and realize long-term patient condition management. At the same time, this invention forms a complete technical system from sample processing and oligosaccharide chain detection to automated data analysis and model prediction, which is conducive to standardized application and large-scale clinical promotion in different medical institutions and has good prospects for clinical translation. Attached Figure Description

[0034] Figure 1 Oligosaccharide chain maps of patients with hepatitis (A), cirrhosis (B), and hepatocellular carcinoma (C) in Example 1;

[0035] Figure 2 The ROC curves are for the training set (A), test set (B), and independent validation set (C) in Example 1. Detailed Implementation

[0036] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0037] Unless otherwise specified, the technical means used in the following embodiments are all conventional means well known to those skilled in the art, and the experimental methods without specific conditions are all conventional methods in the art.

[0038] Unless otherwise specified, all materials and reagents used in the following examples are commercially available.

[0039] Example 1

[0040] 1. Test Sample

[0041] This embodiment collected blood samples from 4149 patients as a dataset, including 1483 hepatitis samples, 891 cirrhosis samples, and 1775 hepatocellular carcinoma samples. These samples came from the Cancer Hospital Affiliated to Guangxi Medical University, Beijing You'an Hospital Affiliated to Capital Medical University, the First Affiliated Hospital of Zhejiang University School of Medicine, Eastern Hepatobiliary Surgery Hospital, Shanghai Changzheng Hospital, Mengchao Hepatobiliary Hospital of Fujian Medical University, and the Fifth People's Hospital of Ganzhou City. All participants were approved by the ethics review committees of the participating centers.

[0042] The 2285 samples from the first three centers (Guangxi Medical University Cancer Hospital, Beijing You'an Hospital affiliated to Capital Medical University, and the First Affiliated Hospital of Zhejiang University School of Medicine) were randomly divided into a training set and a test set, with 1374 samples in the training set and 1309 samples in the validation set. In the training set, there were 463 hepatitis samples, 314 cirrhosis samples, and 597 hepatocellular carcinoma samples. In the test set, there were 305 hepatitis samples, 212 cirrhosis samples, and 394 hepatocellular carcinoma samples.

[0043] The 1,864 samples from the last four centers (Eastern Hepatobiliary Surgery Hospital, Shanghai Changzheng Hospital, Mengchao Hepatobiliary Hospital of Fujian Medical University, and the Fifth People's Hospital of Ganzhou City) were used as an independent validation set, including 715 hepatitis samples, 365 cirrhosis samples, and 784 hepatocellular carcinoma samples.

[0044] 2. Instruments and equipment

[0045] Capillary electrophoresis analyzer, PCR instrument, centrifuge.

[0046] 3. Test reagents

[0047] Reagent A: 5 mM NH4HCO3 solution, with 1% SDS solution added;

[0048] Reagent B: Add 2 U / μL of exoglycoside exonuclease solution to a 1% (w / w) NP-40 solution;

[0049] Reagent C: Add 2 U / μL of sialidase solution to a 100 mM NH4AC solution with a pH of 5;

[0050] Reagent D: ddH2O;

[0051] Reagent E: A solution prepared by mixing 5 mM fluorescent labeling solution (trisodium 8-aminopyrene-1,3,6-trisulfonic acid) with DMSO solution (organic reducing agent NaBH3CN concentration of 1 M).

[0052] 4. Oligosaccharide chain pattern detection and collection

[0053] (1) Release of oligosaccharide chains

[0054] Add 3 μL of reagent A to 5 μL of sample, heat at 95℃ for 5 min to denature, cool to room temperature, add 3 μL of reagent B and 4 μL of reagent C, react at 37℃ for 4 h, and add 80 μL of reagent D.

[0055] (2) Marking of oligosaccharide chains

[0056] Take 10 μL of the sample solution from step (1), dry it at 70℃ for 30 min, then add 3 μL of reagent E, react at 90℃ for 2 h, and finally add 80 μL of reagent D to terminate the reaction.

[0057] (3) Detection of oligosaccharide chains and acquisition of spectra

[0058] Take 10 μL of the oligosaccharide chain sample prepared in step (2), place it in an ABI-specific 96-well plate, and detect it using an ABI3500 sequencer to obtain the oligosaccharide chain map.

[0059] Blood samples underwent protein denaturation, glycosidase treatment, fluorescent labeling, oligosaccharide chain mapping detection, and data acquisition to obtain the relative amounts of 12 specific oligosaccharide chains in each sample. For example... Figure 1 As shown, these 12 specific oligosaccharide chains are: NGA2F (galactosyl α-1,6-core fucosylated biantennary oligosaccharide chain), NGA2FB (galactosyl α-1,6-core fucosylated biantennary oligosaccharide chain), NG1A2F-1 (mono-branched galactosyl α-1,6-core fucosylated biantennary oligosaccharide chain), NG1A2F-2 (mono-branched galactosyl α-1,6-core fucosylated biantennary oligosaccharide chain), NA2 (galactosyl biantennary oligosaccharide chain), NA2F (galactosyl α-1,6-core fucosylated biantennary oligosaccharide chain), and NA2F (galactosylated α-1,6-core fucosylated biantennary oligosaccharide chain). The following oligosaccharides are listed: NA2FB (galactosyl α-1,6 core fucosylated biantennary oligosaccharide), NA3 (galactosylated triantennary oligosaccharide), NA3Fb (galactosylated α-1,3 branched fucosylated triantennary oligosaccharide), NA4 (galactosylated tetraantennary oligosaccharide), NA4Fb (galactosylated α-1,3 branched fucosylated tetraantennary oligosaccharide), and NA4F2b (galactosylated di-α-1,3 branched fucosylated tetraantennary oligosaccharide). Figure 1 It can be seen that the oligosaccharide chain profiles of patients with hepatitis, cirrhosis, and hepatocellular carcinoma are significantly different, indicating that oligosaccharide chains have the potential to identify hepatocellular carcinoma.

[0060] 5. Screening of characteristic oligosaccharide chains

[0061] Oligosaccharide chain data, along with gender and age information, were included in the training set of 1374 cases from patients with hepatitis, cirrhosis, and hepatocellular carcinoma. Hepatitis and cirrhosis samples were used as a non-hepatocellular carcinoma control group. Logistic regression was employed to screen candidate biomarkers for the predictive model. Based on statistical significance criteria, biomarkers were selected... P Indicators with values ​​less than 0.01 were selected as the final feature variables included in the prediction model. The selection results are shown in Table 1 below.

[0062] Table 1. Comparison of various indicators between the non-hepatocellular carcinoma control group and the hepatocellular carcinoma group in the training set.

[0063]

[0064] 6. Construct a classification model (GT2.0)

[0065] Based on 14 selected indicators, including gender, age, and 12 oligosaccharide chains, a classification model named GT2.0 (Glycan Test 2.0) was constructed using a Bootstrap-integrated Support Vector Machine (SVM) algorithm on the training set. During model construction, clinical diagnostic labels were used as the dependent variable, and gender, age, and oligosaccharide chain data were used as independent variables, with males encoded as 1 and females as 0. The optimal parameters of the model were determined through grid search combined with cross-validation, and the kernel function parameter γ and regularization parameter C were ultimately set to 1.45 and 0.09, respectively.

[0066] 7. Performance Evaluation of GT2.0 Model

[0067] like Figure 2 As shown, the area under the receiver operating characteristic (AUC) curves for the GT2.0 model in distinguishing between non-hepatocellular carcinoma and hepatocellular carcinoma in the training set, test set, and independent validation set were 0.984 (95% CI: 0.978-0.990), respectively. Figure 2 Medium A), 0.973 (95% CI: 0.962-0.983, Figure 2 in B) and 0.957 (95% CI: 0.948-0.966, Figure 2 (C). The value output by the GT2.0 model is used as the predicted value. The optimal threshold for distinguishing between hepatocellular carcinoma and non-hepatocellular carcinoma is determined by the optimal Youden index, which is 5.0. That is, a GT2.0 value ≥ 5.0 is judged as hepatocellular carcinoma, and a GT2.0 value < 5.0 is judged as non-hepatocellular carcinoma.

[0068] The specificity, sensitivity, and accuracy of the GT2.0 model on the training set, test set, and independent validation set are shown in Table 2 below.

[0069] Table 2. Specificity, sensitivity, and accuracy of the GT2.0 model in the training, test, and independent validation sets.

[0070]

[0071] As shown in Table 2, in the training set, the GT2.0 model achieved a specificity of 94.38% (437 / 463) for hepatitis, 92.99% (292 / 314) for cirrhosis, and 97.15% (580 / 597) for hepatocellular carcinoma, with an overall accuracy of 95.26% (1309 / 1374). In the test set, the specificity for hepatitis was 95.08% (290 / 305), the specificity for cirrhosis was 91.04% (193 / 212), the sensitivity for hepatocellular carcinoma was 93.91% (370 / 394), and the overall accuracy was 93.63% (853 / 911). In the independent validation set, the specificity for identifying hepatitis was 92.03% (658 / 715), the specificity for identifying cirrhosis was 90.41% (330 / 365), the sensitivity for identifying hepatocellular carcinoma was 86.61% (679 / 784), and the overall accuracy was 89.43% (1667 / 1864).

[0072] 8. Sensitivity of the GT2.0 model in early-stage hepatocellular carcinoma

[0073] Further analysis of the sensitivity of the GT2.0 model in patients with hepatocellular carcinoma at different Barcelona stages (BCLC) was conducted, and the results are shown in Table 3 below.

[0074] Table 3. Sensitivity of AFP, DCP, and GT2.0 models at different stages of BCLC

[0075]

[0076] Table 3 shows that the sensitivity of the GT2.0 model was 74.77% (83 / 111) in stage 0 (very early) patients; 88.48% (430 / 486) in stage A patients; 88.35% (91 / 103) in stage B patients; and 90.63% (58 / 64) in stage C patients. In comparison, the sensitivities of the traditional biomarker alpha-fetoprotein (AFP) in stages 0, A, B, and C were 45.05%, 37.45%, 44.66%, and 68.75%, respectively; and the sensitivities of abnormal prothrombin (DCP) were 36.04%, 69.96%, 75.73%, and 90.63%, respectively. In stages 0, A, and B, the sensitivity of GT2.0 was significantly superior to that of AFP and DCP (P < 0.0001).

[0077] 9. Sensitivity of the GT2.0 model in AFP / DCP-negative hepatocellular carcinoma

[0078] Table 4. Sensitivity of the GT2.0 model in AFP- and DCP-negative patients with hepatocellular carcinoma.

[0079]

[0080] As shown in Table 4, the sensitivity of the GT2.0 model was 87.53% (393 / 449) in 449 patients with AFP-negative (<20 ng / mL) hepatocellular carcinoma; 83.47% (207 / 248) in 248 patients with DCP-negative (<40 ng / mL) hepatocellular carcinoma; and 87.36% (159 / 182) in 182 patients with hepatocellular carcinoma who were negative for both AFP and DCP.

[0081] The experimental results of this embodiment demonstrate that the GT2.0 predictive model for identifying hepatocellular carcinoma, established based on oligosaccharide chains and clinical indicators, exhibits good generalization ability and stability, and has the potential for cross-center and cross-population applications. It increases the detection sensitivity of early-stage hepatocellular carcinoma to 74.47%, significantly improving the detection rate of early-stage patients and potentially significantly improving the overall prognosis of hepatocellular carcinoma patients. For hepatocellular carcinoma patients who are negative for both AFP and DCP, the diagnostic sensitivity of the GT2.0 model can reach over 83%. This technological achievement represents the first time that highly sensitive non-invasive detection has been achieved for patients with double-negative hepatocellular carcinoma, effectively filling the diagnostic gap of existing tumor markers in this group of patients and providing a reliable auxiliary diagnostic method for these "marker-negative" hepatocellular carcinoma patients.

[0082] Example 2 Ablation Experiment

[0083] To verify the independent contribution of specific oligosaccharide chain components (especially NA4, NA4Fb, and NA4F2b) to the predictive performance of the GT2.0 model described in this application, the following ablation experiment was designed.

[0084] 1. Experimental Design

[0085] Using the same training set, test set, and independent validation set as in Example 1, the following five machine learning prediction models were constructed:

[0086] (1) GT2.0 model (complete group): Includes all features, namely 12 oligosaccharide chains (NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb, NA4, NA4Fb, NA4F2b), gender, and age.

[0087] (2) Ablation group 1 (removal of NA4F2b): The oligosaccharide chain feature NA4F2b was removed from the GT2.0 model, and the remaining 11 oligosaccharide chains (NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb, NA4, NA4Fb), gender, and age were retained.

[0088] (3) Ablation group 2 (removal of NA4Fb): The oligosaccharide chain feature NA4Fb was removed from the GT2.0 model, and the remaining 11 oligosaccharide chains (NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb, NA4, NA4F2b), gender, and age were retained.

[0089] (4) Ablation group 3 (removal of NA4): The oligosaccharide chain feature NA4 was removed from the GT2.0 model, and the remaining 11 oligosaccharide chains (NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb, NA4Fb, NA4F2b), gender, and age were retained.

[0090] (5) Ablation group 4 (removal of NA4, NA4Fb, NA4F2b): Remove the three oligosaccharide chain features NA4, NA4Fb, and NA4F2b from the GT2.0 model, and retain the remaining 9 oligosaccharide chains (NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb), gender, and age.

[0091] 2. Model building methods

[0092] All models were constructed using the same Bootstrap ensemble Support Vector Machine (SVM) algorithm as in Example 1. During model construction, clinical diagnostic labels were used as the dependent variable, and the corresponding features (gender, age, and corresponding oligosaccharide chain data) were used as independent variables, with males encoded as 1 and females as 0. Optimal parameters for each model were determined through grid search combined with cross-validation, further confirming the optimal model for each group. The area under the receiver operating characteristic (AUC) curve for each model was output, and the Delong test was used to compare whether the differences in AUC values ​​between each ablation group and the GT2.0 model were statistically significant.

[0093] 3. Results of ablation experiments on the training set

[0094] Table 5 shows the AUC values ​​and Delong test results for each model in the 1374 training sets.

[0095] Table 5. Comparison of AUC values ​​between the GT2.0 model and each ablation group in the training set.

[0096]

[0097] As shown in Table 5, in the training set, the AUC values ​​of ablation groups 1-4 were 0.951, 0.933, 0.916 and 0.908, respectively, all significantly lower than the 0.984 of the GT2.0 model (P < 0.05), indicating that the absence of any oligosaccharide chain will lead to a significant decrease in model performance.

[0098] 4. Test set ablation experiment results

[0099] Table 6 shows the AUC values ​​and Delong test results for each model in the 1309-case test set.

[0100] Table 6. Comparison of AUC values ​​between the GT2.0 model and each ablation group in the test set.

[0101]

[0102] As shown in Table 6, in the test set, the AUC values ​​of ablation groups 1-4 were 0.944, 0.921, 0.914 and 0.895, respectively, all significantly lower than the 0.973 of the GT2.0 model (P < 0.05), which verifies the findings in the training set.

[0103] 5. Independent validation set ablation experiment results

[0104] Table 7 shows the AUC values ​​and Delong test results for each model in the 1864 independent validation sets.

[0105] Table 7 Comparison of AUC values ​​between the GT2.0 model and each ablation group in the independent validation set.

[0106]

[0107] As shown in Table 7, in the independent validation set, the AUC values ​​of ablation groups 1-4 were 0.932, 0.918, 0.894 and 0.868, respectively, all significantly lower than the 0.957 of the GT2.0 model (P < 0.05), further verifying the robustness of the model.

[0108] 6. Analysis of Experimental Results

[0109] Results from the training set, test set, and independent validation set all showed that the AUC of ablation groups 1-4 was statistically significantly lower than that of the complete GT2.0 model group (P < 0.05). This indicates that:

[0110] (1) NA4, NA4Fb and NA4F2b each carry 11 other oligosaccharide chains and unique predictive information that cannot be compensated by gender and age. The three oligosaccharide chains contribute independently and significantly.

[0111] (2) The absence of any oligosaccharide chain will lead to an unacceptable decrease in model performance;

[0112] (3) The AUC of ablation group 4 (simultaneous removal of NA4, NA4Fb and NA4F2b) is lower than that of ablation groups 1-3, and shows a gradual decreasing trend when they are removed one by one. The performance is optimal when all three are present.

[0113] Therefore, fully retaining NA4, NA4Fb, and NA4F2b in the GT2.0 model is a necessary but non-obvious technical means to obtain optimal predictive performance (AUC > 0.95). Any simplification scheme that attempts to reduce or replace any of the oligosaccharide chains cannot achieve the technical effect described in this invention.

[0114] The embodiments described above are only some, not all, of the embodiments of the present invention. The detailed description of the embodiments of the present invention is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments. The scope of protection of the present invention is determined by the scope claimed in the claims. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

Claims

1. A method for constructing a predictive model for identifying hepatocellular carcinoma, characterized in that, Includes the following steps: Step 1) Obtain the training sample set, which includes biological samples, clinical diagnostic labels, and the gender and age of multiple patients with known clinical diagnoses of hepatitis, cirrhosis, and hepatocellular carcinoma. Step 2) Detect the abundance of 12 specific oligosaccharide chains in each sample of the training sample set to obtain oligosaccharide chain abundance data; the 12 specific oligosaccharide chains are: NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb, NA4, NA4Fb, and NA4F2b; Step 3) Using the oligosaccharide chain abundance data, gender, and age as independent variables, and the clinical diagnostic label as the dependent variable, a machine learning algorithm is used to train the model to obtain the prediction model.

2. The method for constructing a predictive model for identifying hepatocellular carcinoma according to claim 1, characterized in that, Step 1) The biological sample is blood, serum or plasma from venous blood or peripheral blood of patients with hepatitis, cirrhosis or hepatocellular carcinoma.

3. The method for constructing a predictive model for identifying hepatocellular carcinoma according to claim 1, characterized in that, Step 3) The machine learning algorithm mentioned is the support vector machine algorithm integrated by Bootstrap.

4. A predictive model for identifying hepatocellular carcinoma, characterized in that, The prediction model is obtained by the construction method as described in any one of claims 1-3.

5. A predictive model for identifying hepatocellular carcinoma according to claim 4, characterized in that, The prediction model uses a predicted value ≥ 5.0 as the threshold for diagnosing hepatocellular carcinoma.

6. A predictive model for identifying hepatocellular carcinoma according to claim 4, characterized in that, The hepatocellular carcinoma mentioned is very early-stage hepatocellular carcinoma, specifically hepatocellular carcinoma in Barcelona clinical liver cancer stage 0 or A.

7. A predictive system for identifying hepatocellular carcinoma, characterized in that, include: The data acquisition module is used to acquire the abundance information of 12 specific oligosaccharide chains in the sample to be tested, as well as the gender and age of the individual to be tested; the 12 specific oligosaccharide chains are composed of the following oligosaccharide chains: NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb, NA4, NA4Fb, and NA4F2b; The prediction module stores the prediction model as described in claim 4. The result output module is used to output the hepatocellular carcinoma prediction results calculated by the prediction model based on the oligosaccharide chain abundance information, gender, and age.

8. A computer-readable storage medium having stored thereon the prediction model as claimed in claim 4.

9. The application of a reagent for detecting 12 specific oligosaccharide chains in the preparation of products for assisting in the identification of hepatocellular carcinoma, characterized in that, The 12 specific oligosaccharide chains consist of the following oligosaccharide chains: NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb, NA4, NA4Fb, and NA4F2b.

10. The application according to claim 9, characterized in that, The product in question is a reagent kit.

11. A kit for assisting in the identification of hepatocellular carcinoma, characterized in that, The kit contains reagents for detecting the abundance of 12 specific oligosaccharide chains and a carrier for recording gender and age information; the 12 specific oligosaccharide chains consist of the following oligosaccharide chains: NGA2F, NGA2FB, NG1A2F-1, NG1A2F-2, NA2, NA2F, NA2FB, NA3, NA3Fb, NA4, NA4Fb, and NA4F2b; the kit is used in conjunction with the prediction model as described in claim 4.