Prognosis model of intrahepatic cholangiocarcinoma and application thereof

By calculating the intratumor heterogeneity and inter-heterogeneity scores in the gene expression levels of patients with intrahepatic cholangiocarcinoma, determining the candidate gene set, and building a machine learning-based prognostic model, the poor repeatability and sampling bias of the existing models were solved, and a more accurate and stable prognostic evaluation was achieved.

CN120220808APending Publication Date: 2025-06-27SHENZHEN HUADA GENE INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311814023.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing prognostic model of intrahepatic cholangiocarcinoma has problems such as poor repeatability, poor robustness and susceptibility to clinical sampling bias, resulting in poor prediction results and difficulty in achieving clinical transformation.

Method used

By calculating the intratumor heterogeneity score and intertumor heterogeneity score in the gene expression level of cancer patients, the candidate gene set is determined, and a cancer prognosis model is constructed based on machine learning models, combining tumor heterogeneity scores to reduce the effect of sampling bias.

Benefits of technology

The constructed prognostic model is more accurate and stable, and is less affected by sampling deviation, which can effectively predict the prognostic risk of patients and improve the reliability and consistency of prognostic evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220808A_ABST
    Figure CN120220808A_ABST
Patent Text Reader

Abstract

The invention relates to the field of biology, in particular to a construction method of a prognosis model of cancer, especially intrahepatic cholangiocarcinoma, and a prognosis model related device, the construction method comprises the following steps: calculating intra-tumor heterogeneity score and inter-tumor heterogeneity score of each gene according to the gene expression level of a cancer patient; determining a candidate gene set according to the intra-tumor heterogeneity score and the inter-tumor heterogeneity score of each gene; and constructing the cancer prognosis model based on the candidate gene set and a machine learning model. The model construction method and the prognosis model provided by the embodiment of the invention have the advantages of being accurate, stable and less affected by sampling deviation, and can promote clinical management of intrahepatic cholangiocarcinoma patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of biology and involves machine learning methods, specifically a prognostic model for intrahepatic cholangiocarcinoma and a method for predicting the prognosis of intrahepatic cholangiocarcinoma based on this model. Background Art

[0002] Intrahepatic cholangiocarcinoma (also known as intrahepatic cholangiocellular carcinoma, ICC) is the second most common primary liver cancer. It has an insidious onset and strong invasiveness, and is one of the tumors with the worst prognosis. At the same time, ICC has a high degree of tumor heterogeneity, and patients with the same clinical stage may have very different prognoses. Currently, seven studies have constructed prognostic classification models for ICC based on the expression levels of prognostic genes (Table 1). However, after evaluation, these prognostic classification models all show problems such as poor reproducibility, poor robustness, and being easily affected by clinical sampling bias, resulting in poor prediction effects ( Figure 1 ), and thus it is difficult to achieve clinical translation.

[0003] Therefore, there is an urgent need to propose a prognostic model that is accurate, stable, and less affected by sampling bias to promote the clinical management of ICC patients and guide the treatment strategies of patients. Summary of the Invention

[0004] This application solves at least one of the problems of the related art from the following aspects.

[0005] To this end, an embodiment of this application provides a method for constructing a cancer prognostic model, including: calculating the intratumor heterogeneity score and intertumor heterogeneity score of each gene according to the gene expression level of cancer patients; determining a candidate gene set according to the intratumor heterogeneity score and the intertumor heterogeneity score of each gene; and constructing the cancer prognostic model based on the candidate gene set and a machine learning model.

[0006] In some embodiments, the calculating the intratumor heterogeneity score and intertumor heterogeneity score of each gene according to the gene expression level of cancer patients includes: calculating the standard deviation of the expression values of the gene in the first sample, and taking it as the intratumor heterogeneity score of the gene, where the first sample is a sample of different regions of the same tumor of the cancer patient; and calculating the standard deviation of the expression values of the gene in the second sample, and taking it as the intertumor heterogeneity score of the gene, where the second sample is a sample of any region in multiple regions of the tumors of all or part of the cancer patients.

[0007] In some embodiments, the second sample is randomly selected, and the calculation of the standard deviation of the expression value of the gene in the second sample is repeated M times. The mean of the M standard deviations is taken as the inter-tumor heterogeneity score of the gene, where M is a positive integer greater than 1. Optionally, M≥3, and preferably M≥10.

[0008] In some embodiments, the first sample and the second sample are taken from a multi-region sampling transcriptome dataset of the cancer.

[0009] In some embodiments, the cancer is a malignant solid tumor, and the malignant tumor is selected from one or more of the following groups: oral cancer, nasopharyngeal cancer, esophageal cancer, colon cancer, rectal cancer, colorectal cancer, lymphoma, thyroid cancer, lung cancer, liver cancer, intrahepatic cholangiocarcinoma, head and neck cancer, pancreatic cancer, gastric cancer, breast cancer, ovarian cancer, prostate cancer, uterine cancer, bladder cancer, medulloblastoma, glioma, and melanoma. In some embodiments, the cancer is intrahepatic cholangiocarcinoma.

[0010] In some embodiments, determining the candidate gene set according to the intra-tumor heterogeneity score and the inter-tumor heterogeneity score of each gene includes: performing a correlation analysis on the intra-tumor heterogeneity score and the inter-tumor heterogeneity score of the gene. Based on the inter-tumor heterogeneity score of the gene being higher than the intra-tumor heterogeneity score, the gene is included in the candidate gene set as a candidate gene.

[0011] In other embodiments, determining the candidate gene set according to the intra-tumor heterogeneity score and the inter-tumor heterogeneity score of each gene includes: determining the median of the inter-tumor heterogeneity scores according to the inter-tumor heterogeneity scores of each gene; based on the inter-tumor heterogeneity score of the gene being higher than the median of the inter-tumor heterogeneity scores, the gene is included in the preselected gene set as a preselected gene; and based on the intra-tumor heterogeneity score and the inter-tumor heterogeneity score of each gene, using the quantile regression method to divide the preselected gene set, where the genes falling into the interval below the lower quartile regression line are included in the candidate gene set as candidate genes.

[0012] In some embodiments, the method further includes: performing a first screening on the candidate gene set based on prognostic significance to obtain a first gene set with significant prognostic differences, where the prognostic significance is a p-value less than 0.05, less than 0.01, or less than 0.001, preferably less than 0.05; performing a second screening on the first gene set based on the high-expression genes of the cancer to obtain a second gene set that is highly expressed in the cancer, where the high expression is a log fold change of the expression value higher than 0.25 and the p-value is less than 0.05. Optionally, the calculation of the high-expression level genes is based on the single-cell transcriptome data of the cancer; and constructing the cancer prognosis model based on the second gene set and a machine learning model.

[0013] In some embodiments, univariate linear regression is used for the statistical analysis of the prognostic significance in the first screening, and the univariate linear regression is Cox regression analysis.

[0014] In some embodiments, constructing the cancer prognosis model based on the candidate gene set and a machine learning model includes: training the machine learning model according to the candidate gene set based on the single-region sampling transcriptome data set and / or multi-region sampling transcriptome data set of the cancer patients to screen out a prognostic feature gene set, and obtaining the cancer prognosis model with the prognostic feature gene set as the model feature. Optionally, the prognostic model is based on a linear model, and the linear model is selected from one or more of a linear regression model, a logistic regression model, a Lasso regression model, a ridge regression model, and a linear discriminant analysis model, preferably a Lasso regression model.

[0015] In some embodiments, the prognostic feature gene set includes one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2.

[0016] In some embodiments, the model feature further includes model coefficients. Optionally, the model coefficients of each gene in the prognostic feature gene set are as follows:

[0017]

[0018]

[0019] Embodiments of the present application also provide a cancer prognosis model, which is constructed according to the construction method of the cancer prognosis model as described in any of the above embodiments. The model includes: a calculation module for using the prognostic feature gene set as model features and calculating a prediction result based on the levels of each gene in the prognostic feature gene set in a biological sample derived from a subject; and an indication module for indicating the prognostic risk of the subject according to the prediction result, where the prognostic feature gene set includes one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2.

[0020] Embodiments of the present application also provide a method for indicating cancer prognostic risk, including: using the prognostic feature gene set as a prediction variable and calculating a prediction result based on the expression levels of each gene in the prognostic feature gene set in a biological sample derived from a subject; and indicating the prognostic risk of the subject according to the prediction result, where the prognostic feature gene set includes one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2.

[0021] In some embodiments, the calculation model for the prognostic risk is:

[0022]

[0023] where x i is the expression level of each gene in the prognostic feature gene set, and Coef is the model coefficient of each gene in the prognostic feature gene set.

[0024] where the model coefficients of each gene are as follows:

[0025]

[0026]

[0027] In some embodiments, the expression level of each gene is RPKM (Reads Per Kilobase per Million mapped reads), FPKM (Fragments Per Kilobase of exon model per Million mapped fragments), or TPM (Transcripts Per Million), preferably FPKM.

[0028] An embodiment of the present application further provides a cancer prognosis system, including: a processor; an input module for inputting the levels of each gene in a prognostic signature gene set from a biological sample derived from a subject, where the prognostic signature gene set includes one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2; a computer-readable medium containing instructions that, when executed by the processor, implement the intrahepatic cholangiocarcinoma prognosis risk indication method as described in any of the above embodiments of the present application; and an output module for indicating the prognosis risk of the subject.

[0029] An embodiment of the present application also provides the use of a reagent for detecting a biomarker in the preparation of a product for cancer prognosis, where the biomarker includes one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2.

[0030] In some embodiments, the reagent is selected from: a probe that specifically recognizes the biomarker; a primer that specifically recognizes the biomarker; or an antibody or ligand that specifically binds to the biomarker.

[0031] The embodiments of the present application achieve the following beneficial effects:

[0032] The method for constructing the cancer prognosis model proposed in this application first takes tumor heterogeneity into consideration. By combining the intratumor heterogeneity score and the intertumor heterogeneity score, it constructs a more accurate, stable prognosis model with less influence from sampling bias, and identifies biomarkers that can be used to indicate the prognosis risk assessment of this cancer. Further, taking intrahepatic cholangiocarcinoma as an example, the method for constructing the intrahepatic cholangiocarcinoma prognosis model is based on a research cohort of ICC patients with a larger data volume and more comprehensive data types (a total of 515 intrahepatic cholangiocarcinoma patients). In the prognosis, it takes into account the intratumor and intertumor heterogeneity of intrahepatic cholangiocarcinoma, and constructs an accurate, stable prognosis model with less influence from sampling bias by combining machine learning methods according to the two heterogeneity scores, and identifies a set of biomarkers that can be used for accurate and stable prognosis of intrahepatic cholangiocarcinoma. The prognosis model and prognosis system developed in this application can reliably predict whether a patient is at high-risk prognosis based on the identified biomarkers. Especially when using different samples of the same patient, the risk levels evaluated are highly consistent. Therefore, it effectively solves the problems of low reproducibility, susceptibility to sampling bias, and low prediction accuracy of existing prognosis models. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0034] Figure 1 It is an evaluation diagram of the effect of the prognosis classification model in the related art;

[0035] Figure 2 It is another evaluation diagram of the effect of the prognosis classification model in the related art;

[0036] Figure 3 It shows the screening of candidate genes according to two heterogeneity scores according to the embodiments of this application;

[0037] Figure 4 It is the method for constructing the intrahepatic cholangiocarcinoma prognosis model proposed in the embodiments of this application;

[0038] Figure 5 It is the effect verification of the prognosis model according to the embodiments of this application;

[0039] Figure 6 It is the consistency of the risk levels evaluated from different samples of the same patient according to the embodiments of this application;

[0040] Figure 7 It is a schematic diagram of multi-region sampling according to the embodiments of this application. Detailed Embodiments

[0041] The present invention will be further described in detail below in conjunction with specific embodiments. The provided embodiments are only for clarifying the present invention and do not limit the scope of the present invention. The following provided embodiments can be used as a guide for those of ordinary skill in the art to make further improvements and do not constitute any limitation to the present invention in any way.

[0042] This application is based on the following understanding of the inventors:

[0043] In the related art, there have been 7 studies that have constructed prognostic classification models for ICC based on the expression levels of prognostic genes (Table 1). However, these models have many problems and cannot be effectively used for clinical translation. Specifically, 1) the existing models show poor applicability to new data or new samples: According to the latest literature reports (as of March 2023), only 2 of the 7 existing prognostic gene expression models for intrahepatic cholangiocarcinoma have prognostic prediction value in a new large ICC public dataset, indicating that the current prognostic classification system has poor reproducibility ( Figure 1 ); 2) the data used in the construction of the existing prognostic models involves a relatively small number of patients, resulting in poor robustness of the prognostic models; 3) the existing model construction is easily affected by clinical sampling bias, resulting in poor prediction effects: Applying these prognostic models to the multi-point sampling patient cohort published by the inventors previously (i.e., the multi-region sampling transcriptome dataset, in which 4-6 different regions of the same tumor tissue of each intrahepatic cholangiocarcinoma patient were extracted for transcriptome sequencing, https: / / www.biosino.org / node / project / detail / OEP002560), it was found that 33.3 - 57.8% of the patients may be misclassified as wrong prognoses due to sampling bias ( Figure 2 ).

[0044] Table 1. Prognostic Classification Models for 7 ICCs

[0045]

[0046] In response to this, after extensive research, the inventors believe that tumors, especially intrahepatic cholangiocarcinoma, are highly heterogeneous tumors, and the intratumoral transcriptome heterogeneity is an important factor affecting prognostic models. Based on this, the prognostic model in the embodiments of this application takes into account the tumor heterogeneity of ICC and constructs a prognostic model for the clonal tumor expression genes of intrahepatic cholangiocarcinoma with little influence from sampling bias, strong reproducibility, operability, and high efficiency based on a research cohort of ICC patients with a larger data volume and more comprehensive data types (a total of 515 intrahepatic cholangiocarcinoma patients), in combination with machine learning methods.

[0047] The prognostic model developed in the embodiments of the present application can reliably predict whether a patient is in a high-risk tumor type by estimating the transcriptome data of cancer patients, especially intrahepatic cholangiocarcinoma patients, effectively solving the problems of low duplication rate, susceptibility to sampling bias, and low prediction accuracy of existing prognostic models.

[0048] Figure 4 This is a method for constructing an intrahepatic cholangiocarcinoma prognostic model proposed in the embodiments of the present application. As Figure 4 shown, the method for constructing an intrahepatic cholangiocarcinoma prognostic model may include the following steps: S1 - S3.

[0049] S1: Calculate the intratumoral heterogeneity score and intertumoral heterogeneity score of each gene according to the gene expression level of cancer patients.

[0050] In the embodiments of the present application, the gene expression level of the cancer patients can be obtained from a multi-region sampling transcriptome dataset (multi-point sampling cohort), where multi-region refers to different positions of the tumor tissue. Taking intrahepatic cholangiocarcinoma as an example, multi-region sampling can refer to randomly taking 4 - 6 regional samples from a piece of tumor tissue and performing bulk whole-transcriptome sequencing respectively (multi-region sampling can be as Figure 7 shown, where R1 - R4 represent multiple different regions). It can be understood that based on the multi-region sampling transcriptome data of intrahepatic cholangiocarcinoma patients, the expression levels of each gene in all regions of the tumor tissue of a certain intrahepatic cholangiocarcinoma patient can be extracted, so as to calculate and consider the tumor heterogeneity of intrahepatic cholangioma. In some embodiments, the multi-region sampling transcriptome dataset can be obtained from https: / / www.biosino.org / node / project / detail / OEP002560, or other cohorts with multi-point sampling data. The present application does not limit the acquisition source of the dataset.

[0051] In the embodiments of the present application, the consideration of the tumor heterogeneity of genes includes two scores, namely the intra-tumor heterogeneity score and the inter-tumor heterogeneity score. Specifically, in some embodiments, the standard deviation of the expression values of all or part of the expressed genes in the first sample in the multi-region sampling transcriptome dataset can be calculated and used as the intra-tumor heterogeneity score of the gene, where the first sample is a sample of different regions of the same tumor of the same patient; and the standard deviation of the expression values of all or part of the expressed genes in the second sample in the multi-region sampling transcriptome dataset can be calculated and used as the inter-tumor heterogeneity score of the gene, where the second sample can be a sample of any region in the multi-regions of the tumors of all or part of the cancer patients in the multi-region sampling transcriptome dataset. In some embodiments, a sample of any region of each patient in the multi-region sampling transcriptome dataset (multi-point sampling cohort) is selected for the calculation of the inter-tumor heterogeneity score. In some embodiments, the selection of any region of the second sample is random.

[0052] In some embodiments, the standard deviation of the expression values of the gene in the second sample can be calculated repeatedly, for example, repeated M times, and the mean of the M standard deviations is taken as the inter-tumor heterogeneity score of the gene, where M is a positive integer greater than 1. In some embodiments, M≥3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30. In some embodiments, M≥35; in other embodiments, M≥40. It can be understood that by repeatedly calculating the standard deviation of the expression values in the second sample, the possible calculation bias between samples can be effectively balanced, so as to obtain a more accurate inter-tumor heterogeneity score.

[0053] S2: Determine a candidate gene set according to the intra-tumor heterogeneity score and the inter-tumor heterogeneity score of each gene.

[0054] In the embodiments of the present application, after obtaining the intra-tumor heterogeneity score and the inter-tumor heterogeneity score of a gene, a correlation analysis can be performed on the two, and genes with an inter-tumor heterogeneity score higher than, preferably significantly higher than, the intra-tumor heterogeneity score are included in the candidate gene set. In some other embodiments, step S2 may include: determining the median of the inter-tumor heterogeneity scores of all genes based on the inter-tumor heterogeneity scores of each gene; if the inter-tumor heterogeneity score of a certain gene is higher than, preferably significantly higher than, the median of the inter-tumor heterogeneity scores of all these genes, then including this gene as a preselected gene in the preselected gene set; and based on the intra-tumor heterogeneity score and the inter-tumor heterogeneity score of each gene, using the quantile regression method to divide this preselected gene set, and including genes falling into the interval below the lower quartile regression line as candidate genes in the candidate gene set. It can be understood that by comparing the intra-tumor heterogeneity score and the inter-tumor heterogeneity score of a gene, and using genes with an inter-tumor heterogeneity score significantly higher than the intra-tumor heterogeneity score as candidate genes, genes less affected by intra-tumor heterogeneity can be effectively screened out, thereby excluding errors caused by sampling bias, and thus effectively ensuring that the subsequent model constructed based on these candidate genes is more stable and accurate.

[0055] In the embodiments of the present application, after determining the candidate gene set, the construction method further includes: performing a first screening on the candidate gene set based on prognostic significance to obtain a first gene set with significant prognostic differences, wherein the discrimination method for genes with significant prognostic differences is: based on the survival status and survival time data of cancer patients (such as intrahepatic cholangiocarcinoma patients) (these data can be from relevant records in a multi-point sampling cohort or other single-point sampling cohorts), if the high or low expression level of a certain gene has a significant correlation with prognosis, that is, the expression level of the gene significantly affects the prognosis of the patient, then it is considered a gene with significant prognostic differences, where the significance is that the p-value is less than 0.05, less than 0.01 or less than 0.001, preferably less than 0.05; performing a second screening on the first gene set based on the high-expression genes of the cancer to obtain a second gene set with high expression in the cancer, and subsequently constructing an intrahepatic cholangiocarcinoma prognosis model based on the second gene set and a machine learning model. In some embodiments, the genes highly expressed in the tumor cells of the cancer are genes with a log fold change in expression value higher than 0.25 compared to normal cells, and at the same time the p-value is less than 0.05, less than 0.01 or less than 0.001, preferably less than 0.05. It can be understood that in the embodiments of the present application, through the first screening, genes (i.e., the first gene set) whose gene levels can affect the prognosis of cancer patients are obtained from the genes expressed by cancer patients, and through the second screening, genes highly expressed in the tumor cells of cancer patients are screened out, making the features used in subsequent model construction more cancer-specific.

[0056] In some embodiments, univariate linear regression is used for the statistical analysis of the prognostic significance in the first screening, such as univariate Cox regression analysis.

[0057] S3: Based on the candidate gene set, construct a cancer prognosis model based on a machine learning model.

[0058] In the embodiments of the present application, the prognosis model can be constructed based on the candidate gene set obtained in step S2, or based on the second gene set after the first and second screenings above.

[0059] Specifically, in some embodiments, step S3 may include: based on the single-region sampling transcriptome dataset and / or multi-region sampling transcriptome dataset of cancer patients, training a machine learning model according to a candidate gene set to screen out a prognostic feature gene set, and obtaining a cancer prognosis model with the prognostic feature gene set as the model feature. In some embodiments, the prognosis model is based on a linear model, and the linear model is selected from one or more of a linear regression model, a logistic regression model, a Lasso regression model, a ridge regression model, and a linear discriminant analysis model, preferably a Lasso regression model. In some embodiments, the single-region sampling transcriptome dataset may be, for example, CPTAC (Clinical Proteomic Tumor Analysis Consortium) or a single-point sampling cohort for intrahepatic cholangiocarcinoma (https: / / www.biosino.org / node / project / detail / OEP001105). It can be understood that based on the transcriptome database, through continuous fitting and optimization of the model in model training and optionally validating the model after model training, model features (prognostic feature gene set / prognostic feature genes) that can be effectively used for prognosis can be determined. These prognostic feature genes can be used as biomarkers for cancer prognosis, and by detecting the expression levels of these biomarkers, the prognostic risk of the subject can be accurately evaluated.

[0060] Accordingly, an embodiment of the present application also proposes an application of a biomarker (i.e., a prognostic feature gene) and a reagent for detecting the biomarker in the preparation of a product for cancer prognosis. In some embodiments, the biomarker (prognostic feature gene set / prognostic feature gene) includes one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2. The cancer is preferably intrahepatic cholangiocarcinoma.

[0061] In some embodiments, the reagent for detecting the biomarker may be a probe that specifically recognizes the biomarker, a primer that specifically recognizes the biomarker, and / or an antibody or ligand that specifically binds to the biomarker. In some embodiments, the reagent for detecting the biomarker may be a reagent used in fluorescence quantitative PCR, western blot, real-time fluorescence quantitative (RT-qPCR), gel shift assay (EMSA), in situ hybridization, immunoprecipitation, immunocytochemistry, subcellular fractionation, two-dimensional gel electrophoresis, microfluidic immunocytometry, surface enhanced laser desorption ionization time-of-flight mass spectrometry (SELDI-TOF), and sequencing for detecting its expression level, where the sequencing may be transcriptome sequencing. It can be understood that the reagent for detecting the biomarker proposed in the embodiments of the present application only needs to be able to detect the expression level of the biomarker in the sample, and the present application does not limit this. In addition, it can be understood that according to the method for constructing a cancer (especially intrahepatic cholangiocarcinoma) prognosis model proposed in the embodiments of the present application and the model constructed based on this method, it is possible to accurately and effectively evaluate the prognosis of cancer, and refer to the prognosis evaluation result to perform relevant treatment on clinical cancer.

[0062] When the method for constructing a prognosis model proposed in the embodiments of the present application identifies the above-mentioned biomarkers (prognostic feature gene set / prognostic feature genes) that can be effectively used for the prognosis of intrahepatic cholangiocarcinoma, the model features of the constructed model also include model coefficients (feature weights) corresponding to the biomarkers. In some embodiments, the model coefficients of each gene in the prognostic feature gene set are selected from one or more of the genes and their model coefficients listed in Table 2.

[0063] Table 2

[0064]

[0065]

[0066] In some embodiments, the calculation model for constructing the prognostic risk may be: the product value of the expression level of the biomarker for the prognosis of intrahepatic cholangiocarcinoma (the biomarker shown in Table 2) and the model coefficients of each biomarker. Optionally, the coefficients in Table 2 may have a perturbation within 5% of the coefficient value.

[0067] In some embodiments, the expression information (level) of the biomarker for intrahepatic cholangiocarcinoma prognosis is denoted as x i , and its model coefficient is denoted as Coef. Calculate the product value of the expression information of the biomarker for intrahepatic cholangiocarcinoma prognosis and the target model coefficient, that is, Coef i *x i , and the tumor prognosis risk assessment model is obtained as follows:

[0068]

[0069] Among them, the output data of the intrahepatic cholangiocarcinoma prognosis risk assessment model is the intrahepatic cholangiocarcinoma prognosis risk score. Among them, the higher the intrahepatic cholangiocarcinoma prognosis risk score, the higher the prognosis risk; the lower the intrahepatic cholangiocarcinoma prognosis risk score, the lower the prognosis risk. This prognosis risk score can be used to indicate the cure situation of the patient and assist in guiding the patient's treatment plan.

[0070] Preferably, the expression levels of the above-mentioned biomarker genes can be RPKM (abbreviation for Reads Per Kilobase per Million mapped reads, representing the number of reads per kilobase length of a certain gene in every million reads), FPKM (Fragments Per Kilobase of exon model per Million mapped fragments, that is, the number of fragments of every thousand-base transcription per million mapped reads), TPM (Transcripts Per Million, the number of transcripts per million), etc., preferably the FPKM value.

[0071] The embodiment of the present application also proposes an intrahepatic cholangiocarcinoma prognosis model. This prognosis model is constructed according to the construction method of the intrahepatic cholangiocarcinoma prognosis model described in any of the above embodiments. This prognosis model may include: a calculation module, which uses the prognosis feature gene set as model features and calculates the prediction result based on the levels of each gene in the prognosis feature gene set in the biological sample derived from the subject; and an indication module, which indicates the prognosis risk of the subject according to the prediction result, where the prognosis feature gene set includes one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2.

[0072] In the embodiments of the present application, "prognostic risk" refers to the assessment of the degree of local recurrence in cancer patients (such as intrahepatic cholangiocarcinoma) after treatment. In some embodiments, "prognostic risk" can be divided into low risk, medium risk and high risk, or can be divided into low risk and high risk. In some embodiments, a preset prognostic risk value is set. If the prognostic risk score of intrahepatic cholangiocarcinoma is higher than the preset prognostic risk value, the subject is determined to have a high prognostic risk. If the prognostic risk score of intrahepatic cholangiocarcinoma is lower than the preset prognostic risk value, the prognostic risk status is determined to be a low prognostic risk status, and the subject is determined to have a low prognostic risk.

[0073] Preferably, the above preset prognostic risk value (i.e., threshold) refers to the average value, median value, upper quartile value, etc. of the prognostic risk values of all intrahepatic cholangiocarcinoma patients in the data cohort used to construct the prognostic model, and is preferably the median value.

[0074] The embodiments of the present application also propose a method for indicating the prognostic risk of cancer, preferably intrahepatic cholangiocarcinoma, including: using a prognostic signature gene set as a predictive variable, calculating a prediction result based on the expression levels of the genes in the prognostic signature gene set in a biological sample derived from a subject; and indicating the prognostic risk of the subject according to the prediction result, where the prognostic signature gene set includes one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2.

[0075] In some embodiments, the calculation model of the prognostic risk is:

[0076]

[0077] where x i is the expression level of each gene in the prognostic signature gene set, and Coef is the model coefficient of each gene in the prognostic signature gene set.

[0078] The model coefficients of each gene are as follows:

[0079]

[0080] Optionally, the expression level of each gene is RPKM (Reads Per Kilobase per Million mapped reads), FPKM (Fragments Per Kilobase of exon model per Million mapped fragments), or TPM (Transcripts Per Million), and is preferably FPKM.

[0081] An embodiment of the present application also provides a prognosis system for cancer, preferably intrahepatic cholangiocarcinoma, comprising: a processor; an input module for inputting the levels of each gene in a prognostic signature gene set from a biological sample derived from a subject, wherein the prognostic signature gene set comprises one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2; a computer-readable medium containing instructions which, when executed by the processor, implement the method for indicating the prognostic risk of intrahepatic cholangiocarcinoma as described in any of the above embodiments of the present application; and an output module for indicating the prognostic risk of the subject.

[0082] It can be understood that the method for constructing the prognostic model proposed in the embodiments of the present application is also applicable to other malignant tumors (especially solid tumors) other than intrahepatic cholangiocarcinoma, such as one or more selected from the following group: oral cancer, nasopharyngeal cancer, esophageal cancer, colon cancer, rectal cancer, colorectal cancer, lymphoma, thyroid cancer, lung cancer, liver cancer, intrahepatic cholangiocarcinoma, head and neck cancer, pancreatic cancer, gastric cancer, breast cancer, ovarian cancer, prostate cancer, uterine cancer, bladder cancer, medulloblastoma, glioma, and melanoma. Based on the multi-point sampling cohort of each malignant tumor, incorporating its tumor heterogeneity into the prognostic risk assessment can obtain a prognostic model that is accurate, stable, and less affected by sampling bias; the application of model coefficients, biomarkers, etc. screened based on this method will also provide a reliable indication for the prognosis of patients, thus effectively promoting the clinical management of patients with malignant tumors.

[0083] It should be noted that the foregoing explanation of the embodiments of the method for constructing the cancer prognostic model is also applicable to the cancer prognostic risk indication method, model, system, and related applications in the above embodiments, and will not be elaborated here.

[0084] The experimental methods in the following embodiments are all conventional methods unless otherwise specified, and are carried out according to the techniques or conditions described in the literature in the art or according to the product specifications. The materials, reagents, etc. used in the following embodiments can be obtained from commercial sources unless otherwise specified.

[0085] Unless otherwise specified, the quantitative analysis tests in the following embodiments are all set up with three repeated experiments, and the results are averaged.

[0086] Example

[0087] In this example, intrahepatic cholangiocarcinoma was taken as an example for model construction and prognostic risk assessment.

[0088] 1.1 Determination of candidate genes

[0089] This example is based on the gene expression data of multi-region (multi-point sampling) transcriptome sequencing of intrahepatic cholangiocarcinoma (https: / / www.biosino.org / node / project / detail / OEP002560). The intra-tumor heterogeneity score and inter-tumor heterogeneity score of all expressed genes were calculated. The calculation method of the intra-tumor heterogeneity score for each gene is: the standard deviation of the expression values of each gene in different regional samples of the same tumor. The calculation method of the inter-tumor heterogeneity score for each gene is: randomly select one sample for each patient, then calculate the standard deviation of the expression values of each gene in these samples, repeat 10 times, and then take the average value, which is the inter-tumor heterogeneity score. Through the correlation analysis of these two heterogeneity indicators, it was found that 1341 genes showed that the inter-tumor heterogeneity score was significantly higher than the intra-tumor heterogeneity score, indicating that these genes were less affected by intra-tumor heterogeneity and could exclude the errors caused by sampling bias( Figure 3 ).

[0090] 1.2 Screening of characteristic genes and construction of models

[0091] Based on the 1341 candidate genes screened out, the screening of characteristic genes and the construction of a prognostic model were carried out using machine learning according to the method as Figure 4 shown. Specifically, first, the univariate Cox regression analysis method was used to select 744 genes with significant prognostic differences (P value less than 0.05); then, using the single-cell transcriptome data of intrahepatic cholangiocarcinoma (https: / / ngdc.cncb.ac.cn / gsa-human / browse / HRA000863), the genes highly expressed in tumor cells in intrahepatic cholangiocarcinoma among the 744 genes were screened out, a total of 209 genes (q value less than 0.05). Then, in the gene expression training data set containing 249 patients (CPTAC cohort), the LASSO regression machine learning method was used to train a prognostic model containing 15 prognostic feature sets (Table 3).

[0092] Table 3. Prognostic model containing 15 prognostic feature sets

[0093]

[0094]

[0095] 1.3 Model validation

[0096] To verify the accuracy and stability of the constructed model, it was further verified using 4 published gene expression data sets of intrahepatic cholangiocarcinoma (GSE89749, E-MTAB-6389, TCGA, GSE107943), and the results are as Figure 5As shown. Refer to Figure 5 It can be seen that for each cohort, by using the intrahepatic cholangiocarcinoma gene expression prognosis model constructed in this embodiment, ICC patients can be effectively divided into two categories: high-risk and low-risk prognosis. Using the univariate Cox regression analysis method, it is found that: patients with high-risk prognosis have a significantly worse prognosis than those with low-risk prognosis, which fully conforms to medical logic, indicating that the model of this embodiment can complete the accurate classification of the prognosis of ICC patients. Among them, 4 published intrahepatic cholangiocarcinoma gene expression data sets are from relevant databases, namely CPTAC (Clinical Proteomic Tumor Analysis Consortium), TCGA (The Cancer Genome Atlas), and NCBI database.

[0097] In addition, based on the gene expression data of multi-point sampling, the intrahepatic cholangiocarcinoma gene expression prognosis model constructed in this embodiment can effectively divide patients into two categories: high-risk and low-risk prognosis (using the median of the risk score as the discrimination threshold), and it is found that using this model, the risks evaluated from different samples of the same patient have high consistency. Among them, only 0.82% of the patients show that the prognosis risk types in all regions of the same tumor are of the same type (both high-risk or both low-risk). This not only shows that the model of this embodiment has high consistency in discrimination, but also indirectly confirms the existence of tumor heterogeneity and the necessity of incorporating tumor heterogeneity into prognosis evaluation ( Figure 6 ). These verification results show that the prognosis model proposed in this embodiment has good classification accuracy, and considering tumor heterogeneity makes it less affected by sampling bias, with strong repeatability and robustness, so it can be effectively applied to clinical practice.

[0098] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0099] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for constructing a cancer prognosis model, characterized in that, Comprising: Calculating the intratumoral heterogeneity score and the intertumoral heterogeneity score of each gene according to the gene expression levels of cancer patients; Determining a candidate gene set according to the intratumoral heterogeneity score and the intertumoral heterogeneity score of each gene; and Constructing the cancer prognosis model based on the candidate gene set and a machine learning model.

2. The method according to claim 1, characterized in that, The calculating the intratumoral heterogeneity score and the intertumoral heterogeneity score of each gene according to the gene expression levels of cancer patients includes: Calculating the standard deviation of the expression values of the gene in a first sample, and taking it as the intratumoral heterogeneity score of the gene, wherein the first sample is samples from different regions of the same tumor of the cancer patient; and Calculating the standard deviation of the expression values of the gene in a second sample, and taking it as the intertumoral heterogeneity score of the gene, wherein the second sample is a sample from any region among multiple regions of the tumors of all or part of the cancer patients, Preferably, the second sample is randomly selected, the calculation of the standard deviation of the expression values of the gene in the second sample is repeated M times, and the mean of the M standard deviations is taken as the intertumoral heterogeneity score of the gene, where M is a positive integer greater than 1, optionally, M≥3, preferably M≥10, Optionally, the first sample and the second sample are taken from the multi-region sampling transcriptome dataset of the cancer.

3. The method according to claim 1 or 2, characterized in that, The cancer is a malignant tumor, and the malignant solid tumor is selected from one or more of the following groups: oral cancer, nasopharyngeal cancer, esophageal cancer, colon cancer, rectal cancer, colorectal cancer, lymphoma, thyroid cancer, lung cancer, liver cancer, intrahepatic cholangiocarcinoma, head and neck cancer, pancreatic cancer, gastric cancer, breast cancer, ovarian cancer, prostate cancer, uterine cancer, bladder cancer, medulloblastoma, glioma, and melanoma, Preferably, the cancer is intrahepatic cholangiocarcinoma.

4. The method according to claim 1 or 2, characterized in that The determining a candidate gene set according to the intratumoral heterogeneity score and the intertumoral heterogeneity score of each gene includes: Performing a correlation analysis on the intratumoral heterogeneity score and the intertumoral heterogeneity score of the gene, and based on the intertumoral heterogeneity score of the gene being higher than the intratumoral heterogeneity score, including the gene as a candidate gene in the candidate gene set, Optionally, the determining a candidate gene set according to the intratumoral heterogeneity score and the intertumoral heterogeneity score of each gene includes: Determining the median of the intertumoral heterogeneity scores according to the intertumoral heterogeneity scores of each gene; Based on the intertumoral heterogeneity score of the gene being higher than the median of the intertumoral heterogeneity scores, including the gene as a preselected gene in the preselected gene set; and Based on the intratumoral heterogeneity score and the intertumoral heterogeneity score of each gene, using the Quantile Regression method to divide the preselected gene set, and including the genes falling into the interval below the lower quartile regression line as candidate genes in the candidate gene set.

5. The method according to claim 1, wherein The method further includes: Perform a first screening on the candidate gene set based on prognostic significance to obtain a first gene set with significant prognostic differences, where the prognostic significance is a p-value less than 0.05, less than 0.01, or less than 0.001, preferably less than 0.05; Perform a second screening on the first gene set based on the genes highly expressed in the tumor cells of the cancer to obtain a second gene set highly expressed in the cancer, where the high expression is a log fold change of the expression value higher than 0.25 and at the same time the p-value is less than 0.

05. Optionally, the calculation of the highly expressed genes is based on the single-cell transcriptome data of the cancer; and Construct the cancer prognosis model based on the second gene set and a machine learning model. Optionally, use univariate linear regression for the statistical analysis of the prognostic significance in the first screening, and the univariate linear regression is Cox regression analysis.

6. The method according to claim 1, characterized in that, The constructing of the cancer prognosis model based on the candidate gene set and a machine learning model includes: Based on the single-region sampling transcriptome dataset and / or multi-region sampling transcriptome dataset of the cancer patients, train the machine learning model according to the candidate gene set to screen out a prognostic feature gene set and obtain the cancer prognosis model with the prognostic feature gene set as the model feature. Optionally, the prognosis model is based on a linear model, and the linear model is selected from one or more of a linear regression model, a logistic regression model, a Lasso regression model, a ridge regression model, and a linear discriminant analysis model, preferably a Lasso regression model.

7. The method according to claim 6, characterized in that, The prognostic feature gene set includes one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2.

8. A cancer prognosis model, characterized in that, The prognosis model is constructed according to the method for constructing the cancer prognosis model according to any one of claims 1 to 7, and the model includes: A calculation module for using the prognostic feature gene set as the model feature and calculating a prediction result based on the levels of the genes in the prognostic feature gene set in a biological sample derived from a subject; and An indication module for indicating the prognostic risk of the subject according to the prediction result. Where the prognostic feature gene set includes one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2. Where the model feature further includes model coefficients. Optionally, the model coefficients of each gene in the prognostic feature gene set are as follows:

9. A method for indicating the prognosis risk of cancer, characterized in that, The method includes: Using the prognostic feature gene set as a predictor variable and calculating a prediction result based on the expression levels of the genes in the prognostic feature gene set in a biological sample derived from a subject; and Indicating the prognostic risk of the subject according to the prediction result wherein the prognostic characteristic gene set includes one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2.

10. The method according to claim 9, characterized in that, The calculation model of the prognostic risk is: where x i is the expression level of each gene in the prognostic signature gene set, and Coef is the gene model coefficient of each gene in the prognostic signature gene set wherein the gene model coefficients are as follows: Optionally, the expression level of each gene is RPKM (Reads Per Kilobase per Million mapped reads), FPKM (Fragments Per Kilobase of exon model per Million mapped fragments), or TPM (Transcripts Per Million), preferably FPKM.

11. A cancer prognosis system, characterized in that, The system includes: a processor; an input module for inputting the levels of each gene in the prognostic characteristic gene set in a biological sample derived from a subject, wherein the prognostic characteristic gene set includes one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2; a computer-readable medium containing instructions that, when executed by the processor, implement the method for indicating the prognostic risk of intrahepatic cholangiocarcinoma as claimed in claim 9; and an output module for indicating the prognostic risk of the subject.

12. Use of a reagent for detecting a biomarker in the preparation of a product for cancer prognosis, characterized in that, The biomarker includes one or more of the following genes: GPRC5A, FXYD2, LMO7, EPHX1, NEURL3, SMAD5, SLC2A1, NET1, ZNF704, TINAGL1, RAPGEF5, REEP3, PKM, OSMR, and CCT2, Optionally, the reagent is selected from: a probe that specifically recognizes the biomarker; a primer that specifically recognizes the biomarker; or an antibody or ligand that specifically binds to the biomarker.