Construction method of gene model for identifying new subtype of hepatocellular carcinoma, and use thereof
By using whole transcriptome sequencing and cluster analysis, the HCC_1 subtype, which is most closely related to intrahepatic cholangiocarcinoma, was identified. Based on differentially expressed genes, a gene model was constructed, which solved the problems of subtype identification and prognostic assessment of hepatocellular carcinoma and enabled personalized treatment and drug response prediction.
Patent Information
- Application Number
- PCT/CN2024/090341
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-28
- Publication Date
- 2025-11-06
AI Technical Summary
Current technologies struggle to effectively identify and differentiate between subtypes of hepatocellular carcinoma and intrahepatic cholangiocarcinoma, leading to unsatisfactory treatment outcomes, especially for patients with advanced, unresectable liver cancer.
Through whole transcriptome sequencing and cluster analysis, the HCC_1 subtype, which is most closely related to intrahepatic cholangiocarcinoma, was identified. A gene model was constructed based on differentially expressed genes, and LASSO regression analysis was used to determine the expression characteristics of key genes. A scoring model for HCC_1 hepatocellular carcinoma was then constructed.
It enables accurate subtype identification and prognostic assessment of hepatocellular carcinoma patients, guides personalized treatment plans, improves treatment outcomes, and predicts drug responsiveness.
Smart Images

Figure CN2024090341_06112025_PF_FP_ABST
Abstract
Description
A gene model construction method for identifying a new subtype of hepatocellular carcinoma and application thereof TECHNICAL FIELD
[0001] The present application belongs to the field of biomedical technology, and particularly relates to a gene model construction method for identifying a new subtype of hepatocellular carcinoma and application thereof. BACKGROUND
[0002] Hepatocellular carcinoma is one of the most common ten malignant tumors worldwide. There are about 500,000 new cases every year in the world, of which 85% are hepatocellular carcinoma. It can be divided into hepatocellular carcinoma (HCC), intrahepatic cholangiocarcinoma (ICC) and mixed hepatocellular carcinoma (HCC / ICC) three types. There is no clear boundary between HCC and ICC, and some subtypes have similar biological characteristics and immune infiltration status. However, compared with HCC, the prognosis of ICC is worse. Therefore, for patients with advanced unresectable hepatocellular carcinoma, the treatment effect is not ideal. Identifying different molecular subtypes is crucial for the best treatment plan.
[0003] SUMMARY
[0004] The present application aims at the deficiencies of the prior art, and provides a gene model construction method for identifying a new subtype of hepatocellular carcinoma and application thereof.
[0005] The scheme adopted by the present application is as follows:
[0006] A gene model construction method for identifying a new subtype of hepatocellular carcinoma, comprising the following steps:
[0007] Performing whole transcriptome sequencing on the collected tumor and adjacent non-tumor tissue samples of hepatocellular carcinoma patients to obtain transcriptome data; calculating the gene expression level of each patient sample based on the obtained transcriptome data; clustering the gene expression levels of all patient samples to divide hepatocellular carcinoma into several subtypes;
[0008] Mapping the several subtypes obtained by division to TCGA LIHC samples respectively, and selecting the subtype which cannot be completely mapped to the TCGA LIHC samples as HCC_1 subtype;
[0009] Obtaining the differential genes of the HCC_1 subtype to other subtype samples, and constructing a gene model for identifying a new subtype of hepatocellular carcinoma based on the differential genes.
[0010] Further, the differential genes are specifically the differential genes of the HCC_1 subtype to other subtype samples with an expression ratio > 4, FDR < 0.01, HCC_1 subtype expression > 1, and other subtype sample expression < 1.
[0011] Further, a gene model for identifying a new subtype of hepatocellular carcinoma is constructed based on the differential genes by using LASSO regression analysis.
[0012] Further, the gene model is specifically:
[0013] HCC_1 hepatocellular carcinoma score = (0.603 x MMP1 expression level) + (0.028 x GCNT3 expression level) + (0.039 x CBX2 expression level) + (0.095 x CSF3R expression level) + (0.255 x TYRO3 expression level) + (0.076 x FZD7 expression level) + (0.437 x GAL3ST4 expression level) + (0.042 x FBXO41 expression level) + (0.065 x LRP12 expression level) + (0.042 x MTHFD2 expression level) + (0.055 x CD38 expression level) + (0.203 x CRACR2B expression level) + (0.078 x NEIL3 expression level) + (0.039 x ELAPOR2 expression level) + (0.035 x LPAR2 expression level) + (0.091 x SLC25A36 expression level) + (0.121 x TRIP13 expression level) + (0.044 x DZIP1L expression level) + (0.147 x NXPE3 expression level) + (0.2 x TAS1R3 expression level) + (0.164 x MX2 expression level) + (0.149 x ARHGAP39 expression level) + (0.184 x PGM2L1 expression level) + (0.339 x ARMC9 expression level) + (0.039 x LDLRAD3 expression level) + (0.504 x CENPE expression level) + (0.252 x FANCE expression level) + (0.08 x KIF18A expression level) + (0.305 x ARMCX1 expression level) - 9.1623.
[0014] A device for constructing a gene model for identifying a new subtype of hepatocellular carcinoma, comprising:
[0015] A data preprocessing module is configured to perform whole transcriptome sequencing on tumor and adjacent non-tumor tissue samples of hepatocellular carcinoma patients to obtain transcriptome data, and calculate gene expression levels of each patient sample based on the obtained transcriptome data.
[0016] A typing module is configured to cluster the gene expression levels of all patient samples and divide the hepatocellular carcinoma into several subtypes.
[0017] A screening module is configured to map the several subtypes obtained by division to TCGA LIHC samples respectively, and select a subtype that cannot be completely mapped to the TCGA LIHC samples as an HCC_1 subtype.
[0018] a model construction module, configured to acquire differential genes of HCC_1 subtype to other subtype samples, and construct a gene model for identifying new subtypes of hepatocellular carcinoma based on the differential genes.
[0019] Further, the differential genes are specifically differential genes of HCC_1 subtype to other subtype samples with an expression ratio of HCC_1 subtype to other subtype samples > 4, FDR < 0.01, expression of HCC_1 subtype > 1, and expression of other subtype samples < 1.
[0020] Further, in the model construction module, LASSO regression analysis is adopted to construct the gene model for identifying new subtypes of hepatocellular carcinoma based on the differential genes.
[0021] The application further discloses an application of the gene model constructed by the gene model construction method for identifying new subtypes of hepatocellular carcinoma in constructing a prognosis system for patients with hepatocellular carcinoma, and the prognosis system for patients with hepatocellular carcinoma comprises:
[0022] a data acquisition module, configured to acquire differential gene expression data of a hepatocellular carcinoma patient to be tested;
[0023] a prognosis judgment module, configured to input the acquired differential gene expression data of the hepatocellular carcinoma patient to be tested into the gene model constructed by the gene model construction method for identifying new subtypes of hepatocellular carcinoma, to obtain a HCC_1 type hepatocellular carcinoma score, and to judge the prognosis of the hepatocellular carcinoma patient based on the HCC_1 type hepatocellular carcinoma score.
[0024] An electronic device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the gene model construction method for identifying new subtypes of hepatocellular carcinoma when executing the computer program.
[0025] A storage medium comprising computer executable instructions, which, when executed by a computer processor, implement the gene model construction method for identifying new subtypes of hepatocellular carcinoma.
[0026] The beneficial effects of the present application are: the construction method of the present application is based on the fact that there is no clear boundary between HCC and ICC, according to the transcriptome sequencing data of hepatocellular carcinoma patients, first, the HCC samples are divided into different subtypes, the subtype closest to ICC is selected and named as HCC_1 subtype, and then a gene model is constructed based on the differential genes of the HCC_1 subtype, the gene model constructed by the method can identify HCC_1 subtype hepatocellular carcinoma patients, and the HCC_1 subtype is closest to ICC, has similar biological behavior and immune infiltration state, and finally has similar prognosis, so the gene model according to the present application can evaluate the prognosis of hepatocellular carcinoma patients and guide clinicians to provide more active treatment plan for patients with HCC subtype, and can also predict drug reactivity to guide treatment; the higher the HCC_1 type hepatocellular carcinoma score output by the gene model of the present application, the greater the probability of belonging to the HCC_1 subtype hepatocellular carcinoma, and the worse the prognosis. BRIEF DESCRIPTION OF DRAWINGS
[0027] The present application will be further described below in combination with the drawings and examples;
[0028] Fig. 1 is a result graph of the reactivity of HCC_1 and other subtype patients to immune checkpoint therapy according to the present application;
[0029] Fig. 2 is a result graph of clustering HCC tumors into four subtypes: HCC_1, HCC_2, HCC_3 and HCC_4 by PCA according to the present application;
[0030] Fig. 3 is a result graph of mapping the clustering results of the present application to TCGA HCC clusters using SubMap;
[0031] Fig. 4 is a result graph of HCC_1 samples and ICC samples showing transcriptional similarity according to the present application;
[0032] Fig. 5 is a metabolic profile of different subtypes of HCC and ICC;
[0033] Fig. 6 is a result graph of immune cell infiltration in different subtypes of HCC and ICC samples; Fig. 6A is a box plot of CD8-positive T cell infiltration scores in different subtypes of HCC and ICC samples; Fig. 6B is a box plot of regulatory T cell infiltration scores in different subtypes of HCC and ICC samples; Fig. 6C is a heat map of CD8-positive exhausted T cell marker gene expression in different subtypes of HCC and ICC samples;
[0034] Fig. 7 is a co-expression gene module of different subtypes of HCC and ICC samples according to the present application;
[0035] Fig. 8 is a result graph of verification in TCGA and ICGC data according to the present application;
[0036] Figure 9 is a survival curve of patients corresponding to HCC_1 subtype and other subtypes of samples obtained by typing of the model of the application in TCGA and ICGC data;
[0037] Figure 10 is a comprehensive single-cell analysis of HCC, HCC_1 and ICC samples; A in Figure 10 is a UMAP of major cell types in all HCC and ICC samples; B in Figure 10 is a ratio heat map of the number of observed and expected NK / T cell subtypes in HCC, HCC_1 and ICC samples; C in Figure 10 is a CD8-positive T cell pseudo-time result graph. DETAILED DESCRIPTION
[0038] The technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0039] The inclusion and exclusion criteria for liver cancer patients (including hepatocellular carcinoma (HCC) and intrahepatic cholangiocarcinoma (ICC)) in the following examples are as follows:
[0040] (1) No other cancer treatment before surgery;
[0041] (2) No history of other malignant tumors;
[0042] (3) Complete clinical and pathological data and follow-up information.
[0043] Embodiment 1: Construction method of gene model for identifying new subtype of hepatocellular carcinoma
[0044] The construction method of the gene model for identifying the new subtype of hepatocellular carcinoma provided by the present application comprises the following steps:
[0045] (1) Whole transcriptome sequencing is performed on the tumor and adjacent non-tumor tissue samples of the collected hepatocellular carcinoma patients to obtain transcriptome data; the gene expression level of each patient sample is calculated based on the obtained transcriptome data; the gene expression levels of all patient samples are clustered, and the hepatocellular carcinoma is divided into several subtypes.
[0046] In this embodiment, whole transcriptome sequencing was performed on tumor and adjacent non-tumor tissue samples of 293 patients with two types of liver cancer, transcriptome data was obtained, and gene expression levels of each sample were calculated; among them, the intrahepatic cholangiocarcinoma patients were used for subsequent verification; then the K-means clustering method was used to cluster the HCC tumor into 4 subtypes: HCC_1, HCC_2, HCC_3 and HCC_4; principal component analysis was used for the 4 subtypes, as shown in Figure 2, and there are obvious differences among the 4 HCC subtypes.
[0047] (2) Map the several subtypes obtained by division to the TCGA HCC cluster respectively, and select the subtype which cannot be completely mapped to the TCGA HCC cluster as the HCC_1 subtype.
[0048] In this embodiment, the SubMap algorithm is used to map the 4 HCC subtypes to the TCGA HCC cluster respectively, as shown in Figure 3, the HCC_2, HCC_3 and HCC_4 subtypes can be completely mapped to the iC3, iC2 and iC1 (P=0.012) in the TCGA HCC cluster respectively, and the HCC_1 subtype can only be partially mapped to the iC1 (P=0.048) in the TCGA HCC cluster, the HCC_1 subtype is the subtype selected by the present application which cannot be completely mapped to the TCGA HCC cluster, and is named as the HCC_1 subtype, which is consistent with the naming result of step (1).
[0049] The transcriptional similarity between each HCC subtype and ICC was analyzed, as shown in Figure 4, it was found that the HCC_1 subtype sample had high transcriptome similarity with the ICC sample and low similarity with other HCC subtype samples, and most of the genes highly expressed in the HCC_1 subtype sample and the ICC sample were enriched in ECM receptor interaction, PI3K-Akt signaling pathway, carbon metabolism and Mucin type O-glycan biosynthesis pathways. Liver is one of the important metabolic organs in mammals, and many studies have reported that there is a very serious metabolic abnormality in liver cancer tissue. In order to analyze the metabolic pathway changes in the HCC and ICC samples, the present application uses the single-sample gene set enrichment analysis (ssGSEA) algorithm to calculate the enrichment score of the metabolic related pathways in the KEGG database for each sample, and compares the metabolic profiles of different liver cancer subtypes, as shown in Figure 5, it is found that HCC_1 and ICC samples show significant metabolic commonality and are different from other HCC subtype samples. HCC_1 and ICC samples both show up-regulation of polysaccharide biosynthesis and down-regulation of fatty acid and glucose metabolism.
[0050] Further, immune cell infiltration also has important influence on the progress of tumor, therefore the present application uses CIBERSORTx algorithm to calculate the percentage of infiltrated immune cells in each tumor sample. As shown in Fig. 6A and Fig. 6B, it is found that CD8+ T cells have higher infiltration in ICC, HCC_1 and HCC_4 samples, and regulatory T cells significantly infiltrate in HCC_1. It is also found that CD8+ exhausted T cells significantly infiltrate in HCC_1, HCC_4 and ICC samples, and T cell exhaustion marker genes such as PDCD1, CTLA4, HAVCR2, ENTPD1, TIGIT, TNFRSF9, CD27, MYO7A, CXCL13, LAYN, PHLDA1 and SNAP47 are significantly highly expressed in these samples, as shown in Fig. 6C.
[0051] The above results all show that the present application selects the subtype in the HCC subtypes that is most different from the TCGA HCC cluster, i.e. the subtype closest to ICC, named HCC_1 subtype, which has similar biological behavior and immune infiltration state, and finally can exhibit similar prognostic characteristics to ICC.
[0052] (3) Obtain the differential genes of HCC_1 subtype to other subtype samples, and construct a gene model for identifying new subtypes of hepatocellular carcinoma based on the differential genes.
[0053] The embodiment analyzes the expression pattern of HCC_1 subtype by differential gene expression, determines HCC_1 as a separate subtype, further calculates the gene expression similarity of all tumor samples using WGCNA algorithm, selects the differential genes of HCC_1 subtype to other subtype samples with the expression ratio of HCC_1 subtype to other subtype samples>4, FDR<0.01, the expression of HCC_1 subtype>1, and the expression of other subtype samples<1; as shown in Figure 7, it is found that the differential genes can be divided into 20 gene modules in total, most of the co-expression gene modules have similar expression profiles in HCC_1 and ICC samples, which are different from other HCC subtype samples. Then, LASSO algorithm is used for analysis, based on R language glmnet package, 1000 times of Cox LASSO regression iteration and 10 times of cross validation, the differential genes are reduced to 29 genes related to HCC prognosis, including: MMP1, GCNT3, CBX2, CSF3R, TYRO3, FZD7, GAL3ST4, FBXO41, LRP12, MTHFD2, CD38, CRACR2B, NEIL3, ELAPOR2, LPAR2, SLC25A36, TRIP13, DZIP1L, NXPE3, TAS1R3, MX2, ARHGAP39, PGM2L1, ARMC9, LDLRAD3, CENPE, FANCE, KIF18A and ARMCX1, and the 29 genes are used as expression characteristics to distinguish this different HCC_1 subtype, and a gene model, i.e., HCC_1 hepatocellular carcinoma typing model, is constructed, as follows:
[0054] HCC_1 hepatocellular carcinoma score = ∑coef i,β × expression i + intercept
[0055] In the formula, i is the serial number of the gene, expression i represents the expression level of the i th gene, coef i,β is the Lasso regression coefficient of the corresponding gene i, and intercept refers to a constant term, both of which are determined by actual conditions. In the embodiment, the coefficients obtained based on the data are shown in Table 1, and the model is specifically represented as follows:
[0056] HCC_1 type hepatocellular carcinoma score = (0.603 x MMP1 expression level) + (0.028 x GCNT3 expression level) + (0.039 x CBX2 expression level) + (0.095 x CSF3R expression level) + (0.255 x TYRO3 expression level) + (0.076 x FZD7 expression level) + (0.437 x GAL3ST4 expression level) + (0.042 x FBX041 expression level) + (0.065 x LRP12 expression level) + (0.042 x MTHFD2 expression level) + (0.055 x CD38 expression level) + (0.203 x CRACR2B expression level) + (0.078 x NEIL3 expression level) + (0.039 x ELAPOR2 expression level) + (0.035 x LPAR2 expression level) + (0.091 x SLC25A36 expression level) + (0.121 x TRIP13 expression level) + (0.044 x DZIP1L expression level) + (0.147 x NXPE3 expression level) + (0.2 x TAS1R3 expression level) + (0.164 x MX2 expression level) + (0.149 x ARHGAP39 expression level) + (0.184 x PGM2L1 expression level) + (0.339 x ARMC9 expression level) + (0.039 x LDLRAD3 expression level) + (0.504 x CENPE expression level) + (0.252 x FANCE expression level) + (0.08 x KIF18A expression level) + (0.305 x ARMCX1 expression level) - 9.1623.
[0057] Table 1 29 typing genes obtained after LASSO regression model
[0058] Based on the constructed gene model, HCC_1 subtype typing can be performed, and further, since HCC_1 subtype has a similar worse prognosis as ICC, the constructed gene model can be further used to judge the prognosis of HCC patients, as follows: obtaining the gene expression levels of MMP1, GCNT3, CBX2, CSF3R, TYRO3, FZD7, GAL3ST4, FBX041, LRP12, MTHFD2, CD38, CRACR2B, NEIL3, ELAPOR2, LPAR2, SLC25A36, TRIP13, DZIP1L, NXPE3, TAS1R3, MX2, ARHGAP39, PGM2L1, ARMC9, LDLRAD3, CENPE, FANCE, KIF18A and ARMCX1 of the HCC patient, calculating the HCC_1 type hepatocellular carcinoma score based on the constructed gene model, if the HCC_1 type hepatocellular carcinoma score > 0, it means that the hepatocellular carcinoma patient belongs to HCC_1 subtype with a higher probability and has a worse prognosis.
[0059] The gene model constructed in Example 1 was tested in TCGA (including 153 Asian HCC patients and 161 Caucasian HCC patients from TCGA) and ICGC (LIRI-JP) data. The formula in the model of the application was applied to score samples in TCGA and ICGC data sets by using the nearest neighbor template prediction algorithm, and HCC_1-like samples similar to the subtyping obtained in the data set collected by the application were identified, which showed similar characteristics to HCC_1 subtype samples, as shown in Figure 8. The survival of HCC_1-like samples and other samples in the two data sets was statistically analyzed, and HCC_1-like samples were associated with poorer survival compared to other samples, as shown in Figure 9.
[0060] In addition, 41 tumor samples from hepatocellular carcinoma patients treated with immune checkpoint inhibitors (ICIs) in Shau Yiu Fai Hospital were subjected to RNA sequencing to obtain gene expression level data, and based on the constructed gene model, they were divided into HCC_1 subtype and other HCC, as shown in Figure 1, and it was found that the response rate of HCC_1 subtype patients to immune checkpoint inhibitor treatment was significantly higher than that of other HCC subtypes (relative risk value: 2.738, 95% confidence interval: 1.514-5.465), combined with the above results, to guide clinicians to provide more aggressive treatment options for patients with this HCC subtype.
[0061] Example 2: HCC 1 subtype tumor microenvironment profiling at single cell level
[0062] To accurately dissect the tumor microenvironment (TME) of HCC 1 subtype, the present application selected fresh primary tumor samples from 15 HCC patients (including 11 ordinary HCC patients and 4 HCC_1 patients) and 8 ICC patients for single cell RNA sequencing (scRNA-seq). The present application used cellranger software to align all scRNA-seq reads to the human reference genome (hg38) to generate a cell-gene count matrix. The scRNA-seq expression profile was analyzed using the R language package Seurat. Cells with a minimum of 500 gene expressions, total UMI count <80000 and mitochondrial gene count percentage <20% were retained for subsequent analysis. The UMI count of all cells was normalized by total expression, multiplied by 10 6After log transformation, single-cell data dimensionality reduction and clustering, find variable features using the function, and after the standardization of the variable genes using the ScaleData function, further principal component analysis of the expression profile. Use the FindNeighbors function to cluster cells in the PCA space on the top 20 principal components (PCs). All cells are visualized by uniform manifold approximation and projection algorithm (UMAP). By looking at the up-regulated genes with the smallest P value and the classic marker genes that have been reported, the cell types are determined. The present application collects a large number of classic marker genes specifically expressed for different cell subtypes. In these tumor samples, all cells are divided into 7 main types according to their specific markers: NK / T cells (CD3D, CD3G, CD3E, CD56, CD247), B cells (CD19, MS4A1), plasma cells (IGHG1, IGHG4, IGHAl), myeloid cells (CD14, FCGR3A, HLA-DRA, LYZ, VCAN), epithelial cells (ALB, EPCAM), fibroblasts (ACTA2, COL1A1, COL1A2, LUM, DCN), and endothelial cells (CLDN5, CDH5, VWF), as shown in Figure 10A. In NK / T cells, 14 different subtypes are further determined according to the highly expressed genes of each subgroup (mainly including CD4+ T cells, CD8+ T cells, Tregs, NK cells, etc.). In order to quantify the enrichment of cell subtypes in different cancer types, the previously reported method was used to calculate the expected number of subtype cells in different cancer types. The formula is as follows:
[0063] In the formula, Observed is the actual number of subtype cells in a specific cancer type, and Expected is the expected number of cells obtained by counting the number of subtype cells and the number of specific cancer types using the chi-square test. When Ro / e > 1, the subtype is considered to be enriched in the cancer type. As shown in Figure 10B, consistent with previous conclusions, in the HCC_1 subtype, Tregs and CD8+ exhausted T cells are significantly enriched compared to ordinary HCC, while effector CD8+ T cells are significantly reduced. CD4+ and CD8+ naive T cells, mucosa-associated invariant T cells (CD8_MAIT), and NK cells are mainly enriched in ICC. In addition, the R language package monocle is used to predict the possible developmental trajectory of CD8+ T cell subgroups. The pseudo-temporal analysis of CD8+ T cells shows that CD8+ exhausted T cells in HCC_1 are more than in ordinary HCC in the later stage of the developmental trajectory, as shown in Figure 10C. These results indicate that there is an exhausted immune microenvironment in the HCC1 subtype, leading to poor prognosis in HCC 1 type patients.
[0064] Embodiment 3: device for constructing gene model for identifying new subtype of hepatocellular carcinoma, electronic device and storage medium
[0065] Corresponding to the foregoing embodiments of the method for constructing a gene model for identifying a new subtype of hepatocellular carcinoma, the present application also provides embodiments of a device for constructing a gene model for identifying a new subtype of hepatocellular carcinoma. The device for constructing a gene model for identifying a new subtype of hepatocellular carcinoma comprises:
[0066] a data preprocessing module configured to perform whole transcriptome sequencing on tumor and adjacent non-tumor tissue samples of hepatocellular carcinoma patients to obtain transcriptome data, and calculate gene expression levels of each patient sample based on the obtained transcriptome data;
[0067] a typing module configured to cluster the gene expression levels of all patient samples, and divide hepatocellular carcinoma into several subtypes;
[0068] a screening module configured to map the several subtypes obtained by division to TCGA HCC clusters respectively, and select a subtype that cannot be completely mapped to the TCGA HCC clusters as an HCC_1 subtype;
[0069] a model construction module configured to obtain differential genes of the HCC_1 subtype with respect to other subtype samples, and construct a gene model for identifying a new subtype of hepatocellular carcinoma based on the differential genes.
[0070] The device for constructing a gene model for identifying a new subtype of hepatocellular carcinoma according to the embodiments of the present application can be applied to any device with data processing capability, which can be a device or apparatus such as a computer.
[0071] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts are described with reference to the parts of the method embodiments. The device embodiments described above are only illustrative, and the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present application according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0072] The present application also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for constructing a gene model for identifying a new subtype of hepatocellular carcinoma when executing the computer program.
[0073] The electronic device is a logically device, and is formed by reading the corresponding computer program instructions in the non-volatile memory into the memory through the processor of any data processing device where the electronic device is located to run; from the hardware level, the electronic device includes the processor, the memory, the network interface, and the non-volatile memory, in addition to this, the electronic device can also include other hardware according to the actual function of the any data processing device, which will not be described here.
[0074] The implementation process of the functions and roles of each unit in the above electronic device is specifically shown in the implementation process of the corresponding steps in the above method, which will not be described here.
[0075] The embodiment of the present application also provides a storage medium containing computer executable instructions, which realize the gene model construction method for identifying new subtypes of hepatocellular carcinoma of liver cells when executed by a computer processor.
[0076] The computer readable storage medium can be an internal storage unit of any data processing device, such as a hard disk or a memory. The computer readable storage medium can also be any data processing device, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. Further, the computer readable storage medium can include both the internal storage unit of any data processing device and the external storage device. The computer readable storage medium is used to store the computer program and other programs and data required by the any data processing device, and can also be used to temporarily store data that has been output or will be output.
[0077] Embodiment 4: Application of the gene model for identifying new subtypes of hepatocellular carcinoma of liver cells
[0078] The gene model for identifying new subtypes of hepatocellular carcinoma of liver cells can be used to construct a hepatocellular carcinoma patient prognosis system, and the hepatocellular carcinoma patient prognosis system includes:
[0079] A data acquisition module is configured to acquire differential gene expression data of a hepatocellular carcinoma patient to be tested.
[0080] A prognosis judgment module is configured to input the acquired differential gene expression data of the hepatocellular carcinoma patient to be tested into the constructed gene model to obtain a HCC_1 hepatocellular carcinoma score, and judge the prognosis of the hepatocellular carcinoma patient based on the HCC_1 hepatocellular carcinoma score: if the HCC_1 hepatocellular carcinoma score is greater than 0, it indicates that the hepatocellular carcinoma patient belongs to the HCC_1 subtype with a higher probability and has a poorer prognosis.
[0081] Obviously, the above embodiments are merely exemplary but not restrictive. Based on the above description, those skilled in the art can make other different forms of changes or modifications. All the embodiments are not required to be exhaustive. The obvious changes or modifications derived therefrom are still within the protection scope of the present application.
Claims
1. A method for constructing a gene model for identifying a new subtype of hepatocellular carcinoma, characterized in that, Comprising the following steps: Whole transcriptome sequencing is performed on the collected tumor and adjacent non-tumor tissue samples of hepatocellular carcinoma patients to obtain transcriptome data; the gene expression level of each patient sample is calculated based on the obtained transcriptome data; the gene expression levels of all patient samples are clustered, and hepatocellular carcinoma is divided into several subtypes; The several subtypes obtained by division are respectively mapped to the TCGA LIHC samples, and the subtype that cannot be completely mapped to the TCGA LIHC samples is named as HCC_1 subtype; Obtain the differential genes of HCC_1 subtype to other subtype samples, and construct a gene model for identifying new subtypes of hepatocellular carcinoma based on the differential genes.
2. The method of claim 1, wherein, The differential genes of HCC_1 subtype to other subtype samples are specifically: HCC_1 subtype to other subtype sample expression ratio>4, FDR<0.01, HCC_1 subtype expression>1, and other subtype sample expression<1 constitute the differential genes of HCC_1 subtype to other subtype samples.
3. The method of claim 1, wherein, LASSO regression analysis is used to construct a gene model for identifying new subtypes of hepatocellular carcinoma based on the differential genes.
4. The method of claim 3, wherein, The gene model is specifically: HCC_1 type hepatocellular carcinoma score=(0.603*MMP1 expression level)+(0.028*GCNT3 expression level)+(0.039*CBX2 expression level)+(0.095*CSF3R expression level)+(0.255*TYRO3 expression level)+(0.076*FZD7 expression level)+(0.437*GAL3ST4 expression level)+(0.042*FBXO41 expression level)+(0.065*LRP12 expression level)+(0.042*MTHFD2 expression level)+(0.055*CD38 expression level)+(0.203*CRACR2B expression level)+(0.078*NEIL3 expression level)+(0.039*ELAPOR2 expression level)+(0.035*LPAR2 expression level)+(0.091*SLC25A36 expression level)+(0.121*TRIP13 expression level)+(0.044*DZIP1L expression level)+(0.147*NXPE3 expression level)+(0.2*TAS1R3 expression level)+(0.164*MX2 expression level)+(0.149*ARHGAP39 expression level)+(0.184*PGM2L1 expression level)+(0.339*ARMC9 expression level)+(0.039*LDLRAD3 expression level)+(0.504*CENPE expression level)+(0.252*FANCE expression level)+(0.08*KIF18A expression level)+(0.305*ARMCX1 expression level)-9.1623.
5. A device for constructing a gene model to identify a new subtype of hepatocellular carcinoma, characterized in that, Comprising: A data preprocessing module for performing whole transcriptome sequencing on the collected tumor and adjacent non-tumor tissue samples of hepatocellular carcinoma patients to obtain transcriptome data; calculating gene expression levels of each patient sample based on the obtained transcriptome data; a typing module for clustering the gene expression levels of all patient samples, and dividing hepatocellular carcinoma into several subtypes; a screening module for mapping the several subtypes obtained by division to TCGA LIHC samples respectively, and selecting a subtype that cannot be completely mapped to the TCGA LIHC samples as HCC_1 subtype; a model construction module for obtaining differential genes of the HCC_1 subtype to other subtype samples, and constructing a gene model for identifying new subtypes of hepatocellular carcinoma based on the differential genes.
6. The apparatus of claim 5, wherein, The differential genes are specifically differential genes of the HCC_1 subtype to other subtype samples with an expression ratio > 4, FDR < 0.01, HCC_1 subtype expression > 1, and other subtype sample expression < 1.
7. The apparatus of claim 5, wherein In the model construction module, LASSO regression analysis is used to construct a gene model for identifying new subtypes of hepatocellular carcinoma based on the differential genes.
8. The application of the gene model constructed by the method for constructing a gene model for identifying a new subtype of hepatocellular carcinoma according to any one of claims 1-4 in constructing a prognosis system for patients with hepatocellular carcinoma, characterized in that, The hepatocellular carcinoma patient prognosis system comprises: a data acquisition module for acquiring differential gene expression data of a hepatocellular carcinoma patient to be tested; a prognosis judgment module for inputting the acquired differential gene expression data of the hepatocellular carcinoma patient to be tested into a gene model constructed by the method for identifying new subtypes of hepatocellular carcinoma according to any one of claims 1-4, obtaining a HCC_1 type hepatocellular carcinoma score, and judging the prognosis of the hepatocellular carcinoma patient based on the HCC_1 type hepatocellular carcinoma score.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the method for identifying new subtypes of hepatocellular carcinoma according to any one of claims 1-4.
10. A storage medium containing computer executable instructions, which, when executed by a computer processor, realize the method for identifying new subtypes of hepatocellular carcinoma according to any one of claims 1-4.
Citation Information
Patent Citations
Method for constructing hepatocellular carcinoma typing system based on ferroptosis process
CN113192560A
Gene model for judging prognosis of hepatocellular carcinoma patient, construction method and application
CN113539376A
Hepatocellular carcinoma prognosis biomarker and application thereof
CN115807089A
Construction method and application of accurate diagnosis model of intrahepatic cholangiocarcinoma patient prone to copper death
CN117373655A
Methods for the detection and treatment of classes of hepatocellular carcinoma responsive to immunotherapy
US20210277480A1