Gene expression scoring algorithm, tumor diagnosis model and construction method

By developing a gene expression scoring algorithm, the problem of differences between different gene expression data sets is solved, the data sets are effectively integrated, and the accuracy and reliability of the liver cancer diagnosis model are improved.

CN120072051APending Publication Date: 2025-05-30GUANGZHOU RXGENE TECH INC +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510133635.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Due to differences in detection platforms, methods, genes and operations, existing gene expression data sets are difficult to achieve effective integration, which limits the diagnosis and research of liver cancer.

Method used

A gene expression scoring algorithm was developed to achieve integration of different data sets by sorting and scoring the original expression values ​​of all genes in each sample and converting them into relative expression values.

Benefits of technology

It effectively eliminates the background differences between different data sets, improves the quality of the integrated data sets, and supports the construction of a more accurate and reliable liver cancer diagnosis model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072051A_ABST
    Figure CN120072051A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of bioinformatics, particularly relates to a gene expression scoring algorithm, a tumor diagnosis model and a construction method, and aims to realize integration and standardization processing of data sets of different sources. Compared with the prior art, the method has the advantages of being simple in principle, convenient to use, high in operation speed and the like, and differences among different data sets can be eliminated. In the field of tumor research, the problems of high heterogeneity among tumors, one-sided single gene research mode, difficulty in integration of sequencing data and the like exist. Therefore, the invention also comprises a method for constructing a novel tumor diagnosis model by applying the relative score value algorithm, integrating a plurality of tumor data sets and based on a large amount of tumor sample data. The model has the characteristics of high accuracy, easy data expansion and high universality, and can be used for discovering potential tumor diagnosis and treatment targets. Based on the liver cancer data set, the feasibility of the scheme is proved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics, and particularly relates to a gene expression scoring algorithm, a tumor diagnosis model and a construction method thereof. Background Art

[0002] With the development of gene sequencing technology and proteome detection technology, a variety of studies have generated a large number of data sets, covering many gene expression data, including DNA, messenger RNA (mRNA), microRNA (miRNA), small nuclear RNA (snRNA), circular RNA (circRNA), long non-coding RNA (lncRNA), small nucleolar RNA (snoRNA), proteins, etc. The gene expression data in different data sets are inconsistent in terms of sample source, species / race, experimental operation, detection method, instrument platform, etc. Therefore, the gene expression data of different data sets often cannot be directly integrated. Although there are some data integration algorithms (for example, ComBat-seq [1] , sva [2] , Ratio-G [3] etc.), they often cannot effectively eliminate the differential effects between data sets (including differences in detection platforms, methods, genes, operations, and data formats).

[0003] Liver cancer is one of the malignant tumors with relatively high incidence and mortality rates globally, seriously threatening human life and health. An excellent biomarker should have a sensitivity and specificity higher than 90%, but in the early screening and clinical treatment practice of liver cancer, no ideal biomarker has been found yet. Therefore, it is necessary to continuously develop more sensitive liver cancer markers. Serum alpha-fetoprotein (AFP) is the most widely used liver cancer screening indicator in China, but its sensitivity and specificity for early liver cancer are poor, limiting its application value in the early screening and diagnosis of liver cancer. Currently, some protein-based liver cancer biomarkers have been discovered, such as abnormal prothrombin (PIVKA II) and alpha-fetoprotein variant 3 (AFP-L3), which have been proven to improve the diagnostic performance when used in combination with AFP, especially for the detection of early liver cancer. Dickkopf-1 (DKK-1) [4] , Lamininγ2Monomer (LG2m) [5] , ST6GAL1 [6] , FGF19 [7]They have been proven to have diagnostic and prognostic value in liver cancer but have not been widely applied clinically. In recent years, novel liver cancer biomarkers based on liquid biopsy techniques have shown great application prospects, including circulating tumor DNA (ctDNA) and its methylation changes, circulating RNA, exosomes, etc. Many studies have confirmed that changes in specific genes in liver cancer tissues can be reflected in ctDNA and serve as potential specific biomarkers, including ctDNA levels, copy number variations, gene integrity, gene mutations, and DNA methylation alterations [8]. Meta-analyses have shown that ctDNA and its methylation alterations can be used as potential biomarkers for screening HCC, and when combined with AFP, they can significantly improve its diagnostic performance [9].

[0004] Single biomarkers are difficult to reflect the multi-step process of tumor development, and diagnostic models combining multiple biomarkers are beneficial for improving overall diagnostic sensitivity and specificity

[10] . Cheng et al.

[10] used a diagnostic model consisting of a panel of 5 proteins (OPN, GDF0, NSE, TRAP892, and OPG) and showed high diagnostic accuracy in predicting early liver cancer, with a higher AUC value and sensitivity compared to AFP. Wu et al.

[11] developed a comprehensive diagnostic model by combining 3 mutation sites (TERT, TP53, and CTNNB1) and 3 serum biomarkers (AFP, AFP-L151, and PIVKA-II), which showed high accuracy in differentiating liver cancer patients who were negative for AFP, AFP-L3, and PIVKA-II. With the progress of high-throughput sequencing and omics technologies, various novel diagnostic models based on metabolomics

[12] , immune microenvironment

[13] , autoantibodies

[14] , hypoxia

[15] have emerged in recent years, and these research results will hopefully help develop personalized treatment plans for liver cancer patients. There is no doubt that identifying excellent biomarkers will contribute to accelerating the establishment of liver cancer diagnostic models and promoting the application of precision medicine in liver cancer clinical practice.

[0005] The choice of modeling method is a key factor in defining the properties of HCC diagnostic models. Currently, the mainstream of research is to use logistic regression or machine learning algorithms. Generally speaking, the HCC diagnostic model based on logistic regression algorithm has the characteristics of simplicity, rapidity and strong interpretability, is suitable for linearly separable problems, and is easy to adjust. However, its diagnostic performance is relatively limited (AUC = 0.862 - 0.94118)[10,16 - 18], and it is sensitive to outliers. In contrast, the HCC diagnostic model based on machine learning demonstrates stronger predictive ability (AUC = 0.893 - 0.99)[19 - 21], can efficiently process diverse data and complex features, but its high complexity requires professional technical knowledge and powerful computing platform support, and at the same time it also faces the problem of insufficient transparency. Therefore, choosing a suitable modeling method requires a balance between diagnostic accuracy and model complexity.

[0006] To sum up, the research on liver cancer faces the following problems: First, the existing research mostly focuses on single-gene research, which limits the in-depth exploration of intra-tumor and inter-tumor heterogeneity

[22] . Second, although single-cell research has made a lot of progress in the field of immune cells, the single-cell level research on the mechanism of hepatocyte carcinogenesis is still relatively scarce, and the specific role of hepatocytes in the occurrence of liver cancer has not been fully revealed. Finally, due to differences in sampling methods, sequencing platforms, regions and ethnic groups, etc., the existing liver cancer sequencing data is difficult to achieve effective integrated analysis.

[0007] To solve the above problems, we plan to develop a general scoring algorithm to integrate the currently publicly available gene expression data. Moreover, based on a large number of liver cancer sample data, a new type of liver cancer diagnostic model is constructed, which has the characteristics of high accuracy, easy data expansion and strong versatility, and can be used to discover potential liver cancer diagnosis and treatment targets. Summary of the Invention

[0008] In order to integrate the gene expression data of different data sets, the present invention develops an algorithm for converting gene expression data into relative scores, which can convert the original absolute value of gene expression into a relative value, thus simply realizing the integration of a large number of data sets. For the integrated data set, the number of cases is greatly increased, which helps to develop a high-performance and general tumor diagnostic model.

[0009] The present invention is realized in the following way:

[0010] In the first aspect of the present invention, a gene expression scoring algorithm is provided, including the following steps:

[0011] S1 Collect data sets, preprocess the gene expression data, and generate an expression matrix with sample names and molecule / gene names as rows and columns or columns and rows respectively for each sample of each data set;

[0012] S2 sorts the original expression values of all genes in each sample, and calculates the position of each gene expression value in the overall ranking of all genes in the sample, that is, the quantile;

[0013] S3 scores all genes in each sample according to the quantile corresponding to each gene;

[0014] S4 combines the scores of the quantiles of all genes of all samples in different data sets to obtain an integrated data set.

[0015] Since the scoring is performed independently in each data set, the background differences between different data sets are effectively eliminated; the integrated data set can be used for subsequent analysis of differential molecules, signaling pathways, model construction, etc.

[0016] This algorithm processes the original absolute gene expression values by sorting and scoring, and then converts them into relative expression values. The converted gene expression data can be effectively integrated, significantly reducing the differences between different data sets. The integrated data set provides large-scale case data for the tumor diagnosis model, supports the training and verification of the model, thereby improving the accuracy, reliability and generalization ability of the model.

[0017] The algorithm of the embodiment of the present invention can be implemented by a computer program, which can be an R language program, but can also be implemented by programs such as python, C, C++, Java, etc.

[0018] The preprocessing of the gene expression data in step S1 can be raw data, or data after normalization processing, or data processed by log2 or log10.

[0019] The quantile is calculated according to the overall data from 2 to 100 classification units (i.e., the 2nd to 100th quantiles), and the original expressions of each gene in each sample are sorted; preferably, according to the deciles, each gene in each sample is given a score of 1-10 (i.e., the relative expression value).

[0020] The algorithm of the embodiment of the present invention can be applied to gene expression data such as DNA, messenger RNA (mRNA), microRNA (miRNA), small nuclear RNA (snRNA), circular RNA (circRNA), long non-coding RNA (lncRNA), small nucleolar RNA (snoRNA), proteins, etc.

[0021] In a second aspect of the present invention, a tumor diagnosis model is provided, and the tumor diagnosis model is constructed based on the above gene expression scoring algorithm.

[0022] The tumor can be any known tumor, including but not limited to liver cancer, lung cancer, gastric cancer, colon cancer, cervical cancer, breast cancer, esophageal cancer, kidney cancer, pancreatic cancer, leukemia, brain tumor, osteosarcoma, lymphoma, etc.

[0023] Samples of the tumor diagnosis model all include a tumor group and corresponding control groups, where the control groups include non-cancer control groups, metastasis, recurrence, or adjacent to cancer, etc.

[0024] In a third aspect of the present invention, a method for constructing the tumor diagnosis model is provided, including the steps:

[0025] Preprocess the gene expression data of all samples to obtain an expression matrix with sample names and gene names defined as row names or column names respectively, and in a column or row adjacent to the sample names, set the column or row name as diagnosis to record the disease status of the samples;

[0026] Set a part of the samples (1 or more independent data sets) as the training set, and set another part of the samples (1 or more independent data sets) as the validation set, and there is at least one validation set;

[0027] On the training set, based on the gene list of interest, screen genes with high diagnostic performance; the screening criteria for the genes with high diagnostic performance are: the correct rate of diagnosing tumors in the original data set is greater than a specified value; for example, 70%, which is set according to the actual situation. The gene list of interest can be differentially expressed genes, prognosis-related genes, inflammatory genes, metabolic genes, characteristic genes of tumor-specific cell subsets, etc., and can be one or more of them.

[0028] Relatively score and integrate the genes with high diagnostic performance in the training set;

[0029] Relatively score and integrate the data sets of the validation set;

[0030] Compare the area under the receiver operating characteristic curve AUROC in the training set with the AUROC in the validation set, judge the ability to diagnose tumors across data sets, and integrate the results of the validation set;

[0031] In an independent integrated validation set, construct a tumor diagnosis model through an algorithm. The algorithms used include but are not limited to: logistic regression algorithm, machine learning algorithm, deep learning algorithm, artificial intelligence large model algorithm, etc. It can be one or more of them.

[0032] The present invention has the following advantages compared with the prior art:

[0033] 1. Simple principle, easy to use, fast running speed, and easy data expansion.

[0034] 2. Since numerical conversion is performed within a single data set, differences between different data sets can be eliminated.

[0035] 3. The constructed disease diagnosis model has characteristics such as high accuracy, reliability, and generalization ability. Brief Description of the Drawings

[0036] Figure 1 Revealed the main cell types and overall transcriptome profiles of HCC based on snRNA-seq; (A) Detection and analysis workflow of HCC; (B) UMAP visualization of the main cell types based on snRNA-seq data, with colors representing different cell types. Samples were taken from liver tissues of HCC patients who underwent surgical resection, including tumor tissues and adjacent tissues; (C) Heatmap visualization of marker genes of the main cell types, with colors representing gene expression levels; (D) Sample correlation between different tissues based on bRNA-seq data; (E) DEGs between different tissues based on bRNA-seq data, with green representing downregulated genes and red representing upregulated genes; (F) KEGG enrichment analysis of co-downregulated and co-upregulated DEGs.

[0037] Figure 2 Revealed the genomic evolution of the pathogenesis of HCC hepatocytes; (A) Identified 10 hepatocyte subtypes from snRNA-seq data, and the circular diagram shows the tissue origin of each hepatocyte subtype; (B) Cell cycle status of hepatocyte subtypes; (C) RNA velocity analysis of hepatocyte subtypes, where the direction and length of the arrows represent the change trend and rate of the unspliced-spliced mRNA ratio; (D) Copy number variations (CNVs) in hepatocyte subtypes, with red representing copy number increase and blue representing copy number loss; (E) CNV regions in Hep1-chr12, Hep2-chr3, Hep9-chr6, and Hep9-chr11, with regions with CNV gain highlighted in red and regions with CNV loss highlighted in blue, circular symbols represent differential genes, and rectangular symbols represent diagnostic genes in this study; colors represent the difference in fold change; (F) CNV scores of 10 hepatocyte subtypes relative to endothelial cells, Note: *p<0.05; ****p<0.0001; (G) Correlation between the proportion of Hep9 in tumor tissues and the serum AFP level of HCC patients; (H) Correlation between the proportion of Hep0 / 9 and the proportion of tregs in the TCGA cohort; (I) Correlation between the proportion of Hep0 / 9 and patient prognosis in the TCGA cohort.

[0038] Figure 3Decoding the molecular pathways of HCC hepatocytes using snRNA-seq; (A) Multi-pathway enrichment analysis of upregulated genes in Hep1-9; (B) Subversion map of the top 20 DEGs in Hep1-9, and the bubble plot in the upper right corner shows the expression levels of these genes in the corresponding hepatocyte subtypes; For the marker genes of the Hep0 group, with Hep2 and Hep9 as controls, (C) KEGG analysis (up / down) was performed, and the asterisks indicate the pathways shared by the two groups; (D) GSEA analysis compared Hep0 with. Hep2 subset (upper figure) and Hep0 vs; Hep9 subset (bottom), the first circle represents the enriched pathways, with colors indicating secondary classification, the second circle represents the number of enriched genes and their p-values, the third circle describes the proportion and number of upregulated / downregulated genes, and the fourth circle represents the enrichment score of the pathway, with the outward direction indicating upregulation and the inward direction indicating downregulation.

[0039] Figure 4 Exploring the carcinogenic landscape of hepatocytes by integrating bRNA-seq with snRNA-seq; (A) Venn analysis of downregulated (left) and downregulated (right), and the asterisks indicate the DEGs included in QDM; (B-C) Metabolic flux analysis derived from snRNA-seq (B) and bRNA-seq (C) data using METAFlux; (D) Identifying differentially expressed regulatory factors in different tissues from bRNA-seq data; (E) Identifying differentially expressed regulatory factors in hepatocyte subtypes from bRNA-seq data; (F) Identifying transcription factors targeting PDE7B using bRNA-seq and snRNA-seq data, and the x-axis represents the activity scores of the regulators.

[0040] Figure 5 Based on the research results of snRNA-seq, a quantile distribution model was constructed based on the HCC cohort; (A) Workflow for constructing the quantile distribution model (QDM); (B) AUC values of different genes for diagnosing liver cancer tissues on 10 datasets; (C) Regression coefficients of each gene in QDM. In training set 1, (D) ROC curve generated using QDM for diagnosing HCC; (E) ROC curve generated using QDM in validation set 1.

[0041] Figure 6Investigation of the biological and clinical significance of the QDM model; (A) Evaluate the RNA expression levels of PDE7B, GHR, GPC3, and SLCO1B3 in the bRNA-seq data of the SMU cohort, and compare them using a paired t-test. Note: nsp>0.05; *p<0.05; **p<0.01; ***p<0.001; ****p<0.0001; (B) Evaluate the RNA expression levels of PDE7B, GHR, GPC3, and SLCO1B3 using the t-test method. The number of samples is indicated above each dataset. Note: nsp>0.05; *p<0.05; **p<0.01; ***p<0.001; ****p<0.0001; (C-D) Detect the protein expression levels of PDE7B (C) and GHR (D) using the HPA database; (E) Correlation between the QDM score and serum AFP and AST levels in the SMU cohort; (F) Correlation between IGF2 and GHR and serum AFP levels in the SMU cohort; (G) Evaluate the tumor microenvironment in the bRNA-seq data from the SMU cohort, and evaluate the immune score, stromal score, and tumor purity using the evaluation method; (H) Correlation between PDE7B and tumor purity in the SMU cohort.

[0042] Figure 7 Classification using the characteristic genes identified in snRNA-seq; (A) Workflow for HCC classification using the characteristic genes in the snRNA-seq data; (B) Mutation frequency of genes in training set 2; (C-D) Kaplan-Meier survival analysis of training set 2 (C) and validation set 2 (D); (E) Immune cell composition in validation group 2; (F) TIDE analysis of training set 2; (G) Enrichment analysis of deg in Class_1, Class_2, and Class_3; (H) Identify sensitive drugs for Class_1, Class_2, and Class_3 by integrating tumor prediction and CMap analysis. Detailed implementation manner

[0043] The following further describes the specific implementation manner of the present invention in detail in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but do not limit the scope of the present invention.

[0044] Example 1 Gene expression scoring algorithm

[0045] This example provides a gene expression scoring algorithm, which includes the following steps:

[0046] S1 Collect the dataset, preprocess the gene expression data, generate an expression matrix with sample names as row names and gene names as column names, and set the column name of the first column as diagnosis to record the disease status (diseased / non-diseased) of the samples.

[0047] S2 In each dataset, sort all the original expression values of each gene and calculate the deciles of each gene's expression values.

[0048] S2 Score the original expressions of all genes in each sample according to the deciles (1-10 points).

[0049] S2 Directly merge the decile scores of all genes of all samples from different datasets to obtain an integrated dataset.

[0050] Since the scoring is performed independently in each dataset, the background differences between different datasets are effectively eliminated. The integrated dataset can be used for subsequent analysis of differential genes, signaling pathways, model construction, etc.

[0051] Example 2 Construction of a tumor diagnosis model

[0052] In this example, based on the decile scoring method of Example 1, data from 29 public HCC (hepatocellular carcinoma) datasets were integrated to establish a decile distribution tumor diagnosis model (QDM) with high diagnostic accuracy. The specific steps are as follows:

[0053] Download the liver cancer datasets from GEO (Gene Expression Omnibus), including the training sets (GSE25097, GSE36376, GSE63898, GSE39791, GSE114564, GSE76427, GSE77314, GSE184733) and the validation sets (GSE14520, GSE102079, GSE144269, GSE89377, GSE17856, GSE45267, GSE36411, GSE57957, GSE6764, GSE124535, GSE105130, GSE41804, GSE17548, GSE60502, GSE135631, GSE84402, GSE29721, GSE22058). At the same time, download the relevant sample information.

[0054] Preprocess the gene expression data to generate an expression matrix with sample names as row names and gene names as column names, and set the column name of the first column as diagnosis to record the disease status (diseased / non-diseased) of the samples.

[0055] Based on the gene list of interest, genes with high diagnostic performance were screened in each dataset of the training set.

[0056] The datasets of the training set were relatively scored and integrated.

[0057] In the integrated training set, diagnostic genes were further screened to construct a logistic regression diagnostic model for liver cancer.

[0058] The datasets of the validation set were relatively scored and integrated.

[0059] In the integrated validation set, the logistic regression diagnostic model for liver cancer was used for prediction.

[0060] The diagnostic AUROC (Area Under the Receiver Operating Characteristic Curve) of this model in the training set was 0.982, and the diagnostic AUROC in the validation set was 0.965, showing excellent cross-dataset ability to diagnose liver cancer.

[0061] And this model was further used to screen potential biomarkers, revealing that PDE7B is a key gene, and its inhibition promotes HCC progression. Guided by Hep0 / 2 / 9 subtype-specific genes, HCC was divided into three categories: metabolic, inflammatory, and stromal, which can be distinguished in terms of gene mutation frequency, survival time, enriched pathways, and immune infiltration. At the same time, three sensitive drugs for the three types of HCC were identified, namely vaabain, teniposide, and TG-101348.

[0062] The specific experiments are as follows:

[0063] Materials and Methods:

[0064] Patient and sample collection

[0065] This clinical study was approved by the Medical Ethics Committee of Zhujiang Hospital, Southern Medical University (SMU), and informed consent was obtained from all participants. All studies were conducted in accordance with the Declaration of Helsinki and the Declaration of Istanbul. The SMU cohort of this study included 77 patients who underwent HCC surgery at Zhujiang Hospital, Southern Medical University from March 2019 to August 2023. A total of 209 pairs of samples, including cancer, adjacent cancer tissues, and normal tissues, were collected and stored in liquid nitrogen for transcriptome analysis. In addition, some tissue samples were fixed in 10% formalin solution and pathologically confirmed to contain HCC.

[0066] Cell mononuclear extraction and snRNA-seq

[0067] The nuclear extraction reagent used in this study was prepared according to the 10x Genomics technical document (https: / / www.10xGenomics.com / support / single-cell-gene-expression / documentation / steps / sample-prep / nucleus-isolation-from-adult-mouse-brain-organization-for-single-cell-RNA-sequencing), and then the method was optimized. Briefly, the liver tissue was minced into small pieces with a surgical blade, 100 μl of cold 0.1X lysis buffer was added, and then gently aspirated. The mixture was incubated on ice for 5 minutes. Next, 1 ml of cold wash buffer was added, mixed well, and the solution was filtered through a 40 μm cell strainer to remove debris and large tissue fragments. The filtrate was centrifuged at 500 rcf for 5 minutes at 4°C, and the supernatant was discarded without disturbing the nuclear pellet. The pellet was resuspended in 1 ml of cold wash buffer, gently mixed, and centrifuged again. This washing step was repeated twice. Finally, the nuclear pellet was resuspended in cold diluted nuclear buffer, and the nuclear concentration was determined using a Countess II FL Automated Cell Counter. The suspension was adjusted to the appropriate concentration, and according to the manufacturer's protocol, a single-cell 3' library was prepared using the 10x Genomics Single Cell 3' Kit (v3), and then sequenced on the Illumina Novaseq X platform.

[0068] Processing of snRNA-seq data

[0069] Preprocess the snRNA-seq data, align it with the human reference genome (GRCh38), and generate the raw gene expression count matrix using the CellRanger (version 6.0.1) pipeline. Use the Scrublet

[12] package with default parameters to calculate and remove doublets. Perform quality control using R software (version 4.0.4) and the Seurat package (version 4.0.5), retaining cells with mitochondrial UMIs < 10%, RNA counts > 200, and RNA features between 200 and 6000. Use functions such as NormalizeData, FindVariableFeatures, ScaleData, and RunPCA in the Seurat package to scale and reduce the dimensionality of the data. The Harmony package further enhances the integration, followed by using the RunUMAP function for non-linear dimensionality reduction and the FindNeighbors and FindClusters functions for clustering analysis. Use the ScType

[13] package for automatic annotation of cell clusters, followed by manual refinement based on established marker genes. The code used in this study is available on the tutorial website (https: / / github.com / suhuanhou / 2024_HCC), which includes detailed documentation for each analysis environment and software package version.

[0070] bRNA program - sequence

[0071] Extract total RNA from tissues using RNAiso Plus reagent (Takara, Japan), and quantify the RNA using a NanoDrop 2000 spectrophotometer (Thermo Scientific, USA). Construct cDNA libraries using the sequed must seq 3’mRNA DEG kit (sequed, China). Sequence the prepared libraries on the Illumina NovaSeq 6000 platform to generate 150bp paired-end reads.

[0072] Analysis of bRNA-seq data

[0073] The raw sequencing reads were quality controlled using trimomatic

[14] , adapters were removed and low-quality data were filtered out to obtain clean reads. The reads were aligned to the human reference genome (GRCh38) using HISAT2

[15] , and the gene expression levels were quantified as read counts based on the alignment results, and then the reads per million mapped reads (RPM) were calculated. DESeq2 was used for read count normalization, and the differential expression analysis applied the thresholds of log2FoldChange > 1.5 and p < 0.05. In addition, the RPM data were used for the analysis using specific algorithms such as ESTIMATE and METAFlux. The detailed procedures are provided on the tutorial website (https: / / github.com / suhuanhou / 2024_HCC).

[0074] Development of the quantile distribution model (QDM) for HCC tissues

[0075] We collected 29 human HCC datasets (total sample number = 4192), including HCC and non-HCC samples. These datasets included 19 microarray datasets and 10 RNA-seq datasets, and the RNA-seq data were converted from FPKM or RPKM to TPM or CPM. A quantile-based scoring algorithm was developed to integrate different datasets of different detection platforms, methods, genes, and operations. Hepatocyte subtypes associated with patient survival were identified in the snRNA-seq data, and their characteristic genes were extracted using machine learning methods. Initially, these characteristic genes were screened in 10 training datasets, resulting in 30 genes with an AUC > 0.7 for HCC diagnosis. A logistic regression model including 15 diagnostic genes was constructed and optimized using the training dataset and validated in an independent validation dataset.

[0076] Cell culture

[0077] The human HCC cell lines used in this study were obtained from Biospecies (Guangzhou, China) and identified by short tandem repeat (STR) profiling to confirm their authenticity. All cell lines were negative for mycoplasma, bacterial, yeast, and fungal contamination tests. Under the optimal growth conditions of 37 °C, in 5% CO 2Cultivate cells in an incubator. Specifically, the Hep G2 and PLC / PRF / 5 cell lines are maintained in MEM medium (Gibco, USA) with 10% fetal bovine serum (FBS, Gibco, USA). MHCC-97H cells are cultured in DMEM medium (Gibco, USA) containing 10% FBS, while SNU-387 cells are cultured in RPMI-1640 medium (Gibco, USA) supplemented with 10% FBS, 1% L-glutamine (Biospecies, Guangzhou, China) and 1% sodium pyruvate (Biospecies, Guangzhou, China).

[0078] Cell transfection

[0079] Harvest cells at approximately 80% confluence in the logarithmic growth phase for siRNA transfection. To ensure the accuracy of the experiment, all consumables are ribonuclease-free. Transfection is performed using Lipofectamine 3000 reagent (Thermo Scientific, USA) and Opti-MEM medium (Gibco, USA) according to the manufacturer's instructions. Six to eight hours after transfection, replace the medium with fresh medium. Further culture or process the cells according to the requirements of subsequent experiments.

[0080] The siRNAs used in this study were synthesized by Tsinghua University (China) and include the following: siRNA negative control (siNeg), siPDE7B-1 (sense: 5’-caccauucaaguagaua(dT)(dT), antisense: 5’-UAUCUAACUUGAAAUGGUG(dT)(dT)) and siPDE7B-2 (sense: 5’-CGCCUACUUAACAGUACAA(dT)(dT), antisense: 5’-uuguacuguuaagugcg(dT)(dT)).

[0081] RNA extraction and qRT-PCR detection

[0082] Total RNA was extracted from cells using RNAiso Plus reagent (Takara, Japan), and RNA quantification was performed using a NanoDrop 2000 spectrophotometer. Reverse transcription was carried out using HiScript II Q RT SuperMix (Vazyme, China), and quantitative PCR was performed using SYBR qPCR Master Mix (Vazyme, China) to detect mRNA expression. Primers were synthesized by Sangon Biotech (China), and ACTB was used as an internal control. The primer sequences were as follows: PDE7B (forward: 5’-ttgacttccgccttacttaacagt-3’, reverse: 5’-taattcccgaagcagccttg-3’). The relative quantification method (2^-δct) was used to calculate the fold change in gene expression.

[0083] Cell scratch assay

[0084] Scratch the confluent monolayer cells in a 6-well plate with a sterile 10 μL pipette tip. After that, wash the cells twice with PBS to remove debris and culture them in medium without FBS. Images of the scratch were captured at 0 h and 24 h, and the scratch width was measured using ImageJ software (version 1.52a) to calculate the rate of cell migration.

[0085] CCK-8 assay

[0086] Approximately 1000 cells were seeded in a 96-well plate, with 5 replicates per group. After overnight culture, siRNA transfection was performed. At 6 - 8 h after transfection, the medium was replaced with fresh medium. After another 48 h, the cells were washed with PBS, and then 90 μL of fresh medium and 10 μL of CCK-8 reagent were added. The cells were cultured for another 4 h, and the absorbance at 450 nm was measured. The cell viability was calculated using the following formula: Cell viability (%) = [(OD_experimental group - OD_blank) / (OD_control group - OD_blank)] × 100%.

[0087] Statistical analysis

[0088] All statistical analyses were performed using R (version 4.2.3) or GraphPad Prism (version 9.0.0). Unless otherwise stated, data are presented as mean ± SEM or individual data points. The Wilcoxon rank-sum test for multiple comparisons was performed using the stat_compare_means function in the ggplot2 package. Kaplan-Meier survival curves were visualized using the survival package, and the log-rank test was performed using the survdiff function. Statistical significance is indicated by asterisks: ns p > 0.05 * p < 0.05 ** p < 0.01 *** p < 0.001 **** p < 0.0001.

[0089] Results

[0090] I. Revealing the major cell types and overall transcriptomic features of HCC based on snRNA-seq

[0091] To comprehensively understand the cellular composition and diversity of HCC, we collected 209 liver tissue samples from 77 HCC patients (SMU cohort, Table S1), including cancer tissues, adjacent tissues, and normal tissues ( Figure 1 A). We performed single-nucleus RNA sequencing (snRNA-seq; 22 cancer and 22 peritumoral samples) and bulk RNA sequencing (bRNA-seq; 165 samples) to investigate the mechanisms of hepatocarcinogenesis. Through computational analysis, we identified the major cell types and hepatocyte subtypes of HCC. Based on the characteristics of hepatocyte subtypes, we integrated 29 HCC datasets to construct a diagnostic model for HCC.

[0092] Using the snRNA-seq data, we obtained a single-cell transcriptomic atlas of 174,756 cells and annotated six major cell types using known lineage-specific marker genes: hepatocytes (Hep), liver immune cells (LIC), endothelial cells (EC), hepatic stellate cells (HSC), Kupffer cells (KC), and cholangiocytes (Cho) ( Figure 1 B - C). The proportion of hepatocytes in scRNA-seq data is approximately 4.68 - 24.48% [5], while the proportion of hepatocytes in our snRNA-seq data reached 85.23% (148182 / 174756), closely reflecting the cellular composition of the liver under normal physiological conditions.

[0093] Using the bRNA-seq data, we performed a correlation analysis between samples and found that adjacent tissues and normal tissues had highly similar molecular characteristics compared to cancer tissues ( Figure 1D). Differential expression gene (DEG) analysis was performed using the |log2FoldChange| > 1.5 and p < 0.05 thresholds, and 543 downregulated genes and 1828 upregulated genes were found in cancer tissues compared to adjacent tissues, and 604 downregulated genes and 2038 upregulated genes were found in cancer tissues compared to normal tissues. Few DEGs were found between adjacent and normal tissues, emphasizing their high similarity. Figure 1 E). KEGG enrichment analysis of the intersecting genes in the cancer vs. para-cancer and cancer vs. normal groups showed that the co-downregulated DEGs were involved in pathways such as retinol metabolism, tryptophan metabolism, cytokine-cytokine receptor interaction, malaria, and chemical carcinogenesis - DNA adducts. In contrast, the co-upregulated DEGs were associated with pathways such as the cell cycle, systemic lupus erythematosus, alcoholism, neutrophil extracellular trap formation, and the muscle cell cytoskeleton. Many of these pathways are related to HCC and are important pathogenic factors in HCC development. It is worth noting that the role of GHR in HCC development has not been fully elucidated; however, some studies have shown that GHR is downregulated in liver cancer tissues.

[0094] II. Interpreting the genomic evolution of hepatocytes in the pathogenesis of HCC

[0095] To elucidate the carcinogenic landscape of hepatocytes in HCC, we extracted 148,182 hepatocytes from the snRNA-seq data. Figure 2 A). These were re-clustered into 10 hepatocyte subtypes (Hep0 - 9, Table S5). We examined the tissue origin of these subtypes and found that most cells in Hep7 / 8 / 9 originated from cancer tissues (87%, 100%, and 89% respectively; Table S6), while most cells in Hep0 originated from adjacent tissues (65%). Further analysis showed that Hep7 and Hep8 were tumor-intermediate heterogeneity subtypes, mainly from individual patients (76% and 87% respectively). Using the Tricycle [8] method based on transfer learning to infer the cell cycle stage, we found that Hep1 was mainly in the G1 / G0 phase, while Hep0 / 2 was mainly in the S phase. Figure 2 B). RNA velocity analysis using scVelo [9] showed that Hep1 exhibited the least active transcriptional kinetics, characterized by the lowest level of unspliced RNA. Figure 2 C). These results indicate that Hep1 represents a quiescent hepatocyte subtype.

[0096] We used intercnv

[10] to identify copy number variations (CNVs) in different hepatocyte subtypes, revealing distinct CNV regional patterns. For example, Hep8 showed large-scale CNV gains on chromosome 1 (position: 146.94 - 235.65 Mb), while Hep7 showed large-scale gains (position: 116.37 - 186.42 Mb) and losses (position: 226.90 - 247.17 Mb)( Figure 2 D). These differences in variant patterns may be related to HCC cell pathological heterogeneity and clonal diversity. We specifically analyzed the relationship between CNV patterns and HCC progression in Hep0 / 2 / 9. Detection of CNV regions in Hep1-chr12, Hep2-chr3, Hep9-chr6, and Hep9-chr11( Figure 2 E) showed that the changes in deg (top 10 genes) were consistent with the CNV patterns, especially the PDE7B gene within the high-frequency loss CNV region on chr6, which was downregulated in Hep9. In subsequent analyses, PDE7B was incorporated as a diagnostic gene into the systematic criteria.

[0097] To explore the patterns of genomic evolution, we constructed subclonal architectures for each patient using the inferCNV pipeline and visualized tumor evolution with UPhyloplot2

[11] . This analysis revealed consistent early CNV events, including increases in chromosomal segments on 1q, 7q, 9q, 11p, 12p, 16p, 19p, and 19q, which were observed in more than 30% of patients. Notably, 19q_gain was present in 72.7% (16 / 22) of patients, highlighting its potential role in HCC pathogenesis.

[0098] CNV scores showed that Hep0 - 9 were significantly higher than reference endothelial cells, with Hep9 ranking the highest( Figure 2 F). Given this, we further investigated the clinical relevance of Hep9 in 22 snRNA-seq HCC cases from the SMU cohort. Patients were divided into the Hep9-high group and the Hep9-low group according to the proportion of Hep9 in cancer tissues. The results showed that the serum AFP level in the Hep9-High group was significantly higher than that in the Hep9-Low group (p = 0.013, Figure 2 G). We used CIBERSORTx

[12] to deconvolute 22 immune cell types in the TCGA-LIHC cohort

[13] and found that the level of regulatory T cells (Tregs) was elevated in HCC tissues with a high proportion of Hep9 (p = 0.013, Figure 2(H) suppresses tumor immune surveillance and promotes tumor development [6c]. Conversely, the lower the proportion of Hep0, the higher the proportion of Tregs (p < 0.0001). Importantly, using the characteristic profiles of hepatocyte subtypes identified in snRNA-seq, we decomposed the cell composition of the TCGA cohort and found that a high proportion of Hep9 (p = 0.045) and a low proportion of Hep0 (p = 0.0013) were associated with poor prognosis of HCC ( Figure 2 (I). Finally, we determined that Hep9 is an immunosuppressive hepatocyte subtype, Hep0 is a beneficial hepatocyte subtype, Hep1 is a quiescent hepatocyte subtype, Hep2 is a major malignant subtype, Hep3-6 are transitional hepatocyte subtypes, and Hep7-8 are tumor inter-heterogeneity subtypes.

[0099] III. Molecular pathways of HCC hepatocytes decoded by snRNA-seq

[0100] Using Hep0 as a reference, we performed multi-pathway enrichment analysis

[14] to identify the pathways with downregulated gene expression in Hep1 to Hep9. The results showed that liver metabolic functions were generally inhibited, especially energy metabolism and amino acid metabolism pathways, including pyruvate metabolism, butyrate metabolism, fatty acid degradation, and tryptophan metabolism. In addition, downregulation of ABC transporters and PPAR signaling pathways was observed in multiple hepatocyte subtypes, indicating their important roles in HCC progression, especially in cases associated with alcoholic liver disease.

[0101] Multi-pathway enrichment analysis of the upregulated genes in Hep1-9 showed obvious differences in pathway activation ( Figure 3 (A). Hep1 mainly showed upregulation of cholesterol metabolism, while Hep2 showed significant alterations in fatty acid metabolism. Hep3 showed strong activation of the PPAR pathway and had characteristics related to bacterial infections, including shigellosis, Salmonella infection, bacterial invasion of epithelial cells, and pathogenic Escherichia coli infection. Hep4 was similar to Hep3 but also showed upregulation of endocytosis and platelet activation. Hep5 showed activation of the Hippo pathway and glycosaminoglycan biosynthesis. Hep6 was upregulated in kinesin, phagosome, and cholinergic synapse pathways. Hep7 showed increased bile secretion, activation of the cAMP signaling pathway, and dysregulation of folate metabolism. Hep8 is a more complex subtype with characteristics of Hep1 / 2 / 3 / 4 / 7, indicating a higher degree of malignancy. Hep9 was similar to Hep8 with upregulation of the oxidative phosphorylation pathway but lacked increased bile secretion and activation of the cAMP signaling pathway. In addition, gene cross-analysis of the top 20 DEGs in Hep1-9 showed that several HCC-related genes were continuously downregulated in more than three hepatocyte subtypes, including CDA, DDI2, IFNLR1, MAN1C1, PEX14, and PHACTR4 (Figure 3 B). Notably, in Hep1 / 3 / 4 / 6, genes such as ERRFI1

[16] and PER3

[17] were significantly downregulated. Although most of these genes were previously associated with HCC, their specific roles in different hepatocyte subtypes remain poorly understood. Further functional validation and comprehensive studies are needed to clarify their roles in the occurrence, progression, and clinical consequences of HCC.

[0102] Carlessi et al.

[18] used snRNA-seq combined with two mouse models to characterize the cellular microenvironment of healthy and pre-cancerous livers. They identified a disease-associated hepatocyte (daHep) subtype, the proportion of which gradually increased with the progression of chronic liver disease. These cells were characterized by elevated ANXA2 expression and were considered potential biomarkers for predicting HCC risk. Our snRNA-seq analysis found that Hep3 / 4 / 5 / 6 were transitional hepatocyte subtypes, which prompted us to investigate their potential link with daHep. Further analysis of the gene expression profiles of 10 hepatocyte subtypes found that Hep5 had the highest ANXA2 expression, while Hep0 had the lowest. In addition, the upregulation of the endocytosis pathway observed in daHep was also found in Hep4 ( Figure 3 A). Pathways downregulated in daHep, such as complement and coagulation cascades, valine, leucine, and isoleucine degradation, fatty acid metabolism, peroxisome, and PPAR signaling pathways, were also enriched in Hep3 / 4 / 5 / 6. These findings strongly suggest that Hep3 / 4 / 5 / 6 represent transitional hepatocytes with potential roles in HCC progression, similar to the characteristics of daHep.

[0103] Next, we investigated the oncogenic landscape and molecular mechanisms of Hep0 compared with Hep2 / 9. Considering the increased tumor heterogeneity in post-cancer hepatocytes, we defined Hep0 as the non-cancer state. Using the FindMarkers function in the Seurat package, we identified the marker genes of Hep0 and Hep2 / 9. Among the top 10 enriched keg pathways, 5 were commonly enriched in both the Hep0 vs. Hep2 and Hep0 vs. Hep9 groups: complement and coagulation cascades, retinol metabolism, peroxisome, biosynthesis of amino acids, and tryptophan metabolism ( Figure 3 C). The downregulation of these pathways promoted the development of liver cancer, especially retinol metabolism and tryptophan metabolism, which was also confirmed by our bRNA-seq data ( Figure 1F). Similarly, the comparisons of Hep9 vs. Hep0 and Hep2 vs. Hep0 revealed the enrichment of tumor-related pathways, including axon guidance

[20] , ecm receptor interaction

[21] , chemical carcinogenesis reactive oxygen species

[22] , and oncoprotein glycan

[23] . However, no shared pathways were found between these two groups, highlighting the heterogeneity of tumor cells in HCC.

[0104] Gene set enrichment analysis (GSEA) further revealed the changes in gene expression and their related biological processes between the Hep0 vs. Hep2 and Hep0 vs. Hep9 groups ( Figure 3 D). We identified key metabolic pathways associated with HCC, such as glycolysis / gluconeogenesis, carbon metabolism, and purine metabolism, providing insights into the biological mechanisms of HCC development and potential therapeutic targets.

[0105] IV. Exploring the oncogenic landscape of hepatocytes by integrating snRNA-seq and bRNA-seq

[0106] To dissect the oncogenic regulatory mechanisms of HCC, we integrated snRNA-seq and bRNA-seq data from the SMU cohort and elucidated the molecular signatures driving the transformation of hepatocytes to HCC. In snRNA-seq, Hep0 was used as the baseline for identifying degs, while in bRNA-seq, adjacent cancerous tissues and normal tissues were used as controls. Venn analysis of these degs identified 74 commonly downregulated genes and 5 commonly upregulated genes associated with HCC ( Figure 4 A). Enrichment analysis using Metascape

[24] showed that the commonly downregulated degs were associated with multiple chronic liver diseases, including fatty liver, drug-induced liver disease, and non-alcoholic fatty hepatitis.

[0107] During HCC progression, we observed the inhibition of the expression of multiple CYP family genes (such as CYP2A7, CYP2B6, CYP2C8, CYP2C19), and significant disturbances in key physiological metabolic processes such as carboxylic acid metabolism, alkene metabolism, and steroid metabolism. Notably, we identified 5 significantly upregulated degs highly associated with HCC progression: ACSL4, GPC3, IGF2BP2, PPP1R9A, and ROBO1. Except for PPP1R9A, the expression of all other genes was elevated in the TCGA cohort, indicating their potential as biomarkers for HCC.

[0108] METAFlux is a computational framework that can predict cancer metabolic fluxes from bulk or single-cell transcriptomic data, revealing metabolic heterogeneity and interactions in the tumor microenvironment. Using METAFlux, we evaluated the metabolic changes in the Hep0 / 2 / 9 subtype and identified progressive and significant differences in metabolic pathways such as glycolysis / gluconeogenesis, nucleotide metabolism, glycosphingolipid biosynthesis - ganglio series, and glycosphingolipid biosynthesis - globo series ( Figure 4 B). Meanwhile, bRNA-seq data showed that these four metabolic pathways were significantly upregulated in cancer tissues, indicating their potential impact on the occurrence and progression of HCC ( Figure 4 C, S4C).

[0109] The pathogenesis of HCC is associated with the abnormal activation of multiple oncogenes, and the regulatory network of transcription factors (TFs) provides an important approach to exploring the molecular mechanisms of this disease. By applying the SCENIC

[26] algorithm to the bRNA-seq data, we found several upregulated regulations in cancer tissues, including PDX1, MAZ, and SOX4 ( Figure 4 D). Then, we evaluated the regulatory activity scores of Hep0 / 2 / 9 and compared them with the regulatory factors identified by bRNA-seq. This analysis revealed five important regulators: E2F1

[27] , SOX4

[28] , BPTF

[29] , EP300

[30] , IKZF1

[31] , and RXRA

[32] ( Figure 4 E, S4E). Although these regulators are known to be involved in the pathogenesis and progression of liver cancer, fully elucidating their regulatory targets remains elusive, and thus further detailed studies are needed.

[0110] Our study aimed to elucidate the gene regulatory network in HCC. Taking PDE7B as the research object, we used bRNA-seq and snRNA-seq to predict the transcription factors targeting PDE7B and ranked them according to their importance coefficients ( Figure 4 F). From the intersection of these analyses, we identified three transcription factors as regulators of PDE7B: ESR1, MSRA, and PLG. Currently, the regulatory relationship between PDE7B and these transcription factors lacks research. Revealing these regulatory mechanisms provides new insights into the pathogenesis and progression of liver cancer.

[0111] V. Construction of a quantile distribution model (QDM) based on the reference snRNA-seq results of the HCC cohort

[0112] To explore the potential clinical value of the snRNA-seq data, we integrated our snRNA-seq results with publicly available RNA-seq data to construct a quantile distribution model (QDM) applicable to HCC (Figure 5 A). First, we decomposed the cellular composition of the TCGA HCC cohort (n = 371) according to the molecular profiles of 10 hepatocyte subtypes (Hep0-9) obtained from snRNA-seq data. Kaplan-Meier survival analysis showed that patients with a higher proportion of the Hep0 subtype had significantly longer survival times (p = 0.0013), while patients with a higher proportion of the Hep9 subtype had significantly shorter survival times (p = 0.045). Similarly, in the Fudan-hcc cohort

[34] (n = 159), we observed a strong positive correlation between the proportion of the Hep0 subtype and survival time (p = 0.0011), and a strong negative correlation between the proportion of the Hep9 subtype and survival time (p = 0.0038). These findings were consistent with our expectations, as Hep0 mainly represents the non-cancerous components of peritumoral tissues, which may be relatively beneficial for long-term survival, while Hep9 represents a highly malignant subtype associated with poor prognosis. In addition, although Hep2 was only weakly negatively correlated with survival time in the TCGA cohort, it was strongly negatively correlated with survival time in the Fudan cohort (p < 0.0001).

[0113] Next, we used the mlr3verse

[35] machine learning R package and the random forest algorithm to extract the Hep0, Hep2, and Hep9 subtypes from all hepatocytes (training samples: test samples = 8:2). In 5-fold cross-validation, the average classification accuracies of Hep0, Hep2, and Hep9 were 89.59%, 93.43%, and 99.44%, respectively. Notably, several established tumor suppressor genes (such as ESR1

[36] , GHR [7a], ALDOB

[37] , BAAT

[38] , ALDH1A2

[39] ) were among the top 10 characteristic genes in Hep0. In contrast, Hep9 was enriched in known pro-tumor genes (such as AFP, PEG10

[40] , GPC3

[41] , IGF2BP2

[42] , ACSL4

[43] ). We extracted the top 30 characteristic genes from Hep0 / 2 / 9 (total = 90) and plotted the receiver operating characteristic (ROC) curves for HCC diagnosis in 10 datasets. Genes with an average area under the curve (AUC) less than 0.70 were excluded, resulting in 29 high-performance diagnostic genes for constructing a diagnostic model ( Figure 5 B).

[0114] In the process of QDM construction, we faced two major challenges. The first was the choice of model construction method. Although machine learning algorithms can usually achieve high diagnostic accuracy, they require high-quality data and have limited generalization ability for new data. Therefore, we chose the logistic regression algorithm because it is simple, easy to apply to new data, and has clinical practicability. The second challenge was data integration. Although there are data integration algorithms (e.g., ComBat-seq

[44] , sva

[45] , Ratio-G

[46] , etc.), they often cannot effectively eliminate the differential effects between datasets (including differences in detection platforms, methods, genes, and operations). For this reason, we adopted the quantile-based scoring algorithm of Example 1: 1) Within each dataset, all the original expression values of each gene were sorted, and the deciles of each gene expression value were calculated. 2) Each original expression was scored according to the deciles (1-10 points). 3) The score values of different datasets were directly merged. Since the scoring was carried out independently in each dataset, the background differences between different datasets were effectively reduced. Using this quantile scoring algorithm, 29 human HCC datasets (4192 samples) from different countries, detection methods, and instrument platforms were integrated, regardless of whether they were infected with hepatitis virus and the type of infection.

[0115] Starting from 29 genes showing high diagnostic performance, we constructed a model using 10 scoring datasets (Training Set1, n = 2314). During this process, we eliminated genes with no significant differences and finally established a quantile distribution model (QDM) containing 16 genes. In the QDM regression formula, we defined the HCC tissue event as 1 and the non-HCC tissue event as 0. The regression coefficient of each gene indicates whether a gene acts as a protective factor (regression coefficient < 0) or a risk factor (regression coefficient > 0) for hepatocarcinogenesis ( Figure 5 c): QDM_score = 5.5942620 - 0.3870143×PDE7B - 0.3409017×GHR + 0.3371936×GPC3 - 0.3187214×SLCO1B3 - 0.2603766×ESRRG + 0.2396385×SLC38A4 + 0.2373530×ROBO1 - 0.2361836×RALYL - 0.1768778×IGF2BP2 - 0.1744291×SERPINA1 - 0.1508238×APOA1 + 0.1467488×MEP1A + 0.1377588××FTL ACSL4 + 0.1109942 - 0.1078583×IGF2. The ROC analysis of Training Set 1 yielded an AUC of 0.982 ( Figure 5D). When applied to the remaining 19 datasets (validation set 1, n = 1878), the overall AUC of QDM for predicting HCC and non-HCC tissues was 0.968( Figure 5 E). Additionally, in our SMU cohort, the diagnostic AUC of QDM reached 0.9868. Apparently, QDM demonstrated an extremely high diagnostic performance in both training set 1 and validation set 1 (total = 4192), supporting the integration, evaluation, and diagnosis of cross-dataset and cross-platform applications.

[0116] VI. Exploring the biological and clinical significance of QDM

[0117] QDM showed strong predictive performance in diagnosing HCC tissues and provided valuable insights into the molecular mechanisms of HCC development, highlighting its important impact in scientific research and clinical practice. The 16 diagnostic genes identified by QDM have largely been confirmed to be crucial for the pathogenesis of HCC, showing tumor-suppressive or oncogenic properties. For example, QDM found protective factors such as GHR[7a], SLCO1B3

[47] , and APOA1

[48] , as well as risk factors such as GPC3

[41] , SLC38A4

[49] , and ROBO1

[50] . The expression patterns of the top 4 genes (PDE7B, GHR, GPC3, SLCO1B3) were analyzed according to the absolute value of the regression coefficients. In the bRNA-seq data of the SMU cohort, the expression levels of these genes in cancer tissues were significantly different from those in adjacent or normal tissues( Figure 6 A). We further detected their expression in 8 large-scale HCC datasets (GSE14520, GSE22058, GSE102079, GSE144269, PMID35382356, GSE89377, GSE17856, GSE45267) and found significant differences between cancer tissues and the control group( Figure 6 B). Additionally, the up- or down-regulation directions matched the regression coefficients. Validation through the Human Protein Atlas (HPA) database

[51] confirmed that the protein level expressions of PDE7B and GHR were down-regulated in HCC cancer tissues compared with normal tissues( Figure 6 C).

[0118] To explore the clinical significance of QDM, we collected the clinical test results of the SMU cohort, including AFP and AST (liver function impairment indicators). We observed a weak correlation between the QDM score and serum AFP level (Spearman coefficient = 0.33, p = 0.02) and serum AST level (Spearman coefficient = 0.3, p = 0.02) Figure 6E). The Spearman coefficient between IGF2 (a diagnostic gene of QDM) and AFP was 0.45 (p = 0.00102), while the Spearman coefficient between GHR (another diagnostic gene of QDM) and AFP was -0.42 (p = 0.00195), suggesting the potential of developing clinical biomarkers through QDM. Figure 6 F). Using the ESTIMATE

[52] algorithm, we evaluated the tumor microenvironment in the bRNA-seq data from the SMU cohort and found that compared with adjacent tissues, the immune and stromal scores were lower and the tumor purity was higher in cancer tissues. Figure 6 G). In addition, we found that PDE7B was negatively correlated with tumor purity (Spearman coefficient = -0.41, p < 0.0001). Figure 6 H). PDE7B was highly expressed in Hep0.

[0119] To verify the reliability and interpretability of QDM, we investigated the tumor biological behavior of PDE7B in HCC cells. PDE7B, a member of the phosphodiesterase family, regulates multiple physiological processes by hydrolyzing cAMP and cGMP, including cell proliferation, differentiation, inflammation, and metabolic functions, and is involved in the tumor development of colorectal cancer

[54] , leukemia

[55] , and breast cancer

[56] ; however, its specific role in the pathogenesis of HCC remains to be further investigated. First, we analyzed the expression of the PDE7B gene in four HCC cell lines (Hep G2, PLC / PRF / 5, MHCC-97H, and SNU-387) and found that the expression of PDE7B was the highest in the SNU-387 cell line, consistent with the data in the HPA database. Therefore, we selected SNU-387 cells for further experimental studies. To evaluate the functional role of PDE7B, we designed two siRNA sequences (siPDE7B-1 and siPDE7B-2) and evaluated their interference efficiency by cell transfection and qRT-PCR. The results showed that the interference efficiency of siPDE7B-1 was 70.0%, while that of siPDE7B-2 was 68.8%. Since there was no significant difference between the two sequences, we selected siPDE7B-1 (hereinafter referred to as siPDE7B) for subsequent experiments. After transfecting SNU-387 cells with siPDE7B, we performed a cell scratch assay 48 hours after transfection and captured images at 0 and 24 hours during the scratch assay. The results showed that the migration rate of the siPDE7B group was significantly higher than that of the control group (p < 0.01). The CCK-8 assay results showed that the cell viability of the siPDE7B group was significantly higher than that of the control group (p < 0.05). In summary, inhibiting PDE7B can significantly enhance the migration ability and viability of HCC cells, which is consistent with the prediction of QDM. Future studies will further explore the molecular mechanism of PDE7B in HCC and evaluate its potential as a clinical biomarker for HCC.

[0120] Overall, QDM exhibited strong diagnostic performance and was significantly correlated with clinical indicators. More importantly, it has great potential in identifying novel biomarkers with clinical significance.

[0121] VII. Classifying HCC Using Signature Genes Identified by snRNA-seq

[0122] To improve the current clinical classification of HCC and utilize advanced scRNA-seq information, we extracted the molecular signatures of three different hepatocyte subtypes (Hep0 / 2 / 9) Figure 7A). We first extracted the characteristic genes from the snRNA-seq data, selecting 30 genes for each subtype, and finally obtained 90 genes. Then, we compiled 8 human HCC datasets (n = 1350), processed them to remove batch effects, and divided them into training set 2 and validation set 2. After excluding non-expressed genes, we retained 50 common genes. Applying unsupervised consensus clustering to training set 2, we identified three distinct HCC classes: Class_1, Class_2, and Class_3.

[0123] Genomic studies have shown that CTNNB1 mutations are associated with well-differentiated HCC tumors and prolonged overall survival

[57] , AXIN1 mutations are associated with activation of the Wnt / β-catenin pathway

[58] , frequent TP53 mutations are associated with poor prognosis

[59] , and TSC2 mutations are often associated with significant stromal fibrosis and sclerosis in HCC

[60] . We examined the gene mutation frequencies of the three HCC types and found that compared with Class_2 and Class_3, the CTNNB1 mutation frequency in Class_1 was significantly higher (27.3%, 15.6%, and 16.3% respectively) ( Figure 7 B). In contrast, the TP53 mutation frequency was significantly lower in Class_1 (30.6%, 48.1%, and 48.8% respectively), suggesting a lower risk of poor prognosis in Class_1. Among them, the AXIN1 mutation frequency was higher in Class_2 (5.4%, 22.0%, and 2.3% respectively), while the TSC2 mutation frequency was higher in Class_3 (2.1%, 6.5%, and 11.6% respectively).

[0124] To further explore the clinical significance of these three classes, we performed survival analysis using the integrated dataset (TCGA and PMID31585088) and observed that Class_1 showed better clinical outcomes ( Figure 7 C). Using TCGA and PMID31585088 as training set 2, we established an HCC classification model with an accuracy of 91.58%. Then, we applied this model to validation set 2 and confirmed that compared with Class_2 and Class_3, the survival time of Class_1 was significantly longer ( Figure 7 D). CIBERSORTx analysis of the immune cell composition of training set 2 showed that compared with Class_1 and Class_3, Class_2 had significantly higher abundances of various immune cell types (memory B cells, monocytes, plasma cells, and follicular helper T cells) ( Figure 7E). Additionally, TIDE

[61] analysis showed that compared with Class_1 and Class_2, Class_3 exhibited higher levels of cancer-associated fibroblasts (CAF) and T cell exclusion scores, indicating that Class_3 has a highly immunosuppressive tumor microenvironment ( Figure 7 F).

[0125] DEGs analysis was performed using Training Set 2, and 164, 68, and 327 upregulated genes were identified in Class_1, Class_2, and Class_3, respectively, with the threshold set at p < 0.05, log2FoldChange > 1.5. Enrichment analysis showed that Class_1 was mainly related to metabolic pathways, including carboxylic acid metabolic process, biological oxidation process, small molecule catabolic process, and steroid metabolic process ( Figure 7 G). Class_2 was enriched in cancer-related pathways, such as Wnt signaling pathway, gastric cancer network 1, MAPK cascade regulation, canonical Wnt signaling pathway, and lncRNA in colorectal cancer ( Figure 7 G). In contrast, Class_3 was related to extracellular matrix-related pathways, including extracellular matrix organization, positive regulation of cell adhesion, and cell-cell adhesion ( Figure 7 G). Based on these findings, we classified Class 1, Class 2, and Class 3 as metabolic HCC, inflammatory HCC, and stromal HCC, respectively.

[0126] The classification of HCC helps to optimize individualized treatment strategies. We evaluated the drug responsiveness of the three HCC types to sorafenib using oncopdict

[62] and found that compared with Class_1 and Class_2, Class_3 had a higher half-maximal inhibitory concentration (IC50), indicating a reduced sensitivity to sorafenib ( Figure 7 H). Connectivity Map (CMap)

[63] is a gene expression-based database designed to clarify the functional relationship between drugs and disease states. We performed CMap analysis on each HCC category and combined these findings with the oncopdict results to determine that the drugs most sensitive to Class_1, Class_2, and Class_3 were ouabain, teniposide, and TG-101348, respectively. The IC50 values of these three drugs were lower than that of sorafenib, making them suitable for combination drug administration regimens for precision treatment of HCC.

[0127] The above embodiments take liver cancer as an example to illustrate the feasibility of the tumor diagnosis model based on the gene expression scoring algorithm of the embodiments of the present invention. The model of the present invention can be used not only for liver cancer prediction, but also for other tumors, including but not limited to liver cancer, lung cancer, gastric cancer, colon cancer, cervical cancer, breast cancer, esophageal cancer, kidney cancer, pancreatic cancer, leukemia, brain tumor, osteosarcoma, lymphoma, etc., such as liver cancer, lung cancer, breast cancer, colorectal cancer, nasopharyngeal cancer, etc. Currently available transcriptome datasets mainly come from technologies such as microarrays and RNA sequencing. These technologies have significant differences in detection principles, platforms, target genes, and data formats, making direct integration complex. By adopting the scoring algorithm of the embodiments of the present invention, these obstacles can be effectively overcome. The algorithm converts gene expression values into relative expression values according to the rankings of genes. This method improves the robustness of the model, is interpretable, general, and iteratively expandable.

[0128] References

[0129] [1] S.L. Wolock, R. Lopez, A.M. Klein Cel Syst 2019, 8, 281.

[0130] [2] A. Ianevski, A.K. Giri, T. Aittokallio Nat Commun 2022, 13, 1246.

[0131] [3] A.M. Bolger, M. Lohse, B. Usadel Bioinformatics 2014, 30, 2114.

[0132] [4] D. Kim, J.M. Paggi, C. Park, C. Bennett, S.L. Salzberg Nat Biotechnol 2019, 37, 907.

[0133] [5] a) Y. Lu, A. Yang, C. Quan, Y. Pan, H. Zhang, Y. Li, C. Gao, H. Lu, X. Wang, P. Cao, H. Chen, S. Lu, G. Zhou Nat Commun 2022, 13, 4594; b) K. Li, R. Zhang, F. Wen, Y. Zhao, F. Meng, Q. Li, A. Hao, B. Yang, Z. Lu, Y. Cui, M. Zhou Hepatology 2024, 79, 1293.

[0134] [6]a) G.N. Ioannou, J Hepatol 2021, 75, 1476; b) J.M. Yuan, Y.T. Gao, C.N. Ong, R.K. Ross, M.C. Yu, J Natl Cancer Inst 2006, 98, 482; c) H. Wang, H. Zhang, Y. Wang, Z.J. Brown, Y. Xia, Z. Huang, C. Shen, Z. Hu, J. Beane, E.A. Ansa-Addo, H. Huang, D. Tian, A. Tsung, J Hepatol 2021, 75, 1271; d) A. Walakira, C. Skubic, N. Nadizar, D. Rozman, T. Rezen, M. Mraz, M. Moskon, Comput Biol Med 2023, 159, 106957. [7] a) C.C. Lin, T.W. Liu, M.L. Yeh, Y.S. Tsai, P.C. Tsai, C.F. Huang, J.F. Huang, W.L. Chuang, C.Y. Dai, M.L. Yu, Clin Mol Hepatol 2021, 27, 313; b) A. Haque, V. Sahu, J.L. Lombardo, L. Xiao, B. George, R.A. Wolff, J.S. Morris, A. Rashid, J.J. Kopchick, A.O. Kaseb, H.M. Amin, J Hepatocel Carcinoma 2022, 9, 823.

[0135] [8] S.C. Zheng, G. Stein-O'Brien, J.J. Augustin, J. Slosberg, G.A. Carosso, B. Winer, G. Shin, H.T. Bjornsson, L.A. Goff, K.D. Hansen, Genome Biol 2022, 23, 41.

[0136] [9] V. Bergen, M. Lange, S. Peidli, F.A. Wolf, F.J. Theis, Nat Biotechnol 2020, 38, 1408.

[0137]

[10] P.A. Kenny, F1000Res 2019, 8, 807.

[0138]

[11] a) J. Zhang, S. S. Spath, S. L. Marjani, W. Zhang, X. Pan, *Precis Clin Med* 2018, 1, 29; b) S. Kurtenbach, A. M. Cruz, D. A. Rodriguez, M. A. Durante, J. W. Harbour, *BMC Genomics* 2021, 22, 419.

[12] A. M. Newman, C. B. Steen, C. L. Liu, A. J. Gentles, A. A. Chaudhuri, F. Scherer, M. S. Khodadoust, M. S. Esfahani, B. A. Luca, D. Steiner, M. Diehn, A. A. Alizadeh, *Nat Biotechnol* 2019, 37, 773.

[0139]

[13] w.b.e. Cancer Genome Atlas Research Network. Electronic address, N. Cancer Genome Atlas Research Cel 2017, 169, 1327.

[0140]

[14] S. Xu, E. Hu, Y. Cai, Z. Xie, X. Luo, L. Zhan, W. Tang, Q. Wang, B. Liu, R. Wang, W. Xie, T. Wu, L. Xie, G. Yu, *Nat Protoc* 2024, 19, 3292.

[0141]

[15] a) Y. Cui, S. Liang, S. Zhang, C. Zhang, Y. Zhao, D. Wu, J. Wang, R. Song, J. Wang, D. Yin, Y. Liu, S. Pan, X. Liu, Y. Wang, J. Han, F. Meng, B. Zhang, H. Guo, Z. Lu, L. Liu J Exp Clin Cancer Res 2020, 39, 90; b) S. Gallage, A. Ali, J. E. Barragan Avila, N. Seymen, P. Ramadori, V. Joerke, L. Zizmare, D. Aicher, I. K. Gopalsamy, W. Fong, J. Kosla, E. Focaccia, X. Li, S. Yousuf, T. Sijmonsma, M. Rahbari, K. S. Kommoss, A. Billeter, S. Prokosch, U. Rothermel, F. Mueller, J. Hetzer, D. Heide, B. Schinkel, T. Machauer, B. Pichler, N. P. Malek, T. Longerich, S. Roth, A. J. Rose, J. Schwenck, C. Trautwein, M. M. Karimi, M. Heikenwalder Cel Metab 2024, 36, 1371; c) L. Gong, F. Wei, F. J. Gonzalez, G. Li Hepatology 2023, 78, 1625.

[0142]

[16] M. Cui, D. Liu, W. Xiong, Y. Wang, J. Mi Cel Death Discov 2021, 7, 274.

[0143]

[17] a) O. Neumann, M. Kesselmeier, R. Geffers, R. Pellegrino, B. Radlwimmer, K. Hoffmann, V. Ehemann, P. Schemmer, P. Schirmacher, J. Lorenzo Bermejo, T. Longerich, Hepatology 2012, 56, 1817; b) D. I. Sanchez, B. Gonzalez-Fernandez, I. Crespo, B. San-Miguel, M. Alvarez, J. Gonzalez-Gallego, M. J. Tunon, J Pineal Res 2018, 65, e12506.

[0144]

[18] R. Carlessi, E. Denisenko, E. Boslem, J. Kohn-Gaone, N. Main, N. D. B. Abu Bakar, G. D. Shirolkar, M. Jones, A. B. Beasley, D. Poppe, B. J. Dwyer, C. Jackaman, M. C. Tjiam, R. Lister, M. Karin, J. A. Fallowfield, T. J. Kendall, S. J. Forbes, E. S. Gray, J. K. Olynyk, G. Yeoh, A. R. R. Forrest, G. A. Ramm, M. A. Febbraio, J. E. E. Tirnitz-Parker, Cel Genom 2023, 3, 100301.

[0145]

[19] a) Z.C. Nwosu, N. Battello, M. Rothley, W. Pioronska, B. Sitek, M.P. Ebert, U. Hofmann, J. Sleeman, S. Wolfl, C. Meyer, D.A. Megger, S. Dooley J Exp Clin Cancer Res 2018, 37, 211; b) W. He, X. Wang, M. Chen, C. Li, W. Chen, L. Pan, Y. Cui, Z. Yu, G. Wu, Y. Yang, M. Xu, Z. Dong, K. Ma, J. Wang, Z. He Clin Transl Med 2023, 13, e1465; c) Q. Zuo, J. He, S. Zhang, H. Wang, G. Jin, H. Jin, Z. Cheng, X. Tao, C. Yu, B. Li, C. Yang, S. Wang, Y. Lv, F. Zhao, M. Yao, W. Cong, C. Wang, W. Qin Hepatology 2021, 73, 644; d) L. Yu, J. Lu, W. Du Cel Commun Signal 2024, 22, 174.

[0146]

[20] J.Q. Liang, N. Teoh, L. Xu, S. Pok, X. Li, E.S.H. Chu, J. Chiu, L. Dong, E. Arfianti, W.G. Haigh, M.M. Yeh, G.N. Ioannou, J.J.Y. Sung, G. Farrell, J. Yu Nat Commun 2018, 9, 4490.

[0147]

[21] Z.H. Liu, B.F. Lian, Q.Z. Dong, H. Sun, J.W. Wei, Y.Y. Sheng, W. Li, Y.X. Li, L. Xie, L. Liu, L.X. Qin Biochim Biophys Acta Mol Basis Dis 2018, 1864, 2360.

[0148]

[22] D. Povero, Y. Chen, S. M. Johnson, C. E. McMahon, M. Pan, H. Bao, X. T. Petterson, E. Blake, K. P. Lauer, D. R. O′Brien, Y. Yu, R. P. Graham, T. Taner, X. Han, G. L. Razidlo, J. Liu J Hepatol 2023, 79, 378.

[23] F. Dituri, G. Gigante, R. Scialpi, S. Mancarella, I. Fabregat, G. Giannelli Cancers (Basel) 2022, 14,

[0149]

[24] Y. Zhou, B. Zhou, L. Pache, M. Chang, A. H. Khodabakhshi, O. Tanaseichuk, C. Benner, S. K. Chanda Nat Commun 2019, 10, 1523.

[0150]

[25] Y. Huang, V. Mohanty, M. Dede, K. Tsai, M. Daher, L. Li, K. Rezvani, K. Chen Nat Commun 2023, 14, 4883.

[0151]

[26] S. Aibar, C. B. Gonzalez-Blas, T. Moerman, V. A. Huynh-Thu, H. Imrichova, G. Hulselmans, F. Rambow, J. C. Marine, P. Geurts, J. Aerts, J. van den Oord, Z. K. Atak, J. Wouters, S. Aerts Nat Methods 2017, 14, 1083.

[0152]

[27] a) F. Gonzalez-Romero, D. Mestre, I. Aurrekoetxea, C. J. O′Rourke, J. B. Andersen, A. Woodhoo, M. Tamayo-Caro, M. Varela-Rey, M. Palomo-Irigoyen, B. Gomez-Santos, D. S. de Urturi, M. Nunez-Garcia, J. L. Garcia-Rodriguez, L. Fernandez-Ares, X. Buque, A. Iglesias-Ara, I. Bernales, V. G. De Juan, T. C. Delgado, N. Goikoetxea-Usandizaga, R. Lee, S. Bhanot, I. Delgado, M. J. Perugorria, G. Errazti, L. Mosteiro, S. Gaztambide, I. Martinez de la Piscina, P. Iruzubieta, J. Crespo, J. M. Banales, M. L. Martinez-Chantar, L. Castano, A. M. Zubiaga, P. Aspichueta Cancer Res 2021, 81, 2874; b) L. N. Kent, S. Bae, S. Y. Tsai, X. Tang, A. Srivastava, C. Koivisto, C. K. Martin, E. Ridolfi, G. C. Miller, S. M. Zorko, E. Plevris, Y. Hadjiyannis, M. Perez, E. Nolan, R. Kladney, B. Westendorp, A. de Bruin, S. Fernandez, T. J. Rosol, K. S. Pohar, J. M. Pipas, G. Leone J Clin Invest 2017, 127, 830; c) T. Su, N. Zhang, T. Wang, J. Zeng, W. Li, L. Han, M. Yang Cancer Res 2023, 83, 4080.

[0153]

[28] a) Z.Z. Chen, L. Huang, Y.H. Wu, W.J. Zhai, P.P. Zhu, Y.F. Gao, Nat Commun 2016, 7, 12598; b) M. Sandbothe, R. Buurman, N. Reich, L. Greiwe, B. Vajen, E. Gurlevik, V. Schaffer, M. Eilers, F. Kuhnel, A. Vaquero, T. Longerich, S. Roessler, P. Schirmacher, M.P. Manns, T. Illig, B. Schlegelberger, B. Skawran, J Hepatol 2017, 66, 1012; c) V. Blanc, J.D. Riordan, S. Soleymanjahi, J.H. Nadeau, I. Nalbantoglu, Y. Xie, E.A. Molitor, B.B. Madison, E.M. Brunt, J.C. Mills, D.C. Rubin, I.O. Ng, Y. Ha, L.R. Roberts, N.O. Davidson, J Clin Invest 2021, 131,

[0154]

[29] X. Zhao, F. Zheng, Y. Li, J. Hao, Z. Tang, C. Tian, Q. Yang, T. Zhu, C. Diao, C. Zhang, M. Chen, S. Hu, P. Guo, L. Zhang, Y. Liao, W. Yu, M. Chen, L. Zou, W. Guo, W. Deng, Redox Biol 2019, 20, 427.

[0155]

[30] Z. Hu, G. Chen, Y. Zhao, H. Gao, L. Li, Y. Yin, J. Jiang, L. Wang, Y. Mang, Y. Gao, S. Zhang, J. Ran, L. Li, Mol Cancer 2023, 22, 55.

[0156]

[31] Q. Huo, C. Ge, H. Tian, J. Sun, M. Cui, H. Li, F. Zhao, T. Chen, H. Xie, Y. Cui, M. Yao, J. Li, Cel Death Dis 2017, 8, e2766.

[0157]

[32] M. Cui, Z. Xiao, Y. Wang, M. Zheng, T. Song, X. Cai, B. Sun, L. Ye, X. ZhangCancer Res 2015, 75, 846.

[0158]

[33] B. Jew, M. Alvarez, E. Rahmani, Z. Miao, A. Ko, K. M. Garske, J. H. Sul, K. H. Pietilainen, P. Pajukanta, E. Halperin Nat Commun 2020, 11, 1971.

[0159]

[34] Q. Gao, H. Zhu, L. Dong, W. Shi, R. Chen, Z. Song, C. Huang, J. Li, X. Dong, Y. Zhou, Q. Liu, L. Ma, X. Wang, J. Zhou, Y. Liu, E. Boja, A. I. Robles, W. Ma, P. Wang, Y. Li, L. Ding, B. Wen, B. Zhang, H. Rodriguez, D. Gao, H. Zhou, J. Fan Cel 2019, 179, 561.

[0160]

[35] M. Lang, M. Binder, J. Richter, P. Schratz, F. Pfisterer, S. Coors, Q. Au, G. Casalicchio, L. Kotthoff, B. Bischl Journal of Open Source Software 2019, 4, 1903.

[0161]

[36] H. Wu, S. Yao, S. Zhang, J. R. Wang, P. D. Guo, X. M. Li, W. J. Gan, L. Mei, T. M. Gao, J. M. Li J Hepatol 2017, 66, 1193.

[0162]

[37] M. Li, X. He, W. Guo, H. Yu, S. Zhang, N. Wang, G. Liu, R. Sa, X. Shen, Y. Jiang, Y. Tang, Y. Zhuo, C. Yin, Q. Tu, N. Li, X. Nie, Y. Li, Z. Hu, H. Zhu, J. Ding, Z. Li, T. Liu, F. Zhang, H. Zhou, S. Li, J. Yue, Z. Yan, S. Cheng, Y. Tao, H. Yin Nat Cancer 2020, 1, 735.

[0163]

[38] L. Xing, Y. Zhang, S. Li, M. Tong, K. Bi, Q. Zhang, Q. Li Int J Mol Sci 2023, 24,

[0164]

[39] Q. Shi, Y. Liu, M. Lu, Q. Y. Lei, Z. Chen, L. Wang, X. He Comput Biol Med 2022, 144, 105376.

[0165]

[40] C. M. Li, A. A. Margolin, M. Salas, L. Memeo, M. Mansukhani, H. Hibshoosh, M. Szabolcs, A. Klinakis, B. Tycko Cancer Res 2006, 66, 665.

[0166]

[41] D. Li, N. Li, Y. F. Zhang, H. Fu, M. Feng, D. Schneider, L. Su, X. Wu, J. Zhou, S. Mackay, J. Kramer, Z. Duan, H. Yang, A. Kolluri, A. M. Hummer, M. B. Torres, H. Zhu, M. D. Hall, X. Luo, J. Chen, Q. Wang, D. Abate-Daga, B. Dropulic, S. M. Hewitt, R. J. Orentas, T. F. Greten, M. Ho Gastroenterology 2020, 158, 2250.

[0167]

[42] F. L. Dong, Z. Z. Xu, Y. Q. Wang, T. Li, X. Wang, J. Li J Nanobiotechnology2024, 22, 298.

[0168]

[43] H. Peng, B. Chen, W. Wei, S. Guo, H. Han, C. Yang, J. Ma, L. Wang, S. Peng, M. Kuang, S. Lin Nat Metab 2022, 4, 1041.

[0169]

[44] Y. Zhang, G. Parmigiani, W. E. Johnson NAR Genom Bioinform 2020, 2, lqaa078.

[0170]

[45] J.T. Leek, W.E. Johnson, H.S. Parker, A.E. Jaffe, J.D. Storey Bioinformatics 2012, 28, 882.

[0171]

[46] J. Luo, M. Schumacher, A. Scherer, D. Sanoudou, D. Megherbi, T. Davison, T. Shi, W. Tong, L. Shi, H. Hong, C. Zhao, F. Elloumi, W. Shi, R. Thomas, S. Lin, G. Tillinghast, G. Liu, Y. Zhou, D. Herman, Y. Li, Y. Deng, H. Fang, P. Bushel, M. Woods, J. Zhang Pharmacogenomics J 2010, 10, 278.

[0172]

[47] S.R. Vavricka, D. Jung, M. Fried, U. Grutzner, P.J. Meier, G.A. Kullak-Ublick J Hepatol 2004, 40, 212.

[0173]

[48] Y. Wang, S. Chen, X. Xiao, F. Yang, J. Wang, H. Zong, Y. Gao, C. Huang, X. Xu, M. Fang, X. Zhang, C. Gao Precis Clin Med 2023, 6, pbad021.

[0174]

[49] H.S. Kim, M.J. Na, K.H. Son, H.D. Yang, S.Y. Kim, E. Shin, J.W. Ha, S. Jeon, K. Kang, K. Moon, W.S. Park, S.W. Nam Exp Mol Med 2023, 55, 95.

[0175]

[50] H. Ito, S. Funahashi, N. Yamauchi, J. Shibahara, Y. Midorikawa, S. Kawai, Y. Kinoshita, A. Watanabe, Y. Hippo, T. Ohtomo, H. Iwanari, A. Nakajima, M. Makuuchi, M. Fukayama, Y. Hirata, T. Hamakubo, T. Kodama, M. Tsuchiya, H. Aburatani Clin Cancer Res 2006, 12, 3257.

[0176]

[51] M. Uhlen, C. Zhang, S. Lee, E. Sjostedt, L. Fagerberg, G. Bidkhori, R. Benfeitas, M. Arif, Z. Liu, F. Edfors, K. Sanli, K. von Feilitzen, P. Oksvold, E. Lundberg, S. Hober, P. Nilsson, J. Mattsson, J. M. Schwenk, H. Brunnstrom, B. Glimelius, T. Sjoblom, P. H. Edqvist, D. Djureinovic, P. Micke, C. Lindskog, A. Mardinoglu, F. PontenScience 2017, 357,

[0177]

[52] K. Yoshihara, M. Shahmoradgoli, E. Martinez, R. Vegesna, H. Kim, W. Torres - Garcia, V. Trevino, H. Shen, P. W. Laird, D. A. Levine, S. L. Carter, G. Getz, K. Stemke - Hale, G. B. Mills, R. G. Verhaak Nat Commun 2013, 4, 2612.

[0178]

[53] J. M. Hetman, S. H. Soderling, N. A. Glavas, J. A. Beavo Proc Natl Acad SciU S A 2000, 97, 472.

[0179]

[54] Z. Chen, W. Song, X. O. Shu, W. Wen, M. Devall, C. Dampier, F. Moratalla-Navarro, Q. Cai, J. Long, L. Van Kaer, L. Wu, J. R. Huyghe, M. Thomas, L. Hsu, M. O. Woods, D. Albanes, D. D. Buchanan, A. Gsur, M. Hoffmeister, P. Vodicka, A. Wolk, L. L. Marchand, A. H. Wu, A. I. Phipps, V. Moreno, P. Ulrike, W. Zheng, G. Casey, X. Guo J Natl Cancer Inst2024, 116, 127.

[0180]

[55] L. Zhang, F. Murray, A. Zahno, J. R. Kanter, D. Chou, R. Suda, M. Fenlon, L. Rassenti, H. Cottam, T. J. Kipps, P. A. Insel Proc Natl Acad Sci U S A 2008, 105, 19532.

[0181]

[56] D. D. Zhang, Y. Li, Y. Xu, J. Kim, S. Huang Oncogene 2019, 38, 1106.

[0182]

[57] Y. Chen, X. Deng, Y. Li, Y. Han, Y. Peng, W. Wu, X. Wang, J. Ma, E. Hu, X. Zhou, E. Shen, S. Zeng, C. Cai, Y. Qin, H. Shen Hepatology 2024, 80, 536.

[0183]

[58] C. Xu, Z. Xu, Y. Zhang, M. Evert, D. F. Calvisi, X. Chen J Clin Invest 2022, 132,

[0184]

[59] Z. Fan, M. Jin, L. Zhang, N. Wang, M. Li, C. Wang, F. Wei, P. Zhang, X. Du, X. Sun, W. Qiu, M. Wang, H. Wang, X. Shi, J. Ye, C. Jiang, J. Zhou, W. Chai, J. Qi, T. Li, R. Zhang, X. Liu, B. Huang, K. Chai, Y. Cao, W. Mu, Y. Huang, T. Yang, H. Zhang, L. Qu, Y. Liu, G. Wang, G. Lv Gut2023, 72, 2149.

[0185]

[60] J. Calderaro, G. Couchy, S. Imbeaud, G. Amaddeo, E. Letouze, J. F. Blanc, C. Laurent, Y. Hajji, D. Azoulay, P. Bioulac - Sage, J. C. Nault, J. Zucman - Rossi J Hepatol2017, 67, 727.

[0186]

[61] P. Jiang, S. Gu, D. Pan, J. Fu, A. Sahu, X. Hu, Z. Li, N. Traugh, X. Bu, B. Li, J. Liu, G. J. Freeman, M. A. Brown, K. W. Wucherpfennig, X. S. Liu Nat Med 2018, 24, 1550.

[0187]

[62] D. Maeser, R. F. Gruener, R. S. Huang Brief Bioinform 2021, 22,

[0188]

[63] J. Lamb, E. D. Crawford, D. Peck, J. W. Modell, I. C. Blat, M. J. Wrobel, J. Lerner, J. P. Brunet, A. Subramanian, K. N. Ross, M. Reich, H. Hieronymus, G. Wei, S. A. Armstrong, S. J. Haggarty, P. A. Clemons, R. Wei, S. A. Carr, E. S. Lander, T. R. GolubScience 2006, 313, 1929.

[0189] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of this invention patent shall be subject to the appended claims.

Claims

1. A gene expression scoring algorithm, characterized in that: The following steps are involved: S1 collects data sets and preprocesses gene expression data, generating an expression matrix for each sample in each data set, where rows and columns or columns and rows are sample names and molecule / gene names respectively; S2 sorts the original expression values ​​of all genes in each sample and calculates the position of each gene expression value in the overall ranking of all genes in the sample, i.e., the quantile; S3 scores all genes in each sample according to the quantile corresponding to each gene; S4 combines the quantile scores of all genes of all samples in different data sets to obtain an integrated data set.

2. The gene expression scoring algorithm according to claim 1, characterized in that The preprocessing of the gene expression data in step S1 may be raw data, normalized data, or log2 or log10 processed data.

3. The gene expression scoring algorithm according to claim 1, characterized in that The quantile is calculated according to the 2 to 100 classification unit quantiles of the overall data (i.e., 2 to 100 quantiles), and the original expression of each gene in each sample is ranked; preferably, each gene in each sample is assigned a score of 1-10 (i.e., relative expression value) according to the decile.

4. The gene expression scoring algorithm according to claim 1, wherein: The data set collected in step S1 includes at least one of DNA, messenger RNA, microRNA, small nuclear RNA, circular RNA, long non-coding RNA, small nucleolar RNA, and protein.

5. A tumor diagnosis model, characterized in that: The tumor diagnosis model is constructed based on the gene expression scoring algorithm described in any one of 1 to 4.

6. The tumor diagnosis model according to claim 5, characterized in that: Tumor samples all included a tumor group and a corresponding control group.

7. The method for constructing a tumor diagnostic model according to any one of claims 5 to 6, characterized in that: Includes steps: Preprocess the gene expression data of all samples to obtain an expression matrix with sample name and gene name defined as row name or column name respectively, and set another column or row name as diagnosis in the column or row next to the sample name to record the disease status of the sample; A portion of the samples is set as a training set, and another portion of the samples is set as a validation set, wherein the validation set is at least one; On the training set, based on the gene list of interest, genes with high diagnostic performance were screened; Genes with high diagnostic performance in the training set were relatively scored and integrated; Relative scoring and integration of the validation set data sets; Compare the area under the receiver operating characteristic curve (AUROC) in the training set with the AUROC in the validation set to determine the ability to diagnose tumors across datasets, and integrate the results of the validation set; In an independent integrated validation set, a tumor diagnosis model was constructed using the algorithm.

8. The method for constructing a tumor diagnostic model according to claim 7, characterized in that: The screening criteria for the genes with high diagnostic performance are: the correct rate of diagnosing tumors in the original data set is greater than a random value or a specified value.

9. The method for constructing a tumor diagnostic model according to claim 7, wherein: The gene list of interest may be at least one of differentially expressed genes, prognosis-related genes, inflammatory genes, metabolic genes, and characteristic genes of tumor-specific cell subpopulations.

10. The method for constructing a tumor diagnostic model according to claim 7, wherein: The algorithm is at least one of a logistic regression algorithm, a machine learning algorithm, a deep learning algorithm, and an artificial intelligence large model algorithm.

Citation Information

Patent Citations

  • Derivatives of 2-aminoethanol, processes for their preparation, pharmaceutical compositions containing such compounds and the use the latter

    EP0030030A1