Digestive tract tumor cross-cancer species molecular typing method and application thereof
By integrating multi-omics data and deep learning methods, we can identify cross-cancer molecular subtypes of gastrointestinal tumors, which solves the shortcomings of cross-cancer subtyping in existing technologies and enables robust gastrointestinal tumor subtyping and personalized treatment guidance.
Patent Information
- Application Number
- CN202511722936.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-02-24
AI Technical Summary
Existing technologies struggle to achieve molecular subtyping of gastrointestinal tumors across cancer types and lack robust multi-omics data integration methods, resulting in treatment strategies that rely on tissue specificity and lack shared molecular characteristics across cancer types.
By integrating copy number variation, DNA methylation, mRNA and miRNA sequencing data, a robust rank aggregation method was used to screen high-confidence functional genes, and a Transformer autoencoder network was used to extract multi-omics features. Consensus clustering was then combined with a consensus clustering method to identify molecular subtypes of pan-gastrointestinal tumors, thus constructing a cross-cancer molecular subtype classifier.
It has achieved a robust and unified classification of gastrointestinal tumors across cancer types, providing a new standardized approach to diagnosis and treatment. It can accurately characterize the biological nature of tumors, be used for in vitro drug response prediction, and support personalized medicine research.
Smart Images

Figure CN121565255A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of tumor molecular biology technology, and in particular relates to a method for molecular subtyping of gastrointestinal tumors across cancer types and its application in predicting in vitro drug responses. Background Technology
[0002] Gastrointestinal tumors mainly include malignant tumors of the esophagus, stomach, pancreas, liver, and colorectal region. Despite significant advances in diagnosis and treatment, effective treatment remains a major challenge for many gastrointestinal tumors.
[0003] Genomic and transcriptomic analyses have elucidated the extensive heterogeneity within gastrointestinal tumors by classifying them into distinct molecular subtypes. Each subtype is marked by specific genetic abnormalities and gene expression signatures, reflecting significant biological differences. Recent evidence further demonstrates the existence of conserved molecular subtype patterns across different gastrointestinal organs. For example, molecular subtypes characterized by epithelial-mesenchymal transition, immune cell infiltration, and metabolic dysregulation are recurring features in almost all gastrointestinal malignancies. Notably, the consensus molecular subtyping (CMS) system, originally developed for colorectal cancer (CRC), has shown applicability in gastric adenocarcinoma (STAD), with the CMS1 subtype predicting a favorable response to pembrolizumab. These conserved molecular patterns suggest that shared treatment strategies across tissue origin boundaries may be feasible. For instance, immunotherapy has shown significant efficacy in treating gastrointestinal tumors with high microsatellite instability (MSI-H) regardless of tissue origin. Similarly, trastuzumab has shown clinical benefit in HER2-positive gastric and colorectal cancers, highlighting the potential of a subtype-driven rather than organ-specific treatment modality.
[0004] Currently, integrating multi-omics data provides a powerful tool for revealing such fine and clinically relevant subtypes. This strategy has been successful in other cancer families, such as providing a unified prognostic classification for gynecological (breast, uterus, cervix, and ovary) and urinary system (bladder, kidney, prostate, and testis) malignancies. For gastrointestinal tumors, the Cancer Genome Atlas (TCGA) project conducted the Pan-GI project, performing fundamental multi-omics characterization of gastrointestinal adenocarcinomas (GIACs), highlighting shared and specific molecular features. However, the final classification still relies on tissue origin and does not employ integrative modeling to define true cross-cancer subtypes. Other studies have explored alternative classification strategies, such as subtype classification based on stromal stiffness using selected gene sets, or multi-omics clustering analysis of gastrointestinal tumors, but these methods are often limited by tissue-dominated clustering, lack robust validation, or fail to truly integrate multi-omics data. Therefore, these existing technologies illustrate the challenge of overcoming tissue-specific signals in the molecular subtyping of gastrointestinal tumors across cancer types, and the necessity of establishing a consistent subtyping framework capable of capturing shared molecular features among gastrointestinal tumors. Summary of the Invention The technical problem to be solved by the present invention is to overcome the deficiencies and defects mentioned in the background art above, and to provide a method for molecular subtyping of gastrointestinal tumors across cancer types, a method for constructing a classifier, and its application in in vitro drug response prediction for non-diagnostic / therapeutic purposes.
[0005] To solve the above-mentioned technical problems, the technical solution proposed by this invention is as follows: A method for molecular subtyping of gastrointestinal tumors across cancer types includes the following steps: (1) Obtain multi-omics data of gastrointestinal tumors, including copy number variation data, DNA methylation data, mRNA sequencing data and miRNA sequencing data; perform batch effect correction on the multi-omics data and merge them to construct a unified multi-omics dataset; (2) For each gene, calculate six quantitative feature scores, as follows: Fcnv: Number of samples with copy number mutations / Total number of samples; Fmethy: The number of co-expressed genes that are significantly associated with the target gene; Ftf: The number of transcription factors that have a significant regulatory effect on the target gene; Fmir: Significantly regulates the number of miRNAs; Fmad: Mean absolute deviation across all samples; Fde: The absolute value of the fold difference in expression levels between cancer and normal tissues is taken as log2; The gene ranking results of the six quantitative features were integrated using a robust rank aggregation method. The comprehensive score of each gene was calculated, and high-confidence functional genes with RRA test p-values <0.05 were screened. The growth and survival dependence of the high-confidence functional genes were verified using genomic-scale RNA interference screening data and CRISPR screening data from the DepMap database of gastrointestinal cancer cell lines, resulting in a high-confidence functional gene set. (3) Based on the high-confidence functional gene set, the copy number variation data, DNA methylation data, mRNA sequencing data and miRNA sequencing data are screened for each omics feature, and a Transformer-based autoencoder network is used to extract high-dimensional representative features from each omics data. The autoencoder network contains a fully connected layer and a Transformer module and is batch normalized. Then, the high-dimensional representative features of each omics are spliced into a unified feature vector, and fused through a second-level autoencoder network to generate a low-dimensional compact feature vector for comprehensively characterizing the multi-omics characteristics of the sample. (4) Based on the low-dimensional compact feature vector, consensus clustering is performed using the Consensus Cluster Plus tool. The feature matrix is clustered using the PAM algorithm in conjunction with the Pearson correlation distance metric. The clustering effect with a k value of 2-10 is evaluated. The consensus matrix is generated through multiple resampling iterations to quantify the sample co-clustering probability, and finally the molecular subtype of pan-gastrointestinal tumors is determined.
[0006] The above method, further, the digestive tract tumors mentioned in step (1) include any one of esophageal cancer, gastric cancer, pancreatic cancer and colorectal cancer; the multi-omics data are from the TCGA-ESCA, TCGA-STAD, TCGA-PAAD, TCGA-COAD and TCGA-READ projects; batch effect correction is performed using the ComBat algorithm.
[0007] Furthermore, in step (2), the robust rank aggregation method of the R package Robust Rank Aggregator is used to integrate the gene sorting results of the six quantitative features; the growth and survival dependence of the high-confidence functional genes in the cell line is verified by quantifying the DEMETER2 score and CERES score respectively.
[0008] Furthermore, in step (3), for the miRNA sequencing data, low-abundance miRNAs are filtered out, and miRNAs with FPKM≥1 in at least 50% of the samples are retained; the input dimensions of each omics data after screening are as follows: copy number variation data 3770 dimensions, DNA methylation data 3826 dimensions, mRNA data 4253 dimensions, and miRNA data 395 dimensions.
[0009] Furthermore, the high-dimensional representative feature mentioned in step (3) is a 64-dimensional representative feature, and the low-dimensional compact feature vector is a 128-dimensional compact feature vector.
[0010] Furthermore, the autoencoder network in step (3) uses a modified linear unit as the activation function, and the calculation formula for the activation function is: , where x represents the original input value; the autoencoder network is trained using the Adam optimizer, with the mean squared error loss function as the optimization objective. The formula for calculating the mean squared error loss function is: Where N represents the number of samples and x represents the original input value. This indicates that the output value is reconstructed, and the learning rate is set to 2.0 × 10. -3 The autoencoder is based on the PyTorch 1.11.0 framework.
[0011] Furthermore, in step (4), the determination of the molecular subtype of pan-gastrointestinal tumors is determined by comprehensively evaluating the cumulative distribution function curve and the consensus matrix, and the multiple resampling iterations are 1000 resampling iterations.
[0012] Furthermore, the molecular subtypes of pan-gastrointestinal tumors finally identified in step (4) include metabolic GIS-A, immune GIS-B, epithelial GIS-C, and mesenchymal GIS-D.
[0013] Based on a general inventive concept, the present invention also provides a method for constructing a molecular subtype classifier for gastrointestinal tumors across cancer types, comprising the following steps: S1. Collect mRNA sequencing data and corresponding molecular subtypes of gastrointestinal tumor samples, wherein the molecular subtypes are determined by the method described above; S2. Feature screening of mRNA sequencing data based on a high-confidence functional gene set, wherein the high-confidence functional gene set is obtained by the method described above; S3. Using the IGP.clusterRepro algorithm, based on the screened mRNA expression features and corresponding molecular subtypes, the probability of each sample belonging to each molecular subtype is calculated by repeated resampling, and a molecular subtype classifier is trained.
[0014] Based on a general inventive concept, the present invention also provides an application of the method, which is used to perform cross-cancer molecular subtyping of gastrointestinal tumors, and the molecular subtypes determined by the method are used for in vitro drug response prediction for non-diagnostic / therapeutic purposes.
[0015] In the above-described applications, the in vitro drugs are the chemotherapy drug 5-fluorouracil and an anti-immune checkpoint PD-L1 inhibitor.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The method of the present invention integrates six-dimensional molecular features and screens high-confidence functional genes through robust rank aggregation algorithm and RNAi / CRISPR dual verification, which solves the problems of single screening dimension and insufficient specificity in traditional screening. The functional gene screening is more accurate and lays a reliable foundation for subtyping. It breaks through the limitation of tissue origin and realizes unified subtyping of pan-gastrointestinal tumors. Four subtypes were identified through 985 training samples and the stability was verified in 2480 multi-cohort and multi-technology platform samples. It realizes robust and unified subtyping across cancer types and provides a new standardized diagnosis and treatment approach.
[0017] (2) This invention uses a two-stage Transformer autoencoder network to deeply integrate four types of omics data, capture cross-omics complementary information, avoid the limitations of a single omics, accurately characterize the biological essence of tumors, and extract multi-omics features more efficiently.
[0018] (3) The molecular subtypes determined by the typing method of the present invention can be used for in vitro drug response prediction for non-diagnostic / therapeutic purposes, such as clarifying the specific in vitro drug response studies of each subtype to 5-fluorouracil and anti-immune checkpoint PD-L1 inhibitors, providing a scientific basis for risk stratification and personalized medicine research, thereby promoting precision treatment research, and having outstanding clinical guidance value.
[0019] (4) The present invention constructs a classifier based on low-cost and easy-to-operate mRNA sequencing data. After algorithm optimization, the classification is clear, robust, practical and easy to promote. Attached Figure Description To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is an analytical diagram identifying molecular subtypes of pan-gastrointestinal tumors; where: Figure 1 Figure A is the comprehensive functional score chart for six-dimensional feature calculation; Figure 1 Figure B shows the data analysis of functional genes at the genome level using RNAi and CRISPR screening. Figure 1 Figure C is a data analysis diagram that comprehensively evaluates the cumulative distribution function curve and consensus matrix to select the optimal clustering K value; Figure 1 The D-plot is a comparison of prognostic differences among GIS subtypes in the training set samples.
[0021] Figure 2 This is a comparison chart of the results of univariate and multivariate Cox analyses.
[0022] Figure 3 This is a comparative chart showing the prognostic differences among molecular subtypes of pan-gastrointestinal tumors across various tumor types; among which: Figure 3 The AD diagram is a comparison of the prognostic results of the four GIS subtypes in CRC, ESCA, STAD, and PAAD.
[0023] Figure 4 This is a sample verification diagram of the GIS type analyzer; where: Figure 4 Figure A is an informational graph of data from a cohort of 2480 independent gastrointestinal tumors; Figure 4 Figure B shows the prognostic validation results of the GIS subtype in 2480 independent gastrointestinal tumor samples.
[0024] Figure 5 This is a graph showing the genomic characterization differences among molecular subtypes of pan-gastrointestinal tumors; where: Figure 5 Figure A shows the scoring data for SNV, Indel, and TMB of the four GIS subtypes. Figure 5 Figure B is a comparison of the enrichment results of tumor mutation marker DNA mismatch repair signals in GIS-C. Figure 5 Figure C shows the distribution data of microsatellite high instability (MSI-H), high CpG island methylation phenotype (CIMP-H), BRAF mutation, KRAS mutation, APC mutation, and TP53 mutation for the four GIS subtypes. Figure 5 The D-plot is the result of TCGA methylation analysis. Figure 5 Figure E shows the copy number variation distribution of the four GIS subtypes; Figure 5 The F-plot shows the homologous recombination defects and dryness scores of the four GIS subtypes.
[0025] Figure 6 This is a diagram analyzing the biological characteristics of molecular subtypes of pan-gastrointestinal tumors; among which: Figure 6 Figure A shows the results of the assessment of cancer-related biological processes and signaling pathways through gene set enrichment analysis. Figure 6 Figure B is a diagram showing the tumor purity analysis of the four GIS subtypes in the training set queue; Figure 6 Figure C is an analysis of the immune scores of the four GIS subtypes in the training set queue; Figure 6 The D-plot is a graph showing the content of interstitial components in the four GIS subtypes in the training set queue. Figure 6 The E-plot is a diagram showing the tumor purity analysis of the four GIS subtypes in the validation set cohort; Figure 6 The F-plot is an analysis of the immune scores of the four GIS subtypes in the validation set cohort; Figure 6 The G-plot is a graph showing the interstitial component content analysis of the four GIS subtypes in the validation set queue.
[0026] Figure 7This is a treatment correlation analysis diagram of molecular subtypes of pan-gastrointestinal tumors; among which: Figure 7 Figure A shows the correlation analysis of drug response among molecular GIS subtype cell lines of pan-gastrointestinal tumors. Figure 7 Figure B shows the response rate of the molecular GIS subtypes of pan-gastrointestinal tumors to 5-FU. Figure 7 Figure C shows the response rate of the molecular GIS subtypes of pan-gastrointestinal tumors to immune checkpoint PD-L1 inhibitors. Figure 7 The D-plot is a TIDE score map of molecular GIS subtypes of pan-gastrointestinal tumors, reflecting differences in immunotherapy.
[0027] Figure 8 This is a statistical chart showing the treatment response of GIS subtypes to 5-FU and anti-immune checkpoint PD-L1 inhibitors in various gastrointestinal tumors; among them: Figure 8 Figure A shows the response rate of the GIS subtype to 5-FU in colorectal cancer; Figure 8 Figure B shows the response rate of the GIS subtype to 5-FU in gastric cancer. Figure 8 Figure C shows the response rate of the GIS subtype to 5-FU in esophageal cancer; Figure 8 The D-plot shows the response rate of the GIS subtype to 5-FU in pancreatic cancer. Figure 8 Figure E shows the response rate of the GIS subtype to immune checkpoint PD-L1 inhibitors in colorectal cancer. Figure 8 The F-plot represents the response rate of the GIS subtype to immune checkpoint PD-L1 inhibitors in gastric cancer. Figure 8 The G-plot represents the response rate of the GIS subtype of esophageal adenocarcinoma to immune checkpoint PD-L1 inhibitors. Figure 8 The H-plot represents the response rate of the GIS subtype to immune checkpoint PD-L1 inhibitors in esophageal squamous cell carcinoma. Detailed Implementation
[0028] This invention provides a method for molecular subtyping of gastrointestinal tumors across cancer types, comprising the following steps: (1) Obtain multi-omics data of gastrointestinal tumors (from TCGA-ESCA, TCGA-STAD, TCGA-PAAD, TCGA-COAD and TCGA-READ projects), the multi-omics data including copy number variation data, DNA methylation data, mRNA sequencing data and miRNA sequencing data; use the ComBat algorithm to perform batch effect correction on the multi-omics data, and merge to construct a unified multi-omics dataset; the gastrointestinal tumors include any one of esophageal cancer, gastric cancer, pancreatic cancer and colorectal cancer; (2) For each gene, calculate six quantitative feature scores, as follows: Fcnv: Number of samples with copy number mutations / Total number of samples; Fmethy: The number of co-expressed genes that are significantly associated with the target gene; Ftf: The number of transcription factors that have a significant regulatory effect on the target gene; Fmir: Significantly regulates the number of miRNAs; Fmad: Mean absolute deviation across all samples; Fde: The absolute value of log2 of expression level between cancer and normal tissues; The robust rank aggregation method of the R package Robust Rank Aggregator was used to integrate the gene ranking results of the six quantitative features, calculate the comprehensive score of each gene, and screen high-confidence functional genes with RRA test p-values <0.05. Then, using genomic-scale RNA interference screening data and CRISPR screening data of gastrointestinal cancer cell lines from the DepMap database, the growth and survival dependence of the cell lines on the high-confidence functional genes was verified by quantifying the DEMETER2 score and CERES score, respectively, to obtain a set of high-confidence functional genes. (3) Based on the high-confidence functional gene set, feature screening of copy number variation data, DNA methylation data, mRNA sequencing data and miRNA sequencing data of each omics was performed; for the miRNA sequencing data, low-abundance miRNAs were filtered out and miRNAs with FPKM≥1 in at least 50% of the samples were retained; the input dimensions of each omics data after screening were as follows: copy number variation data 3770 dimensions, DNA methylation data 3826 dimensions, mRNA data 4253 dimensions, miRNA data 395 dimensions; and a Transformer-based autoencoder network was used to extract high-dimensional (64-dimensional) representative features from each omics data respectively. The autoencoder network contains fully connected layers and Transformer modules and is batch normalized; then the high-dimensional representative features of each omics were spliced into a unified feature vector, and fused through a second-level autoencoder network to generate a low-dimensional (128-dimensional) compact feature vector for comprehensively characterizing the multi-omics characteristics of the samples; The autoencoder network uses a modified linear unit as the activation function, and the activation function is calculated using the following formula: , where x represents the original input value; The autoencoder network is trained using the Adam optimizer, with the mean squared error loss function as the optimization objective. The formula for calculating the mean squared error loss function is as follows: Where N represents the number of samples and x represents the original input value. This indicates that the output value is reconstructed, and the learning rate is set to 2.0 × 10. -3 The autoencoder is implemented based on the PyTorch 1.11.0 framework. (4) Based on the low-dimensional compact feature vector, consensus clustering is performed using the Consensus Cluster Plus tool. The feature matrix is clustered using the PAM algorithm in conjunction with the Pearson correlation distance metric. The clustering effect with a k value of 2-10 is evaluated. The consensus matrix is generated after 1000 resampling iterations to quantify the sample co-clustering probability. The molecular subtypes of pan-gastrointestinal tumors are finally determined by comprehensively evaluating the cumulative distribution function curve and the consensus matrix, including: metabolic GIS-A, immune GIS-B, epithelial GIS-C and mesenchymal GIS-D.
[0029] A method for constructing a cross-cancer molecular subtype classifier for gastrointestinal tumors includes the following steps: S1. Collect mRNA sequencing data and corresponding molecular subtypes of gastrointestinal tumor samples, wherein the molecular subtypes are determined by the method described above; S2. Feature screening of mRNA sequencing data based on a high-confidence functional gene set, wherein the high-confidence functional gene set is obtained by the method described above; S3. Using the IGP.clusterRepro algorithm, based on the screened mRNA expression features and corresponding molecular subtypes, the probability of each sample belonging to each molecular subtype is calculated by repeated resampling, and a molecular subtype classifier is trained.
[0030] An application of a method for molecular subtyping of gastrointestinal tumors across cancer types, wherein the method is used to perform molecular subtyping of gastrointestinal tumors across cancer types, and the molecular subtypes determined by the method are used for in vitro drug response prediction for non-diagnostic / therapeutic purposes, wherein the in vitro drugs are the chemotherapy drug 5-fluorouracil and an anti-immune checkpoint PD-L1 inhibitor.
[0031] To facilitate understanding of the present invention, the present invention will be described more fully and in detail below with reference to the accompanying drawings and preferred embodiments, but the scope of protection of the present invention is not limited to the following specific embodiments.
[0032] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by those skilled in the art. The technical terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of the invention.
[0033] Unless otherwise specified, all raw materials, reagents, instruments and equipment used in this invention can be purchased from the market or prepared by existing methods.
[0034] Example: A molecular subtyping method for gastrointestinal tumors across cancer types includes the following steps: 1. Screening functional genes based on multi-omics data This invention first integrates multi-omics data from the TCGA-ESCA, TCGA-STAD, TCGA-PAAD, TCGA-COAD, and TCGA-READ projects. For each type of omics data, a unified TCGA-GI dataset is constructed by merging the five projects, and batch effects are corrected using the ComBat algorithm. To identify genes with high functional relevance, a computational framework integrating six molecular features is developed, which reflect the activity and regulatory mechanisms of genes at different levels in gastrointestinal tumors. The feature screening and integration process includes the following three steps: (1) Feature calculation: For each gene, six quantitative feature scores are calculated according to the definitions in Table 1 to characterize genome / transcriptome activity and the regulatory role of co-expressed transcription factors and microRNAs; (2) Score integration: The robust rank aggregation method of the R package Robust Rank Aggregator is used to integrate the gene ranking of the six features, and then calculate the comprehensive score of each gene; genes with RRA test p value <0.05 are defined as high-confidence functional genes; (3) Validation analysis: Genomic-scale RNA interference screening data and CRISPR screening data of gastrointestinal cancer cell lines were obtained from the DepMap database. The DEMETER2 score (RNAi) and CERES score (CRISPR) were used to quantify gene dependence, respectively. It should be noted that a lower score indicates that the cell line is more dependent on the growth and survival of the gene.
[0035] Table 1. Description of the six characteristics
[0036] 2. Deep Learning Network Architecture All network nodes use the Rectified Linear Unit (ReLU) activation function, and its calculation formula is shown in Equation (1), where x represents the input value; (1) The model was trained using the Adam optimizer, with the mean squared error loss function as the optimization objective, calculated as shown in equation (2). The learning rate was set to 2.0 × 10⁻⁶. -3 In the formula, N represents the number of samples, and x represents the original input. This indicates the reconstruction output value.
[0037] (2) The autoencoder is based on the PyTorch 1.11.0 framework.
[0038] 3. Multi-omics feature extraction based on deep learning To identify subtypes of pan-gastrointestinal tumors, this invention analyzed 985 samples with complete data from four omics categories, including CNV, DNA methylation, mRNA sequencing, and miRNA sequencing. Based on the previously determined functional gene set, feature screening was performed on CNV, DNA methylation, and mRNA sequencing expression profiles. For miRNA sequencing data, low-abundance miRNAs were filtered out, retaining miRNAs with FPKM≥1 in at least 50% of the samples. The final input dimensions for each omics data were: CNV 3770 dimensions, methylation 3826 dimensions, mRNA 4253 dimensions, and miRNA 395 dimensions. A two-stage deep learning framework is employed to learn integrated latent representations from selected multi-omics data: (1) First stage: Single-omics feature learning Representative features are extracted from each omics dataset using a Transformer-based autoencoder network. Each encoder subnetwork consists of fully connected layers and Transformer modules to enhance the ability to model complex feature relationships. Batch normalization is used to prevent overfitting, and the decoder network has a symmetrical structure to the corresponding encoder. This process generates a 64-dimensional feature representation for each omics dataset. (2) Second stage: Integration of cross-omics features The 64-dimensional feature representations of the four omics are concatenated into a unified feature vector, which is then fused through a second-level autoencoder network to learn a low-dimensional representation that can capture complementary information from all omics patterns. Finally, a fully connected layer autoencoder generates a compact 128-dimensional feature vector to comprehensively represent the multi-omics characteristics of each sample for downstream clustering analysis.
[0039] 4. Identification of subtypes of pan-gastrointestinal tumors Based on the integrated 128-dimensional feature spectrum, this invention uses the R package Consensus Cluster Plus to perform consensus clustering to identify the molecular subtypes of pan-gastrointestinal tumors. To determine the optimal number of clusters, the clustering effect of k values from 2 to 10 was evaluated. The PAM algorithm combined with Pearson correlation distance metric was used to cluster the feature matrix, and the clustering stability was evaluated through 1000 resampling iterations—the probability of sample co-clustering was quantified by generating a consensus matrix.
[0040] 5. Training the GIS classifier In the previous subtype identification study, this invention integrated 985 TCGA samples with complete data from all four omics categories. Given the advantages of low cost and convenient operation of transcriptome sequencing technology, this invention selected mRNA sequencing data to construct a GIS classifier. Based on the previous feature screening results, 4253 genes were finally determined as input features of the autoencoder network. For each gene, its average expression level in all samples of each GIS subtype was calculated.
[0041] The classifier is trained using the `IGP.clusterRepro` function from the R package `clusterRepro`. It quantifies the frequency of each sample co-clustering with a specific subtype during repeated resampling by calculating the within-group proportion statistic, thus assigning samples to the most probable molecular subtype based on consensus principles. The algorithm calculates the assignment proportion of each sample-subtype pair through multiple resampling iterations, ultimately assigning the sample to the subtype that yielded the highest IGP value.
[0042] 6. Genomic and biological characteristic analysis This invention analyzes the biological characteristics of various GIS subtypes by integrating multidimensional molecular profiling data systems. First, GSEA analysis is performed based on subtype-specific differentially expressed genes to identify biological processes and signaling pathways significantly enriched in various subtypes of pan-gastrointestinal tumors. At the genomic variation level, tumor mutation burden and neoantigen burden data from TCGA samples are used in conjunction with the COSMIC database to analyze mutational characteristics and reveal mutational processes specific to immune subtypes.
[0043] To assess genomic characteristics, the GISTIC 2.0 algorithm was used to calculate sample-specific copy number variation scores, and homologous recombination defect scores were obtained from the UCSC PANCAN cohort. At the epigenetic level, subtype-specific high / low methylation genes were identified based on DNA methylation β values from the Illumina Infinium HumanMethylation 450 platform, while referencing CpG island methylation phenotypes determined in previous TCGA studies.
[0044] 7. Tumor Microenvironment Analysis The ESTIMATE method was used to infer the cellular composition of the tumor microenvironment. This algorithm used a single-sample GSEA method to generate matrix scores, immune scores, and tumor environment scores. For each sample, gene expression values were first rank-normalized and sorted, and then the empirical cumulative distribution functions of the characteristic gene set and the remaining genes were calculated separately. Statistical significance was calculated by integrating the difference between the two empirical cumulative distribution functions—this method is similar to GSEA in principle, but focuses on absolute expression levels rather than differential expression analysis.
[0045] 8. Survival Analysis To assess prognostic differences among different GIS subtypes, this invention conducts survival analysis based on overall survival. Patient survival status and follow-up time data were obtained from the TCGA database. Overall survival was defined as the time interval from diagnosis to death from any cause. For patients lost to follow-up, their last follow-up date was used as truncated data. Survival curves for each GIS subtype were plotted using the Kaplan-Meier method, and the statistical significance of survival differences between subtypes was assessed using the log-rank test. Based on this, a univariate Cox proportional hazards regression model was constructed to calculate the hazard ratio of each subtype relative to a reference subtype and its 95% confidence interval. To exclude the influence of potential confounding factors, a multivariate Cox model was further established, adjusting for key clinicopathological features such as microsatellite instability status, BRAF, and KRAS mutations.
[0046] The results are as follows: 1. Conserved molecular subtypes of pan-gastrointestinal tumors transcend tissue origin limitations To define subtypes of pan-gastrointestinal tumors, this invention first screened a group of core function-related genes for subsequent multi-omics integration. A novel evaluation framework was established, prioritizing genes based on six characteristics: genomic / transcriptomic activity, co-expressed transcription factors, and microRNA regulation. Compared to ordinary genes, cancer marker genes showed significantly higher scores across all six dimensions, validating the effectiveness of this method. Figure 1 A). This invention further integrates six-dimensional feature calculation to obtain a comprehensive functional score, and finds that this multi-dimensional scoring system can more effectively distinguish cancer marker genes from ordinary genes. Ultimately, this invention identified 4357 high-confidence functional genes. Analysis of genome-scale RNAi and CRISPR screening data showed that gastrointestinal cancer cell lines have a significantly stronger dependence on these functional genes for growth and survival, confirming their crucial role in gastrointestinal tumor biology. Figure 1 B).
[0047] This invention selected 985 tumor samples with CNV, methylation, mRNA, and miRNA expression profiles for pan-gastrointestinal tumor molecular subtype (GIS) identification. Based on a predefined functional gene set, 128 potential multi-omics features were extracted using a Transformer-based autoencoder network for unsupervised clustering. By comprehensively evaluating the cumulative distribution function curve and consensus matrix, the optimal number of subtypes was determined to be 4, named GIS-A, GIS-B, GIS-C, and GIS-D. Figure 1 C).
[0048] After adjusting for key clinicopathological features such as microsatellite instability and BRAF / KRAS mutations, multivariate Cox regression analysis showed that patients with the GIS-C subtype had the best overall survival, while patients with the GIS-D subtype had the worst prognosis. Figure 1 D; Figure 2 ).
[0049] More importantly, this prognostic difference was consistently validated across various gastrointestinal tumors, indicating that GIS typing captures a biological essence that transcends tissue origin. Figure 3 AD).
[0050] To facilitate clinical application, this invention developed a GIS genotype based on TCGA mRNA sequencing data. This invention collected 2480 gastrointestinal tumor samples from 17 independent cohorts involving 11 technology platforms to validate the GIS classifier. Figure 4 A). The results showed that the genotype reliably reproduced prognostic associations: patients with the GIS-D subtype had significantly shorter overall survival, and this subtype-specific survival pattern was confirmed across all cancer types. These results consistently validated the robust prognostic predictive ability of GIS genotyping across different patient populations and technology platforms. Figure 4 B).
[0051] 2. Unique genomic and biological characteristics define GIS subtypes Through comprehensive (epigeonal)genomic analysis, this invention reveals characteristic aberration patterns in four GIS subtypes. The GIS-B subtype exhibits an immunogenic phenotype, demonstrating a significantly higher tumor mutational burden (TMB) and neoantigen burden derived from SNV and Indel variants. Figure 5 A). This hypermutated state is significantly enriched with defects in DNA mismatch repair (MMR) signaling. Figure 5 B) is further reflected in the high microsatellite instability (MSI-H), high CpG island methylation phenotype (CIMP-H), and a high proportion of BRAF mutant samples in this subtype. Figure 5 C). Although the GIS-A subtype also shows a high MSI-H frequency, its main characteristic is a high KRAS mutation rate ( Figure 5 C). According to the cBioPortal public database, KRAS mutation is one of the most common pathogenic gene mutations in MSI-H type colorectal cancer and gastric adenocarcinoma. Consistent with CIMP-H status, TCGA methylation profiling analysis showed that GIS-B tumors exhibited widespread hypermethylation, while GIS-C tumors showed predominantly hypomethylated status. Figure 5 D), and has a higher frequency copy number gain ( Figure 5 E). Furthermore, GIS-C tumors showed significantly elevated homologous recombination deficiency (HRD) and stemness scores (E). Figure 5 This aligns with the theoretical model that HRD promotes genomic instability and chromosomal abnormalities, thereby inducing a stem cell-like state.
[0052] Through a systematic evaluation of cancer-related biological processes and signaling pathways using gene set enrichment analysis (GSEA), this invention reveals that each GIS subtype exhibits unique biological characteristics. Figure 6 A) The GIS-A subtype is significantly enriched in multiple metabolic pathways and protein secretion features; the GIS-B subtype is characterized by a strong anti-tumor immune response, including extensive infiltration of natural killer cells and cytotoxic T cells and activation of immune signaling pathways; the GIS-C subtype is specifically enriched in cell proliferation processes (such as DNA replication, E2F targets, and TP53 signaling pathways) and exhibits characteristics of a mucinous cell type; the GIS-D subtype is characterized by epithelial-mesenchymal transition (EMT), TGF-β signaling pathway activation, and specific immune-related pathways, consistent with an inflammatory mesenchymal phenotype.
[0053] Analysis of the tumor microenvironment (TME) using the ESTIMATE algorithm revealed that GIS-A and GIS-C tumors exhibited significantly higher tumor purity, consistent with their epithelial characteristics. Figure 6 B); GIS-B tumors show a significantly elevated immune score ( Figure 6 C); while GIS-D tumors are rich in stromal components ( Figure 6 D). These unique TME features were robustly reproduced in independent validation queues. Figure 6 EG).
[0054] Based on genomic and biological characteristic analysis, this invention defines GIS-A as metaboloid, GIS-B as immune type, GIS-C as epithelial type, and GIS-D as mesenchymal type.
[0055] 3. GIS subtypes predict cross-tissue origin drug responses To evaluate the therapeutic guidance value of the classification system, this invention first validated the predictive ability of GIS subtypes for in vitro drug sensitivity using the CCLE database. Analysis showed that the correlation of drug response among cell lines with the same GIS subtype was significantly higher than that among cell lines of the same histological origin. Figure 7 (A) This demonstrates that pan-gastrointestinal subtypes are better predictors of drug response than tissue origins, supporting the rationale for subtype-guided treatment.
[0056] In clinical data validation, this invention focuses on two widely used therapies for gastrointestinal tumors: the core chemotherapy drug 5-fluorouracil (5-FU) and an anti-immune checkpoint PD-L1 inhibitor. The GIS-C subtype, characterized by proliferative epithelial molecules, showed a significantly higher response rate to 5-FU than other subtypes (response rate: 66% vs 47%). Figure 7 B), and this pattern is consistently observed in various types of gastrointestinal tumors. Figure 8AD). Conversely, the GIS-B subtype, rich in immune microenvironment and high TMB, showed significantly better response rates against immune checkpoint PD-L1 inhibitors (response rate: 84% vs 49%). Figure 7 C) indicates that its immunogenicity enhances sensitivity to immune checkpoint blockade. Independent validation using the TIDE algorithm revealed that the GIS-B subtype exhibited significantly lower TIDE scores in most datasets, predicting a better response to immune checkpoint inhibitors. Figure 7 D). This subtype-specific response pattern has been confirmed in all cancer types (except for pancreatic cancer, where data is insufficient). Figure 8 EH), further confirming the clinical application potential of GIS classification in precision treatment.
Claims
1. A method for molecular subtyping of gastrointestinal tumors across cancer types, characterized in that, Includes the following steps: (1) Obtain multi-omics data of gastrointestinal tumors, including copy number variation data, DNA methylation data, mRNA sequencing data and miRNA sequencing data; perform batch effect correction on the multi-omics data and merge them to construct a unified multi-omics dataset; (2) For each gene, calculate six quantitative feature scores, as follows: Fcnv: Number of samples with copy number mutations / Total number of samples; Fmethy: The number of co-expressed genes that are significantly associated with the target gene; Ftf: The number of transcription factors that have a significant regulatory effect on the target gene; Fmir: Significantly regulates the number of miRNAs; Fmad: Mean absolute deviation across all samples; Fde: The absolute value of the fold difference in expression levels between cancer and normal tissues is taken as log2; The gene ranking results of the six quantitative features were integrated using a robust rank aggregation method. The comprehensive score of each gene was calculated, and high-confidence functional genes with RRA test p-values <0.05 were screened. The growth and survival dependence of the high-confidence functional genes were verified using genomic-scale RNA interference screening data and CRISPR screening data from the DepMap database of gastrointestinal cancer cell lines, resulting in a high-confidence functional gene set. (3) Based on the high-confidence functional gene set, the copy number variation data, DNA methylation data, mRNA sequencing data and miRNA sequencing data are screened for each omics feature, and a Transformer-based autoencoder network is used to extract high-dimensional representative features from each omics data. The autoencoder network contains a fully connected layer and a Transformer module and is batch normalized. Then, the high-dimensional representative features of each omics are spliced into a unified feature vector, and fused through a second-level autoencoder network to generate a low-dimensional compact feature vector for comprehensively characterizing the multi-omics characteristics of the sample. (4) Based on the low-dimensional compact feature vector, consensus clustering is performed using the Consensus Cluster Plus tool. The feature matrix is clustered using the PAM algorithm in conjunction with the Pearson correlation distance metric. The clustering effect with a k value of 2-10 is evaluated. The consensus matrix is generated through multiple resampling iterations to quantify the sample co-clustering probability, and finally the molecular subtype of pan-gastrointestinal tumors is determined.
2. The method according to claim 1, characterized in that, The digestive tract tumors mentioned in step (1) include any one of esophageal cancer, gastric cancer, pancreatic cancer, and colorectal cancer; the multi-omics data are from the TCGA-ESCA, TCGA-STAD, TCGA-PAAD, TCGA-COAD, and TCGA-READ projects; batch effect correction is performed using the ComBat algorithm.
3. The method according to claim 1, characterized in that, In step (2), the robust rank aggregation method of the R package Robust Rank Aggregate is used to integrate the gene sorting results of the six quantitative features; the growth and survival dependence of the high-confidence functional genes in the cell line is verified by quantifying the DEMETER2 score and CERES score respectively.
4. The method according to claim 1, characterized in that, In step (3), for the miRNA sequencing data, low-abundance miRNAs are filtered out, and miRNAs with FPKM≥1 in at least 50% of the samples are retained; the input dimensions of each omics data after screening are as follows: copy number variation data 3770 dimensions, DNA methylation data 3826 dimensions, mRNA data 4253 dimensions, and miRNA data 395 dimensions.
5. The method according to claim 1, characterized in that, The high-dimensional representative feature mentioned in step (3) is a 64-dimensional representative feature, and the low-dimensional compact feature vector is a 128-dimensional compact feature vector.
6. The method according to claim 1, characterized in that, The autoencoder network described in step (3) uses a modified linear unit as the activation function, and the formula for calculating the activation function is as follows: , where x represents the original input value; the autoencoder network is trained using the Adam optimizer, with the mean squared error loss function as the optimization objective. The formula for calculating the mean squared error loss function is: Where N represents the number of samples and x represents the original input value. This indicates that the output value is reconstructed, and the learning rate is set to 2.0 × 10. -3 The autoencoder is based on the PyTorch 1.11.0 framework.
7. The method according to claim 1, characterized in that, In step (4), the molecular subtype of pan-gastrointestinal tumors is determined by comprehensively evaluating the cumulative distribution function curve and consensus matrix, and the multiple resampling iterations are 1000 resampling iterations.
8. The method according to claim 1, characterized in that, The molecular subtypes of pan-gastrointestinal tumors finally identified in step (4) include metabolic GIS-A, immune GIS-B, epithelial GIS-C, and stromal GIS-D.
9. A method for constructing a cross-cancer molecular subtype classifier for gastrointestinal tumors, characterized in that, Includes the following steps: S1. Collect mRNA sequencing data and corresponding molecular subtypes of gastrointestinal tumor samples, wherein the molecular subtypes are determined by the method of any one of claims 1-8; S2. Feature screening of mRNA sequencing data based on a high-confidence functional gene set, wherein the high-confidence functional gene set is obtained by the method described in any one of claims 1-8; S3. Using the IGP.clusterRepro algorithm, based on the screened mRNA expression features and corresponding molecular subtypes, the probability of each sample belonging to each molecular subtype is calculated by repeated resampling, and a molecular subtype classifier is trained.
10. An application of the method as described in any one of claims 1-8, characterized in that, The method described above is used for molecular subtyping of gastrointestinal tumors across cancer types, and the molecular subtypes determined by the method are used for in vitro drug response prediction for non-diagnostic / therapeutic purposes.