Graph convolutional networks for identifying and quantifying gene and cancer-specific transcriptome signatures of cancer driver events.

Graph convolutional networks analyze RNA expression data to identify cancer driver events and predict drug efficacy, addressing the limitations of DNA-level mutation analysis by providing accurate and personalized treatment options.

JP2026528719APending Publication Date: 2026-08-25HADASIT MEDICAL RESEARCH SERVICES & DEVELOPMENT LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026504778
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-01
Filing Date
2024-07-31
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Current clinical practices primarily rely on DNA-level mutation analysis, which often fails to identify treatable cancer driver events, leaving many cancer patients without effective targeted therapies.

Method used

A method and system using graph convolutional networks (GCNs) to analyze RNA expression data, integrating biological knowledge, for identifying and scoring genes as mutant or wild-type, and predicting drug efficacy based on gene expression patterns.

Benefits of technology

The system accurately predicts cancer driver events and drug effectiveness, enabling personalized treatment strategies by identifying previously undetected driver mutations and quantifying their impact on patient survival.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026528719000001_ABST
    Figure 2026528719000001_ABST
Patent Text Reader

Abstract

This disclosure describes a machine learning (ML) framework, including a graph convolutional neural network (GCN), for identifying gene expression signatures associated with cancer driver events. The model is trained to identify the TP53 mutation status of cancer samples from gene expression, utilizing a comprehensive, curated graph structure of gene interactions. Quantitative scores are generated to rank the severity of driver events in each sample. Very high AUC results for unknown data across several tumor types are achieved in this method. A strong correlation with protein function exists. The Signature in Transcriptome Associated with Mutant Proteins (STAMP) model can also predict driver events in many combinations of key oncogenes / pathways and several tumor types, based on well-established annotations from the literature. Thus, the STAMP model can identify and quantify driver events, which may lead to improved targeted therapy selection and prioritization in cancer patients.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Cross-reference of related applications) This application claims the benefits and priority of U.S. Provisional Patent Application No. 63 / 516,961, filed on 1 August 2023, which is incorporated in its entirety by reference. [Background technology]

[0002] This invention relates to a method and system for identifying cancer driver events.

[0003] Identifying and targeting cancer-causing (driver) genetic mutations has seen dramatic improvements in recent years, leading to the development of many new targeted therapies. However, identifying, prioritizing, and treating genetic mutations remains insufficient for the vast majority of cancer patients. Current clinical practice relies primarily on DNA-level mutation analysis, which often fails to identify treatable driver events. The use of the transcriptome, a complex and highly informative representation of cellular and tumor conditions, would enhance diagnostic and therapeutic success.

[0004] Therefore, there is a growing need for technologies that provide methods for identifying cancer driver genes in mRNA (transcriptome). [Overview of the project]

[0005] Aspects of this disclosure relate to systems and methods for identifying and scoring samples that contain genes.

[0006] This disclosure describes a method and system for classifying genes as mutant or wild-type and assigning a score associated with the classification to the gene. The method further includes receiving a set of curated data relating to human genes, wherein the curated data is RNA expression data. The curated data is divided into first, second, and third datasets. A classifier is fed into the first dataset to identify the data as mutant or wild-type. The classifier is validated against the second dataset and tested against the third dataset. The method further includes receiving new data and applying a trained classifier to the newly received data to classify the newly received data as mutant or wild-type. The method calculates a score associated with the newly received data and reports the classification and score associated with the newly received data to a user device.

[0007] The classifier may include, for example, a graph convolutional network, a random forest model, or a support vector machine model. The score may be the reciprocal of the activation function of the second-to-last layer of the classifier. The score may be used to predict the efficacy of a drug against a cell line or tumor. The efficacy prediction is based on the drug's informed consent (IC). 50 This method may include estimating the active molecular pathway and the drugs associated with that pathway. This method may further include diagnosing patients with a certain condition based on a score. This method may further include suggesting treatment for patients based on the diagnosis and score.

[0008] A method for applying a trained classifier to suggest treatments for medical conditions is described. This method receives patient genetic data and applies a trained classifier to classify the received data as either mutant or wild-type. The method further calculates a score associated with the received data. This method predicts the effectiveness of a drug against the cell line or tumor associated with the received data. The method further reports the classification, calculated score, drug, and predicted effectiveness to the user device.

[0009] The classifier may include, for example, a graph convolutional network, a random forest model, or a support vector machine model. The score may be the reciprocal of the activation function of the second-to-last layer of the classifier. The method may further include identifying active molecular pathways and drugs associated with those active molecular pathways.

[0010] For example, a system for calculating a score is described, comprising a processor and computer memory. The computer memory contains computer instructions that, when executed by the processor, perform a method comprising a series of steps. One of these steps may include receiving a set of curated data relating to human genes, wherein the curated data is RNA expression data. The curated data is divided into a first dataset, a second dataset, and a third dataset. A classifier is trained on the first dataset to identify data as either mutant or wild-type. The classifier is validated on the second dataset and tested on the third dataset. The system receives new data and applies the trained classifier to the new data to classify the new data as either mutant or wild-type. In addition, the system calculates a score associated with the received new data. The score and classification associated with the new data may be reported to a user device.

[0011] The classifier may include a graph convolutional network, a random forest model, or a support vector machine model. The score may be the reciprocal of the activation function of the second last layer of the classifier. The score may be used to predict the effectiveness of a medicine for a cell line or a tumor. The prediction of effectiveness may include estimating the IC 50 of the drug. The system may further include identifying active molecular pathways and drugs associated with the active molecular pathways. The system may further include diagnosing a patient having a certain condition based on the score. The system may further include suggesting a treatment for the patient based on the diagnosis and the score.

Brief Description of the Drawings

[0012] To better understand the above features of the present invention, a more specific description of the present invention briefly summarized above can be obtained by referring to the embodiments, some of which are shown in the accompanying drawings. However, it should be noted that the accompanying drawings show only typical embodiments of the present invention, and the present invention can be implemented in other equally effective embodiments. [Figure 1] Shows the workflow diagram of the entire method of the present disclosure. [Figure 2A] Shows the results of driver event identification for various samples. [Figure 2B] Shows the results of driver event identification for various samples. [Figure 2C] Shows the results of driver event identification for various samples. [Figure 2D] Shows the results of driver event identification for various samples. [Figure 2E] Shows the results of driver event identification for various samples. [Figure 2F] Shows the results of driver event identification for various samples. [Figure 3A] Shows the verification of the model. [Figure 3B] Shows the verification of the model. [Figure 3C]This demonstrates the validation of the model. [Figure 3D] This demonstrates the validation of the model. [Figure 3E] This demonstrates the validation of the model. [Figure 4A] This shows the gene and tumor type matrices in which the model predicted superior results. [Figure 4B] This shows the gene and tumor type matrices in which the model predicted superior results. [Figure 4C] This shows the gene and tumor type matrices in which the model predicted superior results. [Figure 4D] This shows the gene and tumor type matrices in which the model predicted superior results. [Figure 4E] This shows the gene and tumor type matrices in which the model predicted superior results. [Figure 4F] This shows the gene and tumor type matrices in which the model predicted superior results. [Figure 4G] This shows the gene and tumor type matrices in which the model predicted superior results. [Figure 4H] This shows the gene and tumor type matrices in which the model predicted superior results. [Figure 4I] This shows the gene and tumor type matrices in which the model predicted superior results. [Figure 5A] The results of applying a model to identify driver genes are shown. [Figure 5B] The results of applying a model to identify driver genes are shown. [Figure 5C] The results of applying a model to identify driver genes are shown. [Figure 5D] The results of applying a model to identify driver genes are shown. [Figure 5E] The results of applying a model to identify driver genes are shown. [Figure 5F] The results of applying a model to identify driver genes are shown. [Figure 5G] The results of applying a model to identify driver genes are shown. [Figure 5H] The results of applying a model to identify driver genes are shown. [Figure 6A] An example of a STAMP model for stratifying treatment responses in a clinical cohort is shown. [Figure 6B] An example of a STAMP model for stratifying treatment responses in a clinical cohort is shown. [Figure 6C] An example of a STAMP model for stratifying treatment responses in a clinical cohort is shown. [Figure 6D] An example of a STAMP model for stratifying treatment responses in a clinical cohort is shown. [Figure 6E] An example of a STAMP model for stratifying treatment responses in a clinical cohort is shown. [Figure 6F] An example of a STAMP model for stratifying treatment responses in a clinical cohort is shown. [Figure 7A] An example of a STAMP model for stratifying treatment responses in a clinical cohort is shown. [Figure 7B] An example of a STAMP model for stratifying treatment responses in a clinical cohort is shown. [Figure 7C] An example of a STAMP model for stratifying treatment responses in a clinical cohort is shown. [Figure 7D] An example of a STAMP model for stratifying treatment responses in a clinical cohort is shown. [Figure 7E] An example of a STAMP model for stratifying treatment responses in a clinical cohort is shown. [Figure 7F] An example of a STAMP model for stratifying treatment responses in a clinical cohort is shown. [Figure 8] This shows the downsampling analysis of ELR (MSigDB). [Figure 9A]This paper demonstrates the prediction of TP53 mutation status and protein function based on the concepts of signatures in transcriptome associated with mutant proteins (STAMP) and "false" mutants or WT samples. [Figure 9B] This paper demonstrates the prediction of TP53 mutation status and protein function based on the concepts of signatures in transcriptome associated with mutant proteins (STAMP) and "false" mutants or WT samples. [Figure 9C] This paper demonstrates the prediction of TP53 mutation status and protein function based on the concepts of signatures in transcriptome associated with mutant proteins (STAMP) and "false" mutants or WT samples. [Figure 10] This shows a multivariate analysis of the Cox proportional hazards model for the TCGA pan-cancer cohort. [Figure 11A] Further survival analysis based on mutation or non-mutation status is presented. [Figure 11B] Further survival analysis based on mutation or non-mutation status is presented. [Figure 11C] Further survival analysis based on mutation or non-mutation status is presented. [Figure 11D] Further survival analysis based on mutation or non-mutation status is presented. [Figure 11E] Further survival analysis based on mutation or non-mutation status is presented. [Figure 12A] Further SBS analysis is shown. [Figure 12B] Further SBS analysis is shown. [Figure 12C] Further SBS analysis is shown. [Figure 12D] Further SBS analysis is shown. [Figure 12E]Further SBS analysis is shown. [Figure 12F] Further SBS analysis is shown. [Figure 12G] Further SBS analysis is shown. [Figure 12H] Further SBS analysis is shown. [Figure 12I] Further SBS analysis is shown. [Figure 13A] Further SBS analysis is shown. [Figure 13B] Further SBS analysis is shown. [Figure 13C] Further SBS analysis is shown. [Figure 13D] Further SBS analysis is shown. [Figure 13E] Further SBS analysis is shown. [Figure 13F] Further SBS analysis is shown. [Figure 14A] This shows a comparison between STAMP and TP53_PROF. [Figure 14B] This shows a comparison between STAMP and TP53_PROF. [Figure 14C] This shows a comparison between STAMP and TP53_PROF. [Figure 14D] This shows a comparison between STAMP and TP53_PROF. [Figure 14E] This shows a comparison between STAMP and TP53_PROF. [Figure 14F] This shows a comparison between STAMP and TP53_PROF. [Figure 14G] This shows a comparison between STAMP and TP53_PROF. [Figure 14H] This shows a comparison between STAMP and TP53_PROF. [Figure 14I] This shows a comparison between STAMP and TP53_PROF. [Figure 14J] This shows a comparison between STAMP and TP53_PROF. [Figure 14K] This shows a comparison between STAMP and TP53_PROF. [Figure 15A] This paper demonstrates the application and validation of STAMP trained through various pathways in several tumor type cohorts. [Figure 15B] This paper demonstrates the application and validation of STAMP trained through various pathways in several tumor type cohorts. [Figure 15C] This paper demonstrates the application and validation of STAMP trained through various pathways in several tumor type cohorts. [Figure 15D] This paper demonstrates the application and validation of STAMP trained through various pathways in several tumor type cohorts. [Figure 16A] Further GDSC analysis is presented. [Figure 16B] Further GDSC analysis is presented. [Figure 16C] Further GDSC analysis is presented. [Figure 16D] Further GDSC analysis is presented. [Figure 16E] Further GDSC analysis is presented. [Figure 16F] Further GDSC analysis is presented. [Figure 16G] Further GDSC analysis is presented. [Figure 16H] Further GDSC analysis is presented. [Figure 16I] Further GDSC analysis is presented. [Figure 17] The results of different graph structures used to train STAMP for breast cancer based on TP53 mutation status are shown. [Figure 18A] This shows that there is little correlation between the score and IC50 when identified by the amplification state. [Figure 18B] This shows that there is little correlation between the score and IC50 when identified by the amplification state. [Figure 19] This document describes a computing device that may be used to carry out the teachings of the present invention.

[0013] Other features of this embodiment will become apparent from the following embodiments for carrying out the invention. [Modes for carrying out the invention]

[0014] The following detailed description of preferred embodiments refers to the accompanying drawings, which form part of this specification and illustrate specific embodiments in which the invention may be carried out. It should be understood that other embodiments may be utilized and structural modifications may be made without departing from the scope of the invention. Electrical, mechanical, logical, and structural modifications may be made to embodiments without departing from the spirit and scope of this teaching. Therefore, the following detailed description should not be construed as restrictive, and the scope of this disclosure is defined by the accompanying claims and their equivalents.

[0015] This disclosure relates to a method and system (transcriptome) for identifying driver events in mRNA.

[0016] Cancer diagnosis and treatment have experienced significant breakthroughs in recent years with the advent of precision medicine. Driver mutations in oncogenes have been identified and targeted by novel therapies, leading to significant improvements in clinical outcomes. However, only a small fraction of cancer patients have benefited from such treatments, and many questions have remained unanswered. Three such questions form the core of the invention described herein. First, low-frequency or poorly characterized variants in key oncogenes occur in a significant proportion of cancer patients. The effects of such variants, whether they contribute to tumorigenesis (driver mutations) or not (passenger mutations), are largely unknown. Second, some mutations can result in a spectrum of partial defects in protein activity, further complicating attempts to identify and prioritize their effects. Third, for some patients, mutation analysis identifies only a few mutational events, or none at all. This suggests that other molecular-level analyses, such as epigenetics, transcriptomics, and protein-protein interactions, are needed to complement mutational effect classification.

[0017] Gene expression, measured through mRNA sequencing, simultaneously captures a picture of the activity of all genes in a sample, thus enabling a rich and complex representation of tumor status. Transcriptional profiles have also been shown to hold valuable information about genetic abnormalities, clinical traits, and therapeutic approaches. The high complexity of transcriptomics presents a significant challenge to its analysis, and reaching accurate and unbiased conclusions is by no means easy. To address this complexity, several approaches have been developed, such as path analysis and network analysis. Machine learning (ML) and deep learning (DL) models can be used to address this challenge by focusing on analysis and generating valuable insights. In conventional ML, pre-analysis of the data extracts relevant features for learning (a process called feature selection). In the field of DL, algorithmic architectures targeting specific complex tasks have been developed. For example, convolutional neural networks (CNNs) successfully learn from images by using the spatial proximity of pixels to assume similarity of pixel contents.

[0018] Graph structures such as Funcoup and HumanNet v2 contain human genome representation maps that define various interactions between genes. These graphs accumulate data from metabolic or signaling pathway information, complex and physical protein-protein interactions (PPIs), genomic context, and co-expression patterns into a single graph structure. Graph Convolutional Neural Networks (GCNs) are DL architectures that use graph structures and graph node proximity to focus on learning neural networks in a manner similar to CNNs, which use the spatial proximity of image pixels. Therefore, focusing on learning from transcriptomics by integrating established biological knowledge in graph form could serve as an ideal architecture. Other researchers have investigated the ability of GCNs to predict masked gene expression levels using expression values ​​and connectivity maps of their graph-adjacent genes. A similar approach could be applied to identify driver event signatures from gene expression.

[0019] Multiple ML and DL models for cancer genomics have been developed in recent years. Several studies use mutation or transcriptome data from cell line studies to predict their responses to various drugs. One recent approach used an innovative architecture with a bifurcated neural network on a cell line mutation atlas to predict their responses to treatment and even identify novel drug synergies. Predicting driver events from DNA mutation data is an active research area in itself, and very accurate models have been developed in recent years. Similar studies using only transcriptome data are less frequent. Some of these studies have pointed out that training models to predict driver events can benefit from gene-specific and tumor-specific approaches. Despite the reduced sample size, the benefits of such focus may be explained by the reduced complexity of the learned properties. Whether this approach is beneficial for transcriptome-based models remains to be studied. TP53 is the most frequently mutated and most studied gene in human cancers and is a natural candidate for exploring gene-specific approaches to cancer driver prediction.

[0020] This specification describes a comprehensive supervised deep learning (DL) method for obtaining driver event signatures from RNA-Seq samples using GCN101, a method called Signature Model in Mutant Protein-Associated Transcriptome (STAMP), an overview of which is shown in Figure 1. Figure 1 shows a workflow diagram illustrating the strategies used in the development of STAMP. In step 1 of Figure 1, the data 103 used as input to the model is prepared for training. This data 103 includes gene connectivity graphs 105 such as curated biological databases (e.g., HumanNet v2 or Funcoup). Other databases (e.g., Cancer Genome Atlas = TCGA) may also be referenced. For comparison, in step 2 of Figure 1, features are selected to predict mutational states (107) and used in conjunction with machine learning methods (109) such as random forest (RF) and elastic net regression (ELR). In the STAMP method, a graph convolutional neural network 101 may be used instead to classify mutational states. In step 3 of Figure 3, these methods are combined to predict various metrics such as survival rate and single-nucleotide signature and to classify variants (111). In addition, the metrics or scores output by the model can be used for treatment prioritization (113). Figure 1 outlines the various methods used to validate the analysis and the potential utility that can be obtained from STAMP.

[0021] In one embodiment, machine learning methods used to predict mutational states may include, among others, random forests (RF), elastic network regression (ELR), convolutional neural networks (CNN), support vector machines (SVM), logistic regression (LR), naive Bayes (NB), and k-nearest neighbors (KNN).

[0022] Gene and tumor-specific approaches are applied, first trained and validated against TP53, and then extended to several other important oncogenes and pathways. STAMP predictions may correlate more significantly with the molecular and clinical traits of p53 compared to simple mutation analysis. Models trained for genes and tumor types where relevant targeted therapies exist have been shown to correlate strongly with drug responses in independent cell line databases. Therefore, STAMP successfully identifies signatures in transcriptome data for driver events, improving mutation analysis and potentially leading to the identification of novel, undetected driver events. This enables potential utility for therapeutic prioritization using gene expression information.

[0023] Framework for the STAMP Model Aspects of the present invention include training a machine learning model to predict TP53 mutation status based on RNA-Seq samples from the Cancer Genome Atlas (TCGA) 115 (Figure 1). TP53 mutation status labels were defined in a binary format: mutation (M, positive) for samples with any TP53 mutation, and wild-type (WT (Wild Type), negative) otherwise. A pan-cancer model (n=9,077 samples with both mutation and RNA-seq data) was trained in parallel with a tumor-specific breast cancer (BRCA, n=1,042) model, which is the tumor type with the largest number of samples in the TCGA. Each dataset was randomly split in a 60-20-20 format to form a training set, validation set, and test set, respectively. For each dataset (pan-cancer or BRCA), the splitting was performed before model creation to ensure the same allocation was applied to all models. The model was trained using the training set, the validation set was used for hyperparameter tuning (including identifying the optimal feature selection method), and the test set was used to evaluate the model by its ability to generalize to unknown data. Two conventional learning models, Random Forest (RF) and Elastic Network Regression (ELR), were used for analysis. Each conventional model was run using all features (all genes) and then using the feature selection method.

[0024] Feature Selection A gene expression sample contains nearly 20,000 genes as features, while samples with transcriptome data currently available in publicly accessible cancer databases number only a few hundred to a few thousand. Training on such a sample size (e.g., all 20,000 genes) often leads to overfitting. To address this challenge, a feature selection process can be applied to allow for the reduction of redundant, noisy, or irrelevant features. The reduced feature space also allows for improved computational efficiency and thus speeds up the learning process. Feature selection processes can generally be divided into two approaches: (i) data-driven and (ii) knowledge-based. Data-driven feature selection uses the investigated dataset to identify specific features (genes) that correlate with the studied trait. In contrast, the knowledge-based approach refers to the use of pre-established knowledge to select features that are more relevant to the question at hand.

[0025] For the purposes of this invention, two feature selection methods were used: a data-driven approach using differential expression (DE) and a prior knowledge approach using a literature-curated set of pathway genes. DE was used in RNA-Seq training cohorts to examine the correlation between each gene and TP53 mutation status. 150 genes with the most positive correlation and 150 genes with the most negative correlation were selected (n=300 genes). For the second approach, a set of TP53 pathway-related genes was extracted from the Molecular Signatures Database (MSigDB, n=311 genes).

[0026] The trained conventional models performed very well on the validation set. ELR with MSigDB feature selection performed best (AUC=93.1% for BRCA and AUC=94.1% for pan-cancer). ELR with DE feature selection performed slightly worse (AUC=92.7% for BRCA and AUC=93.2% for pan-cancer). RF performed better with DE than RF with MSigDB (AUC=81.8% for BRCA and 84.1% for pan-cancer), but generally performed worse than ELR (see Table 1 for a complete explanation of AUC for the validation set). Therefore, ELR with MSigDB was selected for further analysis. RF with DE was also investigated further to maintain variability in the methods being compared.

[0027] [Table 1]

[0028] Because the positive label ratio (i.e., the percentage of samples with mutation TP53) varied between cohorts (7%–92.8%, see Table 2 below), several metrics were used in the final stage of comparison against the test set to enable a more comprehensive comparison between imbalanced datasets. Primarily, precision, AUC, and F1 were calculated and compared for different models and cohorts. The models were well generalized on the test set. From conventionally trained models, ELR using MSigDB performed best in both cohorts across all metrics (AUC=93.99%, ACC=88.82%, F1=84.2% for pan-cancer; AUC=97.2%, ACC=93.3%, F1=89.71% for BRCA; see Figures 2a–2c for a complete comparison).

[0029] Figure 2 shows the performance of the STAMP model for different cohorts.

[0030] Figures 2A and 2B -- Test set ROC curves for the following: GCN (green -201), ELR with MSigDB feature selection (red -203), and RF with DE feature selection (purple -205) for pan-cancer (a) and breast cancer (b) cohorts when learning TP53 mutation status. GCN is best for pan-cancer (AUC=94.95%), and ELR is best for breast cancer (AUC=97.2%).

[0031] Figure 2C -- Precision, AUC, and F1 metric for three models (green GCN-207, red ELR-209, and purple RF-211). Comparisons using several metrics are important in cohorts with taxonomic imbalance.

[0032] Figure 2D - Validation of STAMP models on independent METABRIC cohorts of breast cancer tumors. GCN (213) performed best (AUC=90.22%), RF (215) also performed well (AUC=89.28%), and ELR (217) performed significantly worse (AUC=84.12%). The pan-cancer GCN model (219) performed worse than the breast cancer GCN model (AUC=83.99% vs. 90.22%, respectively).

[0033] Figure 2E -- Downsampling analysis of breast cancer illustrates the boundaries for applying STAMP. Models were trained for various combinations of cohort size (x-axis) and the number of samples with positive labels (y-axis). F1 scores are shown (blue - 221 - low, yellow - 223 - medium, red - 225 - high).

[0034] Figure 2F -- Extensions of STAMP models for all eligible tumor types (green GCN

[0227] , red ELR

[0229] , and purple RF

[0231] ) (x-axis, ordered by the logarithm of cohort size as shown by the green bar plot). Most models for tumor types that passed the downsampling criteria achieved a high AUC for the test set (y-axis). Abbreviations: GCN - Graph Convolutional Network, ELR - Elastic Net Regression, RF - Random Forest, MSigDB - Molecular Signatures Database, DE - Differential Expression, AUC - Area Under the (ROC) Curve.

[0035] [Table 2]

[0036] Graph convolutional neural networks integrate prior biological knowledge into the model architecture. Graph Convolutional Networks (GCNs) are graph-based deep learning architectures that enable the integration of both biologically curated and data-driven feature extraction processes into the same model. Genes are included in the model as graph nodes, and different types of gene interactions form graph edges (interactions include protein-protein interactions, co-expression, common pathways, etc.; see the Methods section below in this specification). The edges are then used as constraints on the weight multiplication of the neural network in each layer, allowing for feature modification in non-initial layers. This is both data-driven (modified across the layers of the neural network) and creates features that depend on a biologically curated graph. GCNs also reduce variability in the feature selection process compared to conventional learning models. That is, one graph structure can generate relevant subgraphs for any query gene or tumor type, whereas DE analysis requires re-establishment for each new query, and MSigDB requires finding a new set when dealing with new genes. The graph used by the GCN architecture can vary. For the analysis below, two graph structures were used and compared. HumannetV2 and Funcoup. GCN was applied using primary neighbor genes of TP53 (n=344 for HumannetV2, n=1,425 for Funcoup), and grid search was used for hyperparameter tuning (see Methods). Most analyses were performed with a relatively small number of layers (1-4), but similar experiments (including grid search for hyperparameter tuning) were performed with a 20-layer GCN architecture.

[0037] The best validation results were achieved using GCN, which outperformed conventional learning methods (AUC=94.3% for BRCA, second best being ELR-MSigDB at 93.1%, and for pan-cancer, AUC=94.8%, second best being ELR-MSigDB at 94.1%, Table 1). In some experiments, the GCN with 20 layers failed to converge to normal learning and was therefore omitted from further analysis (AUC=50% for both cohorts, Table 1). In the trial set, GCN generalized well for both cohorts. Compared to conventional models, it performed slightly better in the pan-cancer trial as defined by AUC, but not by other performance measures (AUC=94.95%, second best being ELR-MSigDB at 93.99%, Figures 2A and 2B. Comparison by other metrics is shown in Figure 2C). In the BRCA cohort, performance was worse than conventional models (AUC 91.91%, compared to a better ELR of 97.2% for MSigDB, Figures 2A-2C). The better performance of GCN for pan-cancer may suggest the advantages that deep neural networks are expected to offer as more data is accumulated.

[0038] STAMP model accurately predicts mutation status in an independent cohort of breast tumors. The results from both the validation and test sets confirmed the normal learning of TP53 mutation patterns in gene expression across various TCGA tumor type cohorts. However, this analysis did not demonstrate generalizability to other independent cancer databases, which may be more challenging. To investigate whether the STAMP model could be generalized to samples outside of TCGA, we used the METABRIC database, consisting of 2,000 breast tumors. This study not only validates STAMP against independent data but also tests the generalizability of a model trained on RNA-Seq data in predicting microarray gene expression samples. TCGA also applies unique normalization procedures not used with METABRIC, which further complicates the attempted generalization.

[0039] The model predicted a very high AUC score for the METABRIC cohort, while GCN had a remarkable AUC score of 90.22% (Figure 2D). RF using DE performed next best with an AUC of 89.28%. Interestingly, the ELR model using MsigDB, which performed best on the TCGA breast cancer trial set, performed significantly worse on METABRIC (AUC of 84.12%). This may indicate that the ELR model is overfitted to the TCGA data, as it fails to reproduce a higher AUC score on this independent dataset compared to GCN or RF.

[0040] The METABRIC cohort allowed us to investigate another unresolved issue regarding pan-oncological models compared to tumor type-specific models. GCN models were trained against TCGA pan-oncological models to predict METABRIC. The pan-oncological models performed significantly worse than the tumor type-specific models (AUC = 83.99% vs. 90.22%, respectively). This good performance in METABRIC likely reflects a more accurate measure of the model's potential to predict for unknown data (compared to the TCGA exclusion trial set). Therefore, based on these experiments, GCN is the best-performing architecture for predicting mutational status from gene expression samples, and tumor type-specific models perform better than the pan-oncological approach.

[0041] Downsampling explores tumor type and gene criteria related to learning. To apply this framework to further cancer-specific and gene-specific cohorts, it is necessary to explore the boundaries for normal learning in the case of smaller sample sizes or lower positive / negative label ratios. Therefore, a downsampling analysis was performed on the TP53 breast cancer cohort. Breast cancer was chosen due to its high sample size (n=1,042), the relative abundance of TP53 mutations (32.6% of samples were mutated), and the fact that the model performed well for this cohort. After excluding the validation set (20% of the samples), both the number of samples and the positive label ratio in the training set were gradually reduced using random sampling. GCN and ELR were trained for each pair of downgraded values ​​to determine the sample size and positive label ratio required for normal learning. F1 scores were used for comparison with the test set. For the GCN model, learning did not lose power when the percentage of mutated samples was ≥10% and the number of samples was ≥300. Favorable results were also achieved using only 100 samples when the mutation rate was maintained at 30% (Figure 2E). Similar boundaries were examined for ELR, but this appears to be better for tumor types with fewer samples, as shown in Figure 8. Six tumor types were excluded from further analysis under the criteria suggested by these results (see Methods and Table 2).

[0042] Figure 8 shows the downsampling analysis of ELR (MSigDB). F1 scores (shaded from 0

[0801] to 0.8

[0803] ) are given for models trained under various sample size and positive label ratio scenarios. ELR appears to perform better than GCN for smaller sample sizes and lower positive label ratios (compared to Figure 2D).

[0043] Expansion methodology for all eligible tumor types We selected TCGA tumor type and gene cohorts with sufficient power for training, as defined by the criteria described above. We trained the model and adjusted it for all relevant tumor type cohorts in TCGA (n=17). A comparison of the test set AUC of this extended analysis is given in Figure 2F. The results show that learning the genetic trait of TP53 mutation status from RNA-Seq samples is applicable across a diverse set of tumor types. The RF model generally performed worse than the other models across tumor types.

[0044] Verification of the correlation between STAMP and p53 functionality via genetic and clinical traits. Figure 9 conceptually illustrates how STAMP is trained.

[0045] Figure 9A -- To train STAMP, TCGA tumor samples are first split according to their TP53 mutation status. The TP53 mutation status is reported as mutant (red - 901) or WT (blue - 903).

[0046] Figure 9B – STAMP forms binary predictions for all TCGA samples. These predictions are identified as harmful (D-905) or non-harmful (ND-907). In addition, a linear score is calculated to quantify the impact of the severity of dysfunction.

[0047] Figure 9C -- Samples that have a mutation in TP53 (i.e., red in Figure 9A

[0901] ) but are predicted to be non-toxic by STAMP (i.e., blue in Figure 9B

[0907] ) can be considered "false WT" and are colored yellow (909) here. Their STAMP predictions imply an RNA signature similar to that of samples with WT TP53. Similarly, samples that have WT TP53 (blue in Figure 9A

[0903] ) and are predicted to be D by STAMP (red in Figure 9B

[0905] ) can be considered "false variants" and are colored blue-green (911) here.

[0048] As described, samples in TCGA were divided by binary labeling based on their TP53 mutation status. A STAMP classifier was then trained to predict mutation status based on transcriptome profiles (Figure 9A). Predictions were given for each sample, showing its expected p53 functional effect as either harmful (D, mutant-like RNA pattern -905) or non-harmful (ND, Wt-like RNA pattern -907), as shown in Figure 9b. The transcriptome profiles of specific samples contradict their actual TP53 status at the DNA level, causing the model's prediction to differ from the original mutation status. While these discrepancies may be due to erroneous predictions, they can alternatively indicate the true biological behavior in specific samples, contrary to their mutation status ("false variant"

[0911] or "false Wt"

[0909] samples, Figure 9c). Therefore, the analyses performed in the following sections investigate whether the STAMP classification (D / ND) correlates better with genetic and clinical traits that exhibit different behavior in Wt TP53-derived variants compared to mutation status. Such correlations may suggest that RNA-based models can identify driver events in proteins or pathways that are not detected by gene mutation analysis.

[0049] STAMP predictions form a quantitative score that stratifies patient survival based on the severity of TP53 dysfunction. Most clinical annotations treat the effect of mutations on protein function as either abnormal or functional. However, quantifying the level of dysfunction caused by each mutation can enable the ranking of mutations in tumors and, therefore, the prioritization of gene targeting. To address this objective, the model's predictive scores were modified to score the driver event effect (Figures 9A–9C). This score can be calculated by inverting the sigmoid activation function involved in binary classification. The score is validated in the analysis presented below.

[0050] To validate STAMP's predictions and generated scores, survival data was retrieved from TCGA, and the results are reported in Figures 3A-3E. Patients with mutational TP53 in the pan-cancer cohort showed significantly shorter overall survival (OS) compared to patients with wt TP53 (Figure 3a, p=1.2e-41, log-rank test). Since STAMP forms a score for each tumor based on its TP53 status, we tested whether such scores could rank tumors in terms of their survival prognosis.

[0051] Figures 3A to 3E show the validation of the STAMP model.

[0052] Figure 3A - Survival analysis of the TCGA pan-cancer cohort segmented by TP53 mutation status (red M

[0301] , blue Wt

[0303] , log test p=1.2e-41. See Figure 11a for similar segmentation based on STAMP prediction).

[0053] Figure 3B - TCGA pan-cancer cohort divided into quartiles based on STAMP scores. STAMP's TP53 model predictions stratify tumors for their survival outcomes.

[0054] Figure 3C - Box plots showing the association of SBS4 in LUAD to the actual mutational state of TP53 (left, red M

[0305] and blue Wt

[0307] ) compared with STAMP's TP53 model predictions (right, red D prediction

[0309] and blue ND

[0311] ). STAMP enhances the enrichment of SBS4 (p=3e-05 vs 0.00033).

[0055] Figure 3D - The correlation between STAMP scores and the prevalence of SBS4-derived mutations in LUAD is also statistically significant (Spearman p = 2.16e-07).

[0056] Figure 3E -- Comparison of quantitative STAMP scores between missense mutations predicted to be deleterious (D) (313) and missense mutations predicted to be non-deleterious (ND) (315), as defined by the previously published algorithm TP53_PROF (Wilcoxon p=1.2e-11).

[0057] Abbreviations in Figures 3A-3E: M - mutant, Wt - wild type, Q1-Q4 - quartile 1-4, D - harmful, ND - non-harmful, and LUAD - lung adenocarcinoma.

[0058] STAMP scores for all TCGA samples were divided into quartiles to form four patient groups with distinct predictive values. Statistical differences were then compared among these groups using a log test for Cox regression models. Notably, survival outcomes differed statistically among all four patient groups (p-values ​​given for adjacent quartiles: Q1-Q2, p=2e-06; Q2-Q3, p=2e-15; Q3-Q4, p=0.006, Figure 3B). Multivariate analysis was performed, showing statistically significant differences when controlling for tumor type, age, and mutational burden (Figure 10). The effect of STAMP on survival was also tested in several tumor types where segmentation by TP53 mutation status was statistically significant in survival data. The results of this analysis are presented in detail in Figures 9A-9C-15A-15D.

[0059] Single-base signature (SBS) analysis validates the superiority of gene expression models over mutational states when predicting p53 function. Several mutation processes can generate characteristic signatures of base pair substitutions across DNA sequences. For example, the single-nucleotide signature (SBS)4 is associated with smoking. It is abundant in lung adenocarcinoma (LUAD) and has been shown to correlate with the occurrence of TP53 mutations. Smoking-associated SBS4, and several other signatures, are used here as a novel and strongly supported validation of the TP53 STAMP model. First, the Wilcoxon rank-sum test was used to compare mutant and wt TP53 (model label) for their SBS4 enrichment in LUAD, showing a statistically significant enrichment for mutant TP53 samples (p=0.00033, Figure 3C, left). Next, SBS enrichment was examined based on STAMP predictions. If the model improves the identification of p53 activity compared to mutation analysis, SBS4 enrichment should also be enhanced. Indeed, SBS4 enrichment in LUAD was higher when samples were separated by STAMP prediction compared to separation by mutational status (p=3e-05 vs. p=0.00033, Figure 3Cc). Since sample size remained constant in both analyses, lower p values ​​indicate a stronger effect size. SBS4 enrichment also correlated strongly with STAMP scores (p=2.16e-07, Spearman correlation test; Figure 3D). Correlations of STAMP with APOBEC and HR-DDR signatures were also investigated in pan-cancer or related tumor types, yielding similar results. These analyses are presented in detail in Figures 10 and 11A–11E and elsewhere in this specification. The correlation with SBS enrichment highlights the potential of STAMP RNA-based models in identifying protein function patterns such as those captured by relevant biological processes.

[0060] STAMP identifies patterns of harmful and non-harmful missense variants in TP53. In one embodiment, the binary-labeled STAMP can be trained on TP53 Wt samples isolated from mutants, regardless of the type or effect of the mutation. Therefore, examining the correlation of STAMP with different mutant types and mutation effects can serve as another validation of the model. Thus, we then investigated whether STAMP could identify those with driver mutations from those carrying passenger variants (pseudoWts) that do not affect p53 activity among TP53 mutant tumors. Missense mutations were the focus of this analysis because they are frequent in TP53 and their effects are considered more difficult to identify. TP53_PROF is a recently published gene-specific machine learning algorithm that classifies all TP53 missense variants with very high accuracy (96.5%) for their effects on protein function. In this example, we validated STAMP by comparing its score with the variant classification of TP53_PROF using TP53_PROF. In this analysis, harmful (D) and non-harmful (ND) annotations refer to the variant classification of TP53_PROF for missense mutations.

[0061] A pan-cancer model was examined to include a larger sample size and thus allow for improved statistical power. Samples with missense mutations predicted as ND were given significantly lower STAMP scores compared to missense mutations predicted as D (Figure 3E, p=1.2e-11, Wilcoxon rank-sum test). Eight additional tumor type cohorts with sufficient ND-labeled samples were also examined, and all but one showed significantly lower scores compared to ND samples (Figures 14A-14K). Samples with truncated mutations were also examined and showed similar values ​​to the D missense mutation samples (Figures 14A-14K). Overall, STAMP scores successfully identify patterns in the effect of TP53 missense mutations on protein function.

[0062] This will form a gene expression-based atlas of driver event signatures across important oncogenes and pathways.

[0063] Since many available diagnostic biomarkers and therapies are based on targeting specific oncogenes, extending STAMP to apply to genes other than TP53 is of paramount importance. Many more biomarkers and therapies are in clinical trials and are expected to be available in the future. Identifying the function of key oncological pathways (i.e., whether the entire pathway functions normally), in addition to other genes, can be of great importance for describing the molecular state of tumors and for therapeutic considerations. Therefore, we extended the STAMP pipeline to predict dysfunction of key oncogenes other than TP53 and predict the function of the entire oncological pathway. To implement such an extension, instead of using mutational states as with TP53, we used gene and pathway change maps from the study by Sanchez-Vega et al. as sample labels. The authors utilized prior knowledge from pathway databases, driver databases, mutation and epigenetic analyses, as well as other sources from the scientific literature. The authors integrated these to form comprehensive change maps across all TCGA samples and defined binary function annotations for many key oncogenes and entire pathways. These change maps can be used as highly accurate and rigorously validated labels for STAMP. Such labeling would allow for the training of models to predict dysfunction at the pathway level, in parallel with the training of gene-specific models.

[0064] Figures 4A to 4I show the creation and verification of an atlas of abnormalities across genes, pathways, and tumor types.

[0065] Figure 4A provides a heatmap showing the gene-tumor type crossovers for which the STAMP model was trained. The gray slot (401) represents the gene-tumor type cohort that did not meet the downsampling criteria, as shown in Figure 2E. The slots are colored with gradients of red (high) (403), yellow (medium) (405), and blue (low) (407) according to the AUC of the test set achieved by each model.

[0066] Figures 4B to 4E show survival analyses for LATS2 (b to c) in LGG and BRAF (d to e) in SKCM, similar to those in Figures 3A to 3E.

[0067] Figures 4F and 4H show SBS analyses for MSI-related SBS44 in COAD and SBS20 in STAD, similar to those in Figures 3A to 3E.

[0068] Figures 4G and 4I show that the scores correlate with SBS enrichment for these two models.

[0069] Abbreviations in Figures 4A-4I: SBS - single nucleotide signature, LGG - low-grade glioma, SKCM - cutaneous melanoma, MSI - microsatellite instability, COAD - colon adenocarcinoma, STAD - gastric adenocarcinoma.

[0070] Using the downsampling criteria defined in the preceding sections of this disclosure, we determined the relevant gene-tumor type combinations with sufficient power for training, resulting in 296 trained models across 25 tumor types, 10 pathways, and 44 genes (Figures 4A and 15A). Specifically, the models were trained and adjusted for cohorts containing more than 300 samples with at least 10% positive labeling, or 100 samples with at least 30% positive labeling. Both ELR and GCN models were trained and adjusted. AUC scores for different gene and pathway models are presented in Figures 4A and 15A, respectively (gray slots represent gene-tumor type combinations that did not meet the downsampling criteria).

[0071] The newly trained model enabled further examination of survival and SBS analysis. One example is the effect of the LATS2 gene on survival in patients with low-grade glioma (LGG). LATS2, a core component of the Hippo signaling pathway, is dysregulated via epigenetic silencing in several cancer types. In LGG, LATS2 silencing by promoter hypermethylation has been shown to significantly affect prognosis. Therefore, we investigated the effect of the STAMP LATS2 model on survival outcomes. The survival effect was statistically significant based on LATS2 abnormalities in LGG (labels used for training, log test p=2.5e-16, Figure 4B), and the effect was enhanced when split by STAMP prediction (p=9.3e-19, Figure 15B). Similar to the survival analysis of TP53, STAMP scores for the LATS2 gene in LGG samples were split into quartiles and then compared for differences in survival outcomes. Statistically different outcomes were observed among the three groups (Q1–Q2, p=7.12e-07; Q2–Q3 had a boundary p-value of 0.071; Q2–Q4 were different, p=0.0063; Q3 and Q4 were not different, p=0.68; Figures 4C and 15C). Multivariate analysis was performed, and statistically significant differences were observed when controlling for age and mutational burden features (Figure 15D). Another statistically significant advantage in survival was observed in patients with TCGA skin cancer (SKCM) with BRAF mutations (labeled models, log test p=0.037, Figure 4D), likely attributable to improved treatment opportunities. Patients predicted by STAMP as BRAF dysfunction (i.e., ND = non-adverse) showed a statistically significant enhanced survival advantage compared to labeled patients (p=0.0007 vs. 0.037, Figure 4E).

[0072] SBS analysis was also performed on non-TP53 models. SBS44, a high-frequency signature in colorectal cancer (COAD), is associated with defective DNA mismatch repair and high microsatellite instability (MSI). High MSI is present in 10–15% of COAD cases and is known to affect the prognosis of these patients. The APC gene, a key tumor suppressor in the majority of COAD cases, is significantly more frequently mutated in patients with low MSI. Therefore, we investigated SBS44 enrichment for APC models in COAD. Indeed, COAD patients labeled with functional APC were enriched for SBS44 (Wilcoxon p=3.2e-12, Figure 4F). Similar to TP53 and its associated SBS, the APC model improved the correlation with SBS44 compared to the original label (p<2.2e-16 vs. 3.2e-12, Figure 4F). STAMP scores also strongly correlated with SBS44 enrichment (Spearman correlation, p=4.5e-16, Figure 4G). SBS20 is another signature associated with high MSI and is frequent in stomach adenocarcinoma (STAD). In STAD, high MSI appears in a subset of patients and affects prognosis and response to treatment. NOTCH pathway activation correlates with low MSI in gastric cancer. Indeed, NOTCH pathway abnormalities, labeled by Sanchez-Vega et al., were significantly enriched for SBS20 in gastric adenocarcinoma (p=6.3e-08, Figure 4H). The NOTCH pathway model in STAD improved the correlation with SBS20 compared to the original label (p<2.2e-16 vs. 6.3e-08, Figure 4H). The score, again, was strongly correlated with SBS20 enrichment (Spearman correlation, p=1.3e-15, Figure 4I).

[0073] Testing the correlation between STAMP and drug response in a GDSC cell line database. In the embodiments described herein, STAMP has been successfully applied to many genes and pathways in a tumor-type-specific context. The ability of STAMP to quantify the effects of gene-targeted therapies in cancer has also been investigated. Drug information gathering in cell lines from the Genomics of Drug Sensitivity in Cancer (GDSC) database has also been conducted. 50 We tested the correlation between this and RNA model predictions.

[0074] GDSC is the largest publicly available resource for drug sensitivity in cancer cell lines and was therefore used in this analysis. The first drug tested was nutrin, an MDM2 antagonist that has a strong effect by activating the p53 pathway when TP53 is Wt, and TP53-mutated tumors are resistant to it. Therefore, nutrin can be used here as a strong validation of the predictions of the TP53 model. To improve statistical power, nutrin was tested in a pan-cancer cohort. Five further tumor type, gene, and drug combinations were identified using relevant data from GDSC, including trained STAMP models (i.e., passing the downsampling criteria; all selected models also had a test set AUC > 80%) and potential clinical relevance. These were as follows: Afatinib, which targets ERBB2 in breast cancer; PLX-4720 and dabrafenib, both targeting the BRAF V600E mutation in skin cancer; a KRAS G12C inhibitor, which targets the KRAS G12C mutation in lung adenocarcinoma; and lapatinib, which targets EGFR in glioblastoma. Since lapatinib administration to glioblastoma patients with EGFR mutations has been shown not to improve clinical outcomes, the latter is used as a negative control. Drug informed consent. 50 The values ​​were compared to STAMP scores using Pearson's correlation coefficient test.

[0075] Figures 5A to 5H show IC scores of GDSC cell lines with STAMP scores. 50 This shows the correlation between the values.

[0076] Figure 5A - TP53, the pan - cancer model score (y - axis) correlates with the IC of nutlin in pan - cancer cell lines (x - axis. Pearson correlation coefficient R = 0.5, p < 2.2e - 16). The ellipses are for comparing the region of mutant samples (red ellipse 501, diameter ranging from 2.5 to 8 of IC 50 and scores ranging from - 7 to 6) and the region of Wt samples (blue ellipse 503, diameter ranging from - 1 to 5 of IC 50 and scores ranging from - 15 to - 2). 50 are shown for comparison.

[0077] Figure 5B - ERBB2, the breast cancer model score correlates with the IC of afatinib (R = - 0.6, p = 8.1e - 06). The points are colored by ERBB2 status: red - amplification (505), cyan - loss of function (507), green - gain of function (509), purple - neutral (511). The red ellipse (513) shows cell lines with amplified status examined separately at f. 50 are shown for comparison.

[0078] Figure 5C - Figure 5D - BRAF, the SKCM model score correlates with the IC of cPLX - 4720 50 (R = - 0.63, p = 5.7e - 05), and the IC of dabrafenib 50 (R = - 0.54, p = 0.0021). The samples are colored by BRAF status: Wt - blue (515), mutant - red (517).

[0079] Figure 5E - Similar to Figure 5A, but only Wt TP53 cell lines are examined. The ellipse drawn in a is shown again for comparison. A subset of samples with Wt TP53 shows resistance to nutlin, similar to most mutant samples (as emphasized by the red ellipse 502). The STAMP score statistically significantly correlates with the IC of nutlin when testing samples having only the Wt TP53 state (R = 0.38, p = 2.7e - 10).<00故、 50 is shown for comparison.

[0080] Figure 5F is similar to Figure 5B, but only cell lines with ERBB2 amplification are shown. When examining samples that have only amplified ERBB2 status despite very small sample sizes, the score is the IC of afatinib. 50 This shows a statistically significant correlation (R=-0.65, p=0.011).

[0081] Figure 5G -- EGFR, GBM model score, and lapatinib IC 50 Correlation is tested as a negative control. Samples are colored according to EGFR status: red - mutant (519), blue - Wt (521). As expected, no statistically significant correlation was observed (R=-0.088, p=0.68).

[0082] Figure 5H -- IC of mutant (red) (523) and Wt (blue) (525) EGFR samples. 50 The levels are shown in box plots, and no significant difference was observed (Wilcoxon rank-sum test, p=0.16), confirming this as a negative control.

[0083] Abbreviations in Figures 5A-5H: GDSC - Genomics of Drug Sensitivity in Cancer, Mut - Mutation, Wt - Wild-type, Amp - Amplification, Loss - Loss of function, Gain - Gain of function, BRCA - Breast cancer, SKCM - Cutaneous melanoma, GBM - Glioblastoma.

[0084] Regarding TP53 in the pan-cancer cohort, nutrin IC 50 The scores have a correlation coefficient of R=0.5 (p.val<2.2e-16, Figure 5A), and therefore represent resistance of tumors with abnormal TP53 activity to nutrin. Cell lines are colored for TP53 mutation status, demonstrating the success of the model in predicting abnormal status and nutrin IC. 50 It is emphasized that there is a strong correlation with ERBB2 in breast cancer. 50The correlation coefficient between the score and the drug was R=-0.6 (p=8.1e-06, Figure 5B), representing the sensitivity of cell lines with increased ERBB2 activity to the drug. Cell lines are colored for ERBB2 status (amplification, loss of function, gain of function, and neutral). For BRAF in skin cancer, the correlation coefficient for PLX-4720 was R=-0.63 (p=5.7e-05, Figure 5C), and for dabrafenib it was R=-0.54 (p=0.0021, Figure 5D). BRAF cell lines are colored for the presence of the V600E mutation (coloring of mutants compared to Wt is shown in Figures 16A and 16B). The correlation of STAMP prediction was also examined for samples with the same genetic profile. Notably, even when examining only Wt TP53 cell lines expected to be sensitive to nutrin, the score was IC 50 This showed a strong correlation (Figure 5E, R=0.38, p=2.7e-10). Notably, samples with TP53 Wt that had a similar score to the mutant TP53 (pseudo-mutant samples) also showed a higher IC53. 50 The values ​​were present (as indicated by the red circle (502) in Figure 5E). For samples containing only TP53 variants where nutrin was not expected to have an effect, the correlation was weak but still statistically significant (R=0.095, p=0.032, Figure 16C). The false Wt TP53 showed similar IC values ​​to those of the true Wt samples. 50 The values ​​were obtained (as shown by the blue circle (1601) in Figure 16C). A similar pattern was observed for breast cancer cell lines with amplified ERBB2 (samples inside the red circle in Figure 5B), and their scores were in line with their afatinib IC50. 50We were able to stratify the values ​​(R=-0.65, p=0.011, Figure 5F). When we examined only a subset of skin cancer cell lines with the V600E mutation in BRAF, the results showed a similar trend (R=-0.35, p=0.08 for PLX-4720, R=-0.32, p=0.14 for dabrafenib, Figures 16D and 16E). The statistically significant value of the boundary is likely due to the very small sample size. These analyses demonstrate the important function of STAMP in identifying drug sensitivity even among samples with identical genetic profiles.

[0085] Regarding negative controls, lapatinib IC in glioblastoma 50 Comparing this to the EGFR model score was not statistically significant (R=-0.088, p=0.68, Figure 5G). This is because the mutational IC compared with Wt EGFR in glioblastoma samples. 50 This is consistent with the finding that there was no clear difference in values ​​(Figure 5H, Wilcoxon rank-sum test, p=0.16). The correlation between KRAS G12C inhibitors in LUAD and STAMP scores was also not significant (R=-0.19, p=0.2, Figure 16F, samples colored for the presence of the G12C mutation). This suggests that KRAS-G12C mutation samples and samples without this mutation showed similar relatively high IC scores. 50 This is consistent with having a value (Figure 16G, Wilc. p=0.1). The effect of KRAS G12C inhibitors is similarly not significant for samples based on mutational status in KRAS (any mutation) (Figures 16H and 16I, Wilc. p=0.31). Therefore, these analyses serve as a second negative control.

[0086] Notably, MDM2 expression levels may explain the cell line nutrin response stratification by the TP53 model shown in Figure 5E. If true, this would suggest that STAMP's contribution to the significance of these results is minimal. To investigate this, cell lines from the nutrin 3 cohort with TP53 WT state (shown in Figure 5E) and TP53 mutant state (shown in Supplementary Figure 16C) were colored for their MDM2 amplification state. As shown in Figures 18A and 18B, coloring cell lines for their MDM2 amplification state did not explain the significant stratification of responses to nutrin 3 achieved by STAMP.

[0087] Examples and further details This disclosure describes STAMP, a machine learning algorithm for predicting and quantifying driver events in a gene-specific and cancer-type-specific manner based on gene expression information. Conventional deep learning models were explored along with different strategies for reducing the dimensionality of RNA-Seq data. The models were trained to predict TP53 mutation status and then extended to multiple genes and pathways based on driver event annotation.

[0088] STAMP achieved high accuracy for an unknown TCGA trial set and independent cohort-METABRIC, demonstrating its usefulness in predicting variants for new patients. The fact that the METABRIC samples used for prediction originate from a microarray platform, while the model is trained on TCGA samples derived from an RNA-Seq platform, further validates and establishes the model's robustness. STAMP was shown to have improved predictions compared to simple variant analysis for the TP53 case by achieving stronger correlations with relevant SBS and survival data.

[0089] STAMP compared to the ENLIGHT study STAMP represents an improvement over state-of-the-art tools for predicting treatment response in clinical cohorts and suggests potential biomarkers for cancer treatment.

[0090] A major collection of clinical cohorts containing both mRNA and treatment response data has recently become available through the ENLIGHT study. ENLIGHT is a transcriptome-based tool for predicting drug response in cancer patients, and the dataset used to tune and validate ENLIGHT was used here to examine STAMPs for stratification of drug response in cancer patients. Cohorts targeting genes or pathways with STAMP models associated with the tested drugs were analyzed and compared with ENLIGHT scores.

[0091] Ten cohorts, including 508 cancer patients treated with five different drugs in six different settings, were identified as relevant to evaluate STAMP (Tables 3 and 4). Beyond cohorts with well-established drug targets, PTEN models were also evaluated in a cohort of breast cancer patients treated with anti-PD1. PTEN knockdown in breast cancer cell lines was shown to induce significantly higher PD-L1 expression. PTEN alterations were also shown to be associated with significantly lower overall response rates in patients with metastatic triple-negative breast cancer treated with anti-PD1. Therefore, PTEN models were investigated as potential biomarkers of the anti-PD1 response in breast cancer patients.

[0092] [Table 3]

[0093] [Table 4]

[0094] For five of the six settings, STAMP achieved significant stratification of treatment response (hepatocellular carcinoma patients treated with sorafenib, predicted by the RTK-RAS pathway model, Wilcoxon p=4.5e-07; breast cancer patients treated with sorafenib, predicted by the RTK-RAS pathway model, Wilcoxon p=0.012; five combined cohorts of breast cancer patients treated with trastuzumab, predicted by the ERBB2 gene model). Furthermore, for head and neck cancer patients treated with cetuximab, predicted by the EGFR gene model; Wilcoxon p=0.028; breast cancer patients treated with durvalumab, predicted by the PTEN gene model; Wilcoxon p=2.04e-04; and for breast cancer patients treated with lapatinib, the results predicted by the RTK-RAS pathway model were not statistically significant (Wilcoxon p=0.2). (Figures 6A-6F). All cohorts except the two treated with sorafenib received additional treatments other than the targeted therapies mentioned (Table 3), thus highlighting the robustness of STAMP.

[0095] Compared to ENLIGHT, STAMP improved the AUC of treatment response in three of the six settings (sorafenib in hepatocellular carcinoma, sorafenib in breast cancer, and anti-PD1 in breast cancer), with an AUC improvement of up to 22.76%. In the other settings, STAMP was only 6.15% inferior in AUC (Figures 6A-6F). While ENLIGHT's strengths are described for drugs with more precise targets, it has been noted that its predictive accuracy decreases for drugs with multiple targets, such as sorafenib. The main advantage STAMP shows for sorafenib is likely due to its ability to predict not only gene-level dysfunction but also pathway-level dysfunction, which explains sorafenib's multiple targets.

[0096] The I-SPY2 study is an adaptive clinical trial platform testing drugs for high-risk breast cancer patients in neoadjuvant therapy. One treatment group of I-SPY2 was also used in the validation of ENLIGHT and has already been presented in the above analysis (using the anti-PD1 agent durvalumab in breast cancer patients) (Figure 6F). However, other treatment groups of I-SPY2 tested different treatments in different patients. These included patients treated with chemotherapy in combination with HER2-related therapy, as well as patients treated with pembrolizumab (Tables 3 and 4). These patients were not evaluated by ENLIGHT and provided an opportunity to further expand the clinical analysis. We evaluated the ERBB2 gene model for breast cancer and significantly stratified the response for all patients treated with Her2-related therapy (Wilcoxon p=5.62e-03, Figure 7A. See Tables 3 and 4 for included treatments and cohort sizes, and for excluded samples). ENLIGHT scores were also generated for these patients, and STAMP showed an advantage in AUC (AUC = 59.78% for STAMP, AUC = 57% for ENLIGHT; Figure 7B). Importantly, when examining I-SPY2 patients with similar Her2 status, STAMP was still able to stratify treatment response (Patients below: Her2-positive, Wilcoxon p = 0.068; Her2-negative, Wilcoxon p = 0.038; Figures 7C and 7D).

[0097] I-SPY2 also included patients with Her2-negative status who were tested for their response to the anti-PD1 drug pembrolizumab. This allowed for a second assessment of the predictability of STAMP for the response to anti-PD1. In fact, the STAMP-PTEN gene model for breast cancer successfully stratified the response to pembrolizumab (Wilcoxon p=0.032, Figure 7E). STAMP also showed superiority over ENLIGHT (AUC=65.32% for STAMP vs. AUC=62.39% for ENLIGHT, Figure 7F).

[0098] Overall, STAMP was able to stratify treatment responses in several clinical cohorts by applying both gene-specific and pathway-specific models for relevant targeted therapies. STAMP scores suggest potentially important biomarkers that can be used in clinical settings. In some cohorts, STAMP improved predictions compared to ENLIGHT, the current state-of-the-art method.

[0099] Figures 6A to 6F. STAMP stratifies treatment response within a clinical cohort.

[0100] The box plots in Figures 6A to 6F show the stratification of patients by STAMP and the ROC curves comparing STAMP601 with ENLIGHT603 for the following clinical settings.

[0101] Figure 6A - Hepatocellular carcinoma patients treated with sorafenib, as predicted by the RTK RAS pathway model.

[0102] Figure 6B -- Breast cancer patients treated with sorafenib, as predicted by the RTK RAS pathway model.

[0103] Figure 6C -- Breast cancer patients treated with lapatinib, as predicted by the RTK RAS pathway model.

[0104] Figure 6D -- Breast cancer patients treated with trastuzumab, as predicted by the ERBB2 gene model.

[0105] Figure 6E - Head and neck cancer patients treated with cetuximab, as predicted by the EGFR gene model.

[0106] Figure 6F – Breast cancer patients treated with durvalumab, as predicted by the PTEN gene model. STAMP statistically significant stratified all settings except the lapatinib setting, suggesting that STAMP predictions can be used as a biomarker of treatment response. The ROC curves show that STAMP performed better than ENLIGHT in three of the settings, with an AUC improvement of up to 22.76%. Where STAMP did not perform better, the AUC difference was up to 6.15%.

[0107] Figures 7A to 7F. STAMP stratifies I-SPY2 patients based on their response to Her2-related therapies and their response to anti-PD1 in patients with the same Her2 condition.

[0108] Figures 7A and 7B -- STAMP stratified all patients treated with Her2-related therapy (Wilcoxon p = 5.621e-03) and showed a slight advantage over ENLIGHT (AUC = 59.78% for STAMP, AUC = 57% for ENLIGHT). See also Tables 3 and 4 for further information.

[0109] Figures 7C and 7D -- STAMP stratifies treatment response in patients with both positive and negative Her2 status (Wilcoxon p = 0.068 for Her2-positive patients, Wilcoxon p = 0.0378 for Her2-negative patients).

[0110] Figures 7E and 7F -- Patients in this treatment group with I-SPY2 treated with pembrolizumab allowed for further verification of the PTEN model predictability of the response to anti-PD1 (pembrolizumab, Wilcoxon p=0.0326), again showing an advantage over ENLIGHT (AUC=65.32% for STAMP, AUC=62.39% for ENLIGHT).

[0111] The learned mRNA signatures are designed to provide guidance for gene profiling that complements mutation analysis. STAMP can potentially identify "false" samples, i.e., mutant tumors in which the protein behaves as if it has Wt-like activity, or Wt tumors with abnormal protein activity patterns, as illustrated in Figures 9A–9C.

[0112] In clinical practice, this can help identify driver events based on gene expression profiles, even in samples where no targetable mutations are found. One of the novelties of this study is that, in addition to binary prediction, STAMP forms an effect quantification score. Such a score can guide the prioritization of driver events for treatment in samples where many targetable driver mutations occur. In fact, STAMP predictions have been used to inform the informed consent of many drugs in the GDSC cell line database. 50 It was shown to correlate with the level. Even for samples with the same genetic profile (e.g., all Wt TP53), the variability in drug IC shown for this score was observed. 50 The strong correlation with the values ​​demonstrates the potential for prioritizing patients in ways not feasible with mutation analysis alone. The quantified scores can stratify patients based on their expected drug response. Stratification of survival outcomes is also possible based on the STAMP TP53 pan-cancer model, which further demonstrates the validity of the scores in measuring the severity of the impact of driver events.

[0113] STAMP can be trained on human tumor RNA-Seq samples derived from TCGA. Examples of validation of drug response predictability from human samples and their corresponding responses to therapies are shown, for example, in Figures 6A–6F and 7A–7F of this disclosure. Thus, STAMP has been validated against cell line data from GDSC databases using microarrays for gene expression profiling. Successfully predicting RNA-Seq human-based models for microarray cell line samples provides further evidence of STAMP's robustness. While most studies in the literature examine cell line responses to therapies to infer predictions about human profiles, this disclosure describes the reverse approach. Although STAMP models were trained on human tumors, they were tested against cell line responses to therapies and show strong correlations across multiple drugs, target genes, and tumor types. Because the models were exposed only to human mRNA profiles, the predictability of human tumor responses to therapies may be even better than that presented in this study for cell lines.

[0114] GCN appears to be the best model due to its superiority over the pan-cancer TCGA trial and validation sets when making predictions using METABRIC. As a deep learning architecture, GCN is also expected to gain significant advantages in the future as publicly available data from cancer patients accumulates.

[0115] STAMP was trained on TCGA data normalized as z-scores based on healthy tissue gene expression profiles. To optimally predict new data, STAMP requires similar normalization, but this is rarely applicable. Here, we used a simplified approach, normalizing each new dataset using all available expression profiles in its set. This simplified approach worked remarkably well, particularly with datasets with large sample sizes such as METABRIC, I-SPY2, and GDSC. However, for very small datasets, such as some of the cohorts assembled in ENLIGHT, this approach fails to work properly. Therefore, normalization methodology is a current concern in STAMP's predictability and an important direction for future research. Ideally, STAMP should be applicable even to a single sample. Improving normalization techniques would also likely improve STAMP's currently presented results.

[0116] Each STAMP model has been trained on numerous tumors to obtain gene-specific predictions, but they can also be applied to specific patients to form gene profiles. Multiple pathway and gene-specific models can be applied in parallel to patient RNA-Seq samples to establish a rich explanation of the patient's condition. In a clinical framework, such profiling can aid in direct drug administration and prioritization, including combination therapies.

[0117] Method / Data Source and Availability Gene expression data RNA-Seq gene expression data from TCGA pan-cancer samples were downloaded. IlluminaHiSeq was used for RNA sequencing. The dataset was downloaded from HiSeqV2.gz-Academic Torrents, uploaded by the Xena team, UCSC. Gene expression values ​​are log2(x+1) converted from RSEM (RSEM: accurate quantification of gene and isoform expression from RNA-Seq data, publicly available).

[0118] The TCGAbiolinks package version 2.12.6 in R was also used to download raw gene expression counts, which were then used for differential expression analysis using the Limma package.

[0119] Download and integration of TP53 mutation status and clinical data. Pan-cancer data were downloaded from cBioPortal using the "Pan-cancer TCGA" quick select. Clinical and mutational data were extracted using TCGAbiolinks, allowing for more samples with mutational profiles. Samples with conflicting mutational states were removed from the analysis, resulting in 10,258 profiled samples in the final analysis, including 35.68% mutant samples (3,660). Survival data were also downloaded from cBioPortal where available. Profiled data were integrated with RNA-Seq data to generate the final dataset (9,077 samples).

[0120] Pathway and other genetic data Data used to label training data for genes and pathways other than TP53 was downloaded from Supplementary Table S4 by Sanchez-Vega et al. This table contains a genomic alteration matrix (amplification, deletion, point mutation, epigenetic silencing, and fusion) for key oncogenes in 9,125 TCGA samples. Alterations with low recurrence rates between tumor samples and alterations annotated as likely to be passengers by OncoKB were excluded. Maps of gene alteration states were then defined based on the different alterations recognized in that gene. Gene alteration states in each pathway were integrated to define pathway-level alteration states. For the analyses presented in this study, gene and pathway-level alteration states were used for labeling.

[0121] Training algorithm and feature selection Differential expression using the Limma package Differential expression analysis was performed using the Voom command from the Limma package. All genes in the RNA-Seq data of the training set samples were examined and ranked based on their correlation with TP53 mutation status (expression in the TP53 Wt group was compared to that of the TP53 mutant group). The 150 most positively correlated and 150 most negatively correlated genes were selected (highest / lowest Log2FC).

[0122] MutSigDB set for TP53 From the Molecular Signature Database, we selected the gene set named "FISCHER_DIRECT_P53_TARGETS_META_ANALYSIS". This set contains 311 genes that directly bind to p53 and are regulated by p53.

[0123] Random Forest Random Forest is an ensemble learning method used in this case for binary classification. This is done by using a bagging technique on the decision tree. Bagging means repeatedly selecting n random samples (which are hyperparameters) from the training set before fitting the decision tree to those samples. In addition to bagging, "RF" also randomly selects a subset of features. Thus, the decision tree is trained on a subset of samples and a subset of features from the original data. Each tree is trained as a binary classifier, and the algorithm selects a class according to a majority vote tree vote.

[0124] Elastic Net Regression For ElasticNet regression, I used "ElasticNet" from the scikit-learn package in Python. ElasticNet regression is a specific regularization technique applied in this case alongside logistic regression. It linearly combines the L1 and L2 penalties of the Lasso and Ridge methods, respectively. It can "reduce" features by generating zero-value coefficients, and thus reduce overfitting (the effect of the Lasso penalty). At the same time, it reduces the effect of features that are not very relevant to learning (the effect of the Ridge penalty).

[0125] Graph Convolutional Network Dutil et al. (Dutil F, Cohen JP, Weiss M, Derevyanko G, Bengio Y. Towards Gene Expression Convolutions using Gene Interaction Graphs. 2018, published online) provide a method for performing GCN on gene expression data. This specification employs the gene-specific approach used by Dutil et al. to predict the masked expression level of a gene from the transcription patterns of other genes. The convolution operation is used to extract node information according to its edges. The convolution approximates first-order neighbor nodes, as suggested by Deferrard et al., meaning that only nodes directly connected to the queried gene are used. (Defferrard M, Bresson X, Vandergheynst P. "Convolutional neural networks on graphs with fast localized spectral filtering". Adv Neural Inf Process Syst. 2016; (Nips): 3844-3852), and Kipf & Welling (Kipf TN, Welling M. "Semi-supervised classification with graph convolutional networks". 5th International Conference on Learning Representations, ICLR 2017 - Conference Track Proceedings. 2017.). Both the adjacency matrix of the graph and the gene expression table of the samples are processed before convolution. The identity matrix is ​​added to the adjacency matrix A, forming a self-loop for all genes. A is also normalized using the diagonal node-order matrix D to neutralize any discrepancies that may arise due to the different number of edges each node has. This normalization is performed using the spectral propagation rule proposed by Kipf & Welling (2016). The multiplication for a single convolutional layer is: A'X (I) θ is such that in the formula A' is the normalized adjacency matrix, X(I) θ is the input to the Ith layer, and θ is the weight of the Ith layer. Multiplication of these three matrices is performed at each layer, but the adjacency matrix remains the same for all layers, multiplied by X, which is an n (nodes) × c (features) matrix. Thus, the number of learned parameters per layer is simply c × o, where o is the number of output parameters in a given layer. Expression data is embedded so that each gene is represented by a vector of parameters learned during training. Each convolutional layer also includes the following steps: Skip connections are added to preserve both a given node and its neighbors, and the ReLU activation function is applied. Node aggregation is performed to reduce the number of genes after each layer based on hierarchical clustering. Graph connectivity maps are used for distributed computation of clusters, and max pooling is performed on the resulting clusters. When tested, dropout is used, and each node is excluded with a 40% probability after each convolutional layer. In the final layer, the remaining nodes are connected to a linear layer, followed by an activation function (e.g., sigmoid) that generates binary classification predictions.

[0126] Segmentation into trains, verification, and testing. 20% of the data is retained for the test set to train each tumor type and gene model. 60% of the data is used to train the models, and hyperparameter tuning is performed on the remaining 20% ​​which is used as a validation set. Only the best-performing models for each algorithm (RF, ELR, GCN) are examined on the test set.

[0127] Hyperparameter tuning Different models can be trained using grid search techniques. In the case of GCN, the following hyperparameters can be adjusted: graph structure, number of layers, number of weights in each layer, batch size, and learning rate. An example of comparing two hyperparameters in a grid is shown in Figure 17. For RF, the adjusted hyperparameters were n_estimator, criterion, and maximum feature. For ELR, the learning rate and ratio were adjusted.

[0128] FunCoup graph structure FunCoup inferred functional relationships using Bayesian algorithms and ortholog-based signaling. FunCoup edges are assigned confidence scores to support different interactions, including metabolic, signaling, and complex and physical protein-protein interactions. It integrates data from multiple levels of biological processes, such as co-expression, phylogenetic profile similarity, PPI, intracellular co-localization, and genetic interaction profile similarity. It contains 5,505,787 edges covering 16,765 gene nodes.

[0129] HumanNet v2 graph structure HumanNetV2 includes functional associations inferred from PPIs, gene co-expression, protein co-occurrence, and genomic context. It integrates these diverse types of omics data using a Bayesian framework. HumanNetV2 includes 476,399 edges covering 16,243 gene nodes.

[0130] Analysis and Verification Quantitative score of protein / pathway abnormality severity The GCN model score was generated by extracting the probability score generated by the model and by inverting the activation function involved in the binary pattern of predictions. In one embodiment, the activation function is a sigmoid defined as σ(x) = 1 / (1 + exp(-x)). The inverse logit function can also be used to calculate the score: logit(p) = σ -1 (p) = ln(p / (1-p))

[0131] Tumor types excluded from TP53 mutation status analysis Tumor types from n < 13 samples were excluded from analysis in either the labeling group (mutant or non-mutant samples).

[0132] Downsampling criteria for further analysis of tumor type, genes, and pathways Tumor types were excluded if they had fewer than 100 samples (UCS, KICH, ACC, READ), or if they had a mutation (or non-mutation) percentage below 15% and fewer than 200 samples (LAML, ESCA).

[0133] METABRIC analysis Gene expression samples were downloaded using the MetaGxBreast package, while mutation annotations were downloaded from www.cbioportal.org. To fit the values ​​to the TCGA normalized score, Z-scores were calculated using normalization across all METABRIC samples, and then the GCN model was applied to the normalized score. ROC curves and AUC scores were calculated and visualized in R.

[0134] GCN scores compared by variant and TP53_PROF classification TP53 mutation status was obtained from TCGA annotations downloaded from cBioPortal. Deletions and insertions were combined for in-frame and frameshift mutations. The TP53_PROF classification for harmful or non-harmful TP53 missense variants was extracted from Supplementary Table S5 from Ben Cohen et al. in Briefings in Bioinformatics Vol.23, Iss.2, March 2022 (copied in the appendix).

[0135] Comparison of variants and TP53_PROF Because TP53_PROF predicts only a small number of samples as non-harmful for any particular tumor type, TP53 variant comparisons were performed on pan-cancer data, taking power analysis into consideration. Teardowns included nonsense and fusion mutations. Splice mutations included splice sites, splice regions, and translation sites. In-frame and frameshift mutations included both insertions and deletions. Missense variants were separated into harmful (D) or non-harmful (ND) according to TP53_PROF predictions.

[0136] GDSC cell line drug response analysis Gene expression, drug response, and mutation data for GDSCs were downloaded from appropriate websites. Gene expression was derived using the Affymetrix Human Genome U219 Array and normalized using RMA. To fit to TCGA normalization, Z-scores were calculated using normalization across all GDSC cell lines, and then the GCN model was applied to the normalized scores. Tumor types and gene models were selected if: (i) they had a GCN model trained against them (i.e., they passed the downsampling criteria); or (ii) the corresponding cell line in GDSCs had data on the response to drugs targeting the query gene. Four genes—ERBB2, BRAF, EGFR, and KRAS—fit this threshold. For each, a GCN-trained tumor type was selected, and compounds targeting this gene were selected. Therefore, five analyses were performed: an ERBB2 model in breast cancer using the drug afatinib (the only model for ERBB2), BRAF in SKCM using the drugs dabrafenib and PLX-4720, and KRAS in LUAD using the KRAS(G12C) inhibitor. Since glioblastoma patients with EGFR mutations do not respond to this drug, EGFR in glioblastoma using the drug lapatinib (the only model for EGFR) was used as a negative control.

[0137] Data Availability TCGA gene expression data is available from the following source: Academic Torrents, a product of the Institute for Reproducible Research (a US501(c)3 nonprofit), were used for differential expression analysis using the TCGABiolinks package. Clinical and mutational annotations were downloaded from www.cbioportal.org and compared with the annotations in the TCGABiolinks package. The MSigDB set was downloaded from the Broad Institute's Gene Set Enrichment Site. GCN application code with graph structures for Funcoup 4 and Humannet V2 graph structures can be downloaded from the repository of Dutil et al. (Biorxiv, 2018), which is available at Bertinus / Gene-Graph-Analysis, a public site that also provides a copy of the reference "Analysis of Gene Interaction Graphs as Prior Knowledge for Machine Learning Models". TP53_PROF prediction is available from Ben-Cohen et al, 2022, TP53_PROF: a machine learning model to predict impact of missense mutations in TP53, Briefings in Bioinformatics, Volume 23, Issue 2, March 2022, bbab524.

[0138] GDSC data (gene expression, mutation, and clinical) was downloaded from the Genomics of Drug Sensitivity in Cancer site, jointly sponsored by the Wellcome Sanger Institute and the Massachusetts General Hospital Cancer Center. This site is a repository for approximately 1000 characterized human cancer cell lines screened with 100 compounds, allowing users to find drug response data and genomic markers of sensitivity. METABRIC data (gene expression and clinical) was downloaded using the MetaGxBreast package in R. METABRIC mutation data was downloaded from cBioPortal.

[0139] Figure 10 shows a multivariate analysis of the TCGA pan-cancer cohort against a Cox proportional model, quartiled based on the linear STAMP score. Survival distinctions are statistically significant, even after accounting for tumor type, age, and mutational burden.

[0140] Figures 11A–11E show that survival distinctions are enhanced by STAMP predictions compared to the original mutation state.

[0141] Figure 11A shows the results for pancreatic cancer, and Figures 11B–11E show the results for several tumor subtypes in TCGA (pancreatic adenocarcinoma, PAAD; hepatocellular carcinoma, LIHC; endothelial carcinoma, UCEC; colorectal adenocarcinoma, COAD, respectively). For COAD, STAMP did not improve survival differentiation. For further explanation of this analysis, please refer to the supplementary information below.

[0142] Figures 12A-12I. SBS analysis of APOBEC and HR-DDR signatures in breast cancer (BRCA, all signatures) and head and neck cancer (HNSC, APOBEC signature).

[0143] Figure 12A -- STAMP enhances SBS enrichment for SBS3 in BRCA.

[0144] Figure 12B -- SBS13 in BRCA.

[0145] Figure 12C -- SBS2 in HNSC

[0146] Figures 12D-12F -- When separating true WT (blue-1201), true mutant (red-1203), pseudo-mutant (green-1205), and pseudo-WT (yellow-1207), SBS enrichment is observed even among samples with the same actual mutational state. (Figures 12D-12F relate to SBS3 in BRCA, SBS13 in BRCA, and SBS2 in HNSC, respectively).

[0147] Figures 12G to 12I show that the linear scores of STAMP correlate with SBS enrichment (g to i relate to SBS3 in BRCA, SBS13 in BRCA, and SBS2 in HNSC, respectively). For further explanation of this analysis, please refer to the supplementary information below.

[0148] Figures 13A-13F. SBS analysis of APOBEC and HR-DDR signatures in breast cancer (BRCA, all signatures) and head and neck cancer (HNSC, APOBEC signature).

[0149] Figures 13A-13D -- STAMP enhances SBS enrichment. (Figures 13A-13D relate to SBS13 in pan-cancer, SBS2 in pan-cancer, SBS2 in BRCA, and SBS13 in HNSC, respectively).

[0150] Figures 13E and 13F are similar to Figures 12d to 12f, but relate to SBS2 in BRCA and SBS13 in HNSC, respectively.

[0151] Figures 14A to 14K.

[0152] Figures 14A-14H -- The linear scores of STAMP correlate with the TP53_PROF-based missense mutation classification in the following conditions: low-grade glioma (LGG), colorectal cancer (COAD), lung adenocarcinoma (LUAD), gastric adenocarcinoma (STAD), head and neck cancer (HNSC), breast cancer (BRCA), uterine corpus endothelial carcinoma (UCEC), and bladder cancer (BLCA).

[0153] Figure 14I -- Linear scores of STAMP based on their TP53 mutations for each sample. Frameshift (red), truncated (yellow), and spliced ​​(green) variants all behave similarly to harmful (D) missense variants. Unexpectedly, in-frame (purple) samples also behave similarly to D missense variants. Figures 14J-14K are similar to Figure 14I, but for the following specific tumor type cohorts:

[0154] Figure 14J--BRCA, Figure 14K--LUAD. In-frame variants appear to show a pattern closer to ND in these tumor-type-specific cohorts.

[0155] Figures 15A to 15D.

[0156] Figure 15A -- Applying STAMP to various pathways in several tumor type cohorts. The gray slot (1501) shows combinations of pathways and tumor types that did not meet the downsampling criteria.

[0157] Figure 15B -- STAMP predictions enhance survival discrimination based on the LATS2 model in low-grade gliomas (LGGs) compared to annotations by Sanchez-Vega et al. (labels used for training; compare these with Figure 3B).

[0158] Figure 15C -- The linear scores of STAMP can form three distinct groups in the LATS2 LGG cohort when Q3 and Q4 samples are combined (red - 1503), Q1 is blue (1505), and Q2 is orange (1507).

[0159] Figure 15D -- Multivariate analysis shows that the LATS2 model of STAMP in LGG maintains statistical significance in survival analysis when age and mutational burden are considered.

[0160] Figures 16A to 16I. GDSC analysis.

[0161] Figure 16A -- Skin cancer (SKCM) BRAF model using GDSC cell lines and their PLX-4720 IC 50 The values ​​are stratified. Cell lines are color-coded according to their BRAF mutation status (orange 1603, blue 1605).

[0162] Figure 16B -- Dabrafenib IC 50 Similar to 16A regarding this matter.

[0163] Figure 16C -- Samples containing only the TP53 mutation are further stratified by a linear STAMP score. Most of the mutant samples have high IC53. 50 , and have high linear score values ​​(red circle 1601, similar to the circle in Figure 5A). A small portion of the cell line exhibits WT-like IC5. 50 The levels are presented, and also have lower STAMP linear score values ​​("fake WT" samples given by the blue circle 1601, similar to the circle in Figure 5A).

[0164] Figures 16D-16E show that PLX-4720 and dabrafenib, respectively, correlate with STAMP scores in mutant-only samples, exhibiting boundary-level statistical significance (likely due to small sample size).

[0165] Figure 16F -- KRAS models in lung adenocarcinoma (LUAD) were used as negative controls against KRAS G12C inhibitors, and IC for STAMP scores. 50 There is no significant correlation between the levels. Samples with BRAF mutations are stained red (1607), and WT BRAF samples are stained blue (1609).

[0166] Figure 16G -- Lack of correlation indicates unclear IC for BRAF mutation (red-1611) compared to WT (blue-1613) cell line. 50 It matches the level.

[0167] Similar to Figures 16H, 16I, 16F, and 16G, in this study only, samples containing the G12C mutation in KRAS are colored red (1615), while the remaining samples are colored blue (1617).

[0168] Figure 17. A preliminary investigation of different graph structures for training STAMP against breast cancer based on TP53 mutation status. Numerous experiments were conducted, suggesting that funcoup and humannetV2 are preferred in genemania and regnet graphs. Dropout was also examined and did not show any improvement in performance. Further model adjustments along similar lines may be made as needed.

[0169] Figures 18A and 18B show ICs for WT and M groups in the pan-cancer dataset. 50 This illustrates the correlation between the scores.

[0170] Computer hardware This system and method may include implementations on systems that provide multiprocessor, multitasking, multiprocess, and / or multithreaded computing, as well as implementations on systems that provide only single-processor, single-threaded computing. Multiprocessor computing involves performing computing using two or more processors. Multitasking computing involves performing computing using two or more operating system tasks. A task is an operating system concept that refers to a combination of a program being executed and bookkeeping information used by the operating system. Whenever a program is executed, the operating system creates a new task for it. A task is like an envelope for a program in that it identifies the program by a task number and attaches other bookkeeping information to it. Many operating systems, including Linux, UNIX®, OS / 2®, and Windows®, can run many tasks simultaneously and are called multitasking operating systems. Multitasking is the ability of an operating system to run two or more executable files simultaneously. Each executable file runs within its own address space, meaning that executable files do not have a way of sharing any of their memories. This has advantages because it is impossible for any program to interfere with the execution of any other program running on the system. However, programs have no way of exchanging information other than through the operating system (or by reading files stored in the file system). While the terms task and process are often used interchangeably, and multi-process computing is similar to multi-task computing, some operating systems distinguish between the two.

[0171] The present invention may be a system, method, and / or computer program product in any possible level of technical detail. The computer program product may include a computer-readable storage medium having computer-readable program instructions for causing a processor to perform aspects of the present invention. The computer-readable storage medium may be a tangible device capable of holding and storing instructions for use by an instruction execution device.

[0172] Figure 19 shows a computing device 1900 related to this disclosure. The computing device 1900 may, for example, perform calculations, execute routines and algorithms, process data, communicate with other devices over a network, and display results. For example, the computing device 1900 may include a processor or CPU 1904, a network adapter 1906 for communicating with a network 1908, and the network 1908 may connect the computing device 1900 to other devices 1950 or other data (not shown). The computing device may also include an input / output device 1902. Such an input / output component 1902 may be an input device, an output device, or both, and the computing device 1900 may have several such components. Exemplary input devices 1902 include keyboards, mice, microphones, touchpads, joysticks, etc. Exemplary output devices 1902 include displays, speakers, haptic feedback devices, etc. The computing device 1900 may further include memory 1910 or a computer-readable storage medium 1910. Computer memory 1910 may contain instructions for performing methods and techniques described elsewhere in this disclosure. Computer memory 1910 may also include an operating system 1930 for controlling various parts and components of computing device 1900. Memory 1910 may also store data, such as training data 1912, test data 1914, and validation data 1916. Memory 1910 may also include algorithms, such as graph convolutional network classifier 1918, other classifiers 1920, machine learning algorithms 1922, score calculation algorithms 1924, or other algorithms 1926.

[0173] Computer-readable storage media may include, for example, but are not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any preferred combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or grooved raised structures on which instructions are recorded, and any preferred combination thereof. The computer-readable storage media used herein should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical fiber cables), or electrical signals transmitted through wiring.

[0174] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in the computer-readable storage medium within each computing / processing device.

[0175] The computer-readable program instructions for performing the operation of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk and C++, and procedural programming languages ​​such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by personalizing the electronic circuit using state information of computer-readable program instructions in order to perform aspects of the present invention.

[0176] Aspects of the present invention will be described herein with reference to flowcharts and / or block diagrams, or both, of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It will be understood that each block in a flowchart and / or block diagram, and combinations of blocks in a flowchart and / or block diagram, can be implemented by computer-readable program instructions. These computer-readable program instructions can be provided to a general-purpose computer, a dedicated computer, or a processor of another programmable data processing device in order to produce a machine, such that instructions executed via the processor of the computer or other programmable data processing device create means for implementing functions / operations specified in the flowchart and / or block diagram blocks.

[0177] These computer-readable program instructions may also be stored on computer-readable storage media that can instruct computers, programmable data processing devices, and / or other devices to function in a particular manner, such that the computer-readable storage medium on which the instructions are stored contains products that include instructions that implement the modes of function / operation specified in flowcharts and / or block diagram blocks.

[0178] Computer-readable program instructions may also be loaded onto a computer, another programmable device, or another device to cause the computer, another programmable device, or another device to perform a series of actions in order to generate a computer-executed process that implements a function / operation specified in a flowchart and / or block diagram block.

[0179] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or part of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions described in a block may occur in an order different from the order shown in the diagram. For example, two blocks shown consecutively may actually be executed substantially simultaneously or in reverse order depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs a specified function or operation, or executes a combination of dedicated hardware and computer instructions.

[0180] Exemplary Embodiments Example 1: A method comprising the steps of: receiving a set of curated data relating to human genes, wherein the curated data is RNA expression data; dividing the curated data into a first dataset, a second dataset, and a third dataset; training a classifier using the first dataset to identify data as mutant or wild-type; validating the classifier against the second dataset; testing the classifier against the third dataset; receiving new data; applying the trained classifier to the newly received data to classify the newly received data as mutant or wild-type; calculating a score associated with the newly received data; and reporting the classification and score associated with the newly received data to a user device.

[0181] Example 2: The method according to Example 1, wherein the classifier includes a graph convolutional network, a random forest model, or a support vector machine model.

[0182] Example 3: The method according to Example 1 or 2, wherein the score is the reciprocal of the activation function of the second-to-last layer of the classifier.

[0183] Example 4: The method according to any one of Examples 1 to 3, further comprising using a score to predict the efficacy of a drug against a cell line or tumor.

[0184] Example 5: Efficacy prediction, pharmaceutical informed consent 50 The method according to any one of Examples 1 to 4, including estimating the following.

[0185] Example 6: The method according to any one of Examples 1 to 5, further comprising identifying an active molecular pathway and a pharmaceutical product associated with the active molecular pathway based on a score.

[0186] Example 7: The method according to any one of Examples 1 to 6, further comprising the step of diagnosing a patient having a certain condition based on a score.

[0187] Example 8: The method according to any one of Examples 1 to 7, further comprising treating a patient based on a diagnosis and a prediction of the efficacy of a drug against a cell line or tumor based on a score.

[0188] Example 9: A method for treating a medical condition of interest, comprising: receiving data of the target gene; applying a trained classifier to classify the received data as a mutant or wild type; calculating a score associated with the received data; predicting the efficacy of a drug against a cell line or tumor associated with the received data; reporting the classification, calculated score, drug, and predicted efficacy to a user device; and using the drug to treat the medical condition of interest, if indicated by the calculated score.

[0189] Example 10: The method according to Example 9, wherein the classifier includes a graph convolutional network, a random forest model, or a support vector machine model.

[0190] Example 11: The method according to Example 9 or 10, wherein the calculated score is the reciprocal of the activation function of the second-to-last layer of the classifier.

[0191] Example 12: A system for calculating a score, comprising a processor and computer memory, wherein the computer memory includes computer instructions, and when the computer instructions are executed by the processor, the system implements a method comprising: receiving a set of curated data relating to human genes, wherein the curated data is RNA expression data; dividing the curated data into a first dataset, a second dataset, and a third dataset; training a classifier using the first dataset to identify the data as mutant or wild-type; validating the classifier against the second dataset; testing the classifier against the third dataset; receiving new data; applying the trained classifier to the received new data to classify the new data as mutant or wild-type; calculating a score associated with the received new data; and reporting the classification and score associated with the new data to a user device.

[0192] Example 13: The system according to Example 12, wherein the classifier includes a graph convolutional network, a random forest model, or a support vector machine model.

[0193] Example 14: The system according to Example 12 or 13, wherein the score is the reciprocal of the activation function of the second-to-last layer of the classifier.

[0194] Example 15: The system according to any one of Examples 12-14, further comprising using a score to predict the efficacy of a drug against a cell line or tumor.

[0195] Example 16: A system according to any one of Examples 12 to 15, wherein predicting effectiveness includes estimating the IC of a medicament. 50

[0196] Example 17: A system according to any one of Examples 12 to 16, further comprising identifying an active molecular pathway and a medicament associated with the active molecular pathway based on a score.

[0197] Example 18: A system according to any one of Examples 12 to 17, further comprising diagnosing a patient having a certain condition based on a score.

[0198] Example 19: A system according to any one of Examples 12 to 18, further comprising treating a patient based on a diagnosis and a prediction of the effectiveness of a medicament for a cell line or tumor based on a score.

[0199] Although specific embodiments of the present invention have been described, those skilled in the art will understand that there are other embodiments that are equivalent to the described embodiments. Therefore, it should be understood that the present invention is not limited by the specific exemplary embodiments shown, but only by the appended claims.

[0200] Appendix

[0201] [Table 5-1]

[0202] [Table 5-2]

[0203] [Table 5-3]

[0204] ​Table 5-4

[0205] Table 5-5

[0206] Table 5-6

[0207] Table 5-7

[0208] Table 5-8

[0209] Table 5-9

[0210] Table 5-10

[0211] Table 5-11

[0212] Table 5-12

Claims

1. It is a method, Receiving a curated set of data relating to human genes, wherein the curated data is RNA expression data, The curated data is divided into a first dataset, a second dataset, and a third dataset. Training a classifier using the first dataset to identify data as either variants or wild types, The classifier is validated against the second dataset, The classifier is tested on the third dataset, Receiving new data and, Applying the trained classifier to the newly received data to classify the newly received data as either a variant or wild type, Calculate the score associated with the newly received data, A method comprising reporting the classification and score associated with the newly received data to a user device.

2. The method according to claim 1, wherein the classifier includes a graph convolutional network, a random forest model, or a support vector machine model.

3. The method according to claim 2, wherein the score is the reciprocal of the activation function of the second-to-last layer of the classifier.

4. The method according to claim 1, further comprising using the score to predict the efficacy of a drug against a cell line or tumor.

5. The prediction of the aforementioned efficacy is the informed consent of the pharmaceutical 50 The method according to claim 4, which includes estimating the following.

6. The method according to claim 1, further comprising identifying an active molecular pathway and a pharmaceutical product associated with the active molecular pathway based on the score.

7. The method according to claim 1, further comprising the step of diagnosing a patient having a certain condition based on the score.

8. The method according to claim 7, further comprising treating the patient based on a diagnosis and a prediction of the effectiveness of a drug against a cell line or tumor based on the score.

9. A method of treating the medical condition of the person requiring treatment, The process of receiving data of the target gene, A step of applying a trained classifier to classify the received data as either a variant or a wild type, A step of calculating a score associated with the received data, A step of predicting the efficacy of a drug against a cell line or tumor associated with the received data, A step of reporting the classification, the calculated score, the drug, and the predicted efficacy to the user device, A method comprising the step of using the pharmaceutical to treat the medical condition of the subject, as indicated by the calculated score.

10. The method according to claim 9, wherein the classifier includes a graph convolutional network, a random forest model, or a support vector machine model.

11. The method according to claim 9, wherein the calculated score is the reciprocal of the activation function of the second-to-last layer of the classifier.

12. A system for calculating scores, Processor and The computer memory is modified so that it contains computer instructions, and when the computer instructions are executed by the processor, A step of receiving a set of curated data relating to human genes, wherein the curated data is RNA expression data, The process involves dividing the curated data into a first dataset, a second dataset, and a third dataset. The process involves training a classifier using the first dataset to identify data as either variants or wild-types, A step of validating the classifier against the second dataset, A step of testing the classifier on the third dataset, The process of receiving new data, The steps include applying the trained classifier to the newly received data to classify the new data as either a variant or a wild type, The process of calculating a score associated with the newly received data, A system including computer memory that performs a method comprising the step of reporting the classification and the score associated with the new data to a user device.

13. The system according to claim 12, wherein the classifier includes a graph convolutional network, a random forest model, or a support vector machine model.

14. The system according to claim 13, wherein the score is the reciprocal of the activation function of the second to last layer of the classifier.

15. The system according to claim 12, further comprising using the score to predict the efficacy of a drug against a cell line or tumor.

16. The prediction of the aforementioned efficacy is the informed consent of the pharmaceutical 50 The system according to claim 15, which includes estimating the following.

17. The system according to claim 12, further comprising identifying an active molecular pathway and a pharmaceutical product associated with the active molecular pathway based on the score.

18. The system according to claim 12, further comprising the step of diagnosing a patient having a certain condition based on the score.

19. The system according to claim 18, further comprising treating the patient based on a diagnosis and a prediction of the effectiveness of a drug against a cell line or tumor based on the score.