Systems and methods for detecting cellular pathway dysregulation in cancer specimens - Patents.com

By employing RNA expression data and machine learning models, the system effectively detects cellular pathway perturbations in cancer cells, overcoming limitations of DNA-based methods and enhancing treatment efficacy.

JP2025517450APending Publication Date: 2025-06-05テンパスエーアイインコーポレイテッド
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024568982
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-20
Filing Date
2023-02-28
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Current methods struggle to detect pathway perturbations in cancer cells, especially when genetic mutations are unknown or of unknown significance, leading to incorrect identification of treatment responders.

Method used

The use of RNA expression level information and machine learning models to determine cellular pathway perturbations, allowing for the identification of genetic variants that alter pathway activity and correlation with disease state and progression.

Benefits of technology

This approach enables accurate detection of pathway perturbations, improving the identification of effective treatments and avoiding inappropriate therapies by considering RNA expression data beyond DNA variants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517450000001_ABST
    Figure 2025517450000001_ABST
Patent Text Reader

Abstract

Disclosed herein are systems, methods, and compositions useful for determining cellular pathway perturbations, including the use of RNA expression level information. The determined perturbation levels can aid in identifying genetic variants that alter pathway activity, correlating these variants with disease states and disease progression, and identifying therapeutic agents that are most likely to be effective and those that should be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Patent Application No. 17 / 750,055, filed May 20, 2022, the contents of which are incorporated by reference in their entirety herein. [Background technology]

[0002] Tumor formation and tumor maintenance are believed to be driven primarily by oncogenes and / or the disruption of their signaling pathways. Well-studied examples of such oncogenes and their associated pathways include the receptor tyrosine kinase (RTK) / Ras and phosphoinositide 3-kinase (PI3K) pathways. Many different pathways have been correlated with specific types of cancer, and indeed mutations in genes in these pathways have been identified as drivers of certain cancers. These driver genes and their gene products are therefore important targets of drug development efforts, which have resulted in many life-saving and life-prolonging treatment options for certain patients.

[0003] However, not all cancers are associated with known genetic mutations or known pathways. For example, DNA analysis may detect variants of unknown significance (VUS) in oncogenic signaling pathways. Variants of unknown significance (VUS) are changes with unknown functional consequences and may represent benign passenger mutations (having little or no effect on cellular activity) or may be pathogenic (e.g., novel uncharacterized disease-causing mutations). In some instances, there is no information about the variants because they are rare or difficult to study. These variants may or may not have clinical significance and cannot be distinguished by DNA analysis alone. Thus, some mutations in genes known to interact with or affect a pathway do not alter the activity of the pathway and DNA analysis may yield false positives, i.e., patients who do not respond to targeted therapy may be erroneously identified as responders by DNA analysis. Summary of the Invention [Problem to be solved by the invention]

[0004] Therefore, there is a need in the art to detect pathway perturbations using information other than DNA variants. [Means for solving the problem]

[0005] Disclosed herein are systems, methods, and compositions useful for determining cellular pathway perturbations, including the use of RNA expression level information. By way of example and not limitation, the determined perturbation levels can be used to (1) aid in the identification of genetic variants that alter pathway activity, (2) correlate the identified variants with disease state and disease progression, and (3) identify treatments that are most likely to be effective and treatments that should be avoided.

[0006] In some embodiments, a computer-implemented method of training a machine learning model for detecting dysregulation in a cellular pathway is provided. In some embodiments, the method includes receiving a query including a positive control reference and a negative control reference, the positive control reference including at least one genetic variation condition. In some embodiments, the method includes obtaining a positive control group and a negative control group in an electronic format from a data store including a plurality of cell samples, the positive control group including a cell sample having a genetic variation matching the at least one genetic variation condition of the positive control reference, the negative control group including a cell sample having a genetic attribute matching the negative control reference, and each sample of the plurality of cell samples including genetic data for a plurality of genes of the sample, and transcriptomic data including RNA expression levels. In some embodiments, the method includes training a machine learning model using the positive control group and the negative control group to determine a correlation between at least one genetic variation condition and a pathway dysregulation score. In some embodiments, the method includes generating a score for the machine learning model, the score indicating a degree of accuracy of the machine learning model. In some embodiments, the cell sample of the negative control group does not contain a genetic mutation consistent with at least one genetic mutation status. In some embodiments, training the machine learning model comprises identifying one or more feature genes and determining a weight for each feature gene of the one or more feature genes, the weight indicating the effect of the feature gene on the pathway dysregulation score.

[0007] In some embodiments, a system for training a machine learning model for detecting dysregulation in a cellular pathway is provided. In some embodiments, the system includes a computer including a processor. In some embodiments, the processor is configured to receive a query including a positive control standard and a negative control standard, and obtain a positive control group and a negative control group in electronic format from a data store including a plurality of cell samples, such that the positive control standard includes at least one genetic mutation status, the positive control group includes a cell sample having a genetic mutation that matches the at least one genetic mutation status of the positive control standard, the negative control group includes a cell sample having a genetic attribute that matches the negative control standard, and each sample of the plurality of cell samples includes genetic data for a plurality of genes of the sample, and transcriptome data including RNA expression levels, to train a machine learning model using the positive control group and the negative control group to determine a correlation between at least one genetic mutation status and a pathway dysregulation score, and generate a score for the machine learning model, the score indicating the accuracy of the machine learning model. In some embodiments, the cell sample of the negative control group does not include a genetic mutation that matches the at least one genetic mutation status. In some embodiments, training the machine learning model comprises identifying one or more feature genes and determining a weight for each feature gene of the one or more feature genes, the weight indicating an influence of the feature gene on the pathway dysregulation score.

[0008] In some embodiments, a non-transitory computer-readable storage medium is provided. In some embodiments, the non-transitory computer-readable storage medium has stored thereon instructions that, when executed by a processor, cause the processor to receive a query including a positive control standard and a negative control standard, obtain a positive control group and a negative control group in electronic format from a data store including a plurality of cell samples, such that the positive control standard includes at least one genetic mutation status, the positive control group includes a cell sample having a genetic mutation that matches the at least one genetic mutation status of the positive control standard, and the negative control group includes a cell sample having a genetic attribute that matches the negative control standard, and each sample of the plurality of cell samples includes genetic data for a plurality of genes of the sample, and transcriptome data including RNA expression levels, generate a score for the machine learning model, and generate a score indicating the accuracy of the machine learning model. In some embodiments, the cell sample of the negative control group does not include a genetic mutation that matches the at least one genetic mutation status. In some embodiments, training the machine learning model comprises identifying one or more feature genes and determining a weight for each feature gene of the one or more feature genes, the weight indicating an influence of the feature gene on the pathway dysregulation score. [Brief description of the drawings]

[0009] [Figure 1A] 1 shows examples of signal transduction pathways. [Figure 1B] Show custom routes. [Figure 2A] FIG. 1 is a schematic diagram illustrating an exemplary concept of the systems and methods disclosed herein. [Figure 2B] FIG. 2 is a schematic diagram illustrating another exemplary concept of the systems and methods disclosed herein. [Figure 3A] 1 shows a schematic diagram of a system capable of determining a pathway perturbation state of at least one tissue specimen. [Figure 3B] 1 is a schematic example of a device that can be used in the system. [Figure 3C] 3A and 3B. FIG. 3B illustrates an example of hardware that may be used in some embodiments of the system of FIG. [Figure 4] 1 shows a representation of example data from a data input that may be used to train a route engine. [Diagram 5] 1 illustrates an example of a process by which a route engine can be trained. [Figure 6A] 1 illustrates a process by which an alpha parameter value can be selected for training the route engine. [Figure 6B] 1 illustrates a process by which the pathway engine can be tested using additional test transcriptomes for any given test. [Figure 6C-D] Figure 6C shows an example result of a Wilcoxon rank sum test used to analyze the pathway perturbation scores (used interchangeably with "pathway dysregulation scores") generated by the pathway engine. Figure 6D shows another example result of a Wilcoxon rank sum test used to analyze the pathway perturbation scores generated by the pathway engine. [Figure 6E] 1 illustrates an exemplary process by which a trained pathway engine can be biologically validated. [Figure 6F] We show a process by which a trained route engine can be orthogonally validated. [Figure 6G] 1 illustrates an exemplary process for training a model. [Figure 6H] 1 illustrates a process by which training data can be selected for training a model. [Figure 6I] 1 shows an exemplary model of the RTK-RAS and PI3K pathways with multiple modules. [Figure 6J] Variants of unknown significance (VUS) in the AKT module are shown. [Figure 6K] Pathways with pathogenic variants in the TSC1 module are shown. [Figure 6L]Pathways with pathogenic mutations in the PTEN module are shown. [Figure 6M] We show that genes can be connected to each module involved in the RTK-RAS and PI3K pathways. [Figure 6N] 1 shows the distribution of EGFR pathway dysregulation scores for somatic pathogenic mutations in the EGFR and wild-type cohorts in the holdout set. [Figure 6O-P] Figure 6O shows the scores generated using the TOR model, and Figure 6P shows the probability distributions generated using Gaussian kernel density estimation. [Figure 6Q] The distribution of the cohorts is shown. [Figure 6R] Dysregulation scores in pathways are shown. [Figure 6S] Figure 6R shows pathogenic variants in the pathway and the TSC1 module. [Figure 6T] Figure 6R shows pathogenic mutations in the pathway and the PTEN module. [Figure 6U-V] Figure 6U shows the PIK3C dysregulation score and part of the pathway with pathogenic mutations in EGFR and PTEN. Figure 6V shows the NF1 gene connecting to the RAS pathway. [Figure 6W] Genes for the AKT module are shown individually. [Figure 6X] Genes for the RAS modules are shown individually. [Figure 6Y] 1 illustrates an exemplary data frame that may be generated based on VUS data. [Figure 6Z] 1 shows an exemplary histogram of total global dysregulation scores. [Figure 7A] The results of mutations in NF1 with greater than one cohort for all possible metabolic pathways are shown. [Figure 7B] The results of different mutations in NF1 with more than one cohort for all possible metabolic pathways are shown. [Figure 7C] 1 illustrates an exemplary process by which a trained route engine can be used to generate a route perturbation score. [Figure 8A-D] Figure 8A shows a pie chart of the cancers of interest. Figure 8B shows a pie chart of subsetting the cancer types of Figure 8A by mutation status. Figure 8C shows various graphs of differentially expressed genes (DEGs) between groups. Figure 8D shows the results of validation of the logistic regression model. [Figure 9A-B] Figure 9A shows an example of validation results using an external data set, and Figure 9B shows an example of biological validation results using protein activation data. [Figure 10A] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 10B] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 10C] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 10D] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 10E] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 10F] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 10G] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 10H] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 10I] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 11A] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 11B] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 11C] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 11D] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 11E] 7D shows an example of a route disturbance report generated using the process of FIG. 7C. [Figure 12A] 1 shows the results of a patient transcriptome being analyzed by the multiple pathway engine. [Figure 12B] Showing more results of patient transcriptomes being analyzed by multiple pathway engines. [Figure 12C] 13 shows further results of patient transcriptomes being analyzed by the multiple pathway engine. [Figure 12D] 13 shows further results of patient transcriptomes being analyzed by the multiple pathway engine. [Figure 12E] 13 shows additional results of patient transcriptomes being analyzed by the multiple pathway engine. [Figure 12F] 13 shows additional results of patient transcriptomes being analyzed by the multiple pathway engine. [Figure 13] FIG. 1 is a schematic diagram showing the integration of clinical and molecular data and data science resources with the expertise of drug development companies in translating knowledge into products. [Figure 14] 1 is an example of using the system and method to analyze transcriptomes from a cohort of LUAD patients. [Figure 15A-B] FIG. 13 is an example testing the ability of an alternative method to separate positive from negative controls by dimensionality reduction using DEGs and pathway scores. [Figure 16A-B] Collectively, the systems and methods disclosed herein demonstrate the ability to distinguish between negative and positive controls for a pathway of interest. [Figure 17A-B] 1 shows area under the curve (AUC) and predictive performance graphs illustrating that the systems and methods disclosed herein can distinguish between negative and positive controls for the RAS pathway. [Fig. 17C-D] 1 shows AUC and predictive performance graphs demonstrating that the systems and methods disclosed herein can distinguish between negative and positive controls for the PI3K pathway. [Figure 18] 13 is a performance graph showing that other mutation groups exhibit the expected model output. [Figure 19A] 13 is a performance graph showing the results of validating the KRAS mutant vs. RAS pathway WT model on the TCGA lung adenocarcinoma cohort. [Figure 19B] 13 is a performance graph showing the results of validating the STK11 mutant vs. PI3K pathway WT model on the TCGA lung adenocarcinoma cohort. [Figure 20A] 1 is a graph showing the relationship between pathway perturbation scores generated by the systems and methods and protein expression levels of phosphorylated (ie, activated) MEK1. [Figure 20B] 1 is a graph showing the relationship between pathway perturbation scores generated by the systems and methods and protein expression levels of phosphorylated AMPK. [Figure 21] 1 is a graph showing that the system and method can distinguish between responder and non-responder groups to a particular treatment. [Figure 22] 7D illustrates an exemplary route perturbation report generated by the process of FIG. 7C. [Figure 23] 7D illustrates another example route perturbation report generated by the process of FIG. 7C. [Figure 24] 7D illustrates yet another exemplary route perturbation report generated by the process of FIG. 7C. [Diagram 25] 7D shows a further exemplary route perturbation report generated by the process of FIG. 7C. [Figure 26A] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26B] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26C] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26D]A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26E] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26F] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26G] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Fig. 26H] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26I] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26J] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26K] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26L] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26M] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26N] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26O] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26P]A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26Q] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26R] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26S] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26T] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Fig. 26U] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26V] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Fig. 26W] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Fig. 26X] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 26Y] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Fig. 26Z] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27A] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27B]A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27C] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27D] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27E] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27F] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27G] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Fig. 27H] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27I] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27J] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27K] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27L] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27M] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27N]A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27O] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27P] A table listing anti-neoplastic drugs is compiled, providing the drug name, site of action / tumor type, drug class, and general mechanism of action. [Figure 27Q] A table listing FDA approved anti-neoplastic drugs is compiled, providing the name of the drug, the site of action / tumor type, drug class, and at least one pathway affected by the drug. [Figure 27R] A table listing FDA approved anti-neoplastic drugs is compiled, providing the name of the drug, the site of action / tumor type, drug class, and at least one pathway affected by the drug. [Figure 27S] A table listing FDA approved anti-neoplastic drugs is compiled, providing the name of the drug, the site of action / tumor type, drug class, and at least one pathway affected by the drug. [Figure 27T] A table listing FDA approved anti-neoplastic drugs is compiled, providing the name of the drug, the site of action / tumor type, drug class, and at least one pathway affected by the drug. [Fig. 27U] A table listing FDA approved anti-neoplastic drugs is compiled, providing the name of the drug, the site of action / tumor type, drug class, and at least one pathway affected by the drug. [Figure 27V] A table listing FDA approved anti-neoplastic drugs is compiled, providing the name of the drug, the site of action / tumor type, drug class, and at least one pathway affected by the drug. [Figure 28] 1 shows a Violin plot showing STK11 perturbation score (Y-axis) and progression or no progression of disease (X-axis) after 6 months of immunotherapy regimen. [Figure 29]2B is a graph showing overall survival (Y-axis) versus time (X-axis) for patients with KRAS mutant lung adenocarcinoma with or without STK11 / LKB1 mutations treated with PD-1 inhibitors (Skoulidis et al, Cancer Discov. 2018 DOI:10.1158 / 2159-8290.CD-18-0099, Figure 2B, right panel). [Diagram 30] FIG. 1 is a graph showing two-dimensional clustering of 527 patients based on perturbation scores of component modules of the PI3K and RTK / RAS pathways. [Diagram 31] 1 illustrates an exemplary process for training a model. [Diagram 32] FIG. 1 is a diagram of an exemplary genetic mutation status for defining positive and negative control criteria, according to some embodiments. [Diagram 33] FIG. 1 is a diagram of an exemplary genetic mutation status for defining positive and negative control criteria, which may include thresholds, according to some embodiments. [Diagram 34] 1 is a diagram of example criteria by which a training data set may be filtered, according to some embodiments. [Diagram 35] FIG. 1 is a diagram of exemplary confounder types for which a training dataset can be filtered, according to some embodiments. [Diagram 36] 1 shows the composition of an exemplary training dataset indicated by potential confounder type. [Figure 37] 1 shows the composition of the positive and negative control groups (i.e., wild type) of an exemplary training dataset by cancer type of the samples in the training dataset. [Figure 38] 1 shows the probability distributions of the positive and negative control cancer cohorts for an unweighted model applied to correct for the imbalance in cancer composition between the positive and negative control groups. [Figure 39] 1 illustrates an exemplary process for weighting cancer samples in a training dataset to reduce overfitting of a model. [Diagram 40]FIG. 39 shows probability distributions for the training data set shown in FIGS. 37 and 38 with weightings applied to samples in the training data set according to the process shown in FIG. [Diagram 41] 1 shows an exemplary process for selecting genes to be used as feature genes, or inputs to a machine learning model. [Diagram 42] FIG. 1 is a graph of gene interactions in an exemplary gene regulatory network, where the size of a gene is determined by its decenteredness. [Diagram 43] A graph of gene interactions within the gene regulatory network shown in FIG. 42, where gene size is determined by the eigenvector centrality of a given gene. [Diagram 44] 32 shows probability distributions for the training, threshold, and holdout sets of a model trained according to the process shown in FIG. [Diagram 45] 32 shows the coefficients of feature genes of a model trained according to the process shown in FIG. 31 . [Figure 46] 32 shows a plot of probability distribution of KRAS variants resulting from a model trained according to the process shown in FIG. 31 . [Figure 47] FIG. 1 is an exemplary process for refining cohort definitions by identifying genetic variants significant for generating RNA signatures. [Figure 48] 1 is an exemplary process for generating a grid of models for comparing the impact of parameters used to generate and train the models. [Figure 49] FIG. 49 shows a probability distribution plot for a grid of models generated using the process shown in FIG. 48. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Various aspects of the present disclosure will now be described with reference to the drawings, in which like reference numerals correspond to like elements throughout the several views. It should be understood, however, that the drawings and the following detailed description in conjunction therewith are not intended to limit the claimed subject matter to the particular forms disclosed. Rather, the intent is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the claimed subject matter.

[0011] In the following detailed description, reference is made to the accompanying drawings, which form a part of this specification and which show by way of example specific embodiments in which the present disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the present disclosure. It should be understood, however, that the detailed description and specific examples, while showing examples of embodiments of the present disclosure, are given by way of illustration only and not by way of limitation. From this disclosure, it will be apparent to those skilled in the art that various substitutions, modifications, additional rearrangements, or combinations thereof may be made within the scope of the present disclosure.

[0012] According to common practice, various features shown in the drawings may not be drawn to scale. The figures presented herein are not meant to be actual views of any particular method, device, or system, but are merely idealized representations used to explain various embodiments of the present disclosure. Thus, dimensions of various features may be arbitrarily expanded or reduced for clarity. Furthermore, some of the drawings may be simplified for clarity. Thus, the drawings may not show all of the components of a given apparatus (e.g., device) or method. Moreover, similar reference numbers may be used to denote similar features throughout this specification and the drawings.

[0013] The information and signals described herein can be represented using any of a variety of different technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof. Some figures may show signals as a single signal for clarity of presentation and explanation. Those skilled in the art will understand that a signal may represent a bus of signals, and that the bus may have various bit widths, and that the present disclosure may be implemented with any number of data signals, including a single data signal.

[0014] The various exemplary logic blocks, modules, circuits, and algorithmic operations described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various exemplary components, blocks, modules, circuits, and operations have been described generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementation decisions should not be interpreted as departing from the scope of the embodiments of the present disclosure described herein.

[0015] Further, it should be noted that the embodiments may be described in terms of a process that is depicted as a flowchart, a flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe operational acts as a sequential process, many of these acts may be performed in another sequence, in parallel, or substantially simultaneously. Also, the order of the acts may be rearranged. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. Furthermore, the methods disclosed herein may be implemented in hardware, software, or both. If implemented in software, the functions may be stored or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media includes both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another.

[0016] It should be understood that any reference to an element herein using a designation such as "first," "second," etc. does not limit the quantity or order of those elements unless such limitation is expressly stated. Rather, these designations may be used herein as a convenient way of distinguishing between two or more elements or instances of an element. Thus, a reference to a first and a second element does not imply that only two elements may be used therein, or that the first element must precede the second element in any way. Also, unless otherwise stated, a set of elements may include one or more elements.

[0017] As used herein, terms like "component," "system," and the like are intended to refer to any computer-related entity, whether hardware, a combination of hardware and software, software, or software in execution. For example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable file, a thread of execution, a program, and / or a computer. By way of example, both an application running on a computer and a computer can be a component. One or more components may reside within a process and / or thread of execution, and a component may be localized on one computer and / or distributed between two or more computers or processors.

[0018] The word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects or designs.

[0019] Furthermore, the disclosed subject matter may be implemented as a system, method, apparatus, or article of manufacture using standard programming and / or engineering techniques to generate software, firmware, hardware, or any combination thereof for controlling a computer or processor-based device to implement the aspects detailed herein. As used herein, the term "article of manufacture" (or alternatively, "computer program product") is intended to encompass a computer program accessible from any computer-readable device, carrier, or media. For example, computer-readable media may include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips, etc.), optical disks (e.g., compact disks (CDs), digital versatile disks (DVDs), etc.), smart cards, and flash memory devices (e.g., cards, sticks).

[0020] It should further be understood that a carrier wave may be used to carry computer-readable electronic data such as those used in sending and receiving e-mail or in accessing a network such as the Internet or a local area network (LAN). Of course, those skilled in the art will recognize that many modifications can be made to this configuration without departing from the scope or spirit of the claimed subject matter.

[0021] The terms "polynucleotide", "nucleic acid" and "nucleic acid molecule" are used interchangeably and refer to a covalent sequence of nucleotides (i.e., ribonucleotides in RNA and deoxyribonucleotides in DNA) in which the 3' position of the pentose of one nucleotide is linked to the 5' position of the next pentose by a phosphodiester group, and include sequences of any form of nucleic acid, including but not limited to RNA, DNA and cfDNA molecules. These terms also refer to complementary DNA (cDNA), which is DNA synthesized from a single-stranded RNA (e.g., messenger RNA (mRNA) or microRNA (miRNA)) template in a reaction catalyzed by the enzyme reverse transcriptase. The term "polynucleotide" includes, but is not limited to, single-stranded and double-stranded polynucleotides.

[0022] As used herein, the terms "protein" and "polypeptide" are used interchangeably herein to refer to a series of amino acid residues linked to one another by peptide bonds between the α-amino and carboxy groups of adjacent residues.

[0023] The terms "protein" and "polypeptide" refer to polymers of proteinaceous amino acids, including modified amino acids (e.g., phosphorylated, glycosylated, glycosylated, etc.) and amino acid analogs. Although "protein" and "polypeptide" are often used in reference to relatively large polypeptides and the term "peptide" is often used in reference to small polypeptides, use of these terms in the art overlaps. Exemplary polypeptides or proteins include gene products, naturally occurring proteins, homologs, orthologs, paralogs, fragments and other equivalents, variants, fragments and analogs of the above.

[0024] As used herein, the term "chromosome" refers to the structure of nucleic acids and proteins (i.e., chromatin) found in the nucleus of most living cells that carries genetic information in the form of genes. The conventional internationally recognized human genome chromosome numbering system is used herein.

[0025] As used herein, the term "gene" refers to a nucleic acid sequence that codes for a gene product, either a polypeptide or a functional RNA molecule. The term "gene" should be interpreted broadly herein and encompasses both the genomic DNA form of a gene (i.e., a particular portion of a particular chromosome), as well as the mRNA and cDNA forms of the gene produced therefrom. During gene expression, the genomic DNA is transcribed into RNA, which may be immediately functional or may be translated into a polypeptide that performs a function. In addition to the coding region (i.e., the sequence that codes for the gene product), a gene contains "non-coding regions." The non-coding regions may be immediately adjacent to the coding region (e.g., the 5' and 3' non-coding regions adjacent to the coding region) or may be distant from the coding region (e.g., many kilobases upstream or downstream). Some non-coding regions, including "introns" (i.e., regions removed by RNA splicing before translation) and translational regulatory elements (e.g., ribosome binding sites, terminators, and start and stop codons), are transcribed into RNA but are not translated. Other non-coding regions, including essential transcriptional regulatory regions, are not transcribed. A gene requires a sequence, a "promoter," that recruits RNA polymerase and is recognized and bound by proteins (i.e., transcription factors) that help the RNA polymerase bind and initiate transcription. A gene can have more than one promoter, resulting in messenger RNA (mRNA) that differs in how much it extends at the 5' end. As used herein, a gene can also contain more distally located transcriptional regulatory elements (i.e., "enhancers" and "silencers") that can loop close to the promoter and allow proteins (i.e., "transcription factors") bound to these distal regulatory sites to affect transcription. For example, an "enhancer" increases transcription by binding to activating proteins that help recruit RNA polymerase or initiate transcription. Conversely, a "silencer" binds to a repressor protein that makes DNA less accessible to RNA polymerase or would otherwise inhibit transcription. A gene can also contain an "insulator" element that protects the promoter from inappropriate regulation.Insulators may function by blocking interactions with enhancers or silencers, or by acting as a barrier to prevent the diffusion of condensed chromatin. Enhancers and silencers are generally not considered to be part of the gene itself (given that a single enhancer or silencer may regulate the expression of multiple genes), although as used herein the term gene encompasses distal elements that affect its expression.

[0026] As used herein, the term "promoter" refers to a DNA sequence capable of controlling the expression of a coding sequence or functional RNA. Generally, the coding sequence is located 3' to the promoter sequence. Promoters may be derived in their entirety from a natural gene, or may be composed of different elements derived from different promoters found in nature, or may include synthetic DNA segments. It is understood by those skilled in the art that different promoters may direct the expression of a gene in different tissues or cell types, or at different developmental stages, or in response to different environmental conditions. In many cases, artificial promoters that cause a gene to be expressed in most cell types are commonly referred to as "constitutive promoters". Artificial promoters that allow the selective expression of a gene in most cell types are commonly referred to as "inducible promoters".

[0027] "Genetic analyzer" refers to a device, system, and / or method for determining the characteristics (e.g., sequence) of nucleic acid molecules (i.e., DNA, RNA, cDNA) present in a biological specimen. A "genetic analyzer" may also be used to characterize epigenetic features of nucleic acid molecules, for example, by using methods including bisulfite sequencing, chromatin immunoprecipitation followed by sequencing, assay of transposase-accessible chromatin using sequencing (ATAC-seq), or 3C-based technologies.

[0028] The terms "gene sequence" and "sequence" are used herein to refer to a series of nucleotides present in a DNA, RNA or cDNA molecule. In the context of the present invention, a sequence is determined by sequencing nucleic acids present in a biological specimen.

[0029] The term "read" refers to a DNA sequence of sufficient length (e.g., at least about 30 bp) that it can be used to identify a larger sequence or region, for example, by aligning it to a chromosome, genomic region, or gene.

[0030] As used herein, the term "reference genome" refers to any specific known genomic sequence, whether partial or complete, of any organism or virus that can be used to reference an identified sequence from a subject. Many reference genomes are provided by the National Center for Biotechnology Information at www.ncbi.nlm.nih.gov. "Genome" refers to the complete genetic information of an organism or virus expressed in a nucleic acid sequence.

[0031] As used herein, the terms "alignment", "alignment" or "registration" refer to a process used to identify regions of similarity. In the context of the present invention, alignment refers to matching sequences with positions in a reference genome based on the order of nucleotides in these sequences. Alignment can be performed manually or by computer algorithms using, for example, the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysis pipeline. Alignment can refer to either a 100% sequence match or a match less than 100% (an incomplete match).

[0032] The terms "library" and "sequencing library" as used herein refer to a pool of DNA fragments to which adapters are attached. The adapters are generally designed to interact with the surface of a particular sequencing platform, e.g., a flow cell (Illumina) or beads (Ion Torrent), to facilitate the sequencing reaction.

[0033] The terms "targeted panel" and "targeted gene sequencing panel" are used interchangeably herein to refer to a selected set of genes or genetic regions known or suspected to be associated with a particular disease or phenotype. Targeted panels are useful tools for detecting a specific set of mutations in a given sample because sequencing of targeted panels generates smaller, more manageable data sets compared to more extensive approaches such as whole genome sequencing.

[0034] The term "sequencing probe" or "sequencing primer" is used herein to refer to a short oligonucleotide used to sequence a nucleic acid (i.e., cDNA or DNA). A sequencing probe may hybridize to a target sequence within a nucleic acid or may hybridize to an adapter sequence attached to the nucleic acid to allow non-specific amplification and sequencing.

[0035] The term "RNA read count" is used herein to refer to the number of sequencing reads generated from a genetic analyzer. The term "RNA read count" is often used to refer to the number of reads that overlap with a given feature (e.g., a gene or chromosome).

[0036] The term "bioinformatics pipeline" is used herein to mean a series of processing steps of a pipeline for instantiating a bioinformatics report on next generation sequencing results obtained from a biological specimen. For example, in the context of the present invention, the goal of the pipeline may be to identify variants present in a patient's genome.

[0037] The term "genetic profile" is used herein to refer to information about particular genes in an individual or in a particular type of tissue. This information may include genetic variations (e.g., single nucleotide polymorphisms), gene expression data, other genetic characteristics, or epigenetic characteristics (e.g., DNA methylation patterns) determined, for example, by analysis of next generation sequencing data.

[0038] The term "variant" is used herein to mean a difference in a genetic sequence or genetic profile compared to a reference genome or reference genetic profile.

[0039] The term "expression level" is used herein to describe the copy number of a particular RNA or protein molecule, which may or may not be normalized using standard methods (e.g., counts per million, finding the base 10 logarithm of the raw read counts), generated by a gene or other gene regulatory region (e.g., long non-coding RNA, enhancers), which may be defined by chromosomal location or other genetic mapping indices.

[0040] The term "gene product" is used herein to mean a protein or RNA molecule produced by expression of a gene or other gene regulatory region (i.e., transcription, translation, post-translation modification, etc.).

[0041] As used herein, the terms "biological specimen", "patient sample" and "sample" refer to a specimen taken from a patient. Such samples include, but are not limited to, tumors, biopsies, tumor organoids, other tissues, and bodily fluids. Suitable bodily fluids include, for example, blood, serum, plasma, sputum, lavage fluid, cerebrospinal fluid, urine, semen, sweat, tears, saliva, etc. Samples can be collected, for example, via biopsy, swab, or smear.

[0042] The terms "extracted," "recovered," "isolated," and "separated" refer to a compound (e.g., protein, cell, nucleic acid, or amino acid) that has been removed from at least one component with which it is naturally associated and found in nature.

[0043] The term "enriched" or "enrichment" as used herein in reference to nucleic acids refers to a process that enhances the amount of one or more nucleic acid species in a sample. Exemplary enrichment methods can include chemical and / or mechanical means, as well as amplification of the nucleic acids contained in the sample. Enrichment can be sequence-specific or non-specific (i.e., involving any of the nucleic acids present in the sample).

[0044] As used herein, "cancer" should be construed to mean any one or more of a wide range of benign or malignant tumors, including those capable of invasive growth and metastasis through the human or animal body or parts thereof, such as, for example, through the lymphatic system and / or bloodstream. As used herein, the term "tumor" includes both benign and malignant tumors as well as solid growths. Exemplary cancers include, but are not limited to, carcinomas, lymphomas, or sarcomas, such as human ovarian, colon, breast, pancreatic, lung, prostate, urinary tract, uterine, acute lymphocytic leukemia, Hodgkin's disease, small cell lung cancer, melanoma, neuroblastoma, glioma, and soft tissue sarcoma.

[0045] In the context of the present invention, the term "biomarker" should be taken to mean any genetic variant or molecule that indicates or correlates with a feature of interest, such as the presence of or susceptibility to cancer in a subject, the likelihood that the cancer is of one subtype versus another, the likelihood that a patient will respond or not respond to a particular treatment or class of treatments, the degree of positive response that would be expected to a treatment or class of treatments (e.g., survival and / or progression-free survival), whether a patient is responding to a treatment, or the likelihood that the cancer has progressed or will progress beyond its site of origin (i.e., metastasize).

[0046] As used herein, the terms "cellular pathway," "signal transduction pathway," or "pathway" refer to a communication process that governs fundamental activities of a cell and coordinates multiple cellular operations. A pathway includes biochemical reactions between molecules that control cell functions (e.g., cell division, cell death). A cellular pathway includes the entire sequence of molecular events involved in such a process, including, for example, synthesis and release of a signaling molecule by a cell, transport of a signal to a target cell, binding of a signaling molecule to a specific receptor, receptor activation, and initiation of a signaling pathway.

[0047] As used herein, the terms "cellular pathway dysregulation," "signal transduction pathway dysregulation," and "pathway dysregulation" refer to an abnormality or disorder in the regulation of a cellular pathway. Dysregulation (used herein with the term perturbation) can occur at any stage of the gene expression process, including but not limited to during transcription, RNA splicing, RNA export, translation, and post-translational modification of proteins. Regulation of gene expression provides control over the timing, location, and amount of a given gene product (i.e., protein or ncRNA) present in a cell. Thus, cellular pathway dysregulation can involve over- or under-expression of a gene, as well as changes in protein function or stability. In some cases, genetic variations such as mutations, gene fusions, or changes in DNA copy number, methylation status, etc. contribute to cellular dysregulation. Although cancers are heterogeneous with respect to their genetic mutation profiles, many cancers arise and are maintained through aberrant activation or suppression of molecular signaling pathways. For example, the RAS / receptor tyrosine kinase (RTK) and PI3K pathways, when perturbed, can promote the growth of unregulated cells (and tumors), and often impact cancer. In some cases, dysregulated pathways can be targeted by specific chemotherapeutic agents in an attempt to suppress cancer.

[0048] The terms "treatment", "treating" and the like are used herein to generally mean obtaining a desired pharmacological and / or physiological effect. The effect may be prophylactic, in that it completely or partially prevents the disease or its symptoms, and / or it may be therapeutic, in that it partially or completely cures the disease and / or the adverse effects caused by the disease. As used herein, "treatment" encompasses any treatment of a disease in a mammal, including (a) preventing the disease from occurring in a subject who may have a predisposition to the disease, but has not yet been diagnosed as having it, (b) inhibiting the disease, i.e., arresting its development, or (c) relieving the disease, i.e., causing regression of the disease. Therapeutic agents may be administered before, during, or after the onset of the disease or injury. Treatment of ongoing disease is of particular interest when the treatment stabilizes or reduces undesirable clinical symptoms in the patient. The targeted therapy is desirably administered during, and optionally after, the symptomatic phase of the disease.

[0049] The term "effective amount" as used in the methods of the present disclosure refers to an amount of active agent sufficient to exhibit a detectable therapeutic effect without excessive adverse side effects (such as toxicity, irritation, and allergic reactions) commensurate with a reasonable benefit / risk ratio. The effective amount for a patient will depend on the type of patient, the size and health of the patient, the nature and severity of the condition being treated, the method of administration, the duration of treatment, the nature of concomitant therapy (if any), the particular formulation used, and the like. For this reason, an exact effective amount cannot be specified in advance. However, the effective amount for a given situation can be determined by one of ordinary skill in the art using routine experimentation based on knowledge in the art and the information provided herein. The optimal dosing regimen can be determined by one of ordinary skill in the art without undue experimentation.

[0050] As used herein, the terms "reference sequence," "reference assembly," or "reference genome" refer to one or more nucleic acid databases created using DNA sequencing that are assembled as a representative example of the set of genes in one idealized individual organism of a species. A "reference transcriptome" is similarly defined as a database created using RNA sequencing that reflects the set of expressed sequences in one idealized individual organism of a species. Because they are assembled from sequencing DNA from many individual donors, reference genomes do not accurately represent the set of genes of any single individual organism. The most commonly used human reference genomes are derived from 13 anonymous volunteers, thus providing a haploid mosaic of the different DNA sequences from each donor. The most commonly used human reference genomes are GRCh37 and GRCh38 from the Genome Reference Consortium, with updates released every 1-4 years. A common use of reference genomes is to map transcripts obtained from DNAseq and RNAseq. In the case of a reference transcriptome, because transcription is highly dynamic and varies with tissue type, developmental stage, environmental conditions, and disease states, the reference transcriptome does not reflect gene expression at all times, but rather the entire set of possible transcripts in an organism or species. Commonly used reference transcriptomes include RefSeq and Ensembl, which are themselves integrations of multiple independent sequencing projects. Once RNA is sequenced and aligned to a reference genome, such databases are used to assign reads to specific genes. In some embodiments, one or more reference genomes are used to define wild-type and mutant sequences. In embodiments disclosed herein, a single reference genome and / or a single reference transcriptome is used to define wild-type and mutant sequences in conjunction with building the model. However, embodiments are envisioned in which multiple reference genomes or multiple reference transcriptomes, or updated reference databases are used.

[0051] FIG. 1A shows examples of cellular pathways. (See Sanchez-Vega et.al., 2018, Cell. 173:321-337). This example illustrates pathways curated by The Cancer Genome Atlas (TCGA), including RTK / RAS, Nrf2, TGFβ, PI3K, p53, Wnt, Myc, cell cycle, Hippo, and Notch pathways. Each pathway is surrounded by a box, and elements of each pathway are shown as labeled rectangles within the box. Various interactions (including activation, inhibition, etc.) between pathway elements are indicated by arrows or lines.

[0052] FIG. 1B shows a custom pathway. In the example shown, the custom pathway is a color-coded subset of the PI3K pathway gene list and the RAS pathway gene list. The color codes indicate different functional components of the pathway, meaning that a mutation in any gene in a color group can be predicted to have the same effect on pathway function as a mutation in another gene in the same color group. In this example, the first group is the left column containing PI3KR (PI3KR1 / PI3KR2), the second group is the middle column containing ERBB2, PI3K (PIK3CA / PIK3CB), AKT (AKT1 / AKT2 / AKT3), and MTOR, and the third group is the right column containing EGFR, RAS (KRAS / NRAS / HRAS), RAF (RAF1 / BRAF / ARAF), MEK (MAP2K1 / MAP2K2), and ERK (MAPK3 / MAPK1). In the illustrated example, a "T" shaped line from PTEN to PI3K indicates that PTEN inhibits PI3K, and an arrow indicates activation (eg, EGFR activates both RAS and PI3K).

[0053] Some of the pathways that cause cancer are well characterized, and many examples of perturbations can be traced back to mutations in a handful of "driver" genes, such as KRAS in the RAS / RTK pathway and STK11 in the PI3K pathway. However, there are many cases where driver gene mutations are not present, but one or more pathways nevertheless show signs of perturbation at the transcriptional and / or protein levels. In such cases, DNA analysis alone (including single nucleotide variants, insertions / deletions [in-del], and copy number variants) cannot identify pathway perturbations, resulting in missed opportunities to use therapeutics that target the pathway. Measures of pathway perturbations that are not limited to analysis of DNA may allow for the identification of additional patients who may respond to these therapies.

[0054] (Use of the system / method)

[0055] FIG. 2A is a schematic diagram illustrating an exemplary concept of the systems and methods disclosed herein.

[0056] In one example, the system and method analyzes RNA data to determine the pathway perturbation status of a cancer specimen for at least one cellular pathway. In FIG. 2A, the cellular pathways analyzed for the specimen are the RAS, PI3K, WNT, SHH, and NOTCH pathways. Each pathway has activation range bars of various colors and a black bar that indicates the activity level of the pathway. Black bars located further to the left of the blue or purple region indicate pathways that are not perturbed. Black bars located more centrally in the green region indicate pathways with moderate perturbation. Black bars located further to the right of the red region indicate pathways that are highly perturbed. In this example, the RAS pathway is highly perturbed, the PI3K, WNT, and SHH pathways are not perturbed, and the NOTCH pathway is moderately perturbed.

[0057] The three blue arrows pointing from the pathway perturbation bar to the right portion of Figure 2A indicate downstream uses of the results of the pathway perturbation analysis. At the top, the results of the pathway perturbation analysis can be used to help determine whether a genetic variant or mutation (especially a variant of unknown significance) qualifies as a pathogenic variant, which is a variant that causes cancer, or is more likely to be a benign variant, which is a variant that has little or no effect on the disease. In the middle, the results can determine a therapy that matches the patient or organoid from which the cancer specimen was obtained. For example, if a pathway is perturbed, a therapy that targets the pathway (e.g., by targeting proteins and / or genes in the pathway) can be tailored. The pie chart at the bottom is an example of the portion of cancer cases associated with variants in a given gene, organized by gene name. In this example, approximately 24% of cancer specimens that may have a dysregulated pathway do not have any detected canonical driver mutations in genes associated with that pathway.

[0058] In some embodiments, the systems and methods analyze RNA rather than or in addition to DNA mutation data to evaluate potential pathway perturbations. In some cases, the mutational cause of the pathway perturbation is unknown (e.g., the mechanism of RAS pathway perturbation is unknown in as many as 24% of lung adenocarcinoma cases). However, pathway perturbations may have an RNA signature that is captured by the systems and methods disclosed herein, regardless of the presence of DNA evidence.

[0059] As a result, DNA evidence may suggest pathway perturbations when in fact they are not present. The systems and methods disclosed herein have a more robust ability to correctly classify these potential false positives.

[0060] In various embodiments, the system and method characterize genomic alterations and molecular features into summarized known pathway profiles and associate their relationships with treatment response data from patients, cell lines, and / or tumor organoids.In various embodiments, instead of characterizing a patient's tumor by genomic alterations and RNA expression levels detected at the single gene level, the system and method integrate multiple molecular and genomic profiles into cancer signaling pathways to reveal insights into their relationships with treatment response and disease outcome.

[0061] In various embodiments, the systems and methods also analyze data from the entire gene set (approximately 18,000 genes or more) compared to a smaller subset of genes, making the systems and methods much more flexible than out-of-the-box methods such as single sample gene set enrichment analysis (ssGSEA, see Barbie, et al., 2010, Nature. 462(7269):108-112) in that it allows the ability to explore potential sources of pathway perturbation outside of standard pathway genes and curated gene lists.

[0062] In some embodiments, the system and method utilize the transcriptome together with clinical and DNA variant data or methylation status to detect targetable pathway perturbing events that may not be detected by individual gene expression levels (e.g., a list of genes that are over- or under-expressed in cancer specimens compared to non-cancer specimens), or DNA variants that are currently detected and / or reported to physicians and patients as pathogenic variants. The transcriptome may be captured by whole-exome RNA-seq and is not limited to the expression levels of genes associated with the pathway. This is particularly relevant when dysregulation is caused by genes downstream of the pathway or genes not known to be associated with the pathway. The clinical data may relate to treatments that the patient or organoids have received and the response of the patient or organoids to those treatments (e.g., if the growth rate of cancer cells in the patient or organoids slows after exposure to the treatment). The methylation status may relate to the methylation of genes and / or promoters associated with the pathway.

[0063] In some embodiments, the systems and methods disclosed herein avoid the limitations of DNA analysis in detecting pathway dysregulation. The systems and methods may include orthogonal transcriptome approaches to identify pathway perturbations in cancer patients. The systems and methods may include highly sensitive transcriptome models of oncogenic signaling pathway perturbations that pass several validation tests and identify patients who may respond to targeted therapeutics despite the absence of standard pathway mutations. In certain embodiments, the systems and methods may include machine learning approaches to identify hidden responders who may respond to treatment but whose responder status cannot be detected by standard DNA-based diagnostics.

[0064] In certain embodiments, the systems and methods include identification of pathway perturbations in human cancers by transcriptomics.

[0065] In some embodiments, the systems and methods generate pathway perturbation scores based solely on transcriptome data, providing an orthogonal indicator of pathway perturbation that does not rely on a DNA-based understanding of the underlying mechanism of the perturbation. With sufficient sample size, the same systems and methods can be used to generate models of pathway perturbations for any pathway and any cancer type.

[0066] FIG. 2B is a schematic diagram illustrating another exemplary concept of the systems and methods disclosed herein.

[0067] In some embodiments, the system and method include one or more pathway perturbation models and the results generated by those pathway perturbation models.The training data of the pathway perturbation model includes transcriptome data, and can further include genomic data.The training data and / or biological validation data for determining how the model results reflect the biological state can further include structured clinical data or organoid data, including any evidence of treatments that slow the growth of cancer in patients or tumor organoids, and information from a treatment decision engine, including a list of treatments that target any gene or gene product in a gene set or pathway of interest.

[0068] In one example, the pathway perturbation models include a RAS pathway perturbation model and a PI3K pathway perturbation model, which were developed using transcriptomic and genomic data from lung adenocarcinoma patients, respectively, and were extensively validated in both public and private datasets (second column from the left). In this example, the RAS model assigns similarly strong perturbation scores to patients with mutations in KRAS and BRAF, two neighboring molecules in the RAS pathway. Similarly, strong results were achieved for the PI3K perturbation model (second column from the right). These results demonstrate that the perturbation scores generated by these models can quantitatively estimate the impact of genetic mutations on biological pathways.

[0069] In this example, both models identify candidate target genes or mutations that have unexpected effects on pathway perturbation. For example, the systems and methods disclosed herein can analyze transcriptomes from several specimens that do not have mutations known to cause perturbations in a given pathway and predict that the pathway is perturbed in each of these specimens. The specimens can then be analyzed to determine whether they have a common mutation or mutant gene, even if it is not a mutation or gene known to perturb the pathway, and the common mutation or gene can be identified as the target mutation or gene. This analysis can prioritize genes that produce proteins known to interact with members of the pathway. These protein-protein interactions can be listed in a pathway database 300 (see FIG. 3A).

[0070] The model nevertheless shows that many patients without pathway mutations (pathway normal or wild type) have high perturbation scores (red, blue, and purple dots). These "hidden responders" may benefit from therapies commonly used to target these pathways, and these model results provide further opportunities for biomarker and target discovery. Patients with samples that have variants in these target genes can be matched with one of these treatments.

[0071] In one example, to verify the clinical relevance of the model results, data from patient clinical records or tumor organoid growth experiments can be analyzed for associations between treatment response and the target gene(s) or variants identified by the pathway model. If there is evidence that a treatment can slow the growth of cancer cells in a patient or tumor organoid, and the patient and organoid cancer cells have variants in the target gene(s), the treatment decision engine can be updated with entries for the treatment and the pathway that the target gene(s) alters. In the absence of organoid treatment response data for the identified target gene, organoids can be genetically engineered to have the identified target gene or mutation, and their growth rate can be observed after exposure to pathway-targeted therapy.

[0072] In some embodiments, the cancer patient has lung adenocarcinoma (LUAD). In some embodiments, the cancer patient has breast cancer, colon cancer, or prostate cancer. In some embodiments, the cancer patient has any cancer type. In some embodiments, the system and method refines clinically relevant pathways of interest by characterizing gene expression data, DNA mutation profiles, and immune profiles of PI3K and RTK / RAS pathways across cancer types and testing predictions against clinical response and outcome data. The system and method can extend this approach to other networks / pathways prioritized based on relevance to therapeutic targeting. In some embodiments, the system and method can include algorithm validation and retrospective analysis.

[0073] In some embodiments, the systems and methods disclosed herein include a binomial logistic regression model that uses normalized transcriptome data from a database and pathway scores generated with the same transcriptome data in combination with an algorithm and a molecular pathway gene set. In one example, the molecular pathway gene set is curated. The output of the model can be a single number that indicates the degree to which the sample's transcriptome is consistent with a pathway perturbation.

[0074] In some embodiments, the systems and methods discover integrated multi-omic pathway signatures that predict treatment response and disease outcomes. These multi-omic pathway signatures can include characteristics of patient and / or specimen related data (e.g., data types including clinical, response outcomes, DNA mutations, RNA gene expression, etc.). Machine learning models can be used to analyze these data types, etc., in the context of disease-related gene and protein networks / pathways. Response outcome data can include information regarding survival and progression-free survival of patients or organoids after exposure to various treatments, including over 100 different cancer drugs.

[0075] In various embodiments, the systems and methods can be used to discover novel correlation pathways / networks in DNA alterations, fusions, and RNA-seq gene expression data, as well as molecular patterns associated with treatment response through imaging (including histopathology and radiology images).

[0076] To identify correlated de novo patterns from molecular profiling results, the system and method can include integrative (mutual information, Bayesian networks, neural networks, and other statistical and machine learning methods) to define correlated gene and protein networks associated with disease. Novel disease-associated networks can be tested for association with treatment and outcome data, including data derived from clinical records. Statistically significant associations can be validated using focused data sets to test sensitivity and recall associations with tumor treatment response or patient survival metrics.

[0077] In various embodiments, the systems and methods disclosed herein include artificial intelligence models of pathway perturbations. The systems and methods may be used for biomarker discovery, which may include in silico evaluation of genes and / or variants identified by the model(s) to predict pathway perturbations and the effect of genes and / or variants on cancer.

[0078] The systems and methods may include annotation of new and / or known biomarkers (e.g., genes and / or variants), particularly the likely status of each biomarker as a viable drug target, and may include the use of private and / or public databases. For example, the databases may include descriptions of observed drug interactions with biomarkers, associations between patient responses to drugs and biomarkers observed in patients, and / or protein structures and effects of biomarkers on the protein structure of gene products. These databases may include information for identifying drug targets and prioritizing associations between diseases and drug targets; associations between human diseases and genes, variants, drugs and / or drug targets; information on drugs and their targets (including interactions between drugs and drug targets); interactions between genes and drugs (including the status of genes as drug targets); information on therapeutic protein and nucleic acid targets and associated target diseases (e.g., types of cancer); information on drugs, drug targets and molecules; information on parts of the genome that are druggable (e.g., those that can be targeted by drugs); and associations between chemicals, gene products, phenotypes, diseases and environmental exposures. A drug target can be a gene or protein that is affected by a drug (e.g., a drug can alter, inhibit, or activate the activity or function of a drug target). These databases can contain information based on published research studies.Examples of public databases include DrugBank (see drugbank.ca), ChEMBL (see ebi.ac.uk / chembl), DGIdb (dgidb.org), TTD (see db.idrblab.org / ttd / ), DisGeNET (see disgenet.org), DTC (see drugtargetcommons.fim.fi), Open Targets (see opentargets.org), PHAROS (see pharos.nih.gov), CTD (http: / / ctdbase.org / ), ADReCS-Target (see bioinf.xmu.edu.cn), etc. (for further description of these databases, see Paananen and Fortino, Briefings in Bioinformatics (2019); doi:10.1093 / bib / bbz122), see also Figures 26A-Z and 27A-V.

[0079] The system and method may include in vitro validation of candidate target biomarkers in organoids by genetic engineering and / or drug screening.For example, genetic engineering (e.g., using CRISPR and / or other gene editing tools) can be used to design organoids with candidate biomarkers, and drug screening can be used to determine which treatments can slow the growth of organoids with candidate biomarkers.

[0080] The systems and methods disclosed herein may be used to guide a treatment of a subject. By way of example, a subject sample may be analyzed according to the systems and methods disclosed herein, and a recommended treatment / treatment regimen may be provided by the system. In some embodiments, the method includes treating the subject according to the recommended treatment / treatment regimen. In some embodiments, the recommended treatment includes administering to the subject an effective amount of one or more of the compounds listed in Figures 26A-27P or 27Q-27V.

[0081] Oncogenic signaling pathways are composed of multiple proteins, and it is often useful to subdivide the pathways into modules based on the similarity of the proteins with respect to their sequence or function, their clinical targetability, and the effects of their perturbation. For example, the RAS module of the RTK / RAS parent pathway is composed of KRAS, NRAS, and HRAS. Mutations in these genes are present in different proportions in different cancers, with KRAS mutations being most common in lung adenocarcinoma, NRAS being most common in melanoma, and HRAS being most common in melanoma. However, they have very similar sequences and are characterized by mutations in the same domains that cause uncontrolled growth and result in the activation of the same downstream clinically targetable effectors when perturbed. For the purpose of modeling RTK / RAS pathway perturbations, the grouping of these proteins into modules is logical from a biological and clinical perspective and adds strength to the model generator by allowing a combination of patients with mutations in these genes to form a positive control group.

[0082] Another rationale for grouping into modules can be based solely on the functional effect of proteins such as the PTEN module in the PI3K pathway, which consists of PTEN, PIK3R1 and PIK3R2. Although not structurally similar, each of these proteins is involved in the inhibition of PI3K signaling, potentially providing guidance for treatment. For example, if a disturbance is detected in this module, clinicians can consider treatment with a PI3K inhibitor to block the effect of the dysfunctional inhibitory PTEN module.

[0083] Figures 12A-12E show several such modules for the RTK / RAS and PI3K pathways, each of which was constructed with the above factors in mind. Other oncogenic signaling pathways will have different associated modules. It is also important to note that additional findings regarding the specific target of the pathway under consideration, new treatment recommendations, and / or perturbation models may require redesigning the modules. Thus, the modules shown for the RTK / RAS and PI3K pathways are not intended to, and do not, exemplify the entirety of possible modules that may be used in this method.

[0084] (System and method)

[0085] 3A shows a schematic diagram of a system 10 capable of determining the pathway perturbation status of at least one tissue specimen. The system 10 may include one or more data inputs 100, one or more pathway engines 200, a pathway database 300, a labeled tumor sample database 400, a drug-pathway interaction database 500, a treatment response database 600, a clinical trial database 700, and a patient report generator 800.

[0086] The pathway engine 200 can communicate with a pathway database 300, a labeled tumor sample database 400, a drug-pathway interaction database 500, a treatment response database 600, a clinical trial database 700, and a patient report generator 800 via a communication network 20. One or more pathway engines 200 can receive data inputs 100 and output one or more pathway perturbation scores. The pathway engine 200 can be stored on one or more devices, which are described in more detail below.

[0087] The data input 100 may include a transcriptome value set and one or more dysregulation indicators (described in FIG. 4). The data input 100 may further include DNA variant data, methylation data, cancer type, and / or proteomic data.

[0088] Each of the one or more pathway engines 200 may be trained with a set of data from the data input 100 to determine the likelihood that a pathway associated with a tissue sample has a perturbed state. The system 10 may include 1, 10, 100, or more pathway engines 200. As used herein, the label "200n" is intended to refer to one general-purpose pathway engine of the one or more pathway engines 200.

[0089] In various embodiments, the pathway engine 200n predicts pathway perturbation states based on the RNA data. In various embodiments, the pathway engine 200n includes a predictive model. In various embodiments, the pathway engine 200n includes a support vector machine, a random forest, and / or a k-nearest neighbor model. In some embodiments, the pathway engine 200n includes a logistic regression model.

[0090] In some embodiments, each pathway engine 200n can predict pathway perturbations for specimens having a particular cancer type. In various embodiments, each pathway engine 200n can predict pathway perturbations for a single pathway of interest, a combination of pathways of interest, or several individual pathways of interest.

[0091] In various embodiments, each pathway engine 200n can predict pathway perturbations for a single pathway of interest. The pathway of interest can be a cellular pathway included in the pathway database 300. The pathway of interest can be a TCGA-defined pathway or a custom gene set or gene list. For example, the pathway of interest can include the RAS / RTK, PI3K, and / or WNT pathways. In some embodiments, the pathway includes an oncogenic network / pathway with a known regulatory response to targeted therapy.

[0092] In one example, the pathway engine 200n can predict pathway perturbations of the RTK-RAS / PI3K pathway (see, e.g., FIG. 1B) in patients and / or specimens with lung adenocarcinoma. In one example, the pathway engine 200n can predict pathway perturbations of the WNT pathway in patients and / or specimens with colorectal cancer. In one example, the pathway engine 200n can predict pathway perturbations of the PI3K pathway in patients and / or specimens with breast cancer. In one example, the pathway engine 200n can predict pathway perturbations of the vascular endothelial growth factor (VEGF) pathway.

[0093] In some embodiments, one or more pathways of interest can be examined for each analyte. For example, it may be useful to score multiple pathway dysregulation and / or overall dysregulation of multiple interacting pathways to determine whether treatment may be effective for a patient whose analyte has dysregulation in one or more pathways, particularly if at least one pathway is activated and at least one pathway is inhibited. This may include using multiple trained pathway engines 200a, 200b, ..., 200n to analyze input data associated with each analyte.

[0094] The pathway database 300 may include gene or protein networks, e.g., descriptions and / or lists of sets of genes and / or proteins that interact during the activity of a biological cell. Gene-gene, protein-protein, and gene-protein interactions may involve one gene or protein inhibiting, activating, or altering the activity, expression level, or state of another gene or protein.

[0095] In some embodiments, the pathway is a gene list defined by MSigDB (GSEA) or a TCGA pathway curated list. In some embodiments, the pathway of interest is a custom gene list. A pathway gene list of interest can be selected in collaboration with a team of pathologists or other experts.

[0096] The labeled tumor sample database 400 may include data relating to biological specimens having known pathway perturbation status (e.g., perturbed or unperturbed) for each of one or more pathways. The pathway perturbation status may be based on DNA variants detected in the specimen and located in genes associated with the pathway. The data input 100 may be stored in the labeled tumor sample database 400.

[0097] The drug-pathway interaction database 500 may include data entries that indicate associations between treatments and the genes, gene products, and / or pathways that the treatments target.

[0098] Entries in the treatment response database 600 may include observed instances of a treatment slowing the growth of cancer in a sample from a patient or tumor organoid, and various characteristics of the sample, including an associated list of genetic variants and / or perturbed pathways detected in the sample.

[0099] The clinical trials database 700 may include a list of clinical trials and information about each clinical trial. The clinical trial information may include the trial name, exclusion and / or inclusion criteria, enrollment information, contact information, site name, location, intervention (e.g., treatment, drug, procedure), clinical trial dates (e.g., start and completion dates), and other information (e.g., any information that may be listed on the clinicaltrials.gov website).

[0100] The patient report generator 800 can receive data from the pathway engine 200, the drug-pathway interaction database 500, the treatment response database 600, and the clinical trial database 700. The patient report generator 800 can generate a report to present the pathway perturbation status determined by the pathway engine 200n for a analyte and / or multiple analytes to a patient, the patient's physician, a medical professional, a researcher, etc.

[0101] The patient report generator 800 may include and / or cause to be executed one or more processes for generating a route disturbance score and / or a route disturbance report. In particular, the patient report generator 800 may include and / or cause to be executed processes 502, 602, 630, 650, 660, 670, 750, 702. The processes 502, 602, 630, 650, 660, 670, 750, 702 are described below.

[0102] The patient data store (e.g., labeled tumor sample database 400) can include one or more feature modules that can include a collection of features available to all patients (or tumor organoids) in the system. These features (e.g., data input 100) can be used to generate an artificial intelligence classifier (e.g., pathway engine 200n) in the system. Although the feature range across all patients is informative, a patient's feature set may be sparsely spread across the collective feature range of all features across all patients. For example, the feature range across all patients may extend to tens of thousands of features, but a patient's unique feature set may only include a subset of the collective feature range of hundreds or thousands based on the records available for that patient.

[0103] A feature set (e.g., data entry 100) can include a diverse set of fields available within a patient's health record. Clinical information can be based on fields entered into an electronic medical record (EMR) or electronic health record (EHR) by a doctor, nurse, or other medical professional or representative. Other clinical information may be managed from other sources, such as molecular fields from gene sequencing reports. Sequencing can include next generation sequencing (NGS) and can be long read, short read, or other forms of sequencing the patient's somatic and / or normal genome. A comprehensive set of features in an additional feature module can combine together various features across various areas of medicine, which can include diagnosis, response to treatment regimens, genetic profile, clinical and phenotypic characteristics, and / or other medical, geographic, demographic, clinical, molecular, or genetic features. For example, a subset of features can include molecular data features, such as features derived from RNA feature module or DNA feature module sequencing.

[0104] Another subset of features, imaging features from the imaging feature module, can include features identified through review of specimens, e.g., pathologist review, such as review of stained H&E or IHC slides. As another example, the subset of features can include derived features obtained from analysis of individual and combined results of such feature sets. Features derived from DNA and RNA sequencing can include genetic variants from the variant science module present in the sequenced tissue. Further analysis of genetic variants can include additional steps such as identifying single or multiple nucleotide polymorphisms, identifying whether the mutation is an insertion or deletion event, identifying loss or gain of function, identifying fusions, calculating copy number variation, calculating microsatellite instability, calculating tumor mutation burden (TMB), or calculating other structural variations in DNA and RNA. Analysis of slides for H&E or IHC staining can reveal features such as tumor infiltration, programmed death ligand 1 (PD-L1) status, human leukocyte antigen (HLA) status, or other immunological features.

[0105] Features derived from structured, curated, or electronic medical or health records may include diagnosis, symptoms, treatment, outcome, patient demographics, e.g., patient's name, date of birth, sex, ethnicity, date of death, address, smoking status, date of cancer diagnosis, disease, illness, diabetes, depression, other physical or mental illness, personal medical history, family medical history, clinical diagnosis such as date of initial diagnosis, date of metastasis diagnosis, cancer staging, tumor characterization, tissue of origin, treatment and outcome, e.g., line of treatment, treatment group, clinical trial, medications prescribed or taken, surgery, radiation therapy, imaging, adverse effects, associated outcomes, genetic tests, and laboratory information such as performance scores, clinical tests, pathology results, prognostic indicators, date of genetic testing, testing provider used, testing method used, e.g., gene sequencing method or gene panel, genetic results, e.g., genes included, variants, expression levels / status, or dates corresponding to any of the above.

[0106] Features may be derived from information from additional medical or research-based omics fields, including proteomics, transcriptomics, epigenomics, metabolomics, microbiology, and other multi-omic fields. Features derived from organoid modeling labs may include DNA and RNA sequencing information germane to each organoid, as well as results from treatments applied to those organoids. Features derived from imaging data may further include reports related to stained slides, tumor size, tumor size difference over time, including treatment during change, and machine learning techniques to classify PDL1 status, HLA status, or other characteristics from imaging data. Other features may include additional derived feature sets from other machine learning techniques based at least in part on any new features and / or combinations of features listed above. For example, imaging results may need to be combined with MSI calculations derived from RNA expression to determine further imaging features. In another example, the machine learning model may generate the likelihood that the patient's cancer has metastasized to a particular organ or any other organ. Other features that may be extracted from medical information may also be used. There are thousands of features, and the above list of feature types is merely representative and should not be construed as an exhaustive list of features.

[0107] The modification module may be one or more microservices, servers, scripts, or other executable algorithms that generate modified features related to de-identified patient features from the feature set. The modification module can take input from the feature set and provide modifications for storage. Exemplary modification modules may include one or more of the following modifications as part of the modification module set:

[0108] The IHC (Immunohistochemistry) module can identify antigens (proteins) in cells of tissue sections by utilizing the principle of antibodies specifically binding to antigens in biological tissues. IHC staining is widely used in the diagnosis of abnormal cells such as those found in cancerous tumors. Certain molecular markers are characteristic of certain cellular events such as proliferation or cell death (apoptosis). IHC is also widely used in basic research to understand the distribution and localization of biomarkers and differentially expressed proteins in different parts of biological tissues. Visualization of antibody-antigen interactions can be achieved in a number of ways. In the most common example, antibodies are conjugated to enzymes such as peroxidase that can catalyze a color reaction in immunoperoxidase staining. Alternatively, antibodies can be tagged to fluorophores such as fluorescein or rhodamine in immunofluorescence. Approximations from RNA expression data, H&E slide imaging data, or other data can be generated.

[0109] The treatment module can identify differences in cancer cells that help them grow and proliferate (or other cells nearby), as well as drugs that "target" these differences (see, e.g., Figures 26A-27P or Figures 27Q-27V for exemplary drugs and their targets). Treatment with these drugs is called targeted therapy. For example, many targeted drugs are lethal to cancer cells due to their internal "programming" that makes them different from normal healthy cells, without affecting most healthy cells. Targeted drugs can block or turn off chemical signals that tell cancer cells to grow and divide rapidly, change proteins in cancer cells so that they die, stop creating new blood vessels to supply the cancer cells, trigger the patient's immune system to kill the cancer cells, or deliver toxins to the cancer cells and kill them without affecting normal cells. Some targeted drugs are more "targeted" than others. Some may target only a single change in the cancer cells, while others may affect several different changes. Others promote the way the patient's body fights the cancer cells. This can affect where these drugs act and the side effects they cause. Matching a targeted therapy can include identifying the patient's therapeutic target and meeting any other inclusion or exclusion criteria that may identify patients who may benefit from the treatment.

[0110] The study module can identify and test hypotheses for treating cancers with specific characteristics by matching patient characteristics to clinical trials. These studies must be matched to enroll patients and have inclusion and exclusion criteria that can be taken from and constructed from publications, study reports, or other documents.

[0111] The amplification module can identify genes whose counts are disproportionately increased relative to other genes (e.g., the number of gene products present in the specimen). Amplification can cause genes with increased counts to become dormant, hyperactive, or behave in another unexpected manner. Amplification can be detected at the gene level, variant level, RNA transcript or expression level, or protein level. Detection can be performed across all different detection mechanisms or levels and validated against each other.

[0112] Isoform modules can identify alternative splicing (AS), a biological process in which two or more mRNA types (isoforms) are generated from the same gene transcript through different combinations of exons and introns. Large-scale genomics studies estimate that 30-60% of mammalian genes are alternatively spliced. The possible patterns of alternative splicing for a gene can be highly complex, and the complexity increases rapidly as the number of introns in a gene increases. In silico alternative splicing prediction can find large insertions or deletions within a set of mRNAs that share a large proportion of aligned sequences by identifying genomic loci by searching mRNA sequences against genomic sequences, extracting sequences of genomic loci, extending sequences on both ends up to 20 kb, searching the genomic sequence (repetitive sequences are masked), extracting splicing pairs (three or more expressed sequence tags aligned on two boundaries of an alignment gap with the GT-AG consensus or on both ends of a gap), assembling the splicing pairs according to their coordinates, determining gene boundaries (a splicing pair prediction is generated up to this point), generating predicted gene structures by aligning the mRNA sequences to the genomic template, and comparing the splicing pair predictions with the gene structure predictions to find alternatively spliced ​​isoforms.

[0113] SNP (single nucleotide polymorphism) modules can identify single base substitutions that occur at specific positions in the genome, with each mutation present to some discernible extent in the population (e.g., greater than 1%). For example, at a particular base position or locus in the human genome, a C nucleotide may occur in most individuals, while in a minority of individuals, the position is occupied by an A. This means that an SNP exists at this particular position, and the two possible nucleotide variations, C or A, are said to be alleles at this position. SNPs underlie differences in human susceptibility to a wide range of diseases (e.g., sickle cell anemia, β-thalassemia, and cystic fibrosis are caused by SNPs). The severity of the disease and the way the body responds to treatment are also signs of genetic variation. For example, a single nucleotide mutation in the APOE (apolipoprotein E) gene is associated with a lower risk of Alzheimer's disease. Single nucleotide variants (SNVs) are single nucleotide variations with no frequency restriction that can occur in somatic cells. Somatic single nucleotide mutations (e.g., caused by cancer) may also be referred to as single nucleotide changes. MNP (multiple nucleotide polymorphism) modules can identify substitutions of consecutive nucleotides at specific positions in the genome.

[0114] The Indels module can identify insertions or deletions of bases in the genome of classified organisms during small genetic variations. Indels usually measure 1-10,000 base pairs in length, while microindels are defined as indels that result in a net change of 1-50 nucleotides. Indels can be contrasted with SNPs or point mutations. Indels insert and / or delete nucleotides from a sequence, while point mutations are a form of substitution that replaces one of the nucleotides without changing the overall number in the DNA. Indels, which are insertions and / or deletions, can be used as genetic markers in natural populations, especially in phylogenetic studies. Indel frequencies tend to be significantly lower than those of single nucleotide polymorphisms (SNPs), except near highly repetitive regions including homopolymers and microsatellites.

[0115] The MSI (microsatellite instability) module may identify genetic hypermutability (predisposition to mutations) resulting from DNA mismatch repair defects (MMR). The presence of MSI represents phenotypic evidence that MMR is not functioning normally. MMR corrects errors that arise spontaneously during DNA replication, such as single-base mismatches or short insertions and deletions. Proteins involved in MMR correct polymerase errors by forming a complex that binds to the mismatched portion of DNA, excising the error and inserting the correct sequence in its place. Cells with abnormally functioning MMR are unable to correct errors that arise during DNA replication, which causes the cell to accumulate errors in its DNA. This causes the generation of novel microsatellite fragments. Polymerase chain reaction-based assays can reveal these novel microsatellites and provide evidence for the presence of MSI. Microsatellites are repetitive sequences of DNA. These sequences can be created in repeating units of 1-6 base pairs in length. The length of these microsatellites is highly variable from person to person, contributing to an individual's DNA "fingerprint," but each individual has a set length of microsatellites. The most common microsatellites in humans are dinucleotide repeats of the nucleotides C and A, which occur tens of thousands of times throughout the genome. Microsatellites are also known as simple sequence repeats (SSRs).

[0116] The TMB (Tumor Mutation Burden) module is a predictive biomarker that can identify a measure of mutations carried by tumor cells and is being studied to assess its association with response to Immuno-Oncology (IO) therapy. Tumor cells with high TMB may harbor more neoantigens, with an associated increase in cancer-fighting T cells in the tumor microenvironment and periphery. These neoantigens can be recognized by T cells and induce anti-tumor responses. TMB has more recently emerged as a quantitative marker that can help predict potential response to immunotherapy across a range of cancers, including melanoma, lung and bladder cancer. TMB is defined as the total number of mutations per coding region of the tumor genome. Importantly, TMB is consistently reproducible. It provides a quantitative measure that can be used to better inform treatment decisions, such as selection of targeted or immunotherapy or enrollment in clinical trials.

[0117] The CNV (Copy Number Variation) module can identify deviations from the normal genome, particularly copy numbers of genes, parts of genes, or other parts of the genome not defined by genes, and any subsequent implications from the analysis of genes, variants, alleles, or sequences of nucleotides. CNV is a phenomenon in which structural variations can occur in portions of nucleotides or base pairs, including repeats, deletions, or inversions.

[0118] Fusion modules may identify hybrid genes formed from two previously separate genes. This may occur as a result of translocations, stromal deletions, or chromosomal inversions. Gene fusions can play an important role in tumorigenesis. Fusion genes may contribute to tumorigenesis because fusion genes may produce abnormal proteins that are much more active than non-fusion genes. In many cases, fusion genes are cancer-causing oncogenes, including BCR-ABL, TEL-AML1 (ALL with t(12;21)), AML1-ETO (M2 AML with t(8;21)), and TMPRSS2-ERG with stromal deletions on chromosome 21 that often occur in prostate cancer. In the case of TMPRSS2-ERG, the fusion product regulates prostate cancer by disrupting androgen receptor (AR) signaling and inhibiting AR expression by oncogenic ETS transcription factors. Most fusion genes are found in hematological cancers, sarcomas, and prostate cancer. BCAM-AKT2 is a fusion gene specific and unique to high-grade serous ovarian cancer. Oncogenic fusion genes may result in gene products with new or different functions from the two fusion partners. Alternatively, a proto-oncogene is fused to a strong promoter, whereby the oncogenic function is set to function by upregulation caused by the strong promoter of the upstream fusion partner. The latter is common in lymphomas, where an oncogene is juxtaposed to the promoter of an immunoglobulin gene. Oncogenic fusion transcripts may also be caused by trans-splicing or read-through events. Because chromosomal translocations play such an important role in neoplasia, a specialized database of chromosomal aberrations and gene fusions in cancer has been created. This database is called the Mitelman Database of Chromosome Aberrations and Gene Fusions in Cancer.

[0119] The VUS (Variant of Unknown Significance) module can identify variants that are detected in a patient's genome (especially in a patient's cancer specimen) but cannot be classified as pathogenic or benign upon detection. VUS can be cataloged from publications to identify whether they can be classified as benign or pathogenic.

[0120] The DNA repair pathway module (e.g., Pathway Engine 200n) may identify defects in DNA repair pathways that allow cancer cells to accumulate genomic alterations that contribute to their aggressive phenotype. Cancerous tumors rely on residual DNA repair capacity to overcome damage induced by genotoxic stress, which inactivates isolated DNA repair pathways in cancer cells. DNA repair pathways are generally considered mutually exclusive mechanistic units that deal with different types of damage at different cell cycle phases. However, recent preclinical studies provide strong evidence that multifunctional DNA repair hubs involved in multiple conventional DNA repair pathways are frequently altered in cancer. Identifying potentially affected pathways may lead to important patient treatment considerations.

[0121] The raw count module can identify the count of variants detected from the sequencing data. For DNA, this can be the number of reads from sequencing that correspond to a particular variant in a gene. For RNA, this can be gene expression counts or transcriptome counts from sequencing.

[0122] Structural variant classification can include evaluating features from the feature set, modifications from the modification module, and other classifications from among itself from one or more classification modules. The structural variant classification can provide a classification in a stored classification storage device. An exemplary classification module can include classification of the CNV as "reportable", which can mean that the CNV has been identified in one or more reference databases as affecting tumor cancer characteristics, disease states, or pharmacogenomics, "non-reportable" can mean that the CNV has not been so identified, and "conflict of evidence" can mean that the CNV has evidence suggesting both "reportable" and "non-reportable". Furthermore, classification of therapeutic relevance is similarly confirmed from any reference dataset that mentions treatments that may be affected by detection (or non-detection) of the CNV. Other classifications can include application of machine learning algorithms, neural networks, regression techniques, graphing techniques, inductive reasoning techniques, or other artificial intelligence evaluations within the module. A classifier for clinical trials may include evaluating variants identified from alteration modules identified as significant or reportable, evaluating all available clinical trials to identify inclusion and exclusion criteria, mapping patient variants and other information to the inclusion and exclusion criteria, and classifying clinical trials as applicable or inapplicable to the patient. Similar classifications may be performed for therapeutic, loss of function, gain of function, diagnostic, microsatellite instability, tumor mutational burden, indels, SNPs, MNPs, fusions, and other alterations that may be classified based on the results of alteration modules.

[0123] Each of the feature sets, the modification modules, the structural variants, and the feature store may be communicatively coupled to a data bus to transfer data between each module for processing and / or storage. In some embodiments, each of the feature sets, the modification modules, the structural variants, and the feature store may be communicatively coupled to one another for independent communication without sharing a data bus.

[0124] In addition to the above features and listed modules, the feature modules may further include one or more of the following modules within their respective modules as sub-modules or stand-alone modules:

[0125] The germline / somatic DNA feature module may include a set of features related to DNA origin information of the patient or the patient's tumor. These features may include raw sequencing results, such as those stored in FASTQ, BAM, VCF, or other sequencing file types known in the art; genes; mutations; variant calls; and variant characterizations. Genomic information from the patient's normal sample may be stored as germline, and genomic information from the patient's tumor sample may be stored as somatic.

[0126] The RNA feature module may include a set of features related to RNA-derived information of a patient, such as transcriptome information. These features may include raw sequencing results, transcriptome expression, genes, mutations, variant calls, and variant characterizations.

[0127] The metadata module may include feature collections related to the human genome, protein structures and their effects, such as changes in energy stability based on protein structure.

[0128] The clinical module may include feature sets related to information derived from the patient's clinical records and records from the patient's family. These may be abstracted from unstructured clinical documentation, EMR, EHR, or other sources of the patient's medical history. Information may include the patient's symptoms, diagnoses, treatments, medications, treatments, hospice, response to treatments, laboratory test results, medical history, respective geographic location, demographics, or other characteristics of the patient that may be found in the patient's medical record. Information regarding treatments, medications, therapies, etc. may be ingested as recommendations or prescriptions and / or confirmation that such treatments, medications, therapies, etc. have been administered or taken.

[0129] The imaging module can include feature sets associated with information derived from a patient's imaging records. The imaging records can include H&E slides, IHC slides, radiology images, and other medical imaging that a physician may prescribe in the course of diagnosing and treating various illnesses and diseases. These features can include TMB, ploidy, purity, nuclear-cytoplasmic ratio, large nuclei, altered cell state, biological pathway perturbations, altered hormone receptors, immune cell infiltration, immune biomarkers of MMR, MSI, PDL1, CD3, FOXP3, HRD, PTEN, PIK3CA; collagen or stromal composition, appearance, density, or characteristics; tumor budding, size, aggressiveness, metastasis, immune status, chromatin morphology; and other features of cells, tissues, or tumors for prognostic purposes.

[0130] An epigenomic module, such as an epigenomic module of Omics, can include a set of features related to information derived from DNA modifications that do not change the DNA sequence and regulate gene expression. These modifications are often the result of environmental factors based on what the patient can breathe, eat, or drink. These features can include DNA methylation, histone modifications, or other factors that inactivate genes or change gene function without changing the sequence of nucleotides within the gene.

[0131] A microbiome module, such as a microbiome module from omics, may include a feature set related to information derived from a patient's viruses and bacteria. These features may include viral infections that may affect the treatment and diagnosis of certain diseases, as well as bacteria present in the patient's digestive tract that may affect the effectiveness of medications taken by the patient.

[0132] A proteome module, such as a proteome module from omics, can include a set of features related to information derived from proteins produced in a patient. These features can include the composition, structure, and activity of the protein; when and where the protein is expressed; rates of protein production, degradation, and steady-state abundance; how the protein is modified, such as post-translational modifications such as phosphorylation; trafficking of the protein between subcellular compartments; participation of the protein in metabolic pathways; how proteins interact with each other; or modifications to the protein after translation from RNA, such as phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, or nitrosylation.

[0133] Further omics module(s) also include: cognitive genomics, a feature set including the study of changes in cognitive processes related to genetic profiles; comparative genomics, a feature set including the study of genomic structure and function relationships across different biological species or strains; functional genomics, a feature set including the study of gene and protein function and interactions including transcriptomics; feature sets including studies of large scale analysis of gene-gene interactions, protein-protein interactions or protein-ligand interactions; metagenomics, a feature set including the study of metagenomics such as genetic material recovered directly from environmental samples; neurogenomics, a feature set including the study of genetic influences on nervous system development and function; pangenomics, a feature set including the study of the entire collection of gene families found within a given species; individual genomics, a feature set including the genomics study of the sequencing and analysis of an individual's genome so that once the genotype is known, the individual's genotype can be compared to the published literature to determine the likelihood of trait expression and disease risk to enhance personalized medicine suggestions; epigenomics, a feature set including the study of protein and a collection of features that includes studies supporting the structure of the genome, including RNA binders, alternative DNA structures, as well as chemical modifications on DNA; nucleomics, a collection of features that includes the study of the complete set of genomic components that form the cell nucleus as a complex and dynamic biological system; lipidomics, a collection of features that includes the study of cellular lipids, including modifications made to any particular set of lipids produced by a patient; proteomics, a collection of features that includes the study of proteins, including modifications made to any particular set of proteins produced by a patient; immunoproteomics, a collection of features that includes the study of large sets of proteins involved in the immune response; nutritional proteomics, a collection of features that includes studies identifying molecular targets of nutritional and non-nutritional components of the diet, including the use of proteomic mass spectrometry data for protein expression studies; proteogenomics, a collection of features that includes studies of biological studies at the intersection of proteomics and genomics, including data identifying gene annotations; structural genomics, a collection of features that includes the study of the three-dimensional structures of all proteins encoded by a given genome using a combination of modeling approaches;glycomics, a set of features that includes the study of sugars and carbohydrates and their effects in patients;food science, a set of features that includes the study of the intersection between food and nutrition domains through the application and integration of technologies to improve consumer well-being, health, and knowledge;transcriptomics, a set of features that includes the study of RNA molecules including mRNA, rRNA, tRNA, and other non-coding RNA produced in cells;metabolomics, a set of features that includes the study of chemical processes including metabolites, or the unique chemical fingerprints left by specific cellular processes, and their small molecule metabolite profiles;metabonomics, a set of features that includes the study of quantitative measurements of the dynamic multi-parametric metabolic response of cells to pathophysiological stimuli or genetic modifications;nutrigenetics, a set of features that includes the study of genetic variations on the interaction between diet and health with effects on susceptible subgroups;cognitive genomics, a set of features that includes the study of changes in cognitive processes related to genetic profiles;pharmacogenomics, a set of features that includes the study of the sum effect of variations in the human genome on drugs;pharmacomicrobiology, a set of features that includes the study of the effect of the sum of variations in the human genome on drugs; toxicogenomics, a feature set that includes the study of gene and protein activity within specific cells or tissues of an organism in response to a toxicant;mitogenomics, a feature set that includes the study of the process by which mitochondrial proteins interact;psychogenomics, a feature set that includes the study of the process of applying the powerful tools of genomics and proteomics to achieve a better understanding of the biological substrates of brain diseases that manifest as normal behavior and behavioral abnormalities, including the application of psychogenomics to the study of drug addiction to develop more effective treatments as well as objective diagnostic tools, preventative measures, and cures for these disorders;stem cell genomics, a feature set that includes the study of stem cell biology to establish stem cells as a model system for understanding human biology and disease states;connectomics, a feature set that includes the study of neural connections in the brain;microbial organisms, a feature set that includes the study of the genomes of a collection of microorganisms that inhabit the digestive tract;cellomics, a feature set that includes the study of quantitative cell analysis and studies using bioimaging methods and bioinformatics;These may include tomomics, a set of features that includes the study of tomographic and omics methods to understand tissue or cell biochemistry at high spatial resolution from imaging mass spectrometry data; etomics, a set of features that includes the study of high-throughput instrumental measurements of patient behavior; videoomics, a set of features that includes the study of video analysis paradigms inspired by genomics principles (sequential image sequences or videos can be interpreted as a single image capture evolving through time of mutations, revealing patient insights);

[0134] A sufficiently robust set of features can include all of the features disclosed above. However, models and predictions based on available features can include models trained from a much more restrictive selection of features than the exhaustive feature set. Such constrained feature sets can include only tens to hundreds of features. For example, the constrained feature set of the model can include the genomic results of sequencing the patient's tumor, derived features based on the genomic results, the origin of the patient's tumor, the patient's age at diagnosis, the patient's gender and race, and symptoms that the patient brought to the attention of the physician during routine examination.

[0135] A feature store can enhance a patient's feature set through the application of machine learning and analytics by selecting from any features, modifications, or computed outputs derived from the patient's features or modifications to those features. Such feature stores can generate new features from the original features found in the feature module, or can identify and store key insights or analytics based on the features. The selection of features can be based on the modifications or computations generated and can include computations of single or multiple nucleotide polymorphism insertions or deletions in the genome, tumor mutation burden, microsatellite instability, copy number variations, fusions, or other such computations. Exemplary outputs of modifications or computations generated that can inform future modifications or computations include the discovery of lung cancer and variants in EGFR, an epidermal growth factor receptor gene that is mutated in about 10% of non-small cell lung cancers and about 50% of lung cancers in non-smokers. Here, previously classified variants can be identified in the patient's genome that can inform the classification of novel variants or indicate further risk of disease. Exemplary approaches can include enrichment of variants and their respective classification to identify regions that interact with EGFR and are near or have evidence associated with cancer. Any novel variants detected from patient sequencing that are localized in this region or interact with this region will increase the patient's risk. Features that can be utilized in such alteration detection include the structure of EGFR and the classification of variants therein. Models that focus on enrichment can isolate such variants.

[0136] The above referenced models may be implemented as artificial intelligence engines and may include gradient boosting models, random forest models, neural networks (NNs), regression models, naive Bayes models, or machine learning algorithms (MLAs). The MLAs or NNs may be trained from a training dataset. In an exemplary predictive profile, the training dataset may include patient details such as imaging, pathology, clinical, and / or molecular reports, as well as those managed from EHRs or gene sequencing reports. MLA includes supervised algorithms (e.g., algorithms where features / classifications in the dataset are annotated) using linear regression, logistic regression, decision trees, classification and regression trees, naive Bayes, nearest neighbor clustering; unsupervised algorithms (e.g., algorithms where features / classifications in the dataset are not annotated) using a priori, average clustering, principal component analysis, random forests, adaptive boosting; semi-supervised algorithms (e.g., algorithms where features / classifications in the dataset are annotated) using generative methods (mixtures of Gaussians, mixtures of multinomial distributions, hidden Markov models, etc.), sparse separation, graph-based methods (e.g., min-cuts, harmonic functions, manifold regularization), heuristic methods, or support vector machines. NN includes conditional random fields, convolutional neural networks, attention-based neural networks, deep learning, long short-term memory networks, or other neural models where the training dataset includes pathology reports covering multiple tumor samples, RNA expression data for each sample, and imaging data for each sample. Although MLA and neural network identify distinct approaches to machine learning, these terms may be used interchangeably herein. Thus, unless otherwise specified, a reference to an MLA may include a corresponding NN, or a reference to a NN may include a corresponding MLA. Training may include providing a dataset, labeling these traits as they occur in patient records, and training the MLA to predict or classify based on new inputs.Artificial NNs are efficient computing models that have shown their strength in solving difficult problems in artificial intelligence. They have also been shown to be universal approximators (capable of representing a wide variety of functions given the appropriate parameters). Some MLAs can identify important features and identify coefficients or weights for them. The coefficients can be multiplied with the frequency of occurrence of the feature to generate a score, and a particular classification can be predicted by the MLA when the score of one or more features exceeds a threshold. The coefficient schema can be combined with a rule-based schema to generate more complex predictions, such as predictions based on multiple features. For example, 10 important features across different classifications can be identified. There may be a list of coefficients for the key features, and there may be a rule set for classification. The rule set can be based on the number of occurrences of the feature, scaled weights of the feature, or other qualitative and quantitative evaluation of the feature encoded into logic known to those skilled in the art. In other MLAs, the features may be organized in a binary tree structure. For example, the important features that distinguish most classifications may be present as the root of the binary tree and each subsequent branch in the tree until a classification is given based on reaching a terminal node of the tree. For example, a binary tree can have a root node that tests a first feature. The occurrence or non-occurrence of this feature must be present (a binary decision), and logic can traverse the branch that applies to the item being classified. Additional rules may be based on thresholds, ranges, or other qualitative and quantitative tests. Supervised methods are useful when the training data set has many known values ​​or annotations, but the nature of EMR / EHR documents is that many annotations may not be provided. When exploring large amounts of unlabeled data, unsupervised methods are useful for binning / bucketing instances in the data set. A single instance of the above model, or two or more such instances in combination, can constitute a model for the purposes of a model, artificial intelligence, neural network, or machine learning algorithm, as used herein.

[0137] 3A and 3B, schematic examples of devices that may be used in the system 10 are shown. The pathway engine may be included in a computing device 210 that may be included in the system 10. The computing device 210 may communicate (e.g., wired communication, wireless communication) with a pathway database 300, a labeled tumor sample database 400, a drug-pathway interaction database 500, a treatment response database 600, a clinical trial database 700, and a patient report generator 800 via a communication network 20. The patient report generator 800 may be included in a secondary computing device 210 that may be included in the system and / or the computing device 250. The computing device 210 may communicate with the secondary communication device 250. The computing device 210 and / or the secondary computing device 250 may also communicate with a display 290 that may be included in the system 10 via the communication network 20.

[0138] The communication network 20 can facilitate communication between the computing device 210 and the secondary computing device 250. In some embodiments, the communication network 20 can be any suitable communication network or combination of communication networks. For example, the communication network 20 can include a Wi-Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e.g., a 3G network, a 4G network, a 5G network, etc., conforming to any suitable standard, such as CDMA, GSM, LTE, LTE Advanced, WiMAX, etc.), a wired network, etc. In some embodiments, the communication network 20 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semi-private network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks. Each of the communication links illustrated in FIGS. 3A and 3B can be any suitable communication link or combination of communication links, such as a wired link, an optical fiber link, a Wi-Fi link, a Bluetooth link, a cellular link, etc.

[0139] 3C illustrates an example of hardware that may be used in some embodiments of system 10. Computing device 210 may include a processor 214, a display 216, an input 218, a communication system 220, and a memory 222. Processor 214 may be any suitable hardware processor or combination of processors, such as a central processing device ("CPU"), a graphics processing device ("GPU"), or the like, capable of executing programs that may include the processes described below.

[0140] In some embodiments, the display 216 may present a graphical user interface. In some embodiments, the display 216 may be implemented using any suitable display device, such as a computer monitor, a touch screen, a television, etc. In some embodiments, the input 218 of the computing device 210 may include an indicator, a sensor, an actuatable button, a keyboard, a mouse, a graphical user interface, a touch screen display, etc.

[0141] In some embodiments, communication system 220 may include any suitable hardware, firmware, and / or software for communicating with other systems over any suitable communication network. For example, communication system 220 may include one or more transceivers, one or more communication chips and / or chipsets, etc. In more specific examples, communication system 220 may include hardware, firmware, and / or software that may be used to establish coaxial connections, fiber optic connections, Ethernet connections, USB connections, Wi-Fi connections, Bluetooth connections, cellular connections, etc. In some embodiments, communication system 220 enables computing device 210 to communicate with secondary computing device 250.

[0142] In some embodiments, memory 222 may include any suitable storage device that may be used to store instructions, values, etc. that may be used by processor 214 to communicate with secondary computing device 250 via communication system 220, etc., to present content using display 216, for example. Memory 222 may include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 222 may include RAM, ROM, EEPROM, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, etc. In some embodiments, memory 222 may encode computer programs for controlling the operation of computing device 210 (or secondary computing device 250). In such embodiments, processor 214 may execute at least a portion of the computer programs to present content (e.g., user interfaces, images, graphics, tables, reports, etc.), receive content from secondary computing device 250, transmit information to secondary computing device 250, etc.

[0143] The secondary computing device 250 may include a processor 254, a display 256, an input 258, a communication system 260, and a memory 262. The processor 254 may be any suitable hardware processor or combination of processors, such as a central processing device ("CPU"), a graphics processing device ("GPU"), or the like, capable of executing programs that may include the processes described below.

[0144] In some embodiments, the display 256 may present a graphical user interface. In some embodiments, the display 256 may be implemented using any suitable display device, such as a computer monitor, a touch screen, a television, etc. In some embodiments, the input 258 of the secondary computing device 250 may include an indicator, a sensor, an actuatable button, a keyboard, a mouse, a graphical user interface, a touch screen display, etc.

[0145] In some embodiments, communication system 260 may include any suitable hardware, firmware, and / or software for communicating with other systems over any suitable communication network. For example, communication system 260 may include one or more transceivers, one or more communication chips and / or chipsets, etc. In more specific examples, communication system 260 may include hardware, firmware, and / or software that may be used to establish coaxial connections, fiber optic connections, Ethernet connections, USB connections, Wi-Fi connections, Bluetooth connections, cellular connections, etc. In some embodiments, communication system 260 enables secondary computing device 250 to communicate with computing device 210.

[0146] In some embodiments, memory 262 may include any suitable storage device that may be used to store instructions, values, etc. that may be used by processor 254 to communicate with computing device 210 via communication system 260, etc., to present content using display 256, etc. Memory 262 may include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 262 may include RAM, ROM, EEPROM, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, etc. In some embodiments, memory 262 may encode computer programs for controlling the operation of secondary computing device 250 (or computing device 210). In such embodiments, processor 254 may execute at least a portion of the computer programs to present content (e.g., user interfaces, images, graphics, tables, reports, etc.), receive content from computing device 210, transmit information to computing device 210, etc. Display 290 may be a computer display, a television monitor, a projector, or other suitable display.

[0147] Exemplary Training Data for the Disclosed Systems and Methods

[0148] FIG. 4 illustrates a representation of exemplary data from the data input 100 that may be used to train the pathway engine 200n. Specifically, FIG. 4 illustrates a dataset 410 that may include a number of transcriptome values. Each transcriptome value set (e.g., transcriptome value 1 at 411, transcriptome value 2 at 412, .... transcriptome value N at 413) may be associated with a single tissue specimen. Each transcriptome value 411-413 may represent raw or normalized counts corresponding to the expression levels of all possible RNA products of a gene. Each transcriptome value 411-413 may be associated with a single specimen. The dataset 410 may also include one or more pathway markers associated with each specimen and transcriptome value set. For example, a first specimen may be associated with a first pathway marker 414, a second pathway marker 415, and a third pathway marker 416. Each pathway marker may be associated with a pathway (e.g., a pathway included in the pathway database 300). Each pathway label may be a "positive control" or a "negative control" associated with a pathway alteration detected in a DNA dataset associated with the sample. The transcriptome values ​​and pathway label(s) associated with each sample can be used as training data for training another machine learning model, as described below.

[0149] For example, each transcriptome value set may be generated by sequencing each corresponding tissue specimen using RNA-seq or other sequencing methods. The sequencing may be whole exome sequencing or targeted panel sequencing, and may be next generation sequencing. The transcriptome value sets in the dataset 410 may be stored in a table where each column is a gene, each row is a specimen, and the cell value reflects the expression level value of the specimen-gene pair. The raw expression level values ​​may range from 0 to over 10 million. A column representing a gene may represent the combined expression level of all possible RNA products of that gene (e.g., all possible transcripts, splice variants, or isoforms), or a subset of the RNA products of the gene. In various embodiments, the tissue sample is a biopsy or blood sample from a human patient or a tumor organoid.

[0150] In various embodiments, prior to use with the systems and methods, a set of transcriptome values ​​from a bulk specimen (e.g., a specimen having two or more tissue types) is deconvoluted to remove confounders, including biopsy tissue site. In one example, the deconvolution is performed according to the systems and methods disclosed in U.S. Patent Application No. 62 / 786,756, filed December 31, 2018, and U.S. Patent Application No. 62 / 944,995, filed December 6, 2019, both of which are incorporated herein by reference.

[0151] In various embodiments, the systems and methods include additional strategies to detect known technical and biological covariates and incorporate them into the calculation of the pathway perturbation score. The systems and methods may account for the effects of tissue site and tumor purity when calculating the pathway perturbation score.

[0152] In various embodiments, the values ​​of the transcriptome value set may be normalized. The normalized transcriptome values ​​may range from 0 to 8. In one example, the normalization method is performed according to the systems and methods disclosed in U.S. Patent Application Nos. 16 / 581,706 and USPCT 19 / 52801 (filed 9 / 24 / 2019 and 9 / 24 / 2019, respectively), which are incorporated herein by reference.

[0153] A DNA variant dataset may also be associated with each set of transcriptome values ​​in dataset 410 (not shown in FIG. 4). In one example, each DNA dataset may be generated by sequencing a corresponding tissue specimen using DNA-seq or other sequencing methods. The sequencing may be whole exome sequencing or targeted panel sequencing, and may be next generation sequencing. In another example, the DNA dataset is obtained by microarray or SNP array.

[0154] In one example, the DNA dataset includes pathway mutation data. The pathway mutation data may include data describing genetic variants in the DNA dataset, particularly genetic variants in genes and / or promoters associated with a cellular pathway of interest. In one example, the cellular pathway of interest is one of the oncogenic signaling pathways defined by the TCGA consortium. In another example, the cellular pathway of interest is a custom gene set or list of genes. In one example, the DNA dataset is stored as a variant call format (VCF) file. In another example, the DNA dataset is a list of genetic variants. In various embodiments, a subset of the DNA dataset (e.g., data associated with a cellular pathway of interest) or the entire DNA dataset can be used as features to train the pathway engine 200n. The genetic variants may include any class of variants including single nucleotide polymorphisms, fusions, insertions / deletions, copy number variations, etc.

[0155] Each transcriptome value set in the dataset 410 may be associated with one or more data elements that reflect information about the specimen from which the transcriptome value set was derived. As shown in FIG. 4, each transcriptome value is associated with a specimen ID, a cancer type, and one or more dysregulation indicators. Any or all of the dysregulation indicators may be used as features to train the pathway engine 200n. Each dysregulation indicator may be associated with one or more pathways of interest. If there is no cancer type associated with a transcriptome value set, or the associated cancer type is likely incorrect, a cancer type can be determined for the transcriptome, for example, by analyzing a histopathology slide associated with the transcriptome, or by analyzing the transcriptome and any associated data. An example is described in U.S. Patent Application No. 62 / 855,750, entitled "Systems and Methods for Multi-label Cancer Classification," filed May 31, 2019, and incorporated herein by reference. One example of a transcriptome that may not have an associated cancer type, or may have an inaccurate associated cancer type, is a transcriptome associated with a tumor of unknown origin, a metastatic tumor, or an inaccurately labeled cancer sample.

[0156] In one example, the dataset 410 may be filtered to generate a subset of the dataset 410 for training the pathway engine 200n, and may be filtered based on the cancer type and / or pathway of interest. For example, if the pathway engine 200n is designed to be specific to a cancer type (e.g., lung cancer), rows associated with different cancer types may be removed from the dataset 410 prior to DEG selection and training (as described in connection with FIG. 5). As another example, if the pathway engine 200n is specific to a pathway of interest, dysregulation indicators associated with different pathways may be removed from the dataset 410 prior to selecting DEGs and training the pathway engine 200n. Each transcriptome value set and associated dysregulation indicator selected for training the model is converted into a feature vector.

[0157] In some embodiments, the data in the dataset 410 used to train the pathway engine 200n includes more than 30 transcriptome value sets. In some embodiments, the data in the dataset 410 used to train the pathway engine 200n includes more than 900 transcriptome value sets. In some embodiments, the data in the dataset 410 used to train the pathway engine 200n includes more than 10,000 transcriptome value sets.

[0158] In one example, the data in the dataset 410 used to train the pathway engine 200n may be associated with a primary tumor specimen or a single tissue type to minimize transcriptional heterogeneity, although this is not necessary to generate an accurate pathway engine.

[0159] One type of dysregulation indicator can be a pathway label, as shown in FIG. 4. For example, a pathway label can be a "positive control" or a "negative control." A pathway label can be selected based on any detected pathway alterations in a DNA dataset associated with a sample. In one example, if a DNA dataset contains genetic variants in one or more genes and / or promoters associated with a cellular pathway of interest, the corresponding transcriptome value set is assigned the pathway label positive control for that cellular pathway, while a transcriptome value set associated with a DNA dataset that does not contain genetic variants, or in some embodiments does not contain variants or benign variants, in genes and / or promoters associated with the cellular pathway of interest is assigned the label negative control.

[0160] In another example, only if the DNA dataset contains a pathogenic variant in a gene and / or promoter associated with a cellular pathway of interest, where pathogenic means that the variant is known to contribute to the progression of cancer (or other disease state of interest), the corresponding transcriptome value set is assigned a pathway label positive control for that cellular pathway, while transcriptome value sets associated with DNA datasets that do not contain inherited variants or contain benign variants in genes and / or promoters associated with the cellular pathway of interest are assigned a label negative control.

[0161] In yet another example, the negative control transcriptome value set is wild type for all genes in the pathway, and all positive control transcriptome value sets are associated with genetic variants of one or more genes of the genes in the pathway, or one or more genes of a class of genes in a cellular pathway (e.g., the gene class or module can be all RAS genes-KRAS, NRAS, HRAS, etc.; all RAF genes-RAF1, ARAF, BRAF, etc.; all PI3K genes-PIKCA, PIKCB, etc.), and in one example, the genetic variants are all pathogenic. For example, a transcriptome value set of a patient with known pathway dysregulation (e.g., KRAS G12V mutation for the RAS / RTK pathway) is considered a "positive control," and a transcriptome value set of a patient who is wild type (WT) for all genes and promoters associated with the pathway is considered a "negative control."

[0162] In one example, the negative control does not have variants (including copy number variants and variants of unknown significance) in any pathway genes. In one example, transcriptomes with variants of unknown significance in pathway genes or promoters are excluded from the training data. In another example, only if the DNA dataset contains a pathogenic variant in a gene and / or promoter associated with a cellular pathway of interest, pathogenic meaning that the variant is known to contribute to cancer progression, the corresponding transcriptome value set is assigned a pathway label positive control for that cellular pathway, while the transcriptome value set associated with the DNA dataset that does not contain an inherited variant or contains a benign variant in a gene and / or promoter associated with the cellular pathway of interest is assigned a label negative control.

[0163] In yet another example, the negative control transcriptome value set is wild type for all genes in the pathway, and all positive control transcriptome value sets are associated with genetic variants of a subset of genes in the pathway or only one class of genes in a cellular pathway (e.g., the gene class can be all RAS genes - KRAS, NRAS, HRAS, etc.; all RAF genes - RAF1, ARAF, BRAF, etc.; all PI3K genes - PIKCA, PIKCB, etc.), and in one example, the genetic variants are all pathogenic. For example, a transcriptome value set of a patient with known pathway dysregulation (e.g., KRAS G12V mutation for the RAS / RTK pathway) is considered a "positive control," and a transcriptome value set of a patient who is wild type (WT) for all genes and promoters associated with the pathway is considered a "negative control."

[0164] In one example, the negative control does not have variants in any pathway genes (including copy number variants and variants of unknown significance). In one example, transcriptomes with variants of unknown significance in pathway genes or promoters are excluded from the training data. Non-limiting examples of positive and negative control selection are provided below.

[0165] Exemplary positive and negative control selections for pathways, multigene modules, and single gene modules

[0166] (route)

[0167] Now referring to FIG. 4 and FIG. 12, in some embodiments, to train a model to detect pathway dysregulation, a sample can be labeled as a "positive control" or a "negative control". The pathway can be a well-characterized pathway or a custom pathway. Dysregulation can result in a disease, condition, (e.g., cancer), etc., and in some embodiments, the degree of dysregulation caused by a nucleic acid variant can be indicated by classifying a variant or set of variants in the pathway as "benign", "likely benign", "conflicting evidence", "likely pathogenic", "pathogenic", "unknown significance", and "unknown". In some embodiments, a sample can be labeled as a positive control only if it has a nucleic acid variant or set of variants (e.g., a DNA mutation) that is "pathogenic", i.e., associated with a disease or condition such as cancer. Such variants can be germline or somatic. As an example, to train a model to detect dysregulation in the RTK-RAS pathway as illustrated in FIG. 12, a sample would be labeled as a positive control only if it contains at least one pathogenic nucleic acid variant of a gene included in a pathway module of the RTK-RAS pathway. For example, as shown in Figure 12, the RTK-RAS pathway 1200 includes the RAS module 12110, the RAF module 1215, the EGFR module 1205, the PTEN module 1220, the ERBB2 module 1225, the PI3K module 1230, the AKT module 1235, the TOR module 1240, the MEK module 1245, and the ERK module 1250. Thus, in some embodiments, only specimens containing a pathogenic nucleic acid mutation in one or more genes of one or more of these modules are labeled as positive controls for the model. To illustrate with respect to the RAS and RAF modules, only specimens containing one or more pathogenic mutations in one or more of the KRAS, NRAS, HRAS, RAF1, BRAF, and / or ARAF genes are labeled as positive controls.

[0168] In some embodiments, a sample can be classified as a positive control only if it has at least one pathogenic nucleic acid variant in one or more genes included in the pathway. In some embodiments, a sample can be classified as a positive control only if it has at least one pathogenic and / or potentially pathogenic nucleic acid variant in the pathway. Additionally or alternatively, in some embodiments, a sample can be classified as a positive control if the RNA expression levels of one or more genes in the pathway are abnormal and such abnormal expression levels are pathogenic (i.e., associated with a disease or condition, e.g., cancer).

[0169] In some embodiments, a sample can be labeled as a negative control only if it does not have any type of nucleic acid variant in any gene included in the pathway. In some embodiments, a sample can be labeled as a negative control only if it does not have any variants or only has benign or likely benign nucleic acid variants in one or more genes in the pathway in only germline samples. That is, to qualify as a negative control, benign or likely benign mutations present in one or more genes of the pathway are only acceptable if benign or likely benign mutations are present in non-germline samples if they are germline, and the sample is not eligible as a negative control. In other embodiments, a sample can be labeled as a negative control only if it does not have any variants in one or more genes in the pathway or only has benign or likely benign variants. For example, to train a model to detect dysregulation of the RTK-RAS pathway 1200, a sample can be labeled as a negative control only if it does not have any mutations in the genes of the listed modules of the pathway. In other embodiments, a sample can be labeled as a negative control only if it does not have a mutation in one or more genes of the listed modules, or has a benign or likely benign germline mutation.For example, as shown in FIG. 12, the RTK-RAS pathway 1200 includes a RAS module 12110, a RAF module 1215, an EGFR module 1205, a PTEN module 1220, an ERBB2 module 1225, a PI3K module 1230, an AKT module 1235, a TOR module 1240, a MEK module 1245, and an ERK module 1250. The RAS module includes KRAS, NRAS, and HRAS genes, and the RAF module includes RAF1, BRAF, and ARAF genes.Thus, in one embodiment, the negative control of the RAS module includes a sample that does not have a mutation in any of the KRAS, NRAS, and HRSA genes, and the negative control of the RAF module includes a sample that does not have a mutation in any of the RAF1, BRAF, and ARAF genes. The same is true for other modules in the pathway.Additionally or alternatively, in some embodiments, a negative control for the RAS module includes a specimen that has no mutations in any of the KRAS, NRAS and HRSA genes, or that has germline mutations in the KRAS, NRAS and HRAS genes that are benign or have only the possibility of being benign, and a negative control for the RAF module includes a specimen that has no mutations in any of the RAF1, BRAF and ARAF genes, or that has germline mutations in the RAF1, BRAF and ARAF genes that are benign or have only the possibility of being benign. Similarly for other modules in the pathway. Additionally or alternatively, in some embodiments, a specimen can be classified as a negative control if the RNA expression levels of all genes in the pathway are wild type.

[0170] In some embodiments, samples that cannot be classified as positive or negative controls are removed from the training data.

[0171] (multiple gene modules)

[0172] In some embodiments, specimens can be labeled as "positive controls" or "negative controls" to train a model to detect dysregulation within a module (e.g., a grouping of one or more selected genes). Thus, a model can be associated with a module. In some embodiments, a module can include multiple genes selected from a branch of a single pathway, a subset of genes within a pathway, a collection of genes from different pathways, or other suitable groupings of genes. Thus, a pathway can be a well-characterized pathway or a custom pathway. Dysregulation can result in a disease, condition, etc., and in some embodiments, the degree of dysregulation caused by a nucleic acid variant can be indicated by classifying a variant or set of variants within a module as "benign", "likely benign", "conflicting evidence", "likely pathogenic", "pathogenic", "unknown significance", and "unknown".

[0173] In some embodiments, a specimen can be labeled as a positive control only if it has a nucleic acid variant or set of variants (e.g., DNA mutations) that are "pathogenic", i.e., associated with a disease or condition such as cancer. By way of example and not limitation, a model can be trained to detect dysregulation in the RAS module 1210. The nucleic acid variants can be germline or somatic. In some embodiments, for a pathway engine or model trained to detect dysregulation in a module, a specimen can be labeled as a positive control only if it contains a nucleic acid variant in at least one gene included in the module. For example, for a model trained to detect dysregulation in the RAS module 1210, only specimens that contain a pathogenic nucleic acid variant in one or more of the KRAS, NRAS, and / or HRAS genes of the RAS module 1210 can be labeled as a positive control.

[0174] In some embodiments, a specimen may be classified as a positive control only if it has at least one pathogenic nucleic acid variant included in the module associated with the model. Additionally or alternatively, in some embodiments, a specimen may be classified as a positive control only if it has at least one pathogenic and / or potentially pathogenic nucleic acid variant in the module associated with the model. Additionally or alternatively, in some embodiments, a specimen may be classified as a positive control if the RNA expression level of one or more genes in the module is abnormal and such abnormal expression level is pathogenic (i.e., associated with a disease or condition).

[0175] In some embodiments, a sample can be labeled as a negative control only if the sample does not have any type of nucleic acid mutation in any gene included in the module associated with the model. For example, to train a model for detecting dysregulation in the RAS module 1210, a sample can be labeled as a negative control sample only if the sample does not have mutations in the KRAS, NRAS, and HRAS genes of the RAS module 1210.

[0176] In some embodiments, a sample can be labeled as a negative control only if it does not have any type of nucleic acid variant in any gene included in the module associated with the model or any other module included in the overall pathway that includes the module. For example, in the case of a model trained to detect dysregulation in the RAS module 1210, in some embodiments, a sample can be labeled as a negative control sample only if it does not have mutations in the KRAS, NRAS, and HRAS genes included in the RAS module 1210, and does not have mutations in any genes included in other modules included in the RTK-RAS pathway 1200.

[0177] Additionally or alternatively, the negative control does not include mutations in one or more genes in the module, or includes only benign or likely benign germline mutations. Additionally or alternatively, in some embodiments, the negative control does not include variants, or includes only benign or likely benign germline variants, in one or more genes of the module included in the pathway of interest and / or one or more genes of other modules.

[0178] For example, for a model trained to detect dysregulation in the RAS module 1210, in some embodiments, a sample may be labeled as a negative control sample only if it has no mutations or only benign or benign germline mutations in the KRAS, NRAS, and HRAS genes included in the RAS module 1210, and in some embodiments may have no mutations or only benign or benign mutations in other genes included in other modules included in the RTK-RAS pathway 1200.

[0179] Additionally or alternatively, in some embodiments, a sample can be classified as a negative control only if the RNA expression levels of all genes in a module are wild type and / or if the expression levels of all genes in all modules of a pathway of interest (e.g., a pathway that includes a module) are wild type.

[0180] In some embodiments, samples that cannot be classified as positive or negative controls can be removed from the training data.

[0181] (Single gene module)

[0182] In some embodiments, samples can be labeled as "positive controls" or "negative controls" to train a model to detect dysregulation in a module that includes a single gene. Thus, a model can be associated with a module. In some embodiments, a gene can be referred to as a module. A module can include genes included in a pathway module (e.g., the RAS module 1210). For example, a module can include the KRAS gene. In some embodiments, each gene included in a pathway module can be associated with a model trained to detect dysregulation in the module (e.g., the KRAS gene).

[0183] In some embodiments, the dysregulation may result in a disease, condition, etc., and in some embodiments, the degree of dysregulation may be indicated by classifying a nucleic acid variant or set of variants in a module as "benign," "probably benign," "conflicting evidence," "probably pathogenic," "pathogenic," "unknown significance," and "unknown." In some embodiments, a specimen may be labeled as a positive control only if it has a pathogenic nucleic acid variant or set of variants (e.g., DNA mutation) associated with dysregulation in a module (e.g., KRAS gene). The nucleic acid variants may be germline or somatic. In some embodiments, for a model trained to detect dysregulation in a module with a single gene, a specimen may be labeled as a positive control sample only if the specimen contains a pathogenic nucleic acid variant in the gene. For example, in a model trained to detect dysregulation of the KRAS gene, only specimens containing at least one pathogenic nucleic acid variant in the KRAS gene may be labeled as positive controls.

[0184] In some embodiments, a specimen may be determined to have a mutation and classified as a positive control only if the specimen has at least one pathogenic variant in the DNA contained in a gene contained in a module. In some embodiments, a specimen may be determined to have a mutation and classified as a positive control only if the specimen has at least one pathogenic and / or likely pathogenic variant in the DNA contained in a gene contained in a module. Additionally or alternatively, in some embodiments, a specimen may be classified as a positive control if the RNA expression levels of genes in the module are abnormal and such abnormal expression levels are pathogenic (i.e., associated with a disease or condition).

[0185] In some embodiments, a sample can be labeled as a negative control only if it does not have any type of nucleic acid variant in the gene associated with the model. Additionally or alternatively, in some embodiments, a sample can be labeled as a negative control only if it does not have any mutation in the gene associated with the module, or only has benign or likely benign germline mutations. In some embodiments, a sample can be labeled as a negative control only if it does not have any type of nucleic acid variant in the gene associated with the model, or only has benign or likely benign germline variants associated with the model, and only benign or germline variants in the genes of the entire pathway that includes the gene. For example, in the case of a model trained to detect dysregulation of the KRAS gene, a sample can be labeled as a negative control sample only if it does not have any mutation in the KRAS gene. In some embodiments, negative controls include specimens that do not have mutations in the KRAS, NRAS, and HRAS genes in the RAS module 1210 and only have benign or likely benign germline variants in genes in other modules in the RTK-RAS pathway 1200, or specimens that do not have any type of variants in genes in other modules in the RTK-RAS pathway 1200.

[0186] In some embodiments, a sample can be labeled as a negative control only if it does not have any type of nucleic acid variant in any other gene included in the entire pathway including the gene or gene related to the model. For example, in a model trained to detect dysregulation of the KRAS gene, a sample can be labeled as a negative control sample only if it does not have a mutation in the KRAS, NRAS, and / or HRAS genes included in the RAS module 1210 and does not have a mutation in any gene included in other modules included in the RTK-RAS pathway 1200. Additionally or alternatively, in some embodiments, a sample can be classified as a negative control only if the RNA expression level of the gene in the module is wild type, and / or the expression levels of all genes in the module including the single gene module are wild type, and / or the RNA expression levels of all genes in all modules of the pathway of interest (e.g., the pathway including the single gene module) are wild type.

[0187] In some embodiments, samples that cannot be classified as positive or negative controls can be removed from the training data.

[0188] Using only samples that do not contain nucleic acid variants in a pathway, multigene module, or single gene module as negative control samples for training a model to identify dysregulation in a pathway or module can improve the performance of the model compared to other techniques. The discrimination ability (e.g., ability to correctly identify dysregulated and non-dysregulated modules) of a model trained on transcriptome data from negatively labeled samples that contain nucleic acid variants in other modules in the pathway may be reduced because mutations in the module may dilute the effect of any dysregulation in the module associated with the model. For example, a negative sample can provide a baseline of RNA expression levels to compare with a positive sample that can show the effect of dysregulation on RNA expression levels. If a negative sample has DNA variants in modules other than the module associated with the model, the RNA expression levels of the baseline data may dilute and / or obscure the effect of dysregulation on the RNA expression levels of the positive sample. In other words, a model trained on transcriptome data from negatively labeled samples that do not contain DNA variants in both the module relevant to the model (e.g., the RAS module 1210) and other modules in the pathway can better classify the module as more accurately dysregulated or non-dysregulated because the model can more clearly recognize the exact effect of the mutations in the module without the dilution effects of other pathway modules.

[0189] In particular, some mutations classified as pathogenic or likely pathogenic by the above criteria may not ultimately be considered pathogenic or likely pathogenic based on additional information found during training. For example, a sample with the mutation FGFR2 c.1990~106A>G would not normally be accepted in the negative sample set when determining the perturbation score of a module in the RTK / RAS pathway, since it would be classified as pathogenic or likely pathogenic. However, in the generation of the model, it became clear that a significant proportion of the normal population carried this variant, which is very likely to be benign. Such mutations are identified during model training, and an additional step is included to ignore these mutations when generating the sets of positive and negative samples.

[0190] Another type of dysregulation indicator may be a gene set enrichment analysis result. In some examples, the "positive control" and "negative control" transcriptome value sets in the dataset 410 may be similar. In these examples, to help the pathway engine 200n better distinguish the "positive control" transcriptome value set from the "negative control" transcriptome value set, one or more gene set enrichment analysis scores may be associated with each transcriptome value and used as features during training of the pathway engine 200n. For example, each transcriptome value in the dataset 410 may be associated with one or more such gene set enrichment analysis scores, such as a gene set enrichment analysis (GSEA) or single sample GSEA (ssGSEA) score (not shown in FIG. 4). In one example, ssGSEA is a standard tool in the field of pathway analysis (see Barbie, et al., 2010, Nature. 462(7269):108-112).

[0191] A plurality of ssGSEA scores may be associated with each transcriptome value set in the data set 410. In one example, each ssGSEA score may be an individual dysregulation index in the data set 410. Each ssGSEA pathway score may be associated with one or more pathways of interest. The selection of the gene set from which the ssGSEA score is derived may depend on the pathway that the pathway engine 200n is trained on. For example, if the pathway engine 200n is trained to generate pathway perturbation scores for the RAS pathway, the ssGSEA scores of any associated pathways, including the 43 KRAS-related pathways, may be the most relevant ssGSEA scores.

[0192] In one example, the relevant pathway can be any pathway known to be dysregulated in samples with mutations in the genes used to define the positive control samples, for example, for the RAS / RTK pathway, a KRAS mutation is used to define the positive control samples, so a score is generated for all pathways with names that contain the string "KRAS."

[0193] Another type of dysregulation indicator can be the methylation status of a sample associated with a transcriptome value set. The methylation status can be determined by analyzing the methylation of genes and / or promoters associated with a pathway.

[0194] In various embodiments, a subset of the rows in the dataset 410 is used to train the route engine 200n, and the remaining rows of the dataset 410 that are not used to train the route engine 200n are used to test the route engine 200n.

[0195] A protein expression level dataset may also be associated with each transcriptome value set in dataset 410. In one example (not shown in FIG. 4), each protein expression level dataset may be generated by any known method for measuring protein abundance in a specimen, including proteomic methods.

[0196] In various embodiments, the transcriptome value sets in the dataset 410 may be further associated with imaging data. The imaging data may include histopathology and radiology images generated from the specimen associated with the transcriptome value set, features extracted from these images, and any annotations or information developed by manual or automated analysis of these images.

[0197] In various embodiments, the dataset 410 includes data from The Cancer Genome Atlas (TCGA) Consortium.

[0198] In various embodiments, each transcriptome value set can be generated by processing a patient or tumor organoid sample by RNA whole exome next generation sequencing (NGS) to generate RNA sequencing data, and the RNA sequencing data can be processed by a bioinformatics pipeline to generate an RNA-seq expression profile for each sample. The patient sample can be a tissue sample or a blood sample containing cancer cells.

[0199] More specifically, RNA can be isolated from blood samples or tissue sections using commercially available reagents such as proteinase K, TURBO DNase-I, and / or RNA Clean XP beads. The isolated RNA can be subjected to quality control protocols to determine the concentration and / or quantity of RNA molecules, including the use of fluorescent dyes and fluorescent microplate readers, standard spectrofluorometers or filter fluorometers.

[0200] A cDNA library can be prepared from the isolated RNA, purified, and selected for cDNA molecule size selection using commercially available reagents, such as Roche KAPA Hyper Beads. In another example, a New England Biolabs (NEB) kit can be used. Preparation of a cDNA library can include ligation of adapters onto the cDNA molecules. For example, UDI adapters or UMI adapters (e.g., full-length or robust Y adapters), including Roche SeqCap dual-end adapters, can be ligated to the cDNA molecules. The sequence of nucleotides in the adapters can be sample-specific to distinguish the sequencing data obtained for different samples. In this example, the adapters are nucleic acid molecules that can serve as barcodes to identify cDNA molecules according to the sample they are derived from and / or to facilitate next-generation sequencing reactions and / or downstream bioinformatics processing.

[0201] The cDNA library can be amplified and purified using reagents such as Axygen MAG PCR cleanup beads. The concentration and / or amount of cDNA molecules can then be quantified using fluorescent dyes and fluorescence microplate readers, standard spectrofluorometers or filter fluorometers.

[0202] The cDNA libraries can be pooled and treated with reagents to reduce off-target capture, such as human COT-1 and / or IDT xGen Universal Blocker, and then dried under vacuum. The pools can then be resuspended in a hybridization mixture, such as IDT xGen Lockdown, and a probe, such as IDT xGen Exome Research Panel v1.0 probes, IDT xGen Exome Research Panel v2.0 probes, other IDT probe panels, Roche probe panels, or other probes, can be added to each pool. The pools can then be incubated in an incubator, PCR device, water bath, or other temperature-regulating device to allow the probes to hybridize. The pools can then be treated with streptavidin-coated beads or another means to capture hybridized cDNA probe molecules, particularly cDNA molecules that represent exons of the human genome. In some embodiments, polyA capture can be used. The pools can be amplified and purified using commercially available reagents, such as the KAPA HiFi Library Amplification Kit and Axygen MAG PCR Wash Beads, respectively.

[0203] The cDNA library can be analyzed to determine the concentration or amount of cDNA molecules, for example, using fluorescent dyes (e.g., PicoGreen pool quantification) and a fluorescent microplate reader, a standard spectrofluorometer, or a filter fluorometer. The cDNA library can also be analyzed to determine the fragment size of the cDNA molecules, which can be done by gel electrophoresis techniques and may include the use of a device such as the LabChip GX Touch. The pools can be cluster amplified using a kit (e.g., Illumina Paired-End Cluster Kit with PhiX Spike). In one example, the cDNA library preparation and / or whole exome capture steps can be performed in an automated system using a liquid handling robot (e.g., SciClone NGSX).

[0204] Amplification can be performed on a device such as, for example, an Illumina C-Bot2, and the resulting flow cells containing the amplified target capture cDNA library can be sequenced on a next generation sequencer, for example, an Illumina HiSeq 4000 or an Illumina NovaSeq 6000, to a unique on-target depth selected by the user, for example, 300x, 400x, 500x, 10000x, etc. The next generation sequencer can generate a FASTQ file for each patient sample.

[0205] Each FASTQ file contains reads, which may be paired-end or single-read, and may be short or long reads, with each read representing the sequence of one detected nucleotide in an mRNA molecule isolated from a patient sample, inferred by detecting the sequence of nucleotides contained in a cDNA molecule generated from an mRNA molecule isolated during library preparation using a sequencer. Each read in a FASTQ file is also associated with a quality assessment. The quality assessment may reflect the possibility that an error occurred during the sequencing procedure that affected the associated read. Adapters may facilitate the binding of cDNA molecules to anchor oligonucleotide molecules on a sequencer flow cell and may act as seeds for the sequencing process by providing a starting point for the sequencing reaction. When two or more patient samples are processed simultaneously on the same sequencer flow cell, reads from multiple patient samples may be initially included in the same FASTQ file and then split into separate FASTQ files for each patient. The difference in the order of the adapters used for each patient sample may serve the purpose of a barcode, facilitating the association of each read with the correct patient sample and placement in the correct FASTQ file.

[0206] Each FASTQ file may be processed by a bioinformatics pipeline. In various embodiments, the bioinformatics pipeline may filter the FASTQ data. Filtering the FASTQ data may include correcting sequencer errors and removing low quality sequences or bases, adapter sequences, contaminants, chimeric reads, over-represented sequences, biases caused by library preparation, amplification or capture, and other errors (trimming). Whole reads, individual nucleotides, or multiple nucleotides that are likely to have errors may be discarded based on a quality assessment associated with the reads in the FASTQ file, the known error rate of the sequencer, and / or a comparison of each nucleotide in the read to one or more nucleotides in other reads aligned to the same position in the reference genome. Filtering may be performed in part or in whole by various software tools. FASTQ files can be analyzed for rapid assessment of quality control, see, for example, sequencing data QC software, such as AfterQC, Kraken, RNA-SeQC, FastQC (Illumina, BaseSpace Labs, or https: / / www.illumina.com / products / by-type / informatics-products / basespace-sequence-hub / apps / fastqc.html), or another similar software program. In the case of paired-end reads, the reads can be merged.

[0207] For each FASTQ file, each read in the file may be aligned to the location in the reference genome that has a sequence that most closely matches the sequence of the nucleotides in the read. There are many software programs designed to align reads, such as Bowtie, Burrows Wheeler Aligner (BWA), programs that use the Smith-Waterman algorithm, and the like. Alignment may be performed using a reference genome (e.g., GRCh38, hg38, GRCh37, other reference genomes developed by the Genome Reference Consortium, etc.) by comparing the nucleotide sequence in each read to the portion of the nucleotide sequence in the reference genome to determine the portion of the reference genome sequence that most likely corresponds to the sequence in the read. The alignment may take RNA splice sites into account. The alignment may generate a SAM file, which stores the start and end positions of each read in the reference genome, as well as the coverage (number of reads) of each nucleotide in the reference genome. The SAM file may be converted to a BAM file, the BAM file may be sorted, and duplicate reads may be marked for deletion.

[0208] In one example, Callisto software can be used for alignment and RNA read quantification (see Nicolas L Bray, Harold Pimentel, Pall Melsted and Lior Pachter, Near-optimal probabilistic RNA-seq quantification, Nature Biotechnology 34, 525-527 (2016), doi:10.1038 / nbt.3519). In alternative embodiments, quantification of RNA reads can be performed using other software, such as Sailfish or Salmon (see Rob Patro, Stephen M. Mount, and Carl Kingsford (2014) Sailfish enables alignment-free isoform quantification from RNA-seq reads using lightweight algorithms. Nature Biotechnology (doi:10.1038 / nbt.2862) or Patro, R., Duggal, G., Love, MI, Irizarry, RA, & Kingsford, C. (2017). Salmon provides fast and bias-aware quantification of transcript expression. Nature Methods). These RNA-seq quantification methods may not require alignment. There are many software packages that can be used for normalization, quantitative analysis, and differential expression analysis of RNA-seq data.

[0209] For each gene, the raw RNA read count for a given gene can be calculated. The raw read counts can be stored in a tabular file for each sample, with columns representing genes and each entry representing the raw RNA read count for that gene. In one example, the kallisto alignment software calculates the raw RNA read count as the sum of the probabilities for each read that the read aligns to a gene. Thus, in this example, the raw counts are not integers.

[0210] The raw RNA read counts can then be normalized to correct for GC content and gene length, for example using full quantile normalization, and the sequencing depth can be adjusted, for example using the size factor method. In one example, normalization of the RNA read counts is performed according to the methods disclosed in U.S. Patent Application No. 16 / 581,706 or PCT19 / 52801, entitled Methods of Normalizing and Correcting RNA Expression, filed September 24, 2019. The rationale for normalization is that the copy number of each cDNA molecule in the sequencer may not reflect the distribution of mRNA molecules in the patient sample. For example, during the library preparation, amplification and capture stages, certain portions of the mRNA molecules may be over- or under-represented due to artifacts that occur during various aspects of priming of reverse transcription caused by random hexamers, amplification (PCR enrichment), rRNA depletion, and probe binding, as well as errors that occur during sequencing that may be due to GC content, read length, gene length, and other features of the sequence in each nucleic acid molecule. The number of raw RNA reads for each gene can be adjusted to eliminate or reduce over- or under-representation caused by bias or artifacts in the NGS sequencing protocol. The normalized RNA read counts can be saved for each sample in a tabular file, with columns representing genes and each entry representing the normalized RNA read count for that gene (see also Example 9 for further discussion of RNA preparation methods).

[0211] The transcriptome value set may refer to either normalized RNA read counts or raw RNA read counts, as described above.

[0212] 5 illustrates an example of a process 502 that can train the route engine 200n. The process 502 can be implemented as computer readable instructions on one or more memories or other non-transitory computer readable media and executed by one or more processors in communication with the one or more memories or media. In some embodiments, the process 502 can be implemented as computer readable instructions on the memory 222 and / or memory 262 and executed by the processor 214 and / or processor 254.

[0213] At 505, process 502 can select a route from a plurality of routes, such as route database 300. For example, the route selected can be an RTK / RAS route. In some embodiments, process 502 can select a route based on input from a user.

[0214] (Training data selection)

[0215] At 510, the process 502 can receive a training dataset including transcriptome data. For example, the process 502 can receive the dataset 410. The process can generate a matrix of feature vectors for training the pathway engine 200n based on the training data. The training dataset can include any of the data inputs 100 including DNA variant data, methylation data, cancer type, and / or proteomic data. The methylation data can be formatted as positive / negative controls.

[0216] At 512, the process 502 can generate feature vectors based on the training dataset. The process 502 can filter the training dataset by cancer type or subtype, stage classification, or other genotypic or phenotypic filters (e.g., which cancer type a given specimen is associated with). In some embodiments, the process 502 can generate feature vectors based on specimens associated with multiple cancer types. For example, a first specimen can be associated with lung cancer and a second specimen can be associated with breast cancer. The process 502 can generate a matrix of feature vectors for training based on the filtered or unfiltered dataset. Each feature vector can include at least a portion of any transcriptome data, DNA data, and pathway labels associated with each specimen (e.g., at least a portion of the rows of the dataset 410). For example, the feature vector can include transcriptome data and a single pathway label. The transcriptome can include one or more expression levels associated with one or more genes. The process 502 can reserve a portion of the training dataset for testing the trained pathway engine 200n. In one example, 10% of the matrix of feature vectors can be reserved. In another example, 20% of the matrix of feature vectors can be reserved.

[0217] Pathway labels can be predetermined based on DNA mutation data associated with the transcriptome, as described in Figure 4. For example, if DNA data associated with any gene in the pathway (e.g., EGFR in the RTK / RAS pathway, or any other gene in the RTK / RAS pathway) reflects that a sample associated with that transcriptome contains a genetic variant in one of those genes, then the corresponding feature vector generated from that transcriptome can include a positive control pathway label.

[0218] In some embodiments, at 512, process 502 can generate one or more pathway labels for each feature vector. In this manner, process 502 can receive transcriptome data and raw DNA data associated with each specimen and generate pathway labels for the feature vector. However, it is understood that the training data set can include one or more pathway labels for each specimen. Each specimen with a pathway label, such as a dysregulation indicator as described in FIG. 4. Examples of dysregulation indicators include a positive control or a negative control.

[0219] Process 502 can mark a transcriptome as a positive control if it has DNA mutations in a gene or a subset of genes listed in the pathway selected in 505. For example, the RTK / RAS pathway includes the following genes, among others, as shown in FIG. 1A. For example, if the EGFR gene in the DNA dataset reflects a mutation status, the transcriptome can be marked as a positive control. Similarly for other genes in the RTK / RAS pathway that have a mutation status. In another example, a transcriptome can be marked as a positive control if it has DNA changes only in a particular class of genes or parts in the pathway, for example, RAS genes. In an example, only transcriptomes with pathogenic mutations in the selected genes can be a positive control.

[0220] If all genes in the pathway selected in 505 are considered wild type (e.g., there are no DNA variants, which may include copy number changes and all other classes of DNA variants associated with the gene, or there are no pathogenic DNA variants associated with the gene), the transcriptome can be labeled as a negative control.

[0221] Grouping the positive training data to determine the average expression level and grouping the negative training data to determine the average expression level and calculate a similarity metric

[0222] At 515, process 502 can determine a similarity metric for each gene in the transcriptome included in the training data set. For each gene in the transcriptome, process 502 can compare the expression level associated with a group of positive controls in the training data set (e.g., positive pathway marker value) with the expression level associated with a group of negative controls (e.g., negative pathway marker value) to calculate a similarity metric. A comparison can be made for each gene in the transcriptome. Genes that have statistically different expression levels between the two groups are named differentially expressed genes (DEGs).

[0223] Table 1 shows exemplary information for a positive control sample group and a negative control sample group. In this example, the similarity metric is the calculated fold change in gene expression levels between the two groups. The fold change is calculated by dividing the mean gene expression level of the positive control group by the mean gene expression level of the negative control group and taking the log base 2 logarithm of the quotient. [Table 1]

[0224] In some embodiments, expression level comparisons can be calculated by using edgeR, a publicly available package in the R software environment (see https: / / bioconductor.org / packages / release / bioc / html / edgeR.html).

[0225] (Comparing the similarity metric to a threshold to determine differential expression of genes)

[0226] At 517, process 502 can determine for each gene in the transcriptome whether the gene is differentially expressed. Process 502 can compare the absolute value of the log base 2 of the quotient calculated at 515 to a threshold for each gene. Process 502 can designate the gene as a differentially expressed gene (DEG) based on whether the similarity metric is less than, greater than, or equal to the threshold. In some embodiments, the process can determine whether the absolute value of the similarity metric is higher than a threshold, for example, 0.322 (corresponding to a fold difference of 1.25), 0.585 (corresponding to a fold difference of 1.5), or 1.0 (corresponding to a fold difference of 2). If the absolute value of the similarity metric is higher than the threshold for the gene, process 502 can designate the gene as a differentially expressed gene (i.e., a DEG). The number of DEGs in the training dataset can vary depending on the pathway type, the threshold, and / or the training dataset. In one example, about 1,000 DEGs are selected.

[0227] In some embodiments, process 502 can include running edgeR to calculate the fold change and false discovery rate of each gene to identify DEGs. All DEGs identified by edgeR can be selected as training DEGs. In another example, only high-confidence DEGs are selected as training DEGs. In one example, if the absolute value of the fold change is >1.25 and the false discovery rate (FDR) is <0.05, the DEG is determined to be high-confidence. In another example, the stringency is increased and if the absolute value of the fold change is ≥2 and the FDR is <0.01, the DEG is determined to be high-confidence.

[0228] In particular, the DEGs may include one or more of the genes associated with the model trained to detect dysregulation. For example, for a model trained to detect dysregulation in the RAS module 1210, the associated DEGs may include the KRAS gene, the NRAS gene, and / or the HRAS gene. While other techniques may remove genes associated with the model from consideration as DEGs, in some embodiments, the process 502 may only remove genes associated with the model used for training if the gene is not a DEG. By allowing genes associated with the model to be selected as DEGs, it may be possible for those genes to act as positive controls, which may allow the model to be better trained compared to other techniques that remove genes associated with the model from consideration as DEGs.

[0229] (Create feature vectors for each transcriptome in the training data)

[0230] At 519, process 502 can remove all genes that are not DEGs from each transcriptome contained in the feature vector. Each transcriptome can only contain DEGs. For example, as shown in Table 1, KRAS and MUC2 can be determined to be DEGs, and EGFR, ERBB2, ERBB3, and MET can be determined to be non-DEGs. In this example, process 502 can remove the expression levels of EGFR, ERBB2, ERBB3, and MET genes from each transcriptome while retaining the expression levels of KRAS and MUC2 genes.

[0231] Table 2 shows an exemplary feature vector matrix. As shown, the feature vector may include several expression levels associated with several genes contained in the transcriptome, as well as pathway control values, which may be 1 or 0. The expression levels may be raw levels or normalized levels. In some embodiments, the feature vector may also include DNA variant data, methylation data, cancer type data, and / or proteomic data. The methylation data may be formatted in a binary format, such as 1 (positive, i.e., methylated) or 0 (negative, i.e., unmethylated). [Table 2]

[0232] In an alternative embodiment shown in Table 2B, the RNA expression values ​​of each gene are assigned to their corresponding alleles. One way to achieve this is to use the variant allele fraction (VAF) of each mutation as a surrogate. For example, if the variant allele fraction is 50%, the variant is likely to be present in only one allele. If the VAF is 75%, the associated variant is likely to be present in both alleles, but the sample contained 25% normal non-cancerous tissue and did not have the variant. This is one way to incorporate the VAF into the model. An alternative method (not shown) is to include a VAF in the training data, where each VAF is associated with a variant and is further associated with the RNA expression level calculated for the RNA associated with that variant.

[0233] [Table 3]

[0234] At 520, the process 502 can train the pathway engine 200n based on the training feature vector. In one example, each feature vector entry can represent a gene expression value of a DEG in the training data element, or a positive or negative control label. The feature vector can also include a dysregulation indicator associated with the transcriptome value set.

[0235] In some embodiments, the path engine 200n can include a regression model. In some embodiments, the regression model can be trained based on a predetermined alpha parameter value. In some embodiments, the regression model can be a logistic regression model. In some embodiments, the regression model can be a linear regression model, such as a regularized linear regression model. In some embodiments, the regression model can be trained using an elastic net regularization technique and can be referred to as an elastic net model. In some embodiments, a path perturbation score can be used, and the probability that a path is perturbed can be calculated according to the following formula:

number

[0236] The regression model can be trained using the alpha parameter value. The alpha parameter can be used to penalize (and thus train) the regression model for misclassifying samples (e.g., included training data). The alpha parameter value may range from 0, exclusive, less than or equal to 1. The alpha parameter value can be determined using a process detailed below. In some embodiments, process 502 can receive user input indicating a preferred alpha parameter value and train a logistic regression model based on the preferred alpha parameter value.

[0237] In some embodiments, the regression model can be trained using the alpha parameter and at least one other parameter. For example, in some embodiments, the regression model can be trained using an L1 ratio in addition to the alpha ratio. For certain models, such as elastic net models, the L1 ratio can determine the type of regularization used to train the model. The L1 ratio can be determined using a similar process as the alpha value, for example, by comparing the performance of multiple models with different L1 values ​​in addition to the alpha value.

[0238] In some embodiments, the model used may be an Elastic Net Linear Model from SciKit-Learning. In these embodiments, the model may be trained using an objective function.

number

[0239] The values ​​of the two parameters, the alpha parameter a and the L1 ratio l1, can be determined using a grid search with 10-fold or 15-fold cross-validation, as described below.

[0240] The number of DEGs and / or feature vectors included in each feature vector varies inversely with the alpha parameter. For example, for a larger number of DEGs and / or feature vectors (e.g., 2000 DEGs and 10,000 feature vectors), the alpha parameter value may be 0.1. As another example, for a smaller number of DEGs and / or feature vectors (e.g., 20 DEGs and 2000 feature vectors), the alpha parameter value may be 0.5. The alpha parameter value may be used in methods of regularization, such as elastic net regularization. In some embodiments, process 502 may set the alpha parameter value to 0.2. In some embodiments, process 502 may receive the alpha parameter value from another process, such as process 602, described below.

[0241] At 522, the process 502 may output the trained route engine 200n. In some embodiments, at 522, the process 502 may store the trained route engine 200n in a memory (e.g., memory 222 and / or memory 262). The memory may be included in the computing device 210.

[0242] In some embodiments, process 502 can receive training data that includes only transcriptome data associated with DEGs. In other words, substeps 515, 517, and 519 may have already been performed to remove non-DEGs from the transcriptome data. In these embodiments, the process can proceed to step 520 following step 512.

[0243] 6A, 6B, 6C, 6D, 6E, and 6F relate to an exemplary method for testing and improving the performance of the route engine 200n.

[0244] FIG. 6A illustrates an example process 602 that can select an alpha parameter value for training a pathway engine, such as pathway engine 200n. Process 602 can be implemented as computer-readable instructions on one or more memories or other non-transitory computer-readable media and executed by one or more processors in communication with one or more memories or other media. In some embodiments, process 602 can be implemented as computer-readable instructions on memory 222 and / or memory 262 and executed by processor 214 and / or processor 254. With reference to both FIG. 5 and FIG. 6A, at 610, process 602 can train a pathway engine and determine the performance of the trained pathway engine. The pathway engine can be pathway engine 200n trained using process 502 described above. To evaluate the performance of the pathway engine, the pathway engine can be tested on transcriptomes that were not included in the training data (e.g., reserved for testing as described in step 510).

[0245] In some embodiments, the process 602 can determine the performance of the trained pathway engine by using the trained pathway engine to generate a pathway perturbation score for each stored test transcriptome (see FIG. 7C). The process 602 can provide the reserved feature vector to the trained pathway engine and receive the generated pathway perturbation score from the trained pathway engine. The process 602 can compare the generated pathway perturbation score to the dysregulation index (described in FIG. 4) associated with the transcriptome to determine whether the pathway engine 200n accurately predicted the perturbation status of the pathway of the test transcriptome and calculate a performance metric. In one example, calculating the performance metric includes generating a receiver operating characteristic (ROC) curve and calculating an area under the curve (AUC). In another example, calculating the performance metric includes performing a Wilcoxon rank sum test (see FIG. 6B).

[0246] For example, process 602 can generate a pathway perturbation score using a pathway engine and compare the pathway perturbation score to a threshold to determine a qualitative pathway perturbation score. In one example, the threshold can be selected by selecting a threshold that maximizes the area under the curve (AUC), for example, using reserved transcriptome training data. In another example, the threshold can be selected by selecting a threshold that maximizes a statistical measure defined as the harmonic mean of F1 score, precision (true positive) / (true positive + false positive) and recall (true positive) / (true positive + false negative). In one example, if the distribution of scores returned for the negative control group is irregular for the pathway engine, outliers can be removed before the maximum F1 score is determined. In other embodiments, because the groups are sized unbalanced or one metric of success is more important than another (e.g., precision over recall), a threshold that maximizes another metric may be desirable, including a) Youden's J statistic (specificity + sensitivity - 1), b) precision (true positive + true negative) / (total number of samples), c) precision, or d) recall.

[0247] At 610, process 602 can train multiple pathway engines using a number of different alpha parameter values. Process 602 can then provide test data to each of the trained pathway engines and compare the performance of each trained pathway engine. In one example, the logistic regression parameter α used to train the pathway engines in process 502 can be varied (e.g., from 0.1 to 1 in increments of 0.05). Process 602 can determine the performance of each training pathway engine by calculating either the AUC, Wilcoxon rank sum test, Youden's J statistic (specificity + sensitivity - 1), accuracy (true positives + true negatives) / (total number of samples), precision, or recall of each training pathway engine.

[0248] In one example, at 610, process 602 can perform optional cross-validation of the pathway engine. A possible goal of cross-validation can be to ensure that the pathway engine does not "overfit" the data (e.g., learn specific aspects of the training data set that are not generalizable).

[0249] In one example of cross-validation, for each pathway engine trained in 610, the pathway engine being tested can be trained on a different portion of the data selected in step 510, with the remaining data reserved for testing in step 610. For example, the dataset selected in step 510 can be split into portions having an equal number of transcriptomes, one portion can become the reserved set of testing transcriptomes for each pathway engine trained in 610, with the remaining transcriptomes used to train the pathway engine as described above in connection with FIG.

[0250] In one example, each portion is 10% of the dataset, and step 610 is repeated 10 times, with each portion serving as a reserved test transcriptome for one of the pathway engines trained in step 610, referred to as 10-fold cross-validation. In this example, the pathway engine is run on the reserved 10% samples (out-of-fold) and the AUC is calculated for these reserved samples. The pathway engine 200n output for each reserved (saved) transcriptome is saved, as well as the AUC specific to this test set. This process is repeated 10 times such that the 10x out-of-fold sets do not overlap or cross. That is, each transcriptome in the entire dataset selected in step 510 is included in the reserved 10% test set only once and has only one pathway engine output associated with it. The outputs and AUCs of each of the reserved 10 test sets are collected and, together with their known states in either the positive or negative control sets, a final ROC is generated, which reflects the output of the out-of-fold dataset and is therefore referred to as the out-of-fold ROC.

[0251] In an alternative embodiment, a 5-fold cross-validation with an 80 / 20 split can be performed. In this example, the transcriptomes in the dataset selected in 510 are split into 5 equal parts, and for each of the 5 pathway engines trained in step 610, one of the parts (20% of the dataset) is used to test a pathway engine trained on the remaining 80% of the transcriptomes in the dataset.

[0252] In another example, the path engine is trained on each subset of the data and tested on the remainder as described above using the same alpha parameter value for each training example such that each AUC produced by each test data set is associated with the same alpha parameter value.

[0253] In some embodiments, at 610, process 602 can divide a cohort of similar patients into a training set t1 and a holdout set h1. Process 602 can divide the training set t1 into a training set t2 and a holdout set h2. Process 602 can determine differentially expressed genes in the training set t2 and perform cross-validation to determine a final alpha parameter value and a final L1 parameter value. The final alpha parameter value and the final L1 parameter value can be the alpha parameter value and the L1 parameter value associated with the best cross-validation result. Process 602 can train a final model on the training set t2 using the final alpha parameter value and the final L1 parameter value. Process 602 can apply the final model to the holdout set h2 to select a final threshold for classifying patients as dysregulated / non-dysregulated. Process 602 may determine the final threshold by selecting a threshold such that the maximum number of patients with a perturbation (e.g., true positives) is above the threshold and / or the maximum number of patients without the perturbation (e.g., true negatives) is below the threshold. In some embodiments, process 602 may determine the final threshold by determining a threshold that maximizes the number of correct classifications and / or minimizes the number of incorrect classifications. To validate the final model and final threshold, process 602 may apply the final model and final threshold to the holdout set h1 and calculate the AUC of the final model and final threshold.

[0254] At 615, process 602 can determine a final alpha parameter value based on the performance determined at 610. As described above, process 602 can determine a performance metric for several pathway engines trained using different alpha parameter values. There can be more than one performance metric for a given alpha parameter. In some embodiments, the performance metric can be AUC. In these embodiments, process 602 can select the alpha parameter value associated with the maximum AUC as the final alpha parameter value. In other embodiments, other performance metrics can include Wilcoxon rank sum test, Youden's J statistic (specificity + sensitivity - 1), accuracy (true positives + true negatives) / (total number of samples), precision, or recall for each trained pathway engine. In these embodiments, process 602 can select the alpha parameter value associated with the peak value of the selected performance metric, and process 602 can select the alpha parameter value associated with the highest precision value.

[0255] The AUCs resulting from multiple path engines trained at 610 can be compared to analyze the variance in alpha values ​​caused by different training data subsets and / or the impact of each alpha parameter value on the performance of the path engine. These analyses can facilitate the selection of the final alpha parameter value.

[0256] In one example, process 602 can calculate a standard deviation of the AUCs. In one example, the standard deviation can be calculated for multiple AUCs associated with the same alpha parameter value. In another example, the standard deviation can be calculated for AUCs associated with multiple alpha parameter values.

[0257] In some embodiments, process 602 can determine a final alpha value and a final L1 value that are the alpha value and L1 value associated with the model trained at 610 with the highest AUC or other suitable performance metric (e.g., Wilcoxon rank sum test, accuracy, etc.).

[0258] At 620, the process 602 may determine whether to retrain the path engine. The process 602 may determine whether to retrain the path engine based on the results of 615. The process 602 may compare the selected final alpha parameter value and the associated path engine performance metric with a predefined threshold and determine whether the trained path engine meets the threshold. In one example, a low standard deviation (<≈0.03) and a high AUC (>≈0.80) are generally characteristics of an accurate model. The process 602 may determine whether the standard deviation of the trained path engine is lower than a predefined standard deviation threshold (e.g., 0.03) and whether the AUC of the trained path engine is higher than a predefined AUC threshold (e.g., 0.80). If the process 602 determines that the standard deviation of the trained path engine is lower than the predefined standard deviation threshold and the AUC of the trained path engine is higher than the AUC predefined threshold, the process 602 may determine that the path engine does not need to be retrained. If process 602 determines that the standard deviation of the trained pathway engine is equal to or greater than a predetermined standard deviation threshold or the AUC of the trained pathway engine is equal to or less than an AUC predetermined threshold, process 602 can determine that the pathway engine needs to be retrained. In one example, if the pathway engine needs to be retrained, process 602 can retrain the pathway engine using the original training data and additional features that were not present in the original training data. For example, the additional features can include ssGSEA scores or other dysregulation indicators, as described in FIG. 4.

[0259] If the process 602 determines that the route engine needs to be retrained (i.e., “yes” at 620), the process 602 may return to 610. If the process 602 determines that the route engine does not need to be retrained (i.e., “no” at 620), the process 602 may proceed to 625.

[0260] At 625, the process 602 may output a trained path engine associated with the final alpha parameter value. The process 602 may output a trained path engine that has already been generated, or may train a new path engine using all of the training data and the final alpha parameter value and output the new path engine. The process 625 may store the trained path engine in a memory (e.g., memory 222 and / or memory 262). The memory may be included in the computing device 210.

[0261] 5 and 6B, an exemplary process 630 is shown that can test the pathway engine using additional test transcriptomes for any given test. Process 630 can be implemented as computer readable instructions on one or more memories or other non-transitory computer readable media and executed by one or more processors in communication with the one or more memories or media. In some embodiments, process 630 can be implemented as computer readable instructions on memory 222 and / or memory 262 and executed by processor 214 and / or processor 254.

[0262] At 639, the process 630 can receive a trained route engine, such as route engine 200n. The route engine can be trained using the method 502 of FIG.

[0263] At 640, the process 630 can receive additional test transcriptomes for any tests.

[0264] At 641, process 630 can provide each additional test transcriptome to a pathway engine, such as pathway engine 200n. At 642, process 630 can receive a pathway perturbation score for each additional test transcriptome from the pathway engine. The pathway engine can generate and output a pathway perturbation score for each additional test transcriptome.

[0265] At 644, process 630 can associate each additional test transcriptome with either a positive or negative control label based on the DNA mutation data of the additional test transcriptome. Step 644 can include at least a portion of step 512.

[0266] At 646, process 630 can compare the pathway perturbation scores generated for the positive control transcriptome to the pathway perturbation scores generated for the negative control transcriptome using a predetermined performance metric. In some embodiments, process 630 can compare the pathway perturbation scores generated for the positive control transcriptome to the pathway perturbation scores generated for the negative control transcriptome using an AUC. Process 630 can calculate the AUC of the pathway perturbation scores using a threshold associated with the model included in the pathway engine. In some embodiments, process 630 can compare the pathway perturbation scores generated for the positive control transcriptome to the pathway perturbation scores generated for the negative control transcriptome using a Wilcoxon rank sum test. A significant difference (e.g., p<0.01) when comparing the scores of these groups in the same direction as the training data (e.g., showing that larger scores in the additional test dataset are associated with the same groups as larger scores in the test dataset) can be evidence that the system and method are robust and generalizable to accurately analyze samples other than the original test dataset.

[0267] At 648, process 630 can output the results of the Wilcoxon rank sum test. Process 630 can output the results of the Wilcoxon rank sum test to a display (e.g., display 290, display 256, and / or display 216) for presentation to a user. Process 630 can determine whether the pathway engine is robust and generalizable to accurately analyze samples outside of the original test dataset.

[0268] Figures 6C and 6D show exemplary results of the Wilcoxon rank sum test used to analyze the pathway perturbation scores generated by the pathway engine. In Figures 6C and 6D, the pathway engine was designed to score either the RAS gene group (Figure 6C) or the ERBB2 gene group (Figure 6D). In this example, the RAS gene group includes KRAS, NRAS, and HRAS genes, and the ERBB2 gene group includes only the ERBB2 gene.

[0269] In Figures 6C and 6D, each transcriptome is assigned to a wild type (WT) (left) or positive control (right) group, and pathway engine 200n is used to generate a pathway perturbation score (as described in Figure 7C). The y-axis shows the numerical value of each pathway perturbation score associated with each transcriptome. The x-axis shows the WT or mutant status associated with each transcriptome for all genes in either the RAS pathway in Figure 6C or the ERBB2 pathway in Figure 6D. The horizontal dashed line indicates the threshold (0.85 in Figure 6C and 0.55 in Figure 6D). Transcriptomes with pathway perturbation score values ​​above the threshold are considered to be associated with pathway perturbation.

[0270] 6B and 6C-D, the results shown in FIGS. 6C-D may be determined in step 646 and output in step 648 of the method 630.

[0271] In this example, the boxes in Figures 6C and 6D outline potential "hidden responders" that are WT patients with pathway engines 200n whose outputs exceed the perturbation threshold (dashed line).

[0272] 5 and 6E, an exemplary process 650 is shown in which a trained pathway engine can be biologically validated. Biological validation can be optional. Process 650 can be implemented as computer readable instructions on one or more memories or other non-transitory computer readable media and executed by one or more processors in communication with the one or more memories or media. In some embodiments, process 650 can be implemented as computer readable instructions on memory 222 and / or memory 262 and executed by processor 214 and / or processor 254.

[0273] At 652, the process 650 can receive a trained route engine. The route engine can be route engine 200n. The route engine can be trained using method 502 of FIG.

[0274] At 654, process 650 can biologically validate the pathway engine. For example, process 650 can determine a degree of correlation between the pathway perturbation score generated by the pathway engine and the protein data for each analyte represented by a set of transcriptome values ​​in the test dataset and / or additional test dataset having associated protein data. Process 650 can plot the protein data for each analyte on the x-axis and the pathway perturbation score generated by the pathway engine output on the y-axis. Process 650 can use the plotted data to perform a R 2 A value and an associated p-value can be calculated. Protein data can include measurements of protein expression levels (amount of protein detected in a sample) and / or protein activation levels. For example, protein activation levels can include the total amount of activated protein in a sample or a portion of one or more proteins determined to be present in an activated form, where an example of an activated form of a protein is a phosphorylated protein.

[0275] In one example, a strong correlation (e.g., an R of greater than 0.2) 2 A p-value and / or p-value <1e-5) may indicate that the pathway engine results are biologically significant and reflect a pathway dysregulation affecting protein expression or activation levels. Protein expression or activation levels for an analyte may be predicted by using the pathway engine to generate a pathway perturbation score for the analyte and converting the pathway perturbation score to a protein level based on the correlation determined in 654.

[0276] At 656, the process 650 can output the validation data. 2 The values ​​and / or associated p-values ​​generated in 654 may be output to a display (e.g., display 290, display 256, and / or display 216). The user may then use plots, R 2 You can see the values ​​and / or associated p-values.

[0277] 5 and 6F, an exemplary process 660 is shown that can orthogonally validate a trained route engine. Orthogonal validation can be optional. Process 660 can be implemented as computer readable instructions on one or more memories or other non-transitory computer readable media and executed by one or more processors in communication with one or more memories or media. In some embodiments, process 660 can be implemented as computer readable instructions on memory 222 and / or memory 262 and executed by processor 214 and / or processor 254.

[0278] At 662, the process 660 can receive a trained route engine, such as route engine 200n. The route engine can be trained using the method 502 of FIG.

[0279] At 664, process 660 can orthogonally validate the trained pathway engine. Process 660 can orthogonally validate the trained pathway engine by determining a correlation between the pathway perturbation scores generated by the pathway engine and the output of known pathway analysis methods for each transcriptome in the set of transcriptomes. Known pathway analysis methods can include gene set enrichment analysis (GSEA), gene set variation analysis (GSVA), single sample GSEA (ssGSEA), and / or other pathway analysis methods.

[0280] In 666, process 660 can output any data generated in 664. For example, process 660 can output a correlation between the pathway perturbation scores generated by the pathway engine and the output of a known pathway analysis method for each transcriptome in the set of transcriptomes. Process 660 can output the data to a display (e.g., display 290, display 256, and / or display 216). A user can then view the output data to verify whether the pathway engine has been orthogonally validated.

[0281] Now referring to FIG. 6G, an exemplary process 670 for training a model is shown. The process 670 can train a model to recognize perturbations in modules in a pathway. A module can include one or more genes. For example, as shown in FIG. 12A, the RTK / RAS-PI3K-EGFR pathway, which may be referred to as the RTK-RAS pathway 1200, can include one or more of an EGFR module 1205, a RAS module 1210, a RAF module 1215, a MEK module 1245, an ERK module 1250, a PTEN module 1220, an ERBB2 module 1225, a PI3K module 1230, an AKT module 1235, and a TOR module 1240. The EGFR module 1205 can include the EGFR gene. The RAS module 1210 can include the KRAS gene, the NRAS gene, and the HRAS gene. The RAF module 1215 can include the RAF1 gene, the BRAF gene, and the ARAF gene. In the case of the RTK-RAS pathway, process 670 can be used to train a model associated with the EGFR module 1205, a model associated with the RAS module 1210, and a model associated with the RAF module 1215.

[0282] The process 670 can train a regression model, such as a linear regression model. The linear regression model can be an elastic net linear regression model. The model can be included in a pathway engine, such as pathway engine 200n. In some embodiments, the model can be associated with a type of cancer, such as lung cancer, breast cancer, etc. In some embodiments, the model can be associated with multiple types of cancer. In this manner, the model can detect pathway dysregulation independent of cancer type. The process 670 can be implemented as computer readable instructions on one or more memories or other non-transitory computer readable media and executed by one or more processors in communication with the one or more memories or media. In some embodiments, the process 670 can be implemented as computer readable instructions on memory 222 and / or memory 262 and executed by processor 214 and / or processor 254.

[0283] At 672, the process 670 can receive a number of positively labeled samples and a number of negatively labeled samples. Each sample can include transcriptome data generated based on a tissue sample associated with a patient. The positively labeled samples and the negatively labeled samples can be associated with a particular pathway module (e.g., the RAS module 1210). In the case of a pathway module, the positively labeled samples can be referred to as pathogenic alteration samples, and can be samples that have at least one pathogenic variant, and / or in some embodiments at least one likely pathogenic variant, in at least one of the genes in the module. The negatively labeled samples can be samples that have no somatic mutations, pathogenic (or likely pathogenic) variants, or variants of unknown significance in any gene in the pathway as a whole (i.e., any gene in any module in the entire pathway as defined by TCGA). For example, for a model trained with the RAS module 1210, the positive cohort is samples that have a mutation in at least one of the KRAS, HRAS, or NRAS genes, and the negative cohort is samples that have no somatic mutations, pathogenic (or likely pathogenic) mutations, or variants of unknown significance in any gene throughout the RTK-RAS pathway.

[0284] At 674, process 670 may determine a training set and a holdout set based on the samples received at 672. Process 670 may randomly select a predetermined percentage of both positively and negatively labeled samples to use as the training set. The remaining positively and negatively labeled samples may be used as the holdout set. In some embodiments, process 670 may select approximately 80% of the positively and negatively labeled samples to use as the training set. In other embodiments, process 670 may select approximately 90% of the positively and negatively labeled samples to use as the training set. The training set may be used to train a model, and the holdout set may be used to evaluate the model.

[0285] At 676, process 670 can determine a set for training the model and a set for determining a threshold associated with the model based on the training set. The set for training is referred to as a hyper-parameter set, and the set for determining the threshold is referred to as a threshold set. Process 670 can randomly select a predetermined percentage of both positively and negatively labeled samples in the training set for use as the hyper-parameter set. The remaining positively and negatively labeled samples can be used as the threshold set. In some embodiments, process 670 can select about 80% of the positively and negatively labeled samples in the training set for use as the hyper-parameter set. In other embodiments, process 670 can select about 90% of the positively and negatively labeled samples in the training set for use as the hyper-parameter set. In some embodiments, process 670 may split the training set and select approximately 80% of the positively and negatively labeled samples as the training set, and select two subsets of 10% of the positively and negatively labeled samples, one used to determine the threshold that maximizes the AUC, and one used to validate the model and the selected threshold. In some embodiments, all three sets are selected to contain equal proportions of positive and negative samples. The hyper-parameter set may be a set of parameters, such as an alpha parameter (e.g., a in equation (2) above) and an L1 parameter (e.g., l1 in equation (2) above). 比 ) may be determined. In some embodiments, a set of thresholds may be used to evaluate the model.

[0286] At 678, process 670 can determine differentially expressed genes (DEGs). The process can determine DEGs based on each sample included in the hyper-parameter set. Process 670 can calculate a difference metric between positively labeled samples and negatively labeled samples for each gene included in the transcriptome data. Process 670 can compare the calculated difference metric for each gene to a predefined threshold and retain the gene if the difference metric is below the threshold (or in some embodiments, above the threshold). In some embodiments, process 670 can determine differentially expressed genes using a t-test between positively labeled samples and negatively labeled samples for each gene included in the transcriptome data. Process 670 can correct the P-value generated using the t-test to a Benjamini-Hochberg false discovery rate (FDR). Process 670 can retain genes with Benjamini-Hochberg FDR below a predefined threshold, such as 0.05, for modeling and use as DEGs. Either the P-value or the FDR can be used as a similarity metric.

[0287] At 680, process 670 may determine the final training parameters for the model. In an embodiment where the model is an elastic net linear model, process 670 may determine the final training parameters using equation (2) above). Process 670 may determine the peak in equation (2) using a coordinate descent method. Process 670 may determine the alpha and L1 ratio parameters using a grid search with 10-fold or 15-fold cross-validation on the hyper-parameter set. In some embodiments, the parameter values ​​tested may include alpha values ​​in the range [0.1, 0.5, 1, 2, 5, 10] and L1 ratio values ​​in the range [0, 0.05, 0.1, 0.2, 0.4, 0.6, 0.8, 1]. Process 670 may select the set of alpha and L1 ratio parameters with the highest average AUC from the cross-validation to be the final alpha and L1 ratio parameters.

[0288] At 682, process 670 can train a final model using the final training parameters. In some embodiments, process 670 can train a final elastic net linear model using the final alpha and L1 ratio parameters. Process 670 can then proceed to 684 and 688 in parallel.

[0289] At 684, process 670 can calculate the pathway dysregulation scores of the model against the threshold set to find a probability distribution of the final model. The output of the model may not directly classify patients as dysregulated or non-dysregulated. For example, the output distribution of dysregulated and non-dysregulated patients at the threshold set (not used to train the model) can be graphed as shown in FIG. 6C. The distribution can represent the scores output by the model for positively and negatively labeled samples at the threshold set.

[0290] At 686, process 670 can determine a final threshold based on the distribution. Process 670 can determine the threshold by maximizing the AUC over the distribution. In FIG. 6C, the threshold 649 is about 0.85. Process 670 can determine the threshold based on a set that was not used to train the model and is not the true holdout set, so that process 670 can select an appropriate threshold to approximate what the distribution is on the holdout set and improve performance compared to if the threshold were determined using the true holdout set.

[0291] At 688, process 670 can calculate pathway dysregulation scores for the holdout population using the final model to calculate pathway dysregulation scores for the holdout population. Process 670 can also generate a probability distribution (e.g., the same type of probability distribution generated at 684).

[0292] At 690, process 670 can classify patients included in the holdout set as dysregulated or non-dysregulated based on the final threshold. Process 670 can calculate the AUC across the distribution. The AUC can be the average of the sensitivity and specificity of the model where patients above the final threshold are predicted to be dysregulated and patients below the final threshold are predicted to be non-dysregulated. The AUC can also indicate the overall performance of the final model in the general population since the holdout set was not used to train the model.

[0293] At 692, process 670 can determine the performance of the final model using the AUC calculated at 690. Process 670 can compare the AUC to a predetermined target AUC and decide to retrain the model if the AUC is below the target AUC. Process 670 can display the AUC (e.g., on display 290) for a human practitioner to analyze and / or evaluate the performance of the final model.

[0294] Referring now to FIG. 6H, a process 750 is shown that can select training data for training a model (e.g., a linear regression model) using a model training process such as process 670 of FIG. 6G. More specifically, process 750 can determine whether a sample should be assigned to a group (e.g., a cohort) of positively labeled samples, a group of negatively labeled samples, or excluded from the samples used to train a model associated with either a module (e.g., EGFR module 1205 of FIG. 12A) or an entire pathway (e.g., the entire RTK-RAS pathway 1200 shown in FIG. 12A). The samples can include RNA data, DNA data, cancer type, quality assessment, and other clinically relevant data associated with a tissue sample from a tumor. The model can be associated with a given cancer type.

[0295] In some embodiments, the model may be associated with a pathway (e.g., the RTK-RAS pathway 1200). In some embodiments, the model may be associated with a module included in the pathway (e.g., the RAS module 1210 included in the RTK-RAS pathway 1200). In some embodiments, the model may be associated with a module that includes a single gene included in the pathway (e.g., the KRAS gene included in the RTK-RAS pathway 1200). In some embodiments, a module that includes a gene may have multiple genes.

[0296] At 752, process 750 can receive samples associated with a patient. The samples can be included in a database. Each sample can include RNA data, DNA data, cancer type, methylation status, protein data, ssGSEA data, and / or other clinically relevant data associated with a tissue sample from a tumor. First, process 750 can place all samples into a sample group. Process 750 can then remove unqualified samples from the sample group and label samples included in the group as positive controls (e.g., indicative of dysregulation) or negative controls (e.g., indicative of non-regulation). In some embodiments, the RNA data can include expression values ​​of more than 19,000 genes.

[0297] Each sample can be generated by subjecting tissue samples to targeted panels or whole genome DNA sequencing. Each sample can include a complete list of variants detected, the variant allele fraction (VAF), and the log odds ratio (LOR) of the copy number of each gene in the sample. The list of variants detected for the sample can include single nucleotide variations (SNVs) and insertions / deletions (indels). The sample can include a pathogenicity classification of "benign", "probably benign", "conflicting evidence", "probably pathogenic", "pathogenic", "unknown significance", or "unknown" for each variant in the list of variants detected. The decision of which category a given variant belongs to can be made based on the criteria set forth by the American College of Medical Genetics and Genomics (ACMG). Multiple levels of evidence can be considered, including the frequency of the variant in the population, direct clinical evidence, and the expected effect of the variant on gene expression and / or the function of the translated protein. These levels of evidence are integrated to generate a final decision of the category. Further limited criteria for variant pathogenicity can be generated using DNA variant databases. The sample can include a classification of each variant indicating whether the variant likely originated from the tumor ("somatic") or was present in the patient at birth ("germline"). The VAF can be a measure of how likely an allele is present in a tissue sample compared to the version of the gene present in normal tissue adjacent to the tumor. The log odds ratio of the copy number of each gene can be used by process 750 to determine whether the gene is amplified or deleted. For example, an LOR of 0 can indicate that the copy number of the gene is normal (i.e., 2), an LOR>2 can indicate a strong possibility of amplification, and an LOR<-2 can indicate a strong possibility of deletion.

[0298] Copy number variations can be used to determine the pathogenicity of a sample. Reference databases can include data on whether amplification or deletion indicates that a gene is pathogenic. For example, amplification (i.e., copy number gain) of ERBB2 is considered pathogenic, while deletion (i.e., copy number loss) is not. The opposite is true for gene PTEN. Only these pathogenic copy number changes are considered when determining whether and how a sample is used to generate a pathway perturbation model.

[0299] Whether a given sample has an amplification or deletion in a gene is based on where its copy number log odds ratio (CNLOR) falls in the distribution of the CNLOR of that gene for all samples in the cohort considered. Specifically, if the CNLOR of a gene is 2.0 standard deviations above the average CNLOR of all samples in the cancer cohort considered, the gene is considered to be amplified, and if the CNLOR of a gene is 2.0 standard deviations below the average CNLOR, the gene is considered to be deleted. For example, the average CNLOR of ERBB2 may be 0 for a particular cancer type, with a standard deviation of 1.2. A sample is considered to have an ERBB2 amplification if its ERBB2 CNLOR is greater than 0+(2.0*1.2)=2.4. Alternatively, a cancer may have an average CNLOR of TP53 of -0.1, with a standard deviation of 0.8. A sample is considered to have a TP53 deletion if its TP53 CNLOR is less than -0.1-(2.0*0.8)=-1.7.

[0300] At 754, process 750 can remove any samples in the sample set that are not associated with the same cancer type as the model. For example, process 750 can remove lung cancer samples with a squamous cell diagnosis from the sample set if the model is associated with lung adenocarcinoma. At 756, process 750 can label samples as positive or negative samples and / or remove samples from the sample set based on the LORs of the variants, VAFs and copy numbers of each gene in the sample. In some embodiments, process 750 can determine the positive and negative controls using the criteria described in the "Selection of Exemplary Positive and Negative Controls" section above.

[0301] In some embodiments, for a model trained to detect dysregulation in a pathway (e.g., RTK-RAS pathway 1200), a sample can be labeled as a positive control sample only if it contains either a germline or somatic mutation in the DNA of at least one of the genes included in the pathway module included in the pathway. In some embodiments, a sample can be labeled as a negative control only if it does not have any kind of DNA mutation in any of the genes included in the pathway and / or contains only benign or likely benign germline variants in any of the genes in the pathway.

[0302] In some embodiments, for a model trained to detect dysregulation in a pathway module, a sample can be labeled as a positive control sample only if the sample contains either a germline or somatic mutation in the DNA of at least one gene included in the pathway module. In some embodiments, a sample can be labeled as a negative control only if it does not have any type of DNA mutation in any gene included in the module associated with the model. Furthermore, in some embodiments, a negative control can only contain benign or likely benign germline variants in one or more genes across the pathway including the module.

[0303] In some embodiments, for a model trained to detect dysregulation in a single gene included in a pathway module (e.g., RAS module 1210), a sample can be labeled as a positive control sample only if the sample contains a mutation in the DNA of the gene. In some embodiments, a sample can be labeled as a negative control only if the sample does not have any type of DNA mutation in the gene associated with the model and / or contains only benign or likely benign germline variants in genes throughout the pathway that includes the gene.

[0304] Process 750 may use only genetic data for the pathway on which the model is trained or the pathway that includes the module on which the model is trained when determining which samples should be included in the analysis. For example, if training data for a model of the RAF module in the RTK / RAS pathway is being generated, genetic variants in secondary but unconnected oncogenic pathways (e.g., the WNT pathway) will not be considered in the decision to include the sample in the positive or negative control group or to exclude it from the analysis. Furthermore, mutations in other modules in the parent RTK / RAS pathway, such as the RAS module including HRAS, NRAS, and KRAS, will not affect whether the sample is included in the positive control group RAF. Only pathogenic mutations within the module are considered by process 750 for this determination. For example, when generating a perturbation model of either the RAS or RAF submodule, samples with pathogenic mutations in both BRAF and KRAS (either copy number amplification or deletion depending on the gene, as described above) are included as positive controls. Additionally, process 750 can consider only variants in a sample that have a VAF of at least five percent (i.e., >5%), which can help ensure that any variants that have a perturbing effect on a pathway are present to a sufficient extent that the effect is detectable.

[0305] In some embodiments, for process 750 to mark a sample as a positive sample, the sample must have a detected pathogenic or likely pathogenic variant in any gene in the module, if the model is trained on the module, or in any gene in the pathway on which the model is trained, regardless of whether the variant is somatic or germline. In other words, process 750 marks a sample as positive only if the sample has a somatic and / or germline variant in the pathway on which the model is trained or in the module on which the model is trained.

[0306] In some embodiments, for process 750 to label a sample as a negative sample, the sample must not detect any type of somatic mutation in any gene in the pathway (regardless of whether the model is trained on the pathway or module) and must only have benign or likely benign germline variants in the pathway. In some embodiments, a module may interact with multiple pathways, such as the EGFR and ERBB2 modules. In such cases, the sample must not have somatic mutations in any gene in that module to be labeled as a negative sample. These criteria can help ensure that only samples whose perturbation status can be reliably assessed are included in the model generation. Modeling based on patients in the tail of the pathway perturbation distribution provides an interpretable continuous score that can quantify the effect of the VUS on the pathway perturbation of the patient.

[0307] In some embodiments, process 750 can remove any samples with a quality rating below a predefined threshold. The quality rating can reflect the likelihood that an error occurred during the sequencing procedure that affected the associated read. By way of example, and not limitation, the threshold can be derived by evaluating one or more criteria that may result in poor or unreliable sample quality, such as too few reads, poor read quality, too high read overlap rate, the presence of DNA contamination, contamination with other samples, pathogen contamination, and poor read alignment to the genome assembly.

[0308] Process 750 can remove any samples that are not positively or negatively labeled from the sample population. For example, process 750 can remove samples with pathogenic mutations outside the module on which the model is trained.

[0309] In some embodiments, process 750 may terminate if there are not a sufficient number of positive and negative controls. In some embodiments, the process may terminate if there are not at least 16 positive control samples and the ratio of negative controls to negative controls is at least 5%. In this way, process 750 can ensure that models are trained only when adequate data is available.

[0310] At 758, process 750 can output training data for use in training the model. The training data can include positively labeled and negatively labeled samples in the sample set. Process 750 can output the training data to a database (e.g., labeled tumor sample database 400 of FIG. 3) or to a process such as process 690 of FIG. 6G.

[0311] Examples for classifying individual samples are provided below in Tables 3-7. The examples are intended to illustrate how a decision is made as to whether and how a sample is included in model generation, using the applicable criteria described above in connection with process 750.

[0312] The examples in Table 3 are for samples that are considered for inclusion in the ERBB2 submodule. The sample contains sufficient amplification in the ERBB2 gene to be included as a positive control. Although the sample has other variants, these do not exclude the sample from the positive control group, given that only module-level mutations are considered for this decision. [Table 4]

[0313] The example in Table 4 concerns a sample that is considered for inclusion in the RAF submodule of the RTK / RAS parent pathway. The patient does not have a pathogenic or likely pathogenic mutation in the RAF module, and therefore cannot be included in the positive control group. The patient has a pathogenic mutation in KRAS, which is in the RTK / RAS pathway, the parent pathway of the RAF module. Therefore, this patient cannot be included in the negative control group and is completely excluded from model generation. However, this patient can be included as a positive control for the model of RAS submodule perturbation. [Table 5]

[0314] The example in Table 5 is for another sample considered for inclusion in the RAF submodule of the RTK / RAS pathway: this patient has a pathogenic mutation in BRAF, a member of the RAF module, and therefore can be included in the positive control group. [Table 6]

[0315] The example in Table 6 concerns a sample that is considered for inclusion in the TOR submodule of the PI3K pathway. This sample can be included in the positive control group because it has an amplification in RICTOR, a member of the TOR module. The sample also has an amplification in AKT3, but this does not exclude the sample from the positive control group, given that only module-level mutations are considered for this decision. [Table 7]

[0316] The example in Table 7 is for a sample that is considered for inclusion in the PTEN submodule of the PI3K pathway. This sample has a benign germline mutation in PTEN, which is insufficient to include as a positive control or exclude as a negative control sample. Therefore, this sample is a negative control for generating the PTEN module perturbation model. [Table 8]

[0317] (Classifying variants of uncertain significance)

[0318] Variants of unknown significance (VUS) are mutations where it is unknown whether they are driving cancer (pathogenic) or not (benign). A particular database may have thousands of VUS. It is desirable to characterize the VUS effect on the transcriptome to provide evidence for the classification of pathogenic variants.

[0319] FIG. 6I shows an exemplary model of the RTK-RAS and PI3K pathway 760 with multiple modules. As described above, each module can be associated with a model trained to consider the pathway and identify pathogenic dysregulation of the module. If a VUS causes dysregulation in one of the pathway modules (in which case it should be classified as pathogenic), the composite signal of the models associated with the module can identify patients with that VUS as having a score corresponding to dysregulation. The composite signal can be referred to as a meta-pathway score.

[0320] The above approach relies on the assumption that a pathogenic variant has a direct transcriptional or post-transcriptional mechanism that causes dysregulation of the pathway module it contains and / or pathways downstream of that module. For example, as shown in Figure 6J, a VUS in AKT to be classified as pathogenic would cause perturbations in these modules (numbers are examples of dysregulation scores for patients with that VUS in each module):

[0321] To analyze the effect of a VUS, a global dysregulation score can be calculated that takes into account both the module of origin and all downstream modules. Furthermore, pathogenic variants should cause more dysregulation in modules closer to the module of origin than further away, which can be taken into account when calculating the global dysregulation score.

[0322] (Possible confounding factors)

[0323] The VUS classification score may be confounded by other somatic mutations, pathogenic mutations, or VUS mutations in the same gene as the VUS. If there are other potential pathogenic mutations in the same gene as the VUS (including other VUS), these may explain the calculated pathway dysregulation. The VUS classification score may also be confounded by pathogenic mutations in any gene associated with the pathway with the VUS. Any pathway module that has a pathogenic mutation and is downstream of the origin module should have a high dysregulation score regardless of the pathogenicity of the VUS because patients with such pathogenic mutations were used to train the model. The global dysregulation metapathway score takes into account modules downstream of the origin module that include these patients, thus falsely inflating the global dysregulation score. As seen in Figure 6K, the TSC1 module would be expected to have a high dysregulation score regardless of the pathogenicity of the VUS in AKT.

[0324] Modules with pathogenic mutations in another module upstream of them would also be expected to have a high dysregulation score regardless of the pathogenicity of the VUS, and including these patients as well would falsely inflate the overall dysregulation score. As shown in Figure 6L, PTEN pathogenic mutations would be expected to cause higher dysregulation scores in AKT, TSC1, etc., since they are downstream of PTEN.

[0325] Patients with pathogenic mutations in other upstream modules can be excluded from the analysis. However, some classifiers, such as those involving linear models, can allow for the inclusion of mutation status in other genes in the pathway as covariates to account for the contribution of other gene mutation effects to the meta-pathway score while increasing sample size and power of the analysis.

[0326] Mutations in genes outside a given pathway may affect the pathway of interest. To classify VUS in genes outside the pathway, assume that GENE is connected to each module in the pathway in sequence. For example, GENE 762 can be connected to each module included in the RTK-RAS and PI3K pathway 760 shown in Figure 6M.

[0327] For each connection between an additional GENE and each module in the pathway, a global dysregulation score can be calculated as if the GENE were truly connected to the pathway. The GENE can be assumed to be connected to the pathway at the module connection that results in the highest global dysregulation score in the pathway, and then the VUS is evaluated to see if it has a similar signal to a known pathogenic variant.

[0328] Figure 6N shows the distribution of EGFR pathway dysregulation scores for somatic pathogenic mutations in the EGFR and wild-type cohorts in the holdout set. Even though the AUC threshold of 764 separates pathogenic vs. WT patients well, there are still WT patients with high EGFR scores and pathogenic patients with low scores. Even if a VUS is pathogenic, it may not certainly exceed the threshold (or vice versa). Instead of classifying a VUS by looking at every instance of it individually, we can build a probability distribution using the pathway module dysregulation scores of patients with that VUS, and then compare that distribution to the corresponding pathogenic and WT distributions. If the mutation is pathogenic, its probability distribution will be more similar to the pathogenic cohort distribution, and if it does not dysregulate the pathway, it will be more similar to the WT distribution.

[0329] For example, VUS can use the TOR model to generate the scores shown in Figure 6O. The scores can be converted to a probability distribution using Gaussian kernel density estimation, as shown in Figure 6P. Gaussian kernel density estimation constructs a Gaussian curve at each data point and then adds the Gaussian curves together to get the final result. Note that the final distribution will be highest at the points where the data points are most dense.

[0330] The Gaussian KDE also provides some desirable smoothing properties. For example, in the example shown in Figure 6P, it makes the probability distribution nonzero between 0.55 and 0.6, even if there are no data points in that interval. In addition, the Gaussian KDE can model a Gaussian noise model for each data point, which can improve robustness. The Gaussian distribution can also be normalized for differences in VUS sample sizes, since all probability distributions have an area of ​​1.

[0331] To quantify the pathogenicity of this VUS with the TOR module pathway score, the distribution can be compared to the TOR pathogenicity distribution and the TOR WT distribution using Kullback-Leibler Divergence. In general, KLD measures the difference between two probability distributions. Thus, if the VUS distribution is more similar to the pathogenicity distribution than to the WT, the divergence between the VUS distribution and pathogenicity will be smaller than the divergence between the VUS and the WT.

number

number

[0332] However, taking the Kullback-Leibler Divergens in this way may not work when one distribution is more widely spread than the other, for example in Figure 6Q.

[0333] Using the KLD method above means that the VUS distribution is more similar to WT than to pathogenic (p<0.5), even though it is very similar to the middle of the pathogenic distribution. To correct for this, instead of directly comparing the VUS distribution to WT and pathogenic, the VUS distribution can be added to the WT and pathogenic distributions separately, and then the divergence between the new distribution and their respective original distributions can be measured, which allows us to measure the perturbation that the VUS distribution causes when it is added to the other distribution. If the VUS distribution perturbs (i.e., is more similar to) the pathogenic distribution than the WT, our final result (ratioed and normalized as before) gives a value above 0.5. The value in this example is here p=0.62.

[0334] When constructing the reference distributions for pathogenic and WT, only data that were not used to train the model should be used: using the training data to create the reference distributions will skew them towards the respective extremes.

[0335] A generalized approach to testing the effect of VUS on each pathway model can include all individuals in a linear model and test the effect of each VUS mutation on each pathway module score, similar to expression QTL studies. Single variant effects can then be meta-analyzed across each pathway module of interest. Covariates can be used to control for the effect of other potentially pathogenic mutation effects detected on the pathway. The choice of which modules to meta-analyze can be predefined given known pathway gene lists or identified from RNA data (e.g., network graphs).

[0336] For simplicity, we assume that the graph above is completely accurate, i.e., it represents only all true interactions between pathway modules. This means that a VUS in a pathway module will (and only) affect that module, and possibly downstream pathway modules. For example, if there is a pathogenic mutation in AKT, this should cause dysregulation of AKT, TSC1, TSC2, RHEB, TOR, and STK11. Furthermore, the amount of dysregulation should be greater in pathway modules closer to AKT, and therefore dysregulation in each of these pathways will most likely be ranked in the same order.

[0337] Based on this assumption, a metric can be calculated that quantifies the overall impact of dysregulation on a pathway. As an example, assume there is a VUS in AKT. Define v as the pathway module in which the VUS is located and M as the pathway modules downstream of v, i.e., the pathway module that has the VUS and all pathway modules downstream of it. Then M={AKT, TSC1, TSC2, RHEB, TOR, STK11}. Each pathway module model m in M ​​is scaled from 0 to 1 and assigned a specific dysregulation score DS defined using the Kullback-Leibler Divergence from the section above. m One metric that can be used to quantify the overall effect of dysregulation is Σ m∈M DS m This is the sum of all posterior stream dysregulation scores in M.

[0338] A distance function is introduced to account for the fact that pathogenic mutations affect the pathway modules closest to v more than any other, and affect v more than any other pathway modules. d(m,v)=1+(shortest distance between m and the path module containing the VUS).

[0339] In the present example (v=AKT), d(AKT,v)=1, d(TSC1 / 2,v)=2, d(RHEB,v)=3, etc. To weight the dysregulation scores according to their closeness to v, we use the weighted score

number

[0340] T v may not be normalized for the number of pathway models in M. For example, a pathway may have two VUSs, one in the RAS and one in the RAF. Then,

number

number

number

[0341] The final metrics that can be used to calculate a global dysregulation score are as follows:

number

[0342] Example: VUS in AKT Assume that the VUS under consideration is AKT, and that AKT and its downstream pathways have the dysregulation scores shown in FIG. 6R.

number

[0343] (VUS cohort selection) For any VUS, patients selected for the cohort used to measure its pathogenicity should fulfill two characteristics to make the VUS signal as clear as possible.

[0344] 1) There must be no other somatic mutations, pathogenic mutations or VUS mutations in the VUS gene;

[0345] 2) They should not have pathogenic mutations in any of the pathway modules that link to the pathway module containing the VUS.

[0346] For the first property, if the patient has another somatic, pathogenic, or VUS mutation in the same gene, the perturbation of the downstream pathway module may be due to that mutation and not the VUS of interest.

[0347] For the second trait, if a pathway module had the same score as the VUS in the AKT example above, but TSC1 had a pathogenic mutation as shown in Figure 6S, the high TSC1 score here would be more likely due to the presence of a pathogenic mutation than a VUS in AKT, since the TSC1 model was trained to have a high score for patients with pathogenic mutations in TSC1, thus confounding the perturbation score.

[0348] As another example, suppose there is a pathogenic mutation upstream of AKT, for example in PTEN, as shown in Figure 6T. In that case, dysregulation in AKT and its downstream pathway module scores may be due to a pathogenic mutation in PTEN instead of a VUS in AKT. Again, this would confound the results.

[0349] Patients in the cohort of a VUS of interest should not have pathogenic mutations in any pathway module upstream or downstream of the pathway module containing the VUS of interest. However, this filter is still not strict enough. For example, assume that we are considering a VUS in ERBB2. Given the current rules, patients without pathogenic mutations in the intermediate pathways upstream and downstream of ERBB2 will be selected. Now, say that the PIK3C dysregulation score is high, but there are also pathogenic mutations in EGFR and PTEN, as shown in Figure 6U. The high PIK3C score is likely caused by pathogenic mutations in EGFR and PTEN. Therefore, it is also necessary to exclude patients with pathogenic mutations in any pathway module upstream of any pathway module downstream of the pathway module containing the VUS of interest.

[0350] In summary, the method for determining the pathogenicity of a VUS in a gene in a pathway includes finding a set of patients who have no other somatic, pathogenic, or VUS mutations in the same gene as the VUS, and no pathogenic mutations in any pathway modules upstream of the pathway module containing the VUS, or upstream of any pathway modules downstream of the pathway containing the VUS; generating a probability distribution of the VUS cohort for each of the pathway module models containing the pathway module containing the VUS and the pathway module models containing the downstream pathway modules; calculating the ratio between the similarity of the VUS cohort distribution and the pathogenic distribution and the VUS and WT distributions for each model using Kullback-Leibler Divergence; and calculating a global dysregulation score G by performing a weighted average of the VUS containing module and its downstream modules. v and calculating

[0351] Here, a technique is presented to extend VUS pathogenicity determination to off-pathway genes. The above method can be extended to genes that have known connections to pathways but do not have models trained for them, such as NF1, which connects to the RAS pathway as shown in Figure 6V.

[0352] A method that may be called a whole gene method for classifying VUS in a gene without a trained model involves finding patients without a trained model who have no other somatic, pathogenic, or VUS mutations in the gene (e.g., NF1) and no pathogenic mutations upstream or downstream (e.g., in EGFR, RAS, or RAF), calculating a dysregulation score of this cohort for the downstream modules (e.g., RAS and RAF), and combining the dysregulation scores of this cohort for the downstream modules to produce an overall dysregulation score G v (e.g., RAS and RAF dysregulation scores).

[0353] In particular, how genes are connected to pathways is essential for every part of this process. To properly assess VUS, several metrics need to be known, including knowing which meta-pathways a patient needs to be free of pathogenic mutations, knowing which meta-pathways to calculate the dysregulation score, and knowing how to weight the dysregulation scores to calculate an overall dysregulation score. This cannot be known for genes whose pathway connections are unknown.

[0354] To solve the above problem of VUS in a gene GENE with no known pathway connection, we can calculate all possible global dysregulation scores for GENE (e.g., GENE 762 in Figure 6M) by assuming that the GENE is directly connected to each pathway module in turn.

[0355] In one iteration, we assume that GENE is connected to AKT as shown in Figure 6W.

[0356] The global dysregulation score of VUS in GENE can be calculated in exactly the same way as it was calculated for NF1 connected to RAS. First, a cohort is generated consisting of patients who do not have any other somatic mutations, pathogenic mutations or VUS mutations in GENE and do not have any pathogenic mutations in {EGFR, ERBB2, PTEN, PIK3C, AKT, TSC1 / 2, RHEB, TOR, STK11}. Then, the dysregulation score can be calculated for {AKT, TSC1 / 2, RHEB, TOR, STK11}. Finally, the global dysregulation score can be calculated by weighting the dysregulation score of {AKT, TSC1 / 2, RHEB, TOR, STK11} using the distance of each module from GENE.

[0357] In another iteration, GENE is assumed to be connected to RAS as shown in Figure 6X. The step of finding the global dysregulation score in this case may include generating a cohort composed of patients with no other somatic, pathogenic or VUS mutations in GENE and no pathogenic mutations in {EGFR, RAS, RAF}, calculating the dysregulation score of {RAS, RAF}, and calculating the global dysregulation score by weighting the dysregulation score of {RAS, RAF} using the distance from GENE.

[0358] FIG. 6Y illustrates an exemplary data frame that can be generated using the methods described above.

[0359] (Analysis of the results of all gene analysis) Figure 6Z shows an exemplary histogram of all global dysregulation scores after analyzing all genes (filtering VUS in cohorts >5). A potentially pathogenic VUS threshold 766 is indicated with a perturbation score value of 0.25.

[0360] To test the validity of this method, we calculated the perturbation scores for known NF1 pathogenic mutations using the above all-gene method. Given that NF1 is connected to the RAS pathway module, it is expected that these mutations will result in a higher overall dysregulation score when tested as connected to the RTK_RAS pathway than when tested as connected to the PI3K pathway. Only two mutations in NF1 had cohorts >1 for all possible metabolic pathways, and their results are shown in Figure 7A and Figure 7B, respectively.

[0361] These NF1 mutations, when tested connected to the pathway module of RTK_RAS, result in a higher global dysregulation score than PI3K, suggesting that the method works as expected. It is important to recognize that even the tests with the highest perturbation scores for NF1 LOF are below the proposed p=0.25 cutoff derived to look for tests for all genes, and many of the perturbation scores for NF1 c.3198-2A>G are above the p=0.25 cutoff, even when NF1 is connected to the PI3K pathway. This may suggest that VUS classification should be done at the per-mutation level as well as at the global level.

[0362] 7C illustrates an example process 702 that can use the trained route engine to generate a route perturbation score. Process 702 can be implemented as computer readable instructions on one or more memories or other non-transitory computer readable media and executed by one or more processors in communication with the one or more memories or media. In some embodiments, process 702 can be implemented as computer readable instructions on memory 222 and / or memory 262 and executed by processor 214 and / or processor 254.

[0363] At 705, process 702 can receive transcriptome data. The transcriptome data can include one or more transcriptome value sets. In one example, each transcriptome value set can be a file with a tabular format where each column represents a gene and includes a normalized expression value associated with the gene. In another example, a transcriptome value set can be a file with a tabular format where each column represents a gene and includes a raw expression value (e.g., read counts or copies detected by a next generation sequencer or other genetic analyzer) associated with the gene. The transcriptome value sets can be associated with a specimen and / or a patient.

[0364] A transcriptome may have an associated cancer type that can determine which pathway engine is used to generate a pathway perturbation score for the transcriptome. For example, one or more pathway engines associated with the same cancer type as the transcriptome can be selected. If a transcriptome does not have an associated cancer type or the associated cancer type may be incorrect, the cancer type can be determined for the transcriptome, for example, by analyzing the histopathology slide associated with the transcriptome, or by analyzing the transcriptome and any associated data, for example, as described in U.S. Patent Application Publication No. 62 / 855,750, entitled "Systems and Methods for Multi-label Cancer Classification," filed May 31, 2019, and incorporated herein by reference. One example of a transcriptome that does not have an associated cancer type or may have an incorrect associated cancer type is a transcriptome associated with a tumor of unknown origin, a metastatic tumor, or an inaccurately labeled cancer sample.

[0365] In addition to transcriptome data, process 702 can receive supplemental data including DNA variant data, methylation data, cancer type, and / or proteomic data. All of the data received at 705 may be included in data input 100 described above.

[0366] At 708, the process 702 can provide the transcriptome data to one or more trained pathway engines. The pathway engines can be included in the computing device 210 and can include trained pathway engines. Based on the type of data received at 705, the process 702 can determine which pathway engines to provide the transcriptome data to along with any supplemental data. The transcriptome data can have one or more associated cancer types.

[0367] Process 702 may provide transcriptome data to any pathway engines associated with pathways that may be associated with cancer types. Some pathway engines may be configured to accept only transcriptome data, while others may also accept supplemental data including DNA variant data, methylation data, cancer types and / or proteomic data. Process 702 may provide only transcriptome data to certain pathway engines and provide transcriptome data and supplemental data (e.g., DNA variant data) to other pathway engines. Process 702 may provide data applicable to as many related pathway engines as possible. Trained pathway engines may include engines that accept the same input but are trained on different training datasets.

[0368] At 710, process 702 can receive one or more pathway perturbation scores from one or more trained pathway engines. Each trained pathway engine can generate a pathway perturbation score for each transcriptome value set (and any supplemental data). The pathway perturbation scores can be numerical, graded score outputs, and / or qualitative readouts.

[0369] The trained pathway engine can generate a pathway perturbation score by simultaneously comparing the expression level of each DEG in the transcriptome value set to a range of expected expression levels of that DEG in a positive control and a range of expected expression levels of that DEG in a negative control. The pathway perturbation score can reflect the degree to which a transcriptome value set is similar to the dysregulated positive control transcriptome value set compared to the wild-type negative control transcriptome value set.

[0370] In various embodiments, the system and method generate a graded score output that predicts the degree of pathway perturbation (e.g., a number ranging from negative 2 to 2, or a number ranging from 0 to 1). In such embodiments, a statistical threshold can be generated to generate a qualitative readout of the pathway perturbation (e.g., perturbed or not perturbed, or additional classes such as highly perturbed, mildly perturbed, not perturbed, etc.). This qualitative readout can be a clinician-friendly indicator of pathway perturbation (e.g., "high", "medium", "low"). In one example, the qualitative readout can be determined by comparing the graded score output to a threshold. For example, all graded score outputs below 0 can be labeled as not perturbed, and all graded score outputs above 0 can be labeled as perturbed. In this example, 0 is the selected cutoff threshold. In one example, the threshold can be selected by selecting a threshold that maximizes the F1 score, as described above. In one example, the pathway engine can output a normalized pathway perturbation score ranging from 0 to 1. A "high" route disruption score may include a route disruption score of at least 0.8, a "medium" route disruption score may include a route disruption score of at least 0.6, and all route disruption scores below 0.6 may be considered "low."

[0371] The trained pathway engine can output a score for each module included in a pathway associated with the trained pathway engine. The trained pathway engine can include a trained model (e.g., a trained linear regression model) for each module in the pathway. The score for each module can indicate dysregulation in the associated module. Process 702 can grade each score generated by the model into a qualitative score (e.g., "high," "medium," "low") as described above.

[0372] The pathway perturbation scores can be added to a data set for analysis of pathway perturbation scores in a larger population of specimens. The pathway perturbation scores can be used to determine the confidence in predicting a particular treatment response based on clinical data and / or treatment response data associated with other generated pathway perturbation scores. For example, process 702 can compare the pathway perturbation scores generated by the pathway engine for each specimen in a group of specimens with the clinical data and / or treatment response data associated with the specimen. The pathway perturbation score(s) can be used to develop models for predicting patient outcomes / treatment responses.

[0373] The pathway perturbation score can be used to classify variants of unknown significance (VUS) based on an observed correlation between the pathway perturbation score generated by the systems and methods disclosed herein and the detected VUS in the sample, which predicts the perturbation status of the pathway, especially when no pathogenic variants were detected in the sample. Process 710 can include determining a global dysregulation score using formula (3) above. Process 710 can include performing the whole gene method described above to generate a global dysregulation score.

[0374] Correlation observations may utilize a database of specimen-associated variant calls, which may include all variants detected in a patient, regardless of whether they have clinical significance (i.e., all VUS).

[0375] The pathway perturbation score can be used to rank the therapeutic suitability for a specimen based on the observed correlation between the pathway perturbation score estimated by the systems and methods disclosed herein and clinical response data, particularly data related to the response of a patient or organoid to a treatment.In one example, the systems and methods first robustly correlate pathway perturbation scores with therapeutic response, accounting for several covariates.

[0376] At 715, process 702 can generate a meta-route representation. Exemplary meta-route representations are shown in Figures 12A-12E and described below. The meta-route representation can include one or more routes that can be color coded or otherwise shaded based on the route perturbation scores and / or supplemental data.

[0377] At 718, process 702 can cause the meta-pathway representation to be output to a display (eg, display 290, display 256, and / or display 216) and / or memory (eg, memory 222 and / or memory 262).

[0378] At 720, process 702 can generate any ensemble pathway perturbation score based on the multiple pathway perturbation score outputs. The ensemble model can receive pathway perturbation score outputs from at least two trained pathway engines that are associated with a common pathway and accept the same differentially expressed genes, but trained with different training datasets. Process 702 can provide the pathway perturbation score output to any ensemble model. The ensemble model can convert the pathway perturbation scores into an ensemble pathway score by summing weighted scores, where the weights are determined by training the ensemble model with types of data related to pathway perturbation scores and cancer characteristics, including clinical response data, cancer stage status, consensus molecular subtype (CMS) classification, etc. The ensemble pathway score can reflect the overall cell state and / or biological interactions between at least two gene sets used to train the model. Process 702 can receive the ensemble pathway perturbation score from the ensemble model.

[0379] The ensemble pathway perturbation scores may be added to a data set for analysis of pathway perturbation scores in a larger population of specimens. The ensemble pathway perturbation scores may be used to determine the confidence in predicting a particular treatment response based on the clinical data and / or treatment response data associated with the ensemble pathway perturbation scores generated by the systems and methods, for example, by comparing the ensemble pathway perturbation scores generated by the pathway engine 200n for each specimen in a group of specimens with the clinical data and / or treatment response data associated with the specimen. The ensemble pathway perturbation scores may be used in developing models for predicting patient outcomes / treatment responses.

[0380] The ensemble pathway perturbation score can be used to classify variants of unknown significance (VUS) based on observed correlations between the ensemble pathway perturbation scores generated by the systems and methods disclosed herein and VUS detected in a sample, which predict the perturbation status of the pathway, particularly when no pathogenic variants were detected in the sample.

[0381] Correlation observations may utilize a database of specimen-associated variant calls, which may include all variants detected in a patient, regardless of whether they have clinical significance (i.e., all VUS).

[0382] At 725, the process 702 can output the ensemble pathway perturbation score to a display (e.g., display 290, display 256, and / or display 216) and / or memory (e.g., memory 222 and / or memory 262). The ensemble pathway perturbation score can be used to rank treatment suitability for a specimen based on observed correlations between pathway perturbation scores estimated by the systems and methods disclosed herein and clinical response data, particularly data related to patient or organoid response to treatment. In one example, the systems and methods first robustly correlate the ensemble pathway perturbation score with treatment response, accounting for several covariates.

[0383] At 730, process 702 can generate a pathway perturbation report based on any pathway perturbation scores received at 710. Process 702 can generate a pathway perturbation report further based on the meta-pathway delineation data generated at 715 and / or any ensemble pathway perturbation scores generated at 720. The pathway perturbation report can communicate results from 710 and / or 720, including pathway perturbation scores and / or ensemble pathway perturbation scores generated for patient specimens or organoids associated with the set of transcript values. In one example, the report can include one or more pathway perturbation scores and / or pathway score relationships (e.g., as shown in Figures 10A-H, 11A-D, 12A-E, 22, 23, 24, and 25, described below). For example, if the pathway perturbation scores are -0.5 and -0.5 (one score for each of the two treatable arms or branches of the pathway), reporting the scores for each arm of the pathway may be more informative than an ensemble pathway score of -1 for the entire pathway.

[0384] The pathway report may also include prognostic predictions, including the likelihood of drug sensitivity to drugs targeting the cancer cells in the original specimen, particularly the pathways of interest that are reported to be activated or repressed, and predicted patient survival and / or progression-free survival. The pathway report may include schematics or depictions of cellular pathways or gene sets of interest, and / or meta-pathways (see Figures 10A-H, 11A-D, and / or 12A-E). The pathway report may include citations to references, particularly those related to the pathways of interest and / or therapies targeting the pathways of interest. The numerical values ​​of the pathway score and / or ensemble pathway score may determine which treatments and / or clinical trials match the specimen and are presented in the pathway perturbation report.

[0385] The report may be digital (e.g., available as a digital file such as PDF or JPG or accessible via a user interface such as a portal or website) or hard copy (e.g., printed on paper).

[0386] In one example, for each patient sample in the population undergoing RNA sequencing, their normalized RNA data and, if applicable, the ssGSEA scores of the relevant pathways are subjected to at least one pathway engine to obtain a pathway perturbation score as described above. Patients can receive an indication in the report whether their cancer has any activated or repressed cellular pathways, and if so, can be matched with specific therapies or clinical trials, particularly those with inclusion criteria related to activated or repressed pathways.

[0387] In some embodiments, the pathway perturbation report may include information about which genes in the pathway may cause pathway perturbation as indicated by the pathway perturbation score, even if there are no measurable mutations in the pathway. For example, FIG. 11A shows a pathway graph that may be included in a pathway perturbation report for the PI3K pathway. The PI3K pathway was not detected to have a pathogenic mutation, but a high pathway perturbation score was generated by the pathway engine (e.g., in steps 708 and 710), indicating pathway perturbation. The mutation that causes the high pathway perturbation score (e.g., a pathway perturbation score of 0.85 from the pathway engine that outputs a normalized pathway perturbation score between 0 and 1) may be unknown, but the level of pathway perturbation may be inferred by the pathway perturbation score. In this example, a treatment designed to target CRTC2 may be fitted. The report may indicate that the CRTC2 gene may be targeted by encircling the CRTC2 gene in the pathway, color-coding the CRTC2 gene, or otherwise visually indicating that the CRTC2 gene may be targeted. The pathway perturbation report may include information or links to information (e.g., URL links to NIG web pages) regarding one or more treatments that can be used to target the CRTC2 gene. The pathway perturbation report may include information or links to information regarding clinical trials that can be matched based on study inclusion and / or exclusion criteria. Currently, clinical trials may require pathogenic DNA mutations in the PI3K pathway detected in patients for enrollment, but it is contemplated that clinical trials may be matched to patients based on the pathway perturbation score generated by the pathway engine.

[0388] A particular pathway may have multiple targetable genes or modules. For example, FIG. 22 shows an example of a pathway perturbation report that includes a subset of the MAPK pathway. The pathway perturbation report may include information about where in the MAPK pathway the patient can be treated. The patient may have been determined to have a high pathway perturbation score for the MAPK pathway using one or more pathway engines. Process 702 may determine one or more therapies that can be used to treat the patient. The pathway perturbation report may include one or more treatments that can be used to target one or more genes and / or modules in the MAPK pathway. Additionally, the treatments may be marked (e.g., visually) as potentially more or less effective based on any detected mutations in the pathway (e.g., DNA mutations in the pathway), as well as based on information about the patient, such as treatment history, including any treatments the patient has received.

[0389] The patient may have a detectable mutation in the RAS module (exemplified by the KRAS mutation) as shown in FIG. 22. While certain therapies can be used to treat the RAS module, these therapies may not be approved (e.g., FDA approved) and therefore cannot be used as a treatment unless they are in trials. Furthermore, therapies applied to modules above the RAS module may not treat the mutation at the RAS module level. Other therapies occurring below the RAS module may potentially be less effective or less usable because the treatments are experimental and / or the patient has already been treated without positive results. Thus, potential treatments in the EGFR and RAS modules may be marked with a different color or have a different shading than the other treatments, or may be identified as potentially less effective or less usable treatments. Process 702 may determine one or more treatments that may be more effective for the patient, for example, by determining approved treatments for modules with known mutations, in this example downstream of the RAS module.

[0390] Additionally, process 702 can determine more treatments based on whether treatments applicable to modules downstream of the module with the known mutation have been effective for similar patients. More specifically, the process can compare the transcriptome data, any supplemental data including DNA variant data, methylation data, cancer type, and / or proteomic data received in step 705, and / or any pathway perturbation scores generated for the patient to data for similar patients. Process 702 can receive data for similar patients from one or more databases, such as databases 500, 600, 700 described above. Process 702 can compare one or more pathway perturbation scores received in 710, the transcriptome data, and / or any supplemental data received in step 705 to a database of results from many samples. Process 702 can identify a group of specimens that are most similar to the patient based on the generated pathway scores by identifying which patient's pathway perturbation scores are above / below a threshold identified as indicative of pathway perturbation in the set of other specimens or which scores fall in a quantile (e.g., top 5 quantile) of the scores of the set of other specimens. Process 702 can determine which specimens have transcriptomic data that cluster with the patient when subjected to a dimensionality reduction algorithm (e.g., uniform manifold approximation and projection (UMAP) or principal component analysis (PCA)) and plotted on a two-dimensional Cartesian grid. Process 702 can also compare supplemental data associated with the patient to supplemental data associated with the specimen. Process 702 can determine that specimens that have supplemental data within a predetermined threshold of the patient's supplemental data are similar to the patient.

[0391] In some embodiments, process 702 may include portions of the methods and systems of U.S. patent application Ser. No. 62 / 786,739, entitled "A Method and Process for Predicting and Analyzing Patient Cohort Response," filed 12 / 31 / 18. In step 730, process 702 may compare the data received in step 705 to data in a database of results disclosed in U.S. patent application Ser. No. 62 / 786,739.

[0392] After process 702 determines specimens similar to the patient, process 702 can determine which treatment had the greatest positive effect on the specimen and include that treatment in the pathway perturbation report. In some embodiments, process 702 can determine which treatment was most effective based on information from treatment response database 600.

[0393] Still referring to FIG. 7C, at 735, process 702 can output the pathway perturbation report to at least one of a display or a memory. For example, process 702 can output the pathway perturbation report to a display (e.g., display 290, display 256, and / or display 216) for a user to view. Thus, process 702 can display the pathway perturbation report. As another example, process 702 can output the pathway perturbation report to a memory (e.g., memory 222 and / or memory 262) for storage. In some embodiments, at 735, process 735 can print out the pathway perturbation report. Process 702 can distribute the pathway perturbation report to a physician, medical professional, patient, pharmaceutical designer or manufacturer, or organoid culturing laboratory, particularly to guide treatment decisions and clinical trial or experiment design.

[0394] These systems and methods described above (e.g., system 10 and / or processes 502, 602, 630, 650, 660, 670, 750, and / or 702) can detect more patients with activated or inhibited pathways and match them with potentially beneficial treatments and clinical trials. The patient report generator 800 described above can include and / or perform any number of processes 502, 602, 630, 650, 660, 670, 750, and / or 702.

[0395] Clinicians can benefit from these systems and methods by being able to make more informed treatment choices based on molecular evidence beyond DNA mutation profiles. Patients can also benefit in that they are more likely to respond to treatments selected based on the multiple orthogonal lines of evidence provided by these systems and methods. Pharmaceutical companies can also benefit by using the systems and methods to select patients with specific pathway perturbations for inclusion in relevant clinical trials.

[0396] The systems and methods can help provide insights in clinical and / or pathway perturbation reports, scientific rationale underlying matched treatments, and / or matched clinical trials, and clinically actionable molecular evidence that is validated and driven by the context of oncogenic pathways / networks. Pathway information can also serve as "priors" and / or features in statistical models to correlate integrated -omic and imaging data with treatment and outcomes.

[0397] The systems and methods can facilitate the discovery of novel biomarkers, diagnostic and / or prognostic signatures of pathways (including therapeutically targeted pathways) and enhance the ability to match treatments in reports.

[0398] In various embodiments, systems and methods include a method for detecting cellular pathway dysregulation in a sample that includes receiving a set of genetic data derived from and / or otherwise associated with the sample, and analyzing the set of genetic data to estimate a likelihood of dysregulation for a cellular pathway of interest (a pathway perturbation score).

[0399] A pathway of interest can be any set of genes. A set of genes can represent a cellular pathway. A set of genes can have gene products that interact with each other within a cell during cellular activity. A pathway of interest can be a well-defined cellular pathway (e.g., the RAS / RTK or PI3K pathway). A pathway of interest can be a TCGA-curated pathway.

[0400] Analyzing the genetic dataset may include providing at least a portion of the genetic data to one or more pathway dysregulation engines and receiving results from each pathway dysregulation engine reflecting the likelihood of dysregulation in a cellular pathway. The pathway dysregulation engines may be trained with a training dataset including a training RNA dataset each associated with at least one dysregulation indicator. Each pathway dysregulation engine may be specific to one cellular pathway, and the dysregulation indicators used to train the pathway dysregulation engine may be associated with the cellular pathway.

[0401] The genetic data includes RNA data, and may further include DNA data and protein data.

[0402] The specimen may be a cancer specimen from a human patient or may be an organoid (e.g., an organoid derived from a human cancer specimen).

[0403] The likelihood of dysregulation may be a numerical value or a qualitative indicator. The method may further comprise comparing the likelihood of dysregulation to a threshold to determine a qualitative indicator of the analyte.

[0404] The method may further include estimating many dysregulation likelihoods (e.g., one for each of many cellular pathways of interest) and combining the dysregulation likelihoods to calculate an overall pathway perturbation score, or reporting each pathway perturbation score and optionally reporting relationships between the pathway perturbation scores (e.g., by reporting biological interactions between pathways or pathway portions associated with each pathway perturbation score).

[0405] The method may further include correlating the dysregulation likelihood label or value with the protein expression level and predicting the protein expression level for the sample.

[0406] The method may further include detecting a variant of unknown significance in the set of genetic data and determining that the variant is pathogenic based on the likelihood of dysregulation.

[0407] These systems and methods may include a method of prescribing a treatment that includes receiving a dysregulation likelihood and prescribing a treatment to a patient from whom the specimen was derived based on the dysregulation likelihood.

[0408] These systems and methods may include methods for designing experiments to test therapeutic response in organoids, including receiving a potential for dysregulation of the organoids and, based on the potential for dysregulation, suggesting that the organoids be monitored following exposure to the treatment.

[0409] The systems and methods may include a method of matching a patient to clinical trials, the method including receiving a dysregulated likelihood of a sample from the patient and matching to at least one clinical trial based on the dysregulated likelihood, the method may further include reporting a list of matched clinical trials to the patient or a medical professional caring for the patient.

[0410] These systems and methods may include a method of designing a clinical trial that includes analyzing clinical data for an association between likelihood of dysregulation and response to at least one treatment, and proposing a study of the response to at least one treatment in each of a plurality of patients having likelihood of dysregulation.

[0411] These systems and methods may include a medical device that receives a set of genetic data and detects cellular pathway dysregulation, as described above. In one example, the medical device may include a genetic analysis system and / or a laboratory developed test.

[0412] These systems and methods may include methods of sequencing a cancer specimen, including generating a set of genetic data as described above and detecting cellular pathway dysregulation.

[0413] These systems and methods may include a cloud-based information processing system that receives a set of genetic data and detects cellular pathway dysregulation, as described above.

[0414] 8A-8D collectively show exemplary flow charts of certain methods that can be used to analyze pathway perturbation conditions based on RNA data.

[0415] FIG. 8A shows a pie chart of a cancer of interest. In one example, patients with a particular cancer type are selected (FIG. 8A, one area of ​​the pie chart), and all relevant mutation data for the pathway of interest is obtained, for example, using oncogenic signaling pathways defined by The Cancer Genome Atlas (TCGA) consortium. The mutation data is used to define a set of patients with known pathway perturbations for all members of the pathway (e.g., KRAS G12V mutation in the RAS / RTK pathway, considered as a "positive control") and patients who are wild type (WT) ("negative control"). FIG. 8B shows a pie chart that subsets the selected cancer type by mutation status.

[0416] Figure 8C shows various graphs of differentially expressed genes (DEGs) between groups that can be determined using edgeR, a publicly available package in the R software environment. When applicable, single sample gene set enrichment analysis (ssGSEA) pathway scores are generated for all samples for all relevant pathways. (Figure 8C).

[0417] 8D shows validation results for a logistic regression model trained according to the above-described process 502. Path engine 200n cross-validation is performed according to the above-described process 602.

[0418] Once the final alpha parameter value is determined, the final alpha parameter value can be used to train a final pathway engine (eg, pathway engine 200n) using all samples.

[0419] 9A and 9B collectively display example outputs of particular methods that may be used to test systems and methods in any route engine 200n validation process, as described in FIGS. 6B and 6E, respectively.

[0420] In some embodiments, the pathway engine 200n is validated using publicly available external TCGA data to ensure that the systems and methods have biological validity and that predictive performance is not dependent on specific features of the training dataset.

[0421] In the first step of validation, as described in process 602, TCGA RNA mutation data for the cancer type of interest can be collected and subdivided into positive and negative control samples, as was done with the training data.

[0422] Figure 9A shows an example of validation results using an external data set. All samples are run through the training pathway engine 200n and the output of the positive and negative controls is compared. The significant difference between the scores associated with these groups in the same direction as the training data is evidence of the robustness and generalizability of the pathway engine 200n (Figure 9A).

[0423] FIG. 9B shows an example of biological validation results using protein activation data. Although detectable at the transcriptional level, pathway perturbation / ground truth of perturbation may be defined as the protein state of the pathway's effectors, i.e., the levels of these proteins and / or their activation as indicated by phosphorylation state. For example, RAS / RTK activation can be quantified by the levels of phosphorylated downstream effector kinases MEK, MAPK1, MAP2K2, etc. The degree of correlation between the output of pathway engine 200n and measures of protein activation was determined for TCGA patients as described in 654, with a strong correlation indicating that pathway engine 200n is biologically significant (FIG. 9B).

[0424] As described herein, some embodiments relate to methods and systems for generating and presenting diagnostic and / or treatment data, including fit for clinical trials, to a physician based on patient information, such as genetic, imaging, and clinical information, as described above. In some embodiments, the data provided to the physician may be in the form of a report document presented digitally or in hard copy. In some embodiments, the report includes information such as, but not limited to, an easy to understand stylized visual depiction of the diagnosis and / or treatment pathway in question, the identity of any relevant clinical trials, eligibility criteria for clinical trials or administration of specific therapeutics or combinations of therapeutics, and a treatments section providing additional information regarding any identified treatments.

[0425] 10A-10I collectively show examples of pathway perturbation reports generated at 730 in FIG. 7C, specifically for the MAPK (RAS) pathway. One aspect of the utility of the described embodiments comes from the possibility of communicating treatment options to a physician for a particular patient's cancer condition. That is, for a given cancer condition, there may be a variety of effective or potentially effective treatments (therapies) that target one or more elements in the pathway (i.e., exerting a biological effect on the pathway). For example, various treatment options for KRAS gain-of-function mutations target the ERK module (e.g., ERK inhibitors), the MEK module (e.g., MEK inhibitors), the RAF module (e.g., RAF inhibitors), etc. Thus, even for a particular mutation or pathogen (which may be indicated in a diagnostic pathway), there may be a variety of treatment options, and the report may include a depiction of different effective or potentially effective treatments.

[0426] FIG. 10A shows an example of a pathway perturbation report generated for a hidden responder who did not detect a pathogenic mutation in the RAS pathway but has a high pathway perturbation score generated by the pathway engine 200n. The mutation causing the high pathway perturbation score may be unknown, but the level of pathway perturbation may be inferred by the pathway perturbation score. A therapy that inhibits MEK or ERK may be matched to this patient. A clinical trial may be matched based on the inclusion and / or exclusion criteria of the trial. Currently, clinical trials may require a pathogenic DNA mutation detected in the patient for enrollment, but in the future, clinical trials may be matched to patients based on the pathway perturbation score generated by the pathway engine 200n. In some embodiments, eligibility criteria are added to the report, for example as shown in FIG. 10I. Each treatment may have eligibility criteria related to the efficacy of the treatment and / or to participation in the trial in the case of a clinical trial. Eligibility criteria may include a cancer diagnosis (e.g., type of cancer, cancer stage, type of mutation, presence and / or absence of other mutations), geographic location of the patient, age of the patient, other health conditions, etc. The eligibility criteria can be stored in the database as metadata associated with each mutation or pathogen associated with each therapeutic and / or diagnostic pathway. By way of example and not limitation, the eligibility criteria for the report shown in FIG. 10B can be as follows:

[0427] Eligibility Criteria: a.Diagnosis: pancreatic adenocarcinoma; b. KRAS gain-of-function mutation; c. Clinical trial NCT03051035 matches patient reports; d. No actionable mutations other than TP53 or SMAD4 are present.

[0428] In various embodiments, such as the example provided in Figure 10B, these pathway reports may be generated for patients with pancreatic adenocarcinoma, cancers with KRAS gain-of-function mutations, and no other actionable mutations other than TP53 or SMAD4. Clinical trials for therapies targeting BRAF, MEK, and / or ERK may be matched to the patient reports.

[0429] 11A-11E collectively show examples of pathway perturbation reports generated in 730 of FIG. 7C, specifically for the PI3K pathway.

[0430] FIG. 11A shows an example of a pathway perturbation report generated for a hidden responder who has no detected pathogenic mutations in the PI3K pathway but has a high pathway perturbation score generated by the pathway engine 200n. The mutations causing the high pathway perturbation score may be unknown, but the level of pathway perturbation may be inferred by the pathway perturbation score. In this example, a treatment designed to target CRTC2 may be matched. PD-L1 inhibitors may be contraindicated in this example due to studies showing that PD-L1 inhibitors may be less effective for patients with STK11 mutations. Clinical trials may be matched based on the inclusion and / or exclusion criteria of the trial. Currently, clinical trials may require pathogenic DNA mutations in the PI3K pathway detected in patients for enrollment, but it is contemplated that clinical trials may be matched to patients based on the pathway perturbation score generated by the pathway engine 200n.

[0431] In Figures 11B and 11C, the patient receiving the pathway report may be HER2 positive (eg, HER2 status may be determined by FISH, IHC or NGS).

[0432] In FIG. 11D, the patient's HER2 status may be unknown.

[0433] In various embodiments, these pathway reports can be generated for patients with breast cancer and PI3K gain-of-function mutations. Clinical trials for therapies targeting PIK3CA, AKT, and / or mTOR can be matched with patient reports.

[0434] In some embodiments, a treatment section can be added to any report. Such information can be included, for example, to enhance any treatment information provided in the pathway diagram or to add additional treatment information generally related to a disease state (see, e.g., FIG. 11E).

[0435] Figures 12A, 12B, 12C, 12D, 12E, and 12F summarize the results of metapathway analysis of patient transcriptomes using the systems and methods disclosed herein (see Example 6).

[0436] Figures 12A, 12B, 12C, 12D, 12E and 12F each show a cellular pathway where proteins in the pathway are represented by polygons, arrows indicate activation of one protein by another and "T" shaped lines indicate inhibition of one protein by another.

[0437] Each polygon in a pathway represents a class of genes (e.g., RAS genes including KRAS, NRAS, and HRAS). In this analysis, a pathway engine was trained for each group of genes (each represented by a polygon in each of Figures 14A-14F, as described in process 502, where all positive controls had at least one mutation in a gene of the gene class associated with the polygon and all negative controls were wild type for all genes in the pathway). Each trained pathway engine 200 was then used to analyze the transcriptome associated with a patient to generate a pathway activity score, as described in Figure 7C.

[0438] If a polygon is colored blue, the path engine 200 associated with that polygon has generated a path activity score that indicates no disturbance. If white, the path engine 200 associated with that polygon has generated a medium path disturbance score that indicates the path may be disturbed. If red, the path engine 200 associated with that polygon has generated a path disturbance score that indicates the path has been disturbed.

[0439] In another example, instead of or in addition to color coding of the polygons, a numerical path disturbance score may be added to the image near or within each polygon.

[0440] If a polygon is colored gray, this means that there were too few positive control transcriptome value sets for training and the pathway engine 200 was not trained on that polygon. In one example, at least 30 positive control transcriptome value sets would be desirable to train the pathway engine 200n.

[0441] In these examples, the RTK / RAS-PI 3K-EGFR pathway is depicted. The depictions of the RTK / RAS-PI 3K-EGFR pathway depicted in Figures 12A, 12B, 12C, 12D, 12E, and 12F may be included in a pathway perturbation report to assist a physician in determining one or more treatments to prescribe to a patient. In some embodiments, the report includes a treatment recommendation.

[0442] Each of the pathways can include multiple modules. Each module can be associated with a trained model (e.g., a linear model trained using process 670 of FIG. 6G) that can be included in the pathway engine. The modules can be marked with a color and / or pattern that indicates the level of dysregulation or non-dysregulation in the module. In the example below, a red module has been determined to show signs of dysregulation using the associated trained model. A blue module has been determined to show signs of non-dysregulation using the associated trained model. The darkness of the red or blue can correspond to how dysregulated or non-dysregulated the module is, respectively. White can represent a neutral level of dysregulation.

[0443] In Figure 12A, the patient transcriptome being analyzed by pathway engine 200 does not detect mutations in any of the genes in the pathway (the patient is a wild type negative control). As expected, none of the pathway perturbation scores generated by pathway engine 200 indicate that a pathway perturbation is present.

[0444] In FIG. 12B, the patient had a KRAS mutation and no RAF mutation, but the system and method predicted that the KRAS mutation causes increased activity in the RAF class of proteins. In this example, there are no approved therapies targeting RAS, so the patient is matched with a therapy targeting MEK or ERK. Approved RAS-targeted therapies or clinical trials for RAS-targeted therapies can be matched if they exist. In one example, the therapy is approved by a regulatory agency, for example, the Federal Drug Administrat...

Claims

1. 1. A computer-implemented method for training a machine learning model for detecting dysregulation in a cellular pathway, the method comprising: receiving a query including a positive control reference and a negative control reference, the positive control reference including at least one genetic mutation status; obtaining a positive control group and a negative control group in electronic format from a data store comprising the plurality of cell samples; the positive control group includes a cell sample having a genetic mutation that matches the at least one genetic mutation status of the positive control standard; the negative control group comprises a cell sample having a genetic attribute consistent with the negative control criteria; each sample of the plurality of cell samples comprising genetic data for a plurality of genes of the sample and transcriptomic data comprising RNA expression levels; training a machine learning model using the positive control group and the negative control group to determine a correlation between the at least one genetic variant status and a pathway dysregulation score; generating a score for the machine learning model, the score indicating an accuracy of the machine learning model; and Including, the cell sample of the negative control group does not contain a genetic mutation consistent with the at least one genetic mutation status; training the machine learning model comprises identifying one or more feature genes and determining a weight for each feature gene of the one or more feature genes, the weight indicating an influence of the feature gene on the pathway dysregulation score. A computer-implemented method for training a machine learning model for detecting dysregulation in a cellular pathway.

2. 2. The method of claim 1, wherein a subset of the plurality of cell samples is further associated with confounder data for a confounder type, the confounder type comprising at least one of an assay, a match type, a tissue site, a content of the sample, a cancer type, or a purity level of the cell sample.

3. determining a first positive confounder group and a first negative confounder group comprising a cohort of patients having a common confounder type among the positive control group and the negative control group; Quantifying a difference between the first group of positive confounders and the first group of negative confounders; weighting the first group of positive confounders and the first group of negative confounders if the difference exceeds a threshold; The method of claim 2 further comprising:

4. 2. The method of claim 1, further comprising subjecting a negative control standard comprising at least one negative genetic mutation status, wherein the samples of the negative control group comprise the negative genetic mutation status of the negative control standard.

5. The at least one genetic mutation status is a gene identifier comprising at least one gene; at least one genetic variant; At least one pathogenicity classification; The method of claim 1 , comprising:

6. The method of claim 5 , wherein the genetic identifier comprises a plurality of genes.

7. 6. The method of claim 5, wherein the at least one genetic variation is one of a mutation, a fusion, and a copy number variation.

8. 6. The method of claim 5, wherein the pathogenicity classification is one of benign, likely benign, malignant, likely malignant, of uncertain significance, or conflicting evidence.

9. The method of claim 5 , wherein the positive control criteria comprises multiple genetic mutation statuses.

10. The method of claim 5, further comprising selecting a plurality of feature genes downstream of the at least one gene of the gene identifier in a regulatory network.

11. The method of claim 5, further comprising selecting a plurality of feature genes that are upstream of the at least one gene of the gene identifier in a regulatory network.

12. analyzing genetic information of a cell sample associated with a patient using the machine learning model; classifying the cell sample as having a pathway dysregulation associated with the patient's genetic variant based on the analysis; selecting a treatment for the patient, said treatment being based on the classification of the sample; and The method of claim 10 further comprising:

13. 13. The method of claim 12, further comprising outputting a coefficient of the feature gene based on the correlation between the feature gene and the pathway dysregulation score, the pathway dysregulation score indicative of pathway dysregulation and a coefficient of the feature gene indicative of an influence of the feature gene on the pathway dysregulation.

14. Detecting an insertion of a new cell sample into the data store; and evaluating the new cell sample using the trained machine learning model; and adding a positive or negative label to the new cell sample based on the evaluation, the positive label indicating a potential pathway dysregulation in the sample; and The method of claim 1 further comprising:

15. generating at least one additional positive control standard, the at least one additional positive control standard comprising at least one additional genetic mutation status different from the at least one genetic mutation status; obtaining at least one additional positive control group in electronic format from within the data store containing the plurality of cell samples, the at least one additional positive control group comprising cell samples consistent with the at least one additional genetic mutation status; training an additional machine learning model using the at least one additional positive control group to determine a correlation of the at least one additional genetic variant status to a pathway dysregulation score for the at least one additional positive control group; Further comprising: training the additional machine learning model comprises identifying one or more feature genes corresponding to the at least one additional positive control group; and determining a weight for each feature gene of the one or more feature genes, the weight indicating an influence of the feature gene on the pathway dysregulation score for the at least one additional positive control group. The method according to claim 5.

16. 16. The method of claim 15, wherein the genetic identifier of the at least one additional genetic variation status matches the genetic identifier of the at least one genetic variation status of the positive control standard.

17. 17. The method of claim 16, wherein the one or more feature genes corresponding to the positive control group include a first feature gene and a second feature gene, the second feature gene having a weight less than the weight of the first feature gene, and the gene identifiers of the additional positive control standard do not include the second feature gene.

18. Identifying characteristic genes is receiving a list of genes defining said positive control criteria; Defining selection parameters; generating a network of genes related to the list of genes based on the selection parameters; determining an influence level of each gene in the network of genes; ranking the genes based on the level of influence; Identifying the most influential subset of said network of genes; using the expression levels of the most influential subset as features for a feature vector for the machine learning model; The method of claim 1 , comprising:

19. 13. The method of claim 1, further comprising refining the positive control criteria based on results of the machine learning model, wherein refining the positive control criteria comprises one or more of excluding biomarkers or variants shown to be non-pathogenic or including biomarkers or variants shown to be pathogenic.

20. receiving a seed query, the seed query including a positive control group definition and a negative control group definition, the positive control group definition including at least one genetic variant status; Seed multiple cell samples in electronic format: A seed positive control group containing a cell sample that meets the definition of a positive control group; and Seed negative controls containing cell samples that meet the definition of a negative control and analyzing a plurality of genetic features for each of the positive control group and the negative control group using a decision tree model; refining the seed query based on the analysis to generate the query; and The method of claim 1 further comprising:

21. analyzing the plurality of genetic features includes refining the seed query to generate an intermediate query; and comparing an intermediate RNA signature of a cellular sample that satisfies the intermediate query to a seed RNA signature of a cellular sample that satisfies the seed query. removing cell samples that satisfy the intermediate query from the seed plurality of samples if a difference between the intermediate RNA signature and the seed RNA signature exceeds a significance threshold, and analyzing a plurality of genetic features of the seed plurality of samples using a decision tree model; 21. The method of claim 20, comprising:

22. 2. The method of claim 1, wherein the at least one genetic mutation status comprises a threshold gene expression level and a list of genes, and the positive control group comprises a cell sample having an expression level for each gene in the list of genes that is equal to or greater than the gene expression level.

23. 1. A system for training a machine learning model for detecting dysregulation in a cellular pathway, the system comprising: A computer including a processor, the processor comprising: receiving a query including a positive control reference and a negative control reference, the positive control reference including at least one genetic mutation status; obtaining a positive control group and a negative control group in electronic format from a data store comprising a plurality of cell samples, the positive control group comprising cell samples having a genetic mutation matching the at least one genetic mutation status of the positive control reference, the negative control group comprising cell samples having a genetic attribute matching the negative control reference, each sample of the plurality of cell samples comprising genetic data for a plurality of genes of the sample, and transcriptomic data comprising RNA expression levels; training a machine learning model using the positive control group and the negative control group to determine a correlation between the at least one genetic variant status and a pathway dysregulation score; generating a score for the machine learning model, the score indicating an accuracy of the machine learning model; a computer including a processor, the cell sample of the negative control group does not contain a genetic mutation consistent with the at least one genetic mutation status; training the machine learning model comprises identifying one or more feature genes and determining a weight for each feature gene of the one or more feature genes, the weight indicating an influence of the feature gene on the pathway dysregulation score. A system for training machine learning models to detect dysregulation in cellular pathways.

24. When executed by a processor, the processor: receiving a query including a positive control reference and a negative control reference, the positive control reference including at least one genetic mutation status; obtaining a positive control group and a negative control group in electronic format from a data store comprising a plurality of cell samples, the positive control group comprising cell samples having a genetic mutation matching the at least one genetic mutation status of the positive control reference, the negative control group comprising cell samples having a genetic attribute matching the negative control reference, each sample of the plurality of cell samples comprising genetic data for a plurality of genes of the sample, and transcriptomic data comprising RNA expression levels; training a machine learning model using the positive control group and the negative control group to determine a correlation between the at least one genetic variant status and a pathway dysregulation score; generating a score for the machine learning model, the score indicating an accuracy of the machine learning model; A non-transitory computer-readable recording medium having a program encoding instructions stored thereon, the cell sample of the negative control group does not contain a genetic mutation consistent with the at least one genetic mutation status; training the machine learning model comprises identifying one or more feature genes and determining a weight for each feature gene of the one or more feature genes, the weight indicating an influence of the feature gene on the pathway dysregulation score. A non-transitory computer-readable recording medium.