Systems and methods for processing electronic images to determine carcinogenic signals - Patents.com

JP2024542242A5Pending Publication Date: 2025-10-09PAIGE AI INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024529941
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-11-23
Filing Date
2022-10-31
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Current methods for analyzing oncogenic signaling pathways in cancer treatment are limited by the lack of comprehensive data integration of genetic and epigenetic changes, leading to challenges in predicting patient-specific responses to targeted therapies.

Method used

A system and method that utilizes machine learning to populate gene network graphs with patient-specific gene expression levels based on digital medical images, integrating genomic variants, epigenetic changes, and clinical data to predict oncogenic signaling pathway behavior.

Benefits of technology

Enhances the accuracy of predicting cancer treatment responses by providing patient-specific insights into oncogenic signaling pathways, enabling more effective targeted therapies and personalized treatment strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Systems and methods are disclosed for generating and predicting patient-specific oncogenic signaling pathway or network behavior. In some embodiments, a patient-specific oncogenic signaling pathway or network can be generated by receiving one or more digital medical images associated with a patient, providing a raw gene network graph and the one or more digital medical images as inputs to a trained machine learning system, where the machine learning system is trained to populate the gene network graph with patient-specific gene expression levels based on the one or more digital medical images, and receiving the patient-specific populated gene network graph as output from the trained machine learning system.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Patent Application No. 63 / 264,465, filed November 23, 2021, the contents of which are incorporated herein by reference.

[0002] Various techniques of the present disclosure relate generally to oncogenic signaling pathway analysis. More specifically, certain techniques of the present disclosure relate to systems and methods for generating and predicting patient-specific oncogenic signaling pathway or network behavior. [Background technology]

[0003] A signaling pathway, also called a biochemical cascade, is a series of chemical reactions that occur within a living cell when initiated by a stimulus. For example, a signaling pathway in a particular cancer may include one or more mutated genes and the abnormal molecules produced by the genes. Generally, signaling pathways are represented using gene network graphs. Gene network graphs may include nodes representing genes with expression levels added to graphically depict the signaling pathway. Gene network graphs may be used to show genetic changes in signaling pathways that control cell cycle progression, apoptosis, and cell proliferation, hallmarks of cancer. Identification of actionable changes in gene signaling pathways suggests opportunities for targeted and combination therapies for cancer treatment.

[0004] The background discussion provided herein is intended to generally describe the contents of the present disclosure. Unless otherwise indicated herein, the material described in this section is not prior art to the claims of this application and is not admitted to be prior art or an indication of prior art by its inclusion in this section. Summary of the Invention [Means for solving the problem]

[0005] According to certain aspects of the present disclosure, methods and systems are disclosed for generating and predicting patient-specific oncogenic signaling pathway or network behavior. Each of the aspects of the disclosure herein may include one or more of the features described in connection with any of the other disclosed aspects.

[0006] According to an example of the present disclosure, a method for processing digital medical images to populate a gene network graph, or a data representation of the gene network graph, may be described. An exemplary method may include receiving one or more digital medical images associated with a patient, providing the unpopulated gene network graph and the one or more digital medical images as inputs to a trained machine learning system, where the machine learning system is trained to populate the gene network graph with patient-specific gene expression levels based on the one or more digital medical images, and receiving the patient-specific populated gene network graph as an output from the trained machine learning system.

[0007] According to another example of the present disclosure, a method for training a machine learning system to populate a gene network graph can be described.The exemplary method can include: receiving an unpopulated gene network graph including a gene network graph that does not include expression levels; receiving tumor sequence information associated with each of a plurality of patients; receiving one or more digital medical images associated with each of the plurality of patients; populating the gene network graph for each of the plurality of patients to include expression levels based on the respective tumor sequence information; and training a machine learning system to infer one or more of the populated gene network graphs based on the respective one or more digital medical images.

[0008] According to a further example of the present disclosure, a system for processing digital medical images to populate a gene network graph may be described. The exemplary system may include at least one memory storing instructions and at least one processor configured to execute the instructions to perform operations. The operations may include receiving one or more digital medical images associated with a patient, providing the unpopulated gene network graph and the one or more digital medical images as inputs to a trained machine learning system, the machine learning system being trained to populate the gene network graph with patient-specific gene expression levels based on the one or more digital medical images, and receiving the patient-specific populated gene network graph as an output from the trained machine learning system.

[0009] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed embodiments as claimed.

[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate various exemplary technologies and, together with the description, serve to explain the principles of the disclosed technology. [Brief description of the drawings]

[0011] [Figure 1A] FIG. 1 shows a block diagram of an exemplary system for populating a gene network graph and making change predictions according to one or more techniques.

[0012] [Figure 1B] FIG. 1 shows a block diagram of an exemplary system for populating a gene network graph according to one or more techniques.

[0013] [Figure 1C]FIG. 1 shows a block diagram of an exemplary system for predicting changes in a gene network graph by one or more techniques.

[0014] [Figure 1D] 1 shows an exemplary gene network graph according to one or more techniques.

[0015] [Diagram 2] FIG. 1 shows a schematic diagram of an exemplary system for populating a gene network graph according to one or more techniques.

[0016] [Diagram 3] FIG. 1 shows a schematic diagram of an exemplary system for predicting changes to a gene network graph by one or more techniques.

[0017] [Figure 4] FIG. 1 shows a flow diagram of an exemplary process for generating a gene network graph and predicting changes to a gene network graph according to one or more techniques.

[0018] [Diagram 5] 1 shows a flowchart of an exemplary method for populating a gene network graph according to one or more techniques.

[0019] [Figure 6] 1 illustrates an exemplary method for training a machine learning model of a graph generation system according to one or more techniques.

[0020] [Figure 7] 1 shows a flowchart of an exemplary method for predicting changes in a gene network graph by one or more techniques.

[0021] [Figure 8] 1 illustrates an example method for training a machine learning model for a graph predictive system according to one or more techniques.

[0022] [Figure 9] 1 illustrates an example system or device capable of implementing the techniques presented herein according to one or more techniques. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0023] Reference will now be made in detail to the exemplary technology of the present disclosure, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numbers will be used throughout the drawings to refer to the same or like parts.

[0024] The systems, devices, and methods disclosed herein are described in detail, by way of example, with reference to the drawings. The examples described herein are merely examples and are provided to aid in the explanation of the apparatus, devices, systems, and methods described herein. None of the features or components illustrated in the drawings or described below should be considered essential to any particular implementation of these devices, systems, or methods, unless specifically designated as essential.

[0025] Additionally, for all described methods, whether or not the method is described in conjunction with a flow diagram, unless otherwise specified or required by context, any explicit or implicit ordering of steps performed in the performance of the method should mean that those steps must be performed in the order presented, but may be performed in another order or in parallel.

[0026] As used herein, the term "exemplary" is used in the sense of "example" rather than "ideal." Additionally, the terms "a" and "an" do not denote a limitation of quantity herein, but rather denote the presence of one or more of the referenced item.

[0027] As used herein, the term "gene expression" refers to the process by which information from a gene is used to synthesize a functional gene product that allows the gene to produce an end product, such as a protein or non-coding RNA, and ultimately influences the phenotype as a final effect. Controlling gene expression is the control of the amount and timing of the appearance of a gene's functional product. Controlling expression allows a cell to produce the gene product it needs when it needs it, thereby giving the cell the flexibility to adapt to changing environments, external signals, damage to the cell, and / or other stimuli. A genetic (or genetic) regulatory network (GRN) is a collection of molecular regulators that interact with each other and with other substances in the cell to control gene expression levels of mRNA and proteins, which in turn determine the function of the cell.

[0028] As used herein, the terms "gene network signaling graph", "gene network graph", and the like, can be a weighted and directed graph, or a data structure representing a gene network graph, composed of nodes and edges connecting the nodes. Each node can be a gene with an expression level attached, which can be continuous or thresholded and discrete. The edges of the graph can define a signaling pathway graph based on known discoveries in genetics. The continuous expression levels (continuous weights) indicate the activity levels of the genes. In some examples, the values ​​of the continuous weights can be thresholded to simplify the analysis. In such examples, the thresholds can be obtained from the results of scientific research aimed at determining when an expression level is abnormal. As used herein, the term "populated gene network graph", and the like, can be a gene network graph populated with gene expression levels, or a data structure representing a gene network graph. As used herein, the term "unpopulated gene network graph", and the like, can be a gene network graph that is not populated with gene expression levels. For example, a raw gene network graph may include a structure, organization, or topology of nodes and edges that indicate interactions between genes, but does not include the patient-specific expression levels associated with the genes.

[0029] As used herein, "driver mutations," "driver genetic changes," "driver epigenetic changes," and the like refer to genetic mutations that drive the development of cancer. Driver mutations are mutations that allow cancer to grow and invade human cells, such as somatic cells.

[0030] Epigenetic changes, such as abnormal patterns of DNA acetylation and / or methylation, disrupted patterns of histone post-translational modifications, and / or chromatin remodeling, may work in concert with genetic alterations to generate cancer phenotypes. For example, epigenetic changes as alternative drivers of carcinogenesis may result in gene mutations, and conversely, mutations are frequently observed in genes that modify the epigenome (e.g., microsatellite instability, chromosomal instability, promoter hypermethylation, etc.). There may be associations between genetic and / or epigenetic pathways and patient prognosis, overall survival, and / or response to targeted cancer therapy. Thus, genetic alterations in oncogenic signaling pathways and reversible epigenetic changes may be used to inform precision medicine treatment options.

[0031] For example, in HER2+ breast cancer, successful therapeutic approaches have been developed based on small molecule inhibitors such as anti-HER2 drugs. However, toxic chemotherapy and inevitable resistance to targeted therapeutic agents remain challenges in cancer treatment. Driver genetic and epigenetic changes predictive of response to targeted drugs can result in specific histological phenotypes identifiable by microscopy. Thus, artificial intelligence (AI) systems can be used to predict the presence of predictive biomarkers in whole slide imaging (WSI) of tumor samples and can be used as a screening tool for treatment decision-making. However, one of the challenges in establishing such AI systems can be due to the lack of data and the low prevalence of individual genetic mutations in certain tumor types.

[0032] Furthermore, deep learning can also be implemented to predict oncogenic mutations in individual genes from digital histology images, but relying on predictions related to only individual genes may be oversimplified. Tumor growth is governed not by one gene alone, but by many (or all) genes interacting in a network where one gene influences the activity of other genes. For example, different genetic and epigenetic changes in multiple genes may result in a convergent phenotype with the same signaling pathway disruption. Genetic alterations and epigenetic abnormalities may be considered alternative drivers of carcinogenesis, but the extent, mechanism, and co-occurrence of changes in those oncogenic pathways vary across different tumor types and individual tumor samples. Thus, uncovering the impact of oncogenic signaling pathway disruptions that drive cancer phenotypes may require integration of multi-layered evidence of genetic, epigenetic, and clinical symptoms.

[0033] The techniques discussed herein can use AI techniques, machine learning, and / or image processing tools applied to databases of patient histological images, patient clinical information, genomic data, and / or gene network relationships, among other data types, to generate and predict patient-specific oncogenic signaling pathways and network (e.g., gene network graph) behavior. For example, a system can be established that generates patient-specific gene network graphs populated with the corresponding activity levels of each gene (e.g., representing a signaling pathway). Furthermore, the system can detect changes at the signaling pathway level from digital medical images (e.g., histological WSI) by integrating data from multiple sources, including genomic variants and epigenetic changes, signaling pathways, clinical symptoms, and treatment outcomes. For example, the system can identify computed and / or learned histological features associated with abnormalities in oncogenic signaling pathways or complexes and predict which oncogenic signaling pathways are driving tumor development in a patient, which can be utilized as a therapeutic target.

[0034] FIG. 1A illustrates an exemplary system for generating and predicting patient-specific oncogenic signaling pathway or network behaviors by one or more techniques. Shown in FIG. 1A is an electronic network 120 that may be connected, for example, via one or more computers, servers, and / or handheld mobile devices, to a physician server 121, a hospital server 122, a clinical trial server 123, a laboratory server 124, and / or a laboratory information system 125. According to an exemplary embodiment of the present disclosure, the network 120 may be connected to a server system 110, which may include, for example, one or more processing devices 100 configured to execute or implement the graph generation system 101 and the graph prediction system 102, and a storage device 109. The graph generation system 101 may be configured for population of gene network graphs using one or more trained machine learning systems. The graph prediction system 102 may be configured for prediction of the behavior of gene network graphs, such as, for example, a population of gene network graphs generated by the graph generation system 101 or another system, using one or more trained machine learning systems, according to an exemplary embodiment of the present disclosure. Although the graph generation system 101 and the graph prediction system 102 are shown as separate systems in FIG. 1, it should be understood that in other examples, the graph generation system 101 and the graph prediction system 102 may be subsystems of a larger system.

[0035] The physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125 can create or acquire data such as digital medical images, expression data, genomic variants, and / or clinical data. For example, the digital medical images can include digital pathology images including whole slide image(s), cytology specimen(s), histopathology specimen(s), slide(s) of cytology specimen, digital image(s) of histopathology specimen slide(s), or any combination thereof, of one or more patients that can be created or acquired. Additionally or alternatively, the digital medical images can include other modality type images including digital multiplexed immunofluorescence images, digital multiplexed immunohistochemistry images, magnetic resonance imaging (MRI), computed tomography (CT), x-ray, nuclear medicine imaging, or ultrasound that can be created or acquired.

[0036] Expression data may include patient-specific or non-patient-specific tumor sequence data, protein expression levels, and / or non-coding RNA expression levels. Expression data can be utilized for training purposes by both medical professionals (e.g., pathologists, physicians, etc.) and AI systems alike to improve the accuracy of predicting carcinogenesis patterns, among other tasks. The increasing availability of expression data indicative of specific conditions or diseases will increase the representation variability among expression data, improving the learning capabilities of both medical professionals and AI systems. However, large amounts of expression data are still not available for individual gene mutations in specific tumor types, which necessarily limits the amount of variation that can be learned. For example, treating a patient-specific tumor may be difficult due to genotypic differences compared to another patient with the same phenotype but a different genotype.

[0037] Genomic variants may include mutations in individual genes of a given gene complex or signaling pathway, such as the SWI / SNF complex (e.g., ARID1A, ARID1B, ARID2, PBRM1, SMARCA4, and SMARCB1) or the RTK / RAS pathway (e.g., ERBB2, ERBB3, ERBB4, SOS1, HRAS, BRAF, MAP2K1, and MAPK1). Clinical data may include age, medical history, cancer treatment history, family history, previous biopsy or cytology information, tumor sequence information, mRNA expression levels, gene network graphs (pre- and / or post-treatment), overall survival data, progression-free survival with corresponding censored data, 5-year survival rate, drug treatment outcome data, and the like.

[0038] Digital medical images, expression data, genomic variants, clinical data, and / or other data may be communicated in digital or electronic form between server system 110 and physician server 121, hospital server 122, clinical trial server 123, laboratory server 124, and / or laboratory information system 125 via network 120.

[0039] The server system 110 may include one or more storage devices 109 for storing data, such as digital medical images, expression data, genomic variants, clinical data, etc., received from at least one of a physician server 121, a hospital server 122, a clinical trial server 123, a laboratory server 124, and / or a laboratory information system 125. For example, the one or more populated gene network graphs generated by the graph generation system 101 may be stored in one or more data stores, such as the storage devices 109.

[0040] The server system 110 may include the processing device 100 for processing digital medical images and / or other aforementioned data stored in the storage device 109. The server system 110 may include one or more machine learning tool(s) or functionality. For example, the processing device 100 may execute one or more machine learning systems utilized by the graph generation system 101 and / or the graph prediction system 102 according to one or more techniques. In some examples, the output of the machine learning systems may be stored in the storage device 109 for use in other systems or processes, as described in more detail below. Alternatively or additionally, the present disclosure (or parts of the systems and methods of the present disclosure) may be executed on a local processing device (e.g., a laptop).

[0041] According to an exemplary embodiment of the present disclosure, the graph generation system 101 may be configured to generate a gene network graph using one or more machine learning systems. The populated gene network graph may be patient-specific and may include the corresponding activity level of each gene. According to an exemplary embodiment of the present disclosure, the graph prediction system 102 may be configured to predict how the populated gene network graph may behave over time using one or more machine learning systems with or without one or more treatments. This embodiment may make available patient-specific data, allowing for more accurate prediction of oncogenic changes, e.g., gene expression changes, in response to a particular treatment.

[0042] 1B illustrates an exemplary system for generating a populated gene network graph, such as a graph generation system 101, according to an exemplary embodiment of the present disclosure. The graph generation system 101 may include a training graph generation platform 131 and / or a target graph generation platform 135.

[0043] According to one technique, the training graph generation platform 131 may be implemented to generate or receive one or more data sets of training data that generate and train one or more machine learning models that populate the gene network graph with gene expression levels and / or predicted tumor gene expression levels. According to one technique, the training graph generation platform 131 may include multiple software modules including a training data ingestion module 132, a training data input module 133, and a training data input predictive model 134. Data output by the training graph generation platform 131 and / or the machine learning system may be stored, for example, in the storage device 109, or may be used by other systems, such as the target graph generation platform 135.

[0044] According to one aspect, the training data ingestion module 132 may create or receive training data (e.g., pristine gene network graphs, expression data, digital medical images, optional clinical data, etc.) that may be used to train one or more machine learning methods to generate the populated gene network graph. The training data may be received from any one or any combination of the server system 110, the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125. The training data may be obtained from real sources (e.g., humans, animals, etc.) or synthetic sources (e.g., graphics simulators, graphics rendering engines, 3D models, etc.).

[0045] The training data datasets may include one or more datasets corresponding to a pristine gene network graph, one or more datasets corresponding to expression data (e.g., RNA expression data), one or more datasets corresponding to tumor sequencing information, one or more datasets corresponding to digital medical images, and / or one or more datasets corresponding to clinical data. In some examples, the subsets of training data may overlap between various datasets of gene network graphs, tumor sequencing information, expression data, and / or clinical data. The training datasets may be stored in a digital storage device, such as one of the storage devices 109.

[0046] In some examples, expression data, e.g., gene expression data and / or RNA expression data, may be a direct output of one or more machine learning systems. In other examples, the output of one or more machine learning systems may be used as an input to further processes that allow for the generation of a data-populated gene network graph. In another example, the training WSI may include digitized histology or cytology slides stained with various stains, such as, but not limited to, hematoxylin and eosin, hematoxylin alone, toluidine blue, alcian blue, Giemsa, trichrome, acid-fast, Nissl, etc. Other training data may include genomic variants and / or clinical data, as discussed herein. Clinical data may include histological data, tumor subtype data, tumor grading or staging data, tumor sizing data, patient demographic data, etc.

[0047] The training data input module 133 can populate the gene network graph based on at least the tumor sequencing information. The pristine gene network graph can be generated by one or more systems, such as the training data input module 133, based on interaction data associated with a given gene set obtained from published studies or other similar sources. Additionally or alternatively, the pristine gene network graph can be received from a public database that stores a collection of pristine gene network graphs for various gene sets (e.g., the pristine gene networks can be pre-created by a third party, stored in a public database, and provided as input to the training data input module 133). The tumor sequencing information can be received from any one or any combination of the server system 110, the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, the laboratory information system 125, and / or the training data ingestion module 132. The training data input module 133 can, for example, output a gene network graph populated with patient-specific expression levels, which can be stored, for example, in the storage device 109 and / or utilized by the training data input prediction module 134 during training.

[0048] In some examples, tumor sequence information can be obtained from a gene panel. However, a gene panel may not obtain expression levels from all genes, but only a subset of genes. Thus, the tumor sequence information may be incomplete and may not be able to fully populate the gene network graph. In such cases, genes that are not included in the gene panel may be treated as missing values, which may be determined, for example, using a label propagation algorithm (LPA), so that the gene network graph is fully populated. For example, inference of missing values ​​corresponding to expression levels of unanalyzed genes may be performed using directed LPA.

[0049] In addition to populating the gene network graph using tumor sequencing information, the training data input module 133 can be configured to correlate patient-specific tumor sequence data with phenotypes (e.g., tumor expression with expression of given tumor-regulating genes) and / or add the tumor sequence and gene expression data to a database of other sequencing and expression data (e.g., storage device 109). In some examples, a third party can train one or more machine learning systems of the training data input module 133 and provide the trained machine learning system(s) to the server system 110 for storage (e.g., in the storage device 109) and execution (e.g., by the target graph generation platform 135).

[0050] The training data input prediction module 134 may be trained to infer a populated gene network graph from a digital medical image. In other words, the training data input module 133 may be configured to predict a populated gene network graph generated based on gene expression data (e.g., how a tumor with a given phenotype interacts with other genes) associated with a digital medical image of a given tumor. In some examples, the training data input prediction module 134 may be further trained using clinical data such as overall survival data, progression-free survival with corresponding censoring data, drug treatment outcome data, etc. Exemplary methods for training one or more machine learning systems of the training data input prediction module 134 are described in detail below.

[0051] In some examples, a machine learning system may be generated for each of the different tissues and / or tumor types to learn a corresponding gene network graph. In other examples, one machine learning system may be generated that can learn gene network graphs for more than one tissue and / or tumor type. The training data input prediction module 134 may generate one or more machine learning systems configured to operate via any of a multimodal deep neural network, a graph neural network, a convolutional neural network, a transformer neural network, etc.

[0052] According to one technique, the target graph generation platform 135 can include software modules such as a target data ingestion module 136, a data input module 137, and an output interface 138. According to one aspect, the target graph generation platform 135 can receive a request for expression data and can execute one or more of the machine learning systems trained by the training graph generation platform 131 to generate one or more populated gene network graphs. For example, the request can be received from any one or any combination of the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125. In another example, the request can be automatically received from the graph prediction system 102 in response to the graph prediction system 102 receiving a request to predict a gene network graph and / or receiving a patient-specific unpopulated gene network graph.

[0053] According to one aspect, the target data ingestion module 136 can create or receive target data (e.g., images, optionally clinical data, etc.) that can be used as input for one or more trained machine learning systems to generate a populated gene network graph. For example, the target data ingestion module 136 can receive digital medical images that can be used as input for one or more trained machine learning systems. The target data can be received from any one or any combination of the server system 110, the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125. The target data can be obtained from real sources (e.g., humans, animals, etc.) or synthetic sources (e.g., graphics simulators, graphics rendering engines, 3D models, etc.). The target data ingestion module 136 can create or receive one or more datasets of target data, e.g., digital medical images. For example, the datasets can include one or more datasets corresponding to the digital medical images and / or one or more datasets corresponding to the clinical data, optionally. In some examples, the subsets of target data may overlap between various datasets of image and / or clinical data. The target datasets may be stored in a digital storage device, such as one of storage devices 109.

[0054] The data input module 137 may include any suitable machine learning system, including but not limited to graph neural networks, convolutional neural networks, transformer neural networks, etc. The data input module 137 may execute various machine learning systems generated by the training graph generation platform 131, e.g., the training data input prediction module 134, to facilitate the generation and / or data input of the gene network graph. The data input module 137 may populate the gene network graph with expression levels inferred based on one or more features identified from processing one or more medical images and, optionally, clinical data, as well as associated expression levels learned for those one or more features.

[0055] The output interface 138 may be used to output the populated gene network graph (e.g., to a screen, monitor, storage device, web browser, etc.). According to some techniques, the output interface 138 may output the populated gene network graph to the graph prediction system 102 for use as input in subsequent processes described below. The populated gene network graph and other data generated or used by the graph generation system 101 may be stored in one or more storage devices 109.

[0056] FIG. 1D illustrates an exemplary populated gene network graph 150, such as may be output by the graph generation system 101 and / or the graph prediction system 102. As illustrated in FIG. 1D, the gene network graph 150 may include one or more nodes corresponding to genes in the network, such as one or more transcription factor nodes 152, one or more protein nodes 153, etc. Relationships between the nodes, such as continuous expression levels, may be represented by one or more edges. A visual representation of the edges in the gene network graph 150 may depict characteristics of the relationships, such as the presence of genetic evidence, positive or negative effects, expression and / or regulation. For example, edge 154a may represent genetic evidence of protein 153HOG1 inducing a negative effect on protein 153CHS. In another example, edge 154b may represent transcription factor 152P1F3 inducing a positive effect on expression at transcription factor 152LHY. In another example, edge 154c may represent transcription factor 152P1F3 inducing upregulation of a promoter that binds protein 153CAB1. Any other suitable nodes, edges, and / or combinations of nodes and edges may be represented on gene network graph 150. Although gene network graph 150 may include and depict populated nodes, it should be understood that a gene network graph, such as a non-populated gene network graph received as input by graph generation system 101, may not include or depict populated nodes and resulting signal pathways and / or relationships.

[0057] FIG. 2 shows a schematic diagram 200 of an exemplary system (e.g., graph generation system 101) implemented to generate a populated gene network graph. As shown in FIG. 2, the graph generation system 101 can receive one or more inputs 202, for example, at the target data ingestion module 136. The one or more inputs 202 can include, but are not limited to, one or more gene network graphs 204 without expression levels, patient clinical information 206, digital medical images of the patient 208, or any combination thereof. The one or more inputs 202 can be received from any one or any combination of the server system 110, the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125. The one or more inputs 202 can be processed by the graph generation system 101 using one or more trained machine learning systems (e.g., data input module 137) to output a patient-specific expression level populated gene network graph (populated gene network graph) 212. The output data input gene network graph 212 can be stored, for example, in the storage device 109 and can be received by the graph prediction system 102 for further processing.

[0058] 1C illustrates an example system for predicting the behavior of a gene network graph, such as a graph prediction system 102, in accordance with an example technique of this disclosure. The graph prediction system 102 can include a training graph prediction platform 141 and / or a target graph prediction platform 145.

[0059] According to one technique, the training graph prediction platform 141 can include software modules such as a training data ingestion module 142 and a training prediction module 147. According to one aspect, the training data ingestion module 142 can create or receive training data that can be used to train one or more machine learning systems to generate populated gene network graphs and / or predict treatment outcomes after treatment. The training data can be received from any one or any combination of the server system 110, the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125. The training data can be obtained from real sources (e.g., humans, animals, etc.) or synthetic sources (e.g., graphics simulators, graphics rendering engines, 3D models, etc.). The training data ingestion module 142 can create or receive one or more datasets of training data. For example, the dataset may include (one or more datasets corresponding to) gene network graphs of multiple patients, e.g., pre-treatment and post-treatment, one or more datasets corresponding to treatment data of multiple patients (e.g., type of treatment, dosage, etc.), one or more datasets corresponding to a time delay before and after treatment. In another example, each dataset is patient-specific and may include pre-treatment gene network graph, post-treatment gene network graph, treatment data, and / or a period between pre-treatment and post-treatment gene network graphs for a given patient. In some examples, the subsets of training data may overlap among various datasets for gene network graphs, treatment data, and / or time delay data.

[0060] In some examples, the training data may be the direct output of one or more machine learning systems. In other examples, the output of one or more machine learning systems may be used as input to further processes that allow prediction of changes in the gene network graph. The training dataset may be stored in a digital storage device, such as one of the storage devices 109.

[0061] The training prediction module 143 can use the training data as input to generate one or more machine learning systems that can predict, for example, changes in the gene network graph in response to a proposed treatment regimen. In some examples, a third party can generate one or more trained machine learning systems and provide the trained machine learning system(s) to the server system 110 for storage (e.g., in the storage device 109) and / or execution by the graph prediction system 102. The training prediction module 143 can train a transformer, a graph neural network, or any other suitable type of machine learning system to predict, for example, a post-treatment gene network graph that indicates how gene expression levels may change from a pre-treatment gene network in response to a particular treatment. The training prediction module 143 can store the post-treatment gene network graph in a database, for example, in the storage device 109, along with other gene network graphs, such as, for example, the pre-treatment gene network graph and the populated gene network graph.

[0062] Additionally or alternatively, the training prediction module 143 can use the training data as input to generate one or more machine learning systems capable of predicting treatment outcomes (also referred to herein as patient outcomes). In some examples, a machine learning system may be generated for each of the different tissue types, e.g., tumor types, to learn the corresponding tissue responses to a given treatment. In other examples, one machine learning system may be generated that can predict treatment outcomes for two or more tissue types. Methods for training one or more machine learning systems of the training prediction module 143 are described below.

[0063] According to one technique, the target graph prediction platform 145 can include software modules such as a target data ingestion module 146, a prediction module 147, and an output interface 148. The target data ingestion module 146 can receive one or more target inputs, including, but not limited to, a pre-treatment gene network graph, a treatment regimen, a time delay before and after treatment, etc. For example, the one or more target data can be received from any one or any combination of the server system 110, the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125.

[0064] The target data ingestion module 146 can provide one or more inputs to the prediction module 147 to predict gene network graph changes and / or treatment outcomes. The prediction module 147 can be comprised of one or more components described in more detail below. The prediction module 147 can execute various machine learning models generated by the training graph prediction platform 141 to facilitate predicting gene network graph changes and / or treatment outcomes.

[0065] According to one aspect, the prediction module 147 can receive a request to predict one or more changes in the gene network graph and / or a treatment outcome, and execute one or more of the machine learning systems trained by the training graph prediction platform 141 to predict one or more changes to the gene network graph and / or a treatment outcome in response to the request. For example, the request can be received from any one or any combination of the server system 110, the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125. In another example, the request can be automatically generated by the graph prediction system 102 in response to detection of an output from another system, such as, for example, from the graph generation system 101. In some embodiments, the prediction module 147 can be configured to automatically predict one or more changes in the gene network graph and / or a treatment outcome in response to detection of a significant change or deviation in the data-entered gene network expression levels compared to a baseline, in response to detection of a change in treatment (e.g., drug administration), and the like.

[0066] The prediction module 147 may include any suitable machine learning system, including but not limited to graph neural networks, convolutional neural networks, transformer neural networks, and the like. The prediction module 147 may execute various machine learning systems generated by the training graph prediction platform 141, such as the training prediction module 143, to facilitate the generation and / or data input of post-treatment gene network graphs and / or outcome data. The post-treatment gene network graph may depict a gene network graph having predicted expression values ​​as a function of stimuli, such as, for example, time, treatment regimen, and the like. The outcome data may include clinical data, such as overall survival data, progression-free survival with corresponding censored data, drug treatment outcome data, time delay before treatment (e.g., between diagnosis and treatment), remission rate, and the like.

[0067] The output interface 148 may be used to output (eg, to a screen, monitor, storage device, web browser, etc.) the populated gene network graph and / or the predicted post-treatment outcome data.

[0068] FIG. 3 illustrates a schematic diagram 300 of an exemplary system (e.g., graph prediction system 102) implemented to predict changes to a gene network graph. As illustrated in FIG. 3, the graph prediction system 102 may obtain one or more inputs 302, for example, at the target data ingestion module 146. As described herein, the one or more inputs 302 can be received from any one or any combination of the server system 110, the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125. The one or more inputs 302 may include, but are not limited to, one or more pre-treatment gene network graphs 304, treatment regimens 306, optional pre-treatment and post-treatment time delays 308, and the like. The pre-treatment gene network graph 304 may show a gene network graph with expression values ​​before the occurrence of a stimulus, for example, a time delay, a treatment regimen, and the like. The treatment regimen 306 may include dosage, schedule, timing, and the like. The pre-treatment and post-treatment time delays 308 may include the time between diagnosis and treatment. Using one or more trained machine learning systems as described herein to process the inputs 302, the graph prediction system 102 can provide a post-treatment gene network graph 314 and / or a patient outcome 316 as one or more outputs 312. The outputs 312 can be associated with a period of time indicated by the time delay 308.

[0069] FIG. 4 illustrates a flow diagram of an exemplary process for inputting data into a gene network graph and making change predictions according to one or more techniques. As illustrated in FIG. 4, one or more inputs can be processed by one or more systems (e.g., graph system 412) to generate one or more outputs. Graph system 412 can include a graph data input system 101 and a graph prediction system 102 operating in conjunction with each other. The one or more inputs can include genomic and / or epigenomic data 402 (e.g., genomic variants, epigenetic changes, or gene panels), expression data 404 (e.g., RNA-seq and / or genetic microarrays), clinical data 406 (e.g., survival data and / or treatment response), and / or medical images 408 (e.g., matched images or whole slide images (WSI)).

[0070] The genomic and / or epigenomic data 402 can include, for example, genomic and epigenetic variants such as point mutations, copy number variations, structural variations, histone modifications, and / or hypermethylation. The genomic and / or epigenomic data 402, when available, can be used to identify corresponding gene expression profiles induced by genomic or epigenetic variants.

[0071] For example, the expression data 404, such as RNA-seq expression levels, may include patient gene matrices across multiple cancer types (pan-cancer), patient gene matrices across a single cancer type, patient gene matrices at various stages of treatment such as at different timestamps (e.g., time series), gene matrices of patients after various combination therapies, etc. In some examples, the values ​​may be normalized expression levels, such as normalized FPKM values ​​(fragments per kilobase of transcript per million mapped reads), or microarray gene data. Genes with flat gene expression profiles may be filtered out. In some embodiments, the expression data 404 may be in the form of a populated gene network graph that may include the expression data. The expression data 404 may be used to identify survival-related or drug response-related genes using univariate and / or multivariate Cox Proportional Hazard (CoxPH) regression.

[0072] The clinical data 406 may include age, medical history, cancer treatment history, family history, past biopsy or cytology information, tumor sequence information, mRNA expression levels, gene network graphs (pre- and / or post-treatment), overall survival data, progression-free survival with corresponding censored data, drug treatment outcome data, etc. As discussed herein, the graph generating system 101 may be optionally trained using certain types of clinical data 406 (e.g., cancer treatment history, family history, mRNA expression levels, etc.), while the graph generating system may be trained using other types of clinical data 406 (e.g., overall survival data, progression-free survival with corresponding censored data, drug treatment outcome data, etc.). The medical images 408 may be at the patient-to-sample level. In some examples, the image or WSI may be divided into multiple tiles for analysis and / or processing. However, the predicted expression level may be based on the aggregated values ​​from the multiple tiles.

[0073] The gene expression profile 410 may include a gene network graph before and / or after treatment. The gene expression profile 410 may be based at least in part on expression data 404 and / or clinical data 406 related to treatment and outcome data, and thus the gene expression profile 410 may be associated with an outcome. The gene expression profile 410 may include a continuous value vector indicating normalized expression levels of a list of genes, or an integer vector (e.g., 1, 0, -1) of a list of genes indicating whether the genes may be overexpressed, unchanged, or downregulated. Based on the predicted gene expression profile 410, genes that are highly expressed or downregulated may be identified. For example, pathway enrichment analysis may be used to determine associated pathways, such as T cell receptors, DNA repair pathways, and associate the pathways with potential therapeutic therapies.

[0074] The gene expression profiles 410 may be used to train a machine learning system executed by the graph system 412 to infer a patient outcome given gene expression changes (e.g., changes in gene graph networks) and / or specific treatment regimens and / or time delays, e.g., for different sample subsets, using Bayesian inference, correlation inference, Boolean inference algorithms, or other suitable models. For example, a patient's possible or likely treatment outcomes and / or a patient's prognosis (e.g., good or bad) may be inferred based on how the patient's gene network graph (e.g., oncogenic signaling pathways) is predicted to change based on the treatment received. This process may help identify gene interactions that may contribute to a high risk of cancer or drug resistance, and / or provide insight into various clinical outcomes of the same mutated genes (e.g., identifying HER2+ patients who do not respond well to anti-HER2+ drugs (e.g., lapatinib and trastuzumab) and anti-HER2+ combination drug therapy).

[0075] As discussed herein, the identified subset of inputs may be used to train one or more machine learning systems of the graph system 412, which may include the graph generation system 101 and / or the graph prediction system 102. For example, the genomics and / or epigenomic data 402, the clinical data 406, and / or the medical images 408 may be used to train the graph generation system 101. In another example, the expression data 404, the clinical data 406, and / or the gene expression profiles 410 may be used to train the graph prediction system 102.

[0076] In some techniques, the graph system 412 can generate one or more outputs. For example, the graph system 412 can output one or more predicted patient outcomes 414, predicted expression levels 416, and / or predicted post-treatment graphs 418 (e.g., based on the predicted expression levels 416). In some examples, the predicted patient outcomes 414 can include a heat map of disease and / or histological patterns associated with expression levels, including a matrix (or a matrix of signed values ​​(weights), which may be values ​​indicating associations between genes (i.e., how expression of one gene changes expression of other genes)) for each gene with a value (e.g., 1, 0, and / or -1 for upregulation, no interaction, or downregulation (i.e., inhibition), respectively). The predicted expression levels 416, e.g., predictions including expression levels at future time points, can include different stages / phases of drug treatment (with or without dosage changes), which can be used to infer the predicted post-treatment graph 418 and / or response to treatment. In some examples, the graph system 412 can directly infer the predicted post-treatment graph 418 using one or more inputs, such as, for example, the expression data 404 and / or the medical images 408 .

[0077] 5 shows a flowchart of an exemplary method 500 for inputting data into a gene network graph according to one or more techniques. In step 502, a system, such as, for example, the graph generation system 101, can receive one or more digital medical images associated with a patient. Optionally, in step 504, clinical data associated with the patient can also be received. The input data can be generated and / or stored by a system described herein, such as, for example, the storage device 109, or can be received from one or more of the physician server 121, the hospital server 122, the clinical trial server 123, the laboratory server 124, and / or the laboratory information system 125.

[0078] In step 506, one or more digital medical images and a raw gene network graph associated with the patient may be provided to the trained machine learning system. If clinical data is optionally received, the clinical data may also be provided to the trained machine learning system. The trained machine learning system may process the one or more digital medical images (and optionally the clinical data) to output at least one raw gene network graph. The raw gene network graph may be received from the trained machine learning system in step 508. For example, for a patient with a given cancer genotype, the graph generation system 101 may use a multimodal deep neural network to predict the raw gene network graph for the patient based on the patient's digital medical images, such as digital whole slide images of cancer tissue.

[0079] The machine learning system may be trained as described in Figure 6. The machine learning system may be trained to infer a patient-specific populated gene network graph based on one or more inputs, such as, for example, mRNA and / or tumor sequencing data, patient clinical information, digital medical images, etc.

[0080] 6 illustrates an exemplary method 600 for training a machine learning model implemented by the graph generation system 101 according to one or more techniques. The machine learning model may be trained to infer a populated gene network graph based at least on digital medical images. The exemplary method 600 (e.g., steps 602-610) may be performed by the graph generation system 101. The exemplary method 600 may include one or more of the following steps:

[0081] In step 602, a raw gene network graph can be received. The raw gene network graph may show general relationships between genes, proteins, mRNAs, etc., but may not include any gene expression levels. For example, the raw gene network graph may show that relationships exist between individual genes of the SWI / SNF complex (e.g., ARID1A, ARID1B, ARID2, PBRM1, SMARCA4, SMARCB1 genes), but may not show how those genes may interact within a particular individual cell. In other words, general gene relationships may be shown using the raw gene network graph, but the amount or extent of a particular gene interaction in a patient (e.g., whether a particular gene interaction occurs to a higher or lower degree in a particular patient) is not shown due to the lack of gene expression levels. In some examples, the general relationships between genes, proteins, mRNAs, etc. shown in the raw gene network graph can be received from a public database that can store a collection of general relationship data for various sets of genes.

[0082] At step 604, tumor sequencing information associated with a plurality of patients can be received. The tumor sequencing information can include patient-specific gene sequences (e.g., driver regions, promoter regions, exons, etc.), mutation data based on associated patient populations (e.g., gene mutations associated with Ashkenazi Jewish populations), and the like. The tumor sequencing information can indicate expression levels associated with genes represented by nodes in the unpopulated gene network graph. At step 606, the machine learning system can receive a plurality of patient digital medical images associated with a plurality of patients. As discussed herein, the digital medical images can be any suitable configuration, such as, for example, digital multiplexed immunofluorescence images, digital multiplexed immunohistochemistry images, magnetic resonance imaging (MRI), computed tomography (CT), x-ray, nuclear medicine imaging, ultrasound, and the like. Optionally, at step 608, clinical data associated with a plurality of patients can be received for training. The clinical data can include cancer treatment history, family history, mRNA expression levels, and the like. Steps 602, 604, 606, and 608 can be performed simultaneously and / or separately.

[0083] In step 610, for one or more of the multiple patients, a pre-populated gene network graph may be populated based on the tumor sequencing information of the respective patient. In other words, the training data input module 133 may correlate the tumor sequencing data with the expression level data to generate a populated gene network graph including the expression levels of genes represented by the nodes of the gene network graph. Populating the gene network graph may further include determining whether there are any expression level values ​​missing from the gene network graph and inferring the missing values ​​using a label propagation technique. As discussed herein, the tumor sequence data may not provide or include expression levels of all genes. In such cases, the expression level values ​​of the missing genes may be determined, for example, using a directional label propagation technique, as discussed herein.

[0084] In step 612, the machine learning system may be trained using multiple inputs, such as the unpopulated gene network graph, digital medical images, and / or clinical data, if optionally received. The machine learning system may be trained to infer a populated gene network graph for a patient based on one or more medical images of the patient using supervised or semi-supervised learning. The trained machine learning system may be output to digital storage, such as, for example, storage device 109.

[0085] In some examples, the supervised machine learning system may be trained using classification or regression. In such examples, the supervised machine learning system may include a multimodal deep neural network, a graph neural network, a transformer neural network, a convolutional neural network (CNN), a recurrent neural network (RNN), or a multi-layer perceptron (MLP), among other similar examples. To enable learning, a digital medical image may be provided as an input to the machine learning system. The machine learning system may then output predicted gene sequence data that may be used to populate a gene network graph. The predicted gene sequencing data may be compared to corresponding gene sequencing data to determine a loss or error that may be used to update parameters of the machine learning system to reduce the loss or error. The corresponding gene sequence data may be part of the training gene sequence data that correspond to the digital medical image and are indicative of aspects and / or genotypes of known cancer tissues in the digital medical image. The machine learning system may be modified or altered based on the error (e.g., weights and / or biases associated with one or more nodes and / or layers may be adjusted) to improve the accuracy of the machine learning system. This process may be repeated for each received training digital medical image, or at least until the determined loss or error is below a predetermined threshold. In some examples, a portion of the training images may be retained and used to further validate or test the machine learning system.

[0086] In some examples, the machine learning model may include a sequence-to-sequence ("Seq2Seq") model, e.g., a Transformer Seq2Seq model. The Transformer Seq2Seq model may include an encoder model and a decoder model. The Transformer Seq2Seq model may be configured to receive as input tiles from digital medical images and / or vector embeddings of tiles (and optional clinical data), for example, which the encoder model encodes and / or compresses. The decoder model may receive and decode the encoded and / or compressed vector embeddings from the encoder model to output a populated gene network graph. The decoder output may be compared to an actual populated network graph for a patient (e.g., populated using the patient's tumor sequencing information) to determine a loss or error, which may be used to update parameters of the machine learning system to reduce the loss or error. In some examples, the Transformer Seq2Seq model may receive a variable amount of data as input and generate a fixed-size populated gene network graph.

[0087] In some embodiments, the populated gene network graph output by the trained machine learning system of graph generation system 101 can be a pre-treatment gene network graph (e.g., a gene network graph showing gene expression levels before a patient receives a treatment). Using this graph as input, graph prediction system 102 can be configured to predict an outcome based on a proposed treatment and / or a post-treatment gene network graph.

[0088] 7 shows a flowchart of an exemplary method 700 for predicting changes in a gene network graph according to one or more techniques, e.g., in response to a proposed treatment regimen. The exemplary method 700 (e.g., steps 702-710) may be performed by the graph prediction system 102. The exemplary method 700 may include one or more of the following steps:

[0089] At step 702, a pre-treatment populated gene network graph can be received. As discussed herein, the pre-treatment populated gene network graph can include a gene network graph before the patient receives any treatment, a gene network graph after the patient receives a first treatment, etc. For example, for a patient who has received a first treatment, the pre-treatment populated gene network graph can be populated based on gene expression values ​​after the first treatment but before a proposed second treatment. In some examples, the pre-treatment populated gene network graph can be generated by the graph prediction system 101.

[0090] At step 704, one or more proposed treatment regimens can be received. The proposed treatment regimens can include monotherapy (e.g., chemotherapy, MET inhibitors, etc.), combination therapy (e.g., cisplatin and taxol for pancreatic cancer treatment), timing data (e.g., treatment frequency, treatment duration, etc.), dosage data (e.g., single dose, total dose, etc.), etc. At either or both of steps 702 and 704, the input can be received at the graph prediction system 102, such as, for example, the targeted graph prediction platform 145, and / or can be stored, for example, by the storage device 109. In some examples, the proposed treatment regimens can be vectorized into vector embeddings. At step 706, optionally, time delay data can be received. The time delay data can define the period of time that has elapsed between the pre-treatment data inputted gene network graph received at step 702 and the post-treatment gene network graph, and / or the outcome predicted by the trained machine learning system. For example, if a user is interested in determining how a patient's gene expression levels change after a one-year period based on a proposed treatment regimen, the time delayed data may represent one year.

[0091] In step 708, the pre-treatment populated gene network graph, one or more proposed treatment regimens may be provided as input data to a trained machine learning system, such as the target graph prediction platform 145. Optionally, the input data, if received, may also include time delayed data. The trained machine learning system may process the input data to output at least one post-treatment populated gene network graph, and / or outcome data, in step 710. The post-treatment populated gene network graph may have the same structure or topology as the pre-treatment gene network graph (e.g., the same nodes and edges), but may have different values ​​associated with one or more nodes indicating changes in expression (e.g., changes in behavior) of the genes represented by the nodes. If the optional time delayed data is received, the post-treatment populated gene network graph, and / or outcome data predicted and output by the trained machine learning system may be at a predetermined time delay (e.g., a defined period following the pre-treatment graph). Alternatively, if no time delay data is received, a set of post-treatment populated gene network graphs and / or outcome data at a series of time delays can be predicted and output by the trained machine learning system. For example, the series of time delays can be predefined intervals (e.g., every 6 months, every year, every 3 years, etc.) following the pre-treatment graphs. The outcome data can include, for example, predicted treatment success rates, predicted gene interactions, predicted survival rates, risk of metastasis, immunotherapy resistance of T cell receptors, etc. Any of the inputs to the graph prediction system 102 (e.g., pre-treatment populated gene network graphs), outputs from the graph prediction system 102 (e.g., post-treatment populated gene network graphs), and / or any other data can be stored, for example, in the storage device 109.

[0092] In one example, the post-treatment populated gene network graph can inform the predicted efficacy of a proposed treatment regimen and / or chemotherapy resistance. In another example, for a patient with a given acute lymphoblastic leukemia genotype, the graph prediction system 102 can use a transformer to predict the post-treatment populated gene network graph of the patient based on the pre-treatment populated gene network graph and the proposed treatment of azacitidine. In another example, the graph prediction system 102 can use a graph neural network to predict whether a patient will develop endometrial hyperplasia, which is an overgrowth of normal cells, or atypical endometrial hyperplasia, which is an overgrowth of abnormal cells, based on the changes in pathways at different times.

[0093] The machine learning system may be trained as described in Figure 8. The exemplary method 800 (e.g., steps 802-810) may be performed by the trained graph prediction platform 141 of the graph prediction system 102. The exemplary method 800 may include one or more of the following steps.

[0094] A pre-treatment populated gene network graph for a plurality of patients and a post-treatment populated gene network graph for a plurality of patients may be received at steps 802 and 804, respectively. The post-treatment populated gene network graph may be generated at a predetermined period (e.g., a predetermined time delay) after the pre-treatment populated gene network graph. In some examples, a given patient may have multiple post-treatment populated gene network graphs at different time delays following the pre-treatment populated gene network graph. Time delay data indicating the time elapsed between the pre-treatment populated gene network graph and one or more post-treatment populated gene network graphs may be received for one or more of the plurality of patients for use in training. At step 806, a machine learning system, such as, for example, the training graph prediction platform 141, may receive the treatment regimens received by the plurality of patients. As discussed herein, the treatment regimens may include the type of treatment (e.g., monotherapy treatment, combination treatment), timing data, dosage data, etc. The timing data may include treatment time delay data, such as, for example, the time elapsed from diagnosis to initiation of treatment. Vector embeddings can be generated to describe or represent treatment regimens.

[0095] In step 808, a machine learning system, such as the training graph prediction platform 141, can receive outcome data for a plurality of patients. The outcome data can include clinical data such as overall patient survival, progression-free survival, Response Evaluation Criteria in Solid Tumors (RECIST), pathologic complete response data, or drug treatment outcomes, among other similar data.

[0096] In some cases, any or all of the pre-treatment data input gene network graph, post-treatment gene network graph, proposed treatment regimen, and / or outcome data can be vectorized in a vector format. The vector format of each input can be received by the machine learning system. Steps 802, 804, 806, and 808 can be performed simultaneously and / or separately.

[0097] In step 810, a machine learning system may be trained to infer the data-populated gene network graph and / or at least one treatment outcome following treatment. The machine learning system may be trained using one or more of the inputs from steps 802-808. The machine learning system may use any known method for training, such as, for example, supervised learning. The trained system may be output to a digital storage device, such as, for example, storage device 109.

[0098] In some examples, the supervised machine learning system may be trained using strong annotations (e.g., known patient outcomes from pre-treatment and post-treatment network graphs in response to a given treatment, and / or known changes in gene expression). In such examples, the supervised machine learning system may include a graph neural network, a transform neural network, a convolutional neural network (CNN), or a multi-layer perceptron (MLP), among other similar examples. To enable learning, the patient's pre-treatment gene sequence data (e.g., in the form of a pre-treatment network graph), the patient's post-treatment gene sequence data (e.g., in the form of a post-treatment network graph), the corresponding treatment regimen received by the patient, and outcome data may be provided as input to the machine learning system. Time delay data indicating the period between the pre-treatment gene network graph and the predicted post-treatment gene network graph, i.e., equal to the period between the patient's actual pre-treatment and post-treatment gene network graphs, for example, may also be provided as an input. The machine learning system can then output predicted post-treatment gene sequencing data that can be used to populate the post-treatment gene network graph and / or predicted patient outcome (e.g., with a pre-determined time delay, if optionally received). The predicted post-treatment gene network graph can be compared to the actual post-treatment gene network graph to determine patient losses or errors. Similarly, the predicted patient outcome can be compared to actual patient outcomes (e.g., obtained from clinical data). The actual post-treatment gene network graph and patient outcome can be part of a robust annotation of the training gene sequencing data, which corresponds to the proposed treatment regimen and indicates known gene expression changes (e.g., reduced methylation of driver genes in response to the administered treatment) and / or outcome data from the pre-treatment gene network graph to the post-treatment gene network graph.The machine learning system may be modified or altered based on the error (e.g., weights and / or biases associated with one or more nodes and / or layers may be adjusted) to improve the accuracy of the machine learning system. This process may be repeated until the loss or error falls below a predetermined threshold for each of the received or at least determined training proposed treatment regimens. In some examples, a portion of the training treatment regimens may be withheld and used to further validate or test the machine learning system.

[0099] Exemplary Use: Surrogate for Treatment Outcome Clinical outcomes directly measure whether patients in a clinical trial feel or function better, or live longer. The benefit or potential benefit of a treatment, as measured by the clinical outcome, can be evaluated to determine whether it outweighs the side effects. In some clinical trials, surrogate endpoints may be used instead of clinical outcomes if the clinical outcome will take a long time to study.

[0100] The embodiments disclosed herein can be used to identify oncogenic signaling pathways whose activation or inhibition is associated with or predictive of outcome as a suitable surrogate endpoint (e.g., predicting response to a drug, overall survival, progression-free survival, etc.). This can support clinical trial design by identifying patients who are likely to respond to a drug despite the absence of outcome data for the drug, based on the signaling pathway that the drug targets. This is particularly useful in the early stages of clinical trial design when interim outcome data are not available.

[0101] Exemplary Applications: Biomarker screening and development Screening for biomarkers derived from single genomic mutations can fail due to limited number of positive cases. In contrast, screening signaling pathways derived from multiple genes can increase sample size and allow screening of rare variants and / or rare tumor types. For example, the prevalence of mutations in each of the individual genes of the SWI / SNF complex (ARID1A, ARID1B, ARID2, PBRM1, SMARCA4, SMARCB1) can be low in some tumors, while the prevalence of mutations in the complex is collectively found in approximately 20% of all tumors.

[0102] The aspects disclosed herein can be used to identify signaling pathways as pharmacodynamic biomarkers, and also to identify predictive biomarkers for monotherapy and combination therapy, often when each drug targets different genes but the same pathway. Driver genetic variants and epigenetic variants associated with genes that contribute to functional disruption of the same signaling pathway or network can be integrated, which increases the number of positive cases and contributes to screening biomarkers at the functional signaling pathway or complex level.

[0103] Exemplary Use: Identifying Rare Tumor Subtypes Gene expression assays can be used to classify tumors. Tumor samples can be clustered based on gene expression profiles, and retrospective analysis identifies the clinical significance of each tumor subtype. For example, PAM50 gene expression assays contribute to revealing intrinsic subtypes of breast tumors, such as Luminal A, Luminal B, Basal-like, and normal subtypes of breast cancer. However, different numbers of genes used in the assay may also change the tumor subtype, and it may be difficult to identify rare tumor subtypes based on a limited set of genes.

[0104] The embodiments disclosed herein can be used to detect histological features associated with oncogenic signaling pathways, and pathway detection can be used as a complement or alternative to gene expression assays to aid in patient stratification. Tumor samples can be clustered based on computer-learned histological patterns associated with oncogenic pathway activation. Retrospective analysis combined with clinical information can be used to identify rare tumor subtypes and associated signaling pathways or gene complexes, providing guidance in evaluating treatment strategies for patients of various risk groups.

[0105] Exemplary Use: Estimating the risk of distant metastasis Metastasis is the main cause of cancer treatment failure and death. Adjuvant chemotherapy is often used for distant control. However, not all patients can benefit from adjuvant chemotherapy, especially some patients may further deteriorate after treatment. Assessing the risk of distant metastasis and identifying patients who may benefit from adjuvant chemotherapy for distant control facilitates treatment planning. There are certain mutations and signaling processes that may contribute to metastasis. As an example, Ras mutations are present in about 50% of metastatic tumors, and Ras proteins activate multiple downstream signaling pathways. As another example, epithelial-mesenchymal transition (EMT), a series of transitional steps between epithelial and mesenchymal phenotypes, not only allows cells to acquire a migratory phenotype but also induces the evasion of multiple immunosuppression, drug resistance, and apoptosis mechanisms.

[0106] The embodiments disclosed herein may be used to detect signaling pathways that regulate progression and promote the acquisition of a metastatic phenotype. For example, the above-described systems and methods may be used to identify histological patterns associated with Ras mutations and downstream signaling pathways, and predict activation of the Ras signaling pathway to infer risk of distant metastasis. Furthermore, the above-described systems and methods may be used to correlate the learned histological features with epithelial versus mesenchymal phenotypes, thereby detecting the type of transformation, whether epithelial to mesenchymal transformation (EMT) or mesenchymal to epithelial transformation (MET). These identifications and detections, combined with prognostic data, may help predict risk of distant metastasis after surgery and identify whether a patient would benefit from adjuvant chemotherapy.

[0107] Exemplary uses: Evaluation of therapeutic interventions and synthetic lethality Synthetic lethality is a method to target cancer cells with specific untreatable cancer mutations, a type of genetic interaction that simultaneously disrupts multiple genes, resulting in cell death. Synthetic lethality screening may identify new vulnerabilities caused by specific cancer mutations to develop new therapeutic approaches. Synthetic lethality effects at the pathway level are more reproducible than at the gene level.

[0108] The embodiments disclosed herein can be used to identify visual patterns associated with oncogenic signaling pathways, such as DNA damage response pathways, under normal and disease conditions (functional disruption). By detecting differences in signaling pathways before and after therapeutic intervention, the efficacy of treatment can be estimated and chemotherapy resistance can be further predicted. Furthermore, such identification and detection can facilitate screening of synthetic lethality strategies to develop more effective targeted drugs in cancer treatment.

[0109] Exemplary Use: Prediction of resistance to T cell receptor (TCR)-based immunotherapy TCR-based immunotherapy may have the potential to treat patients with various solid tumors, however, there are multiple pathways associated with TCR resistance, including loss of function, loss of heterozygosity, and epigenetic silencing of key genes involved in antigen processing, presentation, and interferon response pathways.

[0110] The embodiments disclosed herein can be used to identify histological patterns associated with TCR resistance (e.g., T cell receptor signaling pathway or B cell receptor signaling pathway) and to evaluate the efficacy of TCR-based therapy in combination with adjuvant immunotherapy including infusion of immune checkpoint inhibitors and immune stimulatory cytokines to overcome treatment resistance.

[0111] Exemplary Use: Predicting progression of tumors from benign to malignant Oncogenic drivers are found in normal tissues as well as in various benign diseases. Many factors can induce the change from a benign to a malignant state, including the tissue microenvironment, co-loss of genomic driver cofactors or tumor suppressors, size of the mutant clone, etc. For example, atypical endometrial hyperplasia is a precancerous condition that can develop in the lining of the uterus. It can be an overgrowth of abnormal cells or can arise from endometrial hyperplasia, which is an overgrowth of normal cells. Patients with atypical endometrial hyperplasia are at a significantly increased risk of developing endometrial cancer.

[0112] The embodiments disclosed herein can be used to identify histological patterns associated with genomic or epigenetic alterations in normal samples, atypical endometrial hyperplasia, and endometrial tumor samples. For example, pathway changes at different time points can be established to profile tumor progression. Furthermore, the minimal changes required for a particular oncogenic signaling pathway to cause a benign to malignant change can also be determined.

[0113] FIG. 9 illustrates an exemplary system or device 900 capable of implementing the techniques presented herein. The device 900 may include a central processing device (CPU) 920. The CPU 920 may be any type of processor device, including, for example, any type of dedicated or general-purpose microprocessor device. As will be appreciated by those skilled in the art, the CPU 920 may also be a single processor in a multi-core / multi-processor system operating alone or in a cluster of computing devices operating in a cluster or server farm. The CPU 920 may be connected to a data communications infrastructure 910, such as, for example, a bus, a message queue, a network, or a multi-core message passing scheme.

[0114] The device 900 may also include a main memory 940, such as, for example, a random access memory (RAM), and may also include a secondary memory 930. The secondary memory 930, such as, for example, a read only memory (ROM), may be, for example, a hard disk drive or a removable storage drive. Such removable storage drives may include, for example, a floppy disk drive, a magnetic tape drive, an optical disk drive, a flash memory, and the like. The removable storage drive in this example reads and / or writes to the removable storage unit in a well-known manner. The removable storage device may include a floppy disk, a magnetic tape, an optical disk, and the like, which is read and written by the removable storage device drive. As will be appreciated by those skilled in the art, such removable storage units typically include computer usable storage media having computer software and / or data stored thereon.

[0115] In alternative implementations, secondary memory 930 may include similar means for allowing computer programs or other instructions to be loaded into device 900. Examples of such means may include program cartridges and cartridge interfaces (such as those found in video game devices), removable memory chips (such as EPROMs or PROMs) and associated sockets, and other removable storage units and interfaces that allow software and data to be transferred from removable storage units to device 900.

[0116] Device 900 may also include a communications interface (COM) 960. Communications interface 960 allows software and data to be transferred between device 900 and external devices. Communications interface 960 may include a modem, a network interface (such as an Ethernet card), a communications port, a PCMCIA slot and card, or the like. The software and data transferred through communications interface 960 may be in the form of signals, which may be electronic, electromagnetic, optical, or other signals receivable by communications interface 960. These signals may be provided to communications interface 960 via a communications path of device 900, which may be implemented using, for example, wire or cable, fiber optics, a phone line, a cellular phone link, an RF link, or other communications channel.

[0117] The hardware elements, operating systems, and programming languages ​​of such equipment are conventional in nature and are presumed to be sufficiently familiar to those skilled in the art. The device 900 may also include input / output ports 950 for connecting input / output devices such as a keyboard, mouse, touch screen, monitor, display, etc. Of course, various server functions may be implemented in a distributed manner on multiple similar platforms to distribute the processing load. Alternatively, the server may be implemented by appropriate programming of one computer hardware platform.

[0118] Throughout this disclosure, references to components or modules generally refer to items that may be logically grouped together to perform a function or group of related functions. Like reference numbers are generally intended to refer to the same or similar components. The components and / or modules may be implemented in software, hardware, or a combination of software and / or hardware.

[0119] The tools, modules, and / or functions described above may be executed by one or more processors. "Storage" type media may include some or all of the tangible memory of a computer, processor, etc., or associated modules such as various semiconductor memories, tape drives, disk drives, etc. that may provide non-transitory storage for software programming at any time.

[0120] The software may be communicated over the Internet, a cloud service provider, or other telecommunications network. For example, the communication may enable the software to be loaded from one computer or processor to another. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0121] The foregoing general description is exemplary and explanatory only and is not intended to limit the disclosure. Other embodiments will be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only.

Claims

1. receiving one or more digital medical images associated with a patient; providing the pristine gene network graph and the one or more digital medical images as inputs to a trained machine learning system, the system being trained to populate the gene network graph with patient-specific gene expression levels based on the one or more digital medical images; receiving, as output from the trained machine learning system, the gene network graph populated with the gene expression levels specific to the patient; A method for processing digital medical images to populate a gene network graph, comprising:

2. receiving clinical data associated with the patient; providing the clinical data as additional input to the trained machine learning system; The method of claim 1 further comprising:

3. The method of claim 1 , wherein the digital medical image comprises a digital whole slide image, a digital multiplexed immunofluorescence image, or a digital multiplexed immunohistochemistry image.

4. the machine learning system, receiving as training data a plurality of digital medical images associated with a plurality of patients and a populated gene network graph for the plurality of patients; training the machine learning system using the training data to infer one or more of the populated gene network graphs based on the respective one or more digital medical images; The method of claim 1 , wherein the training is performed by

5. 5. The method of claim 4, wherein the training data further comprises clinical data associated with a plurality of patients, the clinical data comprising one or more of age, medical history, cancer treatment history, family history, previous biopsy or cytology information, tumor sequence information, and mRNA expression levels.

6. The gene network graph of the plurality of patients comprises: receiving an unpopulated gene network graph, wherein the unpopulated gene network graph includes a gene network graph without expression levels; receiving tumor sequence information associated with the plurality of patients; populating the gene network graph with expression levels of the plurality of patients based on the respective tumor sequence information; 5. The method of claim 4, wherein the data is input by

7. The data input of the gene network graph of the plurality of patients comprises: determining whether there are missing values ​​of expression levels in each of the populated gene network graphs; upon determining that one or more of the populated gene network graphs have missing values, using one or more label propagation techniques to infer the missing values; The method of claim 6, comprising:

8. The method of claim 7 , wherein the one or more label propagation techniques include directed label propagation.

9. 1. A method for training a machine learning system to populate a gene network graph, comprising: receiving an unpopulated gene network graph, wherein the unpopulated gene network graph includes a gene network graph without expression levels; receiving tumor sequence information associated with each of a plurality of patients; receiving one or more digital medical images associated with each of the plurality of patients; inputting data for each of the plurality of patients such that the gene network graph includes expression levels based on the respective tumor sequence information; training the machine learning system to infer one or more of the populated gene network graphs based on each of the one or more digital medical images; The method comprising:

10. 10. The method of claim 9, wherein the digital medical image comprises a digital whole slide image, a digital multiplexed immunofluorescence image, or a digital multiplexed immunohistochemistry image.

11. For each populated gene network graph, determining whether there are missing values ​​of expression levels within the gene network graph; if determining that there are missing values, using one or more label propagation techniques to infer the missing values; 10. The method of claim 9, further comprising:

12. The method of claim 11 , wherein the one or more label propagation techniques include directed label propagation.

13. 10. The method of claim 9, further comprising receiving clinical data associated with each of the plurality of patients, wherein the machine learning system is further trained to infer the one or more populated gene network graphs based on the respective clinical data, wherein the clinical data further comprises age, medical history, cancer treatment history, family history, previous biopsy or cytology information, tumor sequence information, mRNA expression levels, or a combination thereof.

14. 1. A system for processing digital medical images to populate a gene network graph, comprising: at least one memory for storing instructions; at least one processor configured to execute the instructions to perform operations; The operation is receiving one or more digital medical images associated with a patient; providing the pristine gene network graph and the one or more digital medical images as inputs to a trained machine learning system, the system being trained to populate the gene network graph with patient-specific gene expression levels based on the one or more digital medical images; receiving, as output from the trained machine learning system, the gene network graph populated with the gene expression levels specific to the patient; The system comprising:

15. receiving clinical data associated with the patient; providing the clinical data as additional input to the trained machine learning system; The system of claim 14 further comprising:

16. The system of claim 14 , wherein the digital medical image comprises a digital whole slide image, a digital multiplexed immunofluorescence image, or a digital multiplexed immunohistochemistry image.

17. the machine learning system, receiving as training data a plurality of digital medical images associated with a plurality of patients and a populated gene network graph for the plurality of patients; training the machine learning system using the training data to infer one or more of the populated gene network graphs based on the respective one or more digital medical images; The system of claim 14 , wherein the system is trained by:

18. 20. The system of claim 17, wherein the training data further comprises clinical data associated with a plurality of patients, the clinical data comprising one or more of age, medical history, cancer treatment history, family history, previous biopsy or cytology information, tumor sequence information, and mRNA expression levels.

19. The data input of the gene network graph of the plurality of patients comprises: determining whether there are missing values ​​of expression levels in each of the populated gene network graphs; upon determining that one or more of the populated gene network graphs have missing values, using one or more label propagation techniques to infer the missing values; 20. The system of claim 18, comprising:

20. 20. The system of claim 19, wherein the one or more label propagation techniques include directed label propagation.