Precision combination therapy using prediction of tumor clone response from cell data

An AI-driven prediction model using cell line perturbation data ranks therapeutic responses to develop combination therapies for tumor clones, addressing the challenge of resistance evolution and enhancing treatment efficacy.

DE112024001062T5Pending Publication Date: 2026-02-05INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE112024001062
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-01
Filing Date
2024-02-12
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing treatments for tumors fail to account for the evolution of tumor clones that develop resistance to therapies, necessitating a need for predicting combination therapies that target multiple resistance mechanisms across different clones.

Method used

A prediction model using artificial intelligence platforms and cell line perturbation data to rank responses to combination therapies, enabling the development of targeted treatments for tumor clones.

Benefits of technology

The model effectively predicts the strongest responses to therapeutic interventions, allowing for the design of combination therapies that reduce the risk of disease progression by targeting multiple clones simultaneously.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An AI platform is used to develop a combination therapy for a patient suffering from a tumor that has formed clones. The combination therapy, containing at least two disturbances, is capable of targeting clones (including subclones) that have evaded therapeutic intervention due to resistance and / or evolution. The AI ​​platform is trained on disturbance data obtained from at least one cell line that has similar properties to a relevant clone. The trained AI platform predicts how the relevant clone will respond to disturbances and ranks the disturbance responses from strongest to weakest. At least one cell line can be an existing cell line from an established database or a synthetic cell line generated by the AI ​​platform.The AI ​​platform can include one or more from a machine learning platform, a deep learning platform, an artificial neural network (ANN), a convolutional neural network (CNN), and a generative adversarial network (GAN).
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELDThe present invention relates generally to prediction models, and more particularly to a prediction model that uses artificial intelligence platforms and cell line perturbation data to predict combination therapies for tumor clones.BACKGROUND OF THE INVENTIONIn the course of disease treatment, tumors typically undergo evolution in response to treatments and therapies to avoid and / or resist elimination. The evolution of tumors is often detected by the detection of clones and subclones that have branched from a single tumor in response to therapeutic pressure. The changes in the resistance of the clones (including any subclones) allow them to escape therapy. The replication of such tumor heterogenity is that different clones derived from a single tumor precursor may have unique resistance mechanisms that do not respond to conventional cancer therapies. To prevent progression of the disease in a patient suffering from multiple tumor clones, a specific combination of therapies is required, each therapy within the combination being designed to target one or more of the resistance mechanisms across the different tumor clones. In view of the foregoing, there is a need in the art for the development of a technique that can predict how an existing tumor clone could respond to various therapeutic interventions.SUMMARY OF THE INVENTIONThe present invention fulfills the need of the prior art by providing a prediction model that utilizes perturbation data from existing cell line databases to predict the response of patient tumor clones to combination therapies that have at least two perturbations.In one embodiment, the present invention relates to a method comprising: sequencing a relevant clone obtained from at least one tumor lesion; identifying at least one cell line having properties similar to the relevant clone and composing a dataset comprising perturbation data for the at least one cell line; feeding the dataset to an artificial intelligence (AI) platform, wherein the dataset trains the AI platform to predict responses to disturbances included in the perturbation data; inputting information relating to the relevant clone into the trained AI platform and obtaining as output a ranking of the predicted disorder responses of the relevant clone included in the disorder data, wherein the predicted disorder responses from the strongest disorder response to the weakest disorder response are ranked; and developing a combination therapy for the relevant clone having disorders from at least two of the high-ranking disorder responses.In another embodiment, the present invention relates to a method comprising: sequencing a relevant clone obtained from at least one tumor lesion; identifying at least two existing cell lines that have similar characteristics to the relevant clone and assembling a first dataset comprising perturbation data for the at least two existing cell lines; feeding the first dataset to an AI platform, wherein the first dataset trains the AI platform to generate at least one synthetic cell line comprising perturbation data that is assembled from the at least two existing cell lines, wherein the perturbation data for the at least one synthetic cell line is assembled into a second dataset; applying the first and second data sets as training data for the AI platform to learn how responses to disturbances included in the disturbance data of the first and second data sets can be predicted; inputting information related to the relevant clone into the trained AI platform and obtaining as output a ranking of the predicted disturbance responses of the relevant clone included in the disturbance data of the first and second data sets, wherein the predicted disturbance responses are ranked from the strongest disturbance response to the weakest disturbance response; and developing a combination therapy for the relevant clone that includes disturbances from at least two of the high-ranking disturbance responses.In another embodiment, the present invention relates to a computer program product for ranking tumor clone failure responses, comprising: program instructions on one or more computer readable storage media for training a AI platform to predict responses to failures included in a dataset comprising failure data for at least one cell line having characteristics similar to a relevant clone; program instructions on one or more computer readable storage media for injecting information related to the relevant clone into the trained AI platform, wherein the trained AI platform predicts failure responses for the relevant clone to failures included in the failure data of the dataset; and program instructions on one or more computer readable storage media for issuing a ranking of predicted fault responses for the relevant clone from the strongest fault response to the weakest fault response from the AI platform.In another embodiment, the present invention relates to a computer program product for ranking tumor clone failure responses, comprising: program instructions on one or more computer readable storage media for training a AI platform to generate at least one synthetic cell line comprising failure data merged from failure data for at least two existing cell lines having similar characteristics to a relevant clone, wherein the failure data for the at least two existing cell lines are merged into a first data set and the failure data for the at least one synthetic cell line is merged into a second data set; program instructions on one or more computer readable storage media for training the AI platform to predict responses to the failures included in the failure data for the first and second data sets; Program instructions on one or more computer readable storage media for feeding information relating to the relevant clone to the trained Kl platform, wherein the trained Kl platform predicts failure responses for the relevant clone to the failures contained in the failure data of the first and second data sets; and program instructions on one or more computer readable storage media for outputting a ranking of the predicted failure responses for the relevant clone from the strongest failure response to the weakest failure response from the Kl platform.In another embodiment, the present invention relates to a system comprising: a first computer input data set comprising disruption data relating to at least two existing cell lines having properties similar to a sequence from a relevant clone obtained from at least one tumor lesion; a second computer input data set comprising disruption data relating to a synthetic cell line, wherein the at least one synthetic cell line comprises disruption data fused from the at least two existing cell lines; A AI platform that accepts as input the first dataset, the second dataset, and information related to the relevant clone, and provides as output a ranking of predicted perturbation responses for the relevant clone to perturbations included in the perturbation data of the first and second datasets, wherein the first dataset trains the AI platform to generate the perturbation data for the at least one synthetic cell line, the first and second datasets train the AI platform to predict the perturbation responses of the relevant clone to the perturbations included in the perturbation data of the first and second datasets, and the predicted perturbation responses for the relevant clone are ranked from the strongest perturbation response to the weakest perturbation response of the ranking.In another embodiment, the information relating to the relevant clone is selected from the group consisting of tumor type, tumor site, lesion site, genomics, transcriptomics, proteomics, microbiomics, metabolismomics, lipidomics, epigenomics and combinations thereof.In another embodiment, the characteristics of the cell lines similar to the clone of interest are selected from the group consisting of tumor type, disruption data, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics and combinations thereof.In another embodiment, the AI platform is selected from the group consisting of machine learning, deep learning, artificial neural networks, convolutional neural networks, generating concurrent networks, and combinations thereof.In another embodiment, the disorders for combination therapy are selected from the group consisting of environmental stimuli, drug inhibition, genome editing, disease treatment, and combinations thereof.Additional embodiments and / or aspects of the invention are provided without limitation in the detailed description of the invention as set forth below.BRIEF DESCRIPTION OF THE DRAWINGSFigure 1 is a schematic diagram illustrating a predictive model for tumor clone responses described herein. Figure 2 is a schematic representation of a computing environment that can be used to implement the prediction model for tumor clone responses described herein.DETAILED DESCRIPTION OF THE INVENTIONThe following is a description of what is presently considered to be preferred aspects and / or embodiments of the claimed invention. It is intended that any alternatives or modifications in function, purpose or structure be covered by the appended claims. As used in this specification and the appended claims, the singular forms of terms such as the articles "a / an" and "the / s" also include the plural forms unless the context expressly dictates otherwise. As used in the specification and the appended claims, the terms "comprise," "comprises," "comprises," and / or "comprising" indicate the presence of expressly stated components, elements, features, and / or steps, but do not preclude the presence or addition of one / more other components, elements, features, and / or steps.As used herein, the term "lesion" refers to an abnormal tissue region, and the term "tumor lesion" refers to an abnormal tissue region containing a complete tumor, a portion of a tumor, or tumor cells.As used herein, the term "clone" refers to a collection of cells having the same genomic profile that are considered homogeneous. In the context of tumor lesions, different clonal populations occurring within a single tumor lesion are "subclones" if they share the same precursor clone; precursor tumor clones and subclones are thus identical since they are both homogeneous groups of tumor cells whose therapeutic response would be expected to be identical. Based on the above, the term "relevant clones" as used herein refers to precursor tumor clones and derived subclones.As used herein, the term "disorder" refers to a change in the function of a biological system by external or internal means. Examples of disorders include, but are not limited to, environmental stimuli, drug inhibition, genome editing, and disease treatment.Examples of environmental stimuli include, but are not limited to, temperature changes, osmotic shock, and pressure changes.Drug inhibition alters a biological pathway by binding a small molecule to an active and / or allosteric binding site of an enzyme. Examples of substances that inhibit the growth of cancer cells include, but are not limited to, tyrosine kinase inhibitors, protein kinase inhibitors (e.g., mTOR and PI3K inhibitors), proteosome inhibitors, histone deacetylase inhibitors, hedgehog signaling pathway inhibitors, BRAF inhibitors, and MEK inhibitors.Examples of genome editing include, but are not limited to, gene knockout (gene removal or deactivation), gene knockdown (gene expression gene mitigation), gene knockup (insertion of a protein-encoding cDNA sequence at a gene position), and CRISPR (clustered, regularly interspaced, short palindromic repeats, clustered, regularly distributed, palindromically repeating DNA segments with which DNA can be very specifically verified, removed, and replaced). In the context of the present invention, the biological system is a cell line or cells obtained from a tumor lesion, a tumor clone or a tumor subclone.Examples of disease treatment include, but are not limited to, cell therapy, immunotherapy, and hormone therapy. Examples of cells that may be used in cell therapy include, but are not limited to, autologous cells (derived from the patient), allogeneic cells (derived from a donor), pluripotent stem cells (may differentiate into any other cell in a living organism), multipotent stem cells (may differentiate into all cell types of a particular line), unipotent stem cells (may produce only one cell type but may self-renew), adult stem cells (are responsible for maintaining and repairing the tissue in which they are located), primary cells (final differentiated cells isolated directly from living tissue), secondary cells (cell lines that have been immortalized and may divide infinitely), and combinations thereof. Examples of immunotherapy include, but are not limited to, monoclonal antibodies, non-specific immunotherapy (e.g., cytokines and bacterial therapy), oncolytic viral therapy, T cell therapy, and cancer vaccines. Examples of hormone therapy include, but are not limited to, aromatase inhibitors (Als), estrogen receptor antagonists, selective estrogen receptor modulators (SERMs), luteinizing hormone release hormone (LHRH), anti-androgens, CYP17 inhibitors, progestins and adrenololytics.As used herein, the terms "artificial intelligence", "AI", and "AI platform" refer to a computer algorithm that learns through experience. Examples of AI platforms include, but are not limited to, machine learning, deep learning, and neural networks. Some AI platforms that may be used to implement the present invention are discussed below.As used herein, the term "machine learning" refers to an artificial intelligence (AI) function in which an algorithm learns from training data to make predictions or decisions without having been explicitly programmed therefor. Machine learning is divided into three categories: supervised learning, semi-supervised learning, and unsupervised learning. During supervised learning, a computer algorithm is trained with labeled data. In practice, the computer algorithm receives example inputs and its desired outputs, both provided to the computer algorithm, and the goal of the computer algorithm is to learn a general rule that associates inputs with outputs. In unsupervised learning, a computer algorithm is trained with unlabeled data, leaving the computer algorithm itself left to learn how its input can be structured by its learning algorithm. The goal of unsupervised learning is to identify patterns hidden by feature learning in data to reach a final goal. In semi-supervised learning, a computer algorithm is trained with a small amount of labeled and a large amount of unlabeled data.As used herein, the term "deep learning" refers to an AI function that mimics the functioning of the human brain in processing and classifying data. Deep learning-based AI is capable of learning from unstructured and unlabeled data. In practice, deep learning-based AI programs find correlations between inputs and outputs by learning to approximate an unknown function (f(x)=y) between any input x and any output y, assuming there is a correlation or causality relationship therebetween.As used herein, the term "artificial neural network" or "ANN" refers to a deep learning AI function that simulates the human brain and includes a collection of simulated neurons, all of which are fully interconnected. Each neuron is a node that is connected to other nodes via connections analogous to Axon synapse-dendritic biological connections. Each link has a weight that determines the strength of the influence of one node on another node. Neural networks learn (i.e., are trained) by processing examples, each of which contains a known input and output and forms probability-weighted associations between the input and the output. In practice, a neural network groups unlabeled input data according to similarities between example inputs, automatically extracts features from the groups, clusters groups with similar features, and classifies output data when a labeled dataset is present for training. The patterns recognized by a neural network are numerical in nature and are included in vectors that need to be translated. Examples of vectors of neural networks include, but are not limited to, images, sound, text, time indications, or combinations thereof.As used herein, the term "convolutional network" or "convolutional neural network" (CNN) refers to a neural network that uses convolutional layers to convolutionally fold an input and forward its result to the next layer. While fully connected feedforward neural networks may be used to learn features and classify data, CNNs regularize fully connected layered networks in which each neuron in one layer is connected to all neurons in the next layer. The fully connected layers of a CNN result in an overfitting of the data that must be regularized by their nature. CNNs regularize by utilizing the hierarchical pattern in data to arrive at more complex patterns using smaller and simpler patterns. A CNN is comprised of an input layer, hidden middle layers, and an output layer. The hidden middle layers perform the folds. After traversing a convolution layer, the data is abstracted into a feature map, which is the output of the convolution kernel that was applied to the previous layer.As used herein, the terms "GAN" (generative adversarial network) and "GAN model" refer to a machine learning function having two neural networks, a generator, and an opponent (also referred to herein as a discriminator) that compete with each other in the form of a zero sum game, where the gain of the one neural network is the loss of the other neural network. A GAN is based on indirect training of the generator using the discriminator, wherein the generator generates candidates that the discriminator evaluates. For the evaluation, the discriminator notifies the generator how realistic the generator input seems to be. The training goal of the generator is to increase the error rate of the discriminator by creating new candidates in which the discriminator assumes that they are not synthesized (i.e., they are part of the true data distribution). A known data set serves as the original training data for the discriminator, wherein the training involves presenting samples from the training data set to the discriminator until the discriminator reaches an acceptable accuracy. The generator is trained based on whether it can successfully fürce the discriminator, so that it assumes that the synthetic data is the existing data. A GAN generator network typically begins with a randomized input sampled from a predefined latent space, e.g., a multivariate normal distribution. Thereafter, candidates synthesized by the generator are evaluated by the discriminator. The machine learning function used for a GAN may be associated with one of the three categories: fully supervised, semi-supervised, or unsupervised machine learning. The two neural networks of the GAN may be two ANNs, two CNNs, or a combination of an ANN and a CNN.The present invention trains a Kl platform to predict how a tumor clone will behave in response to various disorders by training the Kl platform with cell line disorder data. Because a tumor usually has multiple clones and a patient may have multiple tumors, the clonal-level predictive response generally requires a combination therapy that has at least two disorders to target the behavior of the multiple clones as they develop from their tumor precursor. To the inventors' knowledge, the present invention is the first instrument for predicting tumor clone perturbation responses.FIG. 1 is a schematic illustration of the workflow required to implement the prediction model described herein. As a starting point, at least one tumor lesion (K=I 1, I 2,... I k) is taken from a single patient and sequenced to determine the clonal composition of the at least one tumor lesion, e.g., precursor tumor clones and subclones derived from the precursor tumor clones. Relevant clones are identified from the sequencing data based on selected characteristics, e.g., but not limited to, tumor type (e.g., tumor type associated with the tumor, e.g., bladder cancer tumor, breast cancer tumor, colon cancer tumor, kidney cancer tumor, lung cancer tumor, etc.), tumor location, lesion location, andomic data. Examples ofomic data include, but are not limited to, genomics data (information on genes), transcriptomics data (information on RNA), proteomics data (information on proteins), microbiomics data (information on microorganisms such as bacteria, fungi and viruses), metabolism data (information on metabolites), lipidomics data (information on lipids) and epigenomics data (information on methylated DNA or modified histone proteins). Thereafter, one or more cell lines are identified from one or more established databases sharing the properties with a relevant clone. It is to be understood that one or more cell lines may share properties with a single clone of interest. Examples of databases that may be used to identify the at least one cell line include, but are not limited to, the Sanger databases (available at https: / / cancer.sanger.ac.uk / cell_Iins), the DepMap portal (available at https: / / depmap.org / portal / ), the CCLE (Cancer Cell Line Encyclopedia, available at https: / / sites.broadinstitute.org / ccle / ), the Library of Network-based Cellular Signatures (LINCS, available at https: / / lincsproject.org / ), and combinations thereof.The Sanger databases include Catalog of Somatic Mutations in Cancer (a somatic mutation database trained by experts), Cell Lines Project (mutation profiles of over 1000 cell lines used in cancer research), COSMIC-3D (including an interactive view of cancer mutations associated with 3D structures), Cancer Genes Census (a catalog of genes with mutations that are causally associated with cancer), Cancer Mutation Census (a classification of genetic variants that promote the development of cancer), and Actionability (mutations that are relevant to precision oncogene).The DepMap portal provides a cancer dependency map containing genetic and pharmacological dependencies, tumor contexts, predictive biomarkers, and more than 2000 cancer models.The CCLE contains cell line annotations for more than 1000 human cancer models, fused mutation identifications for 329 cell lines, RNA expression data for 1019 cell lines, fusion identifications for 1019 cell lines, epigenetic and histone modification data, proteomic data, and metabolomic data. In addition to the above examples, any private cell line database may be used (e.g., databases owned by biotechnology or pharmaceutical companies). In one embodiment, each of the above databases can be used to identify at least one cell line that is similar to the relevant clone by comparing the genomic profile of the relevant clone to the genetic profiles of the at least one cell line.The LINCS database identifies and categorizes molecular signatures that occur when cells are exposed to substances that interfere with their normal function.In some situations, there may be a few cell lines in the known databases (also referred to herein as "existing cell lines") that have a high degree of agreement with a relevant clone. Such situations may include a type of cancer which has a wide variety of tumor clones across all cancer patients or clones which are unique. In such situations, it becomes statistically difficult to identify recurrent features and patterns to determine relevant existing cell lines from which predictions can be made. To identify mechanisms shared between clones as well as disorders that may target multiple clones, existing cell lines similar to two or more relevant clones may be joined into one or more synthetic cell lines. The one or more synthetic cell lines allow the generation of predictions about the tumor lesion as a whole (since a tumor lesion is a representation of one or more relevant clones). It is to be understood that the cell lines used in the prediction model described herein may be existing cell lines and / or synthetic cell lines depending on the type of tumor lesions producing the relevant clone.After the one or more cell lines are selected from the existing databases, information relating to the one or more cell lines is compiled in a table format to be input as training data to an AI platform. Examples of information relating to the one or more cell lines include, but are not limited to, tumor type, perturbation data, andomic profiles. Examples of perturbation data include, but are not limited to, the response of the one or more cell lines to various environmental stimuli, drug inhibition, genome editing, and disease treatment. Examples ofomic profiles include, but are not limited to, genomics, transcriptomics, proteomics, microbiomics, metabolism, lipidomics and epigenomics. Examples of table data formats that may be used for input include, but are not limited to, comma separated value (CSV) files and spreadsheets (e.g., EXCEL ®, Microsoft Corporation, Redmond, WA, USA; NUMBERS ®, Apple Inc., Cupertino, CA, USA).To create a synthetic cell line, the information relating to at least two existing cell lines is placed in a table format and entered into the Kl platform, along with instructions for the Kl platform, how the information relating to the at least two cell lines is to be merged into one or more synthetic cell lines. In one embodiment, the KI platform may be a GAN that uses these two or more cell lines as input to generate one or more synthetic cell lines withomic characteristics similar to the original input.The AI platform is trained with the cell line perturbation data in the table format to learn the responses of the one or more cell lines (both present and synthetic) to the disturbances. Training the cell line disruption data allows the AI platform to predict how the relevant clones (which have similar characteristics to the cell lines) will respond to different disruption. Any AI platform may be used to generate the perturbation responses of the relevant clone, including, but not limited to, a machine learning platform, a deep learning platform, a neural network, a CNN, a GAN, and combinations thereof.After the predicted perturbation reactions for a relevant clone are obtained, the perturbation reactions from the strongest perturbation reaction to the weakest perturbation reaction are ranked. The rankings of the disorder responses are used together to develop a combination therapy designed for treatment of the relevant clone. In one embodiment, the combination therapy for each individual relevant clone generally contains the highest-rank disorders from two or more of the disorders described herein. In another embodiment, the synergy and / or toxicity of drugs is analyzed as needed to ensure that combination therapy avoids deleterious side effects or interactions. By providing a combination therapy designed to safely target all clones of a patient, the risk of disease progression for the patient can be significantly reduced.For example, in a patient with multiple tumor lesions that each elicited their own clones, the presence of the different heterogeneous clones in one and the same patient may require different combinations to target the different resistance mechanisms of the heterogeneous clones. For example, any two heterogeneous clones tested with the prediction model described herein can be predicted for one clone to respond to a combination therapy with the highest-rank genome edit disorder by knockout genome edition and the highest-rank disease treatment disorder by treatment with pluripotent stem cells, while another clone can be predicted to respond to a combination therapy with the highest-rank environmental disorder by osmotic shock, the highest-rank drug inhibition disorder by tyrosine kinase inhibitors, and the highest-rank disease treatment disorder by monoclonal antibody therapy.In one embodiment, the present invention comprises sequencing a relevant clone obtained from at least one tumor lesion; identifying at least one cell line having properties similar to the relevant clone and composing a dataset comprising perturbation data for the at least one cell line; feeding the dataset to an AI platform, the dataset training the AI platform to predict responses to perturbations included in the perturbation data; inputting information relating to the relevant clone to the trained AI platform and obtaining as output a ranking of the predicted perturbation responses of the relevant clone included in the perturbation data, wherein the predicted perturbation responses are ranked from the strongest perturbation response to the weakest perturbation response; and developing a combination therapy for the relevant clone that has disorders from at least two of the high-rank disorder responses.In another embodiment, the present invention comprises sequencing a relevant clone obtained from at least one tumor lesion; identifying at least two existing cell lines that have similar characteristics to the relevant clone and assembling a first dataset comprising perturbation data for the at least two existing cell lines; feeding the first dataset to an artificial intelligence (AI) platform, wherein the first dataset trains the AI platform to generate at least one synthetic cell line comprising perturbation data that is assembled from the at least two existing cell lines, wherein the perturbation data for the at least one synthetic cell line is assembled into a second dataset; applying the first and second data sets as training data for the AI platform to learn how responses to disturbances included in the disturbance data of the first and second data sets can be predicted; inputting information related to the relevant clone into the trained AI platform and obtaining as output a ranking of the predicted disturbance responses of the relevant clone included in the disturbance data of the first and second data sets, wherein the predicted disturbance responses are ranked from the strongest disturbance response to the weakest disturbance response; and developing a combination therapy for the relevant clone having disturbances from at least two of the high-ranking disturbance responses.In another embodiment, the present invention comprises program instructions on one or more computer readable storage media for training a AI platform to predict responses to disturbances included in a dataset comprising disturbance data for at least one cell line having characteristics similar to a relevant clone; program instructions on one or more computer readable storage media for injecting information relating to the relevant clone into the trained AI platform, wherein the trained AI platform predicts disturbance responses for the relevant clone to disturbances included in the disturbance data of the dataset; and program instructions on one or more computer readable storage media for issuing a ranking of predicted fault responses for the relevant clone from the strongest fault response to the weakest fault response from the AI platform.In another embodiment, the present invention comprises program instructions on one or more computer readable storage media for training a AI platform to generate at least one synthetic cell line comprising perturbation data merged from perturbation data for at least two existing cell lines having similar characteristics to a relevant clone, wherein the perturbation data for the at least two existing cell lines are merged into a first data set and the perturbation data for the at least one synthetic cell line is merged into a second data set; program instructions on one or more computer readable storage media for training the AI platform to predict responses to the perturbations included in the first and second data sets; Program instructions on one or more computer readable storage media for feeding information relating to the relevant clone to the trained Kl platform, wherein the trained Kl platform predicts failure responses for the relevant clone to the failures contained in the first and second data sets; and program instructions on one or more computer readable storage media for outputting a ranking of the predicted failure responses for the relevant clone from the strongest failure response to the weakest failure response from the Kl platform.In another embodiment, the present invention comprises a first computer input data set comprising disruption data relating to at least two existing cell lines having properties similar to a sequence from a relevant clone obtained from at least one tumor lesion; a second computer input data set comprising disruption data relating to a synthetic cell line, wherein the at least one synthetic cell line comprises disruption data fused from the at least two existing cell lines; A AI platform that accepts as input the first dataset, the second dataset, and information related to the relevant clone, and provides as output a ranking of predicted perturbation responses for the relevant clone to perturbations included in the perturbation data of the first and second datasets, wherein the first dataset trains the AI platform to generate the perturbation data for the at least one synthetic cell line, the first and second datasets train the AI platform to predict the perturbation responses of the relevant clone to the perturbations included in the perturbation data of the first and second datasets, and the predicted perturbation responses for the relevant clone are ranked from the strongest perturbation response to the weakest perturbation response of the ranking.In another embodiment, the information relating to the clone of interest is selected from tumor type, tumor location, lesion location,omic profiles, and combinations thereof.In another embodiment, the characteristics of the existing cell line that are similar to the clone of interest are selected from tumor type, disruption data,omic profiles, and combinations thereof.In another embodiment, the disorders for combination therapy are selected from the group consisting of environmental stimuli, drug inhibition, genome editing, disease treatment, and combinations thereof.In another embodiment, the AI platform is selected from the group consisting of machine learning, deep learning, ANNs, CNNs, GANs, and combinations thereof.In another embodiment, the AI platform includes a GAN, alone or in combination with an ANN or a CNN.The present invention has applications in personalized medicine and in patient screening. In personalized medicine, the predictive model described herein is useful in developing treatments specific to a person having at least one type of tumor. In patient screening, the predictive model can be used to examine patients for participation in clinical studies based on the clonal composition of the patient's tumor lesions.Various aspects of the present disclosure are described by descriptive text, flowcharts, block diagrams of computer systems and / or block diagrams of machine logic included in computer program product (CPP) embodiments. With respect to any flow charts, depending on the technology involved, the operations may be performed in a different order than shown in a particular flow chart. For example, depending also on the technology involved, two operations shown in successive flowchart blocks may be performed in the reverse order or as a single, integrated step, simultaneously or in a manner that at least partially overlap in time.A computer program product ("CPP" or "CPP") embodiment is a term used in the present disclosure to describe any set of one or more storage media (also referred to as "media") that are commonly included in a set of one or more storage devices that commonly include machine readable code corresponding to instructions and / or data for performing computer operations described in a particular CPP claim. A "storage unit" is any tangible unit that can contain and store instructions for use by a computer processor. The computer readable storage medium may include, but is not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination thereof. Some known types of memory units containing these media include: a floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a static random access memory (SRAM), a compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded unit (e.g., punch cards or pits / bumps on a larger surface of a data carrier) or any suitable combination thereof. A computer readable storage medium as used in the present disclosure is not intended to be storage in the form of transitory signals per se, e.g., radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, light pulses passing through an optical fiber cable, electrical signals transmitted through a wire, and / or other transmission media. It will be appreciated by those skilled in the art that data is typically moved during normal operation of a storage unit at certain times, such as during access, defragmentation or upon performance of a clean-up function, but the storage unit thereby becomes non-transitory, as the data is non-transitory while being stored.The following discussion refers to FIG. 2. a computing environment 100 includes an example of an environment for executing at least a portion of the computer code involved in performing the inventive methods, e.g., the computer code required for executing machine learning, deep learning, ANNs, CNNs, and GANs to predict subclone fault responses, as described herein at 200. In addition to the block 200, the computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end user device (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, the computer 101 includes a processor set 110 (including processing circuitry 120 and cache 121), a communication structure 111, volatile memory 112, persistent storage 113 (including operating system 122 and block 200 identified above), a peripheral unit set 114 (including user interface unit set 123 (UI), memory 124, and internet of things (loT) sensors set 125. Remote server 104 includes remote database 130. The public cloud 105 includes a gateway 140, a cloud orchestration module 141, a host physical machine set 142, a virtual machine set 143, and a container set 144.The COMPUTER 101 may be in the form of a desktop computer, laptop computer, tablet computer, smart phone, smart watch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device that is currently known or developed in the future and that is capable of executing a program, accessing a network, or retrieving a database such as the remote database 130. As is well known in computer technology, and depending on the technology, the performance of a computer-implemented method may be distributed among multiple computers and / or multiple locations. However, in this discussion of the computing environment 100, the detailed discussion will focus on a single computer, specifically the computer 101, to keep the discussion as simple as possible. The computer 101 may be located in a cloud, although it is not shown in a cloud in FIG. 2. However, the computer 101 does not necessarily have to be located in a cloud, unless this is expressly described in this way.The PROCESSOR SET 110 includes one or more computer processors of any type presently known or developed in the future. The processing circuit 120 may be distributed among multiple packages, for example, among multiple coordinated IC chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory residing in the processor chip package and is typically used for data or code that should be available for rapid access by the threads or cores executing in processor set 110. Cache memories are usually organized in multiple levels depending on the relative proximity to the processing circuitry. Alternatively, the cache for the processor set may be entirely or partially "off-chip.". In some computing environments, the processor set 110 may be configured to process qubits and perform quantum computing.Computer readable program instructions are typically read into the computer 101 to cause a sequence of operations to be performed by the processor set 110 of the computer 101 to thereby effect a computer-implemented method such that the instructions performed thereby instantiate the methods performed in flowcharts and / or illustrative descriptions of computer-implemented methods included in this document (collectively referred to as "the inventive methods"). These computer readable program instructions are stored in various types of computer readable storage media, e.g., cache 121 and the other storage media discussed below. The program instructions and associated data are accessed by the processor set 110 to control and direct the inventive methods to be performed. In the computing environment 100, at least some of the instructions for performing the inventive methods may be stored in the persistent storage 113 at the block 200.The DATA TRANSMISSION STRUCTURE 111 constitutes the signal line path that allows the various components of the computer 110 to communicate with each other. This structure usually consists of switches and electrically conductive paths, e.g. the switches and electrically conductive paths forming buses, bridges, physical input / output ports and the like. Other types of signal transmission paths may also be used, e.g., optical fiber data transmission paths and / or wireless data transmission paths.The VOLATILE MEMORY 112 is any type of volatile memory currently known or developed in the future. Examples thereof include a dynamic type random access memory (RAM) or a static type RAM. The volatile working memory 112 is usually characterized by a random access, although this need not necessarily be the case unless expressly described in this way. In the computer 101, the volatile memory 112 is in a single package and is located within the computer 101, but alternatively or additionally the volatile memory may be distributed among a plurality of packages and / or may be located outside the computer 101.PERSISTENT MEMORY 113 is any type of non-volatile memory for computers that is currently known or will be developed in the future. The non-volatility of this memory means that the stored data is maintained regardless of whether the computer 101 and / or directly the persistent memory 113 are powered. The persistent memory 113 may be read-only memory (ROM), but typically at least a portion of the persistent memory allows writing data, erasing data, and rewriting data. Certain common forms of persistent storage include magnetic disks and semiconductor memory devices. The operating system 122 may take several forms, such as various known, proprietary operating systems or open source POSITION (Portable Operating System Interface) type operating systems that use a kernel. The code contained in the block 200 typically includes at least a portion of the computer code involved in performing the inventive methods.The PERIPHERAL UNIT SET 114 includes the set of peripheral units of the computer 101. Communications links between the peripherals and the other components of the computer 101 may be implemented in various ways, such as Bluetooth links, near-field communication (NFC) links, wired links (e.g., USB (Universal Serial Bus) type links), deployment type links (e.g., secure digital card, SD), connections over local communications networks, and even connections over wide area networks such as the Internet. In various embodiments, the UI unit set 123 may include components such as a display screen, a speaker, a microphone, wearable units (e.g., glasses and smart watches), a keyboard, a mouse, a printer, a touchpad, game controllers, and haptic units. The memory 124 is an external memory such as an external hard disk drive or a usable memory such as an SD card. The memory 124 may be persistent and / or volatile. In some embodiments, the memory 124 may be in the form of a quantum computing memory unit for storing data in the form of qubits. In embodiments where the computer 101 needs a large amount of memory (e.g., when the computer 101 locally stores and manages a large database), that memory may then be provided by peripheral storage devices configured to store very large amounts of data, e.g., a storage area network (SAN) shared by numerous geographically distributed computers. The IoT sensor set 125 consists of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer and another sensor may be a motion detector.The NETWORK MODULE 115 is the collection of computer software, hardware, and firmware that allows the computer 101 to communicate with other computers via the WAN 102. The network module 115 may include hardware, e.g., modems or WLAN signal transceivers, software for packaging and / or unpackaging data for transmission over a data transmission network, and / or web browser software for transmitting data over the Internet. In some embodiments, network control and forwarding functions of network module 115 are performed in the same physical hardware unit. In other embodiments (e.g., embodiments using software defined networking (SDN)), the control and forwarding functions of network module 115 are performed in physically separate entities so that the control functions can manage multiple different hardware network entities. Computer readable program instructions for performing the inventive methods may typically be downloaded to the computer 101 from an external computer or storage device via a network adapter card or network interface included in the network module 115.The WAN 102 is a wide area network (e.g., the Internet) capable of transmitting computer data over distances beyond the local frame by any technology for transmitting data currently known or developed in the future. In some embodiments, the WAN 102 may be replaced and / or supplemented with local area networks (LANs) configured to transmit data between entities located in a local area network, e.g., a WLAN network. The WAN and / or LANs may typically include computer hardware such as copper transmission cables, lightwave transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.The END USER UNIT (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of a business operating the computer 101) and may take any of the forms discussed above in connection with the computer 101. The EUD 103 typically receives helpful and useful data from the operation of the computer 101. For example, in a hypothetical case where the computer 101 is configured to provide a recommendation to an end user, this recommendation is typically transmitted from the network module 115 of the computer 101 to the EUD 103 via the WAN 102. In this manner, the EUD 103 may display or otherwise present the recommendation to an end user. In some embodiments, the EUD 103 may be a client device, e.g., a thin client, a heavy client, a mainframe computer, a desktop computer, and so forth.The REMOTE SERVER 104 is any computer system that provides at least some data and / or functionality to the computer 101. Remote server 104 may be controlled and used by the same entity that operates computer 101. Remote server 104 represents the one or more machines that collect and store helpful and useful data for use by other computers, e.g., computer 101. In a hypothetical case where the computer 101 is configured and programmed to provide a recommendation based on historical data, this historical data may be provided to the computer 101 from, for example, a remote database 130 of the remote server 104.The PUBLIC CLOUD 105 is any computer system available for use by multiple entities that provides demand-controlled availability of computer system resources, particularly data storage (cloud storage) and computing power, without direct, active management by the user. Cloud computing typically relies on sharing resources to achieve coherence and scale effects. The direct and active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically realized by virtual computing environments executing on different computers that make up the computers of the host physical machine set 142, which are the entirety of physical computers present in and / or available from the public cloud 105. The virtual computing environments (VCEs) are typically in the form of virtual machines from the set 143 of virtual machines and / or containers from the container set 144. It will be appreciated that these VCEs may be stored as images and transmitted within and between the various host physical machines, either as images or after instantiation of the VCE. The cloud orchestration module 141 manages the transfer and storage of images, provides new instantiations of VCEs, and manages active instantiations of VCE revenues. Gateway 140 is the collection of computer software, hardware, and firmware that enables public cloud 105 to exchange data via WAN 102.A more detailed explanation of virtualized computing environments (VCEs) is provided below. VCEs may be stored as "images.". A new active instance of the VCE may be instantiated from the image. Two common types of VCEs are virtual machines and containers. A container is a VCE that uses operating system level virtualization. This relates to an operating system feature in which the kernel enables the presence of multiple isolated user domain instances, referred to as containers. These isolated user area instances typically behave like real computers from the point of view of programs executed therein. A computer program executing on a common operating system may use all resources of the corresponding computer, e.g., connected devices, files, and folders, shared network areas, CPU power, and quantifiable hardware capabilities. Programs executed within a container, however, may only use the contents of the container and entities assigned to that container, a feature known as containerization.The PRIVATE CLOUD 106 is similar to the public cloud 105, except that the computing resources are only available for use by a single enterprise. Although the private cloud 106 is described as being in communication with the WAN 102, in other embodiments, a private cloud may be completely disconnected from the Internet and only accessible via a local / private network. A hybrid cloud is a composition of several clouds of different types (for example of the private, community or public cloud type), which are frequently implemented by different providers in each case. Each of the plurality of clouds remains a separate and distinct entity, but the larger hybrid cloud architecture is held together by standardized or proprietary technology that enables orchestration, management, and / or data / application portability between the plurality of clouds involved. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.The descriptions of the various aspects and / or embodiments of the present invention have been provided for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Those skilled in the art will appreciate that numerous modifications and variations are possible without departing from the spirit and scope of the described embodiments. The terminology used herein was chosen to best explain the principles of aspects and / or embodiments, the practical application or technical improvement over technologies found in the market, and to enable others of ordinary skill in the art to understand the aspects and / or embodiments disclosed herein.TEST ARRANGEMENTThe following examples are set forth to provide a full disclosure to those skilled in the art of how the aspects and embodiments of the invention set forth herein may be realized and utilized. Although efforts have been made to ensure the accuracy of the variables, trial errors and deviations should be taken into account.EXAMPLE 1Four tumor lesions are taken from the liver (2 samples), brain (1 sample) and subcutaneous soft tissue (1 sample) of a patient afflicted with metastatic colon cancer. Cells from the four tumor lesions are sequenced. The cells of the four tumor lesions are subjected to clonal analysis using the open source software program Contolerti (available at https: / / gitthub.com / ComputationalGenom / Contolerti), and two scrambler clones are identified: one clone with a KRAS p.G12S allele and the other clone with an ELF3 p.S229R allele. The clone with the KRAS p.G12S allele develops two subclones, both of which have the BCLAF p.S496L allele and are specific for liver tissue. The clone with the ELF3 p.S229R allele also develops two subclones, both of which also have the ELF3 p.S229R allele, one however being specific for brain tissue and the other being specific for subcutaneous soft tissue. All four clones are identified as relevant clones.For each of the four relevant clones, tumor type, tumor location, lesion location, genomic profile, and otheromic profiles of the clone are used to identify ten cell lines in the Sanger, DepMap, CCLE, and LINCS databases that are similar to the relevant subclones. The information obtained from the cell line databases is obtained as a csv file. The cell line information contains tumor type, disruption data, genomic profile, and otheromic profiles. The csv file of all ten cell lines is applied as input to a GAN model to generate synthetic cell line data with genomic profiles similar to the identified cell lines to supplement and extend the perturbation dataset. The similarity between the relevant clone and the synthetic and existing cell lines is measured to more accurately determine the match of these cell lines and their disruption data to the relevant clone. Subsequently, a machine learning model such as an ANN is used to predict a malfunction response by training the model with the existing and synthesized data. After the ANN has been trained, the tumor type, tumor location, lesion location, genomic profile and otheromic profiles of the relevant clones are provided to the ANN and the output from the ANN is a ranking of predicted perturbation responses for the relevant clone. This process is repeated for each relevant clone to treat each clone individually.The GAN model is generated using the open source deep learning frameworks PyTor, Tensorflow, and / or KERAS, all of which use the computer programming language Python and allow the generation of the GAN with a training function as well as an ANN and / or CNN generator and discriminator. The GAN may be generated with an ANN generator and an ANN discriminator, a CNN generator and a CNN discriminator, an ANN generator and a CNN discriminator, or a CNN generator and an ANN discriminator. After the generation, the GAN is realized using supervised, unsupervised, or semi-supervised machine learning.EXAMPLE 2Five tumor lesions are taken from the liver (2 samples) and kidney (3 samples) of a patient afflicted with metastatic breast cancer. Cells from the five tumor lesions are sequenced. The cells of the five tumor lesions are subjected to clonal analysis using Concerti, and a total of twelve relevant clones are determined from the five tumor lesions. Each of the twelve relevant clones is found to share certain changes with the other clones, but to be substantially genetically distinct. For all twelve relevant clones, tumor type, tumor location, lesion location, genomic profile and otheromic profiles of the clone are used to identify 30 cell lines in the Sanger, DepMap, CCLE and LINCS databases that are similar to the relevant clones. The information obtained from the cell line databases is obtained as a csv file. The cell line information contains tumor type, disruption data, genomic profile, and otheromic profiles. The csv file of all 30 cell lines is applied as input to a GAN model to generate synthetic cell line data with genomic profiles similar to the identified cell lines to supplement and extend the perturbation dataset. The similarity between the relevant clones and the synthetic and existing cell lines is measured to more accurately determine the match of these cell lines and their disruption data to the relevant clone. Subsequently, a machine learning model such as an ANN is used to predict a malfunction response by training the model with the existing and synthesized data. After the ANN has been trained, the tumor type, tumor location, lesion location, genomic profile and otheromic profiles of the relevant clones are provided to the ANN and the output from the ANN is a ranking of predicted perturbation responses for the relevant clones. This process is repeated for each relevant clone to treat each clone individually.The GAN model is generated using the open source deep learning frameworks PyTor, Tensorflow, and / or KERAS, all of which use the computer programming language Python and allow the generation of the GAN with a training function as well as an ANN and / or CNN generator and discriminator. The GAN may be generated with an ANN generator and an ANN discriminator, a CNN generator and a CNN discriminator, an ANN generator and a CNN discriminator, or a CNN generator and an ANN discriminator. After the generation, the GAN is realized using supervised, unsupervised, or semi-supervised machine learning.References included in the specificationThis list of documents cited by the applicant has been produced in an automated manner and is only included for the better information of the reader. The list is not part of the German patent application or utility model application. The DPMA does not take any adhesion for any faults or omissions.Cited Non-Patent Literaturehttps: / / cancer.sanger.ac.uk / cell_Iines

[0029] https: / / depmap.org / portal / ), the CCLE (Cancer Cell Line Encyclopedia, available at https: / / sites.broadinstitute.org / ccle / ), the Library of Network-based

[0029] https: / / lincsproject.org

[0029]

Claims

A method comprising: sequencing a relevant clone obtained from at least one tumor lesion; identifying at least one cell line having similar characteristics to the relevant clone and composing a dataset comprising perturbation data for the at least one cell line; feeding the dataset to an artificial intelligence, Kl, platform, wherein the dataset trains the KI platform to predict responses to disturbances included in the perturbation data; inputting information relating to the relevant clone to the trained Kl platform and obtaining as output a ranking of the predicted perturbation responses of the relevant clone to the data included in the perturbation data, wherein the predicted perturbation responses are ranked from the strongest perturbation response to the weakest perturbation response of the ranking; and developing a combination therapy for the relevant clone that has disorders from at least two of the high-rank disorder responses.The method of claim 1, wherein the characteristics of the at least one cell line similar to the clone of interest are selected from the group consisting of tumor type, disruption data, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics and combinations thereof.The method of claim 1, wherein the AI platform is selected from the group consisting of machine learning, deep learning, artificial neural networks, convolutional neural networks, generating concurrent networks, and combinations thereof.The method of claim 1, wherein the information relating to the relevant clone is selected from the group consisting of tumor type, tumor site, lesion site, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics and combinations thereof.The method of claim 1, wherein the disorders for combination therapy are selected from the group consisting of environmental stimuli, drug inhibition, genome editing, disease treatment, and combinations thereof.A method comprising: sequencing a relevant clone obtained from at least one tumor lesion; identifying at least two cell lines that have similar characteristics to the relevant clone and composing a first dataset comprising perturbation data for the at least two existing cell lines; feeding the first dataset to an artificial intelligence, Kl, platform, wherein the first dataset trains the Kl platform to generate at least one synthetic cell line comprising perturbation data merged from the at least two existing cell lines, wherein the perturbation data for the at least one synthetic cell line is composed into a second dataset; applying the first and second dataset as training data for the AI platform to learn how responses to disturbances included in the perturbation data of the first and second dataset can be predicted; inputting information relating to the relevant clone into the trained Kl platform and obtaining as output a ranking of the predicted disorder responses of the relevant clone to the data contained in the disorder data of the first and second data sets, wherein the predicted disorder responses from the strongest disorder response to the weakest disorder response are ranked; and developing a combination therapy for the relevant clone having disorders from at least two of the high-ranking disorder responses.The method of claim 6, wherein the characteristics of the at least two existing cell lines similar to the relevant clone are selected from the group consisting of tumor type, disruption data, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics and combinations thereof.The method of claim 6, wherein the AI platform is selected from the group consisting of machine learning, deep learning, artificial neural networks, convolutional neural networks, generating concurrent networks, and combinations thereof.The method of claim 6, wherein the information relating to the relevant clone is selected from the group consisting of tumor type, tumor site, lesion site, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics and combinations thereof.The method of claim 6, wherein the disorders for combination therapy are selected from the group consisting of environmental stimuli, drug inhibition, genome editing, disease treatment, and combinations thereof.A computer program product for ranking tumor clone failure responses, comprising: program instructions on one or more computer readable storage media for training an artificial intelligence, AI, platform to predict responses to failures contained in a dataset comprising failure data for at least one cell line having characteristics similar to a relevant clone; program instructions on one or more computer readable storage media for injecting information relating to the relevant clone into the trained AI platform, wherein the trained AI platform predicts failure responses for the relevant clone to the failures contained in the failure data of the dataset; and program instructions on one or more computer readable storage media for issuing a ranking of the predicted failure responses for the relevant clone from the strongest failure response to the weakest failure response from the AI platform.The computer program product of claim 11, wherein the AI platform is selected from the group consisting of machine learning, deep learning, artificial neural networks, convolutional neural networks, generating concurrent networks, and combinations thereof.The computer program product of claim 11, wherein the AI platform comprises a generating concurrent network.The computer program product of claim 11, wherein the information relating to the relevant clone is selected from the group consisting of tumor type, tumor site, lesion site, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics and combinations thereof.The computer program product of claim 11, wherein the disorders included in the dataset are selected from the group consisting of environmental stimuli, drug inhibition, genome editing, disease treatment, and combinations thereof.A computer program product for ranking tumor clone failure responses, comprising: program instructions on one or more computer readable storage media for training an artificial intelligence, AI, platform to generate at least one synthetic cell line comprising failure data merged from failure data for at least two existing cell lines having similar characteristics to a relevant clone, wherein the failure data for the at least two existing cell lines are merged into a first dataset and the failure data for the at least one synthetic cell line is merged into a second dataset; program instructions on one or more computer readable storage media for training the AI platform to predict responses to the failures included in the failure data for the first and second datasets; Program instructions on one or more computer readable storage media for feeding information relating to the relevant clone to the trained AI platform, wherein the trained AI platform predicts failure responses for the relevant clone to the failures contained in the failure data of the first and second data sets; and program instructions on one or more computer readable storage media for outputting a ranking of the predicted failure responses for the relevant clone from the strongest failure response to the weakest failure response from the AI platform.The computer program product of claim 16, wherein the AI platform is selected from the group consisting of machine learning, deep learning, artificial neural networks, convolutional neural networks, generating concurrent networks, and combinations thereof.The computer program product of claim 16, wherein the AI platform comprises a generating concurrent network.The computer program product of claim 16, wherein the information relating to the relevant clone is selected from the group consisting of tumor type, tumor site, lesion site, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics and combinations thereof.The computer program product of claim 16, wherein the disorders included in the first and second data sets are selected from the group consisting of environmental stimuli, drug inhibition, genome editing, disease treatment, and combinations thereof.A system comprising: a first computer input data set comprising perturbation data relating to at least two existing cell lines having similar characteristics to a sequence from a relevant clone obtained from at least one tumor lesion; a second computer input data set comprising perturbation data relating to a synthetic cell line, wherein the at least one synthetic cell line comprises perturbation data merged from the at least two existing cell lines; and an artificial intelligence, AI, platform accepting as input the first data set, the second data set, and information relating to the relevant clone, and providing as output a ranking of predicted perturbation responses for the relevant clone to perturbations included in the perturbation data of the first and second data sets; wherein the first dataset trains the AI platform to generate the perturbation data for the at least one synthetic cell line, the first and second datasets train the AI platform to predict the perturbation responses of the relevant clone to the perturbation included in the perturbation data of the first and second datasets, and the predicted perturbation responses for the relevant clone are ranked from the strongest perturbation response to the weakest perturbation response.The system of claim 21, wherein the AI platform is selected from the group consisting of machine learning, deep learning, artificial neural networks, convolutional neural networks, generating concurrent networks, and combinations thereof.The system of claim 21, wherein the AI platform comprises a generating concurrent network.The system of claim 21, wherein the information relating to the relevant clone is selected from the group consisting of tumor type, tumor site, lesion site, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics and combinations thereof.The system of claim 21, wherein the disorders contained in the first and second data sets are selected from the group consisting of environmental stimuli, drug inhibition, genome editing, disease treatment, and combinations thereof.