Targeted combination therapy using tumor clone response prediction from cellular data

An AI-based prediction model using perturbation data from cell lines ranks therapeutic interventions to create a combination therapy that addresses the unique resistance mechanisms of tumor clones, effectively targeting multiple clones and reducing disease progression risk.

JP2026510719APending Publication Date: 2026-04-10INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-12
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Tumors evolve in response to treatment, developing unique resistance mechanisms that conventional therapies fail to address, necessitating a specific combination of therapies to target multiple clones within a tumor.

Method used

A prediction model using an AI platform and perturbation data from cell lines to rank potential therapeutic interventions, enabling the creation of a combination therapy tailored to the specific resistance mechanisms of tumor clones.

Benefits of technology

The model effectively predicts the response of tumor clones to different perturbations, allowing for the design of a combination therapy that safely targets all clones, reducing the risk of disease progression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026510719000001_ABST
    Figure 2026510719000001_ABST
Patent Text Reader

Abstract

The AI ​​platform is used to create combination therapies for patients with tumors that have generated clones. Combination therapies involving at least two perturbations can target clones (including subclones) that have evaded therapeutic intervention due to resistance and / or evolution. The AI ​​platform is trained with perturbation data obtained from at least one cell line with characteristics similar to the target clone. The trained AI platform predicts how the target clone will respond to the perturbations and ranks the perturbation responses from highest to lowest. The at least one cell line may be an existing cell line from a well-established database or a synthetic cell line generated by the AI ​​platform. The AI ​​platform may include one or more of the following: machine learning platforms, deep learning platforms, artificial neural networks (ANNs), convolutional neural networks (CNNs), and generative adversarial networks (GANs).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to a prediction model, and more specifically, to a prediction model that uses an artificial intelligence platform and perturbation data of cell lines to predict a combination therapy for tumor clones.

Background Art

[0002] During the course of treatment, tumors typically evolve in response to treatment and therapies, avoiding and / or resisting elimination. Tumor evolution is often confirmed by the detection of clones and subclones that have branched from a single tumor in response to treatment pressure. Resistance changes in the clones (including any subclones) allow them to escape treatment. Such tumor heterogeneity means that different clones arising from a single tumor progenitor cell may have unique resistance mechanisms that do not respond or react to conventional cancer therapies. To avoid the progression of the disease in patients suffering from multiple tumor clones, a specific combination of therapies is required, with each therapy in the combination designed to target one or more of the resistance mechanisms across different tumor clones. In view of the above, there is a need in the art to develop a method for predicting how existing tumor clones may respond to different therapeutic interventions.

Summary of the Invention

[0003] The present invention meets the need in the art by providing a prediction model that utilizes perturbation data from an existing cell line database to predict the response of a patient's tumor clones to a combination therapy comprising at least two perturbations.

[0004] In one embodiment, the present invention relates to a method comprising the steps of: sequencing a target clone obtained from at least one tumor lesion; identifying at least one cell line having similar characteristics to the target clone and compiling a dataset containing perturbation data for the at least one cell line; inputting the dataset into an artificial intelligence (AI) platform, thereby training the AI ​​platform with the dataset to predict responses to perturbations contained in the perturbation data; inputting information related to the target clone into the trained AI platform and obtaining, as output, a ranking of the predicted perturbation responses of the target clone to the perturbations contained in the perturbation data, thereby ranking the predicted perturbation responses from the highest to the lowest perturbation responses; and creating a combination therapy for the target clone, comprising perturbations from at least two of the high-ranking perturbation responses.

[0005] In another embodiment, the present invention comprises the steps of: sequencing a target clone obtained from at least one tumor lesion; identifying at least two existing cell lines having similar characteristics to the target clone and compiling a first dataset having perturbation data for the at least two existing cell lines; inputting the first dataset into an AI platform, where the AI ​​platform is trained with the first dataset to generate at least one synthetic cell line having integrated perturbation data from the at least two existing cell lines, the perturbation data for the at least one synthetic cell line being compiled into a second dataset; and the AI ​​platform The method comprises the steps of: applying the first and second datasets as training data for a form to learn to predict responses to perturbations contained in the perturbation data of the first and second datasets; inputting information relating to the target clone into the trained AI platform and obtaining as output a ranking of the target clone's predicted perturbation responses to the perturbations contained in the perturbation data of the first and second datasets, where the predicted perturbation responses are ranked from the highest to the lowest perturbation response; and creating a combination therapy for the target clone that includes perturbations from at least two high-ranking perturbation responses.

[0006] In a further embodiment, the present invention relates to a computer program product for ranking the perturbation responses of tumor clones, comprising: program instructions on one or more computer-readable storage media for training an AI platform to predict responses to perturbations contained in a dataset comprising perturbation data for at least one cell line having characteristics similar to the clone of interest; program instructions on one or more computer-readable storage media for inputting information relating to the clone of interest into the trained AI platform, wherein the trained AI platform predicts the perturbation response of the clone of interest to the perturbations contained in the perturbation data of the dataset; and program instructions on one or more computer-readable storage media for outputting from the AI ​​platform a ranking of the predicted perturbation responses for the clone of interest from the highest to the lowest perturbation response.

[0007] In another embodiment, the present invention provides one or more program instructions on a computer-readable storage medium for training an AI platform to generate at least one synthetic cell line having integrated perturbation data from perturbation data of at least two existing cell lines having characteristics similar to a target clone, where the perturbation data for the at least two existing cell lines is compiled into a first dataset, and the perturbation data for the at least one synthetic cell line is compiled into a second dataset; and one or more computer-readable storage medium for training the AI ​​platform to predict the response to the perturbations contained in the perturbation data for the first and second datasets. The present invention relates to a computer program product for ranking the perturbation responses of tumor clones, comprising: program instructions on a computer-readable storage medium; program instructions on one or more computer-readable storage media for inputting information relating to the target clone into the trained AI platform, wherein the trained AI platform predicts the perturbation response of the target clone to the perturbations contained in the perturbation data of the first and second datasets; and program instructions on one or more computer-readable storage media for outputting from the AI ​​platform a ranking of the predicted perturbation responses of the target clone from the highest to the lowest perturbation response.

[0008] In a further embodiment, the present invention relates to a system comprising: a first dataset for computer input having perturbation data associated with at least two existing cell lines having properties similar to a sequence from a clone of interest obtained from at least one tumor lesion; a second dataset for computer input having perturbation data associated with a synthetic cell line, wherein the at least one synthetic cell line comprises perturbation data integrated from the at least two existing cell lines; and an AI platform that takes the first dataset, the second dataset, and information related to the clone of interest as input and provides as output a ranking of predicted perturbation responses for the clone of interest to perturbations contained in the perturbation data of the first and second datasets, wherein the AI ​​platform is trained with the first dataset to generate the perturbation data for the at least one synthetic cell line, and the AI ​​platform is trained with the first and second datasets to predict the perturbation response of the clone of interest to the perturbations contained in the perturbation data of the first and second datasets, wherein the predicted perturbation responses for the clone of interest are ranked from the highest perturbation response to the lowest perturbation response.

[0009] In another embodiment, the information relating to the target clone is selected from the group consisting of tumor type, tumor location, lesion location, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics, and combinations thereof.

[0010] In another embodiment, the characteristics of the cell line, which is similar to the clone of the object, are selected from the group consisting of tumor type, perturbation data, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics, and combinations thereof.

[0011] In a further embodiment, the AI ​​platform is selected from the group consisting of machine learning, deep learning, artificial neural networks, convolutional neural networks, generative adversarial networks, and combinations thereof.

[0012] In another embodiment, the perturbation for the combination therapy is selected from the group consisting of environmental stimuli, drug inhibition, gene editing, disease treatment, and combinations thereof.

[0013] Additional embodiments and / or aspects of the invention are provided in the detailed description of the invention below, but are not limited thereto. [Brief explanation of the drawing]

[0014] [Figure 1] This is a schematic diagram illustrating the response perturbation model of tumor clones described herein.

[0015] [Figure 2] This is a schematic diagram of a computer environment that may be used to implement the tumor clone response perturbation model described herein. [Modes for carrying out the invention]

[0016] A description of what is currently considered to be a preferred aspect or embodiment of the present invention or a combination thereof is provided below. Any substitution or modification in function, purpose, or structure is intended to be covered by the appended claims. Where used herein and in the appended claims, the singular form of a word including the articles “a,” “an,” and “the” includes multiple references unless the context makes it clear. Where used herein and in the appended claims, the terms “comprise,” “comprised,” “comprises,” or “comprising,” or any combination thereof, specify the presence of an expressed component, element, feature, or step, or combination thereof, but do not exclude the presence or addition of one or more other components, elements, features, or steps, or combination thereof.

[0017] As used herein, the term "lesion" refers to an area of ​​diseased tissue, and the term "tumor lesion" refers to an area of ​​diseased tissue containing a whole tumor, a part of a tumor, or tumor cells.

[0018] As used herein, the term “clone” refers to a group of cells that are considered homogeneous and have the same genomic profile. In the context of tumor lesions, different clonal populations present within a single tumor lesion are “subclones” if they share the same progenitor clone; therefore, progenitor tumor clones and subclones are identical, as both represent groups of homogeneous tumor cells from which the same therapeutic effect is expected. Based on the above, as used herein, the term “target clone” refers to both progenitor tumor clones and derived subclones.

[0019] As used herein, the term “perturbation” refers to a change in the function of a biological system caused by external or internal means. Examples of perturbations include, but are not limited to, environmental stimuli, drug inhibition, gene editing, and disease treatment.

[0020] Examples of environmental stimuli include, but are not limited to, temperature changes, osmotic shocks, and pressure changes.

[0021] Drug inhibition alters biological pathways through the binding of small molecules to the activity and / or allosteric sites of enzymes. Examples of cancer growth inhibitors include, but are not limited to, tyrosine kinase inhibitors, protein kinase inhibitors (e.g., mTOR and PI3K inhibitors), proteasome inhibitors, histone deacetylase inhibitors, Hedgehog pathway inhibitors, BRAF inhibitors, and MEK inhibitors.

[0022] Examples of gene editing include, but are not limited to, gene knockout (removal or inactivation of a gene), gene knockdown (reduction of gene expression), gene knockup (insertion of a protein encoding a cDNA sequence at a gene site), and CRISPR (clustered, regularly interspaced, short palindromic repeats: clustered, regularly interspaced, short palindromic repeats that modify, remove, and replace DNA in a highly targeted manner). In the context of the present invention, a biological system is a cell line or cells obtained from a tumor lesion, tumor clone, or tumor subclone.

[0023] Examples of disease treatments include, but are not limited to, cell therapies, immunotherapies, and hormone therapies. Examples of cells that may be used in cell therapy include, but are not limited to, autologous cells (derived from the patient), allogeneic cells (derived from a donor), pluripotent stem cells (capable of differentiating into any cell in the body), multipotent stem cells (capable of differentiating into all cell types within a particular lineage), unipotent stem cells (capable of producing only one cell type but self-replicating), adult stem cells (involved in the maintenance and repair of the tissues in which they reside), primary cells (terminally differentiated cells isolated directly from living tissue), secondary cells (immortalized cell lines capable of unlimited division), and combinations thereof. Examples of immunotherapies include, but are not limited to, monoclonal antibodies, nonspecific immunotherapies (e.g., cytokine and bacterial therapies), oncolytic virus therapies, T-cell therapies, and cancer vaccines. Examples of hormone therapies include, but are not limited to, aromatase inhibitors (AIs), estrogen receptor antagonists, selective estrogen receptor modulators (SERMs), luteinizing hormone-release hormone (LHRH), antiandrogens, CYP17 inhibitors, progestins, and anti-adrenergic agents.

[0024] As used herein, the terms “artificial intelligence,” “AI,” and “AI platform” refer to computer algorithms that learn through experience. Examples of AI platforms include, but are not limited to, machine learning, deep learning, and neural networks. Several AI platforms that may be used to implement the present invention are discussed below.

[0025] As used herein, the term "machine learning" refers to an artificial intelligence (AI) function in which an algorithm learns from training data to make predictions or decisions without being explicitly programmed. Machine learning is divided into three categories: supervised learning, semi-supervised learning, and unsupervised learning. In supervised learning, a computer algorithm is trained with labeled data. At the time of application, exemplary inputs and their desired outputs are presented to the computer algorithm, both of which are provided to the computer algorithm, and the goal of the computer algorithm is to learn a general rule for mapping the input to the output. In unsupervised learning, since the computer algorithm is trained with unlabeled data, the computer algorithm is left to its own discretion as to how to find the structure within its input through its learning algorithm. The goal of unsupervised learning is to discover patterns hidden in the data through feature learning and achieve the ultimate goal. In semi-supervised learning, the computer algorithm is trained with a small amount of labeled data and a large amount of unlabeled data.

[0026] As used herein, the term "deep learning" refers to an AI function that mimics the workings of the human brain in data processing and categorization. Deep learning-based AI can learn from unstructured and unlabeled data. During operation, a deep learning-based AI program assumes that any input x and any output y are related by a correlation or causal relationship, and learns to approximate the unknown function (f(x)=y) between them to find the correlation between the input and the output.

[0027] As used herein, the term "artificial neural network" or "ANN (artificial neural network)" refers to a deep learning AI function that models the human brain and has a collection of pseudo-neurons that are all fully connected. Each neuron is a node that is connected to other nodes via links similar to biological axon-synapse-dendrite connections. Each link has a weight that determines the strength of the influence of one node on another node. The neural network learns (i.e., is trained) by processing examples, where each of the neural networks includes known inputs and outputs, and the inputs and outputs form a probabilistically weighted relationship between them. During operation, a group of neural networks with unlabeled input data, based on the similarity between exemplary inputs, automatically extracts features from the group, clusters the groups with similar features, and classifies the output data when there is a labeled dataset for training. The patterns recognized by the neural network are numerical and are contained in vectors, and they must be transformed. Examples of neural network vectors include, but are not limited to, images, sounds, text, times, or combinations thereof.

[0028] As used herein, the term “convolutional neural network” or “CNN (convolutional neural network)” refers to a neural network that uses convolutional layers to convolve an input and pass the result to the next layer. While fully connected feedforward neural networks can be used to learn features and classify data, CNNs regularize a multilayer fully connected network where each neuron in one layer is connected to all neurons on the next layer. The fully connected nature of CNNs means that data can overfit, and it must be regularized. CNNs regularize by leveraging the hierarchical patterns in the data, using smaller, simpler patterns to construct more complex patterns. A CNN consists of an input layer, a hidden hidden layer, and an output layer. The hidden hidden layer performs the convolution. After passing through the convolutional layers, the data is abstracted into a feature map, which is the output of the convolutional kernel applied to the previous layer.

[0029] As used herein, the terms “GAN (generative adversarial network)” and “GAN model” refer to a machine learning function having two neural networks, a generator and an adversarial (also referred to herein as a discriminator), competing with each other in a zero-sum game where the score of one neural network is the score of the other. A GAN is based on indirect training of a generator through a discriminator, where the generator generates candidates that the discriminator evaluates. For evaluation, the discriminator tells the generator how realistic its input seems. The goal of training the generator is to increase the discriminator's error rate by generating novel candidates that the discriminator will determine to be unsynthesized (i.e., part of a real data distribution). A known dataset serves as initial training data for the discriminator, and training involves presenting the discriminator with samples from the training dataset until the discriminator achieves an acceptable accuracy. The generator is trained based on whether it succeeds in misleading the discriminator into believing that the synthesized data is existing data. The GAN generator network is typically seeded with randomly selected inputs, sampled from a defined latent space, such as a multivariate normal distribution. The candidates synthesized by the generator are then evaluated by the discriminator. The machine learning capabilities for GANs can fall into one of three categories: fully supervised, semi-supervised, or unsupervised machine learning. The two neural networks in a GAN can be two ANNs, two CNNs, or a combination of one ANN and one CNN.

[0030] This invention predicts how tumor clones will behave in response to different perturbations by training an AI platform with perturbation data from cell lines. Since tumors typically have multiple clones and patients can have multiple tumors, predicting responses at the clonal level generally requires combination therapies involving at least two perturbations to address the behavior of multiple clones that have evolved from tumor progenitor cells. To the best of the inventors' knowledge, this invention is the first tool to predict the perturbation response of tumor clones.

[0031] Figure 1 is a schematic diagram of the workflow required to implement the predictive model described herein. As a starting point, at least one tumor lesion (K=l1, l2, ...l) is extracted from one patient. kFor the above, sequencing is performed to determine the clonal composition of at least one tumor lesion, including the progenitor tumor clone and subclones derived from the progenitor tumor clone. From the sequencing data, the target clone is identified based on selected characteristics, including but not limited to tumor type (e.g., the type of cancer associated with the tumor, such as bladder cancer, breast cancer, colorectal cancer, kidney cancer, lung cancer), tumor location, lesion location, and omics data. Examples of omics data include but not limited to genomics (gene-related information), transcriptomics (RNA-related information), proteomics (protein-related information), microbiomics (microorganism-related information such as bacteria, fungi, and viruses), metabolomics (metabolites-related information), lipidomics (lipid-related information), and epigenomics (methylated DNA or denatured histone proteins). Next, one or more cell lines with the same characteristics as the target clone are identified from one or more existing databases. It should be understood that one or more cell lines may possess the same characteristics as a single target clone. Examples of databases available to identify at least one cell line include, but are not limited to, the Sanger Database (available at https: / / cancer.sanger.ac.uk / cell_lines), the DepMap portal (available at https: / / depmap.org / portal / ), CCLE (Cancer Cell Line Encyclopedia, available at https: / / sites.broadinstitute.org / ccle / ), the Library of Network-based Cellular Signatures (LINCS) (available at https: / / lincsproject.org / ), and combinations thereof.

[0032] The Sanger Database includes COSMIC (Catalogue of Somatic Mutations in Cancer: a database supervised by experts in somatic mutations), Cellline Projects (mutation profiles of over 1000 cell lines used in cancer research), COSMIC-3D (including interactive diagrams of cancer mutations in 3D structure), Oncogene Survey (a catalog of genes with mutations causally related to cancer), Cancer Mutation Survey (classification of gene mutations that cause cancer progression), and Actionability (mutations that can be addressed with precision oncology).

[0033] The DepMap portal provides cancer-dependent maps including gene and pharmacology dependencies, tumor context, efficacy prediction biomarkers, and over 2000 cancer models.

[0034] CCLE includes cell line annotations for over 1000 human cancer models, integrated mutation calls for 329 cell lines, RNA expression data for 1019 cell lines, fusion calls for 1019 cell lines, epigenetic and histone modification data, proteomics data, and metabolomics data. In addition to the above, it includes any private cell line database (such as those owned by biotechnology or pharmaceutical companies). In one embodiment, at least one cell line similar to the clone of interest can be identified by using any one of the above databases and comparing the genomic profile of the clone of interest with the gene profile of at least one cell line.

[0035] The LINCS database identifies and categorizes molecular signatures that occur when cells are exposed to substances that disrupt their normal functions.

[0036] In some situations, it may be virtually impossible to find a cell line (also referred to herein as an “existing cell line”) that closely matches the target clone in known databases. Such situations may include types of cancer with high tumor clone diversity across all cancer patients, or cancers with unique clones. In such situations, identifying the existing cell line of interest to base predictions on by identifying recurring features and patterns becomes statistically difficult. To identify common mechanisms between clones and perturbations that may affect multiple clones, two or more existing cell lines similar to the target clone are combined into one or more synthetic cell lines. One or more synthetic cell lines, as a whole, enable the generation of predictions about tumor lesions (since tumor lesions are embodiments of one or more target clones). Depending on the type of tumor lesion that generates the target clone, it should be understood that the cell lines used in the predictive models described herein may be existing cell lines and / or synthetic cell lines.

[0037] Once one or more cell lines are selected from an existing database, the information associated with the one or more cell lines is compiled into a tabular format and input into the AI ​​platform as training data. Examples of information associated with one or more cell lines include, but are not limited to, tumor type, perturbation data, and omics profiles. Examples of perturbation data include, but are not limited to, the responses of one or more cell lines to different environmental stimuli, drug inhibition, gene editing, and disease treatment. Examples of omics profiles include, but are not limited to, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, and epigenomics. Examples of tabular data formats that may be used for input include, but are not limited to, comma-separated value (CSV) files and spreadsheets (e.g., Excel® from Microsoft Corporation, Redmond, Washington, USA; Numbers® from Apple Inc., Cupertino, California, USA).

[0038] To generate synthetic cell lines, information related to at least two existing cell lines is prepared in tabular form and input into the AI ​​platform along with instructions for the AI ​​platform to integrate the information from at least two cell lines into one or more synthetic cell lines. In one embodiment, the AI ​​platform may be a GAN that can take these two or more cell lines as input to generate one or more synthetic cell lines having omics (omic) features similar to the initial input.

[0039] The AI ​​platform is trained with tabular cell line perturbation data to learn the responses of one or more cell lines (both existing and synthetic) to perturbations. Training with cell line perturbation data enables the AI ​​platform to predict how a target clone (sharing similar characteristics to the cell line) will respond to different perturbations. Any AI platform can be used to generate the perturbation responses of the target clone, and such platform includes, but is not limited to, machine learning platforms, deep learning platforms, neural networks, CNNs, GANs, and combinations thereof.

[0040] Once predicted perturbation responses are obtained for the target clone, the perturbation responses are ranked from the largest to the smallest. The ranking of perturbation responses is considered comprehensively to create a combination therapy designed to treat the target clone. In one embodiment, for any one target clone, the combination therapy generally includes the highest-ranking perturbation among the two or more perturbations described herein. In another embodiment, the synergistic and / or toxicological effects of the drugs are analyzed as needed to ensure that adverse side effects or drug interactions with the combination therapy are avoided. By providing a combination therapy designed to safely target all clones in a patient, the risk of disease progression to the patient can be significantly reduced.

[0041] For example, in a patient with multiple tumor lesions, each generating its own clone, the presence of different heterologous clones within a single patient may necessitate different combinations to target the different resistance mechanisms of the heterologous clones. For instance, for any two heterologous clones tested in the predictive models described herein, one clone may be predicted to respond to a combination therapy of a top-tier gene editing perturbation, such as knockout gene editing, and a top-tier disease treatment perturbation, such as pluripotent stem cell therapy, while the other clone may be predicted to respond to a combination therapy of a top-tier environmental perturbation, such as osmotic shock, a top-tier drug inhibition perturbation, such as tyrosine kinase drug inhibition, and a top-tier disease treatment perturbation, such as monoclonal antibody therapy.

[0042] In one embodiment, the present invention includes the steps of: sequencing a target clone obtained from at least one tumor lesion; identifying at least one cell line having similar characteristics to the target clone and compiling a dataset containing perturbation data for the at least one cell line; inputting the dataset into an artificial intelligence (AI) platform, thereby training the AI ​​platform with the dataset to predict responses to perturbations contained in the perturbation data; inputting information related to the target clone into the trained AI platform and obtaining, as output, a ranking of the predicted perturbation responses of the target clone to the perturbations contained in the perturbation data, thereby ranking the predicted perturbation responses from the highest to the lowest perturbation responses; and creating a combination therapy for the target clone, comprising perturbations from at least two of the high-ranking perturbation responses.

[0043] In another embodiment, the present invention comprises the steps of: sequencing a target clone obtained from at least one tumor lesion; identifying at least two existing cell lines having similar characteristics to the target clone and compiling a first dataset having perturbation data for the at least two existing cell lines; inputting the first dataset into an AI platform, hereby training the AI ​​platform with the first dataset to generate at least one synthetic cell line having integrated perturbation data from the at least two existing cell lines, the perturbation data for the at least one synthetic cell line being compiled into a second dataset; and the AI ​​platform The steps include: applying the first and second datasets as training data for ratform to learn to predict responses to perturbations contained in the perturbation data of the first and second datasets; inputting information relating to the target clone into the trained AI platform and obtaining as output a ranking of the target clone's predicted perturbation responses to the perturbations contained in the perturbation data of the first and second datasets, where the predicted perturbation responses are ranked from the highest to the lowest perturbation response; and creating a combination therapy for the target clone that includes perturbations from at least two high-ranking perturbation responses.

[0044] In a further embodiment, the present invention includes one or more program instructions on a computer-readable storage medium for training an AI platform to predict responses to perturbations included in a dataset containing perturbation data for at least one cell line having characteristics similar to a clone of interest; one or more program instructions on a computer-readable storage medium for inputting information relating to the clone of interest into the trained AI platform, wherein the trained AI platform predicts the perturbation response for the clone of interest to the perturbations included in the perturbation data of the dataset; and one or more program instructions on a computer-readable storage medium for outputting from the AI ​​platform a ranking of the predicted perturbation responses for the clone of interest from the highest to the lowest perturbation response.

[0045] In another embodiment, the present invention includes one or more program instructions on a computer-readable storage medium for training an AI platform to generate at least one synthetic cell line having integrated perturbation data from perturbation data for at least two existing cell lines having characteristics similar to the clone of interest, wherein the perturbation data for the at least two existing cell lines is compiled into a first dataset, and the perturbation data for the at least one synthetic cell line is compiled into a second dataset; one or more program instructions on a computer-readable storage medium for training the AI ​​platform to predict responses to the perturbations contained in the perturbation data for the first and second datasets; one or more program instructions on a computer-readable storage medium for inputting information relating to the clone of interest into the trained AI platform, wherein the trained AI platform predicts the perturbation response for the clone of interest to the perturbations contained in the perturbation data for the first and second datasets; and one or more program instructions on a computer-readable storage medium for outputting from the AI ​​platform a ranking of the predicted perturbation responses for the clone of interest from the highest to the lowest perturbation response.

[0046] In a further embodiment, the present invention comprises a first dataset for computer input having perturbation data associated with at least two existing cell lines having properties similar to a sequence from a clone of interest obtained from at least one tumor lesion; a second dataset for computer input having perturbation data associated with a synthetic cell line, wherein the at least one synthetic cell line comprises perturbation data integrated from the at least two existing cell lines; and an AI platform that takes the first dataset, the second dataset, and information related to the clone of interest as input and provides as output a ranking of predicted perturbation responses for the clone of interest to perturbations contained in the perturbation data of the first and second datasets, wherein the AI ​​platform is trained with the first dataset to generate the perturbation data for the at least one synthetic cell line, and the AI ​​platform is trained with the first and second datasets to predict the perturbation responses of the clone of interest to perturbations contained in the perturbation data of the first and second datasets, wherein the predicted perturbation responses for the clone of interest are ranked from the highest perturbation response to the lowest perturbation response.

[0047] In another embodiment, information related to the clone of interest is selected from tumor type, tumor location, lesion location, omics profile, and combinations thereof.

[0048] In further embodiments, the characteristics of an existing cell line similar to the clone of interest are selected from tumor type, perturbation data, omics profile, and combinations thereof.

[0049] In another embodiment, the perturbation for combination therapy is selected from the group consisting of environmental stimuli, drug inhibition, gene editing, disease treatment, and combinations thereof.

[0050] In a further embodiment, the AI ​​platform is selected from the group consisting of machine learning, deep learning, ANN, CNN, GAN, and combinations thereof.

[0051] In another embodiment, the AI ​​platform includes GANs alone or in combination with ANNs or CNNs.

[0052] The present invention applies to personalized medicine and patient screening. Regarding personalized medicine, the predictive models described herein are useful in developing treatments specific to individuals with at least one type of tumor. Regarding patient screening, the predictive models can be used to screen patients in clinical trials based on the clonal composition of the patient's tumor lesions.

[0053] Various aspects of this disclosure are described by explanatory text, flowcharts, block diagrams of computer systems, and / or block diagrams of mechanical logic included in embodiments of computer program products (CPPs). With respect to any flowchart, depending on the technology involved, operations may be performed in a different order than those shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive blocks of a flowchart may be performed in reverse order, as a single integrated step, simultaneously, or with at least partial time overlap.

[0054] Embodiments of a computer program product ("CPP Embodiment" or "CPP") are terms used in this disclosure to describe any set of one or more storage media ("mediums") that are collectively comprised of one or more storage devices that collectively contain machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. "Storage device" is any tangible device capable of holding and storing instructions for use by a computer processor. Computer-readable storage media may, but are not limited to, electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, mechanical storage media, or any preferred combination thereof. Some known types of storage devices, including these media, include diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random-access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding devices (punch cards or pits / lands formed on the main surface of a disk), or any suitable combination of the above. When the term "computer-readable storage medium" is used in this disclosure, it shall not be construed as storage in the form of a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses passing through optical fiber cables, electrical signals communicated through wires, and / or other transmission media.As those skilled in the art will understand, data is typically moved at several intermittent points during the normal operation of a storage device, such as during access, defragmentation, or garbage collection; however, data is not transient while it is stored, and therefore the storage device is not transient.

[0055] Refer to Figure 2 below for further discussion. The computing environment 100 includes an example of an environment for executing at least a portion of the computer code associated with performing the method of the invention, such as the computer code required to perform machine learning, deep learning, ANN, CNN, and GAN, and for predicting the perturbed response of subclones 200 as described herein. In addition to block 200, the computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end-user device (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, the computer 101 includes a processor set 110 (including processing circuits 120 and a cache 121), a communication fabric 111, volatile memory 112, persistent storage 113 (including an operating system 122 and the block 200 shown above), a peripheral device set 114 (including a user interface (UI) device set 123, storage 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. The remote server 104 includes the remote database 130. The public cloud 105 includes the gateway 140, the cloud orchestration module 141, the host physical machine set 142, the virtual machine set 143, and the container set 144.

[0056] Computer 101 may take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device that is currently known or may be developed in the future, capable of running programs, accessing networks, or querying databases such as remote database 130. As is well understood in the field of computer technology, and depending on the technology, the execution of a computer implementation may be distributed among multiple computers and / or multiple locations. On the other hand, in this description of the computing environment 100, in order to make the explanation as concise as possible, the detailed discussion will focus on a single computer, specifically computer 101. Computer 101 may be located in the cloud, although it is not shown in the cloud in Figure 2. On the other hand, computer 101 is not required to be located in the cloud, except to any extent that may be definitively shown.

[0057] The processor set 110 includes one or more computer processors of any type currently known or to be developed in the future. The processing circuitry 120 may be distributed across multiple packages, for example, multiple interconnected integrated circuit chips. The processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. The cache 121 is memory located within the processor chip package and is typically used for data or code that should be available for high-speed access by threads or cores running on the processor set 110. The cache memory is typically organized into multiple levels depending on its relative proximity to the processing circuitry. Alternatively, some or all of the cache for the processor set may be located "off-chip". In some computing environments, the processor set 110 may operate using qubits and be designed to perform quantum computing.

[0058] Computer-readable program instructions are typically loaded onto computer 101 and cause the processor set 110 of computer 101 to execute a series of operational steps, thereby realizing the computer implementation method. As a result, the instructions thus executed instantiate the methods specified in the flowcharts and / or descriptions of the computer implementation methods contained herein (collectively referred to as the "Methods of the Invention"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 110 to control and direct the execution of the Methods of the Invention. In computing environment 100, at least some of the instructions for executing the Methods of the Invention may be stored in blocks 200 within persistent storage 113.

[0059] The communication fabric 111 is a signal conduction path that enables various components of the computer 101 to communicate with one another. Typically, this fabric is made up of switches and conductive paths, such as buses, bridges, and physical input / output ports. Other types of signal communication paths, such as optical fiber communication paths and / or wireless communication paths, may be used.

[0060] The volatile memory 112 is any type of volatile memory that is currently known or may be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, but this is not required unless explicitly stated. In computer 101, the volatile memory 112 is located in a single package and resides inside computer 101, but alternatively or additionally, the volatile memory may be distributed across multiple packages and / or located externally to computer 101.

[0061] The persistent storage 113 is any form of non-volatile storage for a computer, currently known or to be developed in the future. The non-volatility of this storage means that the stored data is maintained regardless of whether power is supplied to the computer 101 and / or directly to the persistent storage 113. The persistent storage 113 may be read-only memory (ROM), but typically at least a portion of the persistent storage allows for writing, deleting, and rewriting of data. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. The operating system 122 may take several forms, such as various known proprietary operating systems or open-source portable operating system interface type operating systems using a kernel. The code contained in block 200 typically includes at least some computer code involved in performing the method of the present invention.

[0062] The peripheral device set 114 includes a set of peripheral devices for the computer 101. Data communication connections between the computer 101's peripheral devices and other components may be implemented in various ways, such as Bluetooth® connections, near-field communication (NFC) connections, connections made by cables (such as Universal Serial Bus (USB) type cables), insert-type connections (e.g., Secure Digital (SD) cards), connections made through local area communication networks, and even connections made through wide area networks such as the Internet. In various embodiments, the UI device set 123 may include components such as a display screen, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage 124 is external storage such as an external hard drive, or insertable storage such as an SD card. Storage 124 may be persistent and / or volatile. In some embodiments, storage 124 may take the form of a quantum computing memory device for storing data in the form of qubits. In embodiments where computer 101 requires a large amount of storage (for example, when computer 101 locally stores and manages a large database), this storage may be provided by peripheral storage devices designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 125 consists of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer and another may be a motion detector.

[0063] The network module 115 is a collection of computer software, hardware, and firmware that enables computer 101 to communicate with other computers via the WAN 102. The network module 115 may include hardware such as a modem or Wi-Fi signal transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for transmitting data over the internet. In some embodiments, the network control and network forwarding functions of the network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control and forwarding functions of the network module 115 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for performing the methods of the present invention can typically be downloaded from an external computer or external storage device to computer 101 via a network adapter card or network interface included in the network module 115.

[0064] WAN102 is any wide area network (e.g., the Internet) that can transmit computer data over non-local distances using any currently known or future-developed technology for transmitting computer data. In some embodiments, WAN102 may be replaced and / or complemented by a local area network (LAN), such as a Wi-Fi network, designed to transmit data between devices located in a local area. WANs and / or LANs typically include computer hardware such as copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.

[0065] The end-user device (EUD) 103 is any computer system used and controlled by an end-user (e.g., a customer of the company operating computer 101), and may take any of the forms described above in relation to computer 101. The EUD 103 typically receives useful and valuable data from the operation of computer 101. For example, in a hypothetical case where computer 101 is designed to provide recommendations to an end-user, these recommendations would typically be transmitted from computer 101's network module 115 to the EUD 103 via the WAN 102. Thus, the EUD 103 can display or otherwise present recommendations to the end-user. In some embodiments, the EUD 103 may be a client device such as a thin client, heavy client, mainframe computer, or desktop computer.

[0066] The remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. The remote server 104 may be controlled and used by the same entity that operates computer 101. The remote server 104 represents a machine that collects and stores useful and beneficial data for use by other computers, such as computer 101. For example, in a hypothetical case where computer 101 is designed and programmed to provide recommendations based on historical data, this historical data may be provided to computer 101 from the remote database 130 of the remote server 104.

[0067] The public cloud 105 is any computer system available for use by multiple entities, providing on-demand availability of computer system resources and / or other computing capabilities, particularly data storage (cloud storage) and computing capabilities, without requiring direct and active management by the user. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. Direct and active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments running on various computers that make up the host physical machine set 142, which is a universe of physical computers located within and / or available to the public cloud 105. The virtual computing environment (VCE) typically takes the form of virtual machines from the virtual machine set 143 and / or containers from the container set 144. These VCEs can be stored as images and transferred either as images or after instantiation of the VCEs, among and between hosts on various physical machines. The cloud orchestration module 141 manages the transfer and storage of images, deploys new instantiations of VCEs, and manages the active instantiation of VCE deployments. The gateway 140 is a collection of computer software, hardware, and firmware that enables the public cloud 105 to communicate over the WAN 102.

[0068] Here, some further explanation of virtualized computing environments (VCEs) is provided. A VCE can be stored as an "image." A new active instance of a VCE can be instantiated from an image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows for the existence of multiple isolated user-space instances called containers. These isolated user-space instances typically behave like actual computers in terms of the programs running within them. Computer programs running on a normal operating system can utilize all of that computer's resources, including connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running inside a container can only use the contents of the container and the devices allocated to the container; this feature is known as containerization.

[0069] The private cloud 106 is similar to the public cloud 105, except that its computing resources are available only for use by a single enterprise. While the private cloud 106 is shown as being in communication with the WAN 102, in other embodiments, the private cloud may be completely isolated from the internet and accessible only via a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types), often implemented by different vendors. Each of the multiple clouds remains a separate, discrete entity, but the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the multiple configuration clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.

[0070] The descriptions of various aspects or embodiments or combinations thereof of the present invention are presented for illustrative purposes only and are not intended to be comprehensive or to limit the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best describe the aspects and / or embodiments, the principles of the practical applications or technical improvements of the technology found in the market, or to enable those skilled in the art to understand the aspects and / or embodiments disclosed herein. [experiment]

[0071] The following examples are provided to those skilled in the art to provide a complete disclosure of how to use the aspects and embodiments of the invention described herein. While efforts have been made to ensure accuracy with respect to the variables, experimental errors and deviations should be taken into consideration. [Example 1]

[0072] Four tumor lesions were provided from patients with metastatic colorectal cancer: two liver samples, one brain sample, and one subcutaneous soft tissue sample. Cells from the four tumor lesions were sequenced. Clonal analysis was performed on the cells from the four tumor lesions using the open-source software program Concerti (available at https: / / github.com / ComputationalGenomics / Concerti), and two sibling clones were identified: one clone possessed the KRAS p.G12S allele, and the other possessed the ELF3 p.S229R allele. From the clone possessing the KRAS p.G12S allele, two subclones grew, both of which possessed the BCLAF p.S496L allele, which is specific to liver tissue. Two subclones also grew from the clone possessing the ELF3 p.S229R allele; both of these retain the ELF3 p.S229R allele, but one is specific to brain tissue and the other to subcutaneous soft tissue. All four clones are identified as the target clone.

[0073] For each of the four target clones, ten cell lines showing similarity to the target subclone are identified from within the Sanger, DepMap, CCLE, and LINCS databases, using the clone's tumor type, tumor location, lesion location, genomic profile, and any other omics profile. The information obtained from the cell line databases is obtained as a CSV file. The cell line information includes tumor type, perturbation data, genomic profile, and any other omics profile. All CSV files for the ten cell lines are applied as input to a GAN model to generate synthetic cell line data with genomic profiles similar to the identified cell lines, supplementing and extending the perturbation dataset. The similarity between the target clone and the synthetic and existing cell lines is measured for a more accurate assessment of the fit of these cell lines and their perturbation data to the target clone. Machine learning models such as ANNs are then used to predict perturbation responses by training on existing and synthetic data. Once the ANN is trained, the tumor type, tumor location, lesion location, genomic profile, and any other omics profile of the target clone are provided to the ANN, and the output from the ANN is a ranking of the predicted perturbation response for the target clone. This process is repeated for each target clone, so that they are processed individually.

[0074] GAN models are built using open-source deep learning frameworks such as PyTorch, Tensorflow, and / or KERAS, all of which use the Python computer programming language and enable the construction of ANN and / or CNN generators and discriminators, as well as GANs with training capabilities. GANs can be built with ANN generators and ANN discriminators, CNN generators and CNN discriminators, ANN generators and CNN discriminators, or CNN generators and ANN discriminators. Once built, GANs can be implemented using supervised, unsupervised, or semi-supervised machine learning. [Example 2]

[0075] Five tumor lesions were provided from the liver (2 samples) and kidney (3 samples) of patients with metastatic breast cancer. Cells from the five tumor lesions were sequenced. Clonal analysis was performed on the cells from the five tumor lesions using Concerti, and a total of 12 target clones were found from the five tumor lesions. Each of the 12 target clones was found to exhibit some common changes, but was initially genetically distinct. For all 12 target clones, 30 cell lines showing similarity to the target clones are identified from Sanger, DepMap, CCLE, and LINCS databases, using the clone's tumor type, tumor location, lesion location, genomics, and other omics profiles. The information obtained from the cell line databases is obtained as CSV files. The cell line information includes tumor type, perturbation data, genomic profile, and any other omics profile. All CSV files for the 30 cell lines are applied as input to a GAN model to generate synthetic cell line data with genomic profiles similar to the identified cell lines, supplementing and extending the perturbation dataset. The similarity between the target clones and synthetic and existing cell lines is measured for a more accurate assessment of the fit of these cell lines and their perturbation data to the target clones. Next, machine learning models such as ANNs are used to predict perturbation responses by training on existing and synthetic data. Once the ANN is trained, the tumor type, tumor location, lesion location, genomic profile, and any other omics profile of the target clone are provided to the ANN, and the output from the ANN is a ranking of the predicted perturbation response for the target clone. This process is repeated for each target clone, so that they are processed individually.

[0076] GAN models are built using open-source deep learning frameworks such as PyTorch, Tensorflow, and / or KERAS, all of which use the Python computer programming language and enable the construction of ANN and / or CNN generators and discriminators, as well as GANs with training capabilities. GANs can be built with ANN generators and ANN discriminators, CNN generators and CNN discriminators, ANN generators and CNN discriminators, or CNN generators and ANN discriminators. Once built, GANs can be implemented using supervised, unsupervised, or semi-supervised machine learning.

Claims

1. The step of sequencing the desired clone obtained from at least one tumor lesion; Steps include identifying at least one cell line having characteristics similar to the target clone, and compiling a dataset containing perturbation data for the at least one cell line; The step involves inputting the dataset into an artificial intelligence (AI) platform, where the AI ​​platform is trained with the dataset to predict its response to perturbations contained in the perturbation data; The steps include inputting information related to the target clone into the trained AI platform and obtaining, as output, a ranking of the predicted perturbation responses of the target clone to the perturbations contained in the perturbation data, where the predicted perturbation responses are ranked from the highest to the lowest perturbation response; and Steps to create a combination therapy for the target clone, including perturbations from at least two high-rank perturbation responses. A method that includes this.

2. The method according to claim 1, wherein the characteristics of the at least one cell line similar to the clone of the objective are selected from the group consisting of tumor type, perturbation data, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics, and combinations thereof.

3. The method according to claim 1, wherein the AI ​​platform is selected from the group consisting of machine learning, deep learning, artificial neural networks, convolutional neural networks, generative adversarial networks, and combinations thereof.

4. The method according to claim 1, wherein the information relating to the clone for the objective is selected from the group consisting of tumor type, tumor location, lesion location, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics, and combinations thereof.

5. The method according to claim 1, wherein the perturbation for the combination therapy is selected from the group consisting of environmental stimuli, drug inhibition, gene editing, disease treatment, and combinations thereof.

6. The step of sequencing the desired clone obtained from at least one tumor lesion; Steps include: identifying at least two existing cell lines that have characteristics similar to the target clone, and compiling a first dataset containing perturbation data for the at least two existing cell lines; The first dataset is input into an artificial intelligence (AI) platform, where the AI ​​platform is trained with the first dataset to generate at least one synthetic cell line containing integrated perturbation data from at least two existing cell lines, and the perturbation data for the at least one synthetic cell line is compiled into a second dataset; A step of learning to predict the response to perturbations contained in the perturbation data of the first and second datasets by applying the first and second datasets as training data for the AI ​​platform; The steps include inputting information related to the target clone into the trained AI platform and obtaining, as output, a ranking of the predicted perturbation responses of the target clone to the perturbations contained in the perturbation data of the first and second datasets, where the predicted perturbation responses are ranked from the highest to the lowest perturbation response; and Steps to create a combination therapy for the target clone, including perturbations from at least two high-rank perturbation responses. A method that includes this.

7. The method according to claim 6, wherein the characteristics of the at least two existing cell lines similar to the clone of the objective are selected from the group consisting of tumor type, perturbation data, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics, and combinations thereof.

8. The method according to claim 6, wherein the AI ​​platform is selected from the group consisting of machine learning, deep learning, artificial neural networks, convolutional neural networks, generative adversarial networks, and combinations thereof.

9. The method according to claim 6, wherein the information relating to the clone for the objective is selected from the group consisting of tumor type, tumor location, lesion location, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics, and combinations thereof.

10. The method according to claim 6, wherein the perturbation for the combination therapy is selected from the group consisting of environmental stimuli, drug inhibition, gene editing, disease treatment, and combinations thereof.

11. Program instructions on one or more computer-readable storage media for training an artificial intelligence (AI) platform to predict its response to perturbations contained in a dataset that includes perturbation data for at least one cell line having characteristics similar to a target clone; Program instructions on one or more computer-readable storage media for inputting information related to the target clone into the trained AI platform, wherein the trained AI platform predicts the perturbation response of the target clone to the perturbations contained in the perturbation data of the dataset; and Program instructions on one or more computer-readable storage media for outputting a ranking of the predicted perturbation responses for the target clone from the AI ​​platform, from the highest perturbation response to the lowest perturbation response. A computer program product for ranking the perturbation response of tumor clones, including...

12. The computer program product according to claim 11, wherein the AI ​​platform is selected from the group consisting of machine learning, deep learning, artificial neural networks, convolutional neural networks, generative adversarial networks, and combinations thereof.

13. The computer program product according to claim 11, wherein the AI ​​platform includes a generative adversarial network.

14. The computer program product according to claim 11, wherein the information relating to the clone for the objective is selected from the group consisting of tumor type, tumor location, lesion location, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics, and combinations thereof.

15. The computer program product according to claim 11, wherein the perturbations included in the dataset are selected from the group consisting of environmental stimuli, drug inhibition, gene editing, disease treatment, and combinations thereof.

16. Program instructions on one or more computer-readable storage media for training an artificial intelligence (AI) platform to generate at least one synthetic cell line containing integrated perturbation data from perturbation data of at least two existing cell lines having characteristics similar to a target clone, wherein the perturbation data for the at least two existing cell lines is compiled into a first dataset, and the perturbation data for the at least one synthetic cell line is compiled into a second dataset; Program instructions on a computer-readable storage medium for training the AI ​​platform to predict the response to the perturbations contained in the perturbation data for the first and second datasets; Program instructions on one or more computer-readable storage media for inputting information related to the target clone into the trained AI platform, wherein the trained AI platform predicts the perturbation response of the target clone to the perturbations contained in the perturbation data of the first and second datasets; and Program instructions on one or more computer-readable storage media for outputting a ranking of predicted perturbation responses for the target clone from the AI ​​platform, from the highest perturbation response to the lowest perturbation response. A computer program product for ranking the perturbation response of tumor clones, including...

17. The computer program product according to claim 16, wherein the AI ​​platform is selected from the group consisting of machine learning, deep learning, artificial neural networks, convolutional neural networks, generative adversarial networks, and combinations thereof.

18. The computer program product according to claim 16, wherein the AI ​​platform includes a generative adversarial network.

19. The computer program product according to claim 16, wherein the information relating to the clone for the objective is selected from the group consisting of tumor type, tumor location, lesion location, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics, and combinations thereof.

20. The computer program product according to claim 16, wherein the perturbations included in the first and second datasets are selected from the group consisting of environmental stimuli, drug inhibition, gene editing, disease treatment, and combinations thereof.

21. A first dataset for computer input having perturbation data associated with at least two existing cell lines that have properties similar to the sequence from the target clone obtained from at least one tumor lesion; A second dataset for computer input having perturbation data associated with synthetic cell lines, wherein the at least one synthetic cell line includes perturbation data integrated from the at least two existing cell lines; and An artificial intelligence (AI) platform that takes the first dataset, the second dataset, and information related to the target clone as input, and provides as output a ranking of the predicted perturbation response for the target clone to the perturbations contained in the perturbation data of the first and second datasets. Equipped with, The AI ​​platform is trained with the first dataset to generate the perturbation data for the at least one synthetic cell line. The AI ​​platform is trained on the first and second datasets to predict the perturbation response of the target clone to the perturbations contained in the perturbation data of the first and second datasets. The predicted perturbation responses for the aforementioned target clones are ranked from the highest perturbation response to the lowest perturbation response. system.

22. The system according to claim 21, wherein the AI ​​platform is selected from the group consisting of machine learning, deep learning, artificial neural networks, convolutional neural networks, generative adversarial networks, and combinations thereof.

23. The AI ​​platform includes a generative adversarial network, as described in claim 21.

24. The system according to claim 21, wherein the information relating to the clone of the objective is selected from the group consisting of tumor type, tumor location, lesion location, genomics, transcriptomics, proteomics, microbiomics, metabolomics, lipidomics, epigenomics, and combinations thereof.

25. The system according to claim 21, wherein the perturbations included in the first and second datasets are selected from the group consisting of environmental stimuli, drug inhibition, gene editing, disease treatment, and combinations thereof.