Detection and diagnosis of cancer evolution
A predictive algorithm using tumor gene profiles and patient data addresses treatment resistance in cancer by providing personalized treatment recommendations, enhancing treatment efficacy and survival rates.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- GUARDANT HEALTH INC
- Filing Date
- 2026-01-13
- Publication Date
- 2026-04-14
AI Technical Summary
Current cancer treatments face significant challenges due to the development of resistance, which is often caused by genetic and epigenetic mechanisms, leading to limited effectiveness and high mortality rates, especially in metastasized cancers.
A computer-implemented method for predicting patient response and resistance to cancer treatment by analyzing tumor gene profiles and patient information, generating predictive algorithms to determine the probability of treatment outcomes based on genetic and treatment data, and providing personalized treatment recommendations.
Enables accurate prediction of treatment responses and resistance, allowing for optimized treatment strategies that improve patient survival rates and outcomes.
Smart Images

Figure 2026065085000001 
Figure 2026065085000002 
Figure 2026065085000003
Abstract
Description
Technical Field
[0001] Cross-reference This application claims priority based on U.S. Provisional Patent Application No. 62 / 290,375, filed on February 2, 2016, which is hereby incorporated by reference in its entirety.
Background Art
[0002] Background Cancer is a disease that poses a significant burden worldwide. Every year, tens of millions of individuals worldwide are diagnosed with cancer, and more than half of such individuals are unable to effectively undergo cancer treatment and may ultimately die. In many countries, cancer ranks as the second most common cause of death after cardiovascular disease.
[0003] Drugs targeting genetic vulnerabilities in human tumors are currently being clinically validated as effective cancer treatment methods. However, their usefulness can be significantly limited by the acquisition of resistance to such treatments, remaining a substantial challenge for the clinical management of advanced cancer. Resistance to treatment with anticancer drugs can occur due to various factors, including individual variability in the subject and the emergence and proliferation of gene variants within the tumor. The most common cause of the acquisition of resistance to a wide range of anticancer drugs is the expression of one or more energy-dependent transporters that detect the anticancer drug and excrete it from the cell. Other resistance mechanisms include insensitivity to drug-induced apoptosis and induction of drug detoxification mechanisms.
[0004] The development of resistance to chemotherapeutic drugs is frequent among cancer patients with solid tumors that have metastasized or spread throughout the body, such as those of the breast, prostate, lung, and colon, and often results in death. In some cases, specific mutation mechanisms contribute directly to the acquisition of drug resistance, and in other cases, epigenetic mechanisms are thought to play a significant role as a non-mutational possibility.
[0005] The optimal criteria for mechanically characterizing tumor drug resistance include detailed studies of tumor tissue obtained before treatment and after recurrence, along with experimental confirmation of candidate resistance effectors. [Overview of the project] [Means for solving the problem]
[0006] Abstract As recognized herein, there is a considerable need for alternative tools to predict patient response to and the emergence of resistance to cancer treatment.
[0007] This disclosure provides methods and systems for detecting or monitoring cancer evolution. Such methods and systems can be used to predict the development of patient response and resistance to cancer treatment, as well as for other benefits.
[0008] In one embodiment, the Disclosure provides a computer-implemented method comprising: (a) a step of obtaining information about a plurality of subjects having cancer at a first time point, and determining a first state for each of the plurality of subjects based on the information at the first time point, wherein the information includes, for each of the plurality of subjects, at least a tumor gene profile obtained by performing nucleic acid genotyping from cell-free fluid and any treatment provided to the subject prior to the first time point; (b) a step of obtaining information about the plurality of subjects at one or more second time points after the first time point, and determining a second state for each of the plurality of subjects at one or more second time points based on the information at one of the one or more second time points, in order to generate a set of subsequent states; and (c) a step of using the set of first states from (a) and the set of subsequent states from (b) to generate a predictive algorithm configured to determine the probability that a given first state results in a second state among a set of states at later time points after the given first state. In some embodiments, the method further includes (d) determining the probability that a given first state in a set of states at an early time point leads to a second state in a set of states at a later time point; and (e) generating an electronic output representing the probability determined in (d).
[0009] In one embodiment, the Disclosure provides a computer-implementable method comprising: (a) a step of obtaining information about a plurality of subjects having cancer at a first time point, and determining a first state for each of the plurality of subjects based on the information at the first time point, wherein the information includes, for each of the plurality of subjects, a tumor gene profile obtained by genotyping at least 50 genes and any treatment provided to the subjects prior to the first time point; (b) a step of obtaining information about the plurality of subjects at one or more second time points after the first time point, and determining a second state for each of the plurality of subjects at one or more second time points based on the information at one of the one or more second time points, in order to generate a set of subsequent states; and (c) a step of using the set of first states from (a) and the set of subsequent states from (b) to generate a predictive algorithm configured to determine the probability that a given first state results in a second state among the set of states at later time points after the given first state. In some embodiments, the method further includes (d) determining the probability that a given first state in a set of states at an early time point leads to a second state in a set of states at a later time point; and (e) generating an electronic output representing the probability determined in (d).
[0010] In some embodiments, the step of acquiring information includes sequencing cell-free deoxyribonucleic acid (cfDNA) from multiple subjects and, optionally, conducting a medical history for each of the multiple subjects. In some embodiments, the treatment was provided to the subject before the first time point. In some embodiments, the method includes generating one or more decision trees, each decision tree comprising a root node, one or more decision branches, one or more decision nodes, and one or more terminal nodes, where the state at the root node represents the first time point, one or more decision branches represent alternative treatments, and one or more decision nodes and one or more terminal nodes represent subsequent states. In some embodiments, one or more decision branches include multiple decision branches. In some embodiments, subsequent states include the survival status of the subject, indicating whether the subject is alive or dead. In some embodiments, subsequent states include subject survival rates. In some embodiments, each of the first states includes a common set of one or more somatic mutations. In some embodiments, the information further includes a profile of the subject.
[0011] In some embodiments, the probability is at least partially a function of the selection of a treatment from among multiple treatment options. In some embodiments, one or more second time points include multiple subsequent time points. In some embodiments, the method further includes the step of determining the probability at multiple subsequent time points. In some embodiments, the time points include at least three or at least four time points. In some embodiments, the first time point is before the subject receives the treatment, and the subsequent time point is after the subject receives the treatment. In some embodiments, the second treatment is administered after the subsequent time point, based on the subsequent state at the subsequent time point.
[0012] In some embodiments, information on multiple subjects includes one or more features from the patient profile of the subjects, selected from a group consisting of age, sex, gender, genetic profile, enzyme levels, organ function, quality of life, frequency of medical interventions, remission status, and patient outcomes. In some embodiments, the genetic profile includes the subject's genotype at one or more loci that increase cancer risk, affect pharmacokinetics, or affect drug sensitivity. In some embodiments, information on multiple subjects includes one or more features from the tumor profile of the subjects, selected from a group consisting of one or more gene variants, tissue of origin, tumor volume, tumor drug sensitivity, and tumor stage. In some embodiments, one or more features are determined by assaying cell-free nucleic acid molecules derived from the subjects. In some embodiments, one or more gene variants are quantified to determine the proportion of cell-free nucleic acid molecules containing one or more somatic mutations. In some embodiments, the method further includes the step of determining whether the proportion of one or more somatic mutations has increased or decreased between a first time point and one or more subsequent time points. In some embodiments, the method further includes the step of determining whether the proportion of one or more somatic mutations is increasing or decreasing between one or more subsequent time points. In some embodiments, the proportion of one or more somatic mutations is increasing. In some embodiments, one or more somatic mutations are increasing, and furthermore, the somatic mutations are associated with resistance to treatment. In some embodiments, the assay includes high-throughput sequencing.
[0013] In another embodiment, the Disclosure provides a method comprising: (a) obtaining information about a subject having cancer at a first time point, the information comprising at least one feature of the subject from a patient profile, tumor profile, or treatment; (b) determining the initial state of the subject based on the information at the first time point; (c) determining the probability of each of a plurality of subsequent states at each of one or more subsequent time points based on the initial state of the subject, thereby providing a set of probabilities for the outcome of the state; (d) generating a recommendation for cancer treatment that optimizes the probability that the subject will achieve a particular outcome, at least in part on the set of probabilities for the outcome of the state; and (e) generating an electronic output indicating the recommendation generated in (d). In some embodiments, the probabilities are at least in part a function of treatment selection from a plurality of treatment options. In some embodiments, one or more subsequent time points include a plurality of subsequent time points. In some embodiments, the Method further includes the step of determining probabilities at a plurality of subsequent time points. In some embodiments, the time points include at least three time points. In some embodiments, the time points include at least four time points. In some embodiments, the first time point is before the subject receives treatment, and the subsequent time point is after the subject receives treatment. In some embodiments, the second treatment is administered after the subsequent time point, based on the subsequent condition at the subsequent time point. In some embodiments, at least one feature of the subject is from the patient profile, selected from a group consisting of age, gender, genetic profile, enzyme levels, organ function, quality of life, frequency of medical interventions, remission status, and patient outcomes.
[0014] In some embodiments, the gene profile includes the genotype of the subject at one or more loci that are hereditary oncogenes. In some embodiments, the gene profile includes the genotype of the subject at one or more loci that affect pharmacokinetics. In some embodiments, the gene profile includes the genotype of the subject at one or more loci that affect drug sensitivity. In some embodiments, at least one feature of the subject is from the tumor profile and is selected from the group consisting of one or more somatic mutations, tissue of origin, tumor volume, tumor drug sensitivity, and tumor stage. In some embodiments, at least one feature is determined by assaying cell-free nucleic acid molecules derived from the subject.
[0015] In some embodiments, somatic mutations are quantified to determine the proportion of cell-free nucleic acid molecules derived from tumors containing one or more somatic mutations.
[0016] In some embodiments, the method further includes the step of determining whether the proportion of one or more somatic mutations has increased or decreased between a first time point and one or more subsequent time points. In some embodiments, the method further includes the step of determining whether the proportion of one or more somatic mutations has increased or decreased between one or more subsequent time points. In some embodiments, the assay involves high-throughput sequencing. In some embodiments, the tumor profile does not originate from a tumor tissue biopsy.
[0017] In one embodiment, the Disclosure provides a method comprising: (a) obtaining information relating to a subject, including at least a tumor gene profile and any treatments previously or currently offered to the subject, and determining the initial state of the subject based on the information; (b) providing a decision tree, where the root node represents the initial subject state, the decision branches represent alternative treatments available to the subject, the opportunity nodes represent points of uncertainty, and the decision nodes or terminal nodes represent subsequent states; (c) providing a treatment process for the subject that maximizes the probability that the subject achieves a state of survival at the terminal node; and (d) generating an electronic output indicating the treatment process determined in (c).
[0018] In one embodiment, the Disclosure provides a method comprising: (a) establishing one or more communication links with one or more healthcare providers on a communication network; (b) receiving medical information from one or more healthcare providers on the communication network regarding one or more subjects; (c) receiving one or more samples from healthcare providers, each containing cell-free deoxyribonucleic acid (cfDNA) derived from each of the one or more subjects; (d) sequencing the cfDNA to identify one or more gene variants present in the cfDNA; (e) creating or supplying a database having information about each of the one or more subjects, the information including both the identified gene variants and the received medical information; and (f) generating at least one predictive model that uses the database and a computer-implemented algorithm to predict the probability of a subsequent state for each of a plurality of different therapeutic interventions, based on the initial state of the subjects.
[0019] In one embodiment, the Disclosure provides a non-transient, computer-readable medium, including machine-executable code, which, when executed by one or more computer processors, implements a method comprising: (a) at a first time point, obtaining information about a plurality of subjects having cancer, and determining a first state for each of the plurality of subjects in order to generate a first set of states based on the information at the first time point, wherein the information includes, for each of the plurality of subjects, at least a tumor gene profile obtained by performing nucleic acid genotyping from cell-free fluid and any treatment provided to the subjects prior to the first time point; (b) at one or more second time points after the first time point, obtaining information about the plurality of subjects, and determining a second state for each of the plurality of subjects at one or more second time points in order to generate a set of subsequent states based on the information at one of the one or more second time points; and (c) using the set of first states from (a) and the set of subsequent states from (b), generating a predictive algorithm configured to determine the probability that a given first state results in a second state among a set of states at later time points after the given first state.
[0020] In one embodiment, the Disclosure provides a non-transient, computer-readable medium, including machine-executable code, which, when executed by one or more computer processors, implements a method comprising: (a) at a first time point, obtaining information about a plurality of subjects having cancer and determining a first state for each of the plurality of subjects based on the information at the first time point, wherein the information includes, for each of the plurality of subjects, a tumor gene profile obtained by genotyping at least 50 genes and any treatment provided to the subjects prior to the first time point; (b) at one or more second time points after the first time point, obtaining information about the plurality of subjects and determining a second state for each of the plurality of subjects at one or more second time points, based on information at one of the one or more second time points, in order to generate a set of subsequent states; and (c) using the set of first states from (a) and the set of subsequent states from (b), generating a predictive algorithm configured to determine the probability that a given first state results in a second state among a set of states at later time points after the given first state.
[0021] In one embodiment, the Disclosure provides a method comprising: (a) obtaining information relating to a subject, including at least a tumor gene profile and any treatments previously or currently offered to the subject, and determining the initial state of the subject based on the information; (b) providing a decision tree, where the root node represents the initial state of the subject, the decision branches represent alternative treatments available to the subject, the opportunity nodes represent points of uncertainty, and the decision nodes or terminal nodes represent subsequent states; (c) providing a treatment process for the subject that maximizes the probability that the subject achieves a state of survival at the terminal node; and (d) administering the treatment process to the subject. In some embodiments, the method further includes (e) obtaining information about a subject at a second time point after the initial state, including at least a tumor gene profile and any treatments previously or currently provided to the subject, and determining a second state of the subject from among a plurality of subsequent states based on the information; (f) providing a subsequent treatment process for the subject that maximizes the probability that the subject achieves a state of survival at the terminal node based on the second state; and (g) administering the subsequent treatment process to the subject.
[0022] In one aspect, the present disclosure provides a method comprising the step of providing a treatment course from among a plurality of alternative treatments for a subject having cancer, the subject being characterized by a decision tree comprising a plurality of decision branches, each decision branch representing one alternative treatment from among the plurality of alternative treatments, and the treatment course maximizing the probability that the subject reaches a state of surviving at the terminal node.
[0023] Further aspects and advantages of the present disclosure will be readily apparent to those skilled in the art from the following detailed description, which shows and describes only exemplary embodiments of the present disclosure. As will be understood, the present disclosure is capable of other and different embodiments, and some of the details thereof are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not restrictive. Incorporation by reference
[0024] All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference as if each individual publication, patent, or patent application were specifically and individually indicated to be incorporated by reference.
[0025] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the invention will be obtained from the following detailed description that illustrates exemplary embodiments of the invention and in which the principles of the invention are utilized, as well as from the appended drawings. Brief Description of the Drawings
[0026] [Figure 1] FIG. 1 shows an exemplary method for analyzing mutations in various disease states of a subject.
[0027] [Figure 2A] FIG. 2A shows various exemplary abnormalities in a cancer genome.
[0028] [Figure 2B] Figure 2B shows an exemplary system for detecting evolutionary avoidance routes.
[0029] [Figure 2C] Figure 2C shows an exemplary model generated by the system in Figure 2B.
[0030] [Figure 2D] Figure 2D shows an exemplary heterogeneous population of normal cells and cancer subclones that develop during the evolutionary history of the tumor.
[0031] [Figure 3] Figure 3 shows an exemplary process for reducing error rates and biases in reading deoxyribonucleic acid (DNA) sequences.
[0032] [Figure 4] Figure 4 shows a schematic diagram illustrating how cancer patients can access reports via the internet.
[0033] [Figure 5] Figure 5 shows several genes associated with the gene variant.
[0034] [Figure 6] Figure 6 shows a decision tree that includes a root node (rectangle) representing the initial state, decision branches (arrows) representing different treatment interventions, and opportunity nodes (circles) where opportunity branches (arrows) have occurred at either a terminal node (triangle) or a decision node (square) representing a subsequent state.
[0035] [Figure 7] Figure 7 shows a computer system programmed to implement the methods provided herein, or otherwise configured to do so. [Modes for carrying out the invention]
[0036] Detailed explanation Genetic variants are alternative forms at a gene locus. In the human genome, approximately 0.1% of nucleotide positions are polymorphic, meaning they exist as a second genetic form occurring in at least 1% of the population. Mutations can introduce genetic variants into the germline and into diseased cells such as cancer cells. Reference sequences such as hg19 or NCBI's Build 37 or Build 38 are intended to represent a “wild-type” or “normal” genome. However, they do not distinguish common polymorphisms that can also be considered normal in that they have a single sequence.
[0037] Genetic variants include sequence variants, copy number variants, and nucleotide modification variants. Sequence variants are changes in the nucleotide sequence of a gene. Copy number variants are deviations in the copy number of a portion of the genome from the wild type. Examples of genetic variants include single nucleotide polymorphisms (SNPs), insertions, deletions, reversals, conversions, transpositions, gene fusions, chromosome fusions, gene shortenings, copy number variations (e.g., aneuploidy, partial aneuploidy, polyploidy, gene amplification), abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid methylation.
[0038] The term "polynucleotide," as used herein, generally refers to a molecule comprising one or more nucleic acid subunits. A polynucleotide may include one or more subunits selected from adenosine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), or their variants. A nucleotide may contain A, C, G, T, or U, or their variants. A nucleotide may contain any subunit that can be incorporated into a growing nucleic acid chain. Such a subunit may be A, C, G, T, or U, or specific to one or more complementary A, C, G, T, or U, or complementary to purines (i.e., A or G, or their variants) or complementary to pyrimidines (i.e., C, T, or U, or their variants), or any other subunit. Subunits can allow individual nucleic acid bases or groups of bases (e.g., AA, TA, AT, GC, CG, CT, TC, GT, TG, AC, CA, or their uracil counterparts) to be degraded. In some embodiments, polynucleotides are deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), or derivatives thereof. Polynucleotides may be single-stranded or double-stranded.
[0039] When used herein, the term “subject” generally refers to animals, e.g., mammals (e.g., humans) or birds (e.g., birds), or other living organisms, e.g., plants. More specifically, subjects may be vertebrates, mammals, mice, primates, monkeys, or humans. Animals include, but are not limited to, domesticated animals, athletic animals, and pets. Subjects may be healthy individuals, individuals with or suspected of having a disease or being prone to a disease, or individuals who require or are suspected of requiring treatment. Subjects may be patients.
[0040] The term "genome" generally refers to the entirety of an organism's genetic information. A genome can be encoded in either DNA or RNA. It can include coding regions that code for proteins, as well as non-coding regions. A genome can encompass the sequences of all chromosomes in a single organism. For example, the human genome has a total of 46 chromosomes. All of these sequences together constitute the human genome. A "reference genome" typically refers to a haploid genome. Examples of reference genomes include hg19, or NCBI's Build 37 or Build 38.
[0041] The terms “adaptor,” “adapter,” and “tag” are used synonymously throughout this specification. An adapter or tag can be ligated to a polynucleotide sequence to be “tagged” by any approach, including ligation, hybridization, or other approaches.
[0042] The terms “library adaptor” or “library adapter,” as used herein, generally refer to a molecule (e.g., a polynucleotide) that can be distinguished in a biological sample (also referred to herein as “sample”) by identity (e.g., sequence).
[0043] The term “sequencing adapter,” as used herein, generally refers to a molecule (e.g., a polynucleotide) that is adapted to enable sequencing of a target polynucleotide by a sequencing instrument, such as by interacting with the target polynucleotide in a manner that enables sequencing. The sequencing adapter enables sequencing of the target polynucleotide by a sequencing instrument. In one embodiment, the sequencing adapter includes a nucleotide sequence that hybridizes or binds to a captured polynucleotide attached to a solid support of the sequencing system, e.g., a flow cell. In another embodiment, the sequencing adapter includes a nucleotide sequence that hybridizes or binds to the polynucleotide to form a hairpin loop, which enables sequencing of the target polynucleotide by the sequencing system. The sequencing adapter may include a sequencer motif, which may be a nucleotide sequence that is complementary to the flow cell sequence of another molecule (e.g., a polynucleotide) and is available to the sequencing system for sequencing the target polynucleotide. The sequencer motif may also contain primer sequences for use in sequencing, such as synthetic sequencing (SBS). The sequencer motif may contain sequences required to link the library adapter to a sequencing system and perform sequencing of the target polynucleotide.
[0044] As used herein, the terms “at least,” “at most,” and “about,” when preceding a series, refer to each member of that series unless otherwise indicated.
[0045] In relation to a referenced number, the term “about” and its grammatical equivalents may include values within a range of up to ±10% of that value. For example, the quantity “about 10” may include quantities from 9 to 11. In other embodiments, the term “about” in relation to a referenced number may include values within a range of ±10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of that value.
[0046] Generally, methods for generating predictive models of tumor evolution over time in response to various treatments, and for using these models to select treatments for subjects (e.g., patients), are disclosed herein. The predictive models are based at least on the tumor's genetic profile, and optionally on the patient's profile and / or treatment. The results may be disclosed to patients or healthcare providers to improve care.
[0047] In some cases, the information includes genetic profiles from tumors obtained by genotyping of cell-free fluid (e.g., cfDNA). In some cases, the information further includes treatments and / or therapeutic interventions provided to the subject. In some cases, the information further includes the subject's profile.
[0048] The information can be used to determine the status associated with the subject. The status may include information relevant to predicting the subsequent status of the subject. For example, the status may indicate whether the subject is alive or dead. The status may indicate the median life expectancy of the subject. The status may indicate medically relevant somatic mutations in the tumor (e.g., KRAS variants). The status may indicate drug resistance (e.g., cetuximab resistance).
[0049] Information can be used to generate one or more decision trees that show the probabilities of various endpoints for an object exhibiting a particular state. Decision branches can originate from a root node (which can be thought of as the first decision node). Decision branches can lead to either endpoints (also called terminal nodes) or opportunity nodes. Terminal nodes or endpoints can represent states. Opportunity nodes (or event nodes) can be points of uncertainty from which different outcomes are possible. Uncertainty can be resolved through opportunity branches (event branches) originating from opportunity nodes. Each opportunity branch can lead to either a terminal node or a decision node from which multiple decision branches originate (which itself can represent a state). These decision branches can sequentially lead to endpoints or opportunity nodes in a continuous manner until all branches lead to an endpoint or terminal node.
[0050] The root node in a decision tree can be an initial state. An initial state can be as broad as "cancer diagnosis." More typically, the root node represents some aspect of the genetic profile of the subject. For example, the root node might indicate one or more gene variants detected in cfDNA, e.g., the presence of variants in a particular oncogene and / or their quantity compared to normal DNA. Each decision branch from the root node can represent a different treatment process (or no treatment). For example, a treatment process could represent a different chemotherapy or immunotherapy regimen, a type of surgery, or radiation therapy. Terminal nodes can represent a state, e.g., survival or death, e.g., within a specific time period from diagnosis (e.g., 5-year survival rate). Decision nodes represent new states from which new decisions can be made. For example, a decision node might be the emergence of gene variants resulting in chemotherapy resistance. Such variants might represent avoidance pathways where the tumor evades a response to chemotherapy, requiring a different therapeutic approach.
[0051] Advantageously, the methods disclosed herein can generate predictive algorithms configured to determine the probability that any therapeutic intervention applied to a particular condition (e.g., a specific chemotherapeutic agent for a cancer with a particular genetic profile) will result in a particular condition (e.g., a genetic variant) in which the cancer can evade the therapeutic intervention. Such probabilities may be determined through multiple treatments and evasions. As a result, it can be determined that a particular set of therapeutic interventions will result in a particular mode of evasion, ultimate evasion (e.g., death), or a condition in which the cancer is undetectable at a given frequency or probability.
[0052] This disclosure provides a method for generating a predictive algorithm for assigning probabilities to each branch or each terminal node in a decision tree. The method may utilize a database from which the results for each branch can be calculated from multiple subjects from which data is stored. Probabilities can be determined, for example, by taking a training set of subjects, classifying them into states, recording treatments and / or therapeutic interventions, and then determining the frequency of outcomes (e.g., final states). The probability can then be determined using the frequency of a given outcome in the training set.
[0053] Therefore, for multiple subjects exhibiting a particular condition, multiple decision branches may be identified, and the opportunity for a particular endpoint or the decision node at the end of that branch can be determined. For example, referring to Figure 6, among individuals exhibiting the condition "EGFR mutant," the decision branches may include treatment A and treatment B.
[0054] In Figure 6, treatment A leads to opportunity node A, and treatment B leads to opportunity node B. Opportunity node A leads to a 75% 5-year survival rate (terminal node) and a 25% occurrence of "avoidance A" (decision node A) at that point. Avoidance A may have one decision branch - treatment C. This leads to opportunity node C, from which two opportunity branches arise to the terminal node: 40% 5-year survival and 60% death. In total, this branch has an 85% chance of 5-year survival and a 15% chance of death.
[0055] In Figure 6, treatment B leads to opportunity node B. Opportunity node B leads to a 60% five-year survival rate (terminal node) and a 40% occurrence of "avoidance B" (decision node B) at that point. Avoidance B may have one decision branch - treatment D. This leads to opportunity node D, from which two opportunity branches arise to the terminal node: 40% five-year survival and 60% death. In total, this branch has a 76% chance of five-year survival and a 24% chance of death.
[0056] The reliability of the final probability determined can be increased by adding more data points (objects) at any decision node. In some cases, the initial state can be used to predict subsequent states (e.g., intermediate states (e.g., at a decision node) or final states). In some cases, the initial state is classified as one that leads to a subsequent state (e.g., an intermediate state or final state) with a given frequency. A subsequent state is a state achieved after a decision from a previous state. For example, after state 1, a therapeutic intervention is applied, and the state at a later point in time is the subsequent state. A subsequent state may be a terminal state from which no further decisions are made, or it may be an intermediate state from which another decision is made.
[0057] The initial state can be determined by clustering subjects based on information or a subset of information determined about them. Clusters can be generated using information about subjects or a training set of subjects. For example, the information may be categorical (e.g., whether a KRAS variant is present or absent in a tumor sample), and subjects may be clustered based on common categorical values. In some cases, the information about subjects is quantitative. Subjects may be clustered using quantitative data by any method known in the art. Exemplary methods include, but are not limited to, k-means clustering, hierarchical clustering, or centroid-based clustering. Clustering may be based on a visual inspection of data, including data projected onto a reduced number of dimensions by methods such as principal component analysis. Clustering can be used to create cluster boundaries that define which cluster a subject belongs to.
[0058] A profile includes values (quantitative or qualitative) for each of one or more characteristics. A profile may include information about, for example, phenotypic characteristics, genetic characteristics, demographic characteristics, or medical history (including a history of delivered therapeutic interventions). A genetic profile includes values for various genetic characteristics, such as gene variants at a particular locus (e.g., sequence information for copy number information). For example, a genetic profile may include germline or somatic genotypes at several loci in a diseased (e.g., cancer) cell. A state can be one or more values for a characteristic in a profile.
[0059] The information may include tumor profiles, including the genetic profile of the tumor. The information may include subject profiles, including genetic information about the subject. The information may include previous treatments or therapeutic interventions received by the subject.
[0060] The tumor profile may include the tissue of origin, tumor volume, tumor drug sensitivity, tumor stage, tumor size, tumor metabolic profile, tumor metastasis status, tumor volume, or tumor heterogeneity.
[0061] A tumor profile may include a genetic profile of the tumor, which can be obtained by various methods. For example, a genetic profile of a tumor can be obtained by analyzing nucleic acids derived from a biological sample from the subject by high-throughput sequencing or genotyping arrays. Nucleic acids can be DNA or RNA. Nucleic acids are isolated from the sample. The sample used to obtain the genetic profile may be a tumor biopsy, a microneedle aspiration biopsy, or a cell-free fluid containing nucleic acids derived from tumor cells. For example, the cell-free fluid may be derived from a fluid selected from the group consisting of the subject's blood, plasma, serum, urine, saliva, mucosal secretions, sputum, feces, cerebrospinal fluid, and tears.
[0062] For example, blood from a subject at risk of cancer can be collected and prepared as described herein to generate a cell-free polynucleotide population. In one embodiment, this is cell-free DNA (cfDNA). The systems and methods of this disclosure can be used to detect mutations or copy number variations that may be present in a particular cancer that is present. The methods may be useful in detecting the presence of cancer cells in the body regardless of the absence of disease symptoms or other characteristics.
[0063] Methods for the extraction and purification of nucleic acids are well known in the art. For example, nucleic acids can be purified by organic extraction using phenol, phenol / chloroform / isoamyl alcohol, or similar formulations including TRIzol and TriReagent. Other non-limiting examples of extraction techniques include (1) organic extraction followed by, for example, automated nucleic acid extraction equipment, e.g., Applied Biosystems (Foster Nucleic acid precipitation methods include (1) ethanol precipitation, (2) stationary phase absorption, and (3) salt-induced precipitation methods, which are typically referred to as “salting-out” methods, using phenol / chloroform organic reagents, with or without the use of Model 341 DNA Extractor, available from City, CA. Another example of nucleic acid isolation and / or purification may involve using magnetic particles to which nucleic acids bind specifically or nonspecifically, followed by isolating the beads using a magnet, washing the nucleic acids, and eluting them from the beads. In some embodiments, the isolation methods described above may be preceded by an enzymatic digestion step, e.g., digestion with proteinase K or other similar proteases, which helps to remove undesirable proteins from the sample. If desired, RNase inhibitors may be added to the lysis buffer. For certain cell or sample types, it may be desirable to add a protein denaturation / digestion step to the protocol. Purification methods may be induced to isolate DNA, RNA, or both. If both DNA and RNA are isolated together during or after the extraction procedure, further steps may be taken to purify one or both separately from each other. Partial fractions of the extracted nucleic acids can also be produced by purification based on, for example, size, sequence, or other physical or chemical properties.
[0064] Sequencing can be performed on polynucleotides extracted from a sample to generate sequencing reads. Exemplary sequencing techniques include, for example, emulsion polymerase chain reaction (PCR) (e.g., pyrosequencing from Roche 454, semiconductor sequencing from Ion Torrent, SOLiD sequencing by ligation from Life Technologies, and synthesis sequencing from Intelligent Biosystems), bridge amplification on a flow cell (e.g., Solexa / Illumina), isothermal amplification using Wildfire technology (Life Technologies), or rolony / nanoballs generated by rolling circle amplification (Complete Genomics, Intelligent Biosystems, Polonator). Sequencing techniques that enable direct sequencing of single molecules without prior clonal amplification, such as Heliscope (Helicos), SMRT technology (Pacific Biosciences), or nanopore sequencing (Oxford Nanopore), may be suitable sequencing platforms. Sequencing may be performed with or without enrichment of the target. Exemplary genes and / or regions that may be enriched are found in Figure 5. Enrichment may be performed, for example, by hybridization of a nucleic acid sample or sequencing library to probes arranged in an array or attached to beads. In some cases, polynucleotides derived from the sample are amplified by any preferred approach (e.g., PCR) before and / or during sequencing.
[0065] As a non-limiting example, a sample containing initial genetic material is provided from which cell-free DNA can be extracted. The sample may contain target nucleic acids in low abundance. For example, nucleic acids from a normal genome or germline genome may be dominant in the sample, and the sample may also include nucleic acids from at least one other genome containing genetic alterations, such as a cancer genome, fetal genome, or genome from another individual or species, in amounts not exceeding 20%, not exceeding 10%, not exceeding 5%, not exceeding 1%, not exceeding 0.5%, or not exceeding 0.1%. The initial genetic material can then be converted into a set of tagged parental polynucleotides, sequenced, and sequenced reads can be obtained. In some cases, these sequence reads may contain barcode information. In other examples, barcodes are not used. Tagging may involve attaching sequence tags to molecules in the initial genetic material. Sequence tags may be selected such that all unique polynucleotides mapping to the same reference sequence have a unique identification tag. Sequence tags can be selected such that not all unique polynucleotides mapping to the same reference have unique identification tags. The conversion can be performed with high efficiency on, for example, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the initial nucleic acid molecules. The set of tagged parent polynucleotides can be amplified to obtain an amplified set of progeny polynucleotides. The amplification can be, for example, at least 10, 100, 1,000, or 10,000 times. The amplified set of progeny polynucleotides is sampled for sequencing at both sampling rates such that the resulting sequencing reads (1) cover a target number of unique molecules in the set of tagged parent polynucleotides, and (2) cover unique molecules in the set of tagged parent polynucleotides at a certain target coverage factor (e.g., 5 to 10 times coverage of the parent polynucleotides). The set of sequencing reads can be aggregated to obtain a set of consensus sequences corresponding to unique tagged parent polynucleotides. Sequencing leads may undergo quality checks before being included in the analysis.For example, sequencing leads that do not meet the quality control score may be removed from the pool.
[0066] Sequencing reads can be sorted into families representing reads of progeny molecules derived from a specific, unique parent molecule. For example, a family of amplified progeny polynucleotides may consist of these amplified molecules derived from a single parent polynucleotide. By comparing the sequences of progeny within a family, the consensus sequence of the original parent polynucleotide can be estimated. This yields a set of consensus sequences representing unique parent polynucleotides within a tagged pool. This process may allow sequences to be assigned a confidence score. After sequencing, reads may be assigned a quality score. The quality score can be an indication of whether these reads may be useful in subsequent analyses, based on a threshold. In some cases, some reads are not of sufficient quality or length to perform the subsequent mapping step. Sequencing reads with a predetermined quality score (e.g., above 90%) can be filtered from the data. Sequencing reads that meet a specified quality score threshold can be mapped to a reference genome or a template sequence known to be free of copy number variations. After mapping alignment, sequencing reads may be assigned a mapping score. A mapping score can be a representation or read that maps back to a reference sequence, indicating whether each position is uniquely mapped or not. In some cases, the reads may be sequences unrelated to copy number variation analysis. For example, some sequencing reads may originate from contaminating polynucleotides. Sequencing reads with mapping scores indicating that they are mismapped (e.g., incorrectly mapped) by at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% may be filtered from the dataset. In other cases, sequencing reads assigned a mapping score lower than a given percentage may be filtered from the dataset.
[0067] Sequencing reads that meet a specified quality score threshold can be mapped to a reference genome or a template sequence known to be free of copy number variations. After mapping alignment, the sequencing reads may be assigned a mapping score. In some cases, the reads may be sequences unrelated to copy number variation analysis. After data filtering and mapping, multiple sequencing reads generate chromosomal coverage regions. These chromosomal regions can be divided into windows or bins of variable length. In some cases, each window region may be sized to contain approximately the same number of uniquely mappable bases. In addition, certain windows throughout the genome that are difficult to sequence or have substantially high GC bias may be filtered from the dataset. For example, regions known to fall near the centromere of a chromosome (i.e., centromere DNA) are known to contain high-frequency repeat sequences that can lead to false-positive results. These regions may be filtered. Normalization may be performed to compensate for the effect of GC content on the sequencing reads of a sample. Other regions of the genome, such as regions containing microsatellite DNA, which have abnormally high concentrations, can be filtered out of the dataset.
[0068] For exemplary genomes derived from cell-free polynucleotide sequences, the next step involves determining read coverage for each window region. This can be done using either barcoded or unbarcoded reads. In the case of unbarcoded reads, previous mapping steps may provide coverage for different base positions. Sequencing reads that have sufficient mapping and quality scores and fall within the unfiltered chromosomal window can be counted. The number of coverage reads may be assigned a score for each mappable position. If barcodes are involved, all sequences having the same barcode, physical properties, or a combination of these two may be aggregated into a single read, as they all originate from the sample parent molecule. This step can reduce any bias that may have been introduced during any of the preceding steps, such as steps involving amplification. For example, if one molecule is amplified 10 times while another is amplified 1000 times, each molecule will appear only once after aggregation, thereby eliminating the effects of uneven amplification. Only reads with unique barcodes may be counted at their respective mappable positions, which may influence the assigned score. For this reason, it is important to perform the barcode ligation step in a manner optimized to minimize the amount of bias that may arise. The sequence of each base can be aligned as the most dominant nucleotide read for that particular location. Furthermore, the number of unique molecules can be counted at each location to lead to simultaneous quantification at each location. This step can reduce any bias that may have been introduced during any of the preceding steps, such as steps involving amplification.
[0069] By utilizing the distinct copy number states of each window region, copy number variations in chromosomal regions can be identified. In some cases, all adjacent window regions with the same copy number can be merged into a single segment to report the presence or absence of a copy number variation state. In some cases, various windows can be merged with other segments after filtering.
[0070] Methods for determining gene profiles (e.g., tumor or target gene profiles) may have error rates. For example, sequencing methods may have base-to-base error rates of approximately 0.1%, 0.5%, 1%, or higher. In some cases, nucleic acids derived from tumor cells containing gene variants at a given locus are present in small amounts within the total nucleic acids containing that locus, at a rate similar to or lower than the base-to-base sequencing error rate. In such situations, it can be difficult to distinguish between genotyping or sequencing errors and the presence of low-frequency gene variants. Certain techniques may be employed to reduce error rates, such as those described in WO2014 / 149134 (which is incorporated in its entirety by reference).
[0071] The gene profile of a tumor may include somatic mutations compared to a reference. The reference may be a reference genome, such as a human reference genome. The reference genome may also be the germline genome of the subject. The gene profile may include various gene variants acquired by some or all of the tumor cells. Gene variants may be, for example, single nucleotide variants, whole or small structural variants, or short insertions or deletions. For example, as shown in Figure 2A, common abnormalities in cancer genomes can result in abnormal chromosome number (aneuploidy) and chromosomal structure of the cancer genome. In Figure 2A, the upper line represents the germline genome, and the lower line represents the cancer genome with somatic abnormalities. Double lines are used when it is useful to differentiate between heterozygous and homozygous changes. Dots represent single nucleotide changes, while lines and arrows indicate structural changes.
[0072] The gene profile of a tumor can contain quantitative information about each variant. For example, gene analysis of cell-free DNA by digital sequencing can yield 1,000 reads mapping to the locus of a first oncogene, of which 900 reads correspond to germline sequences and 100 reads correspond to variants present in tumor cells. The same gene analysis can yield 1,000 reads mapping to the locus of a second oncogene, of which 980 reads correspond to germline sequences and 20 reads correspond to variants representing 10% tumor volume. While the total tumor volume is approximately 10% of cell-free DNA based on the locus of the first oncogene, it can be inferred that a small fraction (approximately 20%) of tumor cells may have variants at the locus of the second oncogene. Such quantitative information can be included in the gene profile of a tumor and can be monitored over time or in response to treatment.
[0073] The genetic profile of a tumor may include information about somatic variants. These include, but are not limited to, mutations, insertions and deletions (insertions or deletions), copy number variations, conversions, transpositions, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural alterations, gene fusions, gene shortenings, gene amplifications, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, abnormal changes in nucleic acid methylation, infections, and cancer.
[0074] In some cases, genotyping involves performing genotyping of nucleic acids from cell-free fluid. Such methods can capture genetic information from multiple tumor cells, making it possible to infer information about both tumor heterogeneity and tumor evolution. In some cases, genotyping may be performed on samples obtained from at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten time points. In some cases, genotyping involves determining the genotypes of at least 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 120, 140, 160, 180, or 200, or more, gene loci. In some cases, a gene locus is a gene. In some cases, a gene locus is an oncogene. An oncogene is a gene containing a mutation that drives tumor growth. Exemplary oncogenes can be found in WO2009045443, which is incorporated herein by reference in its entirety. Oncogenes may include the genes listed in Figure 5.
[0075] In some cases, the genetic profile of a tumor may contain information about tumor evolution. For example, if the proportion of cell-free DNA from tumors with KRAS mutations is increasing, it can be inferred that the proportion of tumor cells resistant to specific treatments targeting KRAS is increasing over time. Figure 1 shows an exemplary method for developing a model of tumor evolution in response to treatment. The process in Figure 1 involves collecting genetic profile data from multiple target tumors, as well as tumor treatments and initial treatments (10). The genetic profiles can be used to identify or infer evolutionary avoidance pathways taken by tumor cells that cause resistance to treatment (12). The genetic profiles of individual target tumors can be fitted to a model to obtain the probability that tumor cells acquire gene variants that result in resistance to treatment (14).
[0076] Using more complex models, tumor heterogeneity can be measured, for example, based on the relative incidence of different variants in cell-free DNA. Figure 2B shows an exemplary system for determining the probabilities of various state outcomes. This system may be a Hidden Markov Model (HMM), which is a statistical Markov model in which the system being modeled is assumed to be a Markov process with unobserved (hidden) states. In a simple Markov model (similar to a Markov chain), the states are directly visible to the observer, and therefore the state transition probabilities are the only parameter. In a Hidden Markov model, the states are not directly visible, but the state-dependent outputs are visible. Each state has a probability distribution across all possible output tokens. Therefore, the sequence of tokens produced by the HMM provides information about the sequence of states. A Hidden Markov model can be thought of as a generalization of a mixture model in which hidden (or latent) variables that control which mixture components are selected for each observation are not independent of each other but are related through the Markov process. As shown in Figure 2B, an HMM is typically defined by a set of hidden states, a matrix of state transition probabilities, and a matrix of output probabilities. Common methods for constructing such models include, but are not limited to, Hidden Markov Models (HMMs), Artificial Neural Networks, Bayesian Networks, Support Vector Machines, and Random Forests. Such methods are known to those skilled in the art and are incorporated herein by reference in their entirety by Mohri et al., Foundations of Machine Learning (2012), published by MITPress, and are incorporated herein by reference in their entirety by reference. This is described in detail in MacKay's *Information Theory, Inference, and Learning Algorithms* (2003), published by Cambridge University Press.
[0077] The relative amount of tumor polynucleotides in a cell-free polynucleotide sample is referred to herein as “tumor load.” Tumor load may be related to tumor size. When tested over time, tumor load can be used to determine whether cancer is progressing, stable, or in remission. In some embodiments, the confidence intervals for the estimated tumor load do not overlap and indicate the direction of disease progression. Tumor load and the direction of disease progression may have a diagnostic confidence indication. When used herein, the term “diagnostic confidence indication” refers to an expression, number, rank, degree, or value assigned to indicate the presence of a gene variant and the degree to which its presence is trusted. For example, this expression may, among other things, be a binary value or an alphanumeric rank from A to Z. In yet another embodiment, the diagnostic confidence indication may, among other things, have any value from 0 to 100. In yet another embodiment, the diagnostic confidence indication may be expressed by a range or degree, for example, “low” or “high,” “more” or “less,” “increased” or “decreased.” A low diagnostic confidence index may mean that the presence of a gene variant is not very reliable (the gene variant may be noise). A high diagnostic confidence index may mean that the presence of a gene variant is likely, and in one embodiment, if the diagnostic confidence index is lower than 25-30 out of 100, the result is considered unreliable.
[0078] In one implementation, measurements from multiple samples taken substantially simultaneously or over multiple time points can be used to adjust the diagnostic reliability index for each variant to indicate the reliability of predicting the observation of copy number variation (CNV) or mutation. Reliability can be increased by using measurements at multiple time points to determine whether the cancer is progressing, in remission, or stable. The diagnostic reliability index can be assigned by any of a number of known statistical methods and may be at least in part based on the frequency with which measurements are observed over a period of time. For example, statistical correlation between current and previous results can be performed. Alternatively, a hidden Markov model may be constructed so that for each diagnosis, the maximum likelihood or maximum posterior decision can be made based on the frequency with which a particular test event occurs from multiple measurements or time points. As part of this model, the probability of error for a particular decision and the resulting diagnostic reliability index may also be output. In this form, the parameter measurements may be provided with confidence intervals, regardless of whether they are within the noise range. By testing over time and comparing confidence intervals over time, the predictive reliability of whether the cancer is progressing, stable, or in remission can be increased. The two points in time can be separated by approximately one month to one year, one year to five years, or no more than three months.
[0079] Figure 2C shows an exemplary model generated by the system in Figure 2B for inferring tumor phylogeny from next-generation sequencing data. Subclones are related to one another through the process of evolutionary mutation acquisition. In this example, three clones (leaf nodes) are characterized by different combinations of four single nucleotide variant (SNV) sets A, B, C, and D. The percentages at the edges of the tree indicate the fraction of cells having this particular set of SNVs; for example, 70% of all cells have A, 40% have B, and only 7% have A, B, and D.
[0080] Figure 2D shows an exemplary heterogeneous assortment of normal cells and cancer subclones that develops during the evolutionary history of a tumor. The evolutionary history of the tumor results in a heterogeneous assortment of normal cells (small disks) and cancer subclones (large disks, triangles, squares). Internal nodes that have been completely replaced by their descendants (such as those containing SNV sets A and B but lacking C or D) are no longer part of the tumor.
[0081] A partnership may be established between a medical prognostic provider and one or more healthcare service providers, such as a physician, a hospital, a health insurer (e.g., Blue Cross), or a managed care organization (e.g., Kaiser Permanente). The healthcare service provider may provide the medical prognostic provider with one or more subject samples containing cfDNA, and one or more medical records containing medical information, in addition to or in addition to genetic information about the subject. The medical information may be provided through a secure communication link that enables the medical prognostic provider to access the medical records. The medical prognostic provider may sequence (or have sequenced) cfDNA derived from the sample and create a medical record containing the information to be used in the method of this disclosure. The healthcare service provider may provide new samples containing cfDNA and / or update the information subject through the decision node. The predictive model may be iteratively updated as new information becomes available.
[0082] Figure 3 provides an overview of the process for determining the gene profile. This process receives genetic material derived from a blood sample or other body sample (102). The process converts polynucleotides derived from the genetic material into tagged parental nucleotides (104). The tagged parental nucleotides are amplified to obtain amplified progeny polynucleotides (106). A subset of the amplified polynucleotides is sequenced to obtain sequencing reads (108), which are then grouped into families, each generated from a unique tagged parental nucleotide (110). At selected loci, the process assigns a confidence score to each family (112). Consensus is then determined using previous reads. This is done by considering the previous confidence scores for each family, and if consistent previous confidence scores exist, the current confidence score increases (114). If previous confidence scores exist but are inconsistent, the current confidence score is not modified in one embodiment (116). In other embodiments, the confidence score is adjusted in a predetermined manner with respect to inconsistent previous confidence scores. If it is the first time a family has been detected, the current confidence score may be reduced because it may be a misreading (118). Based on the confidence score, the process can estimate the frequency of the family at loci within the set of tagged parental polynucleotides (120).
[0083] Transient information can enhance information regarding the detection of mutations or copy number variations, but other consensus methods may be applied. In other embodiments, historical comparisons can be used in conjunction with other consensus sequences that map to a specific reference sequence to detect events of gene variation. Consensus sequences that map to a specific reference sequence may be measured and normalized against a control sample. Measurements of molecules that map to a reference sequence can be compared across the entire genome to identify regions within the genome where copy number variations or heterozygosity is lost. Consensus methods include, for example, linear or nonlinear consensus sequence construction methods derived from digital communications theory, information theory, or bioinformatics (e.g., voting, averaging, statistical, maximum posterior probability or maximum likelihood detection, dynamic programming, Bayesian, Hidden Markov, or support vector machine methods). After determining sequence read coverage, a probabilistic modeling algorithm is applied to convert the normalized nucleic acid sequence read coverage for each window region into distinct copy number states. In some cases, this algorithm may include one or more of the following: Hidden Markov models, dynamic programming, support vector machines, Bayesian networks, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering techniques, and neural networks.
[0084] A report may be generated afterward. For example, copy number variations (CNVs) may be reported as graphs showing the corresponding increase, decrease, or maintenance of copy number variations at various locations within the genome. In addition, copy number variations can be used to report a percentage score indicating the amount of disease material (or nucleic acids with copy number variations) present in a cell-free polynucleotide sample.
[0085] Figure 4 shows a schematic diagram of access to reports for cancer patients via the internet. The system in Figure 4 can use a portable or desktop DNA sequencer. A DNA sequencer is a scientific instrument used to automate the DNA sequencing process. In the case of a DNA sample, a DNA sequencer is used to determine the order of four bases: adenine, guanine, cytosine, and thymine. The order of the DNA bases is reported as a string called reads. Some DNA sequencers can also be considered optical instruments, as they analyze light signals originating from fluorescent dyes attached to the nucleotides.
[0086] A tumor profile may contain information about the tissue origin of the tumor. The types and number of cancers that can be detected and profiled include, but are not limited to, blood cancers, brain cancers, lung cancers, skin cancers, nasal cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, colorectal cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, oral cancers, stomach cancers, solid tumors, heterogeneous tumors, and homogeneous tumors.
[0087] A tumor profile may contain information about the drug sensitivity of a tumor. Tumor drug sensitivity can be determined directly by measuring or assessing the response of isolated tumor cells to the drug of interest. Tumor drug sensitivity can also be determined by performing tumor genotyping.
[0088] A tumor profile may include information about the size and / or stage of the tumor. Tumor size can be measured by body scanning techniques, surgical procedures, or any known method. Tumor stage can be determined based on physical examination, imaging studies, laboratory tests, pathology reports, and / or surgical reports.
[0089] The target profile may include the target gene profile. The target gene profile can be determined by assaying non-cancerous tissue from which the target originates. The target gene profile can also be determined by assaying nucleic acids derived from cell-free fluid from which the target originates. Nucleic acids from non-cancerous tissue can be identified, for example, by their frequency in the initial nucleic acid pool or by the length of the nucleic acid molecule. Nucleic acid molecules from tumor cells may have a first mode of 160–180 bases and a second mode of 320–360 bases. Nucleic acid molecules from non-cancerous tissue may have a broader distribution, and many molecules are longer than 400 bases. Molecular size can be controlled by size selection of the initial DNA molecule or library fragment, or by infographically controlling it by mapping paired reads to a reference genome.
[0090] The target gene profile may involve assaying for variants that may alter the effect of a treatment. For example, such variants may affect the pharmacokinetics of a drug. Common variants that affect pharmacokinetics may affect drug transport or drug metabolism. Variants that affect pharmacokinetics are described in *The Handbook of Anticancer Pharmacokinetics and Pharmacodynamics* by MA Rudek et al., published by Springer Science & Business Media in 2014, which is available by reference in its entirety. This specification is incorporated herein.
[0091] The target gene profile may include assaying for variants that affect cancer progression. Such mutations may be hereditary mutations that reduce the efficiency of tumor suppressor products, such as TP53 or BRCA1.
[0092] In some embodiments, the subject profile includes information other than genetic information. Such information may include the subject's age, the effectiveness of other medications the patient has received, clinical information about the subject, and family medical history. Clinical information about the subject may include additional clinical information, such as organ function, e.g., liver and kidney function, blood cell count, cardiac function, lung and respiratory function, and infection status. Clinical information about the subject may include age, sex, gender, genetic profile, enzyme levels, organ function, quality of life, frequency of medical interventions, remission status, and / or patient outcome. The subject profile may include information about previous treatments. Treatments may be, for example, surgical removal, radiation therapy, or chemotherapy. The information may be qualitative (indicating which treatment was received) or quantitative, e.g., including information on dose, duration, and timing. Information about the subject may include whether the subject is alive or dead. By collecting information on the target population at various points in time, it is possible to generate median survival rates, 6-month survival rates, 1-year survival rates, 2-year survival rates, 3-year survival rates, 5-year survival rates, or longer-term survival rates for the target population.
[0093] Determining a state (e.g., the initial state) may involve obtaining information about an object and assigning it to a state based on that information. In some cases, the state is determined based on a subset of information. For example, the state can be determined by clustering objects from a training set, and a new object can be assigned to a state by determining which cluster it is closest to.
[0094] Clustering can be used to transform quantitative data into categorical data. For example, certain cancer medications can cause liver damage. The levels of liver enzymes (e.g., AST and ALT) in the blood of subjects receiving such cancer medications can be measured. Clustering or visual inspection of liver enzyme levels can reveal which subjects have elevated liver enzyme levels and which have normal levels. Liver enzyme levels can be transformed into categorical variables by defining subjects with liver enzyme levels higher than a given level as "elevated" and those with levels lower than a given level as "normal".
[0095] Categorical and quantitative data can be combined. In one exemplary method, categorical data can be transformed for use in methods requiring quantitative data by converting the categorical data into "dummy values." For example, patients with elevated liver enzyme levels may be assigned a value of 1, while patients with normal liver enzyme levels may be assigned a value of 0. Other methods for converting categorical variables to quantitative variables include effect coding, contrast coding, and nonsense coding.
[0096] The state may represent the desired outcome (e.g., survival, remission status, or length of time before tolerance develops), and this can be recorded. A set of subjects (e.g., a training set) can be used to determine the effect size and interaction of the initial state and / or treatment with respect to the desired outcome being determined. These effect sizes and interactions can be used to develop a classifier or predictive model. Methods for determining the effect size and interaction conditions of traits from the initial state include, for example, regression analysis including linear and logarithmic regression analysis; shortest contraction centroid analysis; stabilized linear discriminant analysis; support vector machines; Gaussian processes; conditional inference tree forests; random forests; shortest centroid; naive Bayes; projection-tracking LDA trees; multinomial logistic regression; stamped decision trees; artificial neural networks; binary decision trees; and / or conditional inference trees. The accuracy and sensitivity of the classifier or predictive model can be determined by measuring predictive accuracy in a subset of subjects not used to build the classifier or predictive model (e.g., a test set).
[0097] In some cases, the effect size of predictors is determined, and variables with little influence are removed. Methods for variable selection are known in the art, and examples include filtering and / or wrapping methods for variable selection. Filtering methods are based on general characteristics, such as the correlation between variables and outcomes. Wrapping methods evaluate a subset of variables together to determine the optimal combination of variables. The selected variables can be used to determine a subset of information used to determine the state of the subject.
[0098] In some cases, the training set consists of individuals with the same histological type of tumor. In some cases, the subjects have similar demographic profiles, such as being of the same gender, age, ethnic background, or having the same risk factors. Gender can be male or female. Exemplary risk factors include alcohol consumption, tobacco use and usage, diet, exercise, work-related exposure to carcinogens, frequency of travel, and exposure to ultraviolet light and / or sunburn. In some cases, all subjects in the training set are patients with cancer. In some cases, all subjects in the training set are patients with symptoms consistent with cancer and are undergoing testing for cancer. In some cases, the subjects in the training set are patients with symptoms consistent with cancer and are undergoing treatment for cancer. Subject characteristics may be included in the information about each subject in a group of subjects.
[0099] The probability of a given subsequent state of an object can be determined using its initial state. This probability can be determined using a classifier or predictive model.
[0100] A classifier or predictive model can be used to identify a preferred treatment for a subject with a given profile. For example, using a classifier or predictive model to determine the probability of a given outcome for a subject may involve generating one or more decision trees. The state at a first time point can be represented by a root node (which is the initial decision node), and alternative treatments can be represented by decision branches. In some cases, decision branches may lead to terminal states (from which no further decisions are made) or intermediate state nodes, which themselves may be decision nodes. Intermediate state nodes may represent the appearance of gene variants that confer tumor resistance to the treatment within one or more tumors of the subject; the results of subsequent biopsy or imaging procedures; and / or generally, changes or absences of information from the subject at a given time point. For example, intermediate nodes may include information from the subject one week, two weeks, three weeks, four weeks, one month, two months, three months, six months, one year, two years, three years, four years, or five years after treatment. Intermediate nodes may represent intermediate states where healthcare providers make decisions regarding future treatment options (e.g., after completing a chemotherapy regimen, after surgical intervention to remove a tumor, and at specific points in an active monitoring regimen).
[0101] Intermediate nodes may contain information regarding the development of resistance to the treatment. For example, the presence of a particular variant in a tumor may indicate the development of resistance. An increase in a particular variant over time between treatments may indicate that the variant, or at least a second unidentified variant, is associated with the development of resistance to the treatment. The probability of such a variant appearing may be altered by the presence of a particular variant that predisposes tumor down in a specific evolutionary track. Intermediate nodes may contain information regarding the health of the subject (e.g., the patient).
[0102] Tumor profiles and / or subject profiles can be determined at one or more subsequent time points. Information from the tumor and / or subject profiles at subsequent time points can be used to determine subsequent states. Once a subsequent state is determined, it can be used as a new initial state to update the probabilities of other subsequent nodes. For example, if a subject develops a KRAS variant that does not occur concurrently with a KRAS gene amplification event, the decision tree can be updated to reflect the reduced probability of the KRAS gene amplification event.
[0103] In some cases, the subsequent state is represented by a terminal node (e.g., the subject dies or achieves complete remission). The subsequent state may be a point in time after the procedure. The subsequent state may be a point in time when an additional biopsy is taken. The biopsy may be a liquid biopsy.
[0104] In some cases, the terminal node represents a state where no further medical decisions are made. In some cases, the terminal node represents the death of the subject. In some cases, the terminal node represents the failure to detect cancer in the subject.
[0105] In some cases, recommending an action involves determining which cluster, generated by a classifier or predictive model, the information from the subject belongs to. This determination may be based on cluster boundaries determined by the methods described above. In some cases, the decision may be based on selecting the cluster to which the information from the subject is closest. This selection may be based, at least in part, on distance correlation.
[0106] Using such classifiers or predictive models, a patient's treatment can be selected. For example, for a patient with a given genetic profile and tumor genetic profile, a treatment that maximizes survival (e.g., 5-year survival rate and / or remission rate) may be selected. The patient may be monitored over time. If a genetic mutation occurs that confers resistance to the treatment or increases the risk of developing resistance to the treatment, a second or different treatment that maximizes 5-year survival rate and / or remission may be administered based on the new condition. The treatment appropriate to maximize the patient's survival rate and / or lifespan may be selected.
[0107] The procedure is publicly known to those skilled in the art, and examples are described in the NCCN's Clinical Practice Guidelines in Oncology® or the American Society of Clinical Oncology (ASCO) clinical practice guidelines. Examples of drugs used in the procedure can be found in CMS-approved general overviews, including the National Comprehensive Cancer Network (NCCN) Drugs and Biologics Compendium®, Thomson Micromedex's DrugDex®, Elsevier Gold Standard's Clinical Pharmacology General Overview, and the American Hospital Formulary Service-Drug Information Compendium®. Computer system
[0108] This disclosure provides a computer system programmed to implement the method of this disclosure. Figure 7 shows a computer system 701 programmed, or otherwise configured, to detect or monitor cancer evolution.
[0109] The computer system 701 includes a central processing unit (CPU, also referred to herein as “processor” and “computer processor”) 705, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 701 also includes memory or memory locations 710 (e.g., random-access memory, read-only memory, flash memory), electronic storage units 715 (e.g., hard disks), a communication interface 720 for communicating with one or more other systems (e.g., a network adapter), and peripheral devices 725, e.g., a cache, other memory, a data storage unit, and / or an electronic display adapter. The memory 710, the storage unit 715, the interface 720, and the peripheral devices 725 are in communication with the CPU 705 via a communication bus (solid line), e.g., a motherboard. The storage unit 715 may be a data storage unit (or data repository) for storing data. The computer system 701 may be operably connected to a computer network ("network") 730 by utilizing the communication interface 720. Network 730 may be the Internet, the Internet and / or an extranet, or an intranet and / or extranet communicating with the Internet. In some cases, Network 730 may be a long-distance communication and / or data network. Network 730 may include one or more computer servers that can enable distributed computing, such as cloud computing. In some cases, Network 730 may implement a peer-to-peer network that leverages a computer system 701, allowing devices connected to the computer system 701 to act as clients or servers.
[0110] The CPU 705 can execute a set of machine-readable instructions, which may be embodied in a program or software. Instructions may be stored in memory locations, such as memory 710. Instructions may be directed to the CPU 705, which may then be programmed or otherwise configured to implement the methods of this disclosure. Examples of operations performed by the CPU 705 include fetching, decoding, executing, and writing back.
[0111] The CPU 705 may be part of a circuit, such as an integrated circuit. One or more other components of system 701 may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0112] The storage unit 715 can store files, such as drivers, libraries, and saved programs. The storage unit 715 can also store user data, such as user preferences and user programs. In some cases, the computer system 701 may include one or more additional data storage units located external to the computer system 701, for example, on a remote server that communicates with the computer system 701 via an intranet or the internet.
[0113] Computer system 701 can communicate with one or more remote computer systems via network 730. For example, computer system 701 can communicate with a user's (e.g., a patient or healthcare provider) remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), slate or tablet PCs (e.g., Apple® iPad®, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone®, Android devices, Blackberry®), or personal digital assistants. Users can access computer system 701 via network 730.
[0114] The methods described herein may be implemented using machine-executable code (e.g., a computer processor) stored in an electronic storage location of a computer system 701, such as memory 710 or an electronic storage unit 715. The machine-executable code or machine-readable code may be provided in the form of software. When used, the code may be executed by the processor 705. In some cases, the code may be retrieved from the storage unit 715 and stored in memory 710 for immediate access by the processor 705. In some situations, the electronic storage unit 715 may be omitted, and the machine-executable instructions are stored in memory 710.
[0115] The code is precompiled, and may be configured to be so, or may be compiled at runtime, for use in a machine having a processor adapted to run the code. The code may be supplied in a programming language that can be chosen to allow the code to be executed in a precompiled or immediate compilation manner.
[0116] Embodiments of systems and methods provided herein, such as computer system 701, can be embodied in programming. Various embodiments of this technology can typically be considered as “products” or “manufactured goods” in the form of machine-executable code and / or related data carried or embodied in some kind of machine-readable medium. Machine-executable code can be stored in electronic storage units, such as memory (e.g., read-only memory, random-access memory, flash memory) or hard disks. “Storage” media can include any tangible memory of computers, processors, etc., or their related modules, such as various semiconductor memories, tape drives, disk drives, etc. These can provide non-transient storage units at any point in software programming. All or part of the software can be communicated from time to time via the Internet or various other telecommunication networks. Such communication can, for example, enable the loading of software from one computer or processor to another, for example, from a management server or host computer to an application server computer platform. Therefore, other types of media that can hold software elements include those used with light waves, radio waves, and electromagnetic waves, for example, through physical interfaces between local devices, through wired and optical terrestrial networks, and over various air links. Physical elements with such waves, such as wired or wireless links, optical links, etc., can also be considered media for holding software. As used herein, unless limited to non-transient tangible “storage” media, terms such as computer or machine-readable media refer to any medium involved in providing instructions to a processor for execution.
[0117] Therefore, machine-readable media, such as computer-executable code, can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical disks or magnetic disks, or any of the storage devices in any computer, such as those that can be used to implement a database as shown in the drawing. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wires, and optical fibers, such as wires including buses in computer systems. Carrier media can take the form of electrical or electromagnetic signals or sound or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched card paper tapes, any other physical storage media with perforated patterns, RAM, ROMs, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carriers for transmitting data or instructions, cables or links for transmitting such carriers, or any other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in transporting one or more sequences of one or more instructions to a processor for execution.
[0118] The computer system 701 may include, or communicate with, an electronic display 735 having a user interface (UI) 740 for providing one or more results related to or demonstrating the evolution of cancer. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0119] The methods and systems of this disclosure can be implemented using one or more algorithms. The algorithms can be implemented in software, which are executed by the central processing unit 705. The algorithms can implement, for example, the methods of this disclosure for detecting or monitoring cancer evolution. [Examples]
[0120] (Example 1) Building a model for the development of treatment resistance Patients with cancer undergo a physical screening to determine their patient profile, including their age, gender, cancer type, cancer stage, and organ function. They have blood drawn, processed to remove cells, and obtain cell-free fluid containing nucleic acids. Nucleic acid sequencing is performed to determine the patient's genetic profile and the tumor's genetic profile. Treatment is prescribed by a physician. Patients are followed over time, and the tumor's genetic profile is obtained every three months. Patient outcomes are recorded at each point in time.
[0121] A hidden Markov model is constructed based on the probability that a patient with a given patient profile (including the patient's genetic profile) and tumor genetic profile will have a specific patient outcome at any given time point. (Example 2) Use of a model for the development of treatment resistance
[0122] Subjects with cancer are hospitalized. Subject profiles and tumor profiles are obtained. Subject profiles and tumor profiles are used as the initial state of a model, e.g., the model generated in Example 1. The outcome of the subjects is predicted based on the model, and a treatment is selected to maximize the subject's predicted survival (e.g., measured in months or years). The subject's tumor profile is updated every three months and used as a new initial state input to the model. At a given subsequent time point, the tumor profile indicates that subclones resistant to the current treatment have emerged. In response, a new treatment is selected to maximize the subject's predicted survival. The subjects are given a second treatment (e.g., a second-line treatment) that targets tumor cells resistant to the first treatment (e.g., a first-line treatment). (Example 3) Displaying targets using a decision tree
[0123] The initial node identifies the subject as a 65-year-old male with colon cancer, and the tumor profile indicates the detection of low-frequency KRAS mutations in the subject's cell-free DNA. One branch arising from the initial node represents treatment with panitumumab and cetuximab, and a second branch represents treatment with panitumumab and cetuximab in combination with a mitogen-activated protein kinase (MEK) inhibitor. These branches connect to intermediate nodes indicating resistance development and lack of resistance development. The probability of resistance development is lower in intermediate nodes along branches including co-treatment with a MEK inhibitor than in branches lacking co-treatment with a MEK inhibitor. Each intermediate node is associated with a terminal node indicating death and complete remission. The probability of complete remission is higher in terminal nodes along decision branches including co-treatment with a MEK inhibitor.
[0124] The illustrations of the embodiments described herein are intended to provide a general understanding of the structures of various embodiments. The illustrations are not intended to serve as a complete description of all elements and characteristics of devices and systems utilizing the structures or methods described herein. Numerous other embodiments may be apparent to those skilled in the art by considering this disclosure. Other embodiments may be utilized and derived from this disclosure, and as a result, structural and logical substitutions and modifications may be made without departing from the scope of this disclosure. Accordingly, this disclosure and the drawings are to be construed as illustrative, not restrictive.
[0125] One or more embodiments of this disclosure may be referenced herein, individually and / or collectively, by the term “the present invention,” merely for convenience and without any intention to arbitrarily limit the scope of this application to any particular invention or inventive concept. Furthermore, it should be understood that while certain embodiments are illustrated and described herein, any subsequent arrangements designed to achieve the same or similar objectives may be used instead of the particular embodiments shown. This disclosure is intended to cover all possible subsequent adaptations and variations of various embodiments. Combinations of the embodiments described herein, and other embodiments not specifically described herein, will be apparent to those skilled in the art by considering the description.
[0126] While preferred embodiments of the present invention are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided merely as examples. It is not intended that the present invention be limited by any specific examples provided herein. Although the present invention is described with reference to the preceding specification, the descriptions and illustrations of embodiments herein are not intended to be restrictive. Those skilled in the art will be able to conceive of numerous variations, modifications, and substitutions without departing from the present invention. Furthermore, it should be understood that all aspects of the present invention are not limited to any specific expressions, configurations, or relative proportions described herein, and that they depend on various conditions and variables. It should be understood that various alternative forms to the embodiments of the present invention described herein can be used in the practice of the present invention. Therefore, it is intended that the present invention should also encompass all such alternative forms, modifications, variations, or equivalents. The following claims define the scope of the present invention, and the methods and structures within these claims, and their equivalents, are intended to be encompassed thereby. Examples of embodiments of the present invention include the following: (Item 1) A method implemented by a computer, (a) A step of obtaining information about a plurality of subjects having cancer at a first time point, and determining a first state for each of the plurality of subjects in order to generate a first set of states based on the information at the first time point, wherein the information includes, for each of the plurality of subjects, at least a tumor gene profile obtained by performing nucleic acid genotyping from cell-free fluid and any treatment provided to the subject prior to the first time point. (b) The steps of obtaining the information relating to the plurality of objects at one or more second time points after the first time point, and determining the second state of each of the plurality of objects at each of the one or more second time points in order to generate a set of subsequent states based on the information at one of the one or more second time points, (c) A step of using the set of first states from (a) and the set of subsequent states from (b), wherein a prediction algorithm is generated that determines the probability that a given first state brings about a second state in a set of states at a later time point after the given first state. Methods that include... (Item 2) (d) With respect to the given first state in the set of states at an earlier time, the step of determining the probability that the given first state brings about the second state in the set of states at a later time, (e)(d) A step of generating an electronic output that represents the probability determined in (e)(d) The method described in item 1, further including the method described in item 1. (Item 3) A method implemented by a computer, (a) A step of obtaining information about a plurality of subjects having cancer at a first time point, and determining a first state for each of the plurality of subjects in order to generate a first set of states based on the information at the first time point, wherein the information includes, for each of the plurality of subjects, a tumor gene profile obtained by genotyping at least 50 genes and any treatment provided to the subject prior to the first time point. (b) The steps of obtaining the information relating to the plurality of objects at one or more second time points after the first time point, and determining the second state of each of the plurality of objects at each of the one or more second time points in order to generate a set of subsequent states based on the information at one of the one or more second time points, (c) A step of using the set of first states from (a) and the set of subsequent states from (b), wherein a prediction algorithm is generated that determines the probability that a given first state brings about a second state in a set of states at a later time point after the given first state. Methods that include... (Item 4) (d) With respect to the given first state in the set of states at an earlier time, the step of determining the probability that the given first state brings about the second state in the set of states at a later time, (e)(d) A step of generating an electronic output that represents the probability determined in (e)(d) The method described in item 3, further including the method described in item 3. (Item 5) The method according to item 1 or 3, wherein the step of obtaining the aforementioned information includes sequencing cell-free deoxyribonucleic acid (cfDNA) derived from the multiple subjects and optionally conducting a medical interview with each of the multiple subjects. (Item 6) The method according to item 1 or 3, wherein the treatment was provided to the subject before the first time point. (Item 7) The method according to item 1 or 3, further comprising the step of generating one or more decision trees, each decision tree comprising a root node, one or more decision branches, one or more decision nodes, and one or more terminal nodes, wherein the state at the root node represents the first time point, the one or more decision branches represent alternative actions, and the one or more decision nodes and the one or more terminal nodes represent subsequent states. (Item 8) The method according to item 7, wherein the one or more decision branches include multiple decision branches. (Item 9) The method according to item 1 or 3, wherein the subsequent state includes a survival state of the object indicating whether the object is alive or dead. (Item 10) The method described in item 1 or 3, wherein the subsequent condition includes the target survival rate. (Item 11) The method according to item 1 or 3, wherein each of the first states comprises a common set of one or more somatic mutations. (Item 12) The method described in item 1 or 3, wherein the information further includes the target profile. (Item 13) The method according to item 1 or 3, wherein the probability is at least partially a function of treatment selection from among multiple treatment options. (Item 14) The method according to any one of items 1 to 4, wherein the one or more second time points include a plurality of subsequent time points. (Item 15) The method according to item 14, further comprising the step of determining the probability at multiple subsequent time points. (Item 16) The method according to item 15, wherein the aforementioned time period includes at least three or at least four time periods. (Item 17) The method according to item 1, wherein the first time point is before the subject receives the treatment, and the subsequent time point is after the subject receives the treatment. (Item 18) The method according to item 13, wherein the second treatment is administered after the subsequent time point, based on the subsequent condition at the subsequent time point. (Item 19) The method according to item 1, wherein the information relating to the plurality of subjects includes one or more features from the patient profiles of the subjects, wherein the features are selected from a group consisting of age, sex, gender, genetic profile, enzyme levels, organ function, quality of life, frequency of medical interventions, remission status, and patient outcomes. (Item 20) The method according to item 19, wherein the gene profile includes a target genotype at one or more loci that increases the risk of cancer, affects pharmacokinetics, or affects drug sensitivity. (Item 21) The method according to item 1, wherein the information relating to the plurality of subjects includes one or more features from the tumor profile of the subjects, the features being selected from the group consisting of one or more gene variants, tissue of origin, tumor volume, tumor drug sensitivity, and tumor stage. (Item 22) The method according to item 21, wherein one or more of the above characteristics are determined by assaying cell-free nucleic acid molecules derived from the subject. (Item 23) The method according to item 22, comprising quantifying one or more gene variants and determining the ratio of cell-free nucleic acid molecules containing the one or more somatic mutations. (Item 24) The method according to item 23, further comprising the step of determining whether the ratio of the one or more somatic mutations has increased or decreased between the first time point and the one or more subsequent time points. (Item 25) The method according to item 23, further comprising the step of determining whether the ratio of the one or more somatic mutations is increasing or decreasing between one or more subsequent time points. (Item 26) The method according to item 24 or 25, wherein the ratio of the one or more somatic mutations is increasing. (Item 27) The method according to item 26, wherein the one or more somatic mutations are increased, and further, the somatic mutations are associated with resistance to the treatment. (Item 28) The assay described above is performed according to the method of item 22, which includes high-throughput sequencing. (Item 29) (a) A step of obtaining information about a subject having cancer at a first point in time, wherein the information includes at least one characteristic of the subject from a patient profile, tumor profile, or treatment, (b) A step of determining the initial state of the object based on the information at the first time point, (c) A step of determining the probability of each of a plurality of subsequent states at each of one or more subsequent time points based on the initial state of the subject, thereby providing a set of probabilities relating to the outcome of the state, (d) A step of generating a recommendation for the cancer treatment, which optimizes the probability that the subject will obtain a particular outcome, based at least in part on the set of probabilities relating to the outcome of the condition, (e)(d) A step of generating the electronic output indicating the recommendation generated in (e)(d) Methods that include... (Item 30) The method according to item 29, wherein the probability is at least partially a function of the choice of treatment from among multiple treatment options. (Item 31) The method according to item 29 or 30, wherein the one or more subsequent time points include multiple subsequent time points. (Item 32) The method according to item 31, further comprising the step of determining the probability at multiple subsequent time points. (Item 33) The method described in item 29, wherein the aforementioned time period includes at least three time periods. (Item 34) The method described in item 29, wherein the aforementioned time period includes at least four time periods. (Item 35) The method according to item 29, wherein the first time point is before the subject receives the treatment, and the subsequent time point is after the subject receives the treatment. (Item 36) The method according to item 35, wherein the second treatment is administered after the subsequent time point, based on the subsequent condition at the subsequent time point. (Item 37) The method according to item 29, wherein the at least one feature of the subject is from the patient profile and is selected from a group consisting of age, gender, genetic profile, enzyme levels, organ function, quality of life, frequency of medical interventions, remission status, and patient outcomes. (Item 38) The method according to item 29, wherein the gene profile includes the target genotype at one or more loci that are hereditary oncogenes. (Item 39) The method according to item 29, wherein the gene profile includes the target genotype at one or more loci that affect pharmacokinetics. (Item 40) The method according to item 29, wherein the gene profile includes the target genotype at one or more loci affecting drug sensitivity. (Item 41) The method according to item 29, wherein the at least one feature of the subject is from the tumor profile and is selected from the group consisting of one or more somatic mutations, histological origin, tumor volume, tumor drug sensitivity, and tumor stage. (Item 42) The method according to item 40, wherein at least one of the features is determined by assaying a cell-free nucleic acid molecule derived from the subject. (Item 43) The method according to item 42, comprising quantifying the somatic mutations and determining the ratio of cell-free nucleic acid molecules derived from the tumor containing one or more somatic mutations. (Item 44) The method according to item 43, further comprising the step of determining whether the ratio of the one or more somatic mutations has increased or decreased between the first time point and the one or more subsequent time points. (Item 45) The method according to item 43, further comprising the step of determining whether the ratio of the one or more somatic mutations is increasing or decreasing between one or more subsequent time points. (Item 46) The assay described above is performed according to the method of item 42, which includes high-throughput sequencing. (Item 47) The method described in item 29, wherein the tumor profile is not derived from a tumor tissue biopsy. (Item 48) (a) Obtaining information relating to the subject, including at least the tumor's genetic profile and any treatments previously or currently provided to the subject, and determining the subject's initial state based on the information; (b) A step of providing a decision tree, wherein the root node represents an initial target state, the decision branches represent alternative actions available for the target, the opportunity nodes represent points of uncertainty, and the decision node or terminal node represents a subsequent state. (c) A step of providing a process for treating the subject that maximizes the probability that the subject is alive at the terminal node, (d) A step of generating an electronic output indicating the treatment process determined in (c) and Methods that include... (Item 49) (a) The step of establishing one or more communication links with one or more healthcare service providers on a communication network, (b) The step of receiving medical information relating to one or more subjects from one or more medical service providers on the communication network, (c) The step of receiving one or more samples from the medical service provider, each of which contains cell-free deoxyribonucleic acid (cfDNA) derived from one or more of the subjects, (d) The steps of sequencing the cfDNA and identifying one or more gene variants present in the cfDNA, (e) A step of creating or supplementing a database having information about each of the one or more subjects, wherein the information includes both identified gene variants and received medical information, (f) A step of using the database and an algorithm implemented by a computer, which involves generating at least one predictive model that predicts the probability of a subsequent state for each of several different therapeutic interventions, based on the initial state of the subject. Methods that include... (Item 50) When executed by one or more computer processors, (a) A step of obtaining information about a plurality of subjects having cancer at a first time point, and determining a first state for each of the plurality of subjects in order to generate a first set of states based on the information at the first time point, wherein the information includes, for each of the plurality of subjects, at least a tumor gene profile obtained by performing nucleic acid genotyping from cell-free fluid and any treatment provided to the subject prior to the first time point. (b) The steps of obtaining the information relating to the plurality of objects at one or more second time points after the first time point, and determining the second state of each of the plurality of objects at each of the one or more second time points in order to generate a set of subsequent states based on the information at one of the one or more second time points, (c) A step of using the set of first states from (a) and the set of subsequent states from (b), wherein a prediction algorithm is generated that determines the probability that a given first state brings about a second state in a set of states at a later time point after the given first state. A non-transient, computer-readable medium containing machine-executable code that implements a method including [a specific method]. (Item 51) When executed by one or more computer processors, (a) A step of obtaining information about a plurality of subjects having cancer at a first time point, and determining a first state for each of the plurality of subjects in order to generate a first set of states based on the information at the first time point, wherein the information includes, for each of the plurality of subjects, a tumor gene profile obtained by genotyping at least 50 genes and any treatment provided to the subject prior to the first time point. (b) At one or more second time points after the first time point, the steps of obtaining the information relating to the plurality of objects and determining the second state of each of the plurality of objects at each of the one or more second time points based on the information at one of the one or more second time points, in order to generate a set of subsequent states, (c) A step of using the set of first states from (a) and the set of subsequent states from (b), wherein a prediction algorithm is generated that determines the probability that a given first state brings about a second state in a set of states at a later time point after the given first state. A non-transient, computer-readable medium containing machine-executable code that implements a method including [a specific method]. (Item 52) (a) Obtaining information relating to the subject, including at least the tumor's genetic profile and any treatments previously or currently provided to the subject, and determining the subject's initial state based on the information; (b) A step of providing a decision tree, wherein the root node represents an initial target state, the decision branches represent alternative actions available for the target, the opportunity nodes represent points of uncertainty, and the decision node or terminal node represents a subsequent state. (c) A step of providing a process for treating the subject that maximizes the probability that the subject is alive at the terminal node, (d) A step of administering the treatment process to the subject Methods that include... (Item 53) (e) At a second point in time after the initial state, obtain information relating to the subject, including at least the tumor's genetic profile and any treatments previously or currently provided to the subject, and determine the second state of the subject from among a plurality of subsequent states based on the information; (f) A step of providing a subsequent treatment process for the object that maximizes the probability of the object being in a state of survival at the terminal node based on the second state, (g) A step of administering the subsequent treatment process to the subject. The method described in item 52, further including the method described in item 52. (Item 54) A method comprising the step of providing a treatment process from among a plurality of alternative treatments for a subject having cancer, wherein the subject is characterized by a decision tree comprising a plurality of decision branches, each decision branch representing one alternative treatment from among the plurality of alternative treatments, and the treatment process maximizes the probability that the subject achieves a state of survival at the terminal node.
Claims
[Claim 1] The invention as shown in the drawings.