Detection and diagnosis of cancer evolution

A predictive algorithm using tumor genetic profiles and patient information addresses treatment resistance in cancer by optimizing treatment strategies based on genetic variants, enhancing treatment efficacy and survival rates.

JP7805394B2Active Publication Date: 2026-01-23GUARDANT HEALTH INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024075758
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-02-02
Filing Date
2024-05-08
Publication Date
2026-01-23
Estimated Expiration
2037-02-02

AI Technical Summary

Technical Problem

Current cancer treatments face challenges due to the development of resistance, which is often caused by genetic variants in tumors, leading to limited effectiveness and high mortality rates, especially in metastasized cancers.

Method used

A computer-implemented method for predicting patient response to cancer treatments and resistance development by analyzing tumor genetic profiles and patient information over time, using genotyping of cell-free DNA and generating predictive algorithms to optimize treatment strategies.

Benefits of technology

Enables accurate prediction of treatment outcomes and resistance development, allowing for personalized treatment plans that maximize survival chances and minimize resistance, thereby improving cancer management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007805394000001
    Figure 0007805394000001
  • Figure 0007805394000002
    Figure 0007805394000002
  • Figure 0007805394000003
    Figure 0007805394000003
Patent Text Reader

Abstract

To provide cancer evolution detection and diagnostic.SOLUTION: The present disclosure provides a method for determining a probability that after any of a number of therapeutic interventions, an initial state of a subject, such as somatic cell mutational status of a subject with cancer, will develop a subsequent state. Such probabilities can be used to inform a health care provider as to particular courses of treatment to maximize probability of a desired outcome for the subject. The present disclosure also provides a method and system for detecting or monitoring cancer evolution. Such a method and system can be used to predict the generation of the patient's response to and tolerance for treatment for cancer, and for other advantages.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application claims priority to U.S. Provisional Patent Application No. 62 / 290,375, filed February 2, 2016, which is incorporated herein by reference in its entirety. [Background technology]

[0002] background Cancer is a disease that imposes a significant burden worldwide. Tens of millions of individuals worldwide are diagnosed with cancer each year, and more than half of these individuals will not be able to effectively treat their cancer, which will ultimately lead to death. In many countries, cancer ranks as the second leading cause of death, after cardiovascular disease.

[0003] Drugs that target genetic vulnerabilities in human tumors are currently being clinically tested as effective cancer treatments.However, the development of resistance to such treatments can significantly limit their usefulness, and remains a substantial challenge for the clinical management of advanced cancer.Resistance to anticancer drug treatment can be caused by a variety of factors, including individual variability in subjects and the emergence and proliferation of genetic variants in tumors.The most common cause of resistance to a wide range of anticancer drugs is the expression of one or more energy-dependent transporters that detect and excrete anticancer drugs from cells, while other resistance mechanisms can include insensitivity to drug-induced apoptosis and the induction of drug detoxification mechanisms.

[0004] The development of resistance to chemotherapy drugs is frequent and often fatal for cancer patients with solid tumors that have metastasized or spread throughout the body, such as those of the breast, prostate, lung, and colon. In some cases, specific mutational mechanisms directly contribute to the acquisition of drug resistance, while in other cases, non-mutational, potentially epigenetic mechanisms appear to play a critical role.

[0005] The gold standard for mechanistic characterization of tumor drug resistance involves detailed studies of tumor tissue obtained before treatment and after recurrence, together with experimental confirmation of candidate resistance effectors. Summary of the Invention [Means for solving the problem]

[0006] Abstract As recognized herein, there is a significant need for alternative tools to predict patient response and the emergence of resistance to cancer treatments.

[0007] The present disclosure provides methods and systems for detecting or monitoring the evolution of cancer, which can be used to predict patient response to cancer treatments and the development of resistance, as well as for other benefits.

[0008] In one aspect, the disclosure provides a computer-implemented method comprising: (a) obtaining information about a plurality of subjects having cancer at a first time point, and determining a first state of each of the plurality of subjects based on the information at the first time point to generate a set of first states, the information including, for each of the plurality of subjects, at least a tumor genetic profile obtained by genotyping nucleic acids from acellular body fluid and any treatments provided to the subject prior to the first time point; (b) obtaining information about the plurality of subjects at one or more second time points later than the first time point, and determining a second state of each of the plurality of subjects at each of the one or more second time points to generate a set of subsequent states based on the information at a given one of the one or more second time points; and (c) using the set of first states from (a) and the set of subsequent states from (b) to generate a predictive algorithm configured to determine a probability that a given first state will result in a second state among the set of states at a later time point later than the given first state. In some embodiments, the method further includes (d) determining, for a given first state among the set of states at the earlier time, a probability that the given first state results in a second state among the set of states at the later time; and (e) generating an electronic output indicative of the probability determined in (d).

[0009] In one aspect, the present disclosure provides a computer-implemented method comprising: (a) obtaining information about a plurality of subjects having cancer at a first time point and determining a first state of each of the plurality of subjects based on the information at the first time point to generate a set of first states, the information including, for each of the plurality of subjects, at least, a tumor genetic profile obtained by genotyping at least 50 genes and any treatments provided to the subject prior to the first time point; (b) obtaining information about the plurality of subjects at one or more second time points later than the first time point and determining a second state of each of the plurality of subjects at each of the one or more second time points to generate a set of subsequent states based on the information at a given one of the one or more second time points; and (c) using the set of first states from (a) and the set of subsequent states from (b) to generate a predictive algorithm configured to determine a probability that a given first state will result in a second state among the set of states at a later time point later than the given first state. In some embodiments, the method further includes (d) determining, for a given first state among the set of states at the earlier time, a probability that the given first state results in a second state among the set of states at the later time; and (e) generating an electronic output indicative of the probability determined in (d).

[0010] In some embodiments, the step of obtaining information includes sequencing cell-free deoxyribonucleic acid (cfDNA) from multiple subjects, and optionally conducting a medical interview for each of the multiple subjects. In some embodiments, the treatment was provided to the subjects before the first time point. In some embodiments, the method includes generating one or more decision trees, each decision tree including a root node, one or more decision branches, one or more decision nodes, and one or more terminal nodes, where the state at the root node represents the first time point, the one or more decision branches represent alternative treatments, and the one or more decision nodes and the one or more terminal nodes represent subsequent states. In some embodiments, the one or more decision branches include multiple decision branches. In some embodiments, the subsequent state includes a subject survival state indicating whether the subject is alive or dead. In some embodiments, the subsequent state includes a subject survival rate. In some embodiments, each of the first states includes a common set of one or more somatic mutations. In some embodiments, the information further includes a subject profile.

[0011] In some embodiments, the probability is at least partially a function of treatment selection from among multiple treatment options. In some embodiments, the one or more second time points include multiple subsequent time points. In some embodiments, the method further includes determining the probability at multiple subsequent time points. In some embodiments, the time points include at least three time points or at least four time points. In some embodiments, the first time point is before the subject receives the treatment, and the subsequent time point is after the subject receives the treatment. In some embodiments, a second treatment is administered after the subsequent time point based on the subsequent state at the subsequent time point.

[0012] In some embodiments, the information about the plurality of subjects includes one or more features from the subject's patient profile, the features being selected from the group consisting of age, sex, gender, genetic profile, enzyme levels, organ function, quality of life, frequency of medical interventions, remission status, and patient outcome. In some embodiments, the genetic profile includes the subject's genotype at one or more loci that increase cancer risk, affect pharmacokinetics, or affect drug sensitivity. In some embodiments, the information about the plurality of subjects includes one or more features from the subject's tumor profile, the features being selected from the group consisting of one or more genetic variants, tissue of origin, tumor burden, tumor drug sensitivity, and tumor stage. In some embodiments, the one or more features are determined by assaying cell-free nucleic acid molecules from the subject. In some embodiments, one or more genetic variants are quantified to determine the proportion of cell-free nucleic acid molecules containing one or more somatic mutations. In some embodiments, the method further includes determining whether the proportion of one or more somatic mutations is increased or decreased between the first time point and one or more subsequent time points. In some embodiments, the method further comprises determining whether the rate of one or more somatic mutations increases or decreases between a plurality of the one or more subsequent time points. In some embodiments, the rate of one or more somatic mutations increases. In some embodiments, the one or more somatic mutations increase, and further, the somatic mutations are associated with resistance to treatment. In some embodiments, the assaying comprises high-throughput sequencing.

[0013] In another aspect, the present disclosure provides a method comprising: (a) acquiring information about a subject having cancer at a first time point, the information including at least one characteristic of the subject from a patient profile, a tumor profile, or a treatment; (b) determining an initial state of the subject based on the information at the first time point; (c) determining a probability of each of a plurality of subsequent states at each of one or more subsequent time points based on the subject's initial state, thereby providing a set of probabilities for state outcomes; (d) generating a cancer treatment recommendation that optimizes the probability of the subject achieving a particular outcome based at least in part on the set of probabilities for state outcomes; and (e) generating an electronic output indicative of the recommendation generated in (d). In some embodiments, the probability is at least in part a function of treatment selection from a plurality of treatment options. In some embodiments, the one or more subsequent time points include multiple subsequent time points. In some embodiments, the method further comprises determining the probabilities at multiple subsequent time points. In some embodiments, the time points include at least three time points. In some embodiments, the time points include at least four time points. In some embodiments, the first time point is before the subject receives treatment, and the subsequent time point is after the subject receives treatment.In some embodiments, the second treatment is administered after the subsequent time point based on the subsequent state at the subsequent time point.In some embodiments, at least one characteristic of the subject is from patient profile, and is selected from the group consisting of age, gender, genetic profile, enzyme level, organ function, quality of life, frequency of medical intervention, remission status and patient outcome.

[0014] In some embodiments, the genetic profile comprises the subject's genotype at one or more loci that are hereditary cancer genes. In some embodiments, the genetic profile comprises the subject's genotype at one or more loci that affect pharmacokinetics. In some embodiments, the genetic profile comprises the subject's genotype at one or more loci that affect drug sensitivity. In some embodiments, the subject's at least one characteristic is from a tumor profile and is selected from the group consisting of one or more somatic mutations, tissue of origin, tumor burden, tumor drug sensitivity, and tumor stage. In some embodiments, the at least one characteristic is determined by assaying cell-free nucleic acid molecules from the subject.

[0015] In some embodiments, somatic mutations are quantified to determine the proportion of cell-free nucleic acid molecules derived from the tumor that contain one or more somatic mutations.

[0016] In some embodiments, the method further comprises determining whether the ratio of one or more somatic mutations increases or decreases between the first time point and one or more subsequent time points. In some embodiments, the method further comprises determining whether the ratio of one or more somatic mutations increases or decreases between a plurality of the one or more subsequent time points. In some embodiments, the assaying comprises high-throughput sequencing. In some embodiments, the tumor profile is not derived from tumor tissue biopsy.

[0017] In one aspect, the present disclosure provides a method including: (a) obtaining information about a subject, including at least a genetic profile of the tumor and any treatments previously or currently provided to the subject, and determining an initial state of the subject based on the information; (b) providing a decision tree, wherein a root node represents the initial subject state, decision branches represent alternative treatments available to the subject, opportunity nodes represent points of uncertainty, and decision or terminal nodes represent subsequent states; (c) providing a course of treatment for the subject that maximizes the probability that the subject will achieve a survival state at the terminal node; and (d) generating an electronic output indicative of the course of treatment determined in (c).

[0018] In one aspect, the present disclosure provides a method comprising: (a) establishing one or more communication links with one or more healthcare providers over a communication network; (b) receiving medical information regarding one or more subjects from the one or more healthcare providers over the communication network; (c) receiving from the healthcare provider one or more samples comprising cell-free deoxyribonucleic acid (cfDNA) from each of the one or more subjects; (d) sequencing the cfDNA to identify one or more genetic variants present in the cfDNA; (e) creating or providing a database having information about each of the one or more subjects, the information including both the identified genetic variants and the received medical information; and (f) using the database and a computer-implemented algorithm to generate at least one predictive model that predicts a probability of a subsequent state for each of a plurality of different therapeutic interventions based on the subject's initial state.

[0019] In one aspect, the disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements a method comprising: (a) acquiring information about a plurality of subjects having cancer at a first time point and determining a first state of each of the plurality of subjects based on the information at the first time point, to generate a set of first states, wherein the information includes, for each of the plurality of subjects, at least a tumor genetic profile obtained by genotyping nucleic acids from acellular body fluid and any treatments provided to the subject prior to the first time point; (b) acquiring information about the plurality of subjects at one or more second time points later than the first time point and determining a second state of each of the plurality of subjects at each of the one or more second time points to generate a set of subsequent states based on the information at a given one of the one or more second time points; and (c) using the set of first states from (a) and the set of subsequent states from (b) to generate a predictive algorithm configured to determine a probability that a given first state will result in a second state among the set of states at a later time point later than the given first state.

[0020] In one aspect, the disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements a method comprising: (a) obtaining information about a plurality of subjects having cancer at a first time point and determining a first state of each of the plurality of subjects based on the information at the first time point to generate a set of first states, wherein the information includes, for each of the plurality of subjects, at least, a tumor genetic profile obtained by genotyping at least 50 genes and any treatments provided to the subject prior to the first time point; (b) obtaining information about the plurality of subjects at one or more second time points later than the first time point and determining a second state of each of the plurality of subjects at each of the one or more second time points to generate a set of subsequent states based on the information at a given one of the one or more second time points; and (c) using the set of first states from (a) and the set of subsequent states from (b) to generate a predictive algorithm configured to determine a probability that a given first state will result in a second state among the set of states at a later time point later than the given first state.

[0021] In one aspect, the present disclosure provides a method including: (a) obtaining information about a subject, including at least a genetic profile of the tumor and any treatments previously or currently provided to the subject, and determining an initial state of the subject based on the information; (b) providing a decision tree, wherein a root node represents the initial subject state, decision branches represent alternative treatments available to the subject, opportunity nodes represent points of uncertainty, and decision or terminal nodes represent subsequent states; (c) providing a course of treatment for the subject that maximizes the probability that the subject will achieve a survival state at the terminal node; and (d) administering the course of treatment to the subject. In some embodiments, the method further includes: (e) acquiring information about the subject at a second time point, later than the initial state, including at least a genetic profile of the tumor and any treatments previously or currently being provided to the subject, and determining a second state of the subject from among a plurality of subsequent states based on the information; (f) providing a subsequent course of treatment for the subject based on the second state, the subsequent course of treatment maximizing the probability that the subject will achieve a surviving state at the terminal node; and (g) administering the subsequent course of treatment to the subject. In some embodiments, the method further includes: (e) acquiring information about the subject at a second time point, later than the initial state, including at least a genetic profile of the tumor and any treatments previously or currently being provided to the subject, and determining a second state of the subject from among a plurality of subsequent states based on the information; (f) providing a subsequent course of treatment for the subject based on the second state, the subsequent course of treatment maximizing the probability that the subject will achieve a surviving state at the terminal node; and (g) administering the subsequent course of treatment to the subject.

[0022] In one aspect, the present disclosure provides a method comprising providing a course of treatment from among a plurality of alternative treatments for a subject having cancer, wherein the subject is characterized by a decision tree comprising a plurality of decision branches, each decision branch representing an alternative treatment from among the plurality of alternative treatments, and the course of treatment maximizes the probability that the subject will achieve a survival state at a terminal node.

[0023] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of the present disclosure have been shown and described. As will be understood, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. Incorporation by Reference

[0024] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.

[0025] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description, which sets forth illustrative embodiments, and in which the principles of the invention are utilized, and the accompanying drawings. [Brief explanation of the drawings]

[0026] [Figure 1] FIG. 1 shows an exemplary method for analyzing mutations in various disease states in a subject.

[0027] [Figure 2A] FIG. 2A shows various exemplary abnormalities in cancer genomes.

[0028] [Figure 2B] FIG. 2B illustrates an exemplary system for detecting evolutionary escape paths.

[0029] [Figure 2C] FIG. 2C shows an exemplary model generated by the system of FIG. 2B.

[0030] [Figure 2D] FIG. 2D shows exemplary heterogeneous populations of normal cells and cancer subclones that develop during the evolutionary history of a tumor.

[0031] [Figure 3] FIG. 3 shows an exemplary process for reducing error rates and bias in reading sequences of deoxyribonucleic acid (DNA).

[0032] [Figure 4] FIG. 4 shows a schematic diagram of internet-enabled access to reports of subjects with cancer.

[0033] [Figure 5] Figure 5 shows multiple genes associated with genetic variants.

[0034] [Figure 6] Figure 6 shows a decision tree that includes a root node (rectangle) indicating the initial state, decision branches (arrows) indicating different therapeutic interventions, and opportunity nodes (circles) from which opportunity branches (arrows) arise either at terminal nodes (triangles) or decision nodes (squares) indicating subsequent states.

[0035] [Figure 7] FIG. 7 illustrates a computer system that is programmed or otherwise configured to implement the methods provided herein. DETAILED DESCRIPTION OF THE INVENTION

[0036] Detailed Description Genetic variants are alternative forms at gene loci. In the human genome, approximately 0.1% of nucleotide positions are polymorphic, i.e., exist in a second genetic form that occurs in at least 1% of the population. Mutations can introduce genetic variants into germline cells and into diseased cells such as cancer. Reference sequences such as hg19 or NCBI Build 37 or Build 38 are intended to represent "wild-type" or "normal" genomes. However, they do not identify common polymorphisms, which can also be considered normal in that they have a single sequence.

[0037] Genetic variants include sequence variants, copy number variants, and nucleotide modification variants.Sequence variants are changes in the nucleotide sequence of genes.Copy number variants are the copy number of a part of genome that is deviated from wild type.Genetic variants include, for example, single nucleotide polymorphisms (SNPs), insertions, deletions, inversions, transversions, translocations, gene fusions, chromosome fusions, gene truncations, copy number variations (e.g., aneuploidy, partial aneuploidy, polyploidy, gene amplification), abnormal changes in nucleic acid chemical modification, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid methylation.

[0038] The term "polynucleotide," as used herein, generally refers to a molecule comprising one or more nucleic acid subunits. A polynucleotide may include one or more subunits selected from adenosine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), or variants thereof. Nucleotides may include A, C, G, T, or U, or variants thereof. Nucleotides may include any subunit that can be incorporated into a growing nucleic acid chain. Such subunits may be A, C, G, T, or U, or may be specific to one or more complementary A, C, G, T, or U, or any other subunit that is complementary to a purine (i.e., A or G, or variants thereof) or pyrimidine (i.e., C, T, or U, or variants thereof). The subunits may allow individual nucleobases or groups of bases (e.g., AA, TA, AT, GC, CG, CT, TC, GT, TG, AC, CA, or their uracil counterparts) to be resolved. In some examples, the polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), or derivatives thereof. The polynucleotide may be single-stranded or double-stranded.

[0039] The term "subject," as used herein, generally refers to an animal, such as a mammal (e.g., a human) or an avian (e.g., a bird), or other organism, such as a plant. More specifically, a subject may be a vertebrate, a mammal, a mouse, a primate, a monkey, or a human. Animals include, but are not limited to, livestock animals, sport animals, and pets. A subject may be a healthy individual, an individual having or suspected of having a disease or a propensity for disease, or an individual in need of treatment or suspected of needing treatment. A subject may be a patient.

[0040] The term "genome" generally refers to the entire genetic information of an organism. A genome can be coded by either DNA or RNA. A genome can include coding regions that code for proteins as well as non-coding regions. A genome can include the sequences of all chromosomes in an organism together. For example, the human genome has a total of 46 chromosomes. All of these sequences together make up the human genome. A "reference genome" typically refers to a haploid genome. Examples of reference genomes include hg19, or NCBI Build 37 or Build 38.

[0041] The terms "adaptor," "adapter," and "tag" are used interchangeably throughout this specification. An adapter or tag can be linked to a polynucleotide sequence to be "tagged" by any approach, including ligation, hybridization, or other approaches.

[0042] The term "library adaptor" or "library adapter," as used herein, generally refers to a molecule (e.g., polynucleotide) whose identity (e.g., sequence) can be used to distinguish polynucleotides in a biological sample (also referred to herein as a "sample").

[0043] The term "sequencing adapter," as used herein, generally refers to a molecule (e.g., a polynucleotide) adapted to enable a sequencing instrument to sequence a target polynucleotide, such as by interacting with the target polynucleotide to enable sequencing. The sequencing adapter allows the target polynucleotide to be sequenced by the sequencing instrument. In one example, the sequencing adapter comprises a nucleotide sequence that hybridizes or binds to a capture polynucleotide attached to a solid support, e.g., a flow cell, of a sequencing system. In another example, the sequencing adapter comprises a nucleotide sequence that hybridizes or binds to a polynucleotide to generate a hairpin loop, which allows the target polynucleotide to be sequenced by the sequencing system. The sequencing adapter may comprise a sequencer motif, which may be a nucleotide sequence that is complementary to the flow cell sequence of another molecule (e.g., a polynucleotide) and can be used by the sequencing system to sequence the target polynucleotide. The sequencer motif may also include a primer sequence for use in sequencing, such as sequencing by synthesis (SBS). The sequencer motif may include a sequence required to link the library adapter to a sequencing system and perform sequencing of the target polynucleotide.

[0044] As used herein, the terms "at least," "at most," or "about," when preceding a series, refer to every member of the series, unless otherwise indicated.

[0045] The term "about" and its grammatical equivalents in connection with a referenced numerical value can include values ​​in a range of up to ±10% from that value. For example, the quantity "about 10" can include amounts from 9 to 11. In other embodiments, the term "about" in connection with a referenced numerical value can include values ​​in a range of ±10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of that value.

[0046] Generally, disclosed herein is a method for generating a predictive model of tumor evolution over time in response to various treatments, and using this model to select a treatment for a subject (e.g., patient).The predictive model is based on at least the genetic profile of tumor, and optionally patient profile and / or treatment.Results can be disclosed to patients or healthcare providers to improve care.

[0047] In some cases, the information includes a genetic profile from the tumor obtained by genotyping acellular body fluid (e.g., cfDNA). In some cases, the information further includes the treatment and / or therapeutic intervention provided to the subject. In some cases, the information further includes a profile of the subject.

[0048] The information can be used to determine a condition associated with a subject. The condition can include information relevant to predicting the subsequent condition of a subject. For example, the condition can indicate whether the subject is alive or dead. The condition can indicate the median life expectancy of the subject. The condition can indicate a medically relevant somatic mutation in a tumor (e.g., a KRAS variant). The condition can indicate drug resistance (e.g., cetuximab resistance).

[0049] The information may be used to generate one or more decision trees that indicate the probability of various endpoints of a target exhibiting a particular state. Decision branches may arise from a root node (which may be considered the first decision node). Decision branches may lead to either endpoints (also called terminal nodes) or opportunity nodes. Terminal nodes or endpoints may represent states. Opportunity nodes (or event nodes) may be points of uncertainty from which different outcomes may occur. Uncertainty may be resolved through opportunity branches (event branches) arising from opportunity nodes. Each opportunity branch may lead to either a terminal node or a decision node (which itself may represent a state) from which multiple decision branches arise. These decision branches may, in turn, lead to endpoints or opportunity nodes in a sequential manner until all branches lead to an endpoint or terminal node.

[0050] The root node in a decision tree may be an initial state. The initial state may be as broad as "cancer diagnosis." More typically, the root node represents some aspect of a subject's genetic profile. For example, the root node may represent one or more genetic variants detected in cfDNA, such as the presence of mutations in specific cancer genes and / or their abundance compared to normal DNA. Each decision branch from the root node may represent a different course of treatment (or lack of treatment). For example, a course of treatment may represent a different chemotherapy or immunotherapy regimen, a type of surgery, or radiation therapy. A terminal node may represent a status, such as survival or death, within a certain time period from diagnosis (e.g., 5-year survival rate). A decision node represents a new state from which a new decision can be made. For example, a decision node may be the occurrence of a genetic variant that confers chemotherapy resistance. Such a variant may represent an escape pathway, whereby a tumor avoids responding to chemotherapy and a different therapeutic approach may be required.

[0051] Advantageously, the method disclosed herein can generate a predictive algorithm configured to determine the probability that a certain therapeutic intervention applied to a particular condition (e.g., a specific chemotherapy agent for a cancer with a particular genetic profile) will result in a particular state (e.g., a genetic variant) in which the cancer can evade therapeutic intervention. Such probability can be determined through multiple treatments and evasion. As a result, it can be determined that a particular series of therapeutic interventions will result in a particular evasion mode, ultimate evasion (e.g., death), or a state in which the cancer cannot be detected with a given frequency or probability.

[0052] The present disclosure provides a method for generating a prediction algorithm for assigning a probability to each branch or each terminal node in a decision tree. The method may utilize a database in which the outcome at each branch can be calculated from a plurality of subjects whose data are stored. The probability can be determined, for example, by obtaining a training set of subjects, classifying them into conditions, recording treatments and / or therapeutic interventions, and then determining the frequency of outcomes (e.g., final conditions). The frequency of a given outcome in the training set can be used to determine its probability.

[0053] Thus, for multiple subjects presenting with a particular condition, multiple decision branches may be identified and the chances of a particular endpoint or terminal decision node of that branch may be determined. For example, with reference to Figure 6, among individuals presenting with the condition "EGFR mutant," a decision branch may include treatment A and treatment B.

[0054] In Figure 6, Action A leads to opportunity node A, and Action B leads to opportunity node B. Opportunity node A leads to a 75% five-year survival rate at that time (terminal node) and a 25% occurrence of "Avoid A" at that time (decision node A). Avoid A may have one decision branch - Action C, which leads to opportunity node C, from which two opportunity branches arise to the terminal nodes: 40% five-year survival and 60% death. In total, the branches provide an 85% chance of five-year survival and a 15% chance of death.

[0055] In Figure 6, Treatment B is connected to opportunity node B, which is connected to a 60% five-year survival rate at that time (terminal node) and a 40% occurrence of "Avoid B" at that time (decision node B). Avoid B may have one decision branch - Treatment D, which is connected to opportunity node D, from which two opportunity branches arise to terminal nodes: 40% five-year survival and 60% death. In total, the branches provide a 76% chance of five-year survival and a 24% chance of death.

[0056] Adding more data points (objects) at any decision node may increase the reliability of the final probability determined. In some cases, an initial state can be used to predict a subsequent state (e.g., an intermediate state (e.g., at a decision node) or a final state). In some cases, an initial state is classified as leading to a subsequent state (e.g., an intermediate state or a final state) with a given frequency. A subsequent state is a state reached after a decision from a previous state. For example, after state 1, a therapeutic intervention is applied, and the state at a later time is the subsequent state. A subsequent state may be a terminal state from which no further decisions are made, or it may be an intermediate state from which another decision is made.

[0057] The initial state can be determined by clustering subjects based on the information or a subset of information determined about the subjects.Clusters can be generated using information about subjects or a training set of subjects.For example, the information can be categorical (for example, whether KRAS variants are present or absent in tumor samples), and subjects can be clustered based on common categorical values.In some cases, the information about subjects is quantitative.Subjects can be clustered using quantitative data by any method known in the art.Exemplary methods include, but are not limited to, k-means clustering, hierarchical clustering, or centroid-based clustering.Clustering can be based on visual inspection of data, including data that has been projected into a reduced number of dimensions by methods such as principal component analysis.Clustering can be used to create cluster boundaries that define which clusters subjects fall into.

[0058] A profile includes values ​​(quantitative or qualitative) for each of one or more characteristics. A profile may include, for example, information about phenotypic characteristics, genetic characteristics, demographic characteristics, or medical history (including a history of delivered therapeutic interventions). A genetic profile includes values ​​for various genetic characteristics, for example, gene variants at a locus (e.g., sequence information of copy number information). For example, a genetic profile may include germline or somatic genotypes at several loci in diseased (e.g., cancer) cells. A state may be one or more values ​​of the characteristics in the profile.

[0059] The information may include a tumor profile, which includes a genetic profile of the tumor. The information may include a subject profile, which includes genetic information about the subject. The information may include previous treatments or interventions the subject has undergone.

[0060] The tumor profile may include tissue of origin, tumor burden, tumor drug sensitivity, tumor stage, tumor size, tumor metabolic profile, tumor metastatic status, tumor burden, or tumor heterogeneity.

[0061] The tumor profile can include a tumor genetic profile, which can be obtained by various methods. For example, the tumor genetic profile can be obtained by analyzing nucleic acids from a biological sample from a subject using high-throughput sequencing or genotyping arrays. The nucleic acid can be DNA or RNA. The nucleic acid is isolated from the sample. The sample used to obtain the genetic profile can be a tumor biopsy, a fine needle aspiration biopsy, or an acellular body fluid containing nucleic acids from tumor cells. For example, the acellular body fluid can be derived from a body fluid selected from the group consisting of blood, plasma, serum, urine, saliva, mucosal secretions, sputum, feces, cerebrospinal fluid, and tears of the subject.

[0062] For example, blood from a subject at risk of cancer can be collected and prepared as described herein to generate a cell-free polynucleotide population.In one embodiment, this is cell-free DNA (cfDNA).The system and method of the present disclosure can be used to detect mutations or copy number variations that may exist in certain cancers.This method can be useful for detecting the presence of cancerous cells in the body, regardless of the absence of symptoms or other characteristics of disease.

[0063] Methods for nucleic acid extraction and purification are well known in the art. For example, nucleic acids can be purified by organic extraction using phenol, phenol / chloroform / isoamyl alcohol, or similar formulations including TRIzol and TriReagent. Other non-limiting examples of extraction techniques include: (1) organic extraction followed by automated nucleic acid extraction using, for example, an automated nucleic acid extraction device, such as Applied Biosystems (Foster, MA, USA); These include (1) ethanol precipitation using a phenol / chloroform organic reagent, with or without the use of a Model 341 DNA Extractor available from BioSystems, Inc. (City, CA); (2) stationary phase absorption; and (3) salt-induced nucleic acid precipitation, a precipitation method typically referred to as the "salting out" method. Another example of nucleic acid isolation and / or purification uses magnetic particles to which nucleic acids bind, either specifically or nonspecifically, followed by isolating the beads using a magnet, washing the nucleic acids, and eluting them from the beads. In some embodiments, the above-described isolation methods may be preceded by an enzymatic digestion step, such as digestion with proteinase K or other similar proteases, to help eliminate undesired proteins from the sample. If desired, an RNase inhibitor may be added to the lysis buffer. For certain cell or sample types, it may be desirable to add a protein denaturation / digestion step to the protocol. Purification methods can be directed to isolate DNA, RNA, or both. If both DNA and RNA are isolated together during or after the extraction procedure, additional steps can be used to purify one or both separately from the other. Partial fractions of the extracted nucleic acids can also be generated, for example, by purification by size, sequence, or other physical or chemical properties.

[0064] The polynucleotide extracted from sample can be sequenced to generate sequencing reads.Exemplary sequencing techniques can include, for example, emulsion polymerase chain reaction (PCR) (for example, pyrosequencing from Roche 454, semiconductor sequencing from Ion Torrent, SOLiD sequencing by ligation from Life Technologies, sequencing by synthesis from Intelligent Biosystems), bridge amplification on flow cell (for example, Solexa / Illumina), isothermal amplification by Wildfire technology (Life Technologies), or rolony / nanoball generated by rolling circle amplification (Complete Genomics, Intelligent Biosystems, Polonator).Sequencing technology such as Heliscope (Helicos), SMRT technology (Pacific Biosciences) or nanopore sequencing (Oxford Nanopore) can directly sequence single molecules without prior clonal amplification, and can be a suitable sequencing platform. Sequencing can be performed with or without enrichment of target.The exemplary gene and / or region that can be enriched can be found in Figure 5.Enrichment can be performed, for example, by hybridizing nucleic acid sample or sequencing library to the probe that is arranged on an array or attached to beads.In some cases, the polynucleotide from sample is amplified by any suitable approach (for example, PCR) before and / or during sequencing.

[0065] As a non-limiting example, a sample containing initial genetic material can be provided, and cell-free DNA can be extracted. The sample can contain a target nucleic acid at a low abundance. For example, nucleic acids from a normal genome or a germline genome can be predominant in the sample, and the sample also contains nucleic acids from at least one other genome containing no more than 20%, no more than 10%, no more than 5%, no more than 1%, no more than 0.5%, or no more than 0.1% genetic alterations, such as a cancer genome, a fetal genome, or a genome from another individual or species. The initial genetic material can then be converted into a set of tagged parent polynucleotides, and sequenced to obtain sequencing reads. In some cases, these sequence reads can contain barcode information. In other examples, barcodes are not used. Tagging can include adding sequence tags to molecules in the initial genetic material. The sequence tags can be selected so that all unique polynucleotides that map to the same reference sequence have unique identification tags. The sequence tags can be selected so that not all unique polynucleotides mapped to the same reference have a unique identification tag. The conversion can be performed with high efficiency, for example, on at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the initial nucleic acid molecules. The set of tagged parent polynucleotides can be amplified to obtain a set of amplified progeny polynucleotides. The amplification can be, for example, at least 10, 100, 1,000, or 10,000 times. The set of amplified progeny polynucleotides is sampled for sequencing at both a sampling rate such that the resulting sequencing reads (1) cover a target number of unique molecules in the set of tagged parent polynucleotides, and (2) cover the unique molecules in the set of tagged parent polynucleotides at a target coverage factor (e.g., 5-10 times the coverage of the parent polynucleotides). The set of sequencing reads can be aggregated to obtain a set of consensus sequences corresponding to the unique tagged parent polynucleotides. Sequencing reads may be quality checked for inclusion in the analysis.For example, sequencing reads that do not meet a quality control score can be removed from the pool.

[0066] Sequencing reads can be sorted into families, representing the reads of progeny molecules derived from a specific unique parent molecule. For example, a family of amplified progeny polynucleotides can be composed of these amplified molecules derived from a single parent polynucleotide. By comparing the sequences of the progeny within a family, a consensus sequence of the original parent polynucleotide can be estimated. This results in a set of consensus sequences representing the unique parent polynucleotides in the tagged pool. This process can assign a reliability score to the sequence. After sequencing, the reads can be assigned a quality score. The quality score can be a read indication based on a threshold value, indicating whether these reads can be useful in subsequent analysis. In some cases, some reads are not of sufficient quality or length to perform the subsequent mapping step. Sequencing reads with a predetermined quality score (e.g., greater than 90%) can be filtered from the data. Sequencing reads that meet a specified quality score threshold can be mapped to a reference genome or a template sequence known to not contain copy number variations. After mapping alignment, the sequencing reads can be assigned a mapping score. The mapping score can be a representation or read that is mapped back to the reference sequence, indicating whether each position is uniquely mapped or not.In some cases, the read can be a sequence that is not related to copy number variation analysis.For example, some sequencing reads can originate from contaminating polynucleotides.Sequencing reads that have a mapping score indicating that at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% of the sequencing reads are mismapped (e.g., incorrectly mapped) can be filtered out from the data set.In other cases, sequencing reads that are assigned a mapping score lower than a predetermined percentage can be filtered out from the data set.

[0067] Sequencing reads that meet a specified quality score threshold can be mapped to a reference genome or a template sequence known not to contain copy number variation. After mapping alignment, sequencing reads can be assigned a mapping score. In some cases, the reads can be sequences unrelated to copy number variation analysis. After data filtering and mapping, multiple sequencing reads generate chromosomal coverage regions. These chromosomal regions can be divided into windows or bins of variable length. In some cases, each window region can be sized to contain approximately the same number of uniquely mappable bases. In addition, certain windows that are known to be difficult to sequence or contain substantially high GC bias throughout the genome can be filtered from the data set. For example, regions known to fall near the centromere of a chromosome (i.e., centromeric DNA) are known to contain highly repeated sequences that can lead to false positive results. These regions can be filtered. Normalization can be performed to compensate for the impact of GC content on the sequencing reads of a sample. Other regions of the genome can be filtered from the dataset, such as regions containing unusually high concentrations of other highly repetitive sequences, for example, microsatellite DNA.

[0068] For an exemplary genome derived from cell-free polynucleotide sequences, the next step involves determining the read coverage of each window region. This can be performed using either barcoded or non-barcoded reads. In the non-barcoded case, the previous mapping steps may provide coverage of different base positions. Sequencing reads with sufficient mapping and quality scores that fall within the unfiltered chromosome window can be counted. The number of coverage reads can be assigned a score for each mappable location. In the case of barcodes, all sequences with the same barcode, physical property, or a combination of the two can be aggregated into one read because they all originate from the sample parent molecule. This step can reduce biases that may have been introduced during any of the preceding steps, such as steps involving amplification. For example, if one molecule is amplified 10 times while another is amplified 1000 times, each molecule will only appear once after aggregation, thereby eliminating the effects of unequal amplification. Only reads with unique barcodes can be counted at each mappable location, which can affect the assigned score. For this reason, it is important to carry out the barcode ligation step in an optimized manner so as to minimize the amount of bias that occurs.The sequence of each base can be aligned with the most prevalent nucleotide read for its specific position.Furthermore, the number of unique molecules can be counted at each position, leading to simultaneous quantification at each position.This step can reduce the bias that may be introduced during any of the previous steps, such as the step involving amplification.

[0069] The distinct copy number status of each window region can be used to identify the copy number variation in chromosome region.In some cases, all adjacent window regions with the same copy number can be integrated into one segment to report the existence or absence of copy number variation status.In some cases, various windows can be integrated with other segments after filtering.

[0070] Methods for determining genetic profile (for example, genetic profile of tumor or subject) may have error rate.For example, sequencing method may have about 0.1%, about 0.5%, about 1% or higher base-by-base error rate.In some cases, the nucleic acid from tumor cells that contains genetic variants at a given locus is present in small amounts in the total nucleic acid that contains that locus, at a rate similar to or lower than the base-by-base sequencing error rate.In such a situation, it may be difficult to distinguish between genotyping or sequencing errors and genetic variants that exist at low frequencies.In order to reduce error rate, certain techniques can be implemented, such as those described in WO2014 / 149134 (which is incorporated by reference in its entirety).

[0071] The genetic profile of a tumor can include somatic mutations compared to a reference. The reference can be a reference genome, such as a human reference genome. The reference genome can be the subject's germline genome. The genetic profile can include various genetic variants acquired by some or all of the tumor cells. Genetic variants can be, for example, single nucleotide variants, global or small structural variants, or short insertions or deletions. For example, as shown in Figure 2A, common abnormalities in cancer genomes can result in abnormal chromosome numbers (aneuploidy) and chromosomal structures of cancer genomes. In Figure 2A, the upper line represents the genome of the germline genome, and the lower line represents the cancer genome with somatic abnormalities. Double lines are used when it is useful to distinguish between heterozygous and homozygous alterations. Dots represent single nucleotide changes, while lines and arrows represent structural alterations.

[0072] The genetic profile of a tumor can include quantitative information about each variant. For example, by genetic analysis of cell-free DNA by digital sequencing, 1,000 reads can be obtained that map to the locus of a first cancer gene, of which 900 reads correspond to germline sequences, and 100 reads correspond to variants present in tumor cells. The same genetic analysis can obtain 1,000 reads that map to the locus of a second cancer gene, of which 980 reads correspond to germline sequences, and 20 reads correspond to variants that represent 10% tumor burden. It can be inferred that the overall tumor burden is about 10% in cell-free DNA based on the locus of a first cancer gene, but a small proportion of tumor cells (about 20%) may have variants at the locus of a second cancer gene. Such quantitative information can be included in the genetic profile of a tumor and can be monitored over time or in response to treatment.

[0073] The genetic profile of a tumor may include information about somatic variants, including, but not limited to, mutations, indels (insertions or deletions), copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural alterations, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, abnormal changes in nucleic acid methylation, infections, and cancers.

[0074] In some cases, genotyping involves genotyping nucleic acids from acellular body fluids. Such methods can capture genetic information from multiple tumor cells, allowing information about both tumor heterogeneity and tumor evolution to be inferred. In some cases, genotyping can be performed on samples obtained from at least one time point, at least two time points, at least three time points, at least four time points, at least five time points, at least six time points, at least seven time points, at least eight time points, at least nine time points, or at least ten time points. In some cases, genotyping involves determining the genotypes of at least 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 120, 140, 160, 180, or 200 or more loci. In some cases, the loci are genes. In some cases, the loci are oncogenes. Oncogenes are genes containing mutations that drive tumor growth. Exemplary cancer genes can be found in WO2009045443, which is incorporated herein by reference in its entirety. Cancer genes can include those listed in FIG. 5.

[0075] In some cases, the genetic profile of a tumor may contain information about tumor evolution. For example, if the proportion of cell-free DNA from a tumor containing a KRAS mutation is increasing, it can be inferred that the proportion of tumor cells resistant to a particular KRAS-targeting treatment is increasing over time. Figure 1 shows an exemplary method for developing a model of tumor evolution in response to treatment. The process in Figure 1 involves collecting genetic profile data of multiple subjects' tumors, as well as tumor treatments and initial treatments (10). The genetic profiles can be used to identify or infer evolutionary escape pathways taken by tumor cells that cause resistance to the treatment (12). The genetic profiles of individual subjects' tumors can be fitted to the model to obtain the probability that tumor cells will acquire a genetic variant that confers resistance to the treatment (14).

[0076] More complex models can be used to measure tumor heterogeneity, for example, based on the relative occurrence of different variants in cell-free DNA. Figure 2B shows an exemplary system for determining the probability of various state outcomes. This system can be a hidden Markov model (HMM), which is a statistical Markov model that assumes that the modeled system is a Markov process with unobserved (hidden) states. In a simple Markov model (similar to a Markov chain), the states are directly visible to the observer, and therefore the state transition probability is the only parameter. In a hidden Markov model, the states are not directly visible, but the outputs corresponding to the states are. Each state has a probability distribution over the possible output tokens. Therefore, the sequence of tokens generated by the HMM can provide information about the sequence of states. A hidden Markov model can be considered a generalization of a mixture model, in which the hidden variables (or latent variables) that control the mixture components selected for each observation are related through a Markov process rather than being independent of each other. As shown in Figure 2B, an HMM is typically defined by a set of hidden states, a matrix of state transition probabilities, and a matrix of output probabilities. Common methods for constructing such models include, but are not limited to, hidden Markov models (HMMs), artificial neural networks, Bayesian networks, support vector machines, and random forests. Such methods are known to those skilled in the art and are described in Mohri et al., Foundations of Machine Learning (2012), published by MIT Press, which is incorporated herein by reference in its entirety, and in "Machine Learning: A Practical Guide to Machine Learning," published by MIT Press, which is incorporated herein by reference in its entirety. This is described in detail in MacKay, Information Theory, Inference, and Learning Algorithms (2003), published by Cambridge University Press.

[0077] The relative amount of tumor polynucleotides in a cell-free polynucleotide sample is referred to herein as "tumor burden." Tumor burden can be related to tumor size. When tested over time, tumor burden can be used to determine whether a cancer is progressing, stable, or in remission. In some embodiments, confidence intervals for estimated tumor burden do not overlap, indicating the direction of disease progression. Tumor burden and the direction of disease progression can comprise a diagnostic confidence indication. The term "diagnostic confidence index," as used herein, refers to a representation, number, rank, degree, or value assigned to indicate the presence of a genetic variant and how confident that presence is. For example, this representation can be a binary value or an alphanumeric rank from A to Z, among others. In yet another example, the diagnostic confidence index can have any value between 0 and 100, among others. In yet another example, the diagnostic confidence index can be represented by a range or degree, e.g., "low" or "high," "more" or "less," "increased" or "decreased." A low diagnostic confidence index may mean that the presence of a genetic variant is less reliable (the genetic variant may be noise), whereas a high diagnostic confidence index may mean that the genetic variant is more likely to be present; in one embodiment, if the diagnostic confidence index is lower than 25-30 out of 100, the result is considered unreliable.

[0078] In one implementation, measurements from multiple samples taken substantially simultaneously or across multiple time points can be used to adjust the diagnostic confidence index for each variant to indicate its reliability in predicting the observation of a copy number variation (CNV) or mutation. Confidence can be increased by using measurements at multiple time points to determine whether the cancer is progressing, in remission, or stable. Diagnostic confidence indexes can be assigned by any of a number of known statistical methods and can be based, at least in part, on the frequency with which measurements are observed over a period of time. For example, a statistical correlation between current results and previous results can be performed. Alternatively, a hidden Markov model can be constructed to allow a maximum likelihood or maximum a posteriori decision for each diagnosis to be made based on the frequency with which a particular test event occurs across multiple measurements or time points. As part of this model, the probability of error for a particular decision and the resulting diagnostic confidence index can also be output. In this manner, parameter measurements can be provided along with confidence intervals, regardless of whether they are within the noise range. Testing over time can increase the predictive confidence of whether the cancer is progressing, stable, or in remission by comparing confidence intervals over time. The two time points may be separated by about 1 month to about 1 year, about 1 year to about 5 years, or no more than about 3 months.

[0079] Figure 2C shows an exemplary model generated by the system of Figure 2B for inferring tumor phylogeny from next-generation sequencing data. Subclones are related to each other through the process of evolutionary mutation acquisition. In this example, three clones (leaf nodes) are characterized by different combinations of four single nucleotide variant (SNV) sets A, B, C, and D. The percentages at the edges of the tree indicate the fraction of cells that have this particular set of SNVs; for example, 70% of all cells have A, 40% also have B, and only 7% have A, B, and D.

[0080] Figure 2D shows an exemplary heterogeneous population of normal cells and cancer subclones that develop during the evolutionary history of a tumor. The evolutionary history of a tumor results in a heterogeneous population of normal cells (small disks) and cancer subclones (large disks, triangles, squares). Internal nodes that are completely replaced by their descendants (such as those containing SNV sets A and B, but not C or D) are no longer part of the tumor.

[0081] A partnership may be established between the medical prognosis provider and one or more healthcare providers, such as physicians, hospitals, health insurers (e.g., Blue Cross), or managed care organizations (e.g., Kaiser Permanente). The healthcare provider may provide the medical prognosis provider with one or more subject samples containing cfDNA and one or more medical records containing medical information in addition to or other than genetic information about the subject. The medical information may be provided through a secure communication link that allows the medical prognosis provider to access the medical records. The medical prognosis provider may sequence (or have sequenced) the cfDNA from the sample and create a medical record containing information to be used in the methods of the present disclosure. The healthcare provider may provide a new sample containing cfDNA and / or update the information objects passing through the decision node. The predictive model may be iteratively updated as new information becomes available.

[0082] An overview of the process for determining a genetic profile is provided in Figure 3. In this process, genetic material from a blood sample or other bodily sample is received (102). The process converts polynucleotides derived from the genetic material into tagged parent nucleotides (104). The tagged parent nucleotides are amplified to obtain amplified progeny polynucleotides (106). A subset of the amplified polynucleotides is sequenced to obtain sequencing reads (108), which are grouped into families, each generated from a unique tagged parent nucleotide (110). At the selected locus, the process assigns each family a reliability score (112). Previous reads are then used to determine a consensus. This is done by considering previous reliability scores for each family, and if there are consistent previous reliability scores, the current reliability score is increased (114). If there are previous reliability scores, but they are inconsistent, the current reliability score, in one embodiment, is not modified (116). In other embodiments, the reliability score is adjusted in a predetermined manner with respect to inconsistent previous reliability scores. If this is the first time a family has been detected, the current confidence score may be reduced due to a possible erroneous read (118). Based on the confidence score, the process can infer the frequency of the family at the locus within the set of tagged parent polynucleotides (120).

[0083] Although temporal information can enhance information for detecting mutations or copy number variations, other consensus methods can also be applied. In other embodiments, historical comparison can be used in conjunction with other consensus sequences that map to a specific reference sequence to detect genetic alteration events. The consensus sequences that map to a specific reference sequence can be measured and normalized to a control sample. Measurements of molecules that map to a reference sequence can be compared across the genome to identify regions in the genome where copy number changes or heterozygosity has been lost. Consensus methods include, for example, linear or nonlinear consensus sequence construction methods derived from digital communication theory, information theory, or bioinformatics (e.g., voting, averaging, statistical, maximum a posteriori probability or maximum likelihood detection, dynamic programming, Bayesian, hidden Markov, or support vector machine methods). After determining sequence read coverage, a probabilistic modeling algorithm is applied to convert the normalized nucleic acid sequence read coverage of each window region into distinct copy number states. In some cases, the algorithm may include one or more of the following: hidden Markov models, dynamic programming, support vector machines, Bayesian networks, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering techniques, and neural networks.

[0084] After this, report can be generated.For example, copy number variation (CNV) can be reported as a graph showing various positions in genome and the corresponding increase or decrease or maintenance of copy number variation at each position.In addition, copy number variation can be used to report percentage score, which shows how much disease material (or nucleic acid with copy number variation) exists in cell-free polynucleotide sample.

[0085] Figure 4 shows a schematic diagram of internet-enabled access to reports for subjects with cancer. The system in Figure 4 can use a portable or desktop DNA sequencer. A DNA sequencer is a scientific instrument used to automate the DNA sequencing process. In the case of a DNA sample, a DNA sequencer is used to determine the order of the four bases: adenine, guanine, cytosine, and thymine. The order of the DNA bases is reported as a string of characters called a read. Some DNA sequencers can also be considered optical instruments because they analyze light signals originating from fluorescent dyes attached to the nucleotides.

[0086] The tumor profile may include information about the tissue of origin of the tumor. The types and number of cancers that may be detected and profiled include, but are not limited to, blood cancer, brain cancer, lung cancer, skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, colon cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, etc.

[0087] The tumor profile may include information about the drug sensitivity of the tumor. The drug sensitivity of the tumor can be determined directly by measuring or determining the response of isolated tumor cells to the drug of interest. The drug sensitivity of the tumor can be determined by genotyping the tumor.

[0088] The tumor profile may include information regarding tumor size and / or tumor stage. Tumor size can be measured by body scanning techniques, surgery, or any known method. Tumor stage can be determined based on physical examination, imaging studies, laboratory tests, pathology reports, and / or surgical reports.

[0089] The subject profile may include a genetic profile of the subject. The genetic profile of the subject may be determined by assaying non-cancerous tissue from the subject. The genetic profile of the subject may be determined by assaying nucleic acids from acellular body fluid from the subject. Nucleic acids from non-cancerous tissue may be identified, for example, by their frequency in a pool of initial nucleic acids or by the length of the nucleic acid molecule. Nucleic acid molecules from tumor cells may have a first mode of 160-180 bases and a second mode of 320-360 bases. Nucleic acid molecules from non-cancerous tissue may have a broader distribution, with many molecules greater than 400 bases in length. The size of the molecules may be controlled by size selection of the initial DNA molecules or library fragments, or may be controlled informatically by mapping paired reads to a reference genome.

[0090] The genetic profile of a subject can include assaying for variants that can alter the effect of treatment.For example, such variants can affect the pharmacokinetics of drugs.The common variants that affect pharmacokinetics can affect drug transport or drug metabolism.Variants that affect pharmacokinetics are described in MA Rudek et al., The Handbook of Anticancer Pharmacokinetics and Pharmacodynamics, published by Springer Science & Business Media in 2014, which is incorporated herein by reference in its entirety. incorporated herein.

[0091] The genetic profile of a subject can include assaying for variants that affect cancer progression. Such mutations can be, for example, inherited mutations that reduce the efficiency of tumor suppressor products, such as TP53 or BRCA1.

[0092] In some embodiments, the subject profile includes non-genetic information. Such information may include the subject's age, the effectiveness of other medications the patient has taken, clinical information about the subject, and family medical history. The subject's clinical information may include additional clinical information, such as organ function, e.g., liver and kidney function, blood count, cardiac function, pulmonary and respiratory function, and infection status. The subject's clinical information may include age, sex, gender, genetic profile, enzyme levels, organ function, quality of life, frequency of medical interventions, remission status, and / or patient outcome. The subject profile may include information about previous treatments. Treatments may be, for example, surgical removal, radiation, or chemotherapy administration. The information may be qualitative (indicating which treatments have been received) or quantitative, e.g., including dosage, duration, and timing information. The subject's information may include whether the subject is alive or dead. Subject information can be collected at various time points to generate median survival rates, 6-month survival rates, 1-year survival rates, 2-year survival rates, 3-year survival rates, 5-year survival rates, or longer, for a population of subjects.

[0093] Determining a state (e.g., an initial state) may include obtaining information about the object and assigning the object to a state based on the information. In some cases, the state is determined based on a subset of the information. For example, the state may be determined by clustering objects from a training set, and a new object may be assigned to a state by determining which cluster it is closest to.

[0094] Clustering can be used to convert quantitative data into categorical data.For example, certain cancer medications can cause liver damage.The blood liver enzyme (for example, AST and ALT) levels of subjects who are receiving such cancer medications can be measured.Clustering or visual inspection of liver enzyme levels can reveal some subjects with elevated liver enzyme levels and some subjects with normal liver enzyme levels.Liver enzyme levels can be converted into categorical variables by defining subjects with liver enzymes higher than a given level as "elevated" and those lower than a given level as "normal".

[0095] Categorical data and quantitative data can be combined. In one exemplary method, categorical data can be converted for use in methods requiring quantitative data by converting the categorical data into "dummy values." For example, patients with elevated liver enzyme levels can be assigned a value of 1, while patients with normal liver enzyme levels can be assigned a value of 0. Other methods for converting categorical variables into quantitative variables include effect coding, contrast coding, and nonsense coding.

[0096] The state can represent an outcome of interest (e.g., survival, remission status, or the length of time before resistance occurs), and this can be recorded. A set of subjects (e.g., training set) can be used to determine the effect size and interaction of the initial state and / or treatment on the outcome of interest to be determined. These effect sizes and interactions can be used to develop a classifier or predictive model. Methods for determining the effect size and interaction terms of characteristics from the initial state can include, for example, regression analysis, including linear and logarithmic regression analysis; nearest shrinkage centroid analysis; stabilized linear discriminant analysis; support vector machine; Gaussian process; conditional inference tree forest; random forest; nearest centroid; naive Bayes; projection pursuit LDA tree; multinomial logistic regression; stumped decision tree; artificial neural network; binary decision tree; and / or conditional inference tree. The accuracy and sensitivity of a classifier or predictive model can be determined by measuring the prediction accuracy in a subset of subjects (e.g., test set) that was not used to build the classifier or predictive model.

[0097] In some cases, the effect size of predictors is determined, and the variables with little influence are removed.Methods for variable selection are known in the art, and can include, for example, filter method and / or wrapper method for variable selection.Filter method is based on common characteristics, such as the correlation between variables and outcome.Wrapper method evaluates a subset of variables together to determine the optimal combination of variables.Selected variables can be used to determine the subset of information used to determine the state of subject.

[0098] In some cases, the training set of subjects have tumors of the same histological type. In some cases, the subjects are of similar demographic profile, such as the same gender, age, ethnic background, or risk factors. Gender can be male or female. Exemplary risk factors include alcohol consumption, tobacco use and usage, diet, exercise, occupational exposure to carcinogens, frequency of travel, and ultraviolet light exposure and / or sunburn. In some cases, the subjects in the training set are all patients with cancer. In some cases, the subjects in the training set are all patients with symptoms consistent with cancer and undergoing cancer testing. In some cases, the subjects in the training set are patients with symptoms consistent with cancer and undergoing cancer treatment. Subject characteristics can be included in information about each subject of a plurality of subjects.

[0099] An initial state of the subject can be used to determine the probability of a given subsequent state of the subject. The probability can be determined using a classifier or a predictive model.

[0100] Classifiers or predictive models can be used to identify the preferred treatment for a subject with a given profile.For example, using classifiers or predictive models to determine the probability of a given outcome of a subject can include generating one or more decision trees.The state at a first time point can be represented by a root node (which is the initial decision node), and alternative treatments can be represented by decision branches.In some cases, decision branches can lead to a terminal state (from which no further decisions are made) or an intermediate state node, which can itself be a decision node.Intermediate state nodes can represent the appearance of gene variants in one or more tumors of a subject, which confer tumor resistance to treatment; the results of subsequent biopsy or imaging procedures; and / or generally the change or lack of change in information from a subject at a certain time point. For example, intermediate nodes may include information from a subject at 1 week after treatment, 2 weeks after treatment, 3 weeks after treatment, 4 weeks after treatment, 1 month after treatment, 2 months after treatment, 3 months after treatment, 6 months after treatment, 1 year after treatment, 2 years after treatment, 3 years after treatment, 4 years after treatment, or 5 years after treatment. Intermediate nodes may represent intermediate states where a medical care provider makes decisions regarding future treatment options (e.g., after completing a chemotherapy regimen, after surgical intervention to remove a tumor, and at specific time points during an active monitoring regimen).

[0101] Intermediate nodes may contain information regarding the development of resistance to treatment. For example, the presence of a particular variant in a tumor may indicate that resistance is developing. An increase in a particular variant over time during treatment may indicate that the variant, or at least a second unidentified variant, is associated with the development of resistance to treatment. The probability of such a variant emerging may be altered by the presence of a particular variant that predisposes tumors to a particular evolutionary track. Intermediate nodes may contain information regarding the health of a subject (e.g., a patient).

[0102] Tumor profile and / or subject profile can be determined at one or more subsequent time points.The information from the tumor and / or subject profile at subsequent time points can be used to determine subsequent states.Once the subsequent state is determined, the subsequent state can be used as a new initial state to update the probability of other subsequent nodes.For example, if a subject develops a KRAS variant that does not occur simultaneously with the KRAS gene amplification event, the decision tree can be updated to reflect the reduced probability of the KRAS gene amplification event.

[0103] In some cases, the subsequent state is represented by a terminal node (e.g., the subject has died or gone into complete remission). The subsequent state may be a later time point in treatment. The subsequent state may be a time point when an additional biopsy is taken. The biopsy may be a liquid biopsy.

[0104] In some cases, a terminal node represents a state from which no further medical decisions are made. In some cases, a terminal node represents the death of a subject. In some cases, a terminal node represents the inability to detect cancer in a subject.

[0105] In some cases, recommending a treatment involves determining which cluster generated by a classifier or predictive model the information from the subject belongs to. The determining may be based on cluster boundaries determined by the above-described method. In some cases, the determining may be based on selecting the cluster to which the information from the subject is closest. The selecting may be based, at least in part, on distance correlation.

[0106] Such classifier or predictive model can be used to select the treatment of patient.For example, for a patient with a given genetic profile and the genetic profile of tumor, the treatment that maximizes survival rate (for example, 5-year survival rate and / or remission rate) can be selected.Patient can be monitored over time.If a genetic mutation occurs that confers resistance to the treatment or causes an increased risk of developing resistance to the treatment, a second treatment or different treatment can be administered based on the new condition, which maximizes 5-year survival rate and / or remission rate.The treatment that is suitable for maximizing the survival rate and / or lifespan of the subject can be selected.

[0107] Treatments are known to those skilled in the art, and examples are described in the NCCN's Clinical Practice Guidelines in Oncology™ or the American Society of Clinical Oncology (ASCO) practice guidelines. Examples of drugs used in treatment can be found in CMS-approved compendia, including the National Comprehensive Cancer Network's (NCCN) Drugs and Biologics Compendium™, Thomson Micromedex's DrugDex®, Elsevier Gold Standard's Clinical Pharmacology compendium, and the American Hospital Formulary Service-Drug Information Compendium®. Computer Systems

[0108] The present disclosure provides computer systems programmed to implement the methods of the present disclosure. Figure 7 shows a computer system 701 programmed or otherwise configured to detect or monitor cancer evolution.

[0109] The computer system 701 includes a central processing unit (CPU, also referred to herein as a "processor" and a "computer processor") 705, which may be a single-core or multi-core processor, or may have multiple processors for parallel processing. The computer system 701 also includes memory or memory locations 710 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 715 (e.g., a hard disk), a communication interface 720 (e.g., a network adapter) for communicating with one or more other systems, and peripheral devices 725, such as cache, other memory, a data storage unit, and / or an electronic display adapter. The memory 710, the storage unit 715, the interface 720, and the peripheral devices 725 are in communication with the CPU 705 through a communication bus (solid lines), e.g., a motherboard. The storage unit 715 may be a data storage unit (or data repository) for storing data. The computer system 701 may be operably coupled to a computer network ("network") 730 utilizing the communication interface 720. Network 730 may be the Internet, an Internet and / or extranet, or an intranet and / or extranet in communication with the Internet. In some cases, network 730 is a telecommunications and / or data network. Network 730 may include one or more computer servers, which may enable distributed computing such as cloud computing. Network 730 may, in some cases, leverage computer system 701 to implement a peer-to-peer network, which may enable devices coupled to computer system 701 to act as clients or servers.

[0110] The CPU 705 can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 710. The instructions may be directed to the CPU 705, which may then program or otherwise configure the CPU 705 to implement the methods of the present disclosure. Examples of operations performed by the CPU 705 may include fetch, decode, execute, and writeback.

[0111] The CPU 705 may be part of a circuit, such as an integrated circuit. One or more other components of the system 701 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0112] The storage unit 715 can store files such as drivers, libraries, and saved programs. The storage unit 715 can store user data, such as user preferences and user programs. The computer system 701 may, in some cases, include one or more additional data storage units that are external to the computer system 701, for example, located on a remote server that communicates with the computer system 701 through an intranet or the Internet.

[0113] Computer system 701 can communicate with one or more remote computer systems through network 730. For example, computer system 701 may communicate with a remote computer system of a user (e.g., a patient or a healthcare provider). Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone®, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access computer system 701 via network 730.

[0114] The methods described herein may be implemented using machine-executable (e.g., computer processor) code stored in an electronic storage location of the computer system 701, such as on memory 710 or electronic storage unit 715. The machine-executable or machine-readable code may be provided in the form of software. In use, the code may be executed by the processor 705. In some cases, the code may be retrieved from storage unit 715 and stored in memory 710 for immediate access by the processor 705. In some situations, the electronic storage unit 715 may be excluded, and the machine-executable instructions are stored in memory 710.

[0115] The code may be pre-compiled and configured for use on a machine having a processor adapted to execute the code, or may be compiled at run time. The code may be supplied in a programming language that may be selected to allow the code to be executed in a pre-compiled or on-the-fly compiled manner.

[0116] Aspects of the systems and methods provided herein, such as computer system 701, may be embodied in programming. Various aspects of this technology can be considered "products" or "articles of manufacture," typically in the form of machine- (or processor-) executable code and / or associated data carried or embodied on some type of machine-readable medium. The machine-executable code may be stored in an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" type media may include any and all tangible memory of a computer, processor, etc., or their associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which may provide non-transitory storage units at any time during software programming. All or portions of the software may, from time to time, be communicated over the Internet or various other telecommunications networks. Such communication may, for example, enable the software to be loaded from one computer or processor to another, e.g., from a management server or host computer to an application server computer platform. Thus, other types of media that may carry software elements include light waves, radio waves, and electromagnetic waves, such as those used through physical interfaces between local devices, through wired and optical terrestrial networks, and over various air links. Physical elements carrying such waves, such as wired or wireless links, optical links, etc., may also be considered media that carry software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0117] Thus, machine-readable media, such as computer-executable code, may take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer, such as those that can be used to implement the databases shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire, and optical fiber, such as the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched card paper tape, any other physical storage media with a pattern of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves transmitting data or instructions, cables or links transmitting such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0118] The computer system 701 may include or be in communication with an electronic display 735 that includes a user interface (UI) 740 for providing one or more results, for example, related to or indicative of cancer evolution. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0119] The methods and systems of the present disclosure can be implemented using one or more algorithms. The algorithms can be implemented by software when executed by the central processing unit 705. The algorithms can, for example, implement the methods of the present disclosure for detecting or monitoring the evolution of cancer. [Example]

[0120] Example 1 Building a model for the development of treatment resistance Subjects with cancer undergo physical screening to determine a patient profile, including their age, gender, cancer type, cancer stage, and organ function. The subject undergoes a blood draw, which is processed to remove cells and obtain acellular body fluid containing nucleic acids. The nucleic acids are sequenced to determine the patient's genetic profile and the tumor's genetic profile. The subject is prescribed treatment by a physician. The patient is followed over time, and the tumor's genetic profile is obtained every three months. The patient's outcome is recorded at each time point.

[0121] A hidden Markov model is constructed based on the probability that a patient with a given patient profile (including the patient's genetic profile) and tumor's genetic profile will have a particular patient outcome at any given time point. Example 2 Use of models for the development of treatment resistance

[0122] A subject with cancer is admitted to the hospital. A subject profile and a tumor profile are obtained. The subject profile and the tumor profile are used as the initial state of a model, for example, the model generated in Example 1. The subject's outcome is predicted based on the model, and a treatment is selected to maximize the subject's expected survival time (e.g., measured in months or years). The subject's tumor profile is updated every three months and used as a new initial state input to the model. At a given subsequent time point, the tumor profile indicates the emergence of a subclone that is resistant to the current treatment. In response, a new treatment is selected to maximize the subject's expected survival time. The subject is given a second treatment (e.g., a second-line treatment) that targets tumor cells that are resistant to the first treatment (e.g., a first-line treatment). Example 3 Displaying Objects with Decision Trees

[0123] The subject is associated with an initial node indicating that the subject is a 65-year-old male with colon cancer, and the tumor profile indicates that a low-frequency KRAS mutation is detected in the subject's cell-free DNA. One branch stemming from the initial node indicates treatment with panitumumab and cetuximab, and a second branch indicates treatment with panitumumab and cetuximab administered in conjunction with a mitogen-activated protein kinase (MEK) inhibitor. These branches connect to intermediate nodes indicating resistance development and lack of resistance development. The probability of resistance development is lower in intermediate nodes along the branch including co-treatment with a MEK inhibitor than in the branch lacking co-treatment with a MEK inhibitor. Each intermediate node is associated with a terminal node indicating death and complete remission. The probability of complete remission is higher in the terminal node along the decision branch including co-treatment with a MEK inhibitor.

[0124] The illustrations of the embodiments described herein are intended to provide a general understanding of the structure of various embodiments. The illustrations are not intended to serve as a complete description of all of the elements and features of apparatus and systems that utilize the structures or methods described herein. Numerous other embodiments may be apparent to those skilled in the art upon reviewing the present disclosure. Other embodiments may be utilized and derived from the present disclosure, and as a result, structural and logical substitutions and changes may be made without departing from the scope of the present disclosure. Accordingly, the present disclosure and the figures are to be interpreted as illustrative and not restrictive.

[0125] One or more embodiments of the present disclosure may be referred to herein, individually and / or collectively, by the term "the present invention," merely for convenience and without any intention to arbitrarily limit the scope of the present application to any particular invention or inventive concept. Furthermore, while specific embodiments have been illustrated and described herein, it should be understood that any subsequent arrangement designed to achieve the same or similar purpose may be substituted for the specific embodiment shown. The present disclosure is intended to cover any and all subsequent adaptations and variations of the various embodiments. Combinations of the above-described embodiments, as well as other embodiments not specifically described herein, will be apparent to those skilled in the art upon consideration of the description.

[0126] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided herein. While the present invention has been described with reference to the foregoing specification, the descriptions and illustrations of the embodiments herein are not intended to be limiting. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific expressions, configurations, or relative proportions described herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein can be employed in practicing the invention. It is therefore intended that the present invention shall cover any and all such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the invention, and that methods and structures within the scope of these claims, and their equivalents, be covered thereby. Examples of embodiments of the present invention include the following: (Item 1) 1. A computer-implemented method comprising: (a) obtaining information about a plurality of subjects with cancer at a first time point, and determining a first state of each of the plurality of subjects based on the information at the first time point to generate a set of first states, wherein the information includes, for each subject of the plurality of subjects, at least a tumor genetic profile obtained by genotyping nucleic acids from acellular body fluids and any treatments provided to the subject prior to the first time point; (b) obtaining the information regarding the plurality of objects at one or more second time points subsequent to the first time point, and determining a second state of each of the plurality of objects at each of the one or more second time points based on the information at a given one of the one or more second time points to generate a set of subsequent states; (c) using the first set of states from (a) and the set of subsequent states from (b), generating a prediction algorithm configured to determine a probability that a given first state will result in a second state among the set of states at a later time point after the given first state; A method comprising: (Item 2) (d) determining the probability, for a given first state among a set of states at an earlier time, that the given first state will result in the second state among the set of states at the later time; (e) generating an electronic output indicative of the probability determined in (d); Item 1, the method of claim 1 further comprising: (Item 3) 1. A computer-implemented method comprising: (a) obtaining information about a plurality of subjects having cancer at a first time point, and determining a first state of each of the plurality of subjects based on the information at the first time point to generate a set of first states, wherein the information includes, for each subject of the plurality of subjects, at least a tumor genetic profile obtained by genotyping at least 50 genes and any treatments provided to the subject prior to the first time point; (b) obtaining the information regarding the plurality of objects at one or more second time points subsequent to the first time point, and determining a second state of each of the plurality of objects at each of the one or more second time points based on the information at a given one of the one or more second time points to generate a set of subsequent states; (c) using the first set of states from (a) and the set of subsequent states from (b), generating a prediction algorithm configured to determine a probability that a given first state will result in a second state among the set of states at a later time point after the given first state; A method comprising: (Item 4) (d) determining the probability, for a given first state among a set of states at an earlier time, that the given first state will result in the second state among the set of states at the later time; (e) generating an electronic output indicative of the probability determined in (d); Item 4. The method of item 3, further comprising: (Item 5) 4. The method of claim 1 or 3, wherein the step of obtaining information comprises sequencing cell-free deoxyribonucleic acid (cfDNA) from the plurality of subjects and, optionally, conducting a medical interview with each of the plurality of subjects. (Item 6) 4. The method of item 1 or 3, wherein treatment was provided to the subject prior to the first time point. (Item 7) 4. The method of claim 1, further comprising generating one or more decision trees, each decision tree including a root node, one or more decision branches, one or more decision nodes, and one or more terminal nodes, wherein a state at the root node represents the first time point, the one or more decision branches represent alternative actions, and the one or more decision nodes and the one or more terminal nodes represent subsequent states. (Item 8) 8. The method of claim 7, wherein the one or more decision branches include a plurality of decision branches. (Item 9) 4. The method of claim 1 or 3, wherein the subsequent status comprises a vital status of the subject, indicating whether the subject is alive or dead. (Item 10) 4. The method of item 1 or 3, wherein the subsequent status comprises subject survival. (Item 11) 4. The method of claim 1 or 3, wherein each of the first states comprises a common set of one or more somatic mutations. (Item 12) 4. The method of claim 1 or 3, wherein the information further comprises a subject profile. (Item 13) 4. The method of item 1 or 3, wherein the probability is, at least in part, a function of treatment selection from among a plurality of treatment options. (Item 14) 5. The method of any one of items 1 to 4, wherein the one or more second time points comprises a plurality of subsequent time points. (Item 15) 15. The method of claim 14, further comprising determining the probability at multiple subsequent time points. (Item 16) 16. The method of item 15, wherein the time points include at least three time points or at least four time points. (Item 17) Item 10. The method of item 1, wherein the first time point is before the subject receives the treatment and the subsequent time point is after the subject receives the treatment. (Item 18) 14. The method of claim 13, wherein a second treatment is administered after the subsequent time point based on the subsequent status at the subsequent time point. (Item 19) 2. The method of claim 1, wherein the information about the plurality of subjects comprises one or more features from patient profiles of the subjects, the features being selected from the group consisting of age, sex, gender, genetic profile, enzyme levels, organ function, quality of life, frequency of medical interventions, remission status, and patient outcome. (Item 20) 20. The method of claim 19, wherein the genetic profile comprises the subject's genotype at one or more loci that increase cancer risk, affect pharmacokinetics, or affect drug sensitivity. (Item 21) 2. The method of claim 1, wherein the information about the plurality of subjects comprises one or more features from the subject's tumor profile, the features being selected from the group consisting of one or more genetic variants, tissue of origin, tumor burden, tumor drug sensitivity, and tumor stage. (Item 22) 22. The method of claim 21, wherein the one or more characteristics are determined by assaying cell-free nucleic acid molecules from the subject. (Item 23) 23. The method of claim 22, wherein the one or more genetic variants are quantified to determine the proportion of cell-free nucleic acid molecules that contain the one or more somatic mutations. (Item 24) 24. The method of claim 23, further comprising determining whether the ratio of the one or more somatic mutations increases or decreases between the first time point and the one or more subsequent time points. (Item 25) 24. The method of claim 23, further comprising determining whether the ratio of the one or more somatic mutations increases or decreases between a plurality of the one or more subsequent time points. (Item 26) 26. The method of item 24 or 25, wherein the proportion of the one or more somatic mutations is increased. (Item 27) 27. The method of claim 26, wherein the one or more somatic mutations are increased and further wherein the somatic mutations are associated with resistance to the treatment. (Item 28) 23. The method of claim 22, wherein the assaying comprises high-throughput sequencing. (Item 29) (a) obtaining information about a subject with cancer at a first time point, the information including at least one characteristic of the subject from a patient profile, a tumor profile, or a treatment; (b) determining an initial state of the object based on the information at the first time point; (c) determining a probability of each of a plurality of subsequent states at each of one or more subsequent time points based on the initial state of the object, thereby providing a set of probabilities for state outcomes; (d) generating a treatment recommendation for the cancer that optimizes the probability of the subject achieving a particular outcome based at least in part on the set of probabilities for condition outcomes; (e) generating an electronic output indicative of the recommendations generated in (d); A method comprising: (Item 30) 30. The method of claim 29, wherein the probability is at least in part a function of treatment selection from among a plurality of treatment options. (Item 31) 31. The method of item 29 or 30, wherein the one or more subsequent time points comprises a plurality of subsequent time points. (Item 32) 32. The method of claim 31, further comprising determining the probability at multiple subsequent time points. (Item 33) 30. The method of item 29, wherein the time points include at least three time points. (Item 34) 30. The method of item 29, wherein the time points include at least four time points. (Item 35) 30. The method of claim 29, wherein the first time point is before the subject receives the treatment and the subsequent time point is after the subject receives the treatment. (Item 36) 36. The method of claim 35, wherein a second treatment is administered after the subsequent time point based on the subsequent status at the subsequent time point. (Item 37) 30. The method of claim 29, wherein the at least one characteristic of the subject is from the patient profile and is selected from the group consisting of age, gender, genetic profile, enzyme levels, organ function, quality of life, frequency of medical interventions, remission status, and patient outcome. (Item 38) 30. The method of claim 29, wherein the genetic profile comprises the subject's genotype at one or more loci that are hereditary cancer genes. (Item 39) 30. The method of claim 29, wherein the genetic profile comprises the subject's genotype at one or more genetic loci that affect pharmacokinetics. (Item 40) 30. The method of item 29, wherein the genetic profile comprises the subject's genotype at one or more loci that influence drug sensitivity. (Item 41) 30. The method of claim 29, wherein the at least one characteristic of the subject is from the tumor profile and is selected from the group consisting of one or more somatic mutations, tissue of origin, tumor burden, tumor drug sensitivity, and tumor stage. (Item 42) 41. The method of claim 40, wherein the at least one characteristic is determined by assaying cell-free nucleic acid molecules from the subject. (Item 43) 43. The method of claim 42, wherein the somatic mutations are quantified to determine the proportion of cell-free nucleic acid molecules from the tumor that contain the one or more somatic mutations. (Item 44) 44. The method of claim 43, further comprising determining whether the ratio of the one or more somatic mutations increases or decreases between the first time point and the one or more subsequent time points. (Item 45) 44. The method of claim 43, further comprising determining whether the ratio of the one or more somatic mutations increases or decreases between multiple of the one or more subsequent time points. (Item 46) 43. The method of claim 42, wherein the assaying comprises high-throughput sequencing. (Item 47) 30. The method of claim 29, wherein the tumor profile is not derived from a tumor tissue biopsy. (Item 48) (a) obtaining information about a subject, including at least a genetic profile of the tumor and any treatments previously or currently being provided to said subject, and determining an initial condition of said subject based on said information; (b) providing a decision tree, wherein a root node represents an initial object state, decision branches represent alternative actions available for the object, opportunity nodes represent points of uncertainty, and decision or terminal nodes represent successor states; (c) providing a course of treatment for the subject that maximizes the probability that the subject will achieve a survival state at the terminal node; (d) generating an electronic output indicative of the course of treatment determined in (c); A method comprising: (Item 49) (a) establishing one or more communications links with one or more healthcare providers over a communications network; (b) receiving, over the communications network, medical information regarding one or more subjects from the one or more health care providers; (c) receiving from the health care provider one or more samples containing cell-free deoxyribonucleic acid (cfDNA) from each of the one or more subjects; (d) sequencing the cfDNA and identifying one or more genetic variants present in the cfDNA; (e) creating or supplementing a database with information about each of the one or more subjects, the information including both the identified genetic variants and the received medical information; (f) using the database and a computer-implemented algorithm to generate at least one predictive model that predicts the probability of a subsequent state for each of a plurality of different therapeutic interventions based on the subject's initial state; A method comprising: (Item 50) When executed by one or more computer processors, (a) obtaining information about a plurality of subjects with cancer at a first time point, and determining a first state of each of the plurality of subjects based on the information at the first time point to generate a set of first states, wherein the information includes, for each subject of the plurality of subjects, at least a tumor genetic profile obtained by genotyping nucleic acids from acellular body fluids and any treatments provided to the subject prior to the first time point; (b) obtaining the information regarding the plurality of objects at one or more second time points subsequent to the first time point, and determining a second state of each of the plurality of objects at each of the one or more second time points based on the information at a given one of the one or more second time points to generate a set of subsequent states; (c) using the first set of states from (a) and the set of subsequent states from (b), generating a prediction algorithm configured to determine a probability that a given first state will result in a second state among the set of states at a later time point after the given first state; A non-transitory computer readable medium comprising machine executable code implementing a method comprising: (Item 51) When executed by one or more computer processors, (a) obtaining information about a plurality of subjects having cancer at a first time point, and determining a first state of each of the plurality of subjects based on the information at the first time point to generate a set of first states, wherein the information includes, for each subject of the plurality of subjects, at least a tumor genetic profile obtained by genotyping at least 50 genes and any treatments provided to the subject prior to the first time point; (b) obtaining the information regarding the plurality of objects at one or more second time points subsequent to the first time point, and determining a second state of each of the plurality of objects at each of the one or more second time points based on the information at a given one of the one or more second time points to generate a set of subsequent states; (c) using the first set of states from (a) and the set of subsequent states from (b), generating a prediction algorithm configured to determine a probability that a given first state will result in a second state among the set of states at a later time point after the given first state; A non-transitory computer readable medium comprising machine executable code implementing a method comprising: (Item 52) (a) obtaining information about a subject, including at least a genetic profile of the tumor and any treatments previously or currently being provided to said subject, and determining an initial condition of said subject based on said information; (b) providing a decision tree, wherein a root node represents an initial object state, decision branches represent alternative actions available for the object, opportunity nodes represent points of uncertainty, and decision or terminal nodes represent successor states; (c) providing a course of treatment for the subject that maximizes the probability that the subject will achieve a survival state at the terminal node; (d) administering said course of treatment to said subject; A method comprising: (Item 53) (e) obtaining information about the subject at a second time point, later than the initial state, including at least a genetic profile of the tumor and any treatments previously or currently provided to the subject, and determining a second state for the subject from among a plurality of subsequent states based on the information; (f) providing a subsequent course of treatment for the subject based on the second state, the course of treatment maximizing the probability that the subject will achieve a survival state at a terminal node; (g) administering said subsequent course of treatment to said subject. 53. The method of claim 52, further comprising: (Item 54) 1. A method comprising providing a course of treatment from among a plurality of alternative treatments for a subject having cancer, wherein the subject is characterized by a decision tree comprising a plurality of decision branches, each decision branch representing an alternative treatment from among the plurality of alternative treatments, the course of treatment maximizing the probability that the subject will achieve a survival state at a terminal node.

Claims

1. A method executed by one or more computer processors, comprising: (a) obtaining information about a subject, including at least a genetic profile of a tumor and any treatments previously or currently being provided to the subject, and determining an initial condition of the subject based on the information, wherein the information about the subject includes at least one characteristic of the subject from the genetic profile of the tumor, selected from the group consisting of one or more somatic mutations, tissue of origin, tumor burden, tumor drug sensitivity, and tumor stage, and wherein the at least one characteristic is determined by assaying cell-free nucleic acid molecules from the subject; (b) providing a decision tree, wherein a root node represents an initial object state, decision branches represent alternative actions available for said object, opportunity nodes represent points of uncertainty, and decision or terminal nodes represent successor states; a probability is assigned to each branch of the decision branches and a probability is assigned to each node of the terminal nodes; one or more of the probabilities of each decision branch and one or more of the probabilities of each terminal node are at least in part a function of a treatment selection from among a plurality of treatment alternatives; one or more probabilities for each decision branch and one or more probabilities for each terminal node are generated from a database and one or more machine learning algorithms; (c) providing a course of treatment for the subject that maximizes the probability that the subject will achieve a survival state at the terminal node; (d) generating an electronic output indicative of the treatment course provided in (c); A method comprising:

2. 10. The method of claim 1, wherein the genetic profile comprises the subject's genotype at one or more loci that are hereditary cancer genes.

3. 10. The method of claim 1, wherein the genetic profile comprises the subject's genotype at one or more genetic loci that affect pharmacokinetics.

4. 10. The method of claim 1, wherein the genetic profile comprises the subject's genotype at one or more loci that affect drug sensitivity.

5. 10. The method of claim 1, wherein somatic mutations are quantified to determine the proportion of cell-free nucleic acid molecules derived from the tumor that contain said one or more somatic mutations.

6. 6. The method of claim 5, further comprising determining whether the ratio of the one or more somatic mutations increases or decreases between a first time point and one or more subsequent time points.

7. 6. The method of claim 5, further comprising determining whether the ratio of the one or more somatic mutations is increasing or decreasing between a plurality of the one or more subsequent time points.

8. The method of claim 1 , wherein said assaying comprises high-throughput sequencing.

9. The method of claim 1 , wherein the genetic profile of the tumor is not derived from a tumor tissue biopsy.

10. 2. The method of claim 1, wherein maximizing the probability that the object achieves a surviving state at the terminal node comprises determining that the object corresponds to a cluster of objects in a database comprising a plurality of clusters.

11. The method of claim 1 , wherein the one or more machine learning algorithms are Hidden Markov Models.

12. A method executed by one or more computer processors, comprising: (a) obtaining information about a subject, including at least a genetic profile of a tumor and any treatments previously or currently being provided to the subject, and determining an initial condition of the subject based on the information, wherein the information about the subject includes at least one characteristic of the subject from the genetic profile of the tumor, selected from the group consisting of one or more somatic mutations, tissue of origin, tumor burden, tumor drug sensitivity, and tumor stage, and wherein the at least one characteristic is determined by assaying cell-free nucleic acid molecules from the subject; (b) providing a decision tree, wherein a root node represents an initial object state, decision branches represent alternative actions available for said object, opportunity nodes represent points of uncertainty, and decision or terminal nodes represent successor states; a probability is assigned to each branch of the decision branches and a probability is assigned to each node of the terminal nodes; one or more of the probabilities of each decision branch and one or more of the probabilities of each terminal node are at least in part a function of a treatment selection from among a plurality of treatment alternatives; one or more probabilities for each decision branch and one or more probabilities for each terminal node are generated from a database and one or more machine learning algorithms; (c) providing a course of treatment for the subject that maximizes the probability that the subject will achieve a survival state at the terminal node, the course of treatment being administered to the subject; A method comprising:

13. (e) obtaining information about the subject at a second time point, later than the initial state, including at least a genetic profile of the tumor and any treatments previously or currently being provided to the subject, and determining a second state of the subject from among a plurality of subsequent states based on the information; (f) providing a subsequent course of treatment for the subject based on the second state, the subsequent course of treatment maximizing a probability that the subject will achieve a survival state at a terminal node, the second state being an indicator of whether the subsequent course of treatment is to be administered to the subject; The method of claim 12 further comprising:

14. A method executed by one or more computer processors, comprising: (a) establishing one or more communication links with one or more healthcare service providers over a communications network; (b) receiving, over the communications network, medical information regarding one or more subjects from the one or more health care providers; (c) receiving from the health care provider one or more samples comprising cell-free deoxyribonucleic acid (cfDNA) from each of the one or more subjects; (d) sequencing the cfDNA to identify one or more characteristics of the one or more subjects, wherein the one or more characteristics of the subjects are a genetic profile of a tumor and are selected from the group consisting of one or more somatic mutations, tissue of origin, tumor burden, tumor drug sensitivity, and tumor stage; (e) creating or providing a database having information about each of the one or more subjects, the information including both the identified one or more characteristics and the received medical information; (f) using the database and a computer-implemented algorithm to generate at least one predictive model that predicts the probability of a subsequent state for each of a plurality of different therapeutic interventions based on the subject's initial state; Including, The computer-implemented algorithm comprises: providing a decision tree, wherein a root node represents an initial object state, decision branches represent alternative actions available for said object, opportunity nodes represent points of uncertainty, and decision or terminal nodes represent successor states; a probability is assigned to each branch of the decision branches and a probability is assigned to each node of the terminal nodes; one or more of the probabilities of each decision branch and one or more of the probabilities of each terminal node are at least in part a function of a treatment selection from among a plurality of treatment alternatives; one or more of the probabilities for each decision branch and one or more of the probabilities for each terminal node are generated from one or more machine learning algorithms. A method comprising:

Citation Information

Patent Citations

  • Medical analysis system

    JP2011520206A

  • Compositions and methods for predicting drug sensitivity, drug resistance, and disease progression

    JP2013525786A

  • Medical care treatment decision support system

    US20120047105A1

  • Assays, methods and kits for analyzing sensitivity and resistance to Anti-cancer drugs, predicting a cancer patient's prognosis, and personalized treatment strategies

    WO2014197543A1

  • Method for sensitive detection of target DNA using target-specific nuclease

    WO2015183025A1