System and method for detecting and treating diseases exhibiting disease cell heterogeneity, as well as for communication test results.

By sequencing cancer cells to identify and quantify somatic mutations, a method develops personalized therapeutic interventions for tumor heterogeneity, addressing the challenges of evolving resistance and genetic variant detection in cancer treatment.

JP2026090415APending Publication Date: 2026-06-02GUARDANT HEALTH INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
GUARDANT HEALTH INC
Filing Date
2026-02-13
Publication Date
2026-06-02

Smart Images

  • Figure 2026090415000001_ABST
    Figure 2026090415000001_ABST
Patent Text Reader

Abstract

This invention provides a system and method for detecting and treating diseases exhibiting disease cell heterogeneity, as well as for obtaining communication test results. [Solution] In one embodiment, a method is provided that includes: (a) sequencing polynucleotides from cancer cells derived from a biological sample of interest; (b) identifying and quantifying somatic mutations in the polynucleotides; (c) developing a tumor heterogeneity profile in the subject, indicating the presence and relative amounts of multiple somatic mutations in the polynucleotides, wherein different relative amounts indicate tumor heterogeneity; and (d) determining a therapeutic intervention for a cancer exhibiting tumor heterogeneity, wherein the therapeutic intervention is effective for cancers having the determined tumor heterogeneity profile.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Provisional Application No. 62 / 098,426, filed on December 31, 2014, and U.S. Provisional Application No. 62 / 155,763, filed on May 1, 2015, each of which is hereby incorporated by reference in its entirety.

Background Art

[0002] (Background) Medicine has only just begun to effectively use information from the human genome for the diagnosis and treatment of diseases. Nowhere is this more important than in the treatment of cancer, which causes 7.6 million deaths each year in the United States and costs the country $87 billion annually to treat. Cancer refers to any disorder of various malignant neoplasms characterized by the growth of undifferentiated cells that tend to invade surrounding tissues and metastasize to new body sites, and the pathological conditions characterized by such growth.

[0003] One reason cancer treatment is difficult is that current testing methods may not be useful when physicians try to match a specific cancer to an effective drug treatment. And cancer is a moving target, that is, cancer cells are constantly changing and mutating. Cancer can accumulate genetic variants, for example, due to somatic mutations. Such variants include, for example, sequence variants and copy number variants. Analysis of tumors has shown that different cells in a tumor can have different genetic variants. Such differentiation between tumor cells has come to be called tumor heterogeneity.

[0004] Cancer can evolve over time and become resistant to treatment interventions. It is known that certain variants correlate with responsiveness or resistance to specific treatment interventions. For cancers showing tumor heterogeneity, more effective treatments would be beneficial. Such cancers can be treated with a second, different treatment intervention to which the cancer responds.

[0005] DNA sequencing methods enable the detection of gene variants in DNA derived from tumor cells. Cancerous tumors frequently shed their unique genomic material into the bloodstream. Unfortunately, these evidentiary genomic "signals" are so weak that current genomic analysis techniques, including next-generation sequencing, can only detect them sporadically or in patients with terminally ill, high tumor burdens. The main reason for this is that these techniques suffer from error rates and biases that can be orders of magnitude higher than what is needed for reliable detection of cancer-related de novo genomic alterations.

[0006] In parallel with this trend, in order to understand the clinical significance of genetic testing, treatment professionals must possess a practical knowledge of the fundamental principles of genetics and a reasonable ability to interpret probabilistic data. Some studies suggest that many treatment professionals are not adequately prepared to interpret genetic testing regarding disease susceptibility. Some physicians have difficulty interpreting probabilistic data regarding the clinical usefulness of diagnostic tests, such as positive or negative predictive values ​​for laboratory tests.

[0007] Error rates and biases in detecting cancer-related de novo genomic changes, along with inadequate explanations or implications of cancer-related genetic testing, have reduced the quality of care for cancer patients. (College of American Pathologists (CAP) and the U.S. College of Pathologists and the U.S. College of Clinical Pathologists) Professional societies such as the American College of Medical Genetics (ACMG) publish standards or guidelines for laboratories that provide genetic testing, which require that reports containing genetic information include interpretable content that can be understood by general practitioners. [Overview of the Initiative] [Means for solving the problem]

[0008] (Summary) In some embodiments, a method is provided herein that includes: (a) sequencing polynucleotides from cancer cells derived from a biological sample of interest; (b) identifying and quantifying somatic mutations in the polynucleotides; (c) developing a tumor heterogeneity profile in the interest, indicating the presence and relative amounts of multiple somatic mutations in the polynucleotides, wherein different relative amounts indicate tumor heterogeneity; and (d) determining a therapeutic intervention for cancer exhibiting tumor heterogeneity, wherein the therapeutic intervention is effective for cancer having the determined tumor heterogeneity profile. In some embodiments, the cancer cells are spatially distinct. In some embodiments, the therapeutic intervention is effective for cancer exhibiting multiple somatic mutations rather than for cancer exhibiting one type of somatic mutation rather than all of them. In some embodiments, the method further includes (e) monitoring changes in tumor heterogeneity in the interest over time and determining different therapeutic interventions over time based on the changes. In some embodiments, the method further includes (e) displaying the therapeutic intervention. In some embodiments, the method further includes (e) implementing the therapeutic intervention. In some embodiments, the method further includes (e) generating a phylogenetic tree of tumor evolution based on a tumor profile, wherein the step of determining a therapeutic intervention takes the phylogenetic tree into consideration.

[0009] In some embodiments, the determination step is performed with the help of an algorithm executed by a computer. In some embodiments, the sequence readings generated by sequencing are subjected to noise reduction before the identification and quantification steps. In some embodiments, noise reduction includes molecular tracking of sequences generated from single polynucleotides in the sample.

[0010] In some embodiments, the step of determining a therapeutic intervention takes into account the relative frequencies of tumor-associated genetic alterations. In some embodiments, the therapeutic intervention includes administering multiple drugs in combination or sequentially, each drug being relatively more effective against cancers that present a different somatic mutation occurring at different relative frequencies. In some embodiments, drugs that are relatively more effective against cancers that present somatic mutations occurring at higher relative frequencies are administered in larger doses. In some embodiments, drugs are delivered in stratified doses that reflect the relative amounts of variants in DNA. In some embodiments, cancers presenting at least one genetic variant are resistant to at least one of the drugs. In some embodiments, the step of determining a therapeutic intervention takes into account the tissue of origin of the cancer. In some embodiments, the therapeutic intervention is determined based on a database of interventions that are shown to be treatments for cancers having tumor heterogeneity characterized by each of the somatic mutations.

[0011] In some embodiments, the polynucleotides include cfDNA derived from a blood sample. In some embodiments, the polynucleotides include polynucleotides derived from spatially distinct cancer cells. In some embodiments, the polynucleotides include polynucleotides derived from different metastatic tumor sites. In some embodiments, the polynucleotides include polynucleotides derived from a solid tumor or a diffused tumor. In some embodiments, the polynucleotides are contained in a blood sample or a solid tumor biopsy.

[0012] In some embodiments, the identification step includes generating multiple sequence reads for parent polynucleotides derived from the sample, and reducing the sequence reads to generate a consensus call for the bases in each parent polynucleotide. In some embodiments, the quantification step includes determining the frequency at which somatic mutations are detected in a population of polynucleotides derived from the biological sample. In some embodiments, the biological sample includes biological molecules derived from non-disease cells. In some embodiments, the biological sample includes biological molecules derived from multiple different tissues. In some embodiments, the biomolecules are contained in one biological sample. In some embodiments, the biomolecules are contained in multiple biological samples. In some embodiments, the multiple biological samples are multiple metastatic tumors.

[0013] In some embodiments, the sequencing step includes sequencing all or part of a subset of genes in the genome of interest. In some embodiments, somatic mutations are selected from single-nucleonucleotide mutations (SNVs), insertions, deletions, inversions, transversions, transpositions, copy number variations (CNVs) (e.g., aneuploidy, partial aneuploidy, polyploidy), chromosomal instability, chromosomal structural changes, gene fusions, chromosome fusions, gene truncation, gene amplification, gene duplication, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid methylation. In some embodiments, loci are selected from single nucleotides, genes, and chromosomes.

[0014] In some embodiments, cancer is selected from carcinomas, sarcomas, leukemias, lymphomas, myelomas, and central nervous system cancers (e.g., breast cancer, prostate cancer, colorectal cancer, brain cancer, esophageal cancer, head and neck cancer, bladder cancer, gynecological cancers, liposarcoma, and multiple myeloma). In some embodiments, the cancer cells of the tumor originate from common parent disease cells. In some embodiments, the cancer cells of the tumor originate from different parent cancer cells of the same or different cancer types. In some embodiments, the method further includes the step of determining a measure of somatic mutation against one or more control references to determine relative content relative amounts.

[0015] In some embodiments, polynucleotides are supplied from both circulating cancer polynucleotides and solid tumor biopsies. In some embodiments, profiles are prepared separately for polynucleotides supplied from circulating cancer polynucleotides and solid tumor biopsies.

[0016] In some embodiments, a method is provided herein that provides a therapeutic intervention to a subject having a cancer having a tumor profile in which tumor heterogeneity can be inferred, wherein the therapeutic intervention is effective against the cancer having this tumor profile. In some embodiments, the tumor profile exhibits the relative frequencies of several more somatic mutations. In some embodiments, the method further includes the steps of monitoring changes in relative frequencies in the subject over time and determining different therapeutic interventions over time based on the changes. In some embodiments, the therapeutic intervention is effective against cancers exhibiting each of the somatic mutations rather than against cancers exhibiting any one of the somatic mutations rather than all of them. In some embodiments, the therapeutic intervention involves administering several drugs in combination or sequentially, wherein each drug is relatively more effective against cancers exhibiting different somatic mutations occurring at different relative frequencies. In some embodiments, drugs that are relatively more effective against cancers exhibiting somatic mutations occurring at higher relative frequencies are administered in larger doses. In some embodiments, the drugs are delivered in stratified doses that reflect the relative amounts of variants in DNA. In some embodiments, cancers presenting at least one gene variant are resistant to at least one drug. In some embodiments, cancers are selected from carcinomas, sarcomas, leukemias, lymphomas, myelomas, and central nervous system cancers (e.g., breast cancer, prostate cancer, colorectal cancer, brain cancer, esophageal cancer, head and neck cancer, bladder cancer, gynecological cancers, liposarcomas, and multiple myelomas).

[0017] In one embodiment, the present invention provides a method comprising the step of administering a therapeutic intervention effective against a tumor exhibiting tumor heterogeneity, wherein the therapeutic intervention exhibits different relative amounts of tumor heterogeneity, based on a tumor heterogeneity profile in the subject, which indicates the presence and relative amounts of multiple somatic mutations in polynucleotides.

[0018] In one embodiment, a system is provided herein, including a computer-readable medium containing machine-executable code, that implements a method comprising: (a) receiving sequence readings of polynucleotides mapped to a locus into memory; (b) determining the identity of bases different from the bases of the reference sequence at the locus among the sequence readings for all the sequence readings mapped to the locus; (c) reporting the identity and relative amounts of the determined bases and their locations in the genome; and (d) estimating the heterogeneity of a given sample based on the information in (c). In some embodiments, the implemented method further includes receiving sequence readings from samples at multiple different time points into memory, and calculating the differences in the relative amounts and identities of multiple bases between two samples.

[0019] In some embodiments, a kit comprising a first pharmaceutical and a second pharmaceutical is provided herein, wherein the combination of the first and second pharmaceuticals is therapeutically effective against cancers presenting both the first and second somatic mutations rather than against cancers presenting any one of the somatic mutations rather than all of them. In some embodiments, the combination is contained in a mixture, or each drug is contained in a separate container.

[0020] In some embodiments, a method is provided herein that includes the steps of (a) performing a biomolecular analysis of a biomolecular polymer from disease cells derived from a subject (e.g., spatially distinct disease cells); (b) identifying and quantifying biomolecular variants in the biomolecular polymer; (c) developing a disease cell heterogeneity profile in the subject, indicating the presence and relative amounts of multiple variants in the biomolecular polymer, wherein different relative amounts indicate disease cell heterogeneity; and (d) determining a therapeutic intervention for a disease exhibiting disease cell heterogeneity, wherein the therapeutic intervention is effective for a disease having the determined disease cell heterogeneity profile. In some embodiments, the disease cells are spatially distinct disease cells. In some embodiments, the therapeutic intervention is determined based on a database of interventions, which are shown to be treatments for cancers having tumor heterogeneity characterized by each of somatic mutations.

[0021] In one embodiment, a method for detecting disease cell heterogeneity in a subject is provided herein, comprising the steps of: a) quantifying polynucleotides having sequence variants at each of a plurality of loci in a polynucleotide from a sample derived from a subject, wherein the sample comprises polynucleotides derived from somatic cells and disease cells; b) determining a measure of copy number variation (CNV) for polynucleotides having sequence variants for each locus; c) determining a weighted measure for each locus, as a function of CNV at the locus, of the amount of polynucleotides having sequence variants at the locus; and d) comparing the weighted measures at each of the plurality of loci, wherein different weighted measures indicate disease cell heterogeneity. In some embodiments, the disease cells are tumor cells. In some embodiments, the polynucleotides include cfDNA.

[0022] In some embodiments, the present invention provides a method comprising the steps of a) targeting one or more pulsed therapeutic cycles, each pulsed therapeutic cycle comprising (i) a first period in which one or more drugs are administered in a first amount and (ii) a second period in which one or more drugs are administered in a second reduced amount (e.g., not administered at all), wherein (A) the first period is characterized by a tumor load detected above a first clinical level and (B) the second period is characterized by a tumor load detected below a second clinical level. In some embodiments, the tumor load is measured as a function of the amount of selected somatic variants in tumor polynucleotides. In some embodiments, one or more drugs are multiple drugs, and each amount of each drug in each cycle is determined as a function of the tumor load, which is measured as a function of the respective amounts of multiple different selected somatic variants in tumor polynucleotides. In some embodiments, the method comprises the step of targeting multiple pulsed therapeutic cycles. In some embodiments, the method further includes the step of (b) if the subject is resistant to one or more drugs, subject to one or more pulsed therapy cycles, each pulsed therapy cycle comprising (i) a first period in which one or more different drugs are administered in a first dose and (ii) a second period in which one or more different drugs are administered in a second reduced dose (e.g., not administered at all), wherein (A) the first period is characterized by a tumor load detected above a first clinical level and (B) the second period is characterized by a tumor load detected below a second clinical level.

[0023] In one embodiment, the present invention provides a method comprising: (a) sequencing polynucleotides derived from cancer cells of a subject; (b) identifying and quantifying somatic mutations in the polynucleotides; and (c) developing a profile of tumor heterogeneity in the subject for use in determining effective therapeutic interventions for cancer exhibiting tumor heterogeneity, wherein the profile indicates the presence and relative amounts of multiple somatic mutations in the polynucleotides, and different relative amounts indicate tumor heterogeneity.

[0024] In one embodiment, the present invention provides a method comprising the step of providing a therapeutic intervention to a subject, wherein the therapeutic intervention is determined from a profile of disease cellular heterogeneity in the subject, the profile indicating the presence and relative amounts of multiple somatic mutations in polynucleotides, the different relative amounts indicating disease cellular heterogeneity, and the therapeutic intervention is effective for diseases having the determined disease cellular heterogeneity profile, for example, for diseases that show multiple somatic mutations rather than for diseases that show any one of the somatic mutations rather than all of them.

[0025] In some embodiments, a method is provided herein that includes the steps of: a) determining a measured value (e.g., standard deviation, variance) of deviation from the central trend value of copy numbers in polynucleotides in a sample over a region of at least 1 kb, at least 10 kb, at least 100 kb, at least 1 mb, at least 10 mb, or at least 100 mb of the genome; and b) estimating a measured value of the DNA loading from cells undergoing cell division in the sample based on the measured value of the deviation. In some embodiments, the central trend value is the mean, median, or mode. In some embodiments, the determination step includes partitioning the region into a plurality of non-overlapping intervals, determining a measured value of copy numbers in each interval, and determining a measured value of the deviation based on the measured value of copy numbers in each interval. In some embodiments, this interval is 1 nucleotide, 10 nucleotides, 100 nucleotides, 1 kb nucleotide, or 10 kb or less.

[0026] In one aspect, a method for estimating a measurement value of the load of DNA derived from cells undergoing cell division in a sample, the method including the step of measuring copy number variations induced by the proximity of one or more genomic loci to the replication origin of the cell, wherein an increase in CNV indicates a cell undergoing cell division is provided herein. In some embodiments, the load is measured in cell-free DNA. In some embodiments, the measurement value of the load is related to the fraction of tumor cells in the sample or the genomic equivalent of DNA derived from tumor cells. In some embodiments, the CNV due to proximity to the replication origin is estimated from a set of control samples or cell lines. In some embodiments, a hidden Markov model, a regression model, a model based on principal component analysis, or a genotype modification model is used to approximate the variations due to the replication origin. In some embodiments, the measurement value of the load is the presence or absence of cells undergoing cell division. In some embodiments, the proximity is within 1 kb of the replication origin.

[0027] In one aspect, a method is provided herein for increasing the sensitivity and / or specificity of the determination of gene-related copy number variations by alleviating the effect of variations due to proximity to the replication origin. In some embodiments, the method includes measuring the CNV at a locus, determining the amount of CNV due to the locus being in proximity to the replication origin, and correcting the measured CNV to reflect the genomic CNV, for example, by subtracting the amount of CNV that may be due to cell division. In some embodiments, the genomic data is obtained from cell-free DNA. In some embodiments, the measurement value of the load is related to the fraction of tumor cells in the sample or the genomic equivalent of the DNA. In some embodiments, the variations due to the replication origin are estimated from a set of control samples or cell lines. In some embodiments, a hidden Markov model, a regression model, a model based on principal component analysis, or a genotype modification model is used to approximate the variations due to the replication origin.

[0028] In one aspect, a method is provided herein that includes: a) determining a baseline measurement value of the copy of a DNA molecule at one or more loci from one or more control samples, wherein one or more of the loci include a replication origin and each contains DNA from a cell that is undergoing a defined level of cell division; b) determining a test measurement value of the DNA molecule in a test sample, wherein the measurement value in the test sample is from one or more partitioned one or more loci that include a replication origin; and c) comparing the test measurement value and the baseline measurement value, wherein a test measurement value that exceeds the baseline measurement value indicates DNA in a test sample from a cell that is dividing at a faster rate than the cells providing the DNA in the control sample. In some embodiments, the measurement value is selected from a molecule count value, a measurement of the central tendency of the molecule count values across compartments, or a measurement of the variation of the molecule count values across compartments.

[0029] In one aspect, a method is provided herein that includes: (a) administering to a subject an intervention that increases the amount of tumor-derived DNA in the circulation of the subject; and (b) collecting a sample containing tumor-derived DNA from the subject if the amount has increased. In some embodiments, the intervention preferentially kills tumor cells. In some embodiments, the intervention includes exposing the subject or a suspected affected area of the subject to radiation. In some embodiments, the intervention includes exposing the subject or a suspected affected area of the subject to ultrasound. In some embodiments, the intervention includes exposing the subject or a suspected affected area of the subject to physical agitation. In some embodiments, the intervention includes administering a low dose of chemotherapy to the subject. In some embodiments, the method includes administering the intervention to the subject within one week before sample collection. In some embodiments, the sample is selected from blood, plasma, serum, urine, saliva, cerebrospinal fluid, vaginal secretions, mucus, and semen.

[0030] In some embodiments, a method is provided herein that includes the step of compiling a database, wherein the database includes, for each of a plurality of subjects having cancer, tumor genome testing data including somatic changes collected at two or more time intervals per subject, one or more therapeutic interventions administered to each subject at one or more time intervals, and the effectiveness of the therapeutic interventions, and the database is useful for estimating the effectiveness of therapeutic interventions in subjects having a tumor genome profile. In some embodiments, the plurality is at least 50, at least 500, or at least 5000. In some embodiments, tumor genome testing data is collected by serial biopsies, cell-free DNA, cell-free RNA, or circulating tumor cells. In some embodiments, the relative frequency of detected gene variants is used to classify treatment effectiveness. In some embodiments, additional information, including but not limited to body weight, adverse treatment effects, histological examination, blood tests, radiographic information, previous treatments, and cancer type, is used to aid in classifying treatment effectiveness. In some embodiments, treatment response per patient is quantitatively collected and classified by additional tests. In some embodiments, the additional tests are blood or urine-based tests.

[0031] In one embodiment, the present invention provides a method for identifying one or more effective therapeutic interventions for subjects having cancer, wherein the database includes, for each of a plurality of subjects having cancer, tumor genome testing data including somatic changes collected at two or more time intervals per subject, one or more therapeutic interventions administered to each of the subjects at one or more time intervals, and the effectiveness of the therapeutic interventions. In some embodiments, the identified therapeutic interventions are stratified by effectiveness. In some embodiments, quantitative limits regarding the predicted effectiveness or lack thereof of the therapeutic interventions are reported. In some embodiments, the therapeutic interventions use information on predicted tumor genome evolution or acquired resistance mechanisms in similar patients responding to the treatment.

[0032] In some embodiments, this method uses classification algorithms, such as linear regression processes (e.g., multiple linear regression (MLR), partial least squares (PLS) regression, and principal component regression (PCR)), binary decision trees (e.g., CART-classification and regression trees, recursive partitioning processes), and artificial neural networks such as backpropagation networks for discrimination. The process includes a step of classifying the effectiveness of a treatment using analysis (e.g., Bayesian classifier or Fisher analysis), a logistic classifier, and a support vector classifier (e.g., a support vector machine).

[0033] In some embodiments, a method for reporting the results of one or more genetic tests is disclosed herein, comprising the steps of: using a genetic analyzer to capture genetic information including gene variants and their quantitative measurements across one or more test points; normalizing the quantitative measurements for rendering by one or more test points and generating a scale factor; applying the scale factor to render a tumor response map; and generating a summary of the gene variants. In some embodiments, the method includes the step of analyzing non-CNV (copy number variation) variant allele frequencies. In some embodiments, the method includes the step of converting absolute values ​​to relative metric values ​​for rendering a tumor response map. In some embodiments, the method includes the step of multiplying the variant allele frequencies by a default value and taking the logarithm. In some embodiments, the method includes the steps of multiplying each gene by a converted value to a scale factor to determine a quantity indicator to be rendered on the tumor response map; and assigning a specific visual indicator for each change on a visual panel. In some embodiments, the method includes the step of Y-centering or vertically centering quantity indicators in adjacent panels indicating continuity. In some embodiments, the assignment step further includes the step of assigning a specific color to each change.

[0034] In some embodiments, the method includes the step of analyzing genetic information from a different test point or test point. In some embodiments, if a new test result is the same as a previous test result, the method includes the step of rendering the previous visual panel. In some embodiments, if the changes remain the same but the amount changes, the method includes the steps of maintaining the order and a specific visual indicator for each change, determining a new amount indicator, and generating a new visual panel for all test points. In some embodiments, the method includes the step of determining a new change in the genetic information and adding this change to the top of the existing changes. In some embodiments, the method includes the steps of determining a new change in the genetic information, determining a new transformation value and scale factor, and assigning a specific visual indicator for each new change. In some embodiments, the method includes the step of determining a new change in the genetic information and regenerating a tumor response map that includes the changes from the previous test point that are still detected at the current test point, as well as the new changes. In some embodiments, the method includes the step of determining if the previous change no longer exists, and if so, using a zero height when rendering the amount of change for the previous change for subsequent test points. In some embodiments, the method includes the step of determining whether a previous change no longer exists, and if so, the step of retaining a specific visual indicator associated with the previous change for future use.

[0035] In some embodiments, the method includes the step of analyzing CNV mutant allele frequencies and methylation mutant allele frequencies. In some embodiments, the method first includes grouping of maximum mutant allele frequencies for rendering in a tumor response map. In some embodiments, the method includes the step of rendering gene changes in order of decreasing mutant allele frequencies. In some embodiments, the method includes the step of rendering gene changes in decreasing order. In some embodiments, the method includes the step of selecting the next gene having the next highest mutant allele frequency.

[0036] In some embodiments, for each reported change, the method includes a step of generating a trend indicator of the change across different test points. In some embodiments, the method includes a step of generating a summary of the changes. In some embodiments, the method includes a step of generating a summary of treatment options. In some embodiments, the method includes a step of generating a summary of variant allele frequencies, cell-free amplification, clinically approved indications, and clinical trials. In some embodiments, the method includes a step of generating a panel based on biological pathways. In some embodiments, the method includes a step of generating a panel based on evidence levels. In some embodiments, the genetic information includes one or more of single-nucleotide mutations, copy number variations, insertions and deletions, and gene rearrangements. In some embodiments, the method includes a step of generating a clinical relevance report based on the detected changes. In some embodiments, the method includes a step of generating a treatment outcome summary.

[0037] In one embodiment, the present invention provides a method for generating a gene report, comprising the steps of: generating noncopy number variation (CNV) data using a gene analyzer; determining a scale factor for each non-CNV variant allele frequency; generating a visual panel of each non-CNV variation using the scale factor for a first test; and modifying the non-CNV variations in the visual panel using the scale factor for each subsequent test.

[0038] In some embodiments, the method includes a step of converting absolute values ​​to relative metric values ​​for rendering. In some embodiments, the method includes a step of multiplying the mutant allele frequency by a default value and taking the logarithm of the default value. In some embodiments, the method includes a step of determining the scale factor using the maximum observed value. In some embodiments, for each non-CNV variant, the method includes a step of multiplying the scale factor by a converted value for each gene variant as a quantity indicator for visualizing the gene variant.

[0039] In some embodiments, the method includes the step of assigning a specific visual indicator for each change. In some embodiments, if the test results do not change for subsequent tests, the method includes the step of using a visual panel. In some embodiments, if the changes remain the same in subsequent tests, the method includes the steps of maintaining the order and the specific visual indicator for each change, recalculating the quantity indicator to visualize the variant, and re-rendering the updated values ​​in the existing panel(s) and the new panel for the most recent test. In some embodiments, if a new change is found in a subsequent test, the method includes the steps of adding the change to the top of all existing changes, calculating the transformed value and scale factor, and assigning a specific visual indicator for each new change.

[0040] In some embodiments, the method includes the steps of re-rendering changes at previous inspection points and new changes, and vertically centering images of changes in adjacent panels that show continuity. In some embodiments, if no previous changes are present in subsequent inspections, the method includes the step of using a height of zero as the amount of change for subsequent renderings. In some embodiments, the method includes the step of rendering subject or intervention information related to the change. In some embodiments, the method includes the step of identifying changes that have the highest mutant allele frequency.

[0041] In some embodiments, the method includes the steps of reporting the changes in the gene in order of the mutant allele frequency that reduces non-CNV changes, and reporting the CNV changes of the gene in order of the CNV value that decreases. In some embodiments, the method includes the step of selecting the next gene having the next highest non-CNV mutant allele frequency, and including the steps of reporting the changes in the gene in order of the mutant allele frequency that reduces non-CNV changes, and reporting the CNV changes of the gene in order of the CNV value that decreases.

[0042] In some embodiments, the method includes a step of rendering a trend indicator of change over different test dates. In some embodiments, the method includes grouping of maximum variant allele frequencies and generating annotations including biological pathways or evidence levels. In some embodiments, the method includes a step of generating a panel based on evidence levels. In some embodiments, the method includes a step of generating a panel based on biological pathways. In some embodiments, the genetic information includes one or more of single-nucleotide mutations, copy number variations, insertions and deletions, and gene rearrangements.

[0043] In one embodiment, a method is provided herein that includes: a) providing a plurality of nucleic acid samples derived from a subject, wherein the samples are collected at consecutive time points; b) sequencing polynucleotides derived from the samples to construct and generate sequences; c) determining quantitative measurements of each of a plurality of gene variants among the polynucleotides in each sample; and d) graphing the relative amounts of gene variants at each consecutive time point for somatic mutations present at at least one of the consecutive time points using a computer. In some embodiments, the quantitative measurement is the frequency of gene variants among all sequences mapped to the same locus. In some embodiments, the relative amounts are represented as a stacked area graph. In some embodiments, the relative amounts are stacked from the highest to the lowest at the earliest time point, with gene variants that first appear at non-zero amounts at later time points stacked at the top of the graph. In some embodiments, the areas are represented by different colors. In some embodiments, the graphing further shows quantitative measurements of dominant gene variants at each time point. In some embodiments, the graphing further includes key identifying gene variants represented on the graph. In some embodiments, graphing includes normalizing and scaling quantitative measurements.

[0044] In some embodiments, the polynucleotides include cfDNA. In some embodiments, the locus is located in an oncogene. In some embodiments, multiple gene variants are mapped to different genes in the genome. In some embodiments, multiple gene variants are mapped to the same gene in the genome. In some embodiments, at least 10 different oncogenes are sequenced.

[0045] In some embodiments, determining involves receiving an array into computer memory, using a computer processor to run software, and determining quantitative measurements. In some embodiments, graphing involves using a computer processor to run software that converts the quantitative measurements into a graph format, and representing the graph format on an electronic graphical user interface, for example, a display screen.

[0046] In one embodiment, a method for generating a written or electronic patient test report from data generated by a genetic analyzer is provided herein, comprising the steps of: a) summarizing data from two or more test time points so that a combination of all non-zero test results is reported at each of the subsequent test points after a first test; and b) rendering the test results on a written or electronic patient test report. In some embodiments, the summarizing and rendering are performed on a computer by executing code using a computer processor to (i) identify all non-zero test results, (ii) generate a test report, and (iii) display the test report on a graphical user interface.

[0047] In one embodiment, a method for graphically representing the evolution of tumor gene variants in a subject from data generated by a gene analyzer is provided herein, comprising the steps of: a) generating a stacked representation of gene variants detected at each of several time points in the subject using a computer, wherein the height or width of each layer in the stack corresponding to the gene variants represents the quantitative contribution of the gene variants to the total amount of gene variants at each time point; and b) displaying the stacked representation on a computer monitor or a written report. In some embodiments, the method further includes the step of using a combination of gene variant magnitudes detected in fluid-based tests to estimate the disease burden. In some embodiments, the method further includes the step of using allele fractions, allele imbalances, and gene-specific coverage of detected mutations to estimate the disease burden.

[0048] In some embodiments, the overall stack height represents the overall disease burden or disease burden score in the subject. In some embodiments, a distinct color is used to represent each gene variant. In some embodiments, only a subset of detected gene variants is plotted. In some embodiments, the subset is selected based on its association with the likelihood or reduction of response to treatment, which is the driver change.

[0049] In some embodiments, the method includes the step of generating a test report relating to a genomic test. In some embodiments, a nonlinear scale is used to represent the height or width of each represented gene variant. In some embodiments, a plot of previous test points is depicted on the report. In some embodiments, the method includes the step of estimating disease progression or remission based on the rate and / or quantitative accuracy of change in each test result. In some embodiments, the method includes the step of displaying therapeutic interventions between intervening test points. In some embodiments, the display step includes a) receiving data representing detected tumor gene variants into computer memory; b) executing code by a computer processor to graphically display the quantitative contribution of each gene variant at a given time as a line or area proportional to the relative contribution; and c) displaying the graphical display on a graphical user interface. The present invention provides, for example, the following items: (Item 1) (a) A step of sequencing polynucleotides from cancer cells derived from the target biological sample, (b) A step of identifying and quantifying somatic mutations in the polynucleotides, (c) A step of developing a tumor heterogeneity profile in the subject, which shows the presence and relative amounts of multiple somatic mutations in the polynucleotide, wherein different relative amounts indicate tumor heterogeneity, (d) A step of determining a therapeutic intervention for a cancer exhibiting the tumor heterogeneity, wherein the therapeutic intervention is effective for cancer having the determined tumor heterogeneity profile. A method that includes this. (Item 2) The method according to item 1, wherein the sequence readings generated by sequencing are subjected to noise reduction before identification and quantification. (Item 3) The method according to item 1, wherein the step of determining the therapeutic intervention takes into account the relative frequency of tumor-related genetic changes. (Item 4) The method according to item 1, wherein the therapeutic intervention comprises administering a number of drugs in combination or sequentially, each drug being relatively more effective against cancers that present with a different type of somatic mutation occurring at different relative frequencies. (Item 5) The method according to item 4, wherein a drug that is relatively more effective against cancers exhibiting somatic mutations occurring at a higher relative frequency is administered in a larger dose. (Item 6) The method according to item 4, wherein the drug is delivered in stratified doses that reflect the relative amount of the variant in the DNA. (Item 7) The method according to item 4, wherein a cancer exhibiting at least one of the gene variants is resistant to at least one of the drugs. (Item 8) The method according to item 1, wherein the step of determining the therapeutic intervention takes into account the tissue of the origin of the cancer. (Item 9) A method comprising the steps of providing a therapeutic intervention to a subject having cancer having a tumor profile in which tumor heterogeneity can be estimated, wherein the therapeutic intervention is effective for the cancer having the tumor profile. (Item 10) The method according to item 9, wherein the tumor profile shows the relative frequencies of multiple further somatic mutations. (Item 11) The method according to item 10, further comprising the steps of monitoring changes in the relative frequency in the subject over time, and determining different therapeutic interventions over time based on the changes. (Item 12) The method according to item 10, wherein the therapeutic intervention is more effective against cancers that present each of the somatic mutations than against cancers that present any one of the somatic mutations rather than all of them. (Item 13) The method according to item 10, wherein the therapeutic intervention comprises administering multiple drugs in combination or sequentially, and each drug is relatively more effective against cancers that present a different type of somatic mutation occurring at different relative frequencies. (Item 14) A method comprising the step of administering a therapeutic intervention effective against a tumor exhibiting tumor heterogeneity, wherein the therapeutic intervention exhibits different relative amounts of tumor heterogeneity, based on a tumor heterogeneity profile in the subject, which indicates the presence and relative amounts of multiple somatic mutations in polynucleotides. (Item 15) Execution by a computer processor, (a) The step of receiving the sequence reading of a polynucleotide mapped to a gene locus into memory, (b) A step of determining the identity of all number of sequence readings that are mapped to a locus, of which the bases are different from the bases of the reference sequence at the locus, (c) A step of reporting the identity and relative amount of the determined base and its position in the genome. (d)(c) A step to estimate the heterogeneity of a given sample based on the information in (d)(c) and A system including a computer-readable medium containing machine-executable code that implements a method including the following. (Item 16) The method according to item 15, wherein the implemented method further includes the steps of receiving sequence readings from a sample at several different time points into memory, and calculating the difference in relative amounts and identity of several bases between two samples. (Item 17) A kit comprising a first pharmaceutical product and a second pharmaceutical product, wherein the combination of the first drug and the second drug is therapeutically effective against cancers presenting a first and a second somatic mutation rather than against cancers presenting any one of the somatic mutations rather than all of them. (Item 18) The kit according to item 17, wherein the combination is contained in a mixture, or each drug is contained in a separate container. (Item 19) (a) A step of performing biomolecular analysis of biomolecular polymers derived from disease cells originating from the target (e.g., spatially distinct disease cells), (b) a step of identifying and quantifying biomolecular variants in the biopolymer, and (c) a step of developing a profile of disease cell heterogeneity in the subject, showing the presence and relative amounts of a plurality of the variants in the biopolymer, wherein different relative amounts indicate disease cell heterogeneity. (d) A step of determining a therapeutic intervention for a disease exhibiting disease cell heterogeneity, wherein the therapeutic intervention is effective for a disease having the determined profile of disease cell heterogeneity. A method that includes this. (Item 20) A method for detecting disease cell heterogeneity in a target, a) A step of quantifying polynucleotides having sequence variants at each of multiple gene loci in a polynucleotide from a sample derived from the subject, wherein the sample includes polynucleotides derived from somatic cells and disease cells, b) For each gene locus, the step of determining the measured value of the copy number variation (CNV) of the polynucleotide having the sequence variant, c) For each gene locus, the step of determining a weighted measurement of the amount of polynucleotides having a sequence variant at the gene locus as a function of the CNV at the gene locus, d) a method comprising the step of comparing the weighted measurements at each of the plurality of gene loci, wherein different weighted measurements indicate disease cell heterogeneity. (Item 21) a) A step of targeting one or more pulsed therapy cycles, each pulsed therapy cycle comprising (i) a first period in which one or more drugs are administered in a first dose and (ii) a second period in which the one or more drugs are administered in a second reduced dose (e.g., not administered at all), (A) The first period is characterized by a tumor burden detected at a level exceeding the first clinical level, (B) The second period is characterized by a tumor burden detected below a second clinical level. A method that includes this. (Item 22) The method according to item 21, wherein one or more drugs are multiple drugs, and each amount of each drug in each cycle is determined according to the tumor load, measured as a function of the respective amounts of multiple different selected somatic variants in tumor polynucleotides. (Item 23) b) If the subject exhibits resistance to the one or more drugs, the step of subjecting the subject to one or more pulsed therapy cycles, each pulsed therapy cycle comprising: (i) a first period in which one or more different drugs are administered in a first dose; and (ii) a second period in which one or more different drugs are administered in a second reduced dose (e.g., not administered at all). (A) The first period is characterized by a tumor burden detected at a level exceeding the first clinical level, (B) The second period is characterized by a tumor burden detected below a second clinical level. The method described in item 21, further including the method described in item 21. (Item 24) (a) A step of sequencing polynucleotides derived from cancer cells of the target, (b) A step of identifying and quantifying somatic mutations in the polynucleotides, (c) A step of developing a tumor heterogeneity profile in the subject for use in determining effective therapeutic interventions for cancer exhibiting tumor heterogeneity, wherein the profile shows the presence and relative amounts of multiple somatic mutations in the polynucleotide, and different relative amounts indicate tumor heterogeneity. A method that includes this. (Item 25) A method comprising the step of providing a therapeutic intervention to a subject, wherein the therapeutic intervention is determined from a profile of disease cell heterogeneity in the subject, the profile indicating the presence and relative amounts of multiple somatic mutations in polynucleotides, the different relative amounts indicating disease cell heterogeneity, and the therapeutic intervention is effective for diseases having the determined profile of disease cell heterogeneity, for example, for diseases indicating multiple somatic mutations rather than for diseases indicating one of the somatic mutations rather than all of them. (Item 26) a) A step of determining the measured deviation (e.g., standard deviation, variance) of the deviation from the central trend value of copy number in polynucleotides in the sample over a region of at least 1kb, at least 10kb, at least 100kb, at least 1mb, at least 10mb, or at least 100mb of the genome, b) A step of estimating the measured value of the DNA load from cells undergoing cell division in the sample based on the measured value of the deviation, A method that includes this. (Item 27) The method according to item 26, wherein the value of the central tendency is the mean, median, or mode. (Item 28) The method according to item 26, wherein the determining step includes partitioning the region into a plurality of non-overlapping intervals, determining a measured copy count in each interval, and determining a measured deviation based on the measured copy count in each interval. (Item 29) The method according to item 28, wherein the interval is less than or equal to 1 base, 10 bases, 100 bases, 1 kb base, or 10 kb. (Item 30) A method for estimating a measure of the DNA load derived from a cell undergoing cell division in a sample, comprising the step of measuring copy number variation induced by the proximity of one or more genomic loci to the origin of cell replication, wherein an increase in CNV indicates a cell undergoing cell division. (Item 31) A method for increasing the sensitivity and / or specificity of determining gene-related copy number variations by mitigating the effects of variations caused by proximity to the origin of replication. (Item 32) a) A step of determining a baseline measurement of DNA molecule copies at one or more loci from one or more control samples, wherein one or more of the loci include an origin of replication and each contains DNA from a cell undergoing a predetermined level of cell division, b) A step of determining a test measurement of a DNA molecule in a test sample, wherein the measurement in the test sample originates from one or more partitioned loci, and one or more of the loci include an origin of replication, c) A step of comparing the test measurement value with the baseline measurement value, wherein the test measurement value that is higher than the baseline measurement value indicates that the DNA in the test sample is derived from cells that are dividing at a faster rate than the cells that provide DNA to the control sample. A method that includes this. (Item 33) The method according to item 32, wherein the measurement is selected from molecular counts, measurements of the central trend of molecular counts across compartments, or measurements of variation in molecular counts across compartments. (Item 34) (a) The step of administering an intervention to the subject that increases the amount of tumor-derived DNA in the subject's circulation, (b) If the amount increases, the step of collecting a sample containing tumor-derived DNA from the subject and A method that includes this. (Item 35) A method comprising the step of compiling a database, wherein the database includes, for each of a plurality of subjects having cancer, tumor genome testing data including somatic changes collected at two or more time intervals for each subject, one or more therapeutic interventions administered to each of the subjects at one or more time points, and the effectiveness of the therapeutic interventions, wherein the database is useful for estimating the effectiveness of the therapeutic interventions in subjects having tumor genome profiles. (Item 36) A method comprising using a database to identify one or more effective therapeutic interventions for subjects having cancer, wherein the database comprises, for each of a plurality of subjects having cancer, tumor genome testing data including somatic changes collected at two or more time intervals per subject, one or more therapeutic interventions administered to each of the subjects at one or more time points, and the effectiveness of the therapeutic interventions. (Item 37) The method described in item 36, wherein identified therapeutic interventions are stratified by effectiveness. (Item 38) The method described in item 36, which reports quantitative limitations regarding the predicted effectiveness or lack thereof of the therapeutic intervention. (Item 39) The method according to item 36, wherein the therapeutic intervention uses information on predicted tumor genomic evolution or acquired resistance mechanisms in similar patients responding to the treatment. (Item 40) A method for reporting the results of one or more genetic tests, A step of using a gene analyzer to capture genetic information including gene variants and their quantitative measurements across one or more test points, The steps include: normalizing the quantitative measurements for rendering using one or more inspection points and generating a scale factor; The steps include rendering a tumor response map by applying the aforementioned scaling factor, Steps to generate a summary of gene variants and A method that includes this. (Item 41) The method described in item 40, which includes the step of analyzing the frequency of non-CNV (copy number variation) mutant alleles. (Item 42) A method for generating a gene report, The steps include generating non-copy number variation (CNV) data using a gene analyzer, The steps include determining the scale factor for each non-CNV mutant allele frequency, For the first examination, the steps include generating a visual panel of each non-CNV change using the scale factor, For each of the subsequent tests, the steps include: generating changes in the non-CNV changes for the visual panel using the scale factor; A method that includes this. (Item 43) a) A step of providing multiple nucleic acid samples derived from a target, wherein the samples are collected at successive points in time. b) A step of sequencing the polynucleotide derived from the sample, comprising the step of generating the sequence, c) A step of determining the quantitative measurement values ​​of each of the multiple gene variants among the polynucleotides in each sample, d) A step of using a computer to graphically display the relative amounts of gene variants at each of the consecutive time points with respect to somatic mutations that are present in a non-zero amount at at least one of the consecutive time points. A method that includes this. (Item 44) The method according to item 43, wherein the quantitative measurement is the frequency of the gene variant among all sequences mapped to the same locus. (Item 45) The method described in item 43, wherein the aforementioned relative quantity is represented as a stacked area graph. (Item 46) The method according to item 45, wherein the relative amounts are stacked from highest to lowest at the earliest point in time, and the gene variants that first appear in non-zero amounts at later points in time are stacked at the top of the graph. (Item 47) A method for generating a written or electronic patient examination report from data generated by a gene analyzer, a) A step of summarizing data from two or more test time points, wherein the sum of all non-zero test results is reported at each of the subsequent test points after the first test, b) A method comprising the step of rendering the test results on the written or electronic patient test report. (Item 48) The method according to item 47, wherein the summarizing step and rendering step are performed on a computer by executing code using a computer processor to (i) identify all non-zero check results, (ii) generate the check report, and (iii) display the check report on a graphical user interface. (Item 49) A method for graphically displaying the evolution of tumor gene variants in a subject from data generated by a gene analyzer, a) A step of generating a stacked representation of gene variants detected at each of several time points in the subject by computer, wherein the height or width of each layer in the stack corresponding to the gene variant represents the quantitative contribution of the gene variant to the total amount of gene variants at each time point, b) The step of displaying the stacked display on a computer monitor or a written report. A method that includes this. (Item 50) The steps to display are, a) Receiving data representing the detected tumor gene variant into computer memory, b) Executing a code using a computer processor, which graphically displays the quantitative contribution of each gene variant at a given point in time as a line or area proportional to the relative contribution, c) Displaying the graph on a graphical user interface The method described in item 49, including the method described in item 49. [Brief explanation of the drawing]

[0050] [Figure 1] Figure 1 shows a flowchart illustrating an exemplary method for determining and using therapeutic interventions.

[0051] [Figure 2] Figure 2 shows a flowchart illustrating an exemplary method for determining the frequency of variants in a sample corrected based on CNVs at the gene locus.

[0052] [Figure 3] Figure 3 shows a flowchart illustrating an exemplary method for providing pulsed therapy cycles that can delay drug resistance.

[0053] [Figure 4] Figure 4 shows a flowchart illustrating an exemplary method for detecting tumor burden using CNVs at the origin of replication to detect DNA from dividing cells.

[0054] [Figure 5] Figure 5 shows an exemplary computer system.

[0055] [Figure 6]Figure 6 shows an exemplary scan of CNVs across genomic regions from samples containing cells in a quiescent and dividing state. Genomic CNVs are not observed at loci a and b, but locus c shows gene duplication. In quiescent cells, copy numbers are relatively equal across all segments of the region, except for the segment duplicated at the gene duplication locus. In samples containing DNA from tumor cells in cell division, copy numbers appear to increase immediately after the origin of replication, providing variation in CNVs across the region. The deviation is particularly dramatic at the locus showing CNV at the origin of replication (c).

[0056] [Figure 7] Figure 7 shows an exemplary course of monitoring and treatment for the target disease.

[0057] [Figure 8] Figure 8 shows an exemplary panel of 70 genes that exhibit genetic variation in cancer.

[0058] [Figure 9A] Figure 9A shows an exemplary system for communicating cancer test results.

[0059] [Figure 9B] Figure 9B illustrates an exemplary process for reducing the error rate and bias of DNA sequence readings and generating gene reports for the user.

[0060] [Figure 10A] Figures 10A-10C illustrate an exemplary process for reporting genetic test results to the user. [Figure 10B] Figures 10A-10C illustrate an exemplary process for reporting genetic test results to the user. [Figure 10C] Figures 10A-10C illustrate an exemplary process for reporting genetic test results to the user.

[0061] [Figure 10D]Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10D-1] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10D-2] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10D-3] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10E] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10E-1] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10E-2] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10F] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10F-1] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10F-2] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10G] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10G-1] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10G-2] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10H] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10I] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10I-1] Figures 10D-10I-2 show pages from an exemplary genetic testing report. [Figure 10I-2] Figures 10D-10I-2 show pages from an exemplary genetic testing report.

[0062] [Figure 10J] Figures 10J-10P show various exemplary modified stream graphs. [Figure 10K] Figures 10J-10P show various exemplary modified stream graphs. [Figure 10L] Figures 10J-10P show various exemplary modified stream graphs. [Figure 10M] Figures 10J-10P show various exemplary modified stream graphs. [Figure 10N] Figures 10J-10P show various exemplary modified stream graphs. [Figure 10O] Figures 10J-10P show various exemplary modified stream graphs. [Figure 10P] Figures 10J-10P show various exemplary modified stream graphs.

[0063] [Figure 11A] Figures 11A-11B illustrate an exemplary process for detecting mutations and reporting test results to the user. [Figure 11B] Figures 11A-11B illustrate an exemplary process for detecting mutations and reporting test results to the user. [Modes for carrying out the invention]

[0064] (Detailed explanation) The methods of this disclosure can detect biomolecular mosaicism (e.g., genetic mosaicism) in biological samples such as heterogeneous genomic populations of cells or deoxyribonucleic acid (DNA). Genetic mosaicism can exist at the biological level. For example, gene variants that arise early in development can result in different somatic cells with different genomes. An individual may be a chimera, for example, produced by the fusion of two zygotes. Organ transplantation from an allogeneic donor can result in genetic mosaicism, which can also be detected by testing polynucleotides that detach from the transplanted organ into the blood. Disease cell heterogeneity, in which affected cells have different gene variants, is another form of genetic mosaicism. The methods provided herein can detect mosaicism and, in the case of disease, can provide therapeutic interventions. In certain embodiments, this disclosure provides a method for whole-body profiling of biomolecular mosaicism by using circulating polynucleotides derived from, or potentially originating from, cells at diverse locations in the body of the subject.

[0065] Diseased cells, such as tumors, can evolve over time, giving rise to different clonal subpopulations with novel genetic and phenotypic characteristics. This may result from spontaneous mutations during cell division or be driven by treatments targeting specific clonal subpopulations, which make clones more resistant to the treatment and promote their growth through negative selection. The presence of diseased cell subpopulations with different genotypes or phenotypic characteristics is referred herein to as disease cell heterogeneity, or in the case of cancer, as tumor heterogeneity.

[0066] Currently, cancer is treated based on the variant type found in cancer biopsies. For example, the discovery of Her2+ in even small amounts of breast cancer cells can indicate breast cancer, which can then be treated with anti-Her2+ therapies. As another example, colorectal cancer in which small amounts of KRAS variants are found can be treated with therapies that respond to KRAS.

[0067] Tools for the detailed analysis of affected cells (e.g., tumors) enable the detection of disease cellular heterogeneity. Furthermore, analysis of polynucleotides supplied from affected cells located throughout the body allows for a systemic profile of disease cellular heterogeneity. The use of cell-free DNA or circulating DNA is particularly powerful because polynucleotides in the blood are not supplied from physically localized cells. Rather, these include cells originating from metastatic sites throughout the body. For example, analysis can show that a population of breast cancer cells contains 90% Her2+ and 10% Her2-. This can be determined, for example, by quantifying DNA for each morphology in the sample, e.g., cell-free DNA (cfDNA), and thereby detecting heterogeneity in the tumor.

[0068] Healthcare providers, such as physicians, can use this information to develop therapeutic interventions. For example, a subject with heterogeneous tumors can be treated as if they had two types of tumors, and the therapeutic intervention can treat each of the tumors. The therapeutic intervention may include, for example, a combination therapy involving a first drug effective against the first tumor type and a second drug effective against the second tumor type. The drugs may be given in amounts that reflect the relative amounts of the detected variant types. For example, a drug to treat a variant found in a larger relative amount can be delivered in a larger dose than a drug to treat a variant type found in a smaller relative amount. Alternatively, treatment for a variant in a smaller relative amount can be delayed or staggered compared to a variant in a larger amount.

[0069] Monitoring changes in disease cell heterogeneity profiles over time allows for the calibration of therapeutic interventions for evolving tumors. For example, analysis may reveal an increasing amount of polynucleotides in tumors with drug-resistant variants. In this case, therapeutic interventions can be modified to reduce the amount of drug effective in treating tumors without resistant variants and increase the dose of drug effective in treating tumors with resistance markers.

[0070] Therapeutic interventions may be determined by healthcare providers, computer algorithms, or a combination of both. The database may contain the results of therapeutic interventions for diseases with various profiles of disease cell heterogeneity. The database can be consulted in determining therapeutic interventions for diseases with specific profiles.

[0071] This disclosure provides, in particular, a method for determining therapeutic interventions for subjects with diseases such as cancer, exhibiting disease cell heterogeneity, e.g., tumor heterogeneity. In one embodiment, the method involves a step of analyzing the biological macromolecules (e.g., a step of sequencing polynucleotides) of disease cells (e.g., spatially distinct disease cells) derived from a subject with the disease. A disease cell heterogeneity profile is constructed, showing the presence of gene variants specific to the disease cells and the amounts of these variants relative to each other. This information is subsequently used to determine therapeutic interventions that take this profile into account.

[0072] diseased cells The methods described herein are applicable to any multicellular organism. More specifically, the applications may be plants or animals, vertebrates, mammals, mice, primates, monkeys, or humans. Examples of animals include, but are not limited to, livestock, sport animals, and pets. The subjects may be healthy individuals, individuals with or suspected of having a disease or predisposition to a disease, or individuals who require or are suspected of requiring treatment. The subjects may also be patients, for example, individuals under the care of a professional healthcare provider.

[0073] The subjects may exhibit pathological conditions (diseases). Cells exhibiting disease pathology are referred to as disease cells in this specification.

[0074] In particular, the disease may be cancer. Cancer is a condition characterized by abnormal cells that divide uncontrollably. Cancer includes, without limitation, carcinomas, sarcomas, leukemias, lymphomas, myelomas, and central nervous system cancers. More specific examples of cancer include breast cancer, prostate cancer, colorectal cancer, brain cancer, esophageal cancer, head and neck cancer, bladder cancer, gynecological cancers, liposarcomas, and multiple myelomas.

[0075] Other cancers include, for example, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), adrenocortical carcinoma, Kaposi's sarcoma, anal cancer, basal cell carcinoma, cholangiocarcinoma, bladder cancer, bone cancer, osteosarcoma, malignant fibrous histiocytoma, brainstem glioma, brain cancer, craniopharyngioma, ependymoblastoma, ependymodium, medulloblastoma, medullopeptithelioma, pineal parenchymal cell tumor, breast cancer, bronchial tumor, and Burkitt's cancer. Lymphoma, non-Hodgkin lymphoma, carcinoid tumor, cervical cancer, chordoma, chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), colon cancer, colorectal cancer, cutaneous T-cell lymphoma, intraductal carcinoma in situ, endometrial cancer, esophageal cancer, Ewing's sarcoma, eye cancer, intraocular melanoma, retinoblastoma, fibrous histiocytoma, gallbladder cancer, gastric cancer, glioma, hairy cell leukemia, head and neck cancer, heart cancer, hepatocellular carcinoma (liver) cancer, Hodgkin lymphoma, hypopharyngeal cancer, kidney cancer, laryngeal cancer, lip cancer, oral cancer, lung cancer, non-small cell carcinoma, small cell carcinoma, melanoma, oral cancer, myelodysplasia This includes adult syndromes, multiple myeloma, medulloblastoma, nasal cavity cancer, paranasal sinus cancer, neuroblastoma, nasopharyngeal cancer, oral cancer, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, papilloma, paraganglioma, parathyroid cancer, penile cancer, pharyngeal cancer, pituitary tumor, plasma cell neoplasm, prostate cancer, rectal cancer, renal cell carcinoma, rhabdomyosarcoma, salivary gland cancer, Sézary syndrome, skin cancer, non-melanoma, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, testicular cancer, pharyngeal cancer, thymoma, thyroid cancer, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenström macroglobulinemia, and / or Wilms' tumor.

[0076] A tumor is an aggregate of cancer cells (cancer disease cells). This includes, for example, an aggregate of cells within a single cell mass (e.g., a solid tumor), an aggregate of cells originating from different metastatic tumor sites (metastatic tumors), and a dispersed tumor (e.g., circulating tumor cells). A tumor can contain cells from a single cancer (e.g., colorectal cancer) or multiple cancers (e.g., colorectal cancer and pancreatic cancer). A tumor can contain cells originating from a single somatic cell or cells originating from different somatic cells.

[0077] In certain embodiments, disease cells in a subject are spatially distinct. Disease cells are spatially distinct if they are located within the body, for example, in different tissues or organs, or within the same tissue or organ, at least 1 cm, at least 2 cm, at least 5 cm, or at least 10 cm apart. In the case of cancer, examples of spatially distinct cancer cells include cancer cells originating from a spread cancer (such as leukemia), cancer cells in different metastatic sites, and cancer cells originating from the same tumor cell cluster separated by at least 1 cm.

[0078] Disease cell load (e.g., "tumor load") is a quantitative measure of the amount of disease cells in a subject. One measure of disease cell load is a disease biological macromolecule, a fraction of the total biological macromolecule in a sample, e.g., the relative amount of tumor polynucleotides in a cell-free polynucleotide sample. For example, if cfDNA from a first subject contains 10% cancer polynucleotides, this subject can be said to have a 10% cell-free tumor load. If cfDNA from a second subject contains 5% cancer polynucleotides, the second subject can be said to have half the cell-free tumor load of the first subject. Because the cell-free tumor load in one individual can be much higher or lower than that in another individual, despite the different levels of disease load, these measures are far more relevant on an intra-subject basis than on an inter-subject basis. However, these measures can be very effective in monitoring disease load within an individual; for example, an increase in cell-free DNA tumor load from 5% to 15% may indicate significant disease progression, while a decrease from 10% to 1% may indicate a partial response to treatment.

[0079] Polynucleotides to be sequenced can be sourced from spatially distinct sites. This includes polynucleotides sourced from biopsies at different locations within a single tumor mass. It also includes polynucleotides sourced from cells at different metastatic tumor sites. Cells shed polynucleotides into the bloodstream, where they become detectable as cell-free polynucleotides (e.g., circulating tumor DNA). Cell-free polynucleotides can also be found in other bodily fluids, such as urine. Therefore, cfDNA provides a more accurate profile of tumor heterogeneity across the entire diseased cell population than DNA sourced from a single tumor site. DNA sampled from cells across a diseased cell population in the body is referred to as “disease-loaded DNA,” or in the case of cancer, “tumor-loaded DNA.”

[0080] Disease cells, such as tumors, may share the same or similar biomolecular profiles. For example, tumors may share one, two, three, or more gene variants. Such variants may share the same stratification, e.g., highest frequency, second highest frequency, etc. The profiles may also share similar disease cell loads, e.g., cfDNA loads of 15%, 10%, 5%, or 2% or less.

[0081] Analyte As used herein, a polymer is a molecule formed from monomeric subunits. Monomeric subunits that form a biological polymer include, for example, nucleotides, amino acids, monosaccharides, and fatty acids. Biological polymers include, for example, biopolymers and nonpolymeric polymers.

[0082] Polynucleotides are macromolecules containing polymers of nucleotides. Examples of polynucleotides include polydeoxyribonucleotides (DNA) and polyribonucleotides (RNA). Polypeptides are macromolecules containing polymers of amino acids. Polysaccharides are macromolecules containing polymers of monosaccharides. Lipids are a diverse group of organic compounds, including fats, oils, and hormones, that share the functional characteristic of not interacting with water to a recognizable degree. For example, triglycerides are fats formed from three fatty acid chains.

[0083] Cancer polynucleotides (e.g., cancer DNA) are polynucleotides (e.g., DNA) derived from cancer cells. Cancer DNA and / or RNA can be extracted from tumors, isolated cancer cells, or biological fluids (e.g., saliva, serum, blood, or urine) in the form of cell-free DNA (cfDNA) or cell-free RNA.

[0084] Cell-free DNA is DNA located outside of cells in bodily fluids, such as blood or urine. Circulating nucleic acid (CNA) is nucleic acid found in the bloodstream. Cell-free DNA in the blood is a form of circulating nucleic acid. Cell-free DNA is thought to originate from dying cells that shed their DNA into the bloodstream. Since spatially distinct cancer cells will shed their DNA into bodily fluids such as blood, cancer-derived cfDNA typically contains cancer DNA originating from spatially distinct cancer cells.

[0085] Biological samples The analytes for analysis in the methods of this disclosure may be derived from biological samples, such as samples containing biological macromolecules. Biological samples may be derived from any organ, tissue, or bodily fluid. Biological samples may include, for example, bodily fluids or solid tissue samples. An example of a solid tissue sample is, for example, a tumor sample derived from a solid tumor biopsy. Bodily fluids include, for example, blood, serum, tumor cells, saliva, urine, lymph, prostatic fluid, seminal plasma, milk, sputum, feces, and tears. Bodily fluids are particularly excellent sources of biological macromolecules derived from spatially distinct disease cells, because such cells from many locations in the body can shed such molecules into the bodily fluids. For example, blood and urine are excellent sources of cell-free polynucleotides. Macromolecules derived from such sources can provide a more accurate profile of affected cells than macromolecules derived from localized disease cell clusters.

[0086] The amount of disease polynucleotides in a body fluid sample may increase. Such an increase can increase the detection sensitivity of disease polynucleotides. In one method, an intervention, such as a therapeutic intervention, is administered to the subject, which lyses diseased cells and releases their DNA into the surrounding fluid. Such an intervention may include the administration of chemotherapy. This may also include the step of administering radiation or ultrasound to the subject's whole body or to a part of the subject's body, such as directed to a tumor or affected organ. After the administration of the intervention, and if the amount of disease polynucleotides in the fluid has increased, a fluid sample is collected for analysis. The interval between the administration of the intervention and the collection can be long enough for the disease polynucleotides to increase, but not long enough for them to be eliminated from the body. For example, a low dose of chemotherapy may be administered about one week before sample collection.

[0087] Analysis method This disclosure envisions several types of biomolecular analysis, including, for example, genomics, epigenetics (e.g., methylation), RNA expression, and proteomics. Genomic analysis can be performed, for example, by a gene analyzer using DNA sequencing. Methylation analysis can be performed, for example, by methylated base conversion followed by DNA sequencing. RNA expression analysis can be performed, for example, by polynucleotide array hybridization. Proteomics analysis can be performed, for example, by mass spectrometry.

[0088] As used herein, the term “gene analyzer” refers to a system comprising a DNA sequencer for generating DNA sequence information and a computer containing software for performing bioinformatics analysis on the DNA sequence information. Bioinformatics analysis may include, but is not limited to, the assembly of sequence data, and the detection and quantification of gene variants in a sample, including germline variants (e.g., heterozygosity) and somatic variants (e.g., cancer cell variants).

[0089] The analytical method may include the generation and capture of genetic information. The genetic information may include gene sequence information, ploidy status, identity of one or more gene variants, and quantitative measurements of variants. The term “quantitative measurement” refers to any measurement of a quantity, including absolute and relative measurements. Quantitative measurements may be, for example, numbers (e.g., counts), percentages, frequencies, degrees, or threshold quantities.

[0090] Polynucleotides can be analyzed by any method known in the art. Typically, DNA sequencers are used, or next-generation sequencing (e.g., Illumina, 454, Ion torrent, SOLiD). Sequence analysis can be performed by large-scale parallel sequencing, which simultaneously (or rapidly and sequentially) sequences at least 100,000, million, 10 million, 100 million, or billion polynucleotide molecules. Examples of sequencing methods include, but are not limited to, high-throughput sequencing, pyrosequencing, synthetic sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, RNA-Seq (Illumina), digital gene expression (Helicos), next-generation sequencing, synthetic single-molecule sequencing (SMSS) (Helicos), large-scale parallel sequencing, clonal single-molecule array (Solexa), shot cancer sequencing, Maxam-Gilbert or Sanger sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent, Genius (GenapSys), or nanopore (e.g., Oxford Nanopore) platforms, and any other sequencing method known in the art.

[0091] DNA sequencers can employ either Gilbert's sequencing method, which is based on the chemical modification of DNA and subsequent cleavage of specific bases, or Sanger's technique, which is based on dideoxynucleotide chain termination. Sanger's method has become common due to its increased efficiency and low radioactivity. DNA sequencers can use techniques that do not require DNA amplification (polymerase chain reaction - PCR), thereby accelerating sample preparation before sequencing and reducing errors. In addition, sequencing data is collected from reactions caused by the addition of nucleotides to the complementary strand in real time. For example, DNA sequencers can utilize a method called single-molecule real-time sequencing (SMRT), in which sequencing data is generated by light emitted (captured by a camera) when nucleotides are added to the complementary strand by an enzyme containing a fluorescent dye.

[0092] Genome sequencing may be selective for a specific portion of the genome of interest, for example, for that portion. For instance, many genes (and variants of these genes) are known to be associated with various cancers. Sequencing of a selected gene or portion of a gene may be sufficient for the desired analysis. Polynucleotides mapped to specific loci in the genome of interest can be isolated for sequencing, for example, by sequence capture or site-directed amplification.

[0093] A nucleotide sequence (e.g., a DNA sequence) can refer to a raw sequence reading or a processed sequence reading such as a specific molecular count estimated from a raw sequence reading.

[0094] The sequence readings generated from sequencing are subjected to analysis, for example, that includes identifying gene variants. This may include identifying sequence variants and quantifying the number of base calls at each locus. Quantification may involve, for example, counting the number of readings that map to a particular locus. Different numbers of readings at different loci may indicate copy number variation (CNV).

[0095] Sequencing and bioinformatics methods that reduce noise and distortion are particularly useful when the number of target polynucleotides in a sample is small compared to non-target polynucleotides. When the number of target molecules is small, the signal derived from the target may be weak. This can be a problem, for example, in cell-free DNA where a small number of tumor polynucleotides may be mixed with a much larger number of polynucleotides from healthy cells. Molecular tracking methods can be useful in such situations. Molecular tracking involves tracing sequence readings back from a sequencing protocol to the original sample molecule (e.g., before amplification and / or sequencing) from which the readings originate. Certain methods involve tagging molecules in a way that allows multiple sequence readings generated from the original molecule to be grouped into families of sequences derived from the original molecule. In this way, base calls representing noise can be filtered out. Such methods are, for example, WO2013 / 142389 (Schmitt et al.), US2014 / 0227705 (Vogelstein et al.), and WO2014 / 1491. This is described in more detail in 34 (Talasaz et al.). Up-sampling methods are also useful for more accurately determining the count value of molecules in a sample. In some embodiments, the up-sampling method involves the steps of determining quantitative measurements for individual DNA molecules in which both strands (Watson-Crick strands) are detected; determining quantitative measurements for individual DNA molecules in which only one DNA strand is detected; estimating quantitative measurements for individual DNA molecules in which neither strand was detected from these measurements; and using these measurements to determine quantitative measurements representing a number of individual double-stranded DNA molecules in the sample. This method is described in more detail in PCT / US2014 / 072383, filed on 24 December 2014.

[0096] Genetic variant The methods described herein can be used in the detection of gene variants (also referred to as “genetic changes”). Gene variants are alternative forms at a gene locus. In the human genome, approximately 0.1% of nucleotide positions are polymorphic, i.e., exist as a second genetic form occurring in at least 1% of the population. Mutations can introduce gene variants into the germline and into diseased cells such as cancer cells. Reference sequences, such as hg19 or NCBI Build 37 or Build 38, are intended to represent a “wild-type” or “normal” genome. However, to the extent that they have a single sequence, they do not distinguish common polymorphisms that may be considered normal.

[0097] Genetic variants include sequence variants, copy number variants, and nucleotide modification variants. Sequence variants are variations in the gene nucleotide sequence. Copy number variants are deviations from the wild type in the copy number of a portion of the genome. Genetic variants include, for example, single nucleotide mutations (SNPs), insertions, deletions, inversions, transversions, transpositions, gene fusions, chromosome fusions, gene truncations, copy number variations (e.g., aneuploidy, partial aneuploidy, polyploidy, gene amplification), abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid methylation.

[0098] Genetic variants can be detected by comparing the polynucleotide-derived sequence in a sample with a reference, such as a reference genome sequence, an index, or a database of known variants. In one embodiment, the reference sequence is a publicly available reference sequence, such as the human genome sequence HG-19 or NCBI Build 37. In another embodiment, the reference sequence is a sequence in a non-public database. In yet another embodiment, the reference sequence is the germline sequence of an organism, which is estimated or determined from the sequencing of polynucleotides from that organism.

[0099] Somatic mutations, or somatic alterations, are gene variants that occur in somatic cells. Somatic mutations are distinguished from mutations that occur in the genome of an individual's germline cells (i.e., sperm or egg) or zygote. Somatic mutations, such as those found in cancer cells, can be distinguished from the germline genome of the target from which the cancer originated. These can also be detected by comparing the cancer genome with the germline genome or a reference genome. There are also known gene variants common to cancer cells. A database of SNVs in human cancers can be found on the website:cancer.sanger.ac.uk / cancergenome / projects / cosmic / .

[0100] Figure 8 shows genes known to exhibit point mutations, amplification, fusion, and indels (insertions and deletions) in cancer.

[0101] CNV deviation in rapidly dividing cells During the S phase of the cell cycle, cells replicate their DNA. Diploid cells with replicated 2N chromosomes can correspond to approximately 4X DNA content, while diploid cells with unreplicated 2N chromosomes can correspond to approximately 2X DNA content. Replication proceeds from the origin of replication. In mammals, origins of replication are located at intervals of approximately 15kb to 300kb. During this period, parts of the genome exist in polyploid form. Regions between the origin of replication and polymerase positions overlap, but regions beyond the polymerase positions (or immediately before the origin of replication) still have a single copy number in the strand undergoing replication. When scanning across the genome, the copy number appears uneven or skewed, with regions existing in polyploid form and regions existing in diploid form. Such scanning appears noisy. This is because, in a quiescent state, This is true even for cells whose genomes lack copy number variation. In contrast, scanning for CNVs in G0 phase cells shows a relatively uniform or unstrained copy number profile across the genome. Because cancer cells divide rapidly, their CNV profile across the genome is skewed, regardless of whether the genome also contains CNVs at a particular locus.

[0102] This fact can be used to detect tumor burden in DNA derived from samples containing heterogeneous DNA, such as cfDNA, e.g., a mixture of diseased DNA and healthy DNA. One method for detecting tumor burden involves determining copy number variations resulting from the proximity of a locus(s) under test to various origins of replication. Regions containing origins of replication may have very close numbers of 4 copies of DNA at the locus (in diploid cells), while regions far from origins of replication may have close numbers of 2 copies (in diploid cells). In certain embodiments, the locus(s) under test include at least 1 kb, at least 10 kb, at least 100 kb, at least 1 mb, at least 100 mb, or at least 100 mb across an entire chromosome or genome. A measure of origin of replication CNV (ROCNV) across this region is determined. This may be, for example, a measure of the deviation in copy number from a central tendency value. The central tendency value may be, for example, the mean, median, or mode. The measure of deviation may be, for example, the variance or standard deviation. This measurement can be compared, for example, to ROCNV measurements over the same region in a control sample derived from a healthy individual or quiescent cells. ROCNV can be determined by partitioning the region(s) being analyzed into non-overlapping partitions of varying lengths and adopting CNV measurements in these partitions. This CNV measurement can derive from the number of readings or fragments that are determined to map to these regions after sequencing. Partitions can be at various levels of resolution, e.g., single-base level (base-per-base) DNA can have various sizes to produce 10 nucleotides, 100 nucleotides, 1 kb, 10 kb, or 100 kb. A deviation greater than the control indicates the presence of replicating DNA, which in turn indicates malignancy. The greater the degree of deviation, the greater the amount of DNA from cells undergoing cell division in the sample.

[0103] Various methods can be used to calculate true gene copy number variation, distinct from distortions based on origin of replication. For example, copy number variation can be estimated by calculating deviations from 50% or allelic imbalances at affected CNV loci using heterozygous SNP positions at those loci. Since both copies are generally copied at similar time intervals and are thus self-normalizing (allelic changes may alter the replication of the origin between the two allelic variants), distortions due to proximity of replication origins should not affect this imbalance. For example, duplication of chromosomal segments containing SNPs can be detected at around 67% of readings, while duplication due to ROCNVs will be detected at around 50% of readings. In other methods, counting-based techniques that use the density of detected fragments or readings at a particular locus are used to calculate relative copy numbers. These techniques are generally limited by Poisson noise and systematic bias due to DNA sample preparation and sequencing biases. Combinations of these methods should also yield even better accuracy.

[0104] ROCNV can be calculated for a given sample and can be used to obtain values ​​for cell-free tumor burden despite the lack of detection of traditional somatic variants such as SNVs, gene-specific CNVs, genomic rearrangements, epigenetic variants, and loss of heterozygosity. ROCNV can also be used to increase the sensitivity and / or specificity of a given CNV detection / estimation method by subtracting the distortion of a given sample and removing variations related to origin of replication proximity rather than those caused by true copy number changes in cells. Cell lines with or without known copy number changes across reference can also be used as references for ROCNV for use in estimating their contribution to a given sample.

[0105] In one embodiment, the method involves determining a baseline level of DNA molecule copies at one or more loci derived from one or more control samples, each containing DNA from cells undergoing a predetermined level of cell division, e.g., quiescent cells or rapidly dividing tumor cells. Measurements of DNA molecule copies in the test sample are also determined. Measurements in the test sample may originate from one or more partitioned partitions. In each case, each of the loci contains an origin of replication. The copy measurements from the test sample may be the mean across all partitions or the variance across loci. Measurements of central tendency or variation in copy number (e.g., variance or standard deviation) in the test sample are compared to those in the control sample. Measurements in the test sample that are greater than those in the control (quiescent or slowly dividing cells) indicate that the cells producing DNA in the test sample are dividing more rapidly than the cells supplying DNA to the control sample, e.g., are cancerous. Similarly, similar measurements between test samples and controls of actively dividing cells indicate that the DNA-producing cells in the test sample are dividing at a similar rate to rapidly dividing cells, for example, are cancerous.

[0106] Disease cell heterogeneity Disease cell heterogeneity, such as tumor heterogeneity, is the presence of affected cells with different gene variants. Disease cell heterogeneity can be determined by examining polynucleotides isolated from affected cells and detecting differences in their genomes. Disease cell heterogeneity can also be estimated from examining polynucleotides from samples containing polynucleotides from both affected and healthy cells, based on differences in the relative frequency of somatic mutations. For example, cancer is characterized by changes at the genetic level, for example, due to the accumulation of somatic mutations in different clonal populations of cells. Such changes can contribute to the unregulated growth of cancer cells or serve as markers of responsiveness or unresponsiveness to various therapeutic interventions.

[0107] Tumor heterogeneity is a tumor condition characterized by cancer cells containing different combinations of gene variants, such as different combinations of somatic mutations. That is, a tumor can have different cells containing alterations in different genes, or different alterations in the same gene. For example, a first cell might contain a variant of BRAF, while a second cell might contain variants of both BRAF and ERBB2. Or, a first cancer cell might contain the mononucleotide polymorphism EGRF 55249063 G>A, while a second cell might contain the mononucleotide polymorphism EGRF 55238874 T>A (the numbers refer to nucleotide positions in the genomic reference sequence).

[0108] For example, the initial tumor cells may contain gene variants, such as those in oncogenes. As the cells continue to divide, some of the progeny that carry the initial mutation may independently develop gene variants in other genes or in different parts of the same gene. In subsequent divisions, the tumor cells may accumulate even more gene variants.

[0109] Disease heterogeneity profile The methods of this disclosure enable quantitative and qualitative profiling of disease mosaicism, such as tumor heterogeneity. In one embodiment, the profile includes information derived from polynucleotides from spatially distinct disease cells. In one embodiment, the profile is a whole-body profile containing information derived from cells distributed throughout the body. Analysis of polynucleotides in cfDNA allows for sampling of DNA across the entire geographical range of a tumor, in contrast to sampling of a localized area of ​​the tumor. In particular, this enables sampling of diffuse and metastatic tumors. This contrasts with methods that merely detect the presence of tumor heterogeneity by sampling a localized area of ​​the tumor. The profile may show the precise nucleotide sequence of a variant or simply show a gene with somatic mutations.

[0110] In one embodiment of disease cell heterogeneity profiles, such as tumor cell heterogeneity, the profile identifies gene variations and the relative amounts of each variant. From this information, possible distributions of variants in different cell subpopulations can be estimated. For example, cancer may begin with cells having somatic mutation X. As a result of clonal evolution, some progeny of these cells may develop variant Y. Other progeny may develop variant Z. At the cellular level, after analysis, the tumor can be characterized as 50%X, 35%XY, and 15%XZ. At the DNA level (and considering only DNA derived from tumor cells), the profile may show 100%X, 35%Y, and 15%Z. Both CNVs at the first locus and sequence variants at the second locus can be detected.

[0111] Tumor heterogeneity can be detected from the analysis of cancer polynucleotide sequences, based on the presence of genomic variations at different loci occurring at different frequencies. For example, in a cell-free DNA sample (which may contain germline DNA and cancer DNA), it may be found that the BRAF sequence variant occurs at a frequency of 17%, the CDKN2A sequence variant at 6%, the ERBB2 sequence variant at 3%, and the ATM sequence variant at 1%. These different frequencies of sequence variants indicate tumor heterogeneity. Similarly, gene sequences showing different amounts of copy number variation also indicate tumor heterogeneity. For example, analysis of a sample may show different levels of amplification for the EGFR and CCNE1 genes. This also indicates tumor heterogeneity.

[0112] In the case of cell-free DNA, somatic mutations can be detected by comparing base calls in the sample to a reference sequence, or internally, to more common base calls presumed to be present in the germline sequence as lower-frequency base calls. In either case, the presence of sub-dominant morphs (e.g., less than 40% of total base calls) at different loci at different frequencies indicates disease cellular heterogeneity.

[0113] Cell-free DNA typically contains a dominance of DNA from normal cells with germline genome sequences, and in the case of diseases such as cancer, it contains a small percentage of DNA from cancer cells with cancer genome sequences. The sequence generated from polynucleotides in a cfDNA sample can be compared to the reference sequence to detect differences between the polynucleotides in the reference sequence and the cfDNA. At any gene locus, all or almost all polynucleotides from the test sample may be identical to the nucleotides in the reference sequence. Alternatively, a nucleotide detected at almost 100% frequency in the sample may differ from the nucleotides in the reference sequence. This is most likely to indicate a normal polymorphic morphology at that locus. If a first nucleotide matching the reference nucleotide is detected at approximately 50% and a second nucleotide different from the reference nucleotide is detected at approximately 50%, this is most likely to indicate normal heterozygosity. Heterozygosity can exist at allel ratios different from 50:50, for example, 60:40 or even 70:30. However, if a sample contains detectable nucleotides above the noise at a frequency clearly below (or above) the heterozygous range (e.g., less than 45%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5%), this may be due to the presence of somatic mutations in a certain percentage of cells that contribute the DNA to the cfDNA population. This may originate from diseased cells, e.g., cancer cells (the exact percentage is a function of the tumor burden). If the frequencies of somatic mutations differ at two different loci, e.g., 16% at one locus and 5% at the other, this indicates heterogeneity in diseased cells, e.g., cancer cells.

[0114] In the case of DNA derived from solid tumors, which are expected to predominantly contain tumor DNA, somatic mutations can also be detected by comparison with a reference sequence. Detecting somatic mutations present in 100% of tumor cells may require referencing standard sequences or information on known variants. However, the presence of semi-dominant sequences among the polynucleotide pool at different relative frequencies at different loci indicates tumor heterogeneity.

[0115] The profile may include gene variants in genes for which therapeutic agents are known to be actionable. Knowledge of such variants can contribute to the selection of therapeutic interventions, as therapeutic agents can target them. In the case of cancer, many therapeutic agents are already known for gene variants.

[0116] CNV and SNV in disease cell heterogeneity Generally, the copy number status of a gene should be reflected in the frequency of the gene's morphology in the sample. For example, sequence variants can be detected at frequencies consistent with homozygosity or heterozygosity (e.g., about 100% or about 50%, respectively) without copy number variation. This is consistent with germline polymorphism or mutation. Sequence variants can also be detected at a frequency of about 67% (or about 33%) of polynucleotides at a locus, and in genes measured with increased copy number (generally n=2). This is consistent with germline gene duplication. For example, trisomy may exist in this manner. However, if sequence variants are detected at levels consistent with homozygosity (e.g., about 100%), but in amounts consistent with copy number variation, this is likely to reflect the presence of disease cell polynucleotides with gene amplification. Similarly, if sequence variants are detected at levels not consistent with heterozygosity (e.g., somewhat deviating from 50%), but in amounts consistent with copy number variation, this too is likely to reflect the presence of diseased cellular polynucleotides; affected polynucleotides result in a level of imbalance with allel frequencies deviating from 50:50.

[0117] This observation can be used to estimate whether a sequence variant is likely to exist at the germline level or, for example, is likely to be due to somatic mutations in cancer cells. For instance, a sequence variant in a gene detected at a level almost certainly consistent with germline heterozygosity is more certainly a product of somatic mutations in diseased cells if copy number variations are also detected in the same gene.

[0118] Furthermore, the detection of gene amplification with a sequence variant dose that significantly deviates from the expected amount (e.g., approximately 67% for trisomy at a gene locus) to the extent that the applicants expect germline gene duplication to have variants consistent with increased gene dose (e.g., approximately 67%) indicates that CNVs are likely to be present as a result of somatic mutations.

[0119] Tumor heterogeneity can also be estimated using the fact that somatic mutations at different loci can exist in single or multiple copy numbers within the same diseased cell. More specifically, tumor heterogeneity can be estimated if two genes are detected at different frequencies but their copy numbers are relatively equal. Alternatively, tumor heterogeneity can be estimated if the difference in frequency between two sequence variants coincides with the difference in copy numbers for those two genes. Thus, if EGFR variants are detected at 11% and KRAS variants at 5%, and no CNVs are detected in these genes, the difference in frequency may reflect tumor heterogeneity (e.g., all tumor cells carry the EGFR variant, and half of the tumor cells also carry the KRAS variant). Alternatively, if the EGFR gene carrying the variant is detected at an increased copy number, one consistent interpretation is a homogeneous population of tumor cells where each cell carries variants in both the EGFR and KRAS genes, but the KRAS gene is duplicated. Therefore, both the frequency of sequence variants in the sample and the measured CNV at the locus of the sequence variant can be determined. Next, the frequency can be corrected to reflect the relative number of cells containing the variant by weighting it based on the dose per cell determined from the CNV measurements. This result is then further comparable to sequence variants with no copy number variation in terms of the number of cells containing the variant.

[0120] Communication of test results The results of gene variant analysis (e.g., sequence variants, CNVs, disease cell heterogeneity, and combinations thereof) can be reported by a report generator to healthcare professionals, such as physicians, to assist in the interpretation of test results (e.g., data) and the selection of treatment options. The reports generated by the report generator can provide additional information, such as clinical laboratory results, which may be useful in the diagnosis of disease and the selection of treatment options.

[0121] Referring here to Figure 9A, a system is schematically illustrated that includes a report generation device 1 for reporting, for example, cancer test results and the treatment options derived therefrom. The report generation system may be a central data processing system configured to establish direct communication via a communication link with a remote data site or lab 2, a medical practice / healthcare provider (a specialist in the procedure) 4 and / or a patient / subject 6. Lab 2 may be a medical laboratory, diagnostic laboratory, medical equipment, medical practice, point-of-care testing device, or any other remote data site capable of generating subject clinical information. Subject clinical information may include, but is not limited to, laboratory test data, e.g., genetic variant analysis; imaging and X-ray data; test results; and diagnoses. The healthcare provider or procedure 6 may include medical service providers such as physicians, nurses, home caregivers, technicians, and physician's assistants, and the procedure may be any medical care facility to which the healthcare provider is assigned. In certain specific cases, the healthcare provider / procedure is also a remote data site. If cancer is the disease to be treated, the subject may have cancer among other possible diseases or disorders.

[0122] Other clinical information for cancer subjects6 may include laboratory tests, e.g., results of analyses such as gene variants, metabolic panels, and complete blood counts; medical imaging data; and / or medical procedures targeting the diagnosis of a condition, provision of prognosis, monitoring of disease progression, determination of relapse or remission, or a combination thereof. A list of appropriate sources of clinical information for cancer includes, but is not limited to, CT scans, MRI scans, ultrasound scans, bone scans, PET scans, bone marrow examinations, barium X-rays, endoscopy, lymphangiography, IVU (intravenous urography) or IVP (intravenous pyelography), lumbar puncture, cystoscopy, immunological tests (anti-malignin antibody screening), and cancer marker tests.

[0123] Clinical information for subject 6 can be obtained manually or automatically from lab 2. Where simplicity of system is desired, information can be obtained automatically at predetermined or regular time intervals. Regular time intervals may refer to time intervals in which laboratory data are collected automatically by the methods and systems described herein, based on time measurements such as hours, days, weeks, months, or years. In one embodiment, data collection and processing are performed at least once a day. In one embodiment, data transfer and collection are performed approximately monthly, bi-weekly, weekly, several times a week, or daily. Alternatively, information retrieval can be performed at predetermined time intervals, which do not necessarily have to be regular. For example, the first retrieval step may occur after one week, and the second retrieval step may occur after one month. Data transfer and collection can be made on a case-by-case basis according to the nature of the disorder being managed and the frequency of the required examinations and medical tests for the subject.

[0124] Figure 9B shows an exemplary process for generating a gene report including a tumor response map and a summary of associated changes. The tumor response map is a graphical representation of gene information showing changes in tumor-derived gene information over time, e.g., qualitative and quantitative changes. Such changes can reflect the subject's response to therapeutic intervention. This process can reduce the error rate and bias, which can be orders of magnitude higher than those required for reliable detection of cancer-related de novo gene variants. The process may first include a step of capturing gene information by collecting a bodily fluid sample (e.g., blood, saliva, sweat, urine, etc.) as a source of gene material. Next, the process may include a step of sequencing the material (11). For example, polynucleotides in the sample can be sequenced to generate multiple sequence readings. The tumor burden in a sample containing polynucleotides can be estimated as the relative number of sequence readings containing variants to the total number of sequence readings generated from the sample. Once copy number variants are analyzed, tumor burden can be estimated as a relative excess (e.g., in the case of gene duplication) or relative deletion (e.g., in the case of gene exclusion) of the total number of sequence readings at the test and control loci. For example, one run may generate 1000 readings mapped to an oncogene locus, of which 900 correspond to the wild type and 100 correspond to the cancer variant, indicating copy number variants in this gene. Further details regarding exemplary sample collection and sequencing of genetic material are described later in Figures 10-11.

[0125] Next, the genetic information can be processed (12). Subsequently, gene variants can be identified. The process may include a step of determining the frequency of gene variants in a sample containing genetic material. If the process is noisy, the process may include a step of separating the information from the noise (13).

[0126] Sequencing methods for gene analysis can have error rates. For example, Illumina's mySeq system can produce error rates in the low single digits. For every 1000 sequence readings mapped to a gene locus, approximately 50 readings (about 5%) can be expected to contain errors. Certain methodologies, such as those described in WO2014 / 149134, can significantly reduce error rates. Errors can produce noise that can evoke ambiguous signals from cancer present at low levels in the sample. For example, if a sample has a tumor burden at around 0.1% to 5% of the sequencing system's error rate, it can be difficult to distinguish signals corresponding to gene variants caused by cancer from those caused by noise.

[0127] Analysis of gene variants can be used for diagnosis in the presence of noise. The analysis can be based on the frequency of sequence variants or the level of CNVs (14), and a diagnostic reliability indication or level can be established for detecting gene variants within a noise range (15).

[0128] Next, the process may include steps to increase diagnostic reliability. This can be done using multiple measurements to increase diagnostic reliability (16), or using measurements at multiple time points to determine whether the cancer is progressing, in remission, or stabilized (17). Diagnostic reliability can be used to identify disease states. For example, cell-free polynucleotides collected from a subject may contain polynucleotides derived from affected cells, such as cancer cells, along with polynucleotides derived from normal cells. Polynucleotides derived from cancer cells may have genetic variants, such as somatic mutations and copy number variants. When cell-free polynucleotides from a sample derived from a subject are sequenced, these cancer polynucleotides are detected as sequence variants or copy number variants.

[0129] Regardless of whether the parameter measurements are within the noise range or not, they can have confidence intervals. When examined over time, comparing these confidence intervals over time can determine whether the cancer is progressing, stabilizing, or in remission. If the confidence intervals overlap, there is no statistically significant difference between these measurements, and therefore it is not possible to know whether the disease is increasing or decreasing. However, if the confidence intervals do not overlap, this indicates the direction of the disease. For example, comparing the lowest point in the confidence interval at one time point with the highest point in the confidence interval at a second time point indicates direction.

[0130] Next, the process may include a step of generating a genetic report / diagnosis. The process may include a step of generating a genetic graph of multiple measurements showing mutational tendencies (18) and a step of generating a report showing treatment outcomes and options (19).

[0131] Figures 10A–10C illustrate in more detail one embodiment for generating gene reports and diagnoses (e.g., reports / diagnoses). In one implementation, Figure 10C shows exemplary pseudocode executed by the system in Figure 9A for processing non-CNV reporting mutant allele frequencies. However, the system can also process CNV reporting mutant allele frequencies.

[0132] Samples containing genetic material, such as cfDNA, can be collected from the subject at multiple time points, i.e., sequentially. The genetic material can be sequenced, for example, using a high-throughput sequencing system. Sequencing can target a desired locus to detect gene variants, such as genes with somatic mutations, genes with copy number variations, or genes involved in gene fusion in cancer. At each time point, quantitative measurements of the found gene variants can be determined. For example, in the case of cfDNA, the quantitative measurement may be the frequency or percentage of gene variants among the polynucleotides mapped to the locus, or the sequence readings or absolute number of polynucleotides mapped to the locus. Next, gene variants with a non-zero quantity at at least one time point can be represented graphically throughout the entire time period. For example, in the collection of 1000 sequences, variant 1 may be found at 50, 30, and 0 at time points 1, 2, and 3, respectively. Variant 2 may be found at 0, 10, and 20 at these time points. These quantities can be normalized to 5%, 3%, and 0% for Variant 1, and to 0%, 1%, and 2% for Variant 2. A graph showing the combined total non-zero results can display the quantities of both variants at all time points. The normalized quantities can be scaled so that each percentage is represented by a layer with, for example, a height of 1 mm. For example, in this case, the heights would be: at time point 1: 5 mm (Variant 1) and 0 mm (Variant 2); at time point 2: 3 mm (Variant 1) and 1 mm (Variant 2); and at time point 3: 0 mm (Variant 1) and 2 mm (Variant 2). The graph may be in the form of a stacked area graph, such as a streamgraph. "Zero" time point (Before the first time point) can be represented by a point where all values ​​are 0. The height of the variant amounts in the graph can be, for example, relative or proportional to each other. For example, a variant frequency of 5% at a given time point can be represented by twice the height of a variant with a frequency of 2.5% at the same time point. The stacking order can be chosen for ease of understanding. For example, variants can be stacked from bottom to top, from most to least. Alternatively, in a stream graph, the variant with the highest initial amount can be stacked in the middle, with other variants of decreasing amounts on either side. In certain embodiments, the area can be color-coded based on the variants. Variants in the same gene can be shown with different hues of the same color. For example, KRAS variants can be shown with different shades of blue, and EGFR variants can be shown with different shades of red.

[0133] Next, moving to Figure 10A, the process may include a step of receiving genetic information from a DNA sequencer (30). Subsequently, the process may include a step of determining specific genetic alterations and their amounts (32).

[0134] Next, a tumor response map is generated. To generate the map, the process may include steps to normalize the quantity for each gene change and generate a scale factor for rendering across all test points (34). As used herein, the term “normalize” generally means adjusting values ​​measured on different scales to a conceptually common scale. For example, data measured at different points are transformed / adjusted so that all values ​​can be resized to a common scale. As used herein, the term “scale factor” generally refers to a number that scales or multiplies a quantity. For example, in the equation y=Cx, C is the scale factor for x. C is also the coefficient of x and can be called the constant of the proportionality of y to x. The values ​​are normalized to enable plotting on a visually understandable common scale. The scale factor is also used to know the exact height corresponding to the value being plotted (for example, the 10% mutant allele frequency may represent 1 cm in a report where the total height is 10 cm). The scale factor is considered a universal scale factor because it applies to all test points. For each test point, the process may include a step of rendering information in a tumor response map (36). In operation 36, the process may include a step of rendering the changes and relative heights using a determined scale factor (38) and a step of assigning a specific visual indicator for each change (40). In addition to the response map, the process may include a step of generating a summary of changes and treatment options (42). Information from clinical trials that may be helpful in suggesting specific genetic changes and other useful treatments may also be presented, along with explanations of terminology, test methodology and other information, which may be added to the report and rendered for the user.

[0135] In one implementation, copy number variation can be reported as a graph showing various positions in the genome and the corresponding increase, decrease, or maintenance of copy number variation at each position. Furthermore, copy number variation can be used to report a percentage score indicating how much disease material (or nucleic acids with copy number variation) is present in a cell-free polynucleotide sample.

[0136] In another embodiment, the report includes annotations to help physicians interpret the results and recommend treatment options. The annotations relate to the condition as defined in the NCCN Clinical Practice Guidelines in Oncology (trademark) or the American Society of Clinical Oncology (ASCO) Clinical Practice Guidelines. The report may include annotations for one or more FDA-approved drugs for off-label use, which are listed in the Centers for Medicare and Medicaid Services (CMS) Cancer Treatment Compendium. Annotations may include the inclusion of one or more drugs and / or one or more experimental drugs found in the scientific literature. Annotations may include the step of linking the included drug treatment options to references containing scientific information about the drug treatment options. Scientific information may originate from peer-reviewed articles in medical journals. Annotations may include providing links to information about clinical trials of the drug treatment options in the report. In electronic reports, annotations may include the presentation of information in a pop-up box or fly-over box near the provided drug treatment options. Annotations may include the addition of information to the report, selected from the group consisting of one or more drug treatment options, scientific information related to one or more drug treatment options, one or more links to scientific information about one or more drug treatment options, one or more links to citations of scientific information about one or more drug treatment options, and clinical trial information about one or more drug treatment options.

[0137] Figure 10B illustrates an exemplary process for generating a tumor response map pathway that can be used by healthcare professionals, such as physicians, to make decisions regarding patient care, for example. In this embodiment, the process may first include a step of determining an overall scale factor (43). In one embodiment, for all non-CNV (copy number variation) reported variant allele frequencies, the process may include a step of converting the absolute value to a relative metric / scale that is more acceptable for plotting (e.g., multiplying the variant allele frequency by 100 and taking the logarithm of the value) and a step of determining an overall scale factor using the maximum observed value. The process then involves a step of visualizing information from the earliest test dataset (44). Visualization may include a graphical display of the information in a user interface (e.g., a computer screen) or in a tangible form (e.g., a piece of paper). For each non-CNV variation, the process may include a step of multiplying the scale factor by a gene-specific converted value, using the variant as a quantity indicator for plotting, and subsequently assigning a color / specific visual indicator to each variation. Next, the process may include a step to visualize the information for subsequent inspection points using the following pseudocode (45): If the composition of the test results does not change, the new panel will continue to display the same information as the previous panel date. If the change remains the same, but the quantity changes, To plot the variant, the quantity indicator is recalculated, and all updated values ​​are replotted in the existing panel(s) and new panel(s) for the most recent inspection date. When new changes are added, Add a change to the top level of all existing changes. Calculate the converted value Recalculate the scale factor. Redraw the response map and plot the newly emerging changes along with the changes from previous test dates that are still detected on the current test date. If the existing previous changes are not included in the set of detected changes, Use a zero height and plot the amount of change for all subsequent test days. It still includes colors that are part of an unavailable color set.

[0138] Each of the subsequent panels displaying the examination date may also include additional patient or intervention information that can correlate with the changes seen in the rest of the map. Similar scaling, plotting, and transformations can also be implemented for CNVs and other types of DNA changes (e.g., methylation) to display these quantities in separate or combined charts. These additional annotations can also be quantified themselves and plotted in the map as well.

[0139] The process may then include a step of determining a summary of the changes and treatment options (46). In one embodiment, for the change with the highest mutant allele frequency, the following actions are taken: Report all changes related to the gene in order of the frequency of mutant alleles showing a decrease in non-CNV changes. Report all CNV changes related to the gene in the order of decreasing CNV values. We will repeat the following genes that have the next highest non-CNV mutant allele frequencies, which have not yet been reported. For each reported change, the process may include a step that includes a trend indicator of that change across different inspection date points. The grouping of maximum mutant allele frequencies can also be extended beyond just the genes containing them for better encapsulation annotation, such as biological pathways and evidence levels.

[0140] Figures 10D–10I show an exemplary report generated by the system in Figure 9A. In Figure 10D, the patient identification section 52 provides patient information, reporting date, and physician contact information. The tumor response map 54 includes a modified stream graph 56 that shows tumor activity by a specific color for each variant gene. The graph 56 has an attached summary text box 58. Further details are provided in the summary section 60 of the changes and treatment options. Changes 62 and 64 are presented in section 60 along with mutation trends, variant allele frequencies, cell-free amplification, FDA-approved drug indications, and FDA-approved drugs and clinical drug trial information with other indications. Figures 10D-1, 10D-2, and 10D-3 provide enlarged views of Figure 10D.

[0141] Figure 10E shows an exemplary reporting section providing the definition, comments, and interpretation of the test. Figures 10E-1 and 10E-2 provide enlarged views of Figure 10E. Figure 10F shows an exemplary detailed treatment outcome section of the report. Figures 10F-1 and 10F-2 provide enlarged views of Figure 10F. Figure 10G shows an exemplary consideration of the clinical relevance of the detected changes. Figures 10G-1 and 10G-2 provide enlarged views of Figure 10G. Figure 10H shows potentially available drug therapies under clinical trial. Figure 10I shows the test method and its limitations. Figures 10I-1 and 10I-2 provide enlarged views of Figure 10I.

[0142] Figures 10J–10P show various exemplary modified stream graphs.56 A stream graph, or stream-graph, is a type of stacked area graph that displaces around a central axis, resulting in a flowing, organic shape. A stream graph is a generalization of a stacked area graph with a free baseline. Shifting the baseline minimizes changes in gradient (or "wiggles") within individual series, thereby making it easier to perceive the thickness of any given layer across the data.

[0143] For example, Figure 10J shows seven layers representing at least eight mutants over three periods and a "0" time point (all values ​​are "0"). Figure 10K shows a single mutant over four periods. No mutants are detected at the second, third, and fourth time points. Figure 10L shows the frequency of the dominant allele at each time point. Figure 10M shows a single time point with a total of four mutants in two genes. The mutants are identified by the amino acid at the position where the change occurs (i.e., EGFR T790M).

[0144] One embodiment renders a stream graph so as not to be x-axis reflective. The modified graph applies specific scaling to display proportional properties. The graph can show the addition of new properties over time. The presence or absence of mutations can be reflected in a graph form showing the corresponding increase, decrease, or maintenance of the frequency of mutations at various positions in the genome and at each position. Furthermore, mutations can be used to report a percentage score indicating how much disease material is present in the cell-free polynucleotide sample. Given known statistics of typical variance at reported positions in non-disease reference sequences, a confidence score can be given for each detected mutation. Mutations may be ranked in order of abundance in the subject or by clinically usable importance.

[0145] Mapping genomic positions and copy number variations in subjects with cancer can indicate that a particular cancer is highly aggressive and resistant to treatment. Subjects can be monitored and re-examined for a period of time. If, at the end of this period, the copy number variation profile, for example, depicted in a tumor response map, begins to increase dramatically, this can indicate that the current treatment is no longer effective. Comparisons can also be made using the genetic profiles of other subjects. For example, if it is determined that this increase in copy number variation indicates cancer progression, the initially prescribed treatment regimen will no longer treat the cancer, and a new treatment will be prescribed.

[0146] These reports can be submitted and accessed electronically via the internet. Analysis of sequence data can occur at a location other than the subject's site. Reports can be generated and transmitted to the subject's site. With an internet-connected computer, the subject can access reports reflecting their tumor burden.

[0147] Next, details of an exemplary gene testing process are disclosed. Next, moving to Figure 11A, the exemplary process receives gene material derived from a blood sample or other body sample (1102). The process may include a step of converting polynucleotides derived from the gene material into tagged parental nucleotides (1104). The tagged parental nucleotides are amplified to produce amplified progeny polynucleotides (1106). A subset of the amplified polynucleotides is sequenced to generate sequence readings (1108), which are grouped into families, each generated from a specific tagged parental nucleotide (1110). At a selected locus, the process may include a step of assigning a family-specific confidence score to each family (1112). Next, a consensus is determined using previous readings. This is done by reviewing the previous confidence scores for each family, and if a matching previous confidence score exists, the current confidence score is increased (1114). If a previous confidence score exists but is mismatched, the current confidence score is not modified in one embodiment (1116). In other embodiments, confidence scores are adjusted in a predetermined manner for previous, mismatched confidence scores. If a family is detected for the first time, the current confidence score may be reduced due to the possibility of false readings (1118). The process may include a step of estimating the frequency of the family at a locus in a set of tagged parental polynucleotides based on the confidence score. Subsequently, a genetic test report is generated as described above (1120).

[0148] Temporal information was used in Figures 11A and 11B to enhance information for detecting mutations or copy number variations, but other consensus methods may be applied. In other embodiments, historical comparisons can be used in conjunction with other consensus sequences mapped to specific reference sequences to detect instances of gene mutations. Consensus sequences mapped to specific reference sequences can be measured and normalized against control samples. Measurements of molecules mapped to reference sequences can be compared across the genome to identify regions in the genome where copy number variations or heterozygosity is lost. Consensus methods include, for example, linear or nonlinear methods for constructing consensus sequences derived from digital communications theory, information theory, or bioinformatics (e.g., voting, averaging, statistical, maximum apost-hoc or maximum likelihood detection, dynamic programming, Bayesian, hidden Markov, or support vector machine methods). After sequence read coverage is determined, a stochastic modeling algorithm is applied to convert the normalized nucleic acid sequence read coverage for each window region into separate copy number states. In some cases, this algorithm may include one or more of the following: Hidden Markov models, dynamic programming, support vector machines, Bayesian networks, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering methodology, and neural networks.

[0149] As depicted in Figure 11B, comparing sequence coverage to a control sample or reference sequence can aid in normalization across a window. In this embodiment, cell-free DNA is extracted and isolated from readily available bodily fluids such as blood, sweat, saliva, and urine. For example, cell-free DNA can be extracted using various methods known in the art, including but not limited to isopropanol precipitation and / or silica-based purification. Cell-free DNA can be extracted from any number of subjects, including subjects without cancer, subjects at risk of cancer, or subjects known to have cancer (e.g., by other means).

[0150] After the isolation / extraction step, one of numerous different sequencing operations can be performed on the cell-free polynucleotide sample. The sample can be processed with one or more reagents (e.g., enzymes, specific identifiers (e.g., barcodes), probes, etc.) before sequencing. In some cases, if the sample is processed with a specific identifier such as a barcode, the sample or sample fragments can be tagged with the specific identifier, individually or in subgroups. The tagged sample can then be used in downstream applications, such as sequencing reactions, allowing for the tracking of individual molecules back to their parent molecules.

[0151] Cell-free polynucleotides can be tagged or tracked to enable subsequent identification and origin determination of specific polynucleotides. Assigning identifiers to individual polynucleotides or subgroups of polynucleotides allows a unique identity to be assigned to individual sequences or fragments of sequences. This can enable the acquisition of data from individual samples, not limited to the average of the samples. In some cases, nucleic acids or other molecules originating from a single strand can share a common tag or identifier and thus can be later identified as originating from this strand. Similarly, all fragments originating from a single strand of nucleic acid can be tagged with the same identifier or tag, thereby enabling subsequent identification of the fragment from the parent strand. In other cases, gene expression products (e.g., mRNA) can be tagged to quantify expression. Barcodes or barcodes combined with the sequence to which they are attached can be counted. In yet another case, systems and methods can be used as PCR amplification controls. In such cases, multiple amplification products from a PCR reaction can be tagged with the same tag or identifier. If the products are later sequenced and demonstrate differences in sequences, differences between products with the same identifier may be due to errors in PCR. Furthermore, individual sequences can be identified based on the characteristics of the sequence data of the reading itself. For example, the detection of unique sequence data at the beginning and end of each sequencing reading can be used alone or in combination with the length or number of base pairs of each sequence reading to assign a unique identity to individual molecules. This allows fragments derived from a single strand of nucleic acid that have been assigned a unique identity to subsequently be identified as fragments derived from the parent strand. This can be used in conjunction with the initial starting gene material, which can be a bottleneck, to limit diversity.

[0152] Furthermore, the length of the sequencing read can be used, either alone or in combination with the use of barcodes, using the unique sequence data at the beginning and end of each sequencing read. In some cases, the barcode may be unique, as described herein. In other cases, the barcode itself does not have to be unique. In this case, the use of non-unique barcodes and the length of the sequencing read, combined with the sequence data at the beginning and end of each sequencing read, can enable the assignment of unique identity to individual sequences. Similarly, this can enable the subsequent identification of fragments derived from a single strand of nucleic acid that have been assigned unique identity to the parent strand.

[0153] Generally, the methods and systems provided herein are useful for the preparation of cell-free polynucleotide sequences for downstream sequencing reactions. In many cases, the sequencing method is classical Sanger sequencing. Examples of sequencing methods include, but are not limited to, high-throughput sequencing, pyrosequencing, synthetic sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, RNA-Seq (Illumina), digital gene expression (Helicos), next-generation sequencing, synthetic single-molecule sequencing (SMSS) (Helicos), large-scale parallel sequencing, clonal single-molecule array (Solexa), shot cancer sequencing, Maxham-Gilbert sequencing, primer walking, and any other sequencing method known in the art.

[0154] A sequencing method typically involves sample preparation, sequencing of polynucleotides in the prepared sample to generate sequence readings, and bioinformatics manipulation of the sequence readings to generate quantitative and / or qualitative genetic information about the sample. Sample preparation typically involves converting polynucleotides in the sample to a form compatible with the sequencing platform used. This conversion may involve tagging of polynucleotides. In certain embodiments of the present invention, the tag includes a polynucleotide sequence tag. The conversion methodology used in sequencing may not be 100% efficient. For example, it is not uncommon for polynucleotides in a sample to be converted with an efficiency of about 1-5%, i.e., about 1-5% of the polynucleotides in the sample are converted to tagged polynucleotides. Polynucleotides that are not converted to tagged molecules are not represented in the tagged library for sequencing. Therefore, polynucleotides with gene variants that are expressed at low frequency in the initial genetic material may not be represented in the tagged library and, consequently, may not be sequenced or detected. Increasing conversion efficiency increases the likelihood that polynucleotides in the initial genetic material will be expressed in the tagged library, and consequently, detected by sequencing. Furthermore, rather than directly addressing the problem of low conversion efficiency in library preparation, many protocols currently require more than 1 microgram of DNA as input material. However, when input sample material is limited, or when detection of polynucleotides at low expression is desired, high conversion efficiency allows for efficient sequencing of the sample and / or appropriate detection of such polynucleotides.

[0155] Generally, mutation detection can be performed on selectively enriched regions of the genome or on purified and isolated transcriptomes (1302). Specific regions, including but not limited to genes, oncogenes, tumor suppressor genes, promoters, regulatory elemental sequences, non-coding regions, miRNAs, snRNAs, and others, as described herein, can be selectively amplified from the total population of cell-free polynucleotides. This can be done as described herein. In one example, multiplex sequencing can be used with or without barcode labeling of individual polynucleotide sequences. In another example, sequencing can be performed using any nucleic acid sequencing platform known in the art. This step generates multiple genome fragment sequence readings (1304). Furthermore, a reference sequence is obtained from a control sample taken from another subject. In some cases, the control subject may be one known to be free from known genetic abnormalities or diseases. In some cases, such sequence readings may contain barcode information. In other examples, barcodes are not used.

[0156] After sequencing, the reads can be assigned a quality score. The quality score may be an indication of the reads that, based on a threshold, they may be useful in subsequent analysis. In some cases, some reads are not of sufficient quality or length to perform the subsequent mapping step. Sequenced reads with a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out of the dataset. In other cases, sequenced reads assigned a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out of the dataset. In step 1306, genome fragment reads that meet the specified quality score thresholds are mapped to a reference genome or a reference sequence known to be free of mutations. After mapping alignment, the sequenced reads are assigned a mapping score. The mapping score may be an indication or read that is back-mapped to the reference sequence, indicating whether each position is uniquely mappable. In some cases, the readings may be sequences unrelated to the mutation analysis. For example, some sequence readings may originate from contaminating polynucleotides. Sequence readings with mapping scores of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out of the dataset. In other cases, sequence readings assigned mapping scores of less than 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out of the dataset.

[0157] For each mappable base, any base that does not meet the minimum threshold for mappability or is of low quality can be replaced by the corresponding base found in the reference sequence.

[0158] After verifying the readout coverage and identifying variant bases in each readout compared to the control sequence, the frequency of the variant base can be calculated as the number of readouts containing the variant divided by the total number of readouts (1308). This can then be expressed as a ratio for each mappable position in the genome.

[0159] For each base position, the frequencies of all four nucleotides—cytosine, guanine, thymine, and adenine—can be analyzed in comparison to the reference sequence. Stocastic or statistical modeling algorithms can be applied to transform the normalized ratios for each mappable position to reflect the frequency state for each base variant. In some cases, this algorithm may include one or more of the following: hidden Markov models, dynamic programming, support vector machines, Bayesian or stochastic modeling, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering methodology, and neural networks.

[0160] By utilizing the distinct mutational states of each base position, base variants with high variance compared to the reference sequence baseline can be identified. In some cases, the baseline may represent frequencies of at least 0.0001%, 0.001%, 0.01%, 0.1%, 1.0%, 2.0%, 3.0%, 4.0%, 5.0%, 10%, or 25%. In other cases, the baseline may represent frequencies of at least 0.0001%, 0.001%, 0.01%, 0.1%, 1.0%, 2.0%, 3.0%, 4.0%, 5.0%, 10%, or 25%. In some cases, all adjacent base positions containing base variants or mutations can be merged into a segment to report the presence or absence of the mutation. In some cases, various positions can be filtered before being merged with other segments.

[0161] After calculating the frequency of variance for each base position, variants with the largest deviation from specific positions in the sequence from which the target originates can be identified as mutations by comparing them with the reference sequence. In some cases, the mutation may be a cancer mutation. In other cases, the mutation may correlate with a disease state.

[0162] Mutations or variants may include, but are not limited to, single base substitutions or small indels, transversions, transpositions, inversions, deletions, truncations, or gene truncations, and may include other genetic abnormalities. In some cases, mutations may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 nucleotides long. In other cases, mutations may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 nucleotides long.

[0163] Next, a consensus is determined using previous readings. This is done by reviewing the previous confidence scores for the corresponding bases, and if a matching previous confidence score exists, the current confidence score is increased (1314). If a previous confidence score exists but does not match, in one embodiment, the current confidence score is not modified (1316). In other embodiments, the confidence score is adjusted in a predetermined manner with respect to the mismatched previous confidence score. If the family is detected for the first time, the current confidence score may be reduced due to the possibility of a false reading (1318). The process may then include a step of converting the variance frequencies of each base for each base position into separate variant states (1320).

[0164] Numerous cancers can be detected using the methods and systems described herein. Cancer cells, like most cells, can be characterized by the rate of turnover, in which old cells die and are replaced by newer cells. Generally, dead cells that come into contact with vascular structures in a given object can release DNA or DNA fragments into the bloodstream. This is also true for cancer cells at various stages of disease. Cancer cells can also be characterized by various genetic abnormalities, such as copy number variations and mutations, depending on the stage of the disease. This phenomenon can be used to detect the presence or absence of cancer in an individual using the methods and systems described herein.

[0165] For example, blood can be collected from a subject at risk of cancer and prepared as described herein to generate a population of cell-free polynucleotides. In one example, this may be cell-free DNA. The systems and methods of this disclosure can be used to detect mutations or copy number variations that may be present in certain cancers that exist. The methods can be useful in detecting the presence of cancer cells in the body, even in the absence of disease symptoms or other characteristics.

[0166] The types and number of cancers that may be detected include, but are not limited to, blood cancers, brain cancers, lung cancers, skin cancers, nasal cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, intestinal cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, oral cancers, gastric cancers, solid tumors, heterogeneous tumors, homogeneous tumors, and others.

[0167] This system and method can be used to detect any number of genetic abnormalities that cause or may cause cancer. Such abnormalities include, but are not limited to, mutations, indels, copy number variations, transversions, transpositions, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural changes, gene fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid methylation, all of which are associated with infection and cancer.

[0168] Furthermore, the systems and methods described herein can also be used to help characterize a particular type of cancer. Genetic data generated from the systems and methods of this disclosure can help practitioners better characterize specific forms of cancer. Often, cancers are heterogeneous in both composition and staging. Genetic profiling data can enable the characterization of specific subtypes of cancer, which can be important in the diagnosis or treatment of such specific subtypes. This information can also provide subjects or practitioners with clues about the prognosis of specific types of cancer.

[0169] The systems and methods provided herein can be used to monitor already known cancers or other diseases in specific subjects. This can enable either the subject or the practitioner to adapt treatment options to match the progression of the disease. In this example, the systems and methods described herein can be used to construct a genetic profile of the disease course in a specific subject. In some cases, cancer may progress and become more aggressive and genetically unstable. In other cases, cancer may remain benign, inactive, or dormant. The systems and methods of this disclosure may be useful in determining disease progression.

[0170] Furthermore, the systems and methods described herein may be useful in determining the effectiveness of specific treatment options. In one example, if a treatment is successful in that it can kill more cancers and cause DNA shedding, a successful treatment option can indeed increase the amount of copy number variation or mutation detected in the blood of the subject. In other examples, this may not occur. In yet another example, a particular treatment option may correlate with the genetic profile of cancer over time. This correlation can be useful in selecting a treatment method. Moreover, if the cancer is observed to be in remission after treatment, the systems and methods described herein may be useful in monitoring residual disease or disease recurrence.

[0171] The methods and systems described herein are not limited to the detection of mutations and copy number variations associated solely with cancer. Various other diseases and infections can result in other types of conditions that are suitable for early detection and monitoring. For example, in certain cases, a genetic disorder or infectious disease can cause a specific genetic mosaicism in a subject. This genetic mosaicism can result in observable copy number variations and mutations. In another example, the systems and methods of this disclosure can also be used to monitor the genomes of immune cells in the body. Immune cells, such as B cells, can undergo rapid clonal growth in the presence of a certain disease. This clonal growth can be monitored using copy number variation detection, and a certain immune state can be monitored. In this example, copy number variation analysis can be performed over time to generate a profile of how far a particular disease is progressing.

[0172] Furthermore, the systems and methods of this disclosure can also be used to monitor systemic infections themselves that may be caused by pathogens such as bacteria or viruses. Using copy number variation or even mutation detection, it is possible to determine how much the pathogen population has changed over the course of the infection. This can be particularly important in chronic infections such as HIV / AIDs or hepatitis infections, by which the virus may change its life cycle state and / or mutate into a more pathogenic form over the course of the infection.

[0173] Another example in which the systems and methods of this disclosure can be used is the monitoring of transplanted tissues. Generally, transplanted tissues undergo some degree of rejection by the body after transplantation. Since immune cells attempt to destroy the transplanted tissue, the methods of this disclosure can be used to determine or profile the rejection activity of the host body. This can be useful in monitoring the state of the transplanted tissue and in modifying the course of treatment or prevention of rejection.

[0174] Furthermore, the methods of this disclosure can be used to characterize heterogeneity of abnormal conditions in a subject, the methods comprising the step of generating a gene profile of extracellular polynucleotides in the subject, wherein the gene profile includes multiple data resulting from copy number variation and mutation analysis. Diseases that may be heterogeneous include, but are not limited to, cancer. Disease cells may not be identical. In the example of cancer, it is known that some tumors contain different types of tumor cells, some cells from different stages of cancer. In other examples, heterogeneity may include multiple lesions of the disease. Again, in the example of cancer, there may be multiple tumor lesions, in which case one or more lesions are probably the result of metastasis propagating from the primary site.

[0175] The methods described herein can be used to generate or profile a data fingerprint or set, which is the sum of genetic information derived from different cells in heterogeneous diseases. This dataset may include single or combined copy number variation and mutation analyses.

[0176] Furthermore, the systems and methods of this disclosure can be used to diagnose, predict the prognosis of, monitor or observe cancers or other diseases of fetal origin. That is, such methodologies can be used in pregnant subjects to diagnose, predict the prognosis of, monitor or observe cancers or other diseases in prenatal subjects whose DNA and other polynucleotides may circulate with maternal molecules.

[0177] Furthermore, these reports are submitted and accessed electronically via the internet. Analysis of sequence data occurs at a location other than the subject's site. The reports are generated and transmitted to the subject's site. The subject accesses the reports reflecting their tumor burden via an internet-connected computer.

[0178] Healthcare providers can use the annotated information to select other drug treatment options and / or provide insurance companies with information about drug treatment options. This method may include, for example, a step of annotating drug treatment options in relation to a condition in the NCCN Clinical Practice Guidelines for Oncology (trademark) or the American Society of Clinical Oncology (ASCO) Clinical Practice Guidelines.

[0179] Drug treatment options stratified in the report may be annotated in the report by including additional drug treatment options. These additional drug treatments may be FDA-approved drugs for off-label use. A provision in the Omnibus Budget Reconciliation Act of 1993 (OBRA) requires Medicare to cover off-label use of anticancer drugs included in the standard medicine compendium. Drugs used for annotating the list can be found in CMS-approved compendiums, including the National Comprehensive Cancer Network (NCCN) Drugs and Biologics Compendium®, Thomson Micromedex DrugDex®, Elsevier Gold Standard's Clinical Pharmacology compendium, and American Hospital Formulary Service-Drug Information Compendium®.

[0180] Drug treatment options can be annotated by listing experimental drugs that may be useful in treating cancers with one or more molecular markers for a particular condition. Experimental drugs may be those for which in vitro data, in vivo data, animal model data, preclinical trial data, or clinical trial data are available. Data may be found, for example, in the American Journal of Medicine, Annals of Internal Medicine, and Annals of Oncology. , Annals of Surgical Oncology, Biology of Blood and Marrow Transplantation, Blood, Bone Marrow Transplantation, British Journal of Cancer, British Journal of Hematology, British Medical Journal, Cancer, Clinical Cancer Research, Drugs, European Journal of Cancer (formerly the European Journal of Cancer and Clinical Oncology), Gynecologic Oncology, International Journal of Radiation, Oncology, Biology, and Physics, The Journal of the American Medical Association, Journal of Clinical Oncology, Journal of the National Cancer Institute, Journal of the National Comprehensive Cancer Network (NCCN), Journal of Urology, Lancet, Lancet Oncology, Leukemia, The This may be published in peer-reviewed medical literature found in journals included in the CMS Medicare Benefit Policy Manual, including the New England Journal of Medicine and Radiation Oncology.

[0181] Drug treatment options can be annotated by providing links to electronic reports that connect listed drugs to scientific information about those drugs. For example, a link to information on clinical trials of a drug (clinicaltrials.gov) can be provided. If the reports are provided via computer or computer website, the links may be footnotes, hyperlinks to websites, pop-up boxes, or flyover boxes containing information. Reports and annotated information may be provided in print format, and annotations may be, for example, footnotes to references.

[0182] Information for annotating one or more drug treatment options in a report may be provided by a commercial entity that stores scientific information. Healthcare providers can treat subjects such as cancer patients with experimental drugs included in the annotated information, and healthcare providers can access the annotated drug treatment options, search for scientific information (e.g., printed medical journal articles), and submit this (e.g., printed journal articles) to insurance companies along with a claim for reimbursement to provide the drug treatment. Physicians may use any of the various diagnostic-related group (DRG) codes to enable reimbursement.

[0183] Drug treatment options in a report may also be annotated with information about other molecular components in the pathways affected by the drug (e.g., information about drugs that target kinases downstream of the cell surface receptor that is the drug target). Drug treatment options may also be annotated with information about drugs that target one or more other molecular pathway components. The identification and / or annotation of pathway information may be outsourced or subcontracted to another company.

[0184] Annotated information may include, for example, drug names (e.g., FDA-approved drugs for off-label use; drugs found in the CMS Approval Compendium, and / or drugs listed in scientific (medical) journal articles), scientific information related to one or more drug treatment options, one or more links to scientific information about one or more drugs, clinical trial information about one or more drugs (e.g., information from clinicaltrials.gov / ), and scientific information about the drug. This may be one or more links to quotations, etc.

[0185] Annotated information can be inserted at any location in the report. Annotated information can be inserted at multiple locations in the report. Annotated information can be inserted near the section on stratified drug treatment options in the report. Annotated information can be inserted on a separate page from the section on stratified drug treatment options in the report. Reports that do not contain stratified drug treatment options can be annotated with information.

[0186] This system may also include reporting on the effects of drugs on samples (e.g., tumor cells) isolated from subjects (e.g., cancer patients). In vitro cultures using tumors derived from cancer patients can be established using techniques known to those skilled in the art. This system may also include high-throughput screening of FDA-approved off-label drugs or experimental drugs using the in vitro cultures and / or xenograft models. This system may also include monitoring of tumor antigens for recurrence detection.

[0187] This system can provide internet-connected access to reports of subjects with cancer. The system can use either a handheld or benchtop DNA sequencer. A DNA sequencer is a scientific instrument used to automate the DNA sequencing process. Given a DNA sample, the DNA sequencer is used to determine the order of four bases: adenine, guanine, cytosine, and thymine. The order of the DNA bases is reported as a text string called the reading. Some DNA sequencers can also be considered optical instruments, as they analyze light signals originating from fluorescent dyes attached to nucleotides.

[0188] The data is sent by a DNA sequencer to a computer for processing, either directly or via the internet. The data processing mode of the system can be implemented in digital electronic circuits, or in computer hardware, firmware, software, or a combination thereof. The data processing device of the present invention can be implemented in a computer program product tangibly embodied in a machine-readable storage device for implementation by a programmable processor; the data processing method steps of the present invention can be performed by a programmable processor that executes a program of instructions to perform the functions of the present invention by manipulating input data and generating outputs. The data processing mode of the present invention can be advantageously executed in one or more computer programs executable in a programmable system including at least one programmable processor coupled to receive data and instructions from a data storage system and to transmit data and instructions thereto, at least one input device, and at least one output device. Each computer program can be implemented, if desired, in a high-level procedural or object-oriented programming language or in assembly or machine language; in either case, the language can be a compiled or interpreted language. Suitable processors include, as an example, both general-purpose and special-purpose microprocessors. Generally, a processor will receive instructions and data from read-only memory and / or random-access memory. Suitable storage devices for the tangible embodiment of computer program instructions and data include, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and all forms of non-volatile memory, including CD-ROM disks. Any of the above can be supplemented or incorporated into an ASIC (Application-Specific Integrated Circuit).

[0189] To provide user interaction, the present invention can be implemented using a computer system having a display device such as a monitor or LCD (liquid crystal display) screen for displaying information to the user, and an input device that allows the user to provide input to the computer system, such as a two-dimensional pointing device such as a keyboard, mouse or trackball, or a three-dimensional pointing device such as a data glove or gyroscope mouse. The computer system can be programmed to provide a graphical user interface in which the computer program interacts with the user. The computer system can be programmed to provide virtual reality, a three-dimensional display interface.

[0190] therapeutic intervention The method described herein enables the provision of therapeutic interventions that are more precisely directed towards the disease morphology in the subject, and the calibration of such therapeutic interventions over time. This precision, in part, reflects the ability to profile the systemic tumor state of the subject, as reflected in tumor heterogeneity. Thus, therapeutic interventions are more effective against cancers with this profile than against cancers with any one of these variants.

[0191] A therapeutic intervention is an intervention that produces a therapeutic effect (e.g., is therapeutically effective). A therapeutically effective intervention prevents, slows the progression of, improves (e.g., induces remission) or cures a disease, such as cancer. Therapeutic interventions can include, for example, the administration of procedures such as chemotherapy, radiation therapy, surgery, or immunotherapy, the administration of drugs or nutritional supplements, or behavioral changes such as diet. One measure of therapeutic effectiveness is effectiveness in at least 90% of subjects receiving the intervention across at least 100 subjects.

[0192] Tables 1 and 2 list drug targets in cancer and drugs effective against these targets (excerpted from Bailey et al., Discovery Medicine, Vol. 18, #92, 2 / 7 / 14). [Table 1-1] [Table 1-2] [Table 2-1] [Table 2-2]

[0193] In one embodiment, a therapeutic intervention is determined based on a disease heterogeneity profile, taking into account both the type of gene variant found in diseased cells and its relative amount (e.g., ratio). The therapeutic intervention can treat the subject as if each clonal variant were a different cancer to be treated independently. In some cases, if one or more gene variants are detected at subclinical levels, for example, at least 5 × lower, at least 10 × lower, or at least 100 × lower than the predominantly detected clone, these variants can be excluded from the therapeutic intervention until they rise to a clinical threshold or a significant relative frequency (e.g., exceeding the thresholds mentioned above).

[0194] When multiple different gene variants are found in different amounts, e.g., different numbers or different relative amounts, therapeutic interventions may include treatments effective against the disease having each of the gene variants. For example, in cancer, gene variants can be detected in several genes (e.g., multiple clones and few clones), such as variant forms of the gene or gene amplification. Each of these forms may be available, i.e., a treatment may be known to be responsive to cancer having a particular variant. However, the tumor heterogeneity profile may indicate that one type of variant is present in polynucleotides at, for example, five times the level of each of the other two variants. A therapeutic intervention can be determined involving the delivery of three different drugs to a target, each drug being relatively more effective against the cancer having each of the variants. The drugs can be delivered as a cocktail or sequentially.

[0195] In further embodiments, the drug may be administered in stratified doses that reflect the relative amounts of variants in the DNA. For example, a drug effective against the most common variant may be administered in a larger dose than a drug effective against two less common variants.

[0196] Alternatively, the tumor heterogeneity profile may indicate the presence of a subpopulation of cancer cells with genetic variants resistant to drugs to which the disease typically responds. In this case, therapeutic intervention may involve steps including both a first drug effective against tumor cells without the resistant variant and a second drug effective against tumor cells with the resistant variant. Furthermore, doses can be stratified to reflect the relative amounts of each variant detected in the profile.

[0197] In another embodiment, changes in the profile of tumor heterogeneity are tested over time and treatment interventions are developed to treat the changing tumors. For example, disease heterogeneity can be determined at multiple different time points. Using the profiling methods of the present disclosure, more accurate estimations regarding tumor evolution can be made. In particular, new clonal subpopulations emerge after remission brought about by the first wave of a treatment regimen, which enables the practitioner to monitor the evolution of the disease. In this case, the treatment intervention can be calibrated over time to treat the changing tumors. For example, the profile can indicate that the cancer has a responsive form to a particular treatment. The treatment is delivered and the tumor burden is observed to decrease over time. At some point, genetic variants are found in the tumors that indicate the presence of a population of cancer cells that are not responsive to the treatment. Then, a new treatment intervention targeting cells with the non-responsive marker is determined.

[0198] In response to chemotherapy, the dominant tumor morphology may eventually be replaced by cancer cells harboring variants that render the cancer unresponsive to the treatment regimen, through Darwinian selection. The emergence of such resistant variants can be delayed by the method of this disclosure. In one embodiment of the method, the subject is subjected to one or more pulsed therapy cycles, each comprising a first period in which the drug is administered at a first dose and a second period in which the drug is administered at a second reduced dose. The first period is characterized by a tumor load detected above a first clinical level. The second period is characterized by a tumor load detected below a second clinical level. The first and second clinical levels may differ in different pulsed therapy cycles. For example, the first clinical level may be lower in subsequent cycles. The multiple cycles may include at least two, three, four, five, six, seven, eight or more cycles. For example, the BRAF variant V600E can be detected in diseased cell polynucleotides at a level indicating a 5% tumor burden in cfDNA. Chemotherapy can be initiated with dabrafenib. Subsequent testing may show that the amount of the BRAF variant in cfDNA has decreased to below 0.5% or to undetectable levels. At this point, dabrafenib treatment can be stopped or significantly shortened. Further subsequent testing may find that DNA with the BRAF mutation has increased to 2.5% of the polynucleotides in cfDNA. At this point, dabrafenib treatment can be restarted, for example, at the same level as the initial treatment. Subsequent testing may find that DNA with the BRAF mutation has decreased to 0.5% of the polynucleotides in cfDNA. Again, dabrafenib treatment is stopped or reduced. The cycle can be repeated many times.

[0199] Figure 7 shows an exemplary course of monitoring and treatment of a disease in a subject. The subject examined at the time of Blood Draw 1 has a tumor burden of 1.4% and presents genetic changes in Genes 1, 2, and 3. This subject is treated with Drug A. After some time, the treatment is interrupted. At the time point after the second, the second blood draw shows remission of the cancer. At the time point after the third, the third blood draw shows that the cancer has recurred and, in this case, presents a genetic variant in Gene 4. Therefore, this subject is placed on the course of Drug B for which cancer having this variant is responsive.

[0200] In another embodiment, the therapeutic intervention can be changed after detection of an increase in a variant type that is resistant to the original drug. For example, cancer having the EGFR mutation L858R responds to treatment with erlotinib. However, cancer having the EGFR mutation T790M is resistant to erlotinib. However, this is responsive to luxoritinib. The method of the present disclosure involves monitoring changes in the tumor profile and changing the therapeutic intervention if a gene variant associated with drug resistance rises to a predetermined clinical level.

[0201] Database In another embodiment, a database is constructed recording genetic information derived from sequential samples collected from cancer patients. This database may also include information on mediating treatments and other clinically relevant information, such as body weight, adverse effects, histological examination, blood tests, X-ray examinations, previous treatments, cancer type, etc. Using sequential test results, particularly when used with blood samples, the effectiveness of a treatment can be estimated, which can provide a less biased estimate of tumor burden than self-reported or physician-reported X-ray examinations. Treatment effectiveness can be clustered by those with similar genomic profiles, and vice versa. Genomic profiles can be organized, for example, around primary genetic changes, secondary genetic changes(s), the relative amounts of these genetic changes, and tumor burden. This database can be used to support subsequent patient decisions. Both germline and somatic changes can be used to determine treatment effectiveness. If a treatment that was initially effective begins to fail, acquired resistance changes can also be estimated from the database. This failure can be detected by X-ray, blood, or other means. The primary data used to estimate acquired resistance mechanisms is the genomic tumor profile collected after treatment for each patient. This data can be used to set quantitative limits on possible treatment responses and to predict the time to treatment failure. Based on the possible acquired resistance changes in a given treatment and tumor genomic profile, the treatment regimen can be modified to suppress the acquisition of the most likely resistance changes.

[0202] Computer system The methods of this disclosure can be implemented using or with the help of a computer system. Figure 5 shows a computer system 1501 programmed or otherwise configured to implement the methods of this disclosure. The computer system 1501 includes a central processing unit (CPU, also referred to herein as “processor” and “computer processor”) 1505. The computer system 1501 also includes memory or memory locations 1510 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 1515 (e.g., a hard disk), a communication interface 1520 for communicating with one or more other systems (e.g., a network adapter), and peripheral devices 1525 such as a cache, other memory, data storage and / or an electronic display adapter. The memory 1510, the storage unit 1515, the interface 1520 and the peripheral devices 1525 communicate with the CPU 1505 via a communication bus (solid line). The storage unit 1515 may be a data storage unit (or data repository) for storing data. Computer system 1501 can be operably coupled to computer network ("Network") 1530 with the help of communication interface 1520. Network 1530 may be the Internet, the Internet and / or an extranet, or an intranet and / or extranet connected to the Internet. Network 1530 may, in some cases, be a telecommunications and / or data network. Network 1530 may include one or more computer servers that can enable distributed computing, such as cloud computing. CPU 1505 can execute a set of machine-readable instructions that can be embodied in a program or software. Instructions can be stored in memory locations, such as memory 1510. Storage unit 1515 can store files, such as drivers, libraries, and saved programs.Computer system 1501 can communicate with one or more remote computer systems via network 1530. The methods described herein can be implemented by machine (e.g., computer processor) executable code stored in an electronic storage location of computer system 1501, such as memory 1510 or electronic storage unit 1515. Machine executable or machine-readable code can be provided in the form of software. Embodiments of systems and methods provided herein, such as computer system 1501, can be embodied in programming. Various embodiments of this technology can typically be considered as “products” or “manufactured goods” in the form of machine (or processor) executable code and / or related data performed or embodied in some kind of machine-readable medium. Machine executable code can be stored in an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or hard disk. The “memory” type medium can include all kinds of tangible memory or related modules of a computer, processor or similar, such as various semiconductor memories, tape drives, disk drives, etc., that can provide non-transient memory at any point for software programming. The whole or part of the software can be communicated at some point via the Internet or various other telecommunication networks. The computer system 1501 can include, or communicate with, an electronic display, for example, a user interface (UI) for providing one or more results of sample analysis. [Examples]

[0203] As shown in Figure 2, nucleotide positions in the genome (e.g., loci) can be specified by number. Positions where approximately 100% of base calls are identical to the reference sequence, or where approximately 100% of base calls are different from the reference sequence, are presumed to represent homozygosity of cfDNA (presumed to be normal). Positions where approximately 50% of base calls are identical to the reference sequence are presumed to represent heterozygosity of cfDNA (similarly presumed to be normal). Positions where the percentage of base calls at a locus is substantially less than 50% and is above the detection limit of the base calling system are presumed to represent tumor-related gene variants.

[0204] (Example 1) Method for detecting copy number variations blood sampling Collect a 10-30 mL blood sample at room temperature. Centrifuge the sample to remove cells. After centrifugation, collect the plasma.

[0205] cfDNA extraction The sample is subjected to proteinase K digestion. The DNA is precipitated with isopropanol. The DNA is captured using a DNA purification column (e.g., QIAamp DNA blood mini-kit) and eluted with 100 μl of solution. DNA smaller than 500 bp is selected using Ampure SPRI magnetic beads (PEG / salt). The resulting product is suspended in 30 μl of H2O. The size distribution is checked (major peak = 166 nucleotides; minor peak = 330 nucleotides) and quantified. 5 ng of extracted DNA contains approximately 1700 haploid genome equivalents ("HGE"). The general correlation between the amount of DNA and HGE is as follows: 3 pg DNA = 1 HGE; 3 ng DNA = 1K HGE; 3 μg DNA = 1M HGE; 10 pg DNA = 3 HGE; 10 ng DNA = 3K HGE; 10 μg DNA = 3M HGE.

[0206] "Single-molecule" library preparation High-efficiency DNA tagging (>80%) is performed by sticky-end ligation with two different octomers (i.e., four combinations) using end repair, A-tailing, and overloaded hairpin adapters. 2.5 ng of DNA (i.e., approximately 800 HGE) is used as starting material. Each hairpin adapter contains a random sequence in its non-complementary region. Hairpin adapters are attached to both ends of each DNA fragment. Each tagged fragment can be identified by the combination of the octomer sequence on the hairpin adapter and the endogenous portion of the inserted fragment sequence.

[0207] The tagged DNA is amplified by 12 cycles of PCR to produce approximately 1-7 μg of DNA, each containing roughly 500 copies of the 800 HGE in the starting material.

[0208] To optimize PCR reactions, buffer optimization, polymerase optimization, and cycle reduction can be performed. Amplification biases, such as nonspecific bias, GC bias, and / or size bias, can also be reduced through optimization. Noise(s) (e.g., errors introduced by the polymerase) can be reduced by using high-fidelity polymerases.

[0209] The sequence can be enriched as follows: Biotin-labeled beads are used as probes to the region of interest (ROI) to capture the DNA containing the ROI. The ROI is amplified by 12 cycles of PCR to produce a 2000-fold amplification.

[0210] Large-scale parallel sequencing determination 0.1–1% (approximately 100 pg) of the sample is used for sequencing. The resulting DNA is then denatured, diluted to 8 pM, and added to an Illumina sequencer.

[0211] Digital Bioinformatics The sequence readings are classified into families, with approximately 10 sequence readings in each family. The families are broken down into a consensus sequence by selecting positions within each family (e.g., biased selection). If 8 or 9 members match, a base is called for the consensus sequence. If less than 60% of the members match, no base is called for the consensus sequence.

[0212] The resulting consensus sequence is mapped to a reference genome such as HG19. Each base in the consensus sequence is covered by approximately 3000 different families. A quality score is calculated for each sequence, and sequences are filtered based on these quality scores. The base calls at each position in the consensus sequence are compared to the HG-19 reference sequence. For each position where the base call differs from the reference sequence, the identity of one or more different bases and their percentages corresponding to the total base call of the locus are determined and reported.

[0213] Sequence variations are detected by counting the distribution of bases at each locus. If 98% of the readings have the same bases (homozygous) and 2% have different bases, the locus likely contains a sequence variant from cancer DNA.

[0214] CNV is detected by counting the total number of sequences (bases) mapped to a locus and comparing it with a control locus. To increase CNV detection, perform CNV analysis in specific regions including regions on the ALK, APC, BRAF, CDKN2A, EGFR, ERBB2, FBXW7, KRAS, MYC, NOTCH1, NRAS, PIK3CA, PTEN, RB1, TP53, MET, AR, ABL1, AKT1, ATM, CDH1, CSF1R, CTNNB1, ERBB4, EZH2, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JAK3, KDR, KIT, MLH1, MPL, NPM1, PDGFRA, PROC, PTPN11, RET, SMAD4, SMARCB1, SMO, SRC, STK11, VHL, TERT, CCND1, CDK4, CDKN2B, RAF1, BRCA1, CCND2, CDK6, NF1, TP53, ARID1A, BRCA2, CCNE1, ESR1, RIT1, GATA3, MAP2K1, RHEB, ROS1, ARAF, MAP2K2, NFE2L2, RHOA or NTRK1 genes.

[0215] (Example 2) A method for correcting base calling by determining the total number of invisible molecules in a sample Amplify the fragment, read and align the sequences of the amplified fragments, and then subject the fragments to base calling. Variations in the number of amplified fragments and invisible amplified fragments can introduce errors in base calling. These variations are corrected by calculating the number of invisible amplified fragments.

[0216] When performing base calling for locus A (an arbitrary locus), first assume there are N amplified fragments. Sequence reads can come from two types of fragments: double-stranded fragments and single-stranded fragments. The following is a theoretical example of calculating the total number of invisible molecules in a sample.

[0217] N is the total number of molecules in the sample. Let's assume 1000 is the number of double-stranded fibers detected. Let's assume 500 is the number of single-stranded molecules detected. P is the probability of seeing the chain. Q is the probability that the chain will not be detected.

[0218] Q=1-P 1000 = NP(2) 500 = N2PQ 1000 / P(2)=N 500 ÷ 2PQ = N 1000 / P(2) = 500 ÷ 2PQ 1000 * 2PQ = 500P (2) 2000PQ = 500P(2) 2000Q = 500P 2000(1-P)=500P 2000 - 2000P = 500P 2000 = 500P + 2000P 2000 = 2500P 2000 ÷ 2500 = P 0.8 = P 1000 / P(2)=N 1000 ÷ 0.64 = N 1562=N Therefore, the number of invisible fragments = 62.

[0219] (Example 3) Identification of gene variants in cancer-associated somatic cell variants in patients. Assays are used to analyze a panel of genes in order to identify gene variants with high sensitivity in cancer-associated somatic cell variants.

[0220] Cell-free DNA was extracted from patient plasma and amplified by PCR. Gene variants were analyzed by large-scale parallel sequencing of the amplified target gene. For one set of genes, all exons were sequenced as such sequencing coverage demonstrated clinical utility (Table 3). For another set of genes, sequencing coverage included exons with previously reported somatic mutations (Table 4). The minimum detectable variant allele (detection limit) depended on the cell-free DNA concentration in the patient sample, which was the number of genomic equivalents that could vary from less than 10 to more than 1,000 per 1 mL of peripheral blood. Amplification was undetectable in samples with smaller amounts of cell-free DNA and / or low levels of gene copy amplification. Characteristics of certain samples or variants resulted in reduced analytical sensitivity, e.g., poor sample quality or inadequate collection.

[0221] The percentage of gene variants found in cell-free DNA circulating in the blood is related to the patient's unique tumor biology. Factors influencing the amount / percentage of gene variants detected in circulating cell-free DNA include tumor growth, turnover, size, heterogeneity, vascularization, disease progression, or treatment. Table 5 annotates the percentage (%cfDNA) or allele frequency of altered circulating cell-free DNA detected in this patient. Some of the detected gene variants are listed in descending order by %cfDNA.

[0222] Genetic variants are detected in circulating cell-free DNA isolated from blood samples of this patient. These genetic variants are cancer-associated somatic variants, some of which were associated with increased or decreased clinical response to specific treatments. "Minor changes" are defined as changes detected at less than 10% of the allele frequency of "major changes." Major changes are the dominant changes at the locus. The allele frequencies detected for these changes (Table 5) and the relevant treatments for this patient are noted.

[0223] All genes listed in Tables 3 and 4 are analyzed as part of the test. Amplification is not detected in ERBB2, EGFR, or MET in circulating cell-free DNA isolated from this patient's blood sample.

[0224] The patient test results, including those for gene variants, are shown in Table 6.

[0225] Regarding Table 4, the nucleotide detected at position 13, with a frequency of at least 98.8% in the samples, differs from the nucleotide in the reference sequence, indicating homozygosity of these loci. For example, in the KRAS gene, at position 25346462, T was detected instead of the reference nucleotide C in 100% of cases.

[0226] At position 35, nucleotides detected in 41.4% to 55% of samples differ from the nucleotides in the reference sequence, indicating heterozygosity at these loci. For example, in the ALK gene, at position 29455267, G was detected instead of the reference nucleotide A in 50% of cases.

[0227] Nucleotides detected at three positions with a frequency of less than 9% differ from the nucleotides in the reference sequence. These include variants in BRAF (140453136 A>T, 8.9%), NRAS (115256530 G>T, 2.6%), and JAK2 (5073770 G>T, 1.5%). These are presumed to be somatic mutations from cancer DNA.

[0228] The relative amounts of tumor-associated gene variants are calculated. The ratio of BRAF:NRAS:JAK2 amounts is 8.9:2.6:1.5 or 1:0.29:0.17. From this result, the presence of tumor heterogeneity can be inferred. For example, one possible interpretation is that 100% of tumor cells contain the variant in BRAF, 83% contain the variant in BRAF and NRAS, and 17% contain the variant in BRAF, NRAS, and JAK2. However, CNV analysis may show amplification of BRAF, in which case 100% of tumor cells could have variants in BRAF and NRAS. [Table 3] [Table 4] [Table 5] [Table 6]

[0229] (Example 4) Determining patient-specific detection limits for genes analyzed by assays The method of Example 3 is used to detect genetic alterations in the patient's cell-free DNA. The sequence readings of these genes include exon and / or intron sequences.

[0230] (Example 5) Corrects sequence errors when comparing Watson and Click sequences. Double-stranded cell-free DNA is isolated from patient plasma. Each cell-free DNA fragment is tagged using 16 different bubble-containing adapters, each containing a distinctive barcode. The bubble-containing adapters are attached to both ends of each cell-free DNA fragment by ligation. After ligation, each cell-free DNA fragment can be clearly identified by the distinct barcode sequence and two 20 bp endogenous sequences at each end of the cell-free DNA fragment.

[0231] Tagged cell-free DNA fragments are amplified by PCR. The amplified fragments are enriched using beads containing oligonucleotide probes that specifically bind to a group of cancer-related genes. Thus, cell-free DNA fragments from a group of cancer-related genes are selectively enriched.

[0232] Each method involves attaching a sequencing adapter containing a sequencing primer binding site, sample barcode, and cell flow sequence to a concentrated DNA molecule. The resulting molecule is then amplified by PCR.

[0233] Both strands of the amplified fragment are sequenced. Since each bubble-containing adapter contains a non-complementary region (e.g., a bubble), the sequence of one strand of the bubble-containing adapter differs from the sequence of the other strand (complement). Therefore, the sequence reading of an amplicon derived from the Watson strand of the original cell-free DNA can be distinguished from an amplicon from the click strand of the original cell-free DNA by the sequence of the attached bubble-containing adapter.

[0234] The sequence reading from one strand of the original cell-free DNA fragment is compared to the sequence reading from the other strand of the original cell-free DNA fragment. If the variant is present only in the sequence reading from one strand of the original cell-free DNA fragment and not on the other strand, this variant is identified as an erroneous gene variant (e.g., resulting from PCR and / or amplification) rather than a true one.

[0235] The sequence readings are classified into families. Errors in the sequence readings are corrected. The consensus sequence for each family is generated by decay.

[0236] (Example 6) therapeutic intervention To treat the cancer, a therapeutic intervention is determined. Cancers with BRAF variants respond to treatment with vemurafenib, regorafenib, tranetinib, and dabrafenib. Cancers with NRAS variants respond to treatment with trametinib. Cancers with JAK2 variants respond to treatment with ruxolitinib. A therapeutic intervention including the administration of trametinib and ruxolitinib is determined to be more effective against this cancer than treatment with either of the aforementioned drugs alone. Subjects are treated with a 5:1 dose ratio combination of trametinib and ruxolitinib.

[0237] After several treatments, cfDNA from the subjects is examined again for the presence of tumor heterogeneity. The results show that the BRAF:NRAS:JAK2 ratio is now approximately 4:2:1.5. This indicates that the therapeutic intervention reduced the number of cells with BRAF and NRAS variants and halted the growth of cells with JAK2 variants. A second therapeutic intervention is decided upon, in which trametinib and ruxolitinib are determined to be effective in a 1:1 dose ratio. The subjects are given a course of chemotherapy at this ratio. Subsequent examinations show that BRAF, NRAS, and JAK2 variants are present in the cfDNA at less than 1%.

[0238] (Example 7) therapeutic intervention Blood samples are taken from individuals undergoing melanoma pretreatment, and cell-free DNA analysis is used to determine that the patient has a BRAF V600E mutation at a concentration of 2.8% and no detectable NRAS mutations. The patient is placed on anti-BRAF therapy (dabrafenib). Three weeks later, another blood sample is taken and tested. It is determined that the BRAF V600E level has decreased to 0.1%. Therapy is stopped, and testing is repeated every two weeks. The BRAF V600E level rises again, and therapy is restarted when the BRAF V600E level rises to 1.5%. Therapy is stopped again when the level falls back to 0.1%. This cycle is repeated.

[0239] (Example 8) Correct CNV based on ROCNV measurements. The copy number variations in patient samples are determined. Methods for determination may include molecular tracking and upsampling, as described above. A hidden Markov model based on the expected location of the origin of replication is used to remove the influence of origin proximity from the estimated copy number variations in patient samples. The standard deviation of copy number variations for each gene is then reduced by 40%. The origin of replication proximity model is also used to estimate the acellular tumor burden in patients.

[0240] In many cases, the level of acellular tumors obtained may be low or below the detection limit of certain techniques. This may be because the number of human genome equivalents of tumor-derived DNA in the plasma may be less than one copy per 5 mL. Radiation and chemotherapy have been shown to be more effective in treating patients with advanced cancer because they affect rapidly dividing cells more than stable, healthy cells. Therefore, to preferentially increase the fraction of tumor-derived DNA collected, the patient should be given treatment with minimal adverse effects before blood collection. For example, a low dose of chemotherapy can be administered to the patient, and a blood sample can be collected within 24, 48, 72 hours, or less than one week. For effective chemotherapy, this blood sample will contain a higher concentration of acellular tumor-derived DNA due to a potentially higher rate of cancer cell death. Alternatively, instead of low-dose chemotherapy, low-dose radiotherapy may be applied to the affected area by whole-body radiography or locally. Other treatments are considered, including subjecting the patient to ultrasound, sound waves, exercise, stress, etc.

[0241] All publications, patents, and patent applications pointed out in this specification are incorporated herein by reference to the same extent that each individual publication, patent, or patent application is specifically and individually indicated as being incorporated by reference.

[0242] While preferred embodiments of the present disclosure are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. Without departing from the present disclosure, those skilled in the art can now conceive of many variations, modifications, and substitutions. It should be understood that various substitutions for the embodiments of the present invention described herein can be used in the practice of the invention. The following claims define the scope of the present invention, and methods and structures within the scope of these claims, as well as their equivalents, are intended to be included therein.

Claims

[Claim 1] The method or kit described in the specification.