Population-based treatment recommendations using cell-free DNA

JP7899266B2Active Publication Date: 2026-08-03GUARDANT HEALTH INC
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GUARDANT HEALTH INC
Filing Date
2024-08-07
Publication Date
2026-08-03

AI Technical Summary

Benefits of technology

【0027】 本開示の新規な特徴は、添付の特許請求の範囲に具体的に記載される。本開示の特性および有利点のより良い理解は、本開示の原理が利用される例示的実施形態を記載する以下の詳細な記載および添付の図を参照することによって得られる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007899266000001
    Figure 0007899266000001
  • Figure 0007899266000002
    Figure 0007899266000002
  • Figure 0007899266000003
    Figure 0007899266000003
Patent Text Reader

Abstract

To provide a system and method for classifying cancer patients based on predicted therapeutic response.SOLUTION: A method creates a treatment response prediction or detects a disease by the steps of: creating genetic information using a genetic analyzer; receiving in a computer memory, for each of a plurality of individuals having the disease, a training dataset including (1) genetic information derived from the individual created at a first time point and (2) treatment response of the individual to one or more therapeutic interventions determined at a second, later time point; and implementing a machine learning algorithm using the dataset to create at least one computer-implemented classification algorithm that predicts treatment response for a subject to the therapeutic intervention based on the genetic information derived from the subject.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference This application claims the benefit of U.S. Provisional Application No. 62 / 239,390, filed Oct. 9, 2015, which is hereby incorporated by reference in its entirety.

Background Art

[0002] Background Individual patients respond differently to medical treatments, in part due to genetic and epigenetic differences that affect gene expression. These differences may be present in normal host tissue or may be acquired by cancer cells upon transformation. Such differences can affect various components of treatment response, including the pharmacokinetics (e.g., metabolism or transport) or pharmacodynamics (e.g., target or enzyme modulation) of drugs; host tissue sensitivity to radiation; sensitivity of malignant cells to cytotoxic agents, including drugs and radiation; and the ability of malignant cells to invade and metastasize.

Summary of the Invention

Problems to be Solved by the Invention

[0003] One reason treating cancer is extremely difficult is that current diagnostic methods often fail to help physicians tailor effective drug treatments to specific cancers. Furthermore, the disease state itself can be a fluctuating target—cancer cells are constantly changing and mutating. Cancer tumors continuously release their unique genomic material into the bloodstream, but unfortunately, these definitive genomic "signals" are so weak that current genomic analysis techniques, including next-generation sequencing, can only detect such signals in patients with sporadic or terminally high tumor burdens. The main reason for this is that such techniques suffer from error rates and biases that can be orders of magnitude higher than what is needed to reliably detect de novo genomic changes associated with cancer. Therefore, improvements in systems and methods for determining effective treatments against cancer are needed. [Means for solving the problem]

[0004] Summary of the Invention The present invention relates to a system and method for classifying cancer patients based on their predicted treatment response.

[0005] In one embodiment, the present disclosure provides a method for analyzing a disease state of a subject, comprising the steps of: characterizing the subject's genetic information at two or more time points using a gene analyzer, such as a DNA sequencing machine; and generating a test result adjusted in characterizing the subject's genetic information using information from two or more individuals or at two or more time points.

[0006] In another embodiment, a system and method for detecting a disease are disclosed, comprising the steps of: creating genetic information using a DNA sequencing device; receiving into computer memory a training dataset for each of several individuals having a disease, which includes (1) individual-derived genetic information created at a first time point, and (2) an individual's treatment response to one or more therapeutic interventions determined at a second, later time point; and performing a machine learning algorithm using the dataset to create at least one computer-performed classification algorithm, wherein the classification algorithm predicts the subject's treatment response to a therapeutic intervention based on the subject-derived genetic information. Where used herein, treatment response is a treatment response to a specific therapeutic intervention.

[0007] In another embodiment, the method detects the temporal trend of the amount of cancer polynucleotides in a sample derived from the target by the steps of: determining the frequency of occurrence of cancer polynucleotides at multiple time points; determining the error range for the frequency of occurrence at each of the multiple time points; and determining whether the error ranges between earlier and later time points (1) overlap and indicate stability of the frequency of occurrence, (2) increase outside the error range at the later time point and indicate an increase in the frequency of occurrence, or (3) decrease outside the error range at the later time point and indicate a decrease in the frequency of occurrence.

[0008] In yet another embodiment, the method detects abnormal cellular activity by the steps of: sequencing cell-free nucleic acids with a gene analyzer, e.g., a DNA sequencing machine; comparing a later (e.g., most recent) sequence read with past sequence reads from at least two point in time, thereby updating the diagnostic accuracy index; and detecting the presence or absence of genetic changes and / or the amount of genetic variation in the individual based on the diagnostic accuracy index of the sequence reads. The gene analyzer includes any system for genetic analysis, e.g., sequencing (DNA sequencing machine) or hybridization (microarray, fluorescence in situ hybridization, bionanogenomics).

[0009] In another embodiment, the method involves the steps of creating a consensus sequence by comparing a later (e.g., most recent) sequence read from a gene analyzer, such as a DNA sequencing machine, with a past sequence read from a past period; updating a diagnostic accuracy index based on the past sequence reads, wherein each consensus sequence corresponds to a unique polynucleotide in a set of tagged parent polynucleotides; and creating a gene profile of extracellular polynucleotides in the subject, wherein the gene profile includes multiple data derived from copy number variation or mutation analysis; and detecting mutations in a cell-free or substantially cell-free sample obtained from the subject.

[0010] In another embodiment, a method for detecting abnormal cellular activity by the steps of: providing at least one set of tagged parental polynucleotides; and for each set of tagged parental polynucleotides; amplifying the tagged parental polynucleotides in the set to create a corresponding set of amplified progeny polynucleotides; sequencing a subset of the amplified progeny polynucleotide set using a gene analyzer, e.g., a DNA sequencing machine, to create a set of sequencing reads; and collapsing the set of sequencing reads to generate a set of consensus sequences by comparing the most recent sequence read with past sequence reads from at least one past period; and thereby updating a diagnostic accuracy index, wherein each consensus sequence corresponds to a unique polynucleotide in the set of tagged parental polynucleotides.

[0011] In yet another embodiment, the method detects mutations in cell-free or substantially cell-free samples obtained from a subject by: (b) sequencing extracellular polynucleotides from a body sample derived from the subject using a gene analyzer, such as a DNA sequencing machine; creating multiple sequencing reads for each extracellular polynucleotide; excluding reads that do not meet a set threshold; mapping the sequenced sequence reads to a reference sequence; identifying a subset of mapped sequence reads that are compared with variants of the reference sequence at each mappable base position; for each mappable base position, calculating the ratio of the number of mapped sequence reads containing the variant compared to the reference sequence (a) to the total number of sequence reads for each mappable base position; and comparing the most recent sequence read with past sequence reads from at least one other point in time, thereby updating the accuracy index of the diagnosis.

[0012] In a further embodiment, a method for characterizing heterogeneity of abnormal conditions in a subject is disclosed herein, comprising the steps of: comparing a later (e.g., most recent) sequence read with a past sequence read from at least one other point in time; and thereby updating a diagnostic accuracy index, wherein each consensus sequence corresponds to a unique polynucleotide in a set of tagged parent polynucleotides; and creating a gene profile of extracellular polynucleotides in the subject, wherein the gene profile includes multiple data derived from copy number variation or mutation analysis.

[0013] The implementation of the above system / method may include one or more of the following: The method includes the step of creating a gene profile of extracellular polynucleotides in the subject, wherein the gene profile includes multiple data obtained from copy number variation or mutation analysis. The method includes the step of increasing the diagnostic confidence index in subsequent characterization when information from a first time point confirms information from a second time point. The diagnostic confidence index may be increased in subsequent characterization when information from a first time point confirms information from a second time point. The method includes the step of decreasing the diagnostic confidence index in subsequent characterization when information from a first time point contradicts information from a second time point.

[0014] The advantages of the above system may include one or more of the following: Tumor-derived copy number abnormalities, single nucleotide polymorphisms, and methylation changes can be detected using the present invention, and the information may be applied to a population to improve accuracy. Similarly, the system can identify treatments for genetically similar cases, including analysis of therapeutic targets and gene alterations conferring drug resistance in circulating tumor cells (CTCs) and cell-free circulating tumor DNA (ctDNA) released into peripheral blood. Both CTCs and ctDNA provide complementary information for evaluating novel drugs and drug combinations in a population. The liquid biopsy concept contributes to a better understanding and clinical management of drug resistance in patients with cancer. Counting and characterizing circulating tumor cells (CTCs) in peripheral blood can provide important prognostic information and can help monitor the effectiveness of treatment. Since current assays cannot distinguish between apoptotic and viable CTCs, the fluoroEPISPOT assay, which detects protein secretion / release / efflux from a single epithelial cancer cell, can be used for breast, colon, prostate, head and neck, and ovarian cancers, as well as melanoma. The system enables any high-throughput technology (e.g., planar and bead microarrays, microfluidic quantitative PCR, Luminex bead technology) to meet the specific needs and challenges of diagnostic biomarker discovery and validation in bodily fluids. Diagnostic marker panels based on autoantibodies and DNA methylation for the four major cancer entities (breast, colon, prostate, and lung) can be performed in serum or plasma. The system can handle blood, urine, or saliva as the diagnostic matrix. The genomic analysis technology structure reduces noise and distortion caused by next-generation sequencing to near zero. Digital sequencing enables ultra-high fidelity, single-molecule detection of actionable tumor-specific genomic alterations in cancer with unparalleled specificity and breadth. In other words, clinicians can now non-invasively observe the genomic dimension of cancer throughout a patient's body. The system comprehensively detects resistance and susceptibility mutations outside of tumor biopsies. Simple blood tests can examine numerous genes, including SNVs, CNVs, indels, and rearrangements spanning multiple base pairs, aiding in treatment management.The system also allows the spread of cancer genomes to be used to guide treatment regimens for patients. Correlations between drug treatment efficacy and the presence or absence of molecular markers in patient samples can be used to improve treatment. The resulting recommendations and clinical reports are intuitive for understanding, require a basic level of knowledge of the testing process, and necessitate familiarity with the scientific terminology used to describe the test results. This is done without requiring additional educational materials explaining the instructions for the tests and the interpretation of the test results. The laboratory reports focus on critical information to assist treatment professionals in understanding the information and correctly applying it to clinical practice. The reports facilitate the correct interpretation of complex DNA testing information. Improved communication of genetic test results leads to a reduction in misinterpretation of genetic test results and improves the delivery of interventions or treatments based on DNA sequencing results.

[0015] In one aspect, the present disclosure provides a method for creating a therapeutic response predictor, comprising the steps of: creating genetic information using a gene analyzer; receiving into computer memory a training dataset for each of several individuals having a disease, which includes (1) individual-derived genetic information created at a first time point, and (2) the individual's treatment response to one or more therapeutic interventions determined at a second, later time point; and performing a machine learning algorithm using the dataset to create at least one computer-performed classification algorithm, wherein the classification algorithm predicts the therapeutic response of the subject based on the subject-derived genetic information.

[0016] In some embodiments, the machine learning algorithm is selected from a group consisting of supervised or unsupervised learning algorithms, selected from support vector machines, random forests, nearest neighbor analysis, linear regression, binary decision trees, discriminant analysis, logistic classification metrics, and cluster analysis. In some embodiments, the method includes a step of predicting the direction of tumor development based on examinations at two or more time points. In some embodiments, the predictions made include a step of determining the probability of distant metastasis development. In some embodiments, the training dataset further includes clinical data selected from a group consisting of cancer stage, type of surgical procedure, age, tumor grade, depth of tumor invasion, occurrence of postoperative complications, and presence of vascular invasion. In some embodiments, genetic information includes variables that define the genomic composition of cancer cells. In some embodiments, genetic information includes variables that define the genomic composition of a single disseminated cancer cell. In some embodiments, the method includes a step of preprocessing the training dataset. In some embodiments, the step of preprocessing the training dataset includes a step of converting the provided data into class-conditional probabilities.

[0017] In some embodiments, the genetic information includes sequence or abundance data from one or more gene loci in cell-free DNA derived from the individual. In some embodiments, the treatment response includes individual-derived genetic information created at a second, later point in time. In some embodiments, the disease state is cancer, and the gene analyzer is a DNA sequencing device.

[0018] In one aspect, the disclosure provides a method comprising the steps of: using a gene analyzer to create genetic information about a subject; receiving a test dataset containing the genetic information into computer memory; and performing a computer-implemented classification algorithm, wherein the classification algorithm predicts the subject's therapeutic response to a therapeutic intervention based on the genetic information.

[0019] In some embodiments, the method includes a step of predicting tumor development. In some embodiments, the method includes a step of predicting the development of distant metastases. In some embodiments, the training dataset further includes variables selected from the group consisting of cancer stage, type of surgical procedure, age, tumor grade, depth of tumor invasion, occurrence of postoperative complications, and presence of vascular invasion. In some embodiments, genetic information includes variables that define the genomic composition of cancer cells. In some embodiments, genetic information includes variables that define the genomic composition of a single disseminated cancer cell. In some embodiments, the method includes a step of preprocessing the test dataset. In some embodiments, the step of preprocessing the test dataset includes a step of converting the provided data into class-conditional probabilities. In some embodiments, 20 or fewer variables are selected. In some embodiments, 10 or fewer variables are selected. In some embodiments, the classification algorithm uses an artificial neural network. In some embodiments, the artificial neural network is trained using a Bayesian framework.

[0020] In one embodiment, the Disclosure provides a method for analyzing a disease state of a subject, comprising the steps of: receiving data from a gene analyzer about the subject's genetic information at two or more time points; using the information from two or more time points to create a test result adjusted in characterizing the subject's genetic information; identifying subjects from a population that fit the genetic information; and recommending treatment based on past treatments of subjects that fit the genetic information. In some embodiments, the method includes the step of comparing the most recent sequence read with past sequence reads and thereby updating a diagnostic confidence index. In some embodiments, the method includes the step of creating a confidence interval for the most recent sequence read. In some embodiments, the method includes the step of comparing the confidence interval with one or more past confidence intervals and determining disease progression based on overlapping confidence intervals.

[0021] In some embodiments, the method includes a step of increasing the accuracy index of the diagnosis in subsequent or earlier characterization when information from a first time point confirms information from a second time point. In some embodiments, characterization includes a step of determining the frequency of occurrence of one or more genovariates detected in a collection of sequence reads from DNA in a sample derived from a subject, and the step of creating a adjusted test result includes a step of comparing the frequency of occurrence of one or more genovariates at two or more time points for subjects that fit the genetic information. In some embodiments, characterization includes a step of determining the amount of copy number variation at one or more loci detected in a collection of sequence reads from DNA in a sample derived from a fitted subject, and the step of creating a adjusted test result includes a step of comparing the amounts at two or more time points. In some embodiments, characterization includes a step of making a diagnosis of health or disease.

[0022] In some embodiments, the genetic information includes sequence data derived from a portion of the genome containing disease-related or cancer-related gene variants. In some embodiments, the method includes a step of increasing the sensitivity to detect gene variants by increasing the read depth of polynucleotides in the sample from the target at two or more time points. In some embodiments, characterization includes a step of diagnosing the presence of disease polynucleotides in the sample from the target, and adjustment includes a step of adjusting the diagnosis from negative or uncertain to positive if the same gene variant is detected within the noise range of multiple sample collections or time points. In some embodiments, characterization includes a step of diagnosing the presence of disease polynucleotides in the sample from the target, and adjustment includes a step of adjusting the diagnosis from negative or uncertain to positive in characterization from an earlier time point if the same gene variant is detected within the noise range at an earlier time point and beyond the noise range at a later time point. In some embodiments, characterization includes a step of diagnosing the presence of disease polynucleotides in the sample from the target, and adjustment includes a step of adjusting the diagnosis from negative or uncertain to positive in characterization from an earlier time point if the same gene variant is detected within the noise range at an earlier time point and beyond the noise range at a later time point.

[0023] In one embodiment, the disclosure provides a method comprising the steps of: a) providing a plurality of nucleic acid samples derived from a subject, wherein the samples are collected at consecutive time points; b) sequencing polynucleotides derived from the samples; c) determining the quantitative measurement of each of a plurality of somatic variants within the polynucleotides in each sample; d) graphing the relative amounts of these somatic variants at each of the consecutive time points for those present in non-zero amounts at at least one of the consecutive time points; and e) relating variants from a group of genetically similar subjects and formulating treatment recommendations based on past treatment data for genetically similar subjects.

[0024] In one aspect, the present disclosure provides a method for recommending cancer treatment from data generated by a genetic analyzer, the method comprising: identifying one or more subjects whose genetic profiles match from a population of people suffering from cancer, and retrieving past treatment data from the matching subjects; and identifying the best treatment option based on the medical history of the matching subjects; providing the recommendation in a paper or electronic patient examination report. In some embodiments, the method includes using a combination of the magnitudes of genomic changes detected in a body fluid-based test to estimate the disease burden. In some embodiments, the method includes using the allele frequency, allelic imbalance, or gene-specific scope of the detected mutations to estimate the disease burden. In some embodiments, the overall stack height represents the overall disease burden or disease burden score in an individual. In some embodiments, different colors are used to represent each genomic change. In some embodiments, only a subset of the detected changes is plotted. In some embodiments, the subset is selected based on the likelihood of being a driver change or its association with an increase or decrease in response to treatment. In some embodiments, the method includes creating an examination report for the genomic test. In some embodiments, a non-linear scale is used to represent the height or width of each representative genomic change. In some embodiments, a plot of points from previous tests is presented in the report. In some embodiments, the method includes estimating the progression or remission of the disease based on the rate of change and / or quantitative accuracy of each test result. In some embodiments, the method includes displaying treatment interventions between intervention examination points. The present invention provides, for example, the following items. (Item 1) A method for creating a treatment response predictor, comprising: using a genetic analyzer to create genetic information; For each of a plurality of individuals having a disease, a step of receiving in a computer memory a training data set including (1) gene information derived from the individual created at a first time point, and (2) the treatment response of the individual to one or more treatment interventions determined at a second, later time point; and A step of performing a machine learning algorithm using the data set to create at least one computer-performed classification algorithm, wherein the classification algorithm predicts the treatment response of the subject based on gene information from the subject. A method comprising. (Item 2) The method according to item 1, wherein the machine learning algorithm is selected from the group consisting of supervised or unsupervised learning algorithms selected from support vector machines, random forests, nearest neighbor analysis, linear regression, binary decision trees, discriminant analysis, logistic classification metrics, and cluster analysis. (Item 3) The method according to item 1, comprising a step of predicting the direction of tumor development based on examinations at two or more time points. (Item 4) The method according to item 1, wherein the prediction made includes a step of determining the probability of distant metastasis development. (Item 5) The method according to item 1, wherein the training data set further includes clinical data selected from the group consisting of cancer stage, type of surgical procedure, age, tumor grade, depth of tumor invasion, occurrence of postoperative complications, and presence of vascular invasion. (Item 6) The method according to item 1, wherein the gene information includes variables defining the genomic constitution of cancer cells. (Item 7) The method according to item 1, wherein the gene information includes variables defining the genomic constitution of a single seeding cancer cell. (Item 8) The method according to item 1, comprising a step of preprocessing the training data set. (Item 9) The method according to item 8, wherein the step of preprocessing the training dataset includes the step of converting the provided data into class-conditional probabilities. (Item 10) The method according to item 1, wherein the genetic information includes sequence or abundance data from one or more gene loci in cell-free DNA derived from the individual. (Item 11) The method according to item 1, wherein the treatment response includes a second, later-prepared genetic information derived from the individual. (Item 12) The method according to item 1, wherein the disease state is cancer and the gene analyzer is a DNA sequencing device. (Item 13) A step of creating genetic information about the subject using a gene analyzer; The steps of receiving the test dataset containing the aforementioned genetic information into computer memory; and A step of implementing a computer-based classification algorithm, wherein the classification algorithm predicts the therapeutic response of the subject to a therapeutic intervention based on the genetic information. A method that includes this. (Item 14) The method described in item 13, including a step for predicting tumor development. (Item 15) The method described in item 13, which includes a step to predict the development of distant metastasis. (Item 16) The method according to item 13, wherein the training dataset further includes variables selected from groups consisting of cancer stage, type of surgical procedure, age, tumor grade, depth of tumor invasion, occurrence of postoperative complications, and presence of vascular invasion. (Item 17) The method according to item 13, wherein the genetic information includes variables that define the genomic composition of cancer cells. (Item 18) The method according to item 13, wherein the genetic information includes variables that define the genomic composition of a single disseminated cancer cell. (Item 19) The method according to item 13, comprising the step of preprocessing the aforementioned test dataset. (Item 20) The method according to item 19, wherein the step of preprocessing the test dataset includes the step of converting the provided data into class conditional probabilities. (Item 21) The method described in item 13, wherein 20 or fewer variables are selected. (Item 22) The method described in item 13, wherein 10 or fewer variables are selected. (Item 23) The method described in item 13, wherein the classification algorithm uses an artificial neural network. (Item 24) The method described in item 23, wherein the artificial neural network is trained using a Bayesian framework. (Item 25) A method for analyzing the disease state of a target, A step of receiving data from a gene analyzer about the subject's genetic information at two or more time points; A step of creating a test result that is adjusted in characterizing the genetic information of the subject using the information from the aforementioned two or more time points; Steps to identify individuals from a population that match the genetic information; and A step to recommend treatment based on past treatments of the subject that match the genetic information. Methods that include... (Item 26) The method described in item 25, comprising the steps of comparing the most recent sequence reads with past sequence reads and thereby updating the diagnostic accuracy index. (Item 27) The method described in item 25, which includes the step of creating a confidence interval for the most recent sequence read. (Item 28) The method according to item 25, comprising the steps of comparing the confidence interval with one or more past confidence intervals and determining disease progression based on overlapping confidence intervals. (Item 29) The method of item 25, which includes a step of increasing the diagnostic accuracy index in subsequent or prior characterization when information from a first time point confirms information from a second time point. (Item 30) The method according to item 25, wherein characterization includes the step of determining the frequency of occurrence of one or more gene variants detected in a collection of sequence reads from DNA in a sample derived from the subject, and the step of creating a reconciled test result includes the step of comparing the frequency of occurrence of the one or more gene variants at two or more time points for the subject that fits the genetic information. (Item 31) The method according to item 25, wherein characterization includes the step of determining the amount of copy number variation at one or more loci detected from a collection of sequence reads from DNA in a sample derived from a matched subject, and the step of creating a refined test result includes the step of comparing the amounts at the two or more time points. (Item 32) The method described in item 25, wherein the characterization includes steps for making a diagnosis of health or disease. (Item 33) The method described in item 25, wherein the genetic information includes sequence data derived from a portion of the genome containing disease-related or cancer-related gene variants. (Item 34) The method according to item 25, comprising the step of increasing the sensitivity for detecting gene variants by increasing the polynucleotide read depth in the sample derived from the subject at two or more time points. (Item 35) The method according to item 25, wherein characterization includes the step of diagnosing the presence of disease polynucleotides in a sample derived from the subject, and adjustment includes the step of adjusting the diagnosis from negative or indeterminate to positive if the same gene variant is detected in a noise range of multiple sample collections or time points. (Item 36) The method of item 25, wherein characterization includes the step of diagnosing the presence of disease polynucleotides in a sample derived from the subject, and adjustment includes the step of adjusting the diagnosis from negative or uncertain to positive in the characterization from the earlier point in time if the same gene variant is detected within the noise range at an earlier point in time and beyond the noise range at a later point in time. (Item 37) a) A step of providing multiple nucleic acid samples derived from a target, wherein the samples are collected at consecutive points in time; b) A step of sequencing the polynucleotides derived from the sample; c) A step of determining the quantitative measurement of each of the multiple somatic cell variants within the polynucleotide in each sample; d) A step of graphing the relative amounts of somatic mutations at each of the consecutive time points for somatic mutations that are present in a non-zero amount at at least one of the consecutive time points; and e) A step of associating variants from groups of genetically similar subjects and creating treatment recommendations based on past treatment data for the genetically similar subjects. A method that includes this. (Item 38) A method for recommending cancer treatment based on data created by a gene analyzer, The steps include identifying one or more individuals from a population of people with cancer whose genetic profiles match, and searching for historical treatment data derived from the matched individuals; and A step to identify the best treatment option based on the prior medical history of the aforementioned suitable subject; Steps to provide recommendations on paper or electronic patient examination reports. Methods that include... (Item 39) The method of item 38, which includes the step of using a combination of the magnitudes of genomic changes detected by fluid-based tests to estimate the disease burden. (Item 40) The method according to item 38, comprising the step of using the allele proportion, allele imbalance, or gene-specific scope of the detected mutations to estimate the disease burden. (Item 41) The method described in item 38, where the total stack height represents the overall disease burden or disease burden score for the individual. (Item 42) The method described in item 38, where different colors are used to represent each genomic change. (Item 43) The method described in item 38, in which only a subset of the detected changes is plotted. (Item 44) The method described in item 43, wherein a subset is selected based on the possibility that it is a change in the driver, or in relation to an increase or decrease in response to the treatment. (Item 45) The method described in item 38, including the step of creating a test report for genomic testing. (Item 46) The method described in item 38, in which a nonlinear scale is used to represent the height or width of each representative genomic change. (Item 47) The method described in item 38, in which the plot of points from previous inspections is shown in the report. (Item 48) The method according to item 38, comprising the step of estimating disease progression or remission based on the rate of change or quantitative accuracy of each test result. (Item 49) The method described in item 38, which includes a step of displaying a therapeutic intervention between intervention test points.

[0025] Other purposes of this disclosure may become apparent to those skilled in the art by reading the subsequent specification and claims. Embedding by reference

[0026] All publications, patents, and patent applications described herein are incorporated herein by reference to the same extent as each individual publication, patent, or patent application is specifically and individually indicated as being incorporated by reference.

[0027] Novel features of this disclosure are specifically described in the appended claims. A better understanding of the characteristics and advantages of this disclosure can be obtained by referring to the following detailed description and appended drawings which describe exemplary embodiments in which the principles of this disclosure are utilized. [Brief explanation of the drawing]

[0028] [Figure 1A] Figure 1A shows a gene-based cancer treatment system based on an exemplary population.

[0029] [Figure 1B] Figure 1B shows a diagram of gene-based cancer treatments based on exemplary recommended populations.

[0030] [Figure 1C] Figure 1C shows an exemplary system for recommending treatment based on test results.

[0031] [Figure 1D] Figure 1D illustrates an exemplary process for reducing error rates and biases in DNA sequence reading and generating genetic reports for users based on population test results.

[0032] [Figure 2A] Figures 2A and 2B illustrate exemplary steps for reporting genetic test results to users and recommending actions based on population data. [Figure 2B] Figures 2A and 2B illustrate exemplary steps for reporting genetic test results to users and recommending actions based on population data.

[0033] [Figure 2C]Figures 2C-2I show the contents of an example genetic testing report. [Figure 2D] Figures 2C-2I show the contents of an example genetic testing report. [Figure 2E] Figures 2C-2I show the contents of an example genetic testing report. [Figure 2F] Figures 2C-2I show the contents of an example genetic testing report. [Figure 2G] Figures 2C-2I show the contents of an example genetic testing report. [Figure 2H] Figures 2C-2I show the contents of an example genetic testing report. [Figure 2I] Figures 2C-2I show the contents of an example genetic testing report.

[0034] [Figure 3A] Figures 3A-3B illustrate an exemplary process for detecting mutations and reporting test results to the user. [Figure 3B] Figures 3A-3B illustrate an exemplary process for detecting mutations and reporting test results to the user. [Modes for carrying out the invention]

[0035] Detailed description of the invention Cancer is a particularly heterogeneous disease, both in terms of the sheer number of cancer types and how specific types manifest in individuals. Therefore, predicting the best course of treatment for a given patient is difficult. This disclosure provides systems and methods for improving treatment outcomes for cancer patients.

[0036] Referring to Figure 1A, a population-based gene-based cancer treatment system is shown. In one embodiment, the system extracts cell-free DNA (cfDNA) from a population of cancer subjects or patients (2). Extraction is performed using genetic data captured from patients receiving treatment or from healthy individuals. Once the data is extracted, the system can recommend treatment based on past outcomes and by adapting the treatment to the genetic characteristics of the subject / patient. First, the system obtains subject criteria, including genetic characteristics (4). Next, the system identifies similar subjects with similar genetic characteristics (6). Then, the system identifies successful treatments from these similar subjects (8). Based on past treatments and outcomes for similar subjects, the system identifies a recommended treatment for the current subject (10).

[0037] Next, the system iteratively monitors the treatment process, which is done through subsequent gene reading (12). Based on the reading, the system identifies the best-suited treatment and recommends treatment based on the outcome and subsequent gene analysis (14). The system then tracks whether the patient has a positive outcome (16). If the patient does not heal, additional treatment is performed based on recommendations, which loop back to 12, or the patient is released (18).

[0038] Figure 1B shows an exemplary recommender.290 In this system, clinical information210 is stored in a database array. For example, the system can store patient information from physicians and laboratories in this database. Text data220, such as genome sequences, and histological reports for each patient are also captured. For example, the data may come from the cBio Cancer Genomics Portal (http: / / cbioportal.org), an open-access resource for interactive exploration of multidimensional cancer genome datasets, which currently provides access to data from over 5,000 tumor samples from 20 cancer studies. The cBio Cancer Genomics Portal significantly reduces the barriers between complex genomic data and cancer researchers who need rapid, intuitive, and high-quality access to molecular profiles and clinical characteristics from large-scale cancer genomics projects, enabling researchers to translate these rich datasets into biological insights and clinical applications. Image data 230, such as CT scans, may be captured along with other information such as MRI scans, ultrasound scans, bone scans, PET scans, bone marrow examinations, barium X-rays, endoscopy, lymphangiography, IVU (intravenous urography) or IVP (intravenous pyelography), lumbar puncture, cystoscopy, immunological tests (anti-malignin antibody screening), and cancer marker tests. Features are then extracted by an extractor 240. The features may then be used by one or more classification indices, such as a neural network 250, a vector machine 260, and a hidden Markov machine (HMM) 270. In some embodiments, the neural network is trained using a Bayesian framework. The output of the classification indices is then provided to an inference unit or engine 280. The results are provided as the output of a recommender 290, and the results are used in a report by the report generator 21 in Figure 1C. In some embodiments, data from two or more of the above categories may be used to create a more stable classification than data from a single category.

[0039] In some embodiments, text not constructed is obtained from available histological reports. The text is first normalized to reduce nucleotide polymorphisms: acronyms, number and size formats are standardized, related abbreviations are expanded, variants are mapped to common forms, and any non-informative strings are removed. The set of normalization rules is coded using ordinary expressions and implemented using simple search and replace operations.

[0040] Some embodiments utilize classification indices based on characteristics trained on validated somatic mutation samples, while benefiting from other available information such as base quality, mapping quality, strand bias, and tail distance. Given paired normal / tumor BAM files, the embodiments output the probability that each candidate site is somatic. Through the systems and methods described herein, the disclosure provides methods for classifying treatment responses to therapeutic interventions and then determining whether a given individual falls into a specific classification (e.g., a specific level of response such as responsive to treatment, unresponsive to treatment, fully responsive, or partially responsive).

[0041] In some embodiments, the method is provided for creating a trained classification index, comprising the steps of: (a) providing a plurality of different classes, each representing a set of subjects that share a common feature (e.g., derived from one or more cohorts); (b) providing a representative multi-parameter model of cell-free DNA molecules derived from each of a plurality of samples belonging to each of the classes, thereby providing a training dataset; and (c) training a learning algorithm on the training dataset to create one or more trained classification indices, each trained classification index classifying a test sample into one or more of the plurality of classes.

[0042] For example, the trained classification metric can use a learning algorithm selected from a group consisting of random forests, neural networks, support vector machines, and linear classification metrics. Each of the multiple different classes may be selected from a group consisting of healthy individuals, breast cancer, colon cancer, lung cancer, pancreatic cancer, prostate cancer, ovarian cancer, melanoma, and liver cancer.

[0043] A trained classification index may be applied to a method for classifying samples derived from a subject. This classification method may include the steps of (a) providing a representative multi-parameter model of cell-free DNA molecules from a test sample derived from a subject; and (b) classifying the test sample using the trained classification index. After the test samples have been classified into one or more classes, therapeutic interventions in the subject may be carried out based on the classification of the samples.

[0044] In some embodiments, the training set is provided to a machine learning unit, such as a neural network or a support vector machine. Using the training set, the machine learning unit can create a model for classifying samples according to their treatment response to one or more therapeutic interventions. This is also referred to as "calling." The developed model can incorporate information from any portion of the test vector.

[0045] In some embodiments, DNA from a population of several individuals may be analyzed by a set of multiplex arrays. The data for each multiplex array may be self-normalized using the information contained in that specific array. This normalization algorithm may be adjusted for nominal intensity variations observed in two-color channels, background differences between channels, and the possibility of crosstalk between pigments. The behavior of each base position may then be modeled using a clustering algorithm that incorporates several biological heuristics in SNP genotyping. If fewer than three clusters are observed (e.g., due to low minor allele occurrence frequencies), the location and shape of the missing clusters can be estimated using a neural network. Depending on the cluster shapes and their relative distances from one another, a statistical score (training score) can be devised. Scores such as the GenCall Score are designed to mimic assessments made by human experts using visual and empirical systems. Furthermore, this is evolved using genotyping data from both upper and lower strands. This score may be combined with several penalty terms (e.g., low intensity, mismatch between existing and predicted clusters) to create a training score. The training score, along with the cluster location and shape for each SNP, is stored for use by the calling algorithm.

[0046] To recall treatment responses, the recall algorithm may utilize genetic information and treatment responses from multiple individuals with the disease or condition. The data may be normalized first (using the same procedure as for the clustering algorithm). The recall operation (classification) may be performed using, for example, a Bayesian model. For each recall, a score is given; the recall score may be the product of the training score and the data versus model fit score. After scoring all treatment responses, the application can calculate a composite score.

[0047] In some embodiments, the training dataset includes clinical data selected from a group consisting of cancer stage, type of surgical procedure, age, tumor grade, depth of tumor invasion, occurrence of postoperative complications, and presence of vascular invasion. In some embodiments, the training dataset is preprocessed by converting the provided data into class-conditional probabilities.

[0048] Another embodiment uses machine learning techniques to train a statistical classification index, specifically a support vector machine, for each cancer stage category based on word occurrences in a corpus of histological reports for each patient. The new reports may then be classified according to the most likely stage, facilitating the collection and analysis of population staging data. • Converting data to a Support Vector Machine (SVM) package format • Performing scaling on data • Consider the RBF kernel • Use cross-validation to find the best parameters C and γ. • Use the best parameters C and γ to train the entire training set. · inspection • Starting with patient data This embodiment is an open-source implementation of the Support Vector Machine (SVM) in C. light It uses the following: The main characteristics of the program are as follows: High-speed optimization algorithm Selecting work settings based on the steepest possible descent. "Reduction" heuristic Kernel evaluation caching Using folding in the linear case Solving classification and regression problems. Using SVMstruct for multivariate and structured outputs. Solving ranking problems (e.g., learning the search functionality of the STRIVER search engine). Calculate XiAlpha estimates of error rate, precision, and recall. Efficiently calculate Leave-One-Out estimates of error rate, precision, and recall. This includes an algorithm for approximately training a Large Transform SVM (TSVM) (see also Spectral Graph Transformer). Allows restarting from a specific vector of dual variables so that SVMs can be trained with cost models and example-dependent costs. Handling thousands of support vectors It handles hundreds of thousands of training examples. Supports standard kernel functions and allows users to define their own. Use sparse vector presentation

[0049] In some embodiments, the machine learning algorithm is selected from a group of supervised or unsupervised learning algorithms, which include support vector machines, random forests, nearest neighbor analysis, linear regression, binary decision trees, discriminant analysis, logistic classification metrics, and cluster analysis.

[0050] Referring next to Figure 1C, a system including a report generator 21 for reporting cancer test results and the resulting treatment options is schematically illustrated. The report generator system may be a central data processing system configured to establish direct communication with a remote data site or laboratory 22, a medical practitioner / healthcare provider (treatment specialist) 24 and / or a patient / subject 26 via a communication line. The laboratory 22 may be a medical laboratory, diagnostic laboratory, medical facility, medical practice, point-of-care testing device, or any other remote data site capable of generating subject clinical information. Subject clinical information may include, but is not limited to, clinical test data, X-ray data, examinations, and diagnoses. The healthcare provider or clinic 26 includes medical service providers such as physicians, nurses, home healthcare providers, technicians, and physician's assistants, and the clinic is any medical care facility to which the healthcare provider is assigned. In certain cases, the healthcare provider / clinic is also a remote data site. In the cancer treatment embodiment, the subject may specifically have cancer.

[0051] Other clinical information regarding cancer subject 26 includes the results of clinical examinations, imaging or medical procedures directed to specific cancers that can be easily identified by those skilled in the art. Appropriate sources of clinical information about cancer include, but are not limited to, CT scans, MRI scans, ultrasound scans, bone scans, PET scans, bone marrow examinations, barium X-rays, endoscopy, lymphangiography, IVU (intravenous urography) or IVP (intravenous pyelography), lumbar puncture, cystoscopy, immunological tests (anti-malignin antibody screening), and cancer marker tests.

[0052] Clinical information for subject 26 can be obtained manually or automatically from the laboratory 22. For the simplicity of the system, the information is obtained automatically at predetermined or fixed time intervals. Fixed time intervals refer to time intervals at which laboratory data collection is performed automatically by the method and system described herein based on time measurements such as hours, days, weeks, months, or years. In one embodiment of the present invention, data collection and processing are performed at least once a day. In one embodiment, data transmission and collection are performed once a month, every two weeks, once a week, or once every two or three days. Alternatively, information retrieval may be performed at predetermined but non-fixed time intervals. For example, the first retrieval step may be performed after one week, and the second retrieval step may be performed after one month. Data transmission and collection may be customized according to the nature of the disorder being managed and the required frequency of examinations and consultations for the subject.

[0053] Figure 1D shows an exemplary process for creating a genetic report, including a tumor response map and an overview of associated changes. This process reduces error rates and biases that may be orders of magnitude higher than necessary to reliably detect new cancer-related genomic changes. The process first captures genetic information by collecting bodily fluid samples (e.g., blood, serum, plasma, urine, cerebrospinal fluid, saliva, feces, lymph, synovial fluid, cystic fluid, ascites, pleural fluid, amniotic fluid, chorionic villi samples, preimplantation embryo-derived fluid, placental samples, lavage and cervical vaginal fluid, interstitial fluid, cheek swab samples, sputum, bronchial lavage, Pap smear samples, or ocular fluid) as sources of genetic material, and then the process sequences the material (71). For example, polynucleotides in a sample may be sequenced to produce multiple sequence reads. The tumor burden in a sample containing polynucleotides can be estimated as the relative number of sequence reads containing variants to the total number of sequence reads generated from the sample. Similarly, in the case of copy number variants, tumor burden can be estimated as a relative excess (in the case of gene duplication) or relative deficiency (in the case of gene reduction) of the total number of sequence reads at the test and control loci. For example, Ran can generate 1000 reads mapped to an oncogene locus, of which 900 correspond to the wild type and 100 correspond to cancer variants exhibiting copy number variants in this gene. In some embodiments, the genetic information includes variables that define the genomic composition of cancer cells or the genomic composition of a single disseminated cancer cell. In some embodiments, the genetic information includes sequence or abundance data from one or more loci in cell-free DNA derived from an individual. Further details on the sequencing of exemplary specimen collections and genetic material are discussed in Figures 3A–3B below.

[0054] Next, the genetic information is processed (72). Then, genetic variants are identified. Genetic variants include sequence variants, copy number variants, and nucleotide modification variants. Sequence variants are variations in the gene nucleotide sequence. Copy number variants are deviations from the wild type in the copy number of a portion of the genome. Genetic variants include, for example, single nucleotide polymorphisms (SNPs), insertions, deletions, inversions, transpositions, migrations, gene fusions, chromosome fusions, gene shortenings, copy number polymorphisms (e.g., aneuploidy, partial aneuploidy, polyploidy, gene amplification), abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid methylation. The process then determines the frequency of occurrence of genetic variants in a sample containing genetic material. Since this process is noisy, the process separates the information from the noise (73). The sensitivity of genetic variant detection can be increased by increasing the polynucleotide read depth (e.g., by sequencing at a greater read depth in a sample derived from the target at two or more time points).

[0055] Sequencing methods have an error rate. For example, Illumina's mySeq system can produce error rates in the low single digits. Therefore, for the mapping of 1000 sequence reads to a gene locus, it can be predicted that approximately 50 reads (about 5%) will contain errors. WO2014 / 149134 (Talasaz and Eltoukhy) Certain methodologies, such as those described, can significantly reduce the error rate. This error creates noise that can obscure low levels of cancer-derived signals present in the sample. Therefore, if a sample has a tumor burden at a level near the sequencing system error rate, e.g., approximately 0.1%–5%, it may be difficult to distinguish signals corresponding to cancer-induced genetic mutations from those caused by noise.

[0056] Cancer can be diagnosed by analyzing genetic variants, even in the presence of noise. The analysis can be based on the frequency of occurrence of sequence variants or the level of CNVs (74), and diagnostic accuracy indices or levels for detecting genetic variants within a noise range can be established (75).

[0057] The next step is to increase diagnostic accuracy. This can be done by using multiple measurements to increase diagnostic accuracy (6), or alternatively by using measurements at multiple time points (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more time points) to determine whether the cancer is progressing, in remission, or stable (77). Diagnostic accuracy can be used to identify the disease state. For example, cell-free polynucleotides taken from a subject may contain polynucleotides from normal cells and polynucleotides from disease cells such as cancer cells. Polynucleotides from cancer cells may harbor genetic variants such as somatic mutations and copy number variants. When cell-free polynucleotides from a sample derived from a subject are sequenced, these cancer polynucleotides are detected as sequence variants or copy number variants. The relative amount of tumor polynucleotides in a cell-free polynucleotide sample is referred to as the "tumor load."

[0058] Parameter measurements may be provided with confidence intervals, regardless of whether they are within the noise range. When examined over time, the confidence intervals can be compared over time to determine whether the cancer is progressing, stable, or in remission. If the confidence intervals do not overlap, this indicates the direction of the disease.

[0059] The next step is to create a genetic report / diagnosis. First, the step searches for past treatments from populations with similar genetic profiles (78). The step includes creating a genetic graph for several measures showing mutation tendencies (79) and creating a report showing treatment outcomes and options (80).

[0060] One application is cancer detection. Numerous cancers can be detected using the methods and systems described herein. Cancer cells, like many cells, can be characterized by their metabolic turnover rate, in which old cells die and are replaced by newer cells. Generally, dead cells may release DNA or DNA fragments into the bloodstream upon contact with the vascular system of a given object. This is also true of cancer cells at various stages of the disease. Depending on the stage of the disease, cancer cells can also be characterized by various genetic abnormalities, such as copy number variations and mutations. This phenomenon can be used to detect the presence or absence of individual cancers using the methods and systems described herein.

[0061] For example, blood from a subject at risk of cancer may be collected and prepared as described herein to generate a population of cell-free polynucleotides. In one example, this may be cell-free DNA. The systems and methods of this disclosure may be used to detect mutations or copy number variations that may be present in a particular cancer. The methods can assist in detecting the presence of cancerous cells in the body despite the absence of disease symptoms or other hallmarks.

[0062] As used herein, the term “cancer” includes, but is not limited to, a variety of malignant neoplasms, most of which can invade surrounding tissues and metastasize to different sites (see, for example, the PDR Medical Dictionary, 1st edition (1995), which is incorporated herein by reference in its entirety for all purposes). The terms “neoplasm” and “tumor” refer to abnormal tissue that grows by more rapid cellular proliferation than normal and continues to grow after the stimulus that initiated the proliferation has been removed. Such abnormal tissue exhibits a partial or complete lack of structural organization and functional coordination with normal tissue and may be either benign (e.g., a benign tumor) or malignant (e.g., a malignant tumor). Examples of general categories of cancer include, but are not limited to, carcinoma (e.g., malignant tumors of epithelial cell origin, such as common forms of breast, prostate, lung, and colon cancer), sarcoma (malignant tumors of connective tissue or mesenchymal cell origin), lymphoma (malignant lesions of hematopoietic cell origin), leukemia (malignant lesions of hematopoietic cell origin), and germ cell tumors (tumors of totipotent cells, most often found in the testes or ovaries in adults; most often found in the midline of the body, particularly at the tip of the coccyx, in fetuses, infants, and young children), and blastoma (malignant tumors typically resembling immature or embryonic tissue). Examples of types of neoplasms intended to be encompassed by the present invention include, but are not limited to, neoplasms associated with cancers of nerve tissue, hematopoietic tissue, breast, skin, bone, prostate, ovary, uterus, cervix, liver, lung, brain, larynx, gallbladder, pancreas, rectum, parathyroid gland, thyroid gland, adrenal gland, immune system, head and neck, colon, stomach, bronchi, and / or kidneys. In specific embodiments, the types and number of cancers that can be detected include, but are not limited to, blood cancer, brain cancer, lung cancer, skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, and homogeneous tumors.

[0063] In the early detection of cancer, any of the systems and methods described herein can be used to detect cancer, including mutation detection or copy number variation detection. These systems and methods can be used to detect a number of genetic abnormalities that cause or may arise from cancer. These include, but are not limited to, mutations, indels, copy number variations, translocations, migrations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural changes, gene fusions, gene shortenings, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, abnormal changes in nucleic acid methylation, infection, and cancer.

[0064] Additionally, the systems and methods described herein can also be used to assist in the characterization of certain cancers. Genetic data derived from the systems and methods of this disclosure may enable practitioners to better characterize specific forms of cancer. Cancers are often heterogeneous in both composition and stage. Genetic profiling data may enable the characterization of specific subtypes of cancer, which may be important in the diagnosis or treatment of specific subtypes. This information may also provide subjects or practitioners with clues about the prognosis of specific types of cancer.

[0065] The systems and methods provided herein can be used to monitor known cancers or other diseases in specific subjects. This may allow either the subject or the practitioner to adapt treatment options as the disease progresses. In this example, the systems and methods described herein can be used to construct a genetic profile of the disease course in a specific subject. In some cases, cancer may progress and become more aggressive and genetically unstable. In other cases, cancer may remain benign, inactive, or quiescent. The systems and methods of this disclosure may be useful in determining disease progression.

[0066] Furthermore, the systems and methods described herein may be useful in determining the effectiveness of specific treatment options. In one example, a successful treatment may result in the death of more cancer cells and the shedding of DNA, thus demonstrably increasing the amount of copy number variations or mutations detected in the target blood. In other examples, this may not occur. In yet another example, a particular treatment option may correlate over time with the genetic profile of the cancer. This correlation may be useful in selecting a treatment. Additionally, if the cancer is observed to be in remission after treatment, the systems and methods described herein may be useful in monitoring residual disease or disease recurrence.

[0067] The methods and systems described herein are not limited to the detection of mutations and copy number variations solely related to cancer. Various other diseases and infections may produce other types of conditions for which early detection and monitoring may be preferable. For example, in certain cases, hereditary disorders or infectious diseases may produce certain genetic mosaicisms that are within the scope of interest. These genetic mosaics may produce observable copy number variations and mutations. In another example, the systems and methods of this disclosure can also be used to monitor the genomes of immune cells in the body. Immune cells, such as B cells, may undergo rapid clonal proliferation in the presence of certain diseases. Clonal proliferation can be monitored using copy number variation detection, and certain immune conditions can be monitored. In this example, copy number variation analysis may be performed over time to profile how a particular disease may progress.

[0068] Furthermore, the systems and methods of this disclosure can also be used to monitor systemic infections themselves when they are caused by pathogens such as bacteria or viruses. Copy number variation or further mutation detection can be used to determine how a population of pathogens changes over the course of an infection. This may be particularly important in chronic infections such as HIV / AIDs or hepatitis infections, as viruses may change their lifecycle status and / or mutate into pathogenic forms over the course of an infection.

[0069] Another example in which the systems and methods of this disclosure may be used is the monitoring of transplanted tissues. Transplanted tissues generally undergo some degree of rejection by the body during transplantation. The methods of this disclosure may be used to determine or profile the rejection activity of the host body when immune cells attempt to destroy the transplanted tissue. This may be useful in monitoring the status of the transplanted tissue and in altering the course of treatment or preventing rejection.

[0070] Furthermore, the methods of this disclosure can be used to characterize heterogeneity of an abnormal condition in a subject, the methods comprising the step of creating a gene profile of extracellular polynucleotides in the subject, wherein the gene profile includes multiple data resulting from copy number variation and mutation analysis. In some cases, including cancer, the disease may be heterogeneous. Disease cells may not be identical. In the case of cancer, it is known that some tumors contain different types of tumor cells, some cells in different stages of cancer. In other cases, heterogeneity may include multiple lesions of the disease. Furthermore, in the case of cancer, there may be multiple tumor lesions, one or more of which may be the result of metastasis (also known as distant metastasis) that has spread from the primary site.

[0071] The methods described herein can be used to create a profile, fingerprint, or set of data which is the sum of genetic information from different cell origins in heterogeneous diseases. This set of data may include copy number variation and mutation analysis, either alone or in combination.

[0072] Additionally, the systems and methods of this disclosure may be used for the diagnosis, prediction, monitoring, or observation of cancer or other diseases of fetal origin. That is, these methodologies may be used in pregnant subjects for the diagnosis, prediction, monitoring, or observation of cancer or other diseases in prenatal subjects whose DNA and other polynucleotides may be circulating with maternal molecules.

[0073] Furthermore, these reports are submitted and accessed electronically via the internet. Analysis of sequence data is performed at a site other than the subject's location. The reports are created and transmitted to the subject's location. Subjects access reports showing their tumor burden via an internet-enabled computer.

[0074] The annotated information may be used by healthcare providers to select other drug treatment options and / or to provide insurance companies with information about drug treatment options. For example, the NCCN Clinical Practice Guidelines in Oncology ( This may include a step of annotating drug treatment options for conditions listed in the trademark or the American Society of Clinical Oncology (ASCO) clinical practice guidelines.

[0075] Drug treatment options stratified in the report may be annotated in the report by listing additional drug treatment options. Additional drug treatments may be FDA-approved drugs for off-label use. A provision of the Comprehensive Financial Adjustment Act of 1993 (OBRA) requires Medicare to bear the cost of off-label use of anticancer drugs included in the standard medical compendia. Drugs used for the annotation list include the National Comprehensive Cancer Network (NCCN) Drugs and Biologics Compendium®, Thomson Micromedex DrugDex®, Elsevier Gold Standard's Clinical Pharmacology Compendium, and American Hospital Formulary Found in the CMS approval list, including the Service-Drug Information Compendium (registered trademark). It is possible.

[0076] Drug treatment options can be annotated by listing experimental drugs that may be useful in treating cancers having one or more molecular markers of a specific situation. Experimental drugs may be drugs for which in vitro data, in vivo data, animal model data, preclinical trial data, or clinical trial data are available. The data can be found in academic journals listed in the CMS Medicare Benefit Policy Manual, including, for example, the American Journal of Medicine, Annals of Internal Medicine, Annals of Oncology, Annals of Surgical Oncology, Biology of Blood and Marrow Transplantation, Blood, Bone Marrow Transplantation, British Journal of Cancer, British Journal of Hematology, British Medical Journal, Cancer, Clinical Cancer Research, Drugs, European Journal of Cancer (formerly the European Journal of Cancer and Clinical Oncology), Gynecologic Oncology, International Journal of Radiation, Oncology, Biology, and Physics, The Journal of the American Medical Association, Journal of Clinical Oncology, Journal of the National Cancer Institute, Journal of the National Comprehensive Cancer Network (NCCN), Journal of Urology, Lancet, Lancet Oncology, Leukemia, The New England Journal of Medicine, and Radiation Oncology. It is acceptable if it has been published in peer-reviewed medical literature.

[0077] Drug treatment options may be annotated by providing links to electronic reports that connect the listed drugs to scientific information about those drugs. For example, links may be provided for information on clinical trials of a drug (clinicaltrials.gov). If the report is provided via computer or computer website, links may be footnotes, hyperlinks to websites, pop-up boxes, or flyover boxes containing information. The report and annotated information may be provided in printed form, and annotations may be, for example, footnotes to references.

[0078] Information for annotating one or more drug treatment options in a report may be provided, for example, by a commercial entity that stores scientific information. Healthcare providers can treat subjects, such as cancer subjects, with the experimental drugs listed in the annotated information, and can access the annotated drug treatment options, search for scientific information (e.g., print medical journal articles), and submit it to insurance companies along with a request for reimbursement for the provision of drug treatment (e.g., printed journal articles). Physicians can use any of the various diagnostic-related group (DRG) codes to enable reimbursement.

[0079] Drug treatment options in the report may be annotated with information about other molecular components in the pathways affected by the drugs (e.g., information about drugs that target kinases downstream of the cell surface receptor that is the drug target). Drug treatment options may be annotated with information about drugs that target one or more other molecular pathway components. Identification and / or annotation of pathway information may be outsourced or subcontracted to another company. Annotated information may include, for example, drug names (e.g., FDA-approved drugs for off-label use; drugs found in the CMS approval list, and / or drugs described in scientific (medical) journal articles), scientific information about one or more drug treatment options, one or more links to scientific information about one or more drugs, clinical trial information about one or more drugs (e.g., information from clinicaltrials.gov / ), and one or more links for citations of scientific information about drugs. Annotated information may be inserted anywhere in the report. Annotated information may be inserted in multiple locations in the report. Annotated information may be inserted into the report near the section on hierarchical drug treatment options. Annotated information may be inserted into the report on pages separate from the hierarchical drug treatment options. Reports that do not contain hierarchical drug treatment options may be annotated with information.

[0080] The methods provided may also include means for investigating the effects of a drug on a sample (e.g., tumor cells) isolated from a subject (e.g., a cancer subject). In vitro cultures using tumors from cancer subjects can be established using techniques known to those skilled in the art. The methods provided may also include high-throughput screening of FDA-approved off-label or experimental drugs using the in vitro cultures and / or xenograft models. The methods provided may also include a step of monitoring tumor antigens for the detection of recurrence.

[0081] A report is prepared, mapping genomic location and copy number variations for subjects with cancer. These reports can indicate that a particular cancer is aggressive and resistant to treatment, compared to other profiles of subjects with known outcomes. Subjects are monitored and retested for a period of time. If the copy number variation profile begins to increase dramatically at the end of that period, this may indicate that the current treatment is not working. Comparisons are made with the genetic profiles of other prostate subjects. For example, if this increase in copy number variations is determined to indicate that the cancer is progressing, the original treatment regimen prescribed is no longer treating the cancer, and a new treatment is prescribed.

[0082] Figures 2A-2B illustrate in more detail one embodiment for generating genetic reports and diagnoses. In one implementation, Figure 2B shows exemplary pseudocode executed by the system in Figure 1A to process non-CNV reporting mutant allele occurrence frequencies. However, the system can also process CNV reporting mutant allele occurrence frequencies.

[0083] Next, considering Figure 2A, the process receives genetic information from a DNA sequencing machine (30). The process then determines specific gene changes and their amounts (32). A tumor response map is then created. To create the map, the process normalizes the amounts for each gene change to render across all test points, and then creates a magnification (34). As used herein, the term “normalize” generally refers to a means of adjusting values ​​measured on different scales to a nominally common scale. For example, data measured at different points are transformed / adjusted so that all values ​​can be resized to a common scale. As used herein, the term “magnification” generally refers to a number by which a quantity is measured or magnified. For example, in the equation y=Cx, C is the magnification with respect to x. C is also the coefficient of x and may be referred to as the constant of the proportionality between y and x. Values ​​are normalized so that they can be plotted on a visually understandable common scale. The magnification is used to know the exact height corresponding to the value being plotted (for example, the 10% mutant allele frequency may mean, for example, 1 cm on the report). The magnification is applied to all test points and is therefore considered the overall magnification. For each test point, the process renders the information onto a tumor response map (36). In operation 36, the process renders the changes and relative heights using the determined magnification (42) and assigns a unique visual indicator to each change. In addition to the response map, the process creates an overview of the changes and treatment options. Similarly, information from clinical trials that can assist in identifying specific genetic changes and other helpful treatment suggestions are presented along with explanations of terminology and testing methodology, and other information is added to the report and provided to the user.

[0084] In one implementation, copy number variations may be reported as a graph, showing the various locations in the genome and the corresponding increase, decrease, or maintenance of copy number variations at each location. Additionally, copy number variations may be used to report a percentage score indicating the amount of disease material (or nucleic acids with copy number variations) present in a cell-free polynucleotide sample.

[0085] These reports are submitted and accessed electronically via the internet. Analysis of sequence data is performed at a site other than the subject's location. The reports are created and transmitted to the subject's location. Subjects access reports showing their tumor burden via an internet-enabled computer.

[0086] Next, details of an exemplary gene testing process are disclosed. Next, considering Figure 3A, the exemplary process receives gene material from a blood sample or other body sample (102). The process converts polynucleotides from the gene material into tagged parental nucleotides (104). The tagged parental nucleotides are amplified to produce amplified progeny polynucleotides (106). A subset of amplified polynucleotides is sequenced to create sequence reads (108) and grouped into families, each generated from a unique tagged parent nucleotide (110). At selected loci, the process assigns each family a confidence score (112). A consensus is then determined using past readings, which is done by reviewing past confidence scores for each family, and where matching past confidence scores exist, the current confidence score is increased accordingly (114). If past confidence scores exist but they do not match, in one embodiment the current confidence score is not changed (116). In other embodiments the confidence scores are adjusted in a predetermined manner for the mismatched past confidence scores. In the case of the first detection of a family, the current confidence scores may be reduced because they may be erroneous readings (118). The process can infer the frequency of occurrence of families at loci in the set of tagged parent polynucleotides based on the confidence scores. A genetic testing report is then prepared as discussed above (120).

[0087] In some embodiments, only a subset of detected changes is plotted. In some embodiments, the subset is selected based on its likelihood of being a driver change or its association with increased or decreased response to treatment. In some embodiments, a combination of magnitudes of genomic changes detected by fluid-based testing is used to estimate disease burden. In some embodiments, the allele proportion, allele imbalance, or gene-specific scope of detected mutations is used to estimate disease burden. In some embodiments, the overall stack height represents the overall disease burden or disease burden score in the subject. In some embodiments, different colors are used to represent each gene variant. In some embodiments, only a subset of detected gene variants is plotted. In some embodiments, the subset is selected based on its likelihood of being a driver change or its association with increased or decreased response to treatment.

[0088] While temporal information was used in Figures 3A-3B to enhance information for mutation or copy number variation detection, other common methods can also be applied. In other embodiments, historical comparisons may be used in conjunction with other consensus sequence mappings to specific reference sequences to detect examples of genetic polymorphism. Consensus sequence mappings to specific reference sequences may be measured and normalized against control samples. Measurements of molecular mappings to reference sequences may be compared across the genome to identify regions in the genome where copy number variation or heterozygosity is lost. Common methods include, for example, linear or nonlinear methods of consensus sequence construction derived from digital communications theory, information theory, or bioinformatics (such as voting, averaging, statistical, maximum apost-hoc or maximum likelihood detection, dynamic programming, Bayesian, Hidden Markov, or Support Vector Machine methods). After the sequence read scope is determined, a probabilistic modeling algorithm is applied to translate the normalized nucleic acid sequence read scope for each window region into separate copy number states. In some cases, this algorithm may include one or more of the following: hidden Markov models, dynamic programming, support vector machines, Bayesian networks, trellis decoding, viter-vide decoding, expectation maximization, Kálmán filtering methodology, and neural networks.

[0089] Artificial neural networks (NNets) mimic networks of "neurons" based on the neural structure of the brain. They process recordings one at a time or in batches and "learn" by comparing their classifications of the recordings (which are initially largely random) to known actual classifications of the recordings. In MLP-NNets, the error from the initial classification of the first recording is fed back into the network and used to modify the network's algorithm for the second time, and similarly for numerous iterations.

[0090] The neural network uses an iterative learning process in which data cases (columns) are presented to the network one at a time, and the weights associated with the input values ​​are adjusted each time.

[0091] After all cases have been presented, the process is often repeated from the beginning. During this learning phase, the network learns by adjusting its weights so that it can predict the correct classification label for the input sample. Neural network learning is also called "connectionist learning" due to the connections between units. Advantages of neural networks include their high tolerance for noisy data and their ability to classify patterns they have not been trained on. One neural network algorithm is the backpropagation algorithm, such as the Levenberg-Macart algorithm. Once a network is built for a specific application, it is ready to be trained. Initial weights are randomly selected to start this process. Training or learning then begins.

[0092] The network uses weights and functions in a hidden layer to process one record at a time from the training data, then compares the resulting output to the desired output. The error is then propagated back through the system, causing the system to adjust the weights for application to the next record processed. This process is repeated as the weights are continuously fine-tuned. The same set of data is processed many times during network training so that the connected weights continuously improve.

[0093] In one embodiment, the training step of machine learning units on a training dataset can create one or more classification models to be applied to test samples. These classification models may be applied to test samples to predict the subject's response to a therapeutic intervention.

[0094] As shown in Figure 3B, comparing the sequence coverage with a control sample or reference sequence can aid in normalization across a window. In this embodiment, cell-free DNA is extracted and isolated from readily available bodily fluids such as blood. For example, cell-free DNA may be extracted using various methods known in the art, including, but not limited to, isopropanol precipitation and / or silica-based purification. Cell-free DNA may be extracted from any number of subjects, including subjects without cancer, subjects at risk of cancer, or subjects known to have cancer (e.g., through other means).

[0095] Following the isolation / extraction step, any number of different sequencing operations may be performed on the cell-free polynucleotide sample. The sample may be treated with one or more reagents (e.g., enzymes, unique identifiers (e.g., barcodes), probes, etc.) before sequencing. In some cases where the sample is treated with unique identifiers such as barcodes, the sample or sample fragments may be tagged individually or in subgroups with the unique identifiers. The tagged sample can then be used in downstream applications such as sequencing reactions in which individual molecules can be traced back to their parent molecules.

[0096] Cell-free polynucleotides may be tagged or tracked to enable the identification and origin of subsequent specific polynucleotides. Assigning identifiers (e.g., barcodes) to individual or subgroup polynucleotides may allow for the assignment of unique identity to individual sequences or fragments of sequences. This may enable the acquisition of data derived from individual samples, not limited to the average of the samples. In some cases, nucleic acids or other molecules derived from a single strand may share a common tag or identifier, thereby allowing for later identification of their origin from that strand. Similarly, all fragments derived from a single strand of nucleic acid may be tagged with the same identifier or tag, thereby enabling the identification of subsequent fragments from the parent strand. In other cases, gene expression products (e.g., mRNA) may be tagged to quantify expression, as the barcode, or the barcode in combination with the sequence to which it is attached, may be counted. In yet another case, systems and methods can be used as PCR amplification controls. In such cases, multiple amplified products derived from a PCR reaction may be tagged with the same tag or identifier. If the products are later sequenced and demonstrate sequence differences, differences within products with the same identifier may be due to PCR errors. Additionally, individual sequences can be identified based on the characteristics of the sequence data of the read itself. For example, the detection of unique sequence data at the beginning (start) and end (stop) portions of individual sequencing reads may be used alone or in combination with the length or number of base pairs of the sequence-specific sequences for each sequence read to assign a unique identity to individual molecules. Fragments derived from a single strand of nucleic acid may be assigned a unique identity, thereby enabling the identification of subsequent fragments derived from the parent strand. This can be used in conjunction with constraints on the initial starting gene material to limit variability.

[0097] Furthermore, the use of unique sequence data at the beginning (start) and end (stop) portions of individual sequencing reads, as well as the sequencing read length, may be used alone or in combination with the use of barcodes. In some cases, the barcodes may be unique as described herein. In other cases, the barcodes themselves may not be unique. In this case, the use of non-unique barcodes in combination with sequence data at the beginning (start) and end (stop) portions of individual sequencing reads, as well as the sequencing read length, may enable the assignment of unique identity to individual sequences. Similarly, a single-stranded nucleic acid fragment to which a unique identity has been assigned may thereby enable the subsequent identification of a fragment derived from the parent strand.

[0098] Generally, the methods and systems provided herein are useful for the preparation of cell-free polynucleotide sequences for downstream sequencing reactions. Often, the sequencing method is classical Sanger sequencing.

[0099] As used herein, the term “sequencing” refers to any number of techniques used to determine the sequence of a biomolecule, such as nucleic acids, including DNA or RNA. Examples of sequencing methods, though not limited to these, include targeted sequencing, single-molecule real-time sequencing, exon sequencing, electron microscope-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxytermination sequencing, whole-genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, gel electrophoresis, double-strand sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, ultra-parallel signature sequencing, emulsion PCR, simultaneous amplification at low-denaturation temperature PCR (COLD-PCR), multiplex PCR, reversible diterminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single-molecule sequencing, synthesis sequencing, real-time sequencing, reverse-terminator sequencing, nanopore sequencing, 454 sequencing, and Solexa Genome sequencing. Examples include Analyzer sequencing, SOLiD® sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing may be performed by a gene analyzer, such as a gene analyzer commercially available from Illumina or Applied Biosystems. In some embodiments, the sequencing method may be massively parallel sequencing, i.e., simultaneous (or high-speed sequential) sequencing of at least 100, 1,000, 10,000, 100,000, 1 million, 10 million, 100 million, or 1 billion polynucleotide molecules.

[0100] A sequencing method typically includes sample preparation, sequencing of polynucleotides in the prepared sample to create sequence reads, and bioinformatics manipulation of the sequence reads to generate quantitative and / or qualitative genetic information about the sample. Sample preparation typically includes converting the polynucleotides in the sample into a form compatible with the sequencing platform used. This conversion may include tagging the polynucleotides. In certain embodiments of the present invention, the tags include polynucleotide sequence tags. The conversion methodologies used in sequencing may not be 100% efficient. For example, it is not uncommon for polynucleotides in a sample to be converted with a conversion efficiency of about 1-5%, i.e., about 1-5% of the polynucleotides in the sample to be converted into tagged polynucleotides. Polynucleotides that are not converted into tagged molecules are not presented in the tagged library for sequencing. Therefore, polynucleotides with gene variants that are presented at low frequency in the start gene material may not be presented in the tagged library and may therefore not be sequenced or detected. Increasing conversion efficiency increases the probability that polynucleotides in the start gene material will be presented in the tagged library and consequently detected by sequencing. Furthermore, most protocols to date require more than 1 microgram of DNA as input material, rather than directly addressing the low conversion efficiency challenge of library preparation. However, when input sample material is limited or detection of polynucleotides at low presentation is desired, high conversion efficiency allows for efficient sequencing of the sample and / or proper detection of such polynucleotides.

[0101] Mutation detection may generally be performed on selectively enriched regions of a purified and isolated genome or transcriptome (302). Certain regions described herein, but not limited to, may include genes, oncogenes, tumor suppressor genes, promoters, regulatory sequence elements, non-coding regions, miRNAs, snRNAs, etc., can be selectively amplified from the entire population of cell-free polynucleotides. This may be carried out as described herein. In one example, multiple sequencing may be performed with or without barcode labeling for individual polynucleotide sequences. In other examples, sequencing may be performed using any nucleic acid sequencing platform known in the art. This step generates multiple genome fragment sequence reads (304). Additionally, reference sequences are obtained from control samples taken from another subject. In some cases, the control subject may be a subject known not to have any known genetic abnormalities or diseases. In some cases, these sequence reads may contain barcode information. In other examples, barcodes are not used.

[0102] After sequencing, each read is assigned a quality score. The quality score may be a presentation of the read that indicates, based on a threshold, whether these reads are useful in subsequent analyses. In some cases, some reads are not of sufficient quality or length to perform the subsequent mapping step. Sequenced reads with quality scores of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% may be excluded from the dataset. In other cases, sequenced reads assigned a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% may be excluded from the dataset. In step 306, genome fragment reads that meet a specified quality score threshold are mapped to a reference genome or a reference sequence that is known not to contain mutations. After mapping the sequence comparison, the sequence reads are assigned a mapping score. The mapping score may be a presentation or read that maps back to a reference sequence indicating whether each position is uniquely mappable. In some cases, reads may be sequences irrelevant to the mutation analysis. For example, some sequence reads may be derived from contaminating polynucleotides. Sequence readings with mapping scores of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% may be excluded from the dataset. In other cases, sequence readings assigned mapping scores of less than 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% may be excluded from the dataset.

[0103] For each mappable base, bases that do not meet the minimum threshold for mappability or are of low quality may be replaced by the corresponding base found in the reference sequence.

[0104] Once the read coverage has been confirmed and variant bases have been identified in each read compared to the control sequence, the frequency of occurrence of the variant base can be calculated by dividing the number of reads containing the variant by the total number of reads. This can then be expressed as a ratio for each mappable location in the genome.

[0105] The frequencies of all four nucleotides—cytosine, guanine, thymine, and adenine—for each base position are analyzed in comparison to a reference sequence. Probabilistic or statistical modeling algorithms are applied to convert each mappable position into a normalized ratio to reflect the frequency state for each base variant. In some cases, this algorithm may include one or more of the following: hidden Markov models, dynamic programming, support vector machines, Bayesian or probabilistic modeling, trellis decoding, Viter-Vi decoding, expectation maximization, Kálmán filtering methodology, and neural networks.

[0106] In step 312, the distinct variant status of each base position can be used to identify base variants that have high frequency variations compared to the baseline of the reference sequence. In some cases, the baseline may represent frequency occurrences of at least 0.0001%, 0.001%, 0.01%, 0.1%, 1.0%, 2.0%, 3.0%, 4.0%, 5.0%, 10%, or 25%. In other cases, the baseline may represent frequency occurrences of at least 0.0001%, 0.001%, 0.01%, 0.1%, 1.0%, 2.0%, 3.0%, 4.0%, 5.0%, 10%, or 25%. In some cases, all adjacent base positions with base variants or mutations may be merged into a segment to report the presence or absence of the mutation. In some cases, various positions may be filtered before they are merged with other segments.

[0107] After calculating the frequency of variation for each base position, the variant with the largest deviation at specific positions in the target sequence compared to the reference sequence is identified as a mutation. In some cases, the mutation may be a cancer mutation. In other cases, the mutation may correlate with a disease state.

[0108] Mutations or variants may include, but are not limited to, single nucleotide substitutions or genetic abnormalities such as small indels, transpositions, transitions, inversions, deletions, shortenings, or gene shortenings. In some cases, mutations may be up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 nucleotides in length. In other cases, mutations may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 nucleotides in length.

[0109] Next, consensus is determined using past readings. This is done by reviewing past confidence scores for the corresponding bases, and if matching past confidence scores exist, the current confidence score is increased accordingly (314). If past confidence scores exist but they do not match, in one embodiment the current confidence score is not changed (316). In other embodiments the confidence score is adjusted in a predetermined manner for the mismatched past confidence scores. In the case of the first detection of a family, the current confidence score may be reduced because they may be false readings (318). The process then converts the frequency of occurrence of variations per base into separate variant states for each base position (320).

[0110] Numerous cancers can be detected using the methods and systems described herein. Cancer cells, like many cells, can be characterized by their metabolic turnover rate, in which old cells die and are replaced by newer cells. Generally, dead cells may release DNA or DNA fragments into the bloodstream upon contact with the vascular system of a given object. This is also true of cancer cells at various stages of the disease. Depending on the stage of the disease, cancer cells may also be characterized by various genetic abnormalities, such as copy number variations and mutations. This phenomenon can be used to detect the presence or absence of individual cancers using the methods and systems described herein.

[0111] For example, blood from a subject at risk of cancer may be collected and prepared as described herein to generate a population of cell-free polynucleotides. In one example, this may be cell-free DNA. The systems and methods of this disclosure may be used to detect mutations or copy number variations that may be present in a particular cancer. The methods can assist in detecting the presence of cancerous cells in the body despite the absence of disease symptoms or other hallmarks.

[0112] The types and number of cancers that can be detected are not limited to these, but may include blood cancers, brain cancers, lung cancers, skin cancers, nasal cancers, pharyngeal cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, intestinal cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, oral cancers, stomach cancers, solid tumors, heterogeneous tumors, and homogeneous tumors.

[0113] The system and method can be used to detect a number of genetic abnormalities that may cause or result from cancer. These may include, but are not limited to, mutations, indels, copy number variations, translocations, migrations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural changes, gene fusions, gene shortenings, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, abnormal changes in nucleic acid methylation, infection, and cancer.

[0114] Additionally, the systems and methods described herein can also be used to assist in the characterization of certain cancers. Genetic data derived from the systems and methods of this disclosure may enable practitioners to better characterize specific forms of cancer. Cancers are often heterogeneous in both composition and stage. Genetic profiling data may enable the characterization of specific subtypes of cancer, which may be important in the diagnosis or treatment of specific subtypes. This information may also provide subjects or practitioners with clues about the prognosis of specific types of cancer.

[0115] The systems and methods provided herein can be used to monitor known cancers or other diseases in specific subjects. This may allow either the subject or the practitioner to adapt treatment options as the disease progresses. In this example, the systems and methods described herein can be used to construct a genetic profile of the disease course in a specific subject. In some cases, cancer may progress and become more aggressive and genetically unstable. In other cases, cancer may remain benign, inactive, or quiescent. The systems and methods of this disclosure may be useful in determining disease progression.

[0116] Furthermore, the systems and methods described herein may be useful in determining the effectiveness of specific treatment options. In one example, a good treatment option can actually increase the amount of copy number variations or mutations detected in the target blood, as a successful treatment may result in the death of more cancer cells and the shedding of DNA. In other examples, this may not occur. In yet another example, a particular treatment option may correlate over time with the genetic profile of the cancer. This correlation may be useful in selecting a treatment. Additionally, if the cancer is observed to be in remission after treatment, the systems and methods described herein may be useful in monitoring residual disease or disease recurrence.

[0117] The methods and systems described herein are not limited to the detection of mutations and copy number variations solely related to cancer. Various other diseases and infections may produce other types of conditions for which early detection and monitoring may be preferable. For example, in certain cases, hereditary disorders or infectious diseases may produce certain genetic mosaicisms that are within the scope of study. These genetic mosaics may produce observable copy number variations and mutations. In another example, the systems and methods of this disclosure can also be used to monitor the genomes of immune cells in the body. Immune cells, such as B cells, may undergo rapid clonal proliferation in the presence of certain diseases. Clonal proliferation can be monitored using copy number variation detection, and certain immune conditions can be monitored. In this example, copy number variation analysis may be performed over time to profile how a particular disease may progress.

[0118] Furthermore, the systems and methods of this disclosure can also be used to monitor systemic infections themselves when they are caused by pathogens such as bacteria or viruses. Copy number variation or further mutation detection can be used to determine how a population of pathogens changes over the course of an infection. This may be particularly important in chronic infections such as HIV / AIDs or hepatitis infections, as viruses may change their lifecycle status and / or mutate into pathogenic forms over the course of an infection.

[0119] Another example in which the systems and methods of this disclosure may be used is the monitoring of transplanted tissues. Transplanted tissues generally undergo some degree of rejection by the body during transplantation. The methods of this disclosure may be used to determine or profile the rejection activity of the host body when immune cells attempt to destroy the transplanted tissue. This may be useful in monitoring the status of the transplanted tissue and in altering the course of treatment or preventing rejection.

[0120] Furthermore, the methods of this disclosure can be used to characterize heterogeneity of an abnormal condition in a subject, the methods comprising the step of creating a gene profile of extracellular polynucleotides in the subject, wherein the gene profile includes multiple data resulting from copy number variation and mutation analysis. In some cases, including cancer, the disease may be heterogeneous. Disease cells may not be identical. In the case of cancer, it is known that some tumors contain different types of tumor cells, some cells in different stages of cancer. In other cases, heterogeneity may include multiple lesions of the disease. Furthermore, in the case of cancer, there may be multiple tumor lesions, one or more of which may be the result of metastasis spreading from the primary site.

[0121] The methods described herein can be used to create or profile a fingerprint or set of data, which is the sum of genetic information from different cell origins in heterogeneous diseases. This set of data may include copy number variation and mutation analysis, either alone or in combination.

[0122] Additionally, the systems and methods of this disclosure can be used for the diagnosis, prediction, monitoring, or observation of cancer or other diseases of fetal origin. Specifically, these methods may be used in pregnant subjects for the diagnosis, prediction, monitoring, or observation of cancer or other diseases in prenatal subjects in whom their DNA and other polynucleotides may be circulating with maternal molecules.

[0123] Furthermore, these reports are submitted and accessed electronically via the internet. Analysis of sequence data is performed at a site other than the subject's location. The reports are created and transmitted to the subject's location. Subjects access reports showing their tumor burden via an internet-enabled computer.

[0124] The annotated information may be used by healthcare providers to select other drug treatment options and / or to provide insurance companies with information about drug treatment options. For example, the NCCN Clinical Practice Guidelines in Oncology ( This may include a step of annotating drug treatment options for conditions listed in the trademark or the American Society of Clinical Oncology (ASCO) clinical practice guidelines.

[0125] Drug treatment options stratified in the report may be annotated in the report by listing additional drug treatment options. Additional drug treatments may be FDA-approved drugs for off-label use. A provision of the Comprehensive Financial Adjustment Act of 1993 (OBRA) requires Medicare to bear the cost of off-label use of anticancer drugs included in the Standard Medical List. Drugs used for the annotation list are those from the National Comprehensive Cancer Network (NCCN) Drugs and Biologics Compendium®, Thomson Micromedex DrugDex®, and Elsevier Gold Standard's Clinical Pharmacology. It can be found in the CMS approval list, including the compendium and the American Hospital Formulary Service-Drug Information Compendium (registered trademark).

[0126] Drug treatment options can be annotated by listing experimental drugs that may be useful in treating cancers having one or more molecular markers of a specific situation. Experimental drugs may be drugs for which in vitro data, in vivo data, animal model data, preclinical trial data, or clinical trial data are available. Data may be found in, for example, the American Journal of Medicine, Annals of Internal Medicine, Annals of Oncology, Annals of Surgical Oncology, Biology of Blood and Marrow Transplantation, Blood, Bone Marrow Transplantation, British Journal of Cancer, British Journal of Hematology, British Medical Journal, Cancer, Clinical Cancer Research, Drugs, European Journal of Cancer (formerly the European Journal of Cancer and Clinical Oncology), Gynecologic Oncology, International Journal of Radiation, Oncology, Biology, and Physics, The Journal of the American Medical Association, and Journal of Clinical Oncology. This may include publications in peer-reviewed medical literature found in academic journals listed in the CMS Medicare Benefit Policy Manual, including the Journal of the National Cancer Institute, the Journal of the National Comprehensive Cancer Network (NCCN), the Journal of Urology, The Lancet, The Lancet Oncology, Leukemia, The New England Journal of Medicine, and Radiation Oncology.

[0127] Drug treatment options may be annotated by providing links to electronic reports that connect the listed drugs to scientific information about those drugs. For example, links may be provided for information on clinical trials of a drug (clinicaltrials.gov). If the report is provided via computer or computer website, links may be footnotes, hyperlinks to websites, pop-up boxes, or flyover boxes containing information. The report and annotated information may be provided in printed form, and annotations may be, for example, footnotes to references.

[0128] Information for annotating one or more drug treatment options in a report may be provided by a commercial entity that stores scientific information. Healthcare providers can treat subjects such as cancer patients with the experimental drugs listed in the annotated information, and can access the annotated drug treatment options, search for scientific information (e.g., print medical journal articles), and submit it to insurance companies along with a request for reimbursement for the provision of drug treatment (e.g., printed journal articles). Physicians can use any of the various diagnostic-related group (DRG) codes to enable reimbursement.

[0129] Drug treatment options in the report may be annotated with information about other molecular components in the pathways affected by the drugs (e.g., information about drugs that target kinases downstream of the cell surface receptor that is the drug target). Drug treatment options may be annotated with information about drugs that target one or more other molecular pathway components. Identification and / or annotation of pathway information may be outsourced or subcontracted to another company.

[0130] Annotated information may include, for example, drug names (e.g., FDA-approved drugs for off-label use; drugs found in the CMS approval list; and / or drugs described in scientific (medical) journal articles), scientific information on one or more drug treatment options, one or more links to scientific information on one or more drugs, clinical trial information on one or more drugs (e.g., information from clinicaltrials.gov / ), and one or more links to citations of scientific information on drugs.

[0131] Annotated information may be inserted anywhere in the report. Annotated information may be inserted in multiple locations in the report. Annotated information may be inserted in the report near the section on hierarchical drug treatment options. Annotated information may be inserted in the report on pages away from the hierarchical drug treatment options. Reports that do not contain hierarchical drug treatment options may be annotated with information.

[0132] The system may also include reporting on the effects of drugs on samples (e.g., tumor cells) isolated from subjects (e.g., cancer patients). In vitro cultures using tumors from cancer patients can be established using techniques known to those skilled in the art. The system may also include high-throughput screening of FDA-approved off-label or experimental drugs using the in vitro cultures and / or xenograft models. The system may also include monitoring of tumor antigens for the detection of recurrence.

[0133] The system can provide internet-connected access to reports of subjects with cancer. The system can use portable or benchtop DNA sequencers. A DNA sequencer is a scientific instrument used to automate the DNA sequencing process. Given a DNA sample, a DNA sequencer is used to determine the order of four bases: adenine, guanine, cytosine, and thymine. The order of the DNA bases is reported as a text string called a read. Some DNA sequencers may also be considered optical instruments because they analyze light signals produced from fluorescent dyes attached to the nucleotides.

[0134] DNA sequencing machines can utilize Gilbert's sequencing method, which is based on cleavage at specific bases following chemical modification of DNA, or Sanger's technique, which is based on dideoxynucleotide chain termination. Sanger's method became popular due to its increased efficiency and low radioactivity. DNA sequencing machines can use techniques that do not require DNA amplification (polymerase chain reaction - PCR), accelerating sample preparation before sequencing and reducing errors. Additionally, sequencing data is collected in real time from the reaction resulting from nucleotide addition on the complementary strand. For example, DNA sequencing machines can utilize a method called single-molecule real-time sequencing (SMRT), where sequencing data is generated by light emitted (captured by a camera) when nucleotides are added to the complementary strand by an enzyme containing a fluorescent dye. Alternatively, DNA sequencing machines can use electronic systems based on nanopore sensing technology.

[0135] The data is sent by the DNA sequencing instrument to a computer for processing via a direct connection or via the internet. The data processing mode of the system may be implemented by digital electronic circuits, i.e., computer hardware, firmware, software, or a combination thereof. The data processing device of the present invention may be implemented by a computer program product substantially embedded in a machine-readable storage device for execution by a programmable processor; the data processing method steps of the present invention may be implemented by a programmable processor that runs an explanatory program for performing the function of the present invention by manipulating input data and producing an output. The data processing mode of the present invention can be advantageously implemented in one or more computer programs executable in a programmable system including at least one programmable processor, at least one input device, and at least one output device connected to receive data and explanations from and to the data storage system. Each computer program may be implemented in a high-level procedural or object-oriented programming language, or in assembly or machine language, as desired; in any case, the language may be a compiled or interpreted language. Suitable processors include, in the manner exemplified, general-purpose and special-purpose microprocessors. Generally, processors receive descriptions and data from read-only memory and / or random-access memory. Suitable storage devices for substantially embedded computer program descriptions and data include, in the exemplary manner, semiconductor memory devices such as EPROMs, EEPROMs, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM disks, and all forms of non-volatile memory. Any of the foregoing may be complemented by or incorporated into an ASIC (Application-Specific Integrated Circuit).

[0136] To provide user interaction, the present invention may implement the use of a computer system having a display device such as a monitor or LCD (liquid crystal display) screen for displaying information to the user, and an input device such as a keyboard, mouse or trackball (two-dimensional pointing device) or a data glove or gyro mouse (three-dimensional pointing device) for enabling the user to input into the computer system. The computer system may be programmed to provide a graphical user interface through which the computer program interacts with the user. The computer system may be programmed to provide virtual reality, a three-dimensional display interface.

[0137] Preferred embodiments of the present invention are shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided only by illustrative methods. Many variations, changes, and substitutions can be found to those skilled in the art without departing from the present invention. It should be understood that various modifications to the embodiments of the present invention described herein can be used when carrying out the present invention. It is intended that subsequent claims define the scope of the present invention, and that the methods and structures within these claims, as well as their equivalents, are thereby covered.

Claims

1. A method for creating a predictor of treatment response, Steps to create genetic information using a gene analyzer; Steps include receiving into computer memory a training dataset for each of several individuals with cancer, which includes (1) genetic information derived from the individual created at a first time point, and (2) a second training dataset for each individual, which includes the individual's treatment response to one or more therapeutic interventions determined at a later time point; The step of training a computer classifier using the aforementioned training dataset and obtaining a trained computer classifier, wherein the trained computer classifier is as follows: The tumor burden of a patient's sample is determined based on the number of consensus sequences corresponding to the polynucleotides in the sample that have variants with respect to the total number of consensus sequences corresponding to the total number of polynucleotides detected in the sample, wherein cancer is present in the patient and the patient is independent of multiple individuals having the cancer. Determine the confidence interval for the tumor burden. Determine the change between the confidence interval of the tumor burden for the patient's sample and one or more previous confidence intervals of the tumor burden associated with the patient, and determine if the one or more previous confidence intervals correspond to one or more previous samples of the patient. Based on the aforementioned change, the rate of cancer progression in the patient is determined, and Predicting the patient's treatment response based at least partially on the extent of cancer progression. The steps are configured as follows: Methods that include...

2. The method according to claim 1, wherein the computer classifier is selected from the group consisting of supervised or unsupervised learning algorithms selected from support vector machines, random forests, nearest neighbor analysis, linear regression, binary decision trees, discriminant analysis, logistic classifiers, and cluster analysis.

3. The method according to claim 1, further comprising the step of using the trained computer classifier to predict the direction of tumor development based on examinations at two or more time points.

4. The method according to claim 1, wherein the prediction created comprises determining the probability of distant metastasis development.

5. The method according to claim 1, wherein the training dataset further comprises clinical data selected from the group consisting of cancer stage, type of surgical procedure, age, tumor grade, depth of tumor invasion, occurrence of postoperative complications, and presence of vascular invasion.

6. The method according to claim 1, wherein the genetic information includes variables that define the genomic composition of cancer cells.

7. The method according to claim 1, wherein the genetic information includes variables that define the genomic composition of a single disseminated cancer cell.

8. The method according to claim 1, further comprising the step of preprocessing the training dataset.

9. The method according to claim 8, wherein the step of preprocessing the training dataset includes the step of converting the data into class-conditional probabilities.

10. The method according to claim 1, wherein the genetic information includes sequence or abundance data from one or more gene loci in cell-free DNA derived from the individual.

11. The method according to claim 1, wherein the treatment response includes at least a second, later-prepared genetic information derived from the individual.

12. The method according to claim 6, wherein 20 or fewer variables are selected.

13. The method according to claim 6, wherein 10 or fewer variables are selected.

14. The method according to claim 1, wherein the trained computer classifier uses an artificial neural network.

15. The method according to claim 14, wherein the artificial neural network is trained using a Bayesian framework.

16. A computer execution method for analyzing the disease state of a target, A computer system comprising one or more processors and computer memory receives genetic information of the subject from a gene analyzer, wherein the genetic information includes data acquired at two or more time points, and cancer is detected in the subject; A step of extracting one or more features from the genetic information using an extractor performed by the computer system, wherein the one or more features include one or more genetic variants identified from a plurality of samples obtained from the subject; A step of creating a first output using one or more features as input to a first classifier, which is performed by the computer system using a first machine learning algorithm, wherein the first output represents a first classification of the object, and the first machine learning algorithm is selected from one of a neural network, a support vector machine, a hidden Markov model, or a random forest model; A step of creating a second output using one or more features as input to a second classifier, which is performed by the computer system using a second machine learning algorithm, wherein the second output represents a second classification of the object, and the second machine learning algorithm is selected from another one of neural networks, support vector machines, hidden Markov models, or random forest models; The computer system identifies additional subjects from the population that have genetic information matching the genetic information of the subject, based on the first and second classifications; A step of determining a plurality of scores relating to an additional subject based on past treatments of the subject that are matched to the genetic information, wherein each of the plurality of scores corresponds to one level of responsiveness to a treatment from a plurality of levels of responsiveness to a treatment; The computer system determines a composite score using the plurality of scores; and A step in which a recommender performed by the computer system determines a recommendation indicating a course of action for the subject based on the composite score indicating a high level of responsiveness by the subject to the course of action, wherein the recommendation is the output of the recommender. Methods that include...

17. The method according to claim 16, comprising the steps of comparing the most recent sequence reads with past sequence reads and thereby updating the diagnostic accuracy index.

18. The method according to claim 16, comprising the step of creating a confidence interval for the most recent sequence read.

19. The method according to claim 18, comprising the steps of comparing the confidence interval with one or more past confidence intervals and determining disease progression based on overlapping confidence intervals.

20. The method according to claim 16, further comprising the step of increasing the accuracy index of the diagnosis in a subsequent or prior characterization when information from a first time point confirms information from a second time point.

21. A step of characterizing the genetic information of the subject by determining the frequency of occurrence of one or more gene variants detected in a collection of sequence reads from DNA in a sample derived from the subject, and A step of generating a test result adjusted using information from two or more time points by comparing the frequency of occurrence of one or more gene variants at two or more time points for the subject having matching genetic information. The method according to claim 16, including the method described in claim 16.

22. The steps include: characterizing the genetic information of the subject by determining the amount of copy number variation at one or more loci detected from a collection of sequence reads from DNA in a sample derived from a suitable subject; and Steps to create adjusted test results by comparing the amounts at the two or more time points using information from those two or more time points. The method according to claim 16, including the method described in claim 16.

23. The individual scores of the plurality of scores are The computer system analyzes the genetic information of the additional target using one or more clustering algorithms and determines one or more candidate genomic locations corresponding to at least one level of responsiveness to the treatment. The computer system analyzes the genetic information of the target and determines the probability of the gene variant present at one or more candidate genome locations. The method according to claim 16, as determined by...

24. The method according to claim 16, wherein the genetic information includes sequence data derived from a portion of the genome containing disease-related or cancer-related gene variants.

25. The method according to claim 16, comprising the step of increasing the sensitivity for detecting gene variants by increasing the polynucleotide read depth in the sample derived from the subject at two or more time points.

26. The steps include using the genetic information of the subject as an indicator of the presence of disease polynucleotides in a sample derived from the subject, and If the same gene variant is detected in the noise range of multiple sample collections or time points, the step of adjusting the presence of the disease polynucleotide from negative or uncertain to positive. The method according to claim 16, including the method described in claim 16.

27. The steps include using the genetic information of the subject as an indicator of the presence of disease polynucleotides in a sample derived from the subject, and If the same gene variant is detected within the noise range at an earlier time and beyond the noise range at a later time, the step of adjusting the presence of the disease polynucleotide in the characterization from the earlier time point from negative or uncertain to positive. The method according to claim 16, including the method described in claim 16.