Methods for determining tumor gene copy number by analysis of cell-free DNA

The method corrects for biases in sequencing data by using saturation and probe efficiency corrections to accurately detect copy number variations in tumors from cell-free DNA, overcoming the limitations of invasive biopsies and data biases.

JP7802731B2Active Publication Date: 2026-01-20GUARDANT HEALTH INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023108348
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2015-12-17
Filing Date
2023-06-30
Publication Date
2026-01-20
Estimated Expiration
2036-12-16

AI Technical Summary

Technical Problem

Conventional methods for detecting copy number variations in tumors require invasive biopsies and are limited by biases in sequencing data due to factors like amplification efficiency and GC content, leading to inaccurate results.

Method used

A method for analyzing cell-free DNA from bodily fluids that corrects for saturation equilibrium and probe efficiency using saturation and probe efficiency corrections, followed by removing high variability loci, to determine accurate copy number states.

Benefits of technology

Provides a non-invasive, rapid, and accurate method for detecting copy number variations in tumors by correcting for biases in sequencing data, improving the reliability of copy number detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007802731000002
    Figure 0007802731000002
  • Figure 0007802731000003
    Figure 0007802731000003
  • Figure 0007802731000004
    Figure 0007802731000004
Patent Text Reader

Abstract

To provide improved methods for detecting copy number variation in tumor cells from samples derived from cell-free bodily fluids.SOLUTION: In one aspect, the present disclosure provides a method including the steps of: (a) obtaining sequencing reads of DNA molecules of a cell-free bodily fluid sample of a subject; (b) generating from the sequence reads a first data set comprising for each genetic locus in a plurality of genetic loci a quantitative measure related to sequencing read coverage ("read coverage"); (c) correcting the first data set by performing saturation equilibrium correction and probe efficiency correction; (d) determining a baseline read coverage for the first data set, wherein the baseline read coverage relates to saturation equilibrium and probe efficiency; and (e) determining a copy number state for each genetic locus in the plurality of genetic loci relative to the baseline read coverage.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application claims priority to U.S. Provisional Application No. 62 / 269,051, filed December 17, 2015, which is incorporated herein by reference in its entirety. [Background technology]

[0002] Cancer results from the accumulation of mutations in an individual's normal cells, at least some of which result in improperly regulated cell division. Such mutations typically involve copy number alterations, in which the number of copies of a gene in the tumor genome is increased or decreased compared to the subject's non-cancerous cells.

[0003] The detection and characterization of copy number mutations in tumor cells are used to monitor tumor progression, predict patient outcome, and refine treatment selection.However, conventional methods are carried out on cell samples, which are often obtained by painful and time-consuming biopsy.Such biopsies also often only test a small fraction of tumor cells in a subject, and therefore do not necessarily represent the tumor cell population.There is a need for a simpler and more rapid test for copy number mutations in tumors, which does not require cell biopsy, fluorescence in situ hybridization (FISH), comparative genomic hybridization array, or quantitative fluorescence polymerase chain reaction (PCR) assay.

[0004] A particularly difficult problem in determining copy number variation using sequencing data is that loci show variance in their coverage depth due to reasons unrelated to true copy number.For example, amplification efficiency, PCR efficiency and guanine-cytosine content can cause different coverage depths even for individual loci present in samples with the same copy number.Improved methods are needed to eliminate the bias caused by this effect in order to improve copy number detection. Summary of the Invention [Means for solving the problem]

[0005] There is a significant need for improved methods for detecting copy number variations in tumor cells from a sample derived from an acellular body fluid. The present invention addresses this need and provides additional advantages. In one aspect, the present disclosure provides a method including: (a) obtaining sequencing reads for deoxyribonucleic acid (DNA) molecules of an acellular body fluid sample from a subject; (b) generating a first dataset from the sequence reads, the first dataset including a quantitative measure related to sequencing read coverage ("read coverage") for each locus at a plurality of loci; (c) correcting the first dataset by performing a saturation equilibrium correction and a probe efficiency correction; (d) determining a baseline read coverage for the first dataset, the baseline read coverage being related to the saturation equilibrium and probe efficiency; and (e) determining a copy number state for each locus at the plurality of loci relative to the baseline read coverage. In some embodiments, the first dataset includes, for each locus in the plurality of loci, (i) a quantitative measure related to the guanine-cytosine content ("GC content") of the locus. In some embodiments, the method includes, prior to (c), removing loci from the first dataset that are high variability loci, the removing step comprising: (i) fitting a model related to the quantitative measure related to the guanine-cytosine content and the quantitative measure of sequencing read coverage of the locus; and (ii) removing at least 10% of the loci from the loci, wherein removing the loci comprises removing loci that are most divergent from the model, thereby providing a first dataset of baselined genetic loci. In some embodiments, the method includes removing at least 45% of the loci.

[0006] In some embodiments, performing the saturation equilibrium correction comprises: (i) determining, for each locus from the first dataset of baselined loci, a quantitative measure related to the probability that a strand of a DNA molecule from the sample that originates from the locus is represented in a sequencing read; (ii) determining a first transformation for read coverage by relating the read coverage of the first dataset of baselined loci to both the GC content of the first dataset of baselined loci and the quantitative measure related to the probability that a strand of DNA from each locus in the first dataset of baselined loci is represented in a sequencing read; and (iii) transforming the first dataset of baselined data loci into a saturation-corrected dataset by applying the first transformation to the read coverage of each locus from the first dataset of baselined loci to provide a saturation-corrected dataset comprising a first set of transformed read coverages of the first dataset of baselined loci.

[0007] In some embodiments, determining the first transformation comprises: (i) determining a measure related to the central tendency of the read coverage of the first dataset of baselined loci; (ii) determining a function that fits the measure related to the central tendency of the read coverage of the first dataset of baselined loci based on the GC content of the loci and a quantitative measure related to the probability that a strand of DNA from the locus is represented in the sequencing reads; and (iii) determining, for each locus of the first dataset of baselined loci, the difference between the read coverage predicted by the function and the read coverage, wherein the difference is the transformed read coverage. In some embodiments, the function is a surface approximation. In some embodiments provided herein, the surface approximation is a two-dimensional quadratic polynomial.

[0008] In some embodiments, performing probe efficiency correction comprises converting the saturation-corrected dataset into a probe efficiency-corrected dataset by: (i) removing loci from the saturation-corrected dataset that are high variability loci with respect to the first set of transformed read coverages, thereby providing a second dataset of baselined loci; (ii) determining a second transformation for the first set of transformed read coverages related to the probe efficiency of the second dataset of baselined loci; and (iii) transforming the first set of transformed read coverages of the second dataset of baselined loci using the second transformation, thereby providing a probe efficiency-corrected dataset comprising the second set of transformed read coverages of the second dataset of baselined loci. In some embodiments, removing loci that are high variability loci from the first dataset comprises: (i) fitting a model associated with the first set of transformed read coverages of the GC content and saturation-corrected dataset; and (ii) removing at least 10% of the loci from the saturation-corrected dataset, wherein removing loci comprises removing loci that are most divergent from the model, thereby providing a second dataset of baselined loci. In some embodiments provided herein, the removal is at least 45% of the loci.

[0009] In some embodiments, probe efficiency is determined by performing saturation equilibrium correction in one or more reference samples, and probe efficiency is the converted read coverage obtained by performing saturation equilibrium correction.In some embodiments, one or more reference samples are acellular body fluid samples from a subject who does not have cancer.In some embodiments provided herein, one or more reference samples are acellular body fluid samples from a subject who has cancer, and corresponding loci do not undergo copy number alteration.

[0010] In some embodiments, determining the second transformation comprises (i) adapting the probe efficiencies determined for the loci from the one or more reference samples to the first set of read coverages from the second dataset of baselined loci, and (ii) dividing the transformed read coverage for each locus in the second dataset of baselined loci by the predicted probe efficiency based on the adaptation in (i). In some embodiments, the method further comprises (f) determining a third transformation for the second set of transformed read coverages by relating the transformed read coverages of the second dataset of baselined loci to both the GC content of the second dataset of baselined loci and a quantitative measure related to the probability that a strand of DNA from each locus in the second dataset of baselined loci is represented in the sequencing reads, and (g) applying the third transformation to the second set of transformed read coverages to provide a fourth dataset comprising a third set of transformed quantitative read coverages.

[0011] In some embodiments, DNA from the acellular body fluid sample is enriched for a set of loci using one or more oligonucleotide probes complementary to at least a portion of the loci from the set of loci. In some embodiments, the GC content of each locus from the set of loci is a measure related to the central tendency of the guanine-cytosine content of one or more oligonucleotide probes complementary to at least a portion of the loci from the set of loci. In some embodiments, the read coverage of a locus is a measure related to the central tendency of the read coverage of the region of the locus corresponding to one or more oligonucleotide probes. In some embodiments, performing saturation equilibrium correction and performing probe efficiency correction comprises fitting a Langmuir model, wherein the Langmuir model comprises probe efficiency (K) and saturation equilibrium constant (Isat). In some embodiments, K and Isat are empirically determined for each oligonucleotide probe in the one or more oligonucleotide probes. In some embodiments, performing saturation equilibrium correction and performing probe correction comprises fitting the read coverage of the locus to a Langmuir model, assuming that the locus exists in the same copy number state, thereby providing a baseline read coverage. In some embodiments, the same copy number state is diploid. In some embodiments, the baseline read (rad) coverage is a function that depends on probe efficiency and saturation equilibrium.

[0012] In some embodiments, determining the copy number state comprises comparing the read coverage of the locus with a baseline read coverage. In some embodiments, the acellular body fluid is selected from the group consisting of serum, plasma, urine, and cerebrospinal fluid. In some embodiments, the read coverage is determined by mapping the sequencing reads to a reference genome. In some embodiments, obtaining the sequencing reads comprises ligating adapters to DNA molecules from the acellular body fluid from the subject. In some embodiments, the DNA molecules are double-stranded DNA molecules, and the adapters are ligated to the double-stranded DNA molecules such that each adapter differentially tags a complementary strand of the DNA molecule to provide a tagged strand. In some embodiments, determining a quantitative measure associated with the probability that a strand of DNA from a locus is represented in a sequencing read comprises sorting sequencing reads into paired reads and unpaired reads, where (i) each paired read corresponds to a sequence read generated from a first tagged strand and a second differently tagged complementary strand from a double-stranded polynucleotide molecule in the set, and (ii) each unpaired read represents a first tagged strand that does not have a second differently tagged complementary strand from a double-stranded polynucleotide molecule represented in the sequence read in the set of sequence reads. In some embodiments, the method further comprises determining quantitative measures of (i) the paired reads and (ii) the unpaired reads that map to each of one or more loci, and determining a quantitative measure associated with the total double-stranded DNA molecules in the sample that map to each of the one or more loci based on the quantitative measures associated with the paired reads and unpaired reads that map to each locus. In some embodiments, the adapter comprises a barcode sequence.

[0013] In some embodiments, determining read coverage comprises collapsing the sequencing reads based on the mapping positions of the sequencing reads to the reference genome and barcode sequences. In some embodiments, the loci comprise one or more cancer genes. In some embodiments, the method comprises determining that at least a subset of the baselined loci have undergone copy number alterations in the subject's tumor cells by determining the relative abundance of variants within the baselined loci at which the subject's germline genome is heterozygous. In some embodiments, the relative abundance of variants is not approximately equal. In some embodiments, baselined loci at which the relative abundance of variants is not approximately equal are removed from the baselined loci, thereby providing allele frequency-corrected baselined loci. In some embodiments, the allele frequency-corrected baselined loci are used as baselined loci in the method of any one of the preceding claims.

[0014] In another aspect, the disclosure provides a method comprising receiving, into memory, sequencing reads of deoxyribonucleic acid (DNA) molecules of an acellular body fluid sample of a subject; and executing code using a computer processor to perform the following steps: generating from the sequence reads a first dataset comprising a quantitative measure related to sequencing read coverage ("read coverage") for each locus at a plurality of loci; correcting the first dataset by performing a saturation equilibrium correction and a probe efficiency correction; determining a baseline read coverage for the first dataset, wherein the baseline read coverage is related to the saturation equilibrium and probe efficiency; and determining a copy number state for each locus at the plurality of loci compared to the baseline read coverage.

[0015] In another aspect, the disclosure provides a system including a network; a database including a computer memory connected to the network and configured to store nucleic acid (e.g., DNA) sequence data; and a bioinformatics computer including the computer memory and one or more computer processors, the computer connected to the network, the system further including machine-executable code that, when executed by the one or more computer processors, performs steps including: copying the nucleic acid (e.g., DNA) sequence data stored in the database and writing the copied data to memory in the bioinformatics computer; generating from the nucleic acid (e.g., DNA) sequence data a first dataset including a quantitative measure related to sequencing read coverage ("read coverage") for each locus at a plurality of loci; correcting the first dataset by performing a saturation equilibrium correction and a probe efficiency correction; determining a baseline read coverage for the first dataset, the baseline read coverage related to the saturation equilibrium and probe efficiency; and determining a copy number state for each locus at the plurality of loci relative to the baseline read coverage. In some embodiments, the database is connected to a DNA sequencer.

[0016] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.

[0017] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 illustrates exemplary cancer genes and targets for sequence capture probes.

[0019] [Figure 2] FIG. 2 illustrates gene-level signal across three spike-ins versus theoretical copy number and probe-level signal variation across spike-in genes.

[0020] [Figure 3] FIG. 3 illustrates a bait optimization experiment relating bait amounts with specific molecular counts.

[0021] [Figure 4A] Figures 4A and 4B illustrate the nonlinear effects of p (Figure 4A) and GC content (Figure 4B) on unique molecule counts. [Figure 4B] Figures 4A and 4B illustrate the nonlinear effects of p (Figure 4A) and GC content (Figure 4B) on unique molecule counts.

[0022] [Figure 5] FIG. 5 illustrates the unique molecule counts per probe with no saturation or probe efficiency corrections performed.

[0023] [Figure 6] FIG. 6 illustrates the saturation-corrected unique molecule counts per probe.

[0024] [Figure 7] FIG. 7 illustrates the unique molecule counts per probe after saturation correction and probe efficiency correction.

[0025] [Figure 8] FIG. 8 illustrates the proposed Langmuir model of the interplay between true copy number and unique molecule counting in relation to probe saturation and probe efficiency.

[0026] [Figure 9] FIG. 9 illustrates the probe signal-to-noise reduction of baselined loci after saturation correction, probe efficiency correction, and second round of probe efficiency correction in a typical clinical sample.

[0027] [Figure 10A] Figures 10A and 10B illustrate saturation-corrected UMC plotted against probe efficiencies determined in reference samples to perform probe efficiency correction. Figure 10A is from a subject without copy number alterations in tumor cells. Figure 10B is from a subject with copy number alterations in tumor cells. [Figure 10B] Figures 10A and 10B illustrate saturation-corrected UMC plotted against probe efficiencies determined in reference samples to perform probe efficiency correction. Figure 10A is from a subject without copy number alterations in tumor cells. Figure 10B is from a subject with copy number alterations in tumor cells.

[0028] [Figure 11] Figure 11 illustrates the final report of saturation and probe efficiency corrected copy number variation detection in patient samples. The asterisk above the sample indicates the gene amplification detected based on the corrected signal and minor-allele frequency corrected baseline optimization.

[0029] [Figure 12] FIG. 12 illustrates a computer system 1201 that is programmed or otherwise configured to implement the methods of the present disclosure.

[0030] [Figure 13]13 illustrates the observed copy number (CN) versus theoretical CN of the gene ERBB2 measured using the methods of the present disclosure. Solid dots represent the observed copy number of approximately 2 (diploid sample), open dots represent detected amplification events, and the thick horizontal dashed line marks the average gene CN cutoff.

[0031] [Figure 14] 14 illustrates the observed copy number (CN) versus theoretical CN of the gene ERBB2 measured using the disclosed method (dots) compared to a control method (squares). Filled dots represent an observed copy number of approximately 2 (diploid sample), open dots represent detected amplification events, and the thick horizontal dashed line marks the average gene CN cutoff.

[0032] [Figure 15] FIG. 15 illustrates probe copy numbers plotted for probes used in validation studies for the disclosed method (triangles) versus the control method (X). DETAILED DESCRIPTION OF THE INVENTION

[0033] definition The term "genetic variant" as used herein generally refers to an alteration, variant, or polymorphism in a nucleic acid sample or genome of a subject. Such alteration, variant, or polymorphism may be relative to a reference genome, which may be the reference genome of the subject or another individual. Single nucleotide polymorphism (SNP) is one form of polymorphism. In some examples, one or more polymorphisms include one or more single nucleotide variations (SNVs), insertions, deletions, repeats, small insertions, small deletions, small repeats, structural variant junctions, variable length tandem repeats, and / or flanking sequences. Copy number variants (CNVs), transversions, and other rearrangements are also forms of genetic variation. Genomic alterations may be base changes, insertions, deletions, repeats, copy number variations, or transversions.

[0034] The term "polynucleotide" as used herein generally refers to a molecule comprising one or more nucleic acid subunits. A polynucleotide can comprise one or more subunits selected from adenosine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), or variants thereof. A nucleotide can comprise A, C, G, T, or U, or variants thereof. A nucleotide can comprise any subunit that can be incorporated into a growing nucleic acid chain. Such subunits can be A, C, G, T, or U, or any other subunit specific to one or more complementary A, C, G, T, or U, or complementary to a purine (i.e., A or G, or variants thereof) or pyrimidine (i.e., C, T, or U, or variants thereof). A subunit can allow for the resolution of individual nucleic acid bases or base groups (e.g., AA, TA, AT, GC, CG, CT, TC, GT, TG, AC, CA, or their uracil counterparts). In some examples, the polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), or a derivative thereof. The polynucleotide can be single-stranded or double-stranded.

[0035] The term "subject," as used herein, generally refers to animals, such as mammalian species (e.g., humans) or avian (e.g., avian) species, or other organisms, such as plants. More specifically, a subject can be a vertebrate, mammal, mouse, primate, monkey, or human. Animals include, but are not limited to, farm animals, sport animals, and pets. A subject can be a healthy individual, an individual having or suspected of having a disease or predisposition to a disease, or an individual in need of treatment or suspected of needing treatment. A subject can be a patient.

[0036] The term "genome" generally refers to the entirety of an organism's genetic information. A genome can be coded in either DNA or RNA. A genome can include coding regions that code for proteins as well as non-coding regions. A genome can include all chromosomal sequences in an organism together. For example, the human genome has a total of 46 chromosomes. All of these sequences together make up the human genome.

[0037] The terms "adapter(s)," "adapter(s)," and "tag(s)" are used interchangeably throughout this specification. An adapter or tag can be coupled to a polynucleotide sequence to be "tagged" by any approach, including ligation, hybridization, or other approaches.

[0038] The term "library adaptor" or "library adaptor," as used herein, generally refers to a molecule (e.g., polynucleotide) whose identity (e.g., sequence) can be used to distinguish polynucleotides in a biological sample (also referred to herein as a "sample").

[0039] The term "sequencing adapter," as used herein, generally refers to a molecule (e.g., a polynucleotide) adapted to enable a sequencing instrument to sequence a target polynucleotide, such as by interacting with the target polynucleotide to enable sequencing. A sequencing adapter enables a target polynucleotide to be sequenced by a sequencing instrument. In one example, a sequencing adapter comprises a nucleotide sequence that hybridizes or binds to a capture polynucleotide attached to a solid support of a sequencing system, such as a flow cell. In another example, a sequencing adapter comprises a nucleotide sequence that hybridizes or binds to a polynucleotide to generate a hairpin loop that enables the target polynucleotide to be sequenced by a sequencing system. A sequencing adapter can comprise a sequencer motif, which can be a nucleotide sequence that is complementary to the flow cell sequence of another molecule (e.g., a polynucleotide) and can be used to sequence the target polynucleotide by a sequencing system. A sequencer motif can also comprise a primer sequence for use in sequencing, such as sequencing by synthesis. A sequencer motif can comprise a sequence(s) required for coupling the library adapter to a sequencing system and sequencing the target polynucleotide.

[0040] As used herein, the terms "at least," "at most," or "about," when preceding a series, refer to every member of the series, unless otherwise specified.

[0041] The term "about" and its grammatical equivalents in reference to a reference numerical value can include a range of values ​​from that value up to plus or minus 10%. For example, the amount "about 10" can include an amount of 9 to 11. In other embodiments, the term "about" in reference to a reference numerical value can include a range of values ​​from that value plus or minus 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%.

[0042] The term "at least" and its grammatical equivalents in reference to a reference numerical value can include the reference numerical value and more than that value. For example, the amount "at least 10" can include the value 10 as well as any number greater than 10, such as 11, 100, and 1,000.

[0043] The term "at most" and its grammatical equivalents in reference to a reference numerical value can include the reference numerical value and less than that value. For example, the amount "at most 10" can include the value 10 as well as any number less than 10, such as 9, 8, 5, 1, 0.5, and 0.1.

[0044] The term "quantitative measure" refers to any measure of quantity, including absolute and relative measures. A quantitative measure can be, for example, a number (e.g., a count), a percentage, a degree, or a threshold.

[0045] The term "read coverage" refers to coverage by raw or processed sequence reads, such as unique molecule counts inferred from raw sequence reads.

[0046] The term "baseline read coverage" refers to the expected read coverage of a probe in a sample containing a diploid genomic environment based on given probe parameters, such as GC content, probe efficiency, ligation efficiency, or pull-down efficiency.

[0047] "Probe" as used herein refers to a polynucleotide that includes a functionality, which can be a detectable label (fluorescence), a binding moiety (biotin), or a solid support (magnetically attractable particle or chip).

[0048] "Complementarity" refers to the ability of a nucleic acid to form hydrogen bond(s) with another nucleic acid sequence, either by traditional Watson-Crick or other non-traditional types. Percent complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (Watson-Crick base pairing) with a second nucleic acid sequence (5, 6, 7, 8, 9, 10 out of 10 being 50%, 60%, 70%, 80%, 90%, and 100% complementary, respectively). "Fully complementary" means that all contiguous residues of a nucleic acid sequence will hydrogen bond with the same number of contiguous residues in a second nucleic acid sequence.

[0049] "Substantially complementary," as used herein, refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions. Sequence identity, such as for purposes of assessing percent complementarity, can be measured by any suitable alignment algorithm, including, but not limited to, the Needleman-Wunsch algorithm (see, e.g., the EMBOSS Needle aligner, optionally with default settings, available at the World Wide Website: ebi.ac.uk / Tools / psa / emboss_needle / nucleotide.html), the BLAST algorithm (see, e.g., the BLAST alignment tool, optionally with default settings, available at blast.ncbi.nlm.nih.gov / Blast.cgi), or the Smith-Waterman algorithm (see, e.g., the EMBOSS Water aligner, optionally with default settings, available at the World Wide Website: ebi.ac.uk / Tools / psa / emboss_water / nucleotide.html). Optimal alignment can be assessed using any suitable parameters for the selected algorithm, including default parameters.

[0050] "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex stabilized by hydrogen bonding between the bases of the nucleotide residues. Hydrogen bonding can occur by Watson-Crick base pairing, Hoogstein binding, or in any other sequence-specific manner according to base complementarity. The complex can comprise two strands forming a duplex structure, three or more strands forming a multistranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction can constitute a step in a larger process, such as the initiation of PCR or the enzymatic cleavage of a polynucleotide with an endonuclease. A second sequence that is complementary to a first sequence is called the "complement" of the first sequence. The term "hybridizable" as applied to polynucleotides refers to the ability of the polynucleotide to form a complex stabilized by hydrogen bonding between the bases of the nucleotide residues in a hybridization reaction.

[0051] The term "stringent hybridization conditions" refers to conditions under which a polynucleotide will hybridize preferentially to its target subsequence and will hybridize to a lesser extent or not at all to other sequences. "Stringent hybridization" in the context of nucleic acid hybridization experiments is sequence-dependent and will be different under different environmental parameters. An extensive guide to nucleic acid hybridization is found in Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology--Hybridization with Nucleic Acid Probes, Part I, Chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assays," Elsevier, New York.

[0052] Generally, highly stringent hybridization and washing conditions are selected to be about 5°C lower than the thermal melting point (Tm) for the specific sequence at a defined ionic strength and pH. The Tm is the temperature (under defined ionic strength and pH) at which 50% of the target sequence hybridizes to a perfectly matched probe. Very stringent conditions are selected to be equal to the Tm for a particular probe.

[0053] Stringent hybridization conditions include water, buffers (phosphate, tris, SSPE, or SSC buffers at pH 6-9 or pH 7-8), buffers containing salts (sodium or potassium) and denaturants (SDS, formamide, or tween), and temperatures of 37°C to 70°C, 60°C to 65°C.

[0054] An example of stringent hybridization conditions for hybridization of complementary nucleic acids with more than 100 complementary residues on Southern or Northern blot filters is 50% formalin and 1 mg heparin at 42°C, with hybridization overnight. An example of highly stringent wash conditions is 0.15 M NaCl at 72°C for approximately 15 minutes. An example of stringent wash conditions is a 0.2x SSC wash at 65°C for 15 minutes (see Sambrook et al. for a description of SSC buffers). A low stringency wash often precedes a high stringency wash to remove background probe signal. An example of an intermediate stringency wash for duplexes of more than 100 nucleotides is 1x SSC at 45°C for 15 minutes. An example of a low stringency wash for duplexes of more than 100 nucleotides is 4-6x SSC at 40°C for 15 minutes. In general, a signal-to-noise ratio of 2x (or higher) of that observed for an unrelated probe in the particular hybridization assay indicates detection of a specific hybridization.

[0055] In one aspect, the disclosure provides a method for analyzing a subject's acellular body fluid sample, comprising: (a) obtaining sequencing reads derived from deoxyribonucleic acid (DNA) molecules of the subject's acellular body fluid sample; (b) generating a first dataset including, for each locus in a plurality of loci, (i) a quantitative measure related to the guanine-cytosine content of the locus and (ii) a quantitative measure related to sequencing read coverage of the locus from the sequencing reads; and (c) (i) removing from the first dataset loci that are high variability loci with respect to the quantitative measure related to sequencing read coverage, thereby providing a first set of remaining loci; and (ii) determining, for each locus from the first set of remaining loci, a quantitative measure related to the probability that a strand of DNA from the sample derived from the locus is represented in the sequencing reads; ii) determining a first transformation for the quantitative measures associated with sequencing read coverage of the first set of remaining loci by relating the quantitative measures to both a quantitative measure associated with the GC content of the first set of remaining loci and a quantitative measure associated with the probability that a strand of DNA from each locus in the first set of remaining loci is represented in a sequencing read; and (iv) transforming the first dataset into a second dataset by applying the first transformation to the sequence read coverage of each locus from the first set of remaining loci to provide a second dataset comprising a first set of transformed quantitative measures of sequencing read coverage of the first set of remaining loci.

[0056] In some embodiments, the method further comprises transforming the second dataset into a third dataset by: (d) removing loci from the second dataset that are high variability loci with respect to the first set of transformed quantitative measures of sequencing read coverage, thereby providing a second set of remaining loci; (e) determining a second transformation for the first set of transformed quantitative measures of sequencing read coverage associated with the efficiency of the second set of remaining loci; and (f) transforming the first set of transformed quantitative measures of sequencing read coverage of the second set of remaining loci using the second transformation, thereby providing a third dataset comprising the second set of transformed quantitative measures associated with sequencing read coverage of the second set of remaining loci in (d, i).

[0057] obtaining sequencing reads from DNA molecules of the cell-free body fluid from the subject; Obtaining sequencing reads from DNA molecules in acellular body fluids of a subject can include obtaining acellular body fluids. Exemplary acellular body fluids can be or be derived from serum, plasma, blood, saliva, urine, synovial fluid, whole blood, lymph, ascites, interstitial fluid or extracellular fluid, gingival crevicular fluid, bone marrow, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, fluid in the space between cells including urine, or any other body fluid. The acellular body fluid can be selected from the group consisting of plasma, urine, or cerebrospinal fluid. The acellular body fluid can be plasma. The acellular body fluid can be urine. The acellular body fluid can be cerebrospinal fluid.

[0058] Nucleic acid molecules, including DNA molecules, can be extracted from acellular body fluids. The DNA molecules can be genomic DNA. The DNA molecules can be derived from cells of a subject's healthy tissue. The DNA molecules can be derived from non-cancerous cells that have undergone somatic mutations. The DNA molecules can be derived from a fetus in a maternal sample. Those skilled in the art will understand that in embodiments in which the DNA molecules are derived from a fetus in a maternal sample, the subject can refer to the fetus even if the sample is maternal. The DNA molecules can be derived from precancerous cells of the subject. The DNA molecules can be derived from cancerous cells of the subject. The DNA molecules can be derived from cells within a subject's primary tumor. The DNA molecules can be derived from a secondary tumor of the subject. The DNA molecules can be circulating DNA. The circulating DNA can include circulating tumor DNA (ctDNA). The DNA molecules can be double-stranded or single-stranded. Alternatively, the DNA molecules can include a combination of double-stranded and single-stranded portions. The DNA molecules need not be cell-free. In some cases, the DNA molecules can be isolated from a sample. For example, the DNA molecule can be cell-free DNA isolated from a bodily fluid, such as serum or plasma.

[0059] Sample can contain various amounts of genome equivalent nucleic acid molecules.For example, the sample of about 30ng DNA can contain about 10,000 haploid human genome equivalents, and for cfDNA, can contain about 200 billion individual polynucleotide molecules.Similarly, the sample of about 100ng DNA can contain about 30,000 haploid human genome equivalents, and for cfDNA, can contain about 600 billion individual molecules.

[0060] Cell-free DNA molecules can be isolated and extracted from bodily fluids using various techniques known in the art. In some cases, cell-free nucleic acids can be isolated, extracted, and prepared using commercially available kits, such as the Qiagen Qiamp® Circulating Nucleic Acid Kit protocol. In other examples, nucleic acids can be quantified using the Qiagen Qubit™ dsDNA HS Assay Kit protocol, the Agilent™ DNA 1000 Kit, or the TruSeq™ sequencing library preparation; low-throughput (LT) protocol. Cell-free nucleic acids can originate from fetuses (through fluids collected from pregnant subjects) or from the subject's own tissues. Cell-free nucleic acids can be derived from neoplasms (e.g., tumors or adenomas).

[0061] Generally, cell-free nucleic acids are extracted and isolated from biological fluids by a partitioning step, in which cell-free nucleic acids present in solution are separated from cells and other insoluble components of the biological fluid. Partitioning can include, but is not limited to, techniques such as centrifugation or filtration. In other cases, cells are first dissolved rather than partitioned from the cell-free nucleic acids. In one example, intact cellular genomic DNA is partitioned by selective precipitation. Cell-free nucleic acids, including DNA, can remain soluble and can be separated and extracted from the insoluble genomic DNA. Generally, nucleic acids can be precipitated using isopropanol precipitation, followed by the addition of buffers specific to different kits and other washing steps. Additional purification steps, such as silica-based columns, can be used to remove contaminants or salts. General steps can be optimized for specific applications. For example, non-specific bulk carrier nucleic acids can be added throughout the reaction to optimize certain aspects of the procedure, such as yield.

[0062] The cell-free DNA molecule can be at most 500 nucleotides in length, at most 400 nucleotides in length, at most 300 nucleotides in length, at most 250 nucleotides in length, at most 225 nucleotides in length, at most 200 nucleotides in length, at most 190 nucleotides in length, at most 180 nucleotides in length, at most 170 nucleotides in length, at most 160 nucleotides in length, at most 150 nucleotides in length, at most 140 nucleotides in length, at most 130 nucleotides in length, at most 120 nucleotides in length, at most 110 nucleotides in length, or at most 100 nucleotides in length.

[0063] The cell-free DNA molecule can be at least 500 nucleotides in length, at least 400 nucleotides in length, at least 300 nucleotides in length, at least 250 nucleotides in length, at least 225 nucleotides in length, at least 200 nucleotides in length, at least 190 nucleotides in length, at least 180 nucleotides in length, at least 170 nucleotides in length, at least 160 nucleotides in length, at least 150 nucleotides in length, at least 140 nucleotides in length, at least 130 nucleotides in length, at least 120 nucleotides in length, at least 110 nucleotides in length, or at least 100 nucleotides in length. In particular, the cell-free nucleic acid can be between 140 and 180 nucleotides in length.

[0064] Cell-free DNA can contain various amounts of DNA molecules from healthy tissues and tumors.Tumor-derived cell-free DNA can be at least 0.1% of the total amount of cell-free DNA in the sample, at least 0.2% of the total amount of cell-free DNA in the sample, at least 0.5% of the total amount of cell-free DNA in the sample, at least 0.7% of the total amount of cell-free DNA in the sample, at least 1% of the total amount of cell-free DNA in the sample, at least 2% of the total amount of cell-free DNA in the sample, at least 3% of the total amount of cell-free DNA in the sample, at least 4% of the total amount of cell-free DNA in the sample, at least 5% of the total amount of cell-free DNA in the sample, at least 10% of the total amount of cell-free DNA in the sample, at least 15% of the total amount of cell-free DNA in the sample, at least 20% of the total amount of cell-free DNA in the sample, at least 25% of the total amount of cell-free DNA in the sample, or at least 30% of the total amount of cell-free DNA in the sample, or more.

[0065] In some cases, the DNA molecules may be sheared during the extraction process and comprise fragments between 100 and 400 nucleotides in length. In some cases, the nucleic acids may be sheared after extraction and comprise nucleotides between 100 and 400 nucleotides in length. In some cases, the DNA molecules are already between 100 and 400 nucleotides in length and no additional shearing is intentionally performed.

[0066] The subject can be an animal. The subject can be a mammal, such as a dog, horse, cat, mouse, rat, or human. The subject can be a human. The subject can be suspected of having cancer. The subject can have a previous cancer diagnosis. The subject's cancer status can be unknown. The subject can be male or female. The subject can be at least 20 years old, at least 30 years old, at least 40 years old, at least 50 years old, at least 60 years old, or at least 70 years old.

[0067] Sequencing can be done by any method known in the art. For example, sequencing techniques include classical techniques (e.g., dideoxy sequencing reactions (Sanger's method) using labeled terminators or primers and gel separation in slabs or capillaries) and next-generation techniques. Exemplary techniques include sequencing-by-synthesis using reversibly terminated labeled nucleotides, pyrosequencing, 454 sequencing, Illumina / Solexa sequencing, allele-specific hybridization to a library of labeled oligonucleotide probes, sequencing-by-synthesis using allele-specific hybridization to a library of labeled clones followed by ligation, real-time monitoring of labeled nucleotide incorporation during the polymerization step, polony sequencing, and SOLiD sequencing. These include targeted sequencing, single molecule real-time sequencing, exon sequencing, electron microscope-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, whole genome sequencing, sequencing by hybridization, capillary electrophoresis, gel electrophoresis, duplex sequencing, cycle sequencing, single base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification in low denaturing temperature-PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single molecule sequencing, real-time sequencing, reverse terminator sequencing, nanopore sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, the sequencing method is massively parallel sequencing, i.e., simultaneously (or in rapid succession) sequencing any of at least 100, 1000, 10,000, 100,000, 1 million, 10 million, 100 million, or 1 billion polynucleotide molecules.In some embodiments, sequencing can be performed by a genetic analyzer, such as those commercially available from Illumina or Applied Biosystems. Sequencing of separated molecules has more recently been demonstrated by sequential or single extension reactions using polymerases or ligases, or by single or sequential differential hybridization with a library of probes. Sequencing can be performed by a DNA sequencer (e.g., a machine designed to perform sequencing reactions). In some embodiments, the DNA sequencer can include or be connected to, for example, a database containing DNA sequence data.

[0068] Sequencing techniques that can be used include, for example, the use of sequencing-by-synthesis systems. In the first step, DNA is sheared into fragments of approximately 300-800 base pairs, and the fragments are blunt-ended. Next, oligonucleotide adapters are ligated to the ends of the fragments. The adapters serve as primers for fragment amplification and sequencing. The fragments can be attached to DNA capture beads, e.g., streptavidin-coated beads, using, for example, adapter B containing a 5'-biotin tag. The bead-attached fragments are PCR-amplified within droplets of an oil-water emulsion. The result is multiple copies of clonally amplified DNA fragments on each bead. In the second step, the beads are captured within wells (picoliter size). Pyrosequencing is performed in parallel on each DNA fragment. The addition of one or more nucleotides generates a light signal, which is recorded by a CCD camera in the sequencing instrument. The signal intensity is proportional to the number of nucleotides incorporated. Pyrosequencing utilizes pyrophosphate (PPi), which is released upon nucleotide addition. PPi is converted to ATP by ATP sulfurylase in the presence of adenosine 5' phosphosulfate. Luciferase uses ATP to convert luciferin to oxyluciferin, a reaction that produces light, which is detected and analyzed.

[0069] Another example of a DNA sequencing technique that can be used is Applied Math™ from Life Technologies Corporation (Carlsbad, Calif.). This is the SOLiD technology from Biosystems. In SOLiD sequencing, genomic DNA is sheared into fragments, and adapters are attached to the 5' and 3' ends of the fragments to generate a fragment library. Alternatively, internal adapters can be introduced by ligating adapters to the 5' and 3' ends of the fragments, circularizing the fragments, digesting the circularized fragments to generate internal adapters, and attaching adapters to the 5' and 3' ends of the resulting fragments, generating a library of mate pairs. Next, a clonal bead population is prepared in a microreactor containing beads, primers, templates, and PCR components. After PCR, the templates are denatured and the beads are enriched to isolate beads with extended templates. The templates in selected beads are subjected to a 3' modification that allows them to be attached to a glass slide. Sequences can be determined by sequential hybridization and ligation of partially random oligonucleotides with a central determinant base (or base pair) identified by a specific fluorophore. After the color is recorded, the ligated oligonucleotides are removed, and the process is then repeated.

[0070] Another example of a DNA sequencing technique that can be used is ion semiconductor sequencing, using, for example, a system sold under the trademark ION TORRENT by Ion Torrent of Life Technologies (South San Francisco, Calif.). Ion semiconductor sequencing is described, for example, in Rothberg et al., An integrated semiconductor device enabling non-optical genome sequencing, Nature 475:348-352 (2011); U.S. Patent Application Publication No. 2010 / 0304982; U.S. Patent Application Publication No. 2010 / 0301398; U.S. Patent Application Publication No. 2010 / 0300895; U.S. Patent Application Publication No. 2010 / 0300559; and U.S. Patent Application Publication No. 2009 / 0026082, the contents of each of which are incorporated herein by reference in their entirety.

[0071] Another example of a sequencing technology that can be used is Illumina sequencing. Illumina sequencing is based on the amplification of DNA on a solid surface using fold-back PCR and tethered primers. Genomic DNA is fragmented, and adapters are added to the 5' and 3' ends of the fragments. DNA fragments attached to the surface of a flow cell channel are extended and cross-linked amplified. The fragments become double-stranded, and the double-stranded molecules are denatured. Multiple cycles of solid-phase amplification and subsequent denaturation can generate millions of clusters of approximately 1,000 copies of single-stranded DNA molecules of the same template in each channel of the flow cell. Primers, DNA polymerase, and four fluorophore-labeled reversibly terminating nucleotides are used to perform sequential sequencing. After nucleotide incorporation, a laser is used to excite the fluorophore, an image is captured, and the identity of the first base is recorded. The 3' terminator and the fluorophore of each incorporated base are removed, and the incorporation, detection, and identification steps are repeated. Sequencing according to this technique is described in U.S. Patent Nos. 7,960,120; 7,835,871; 7,232,656; 7,598,035; 6,911,345; 6,833,246; 6,828,100; 6,306,597; 6,210,891; U.S. Patent Application Publication No. 2011 / 0009278; U.S. Patent Application Publication No. 2007 / 0114362; U.S. Patent Application Publication No. 2006 / 0292611; and U.S. Patent Application Publication No. 2006 / 0024681, each of which is incorporated herein by reference in its entirety.

[0072] Another example of a sequencing technology that can be used includes Pacific Biosciences' (Menlo Park, Calif.) single-molecule, real-time (SMRT) technology. In SMRT, each of the four DNA bases is attached to one of four different fluorescent dyes. These dyes are phosphate-linked. A single DNA polymerase is immobilized with a single molecule of template single-stranded DNA at the bottom of a zero-mode waveguide (ZMW). Incorporation of the nucleotide into the growing strand takes several milliseconds. During this time, the fluorescent label is excited, producing a fluorescent signal, and the fluorescent tag is cleaved off. Detection of the corresponding fluorescence of the dye indicates which base has been incorporated. This process is repeated.

[0073] Another example of a sequencing technique that can be used is nanopore sequencing (Soni and Meller, 2007; Progress toward ultrafast sequencing). DNA sequencing using solid-state nanopores, Clin Chem Vol. 53(11):1996-2001. A nanopore is a small hole with a diameter on the order of 1 nanometer. Immersion of a nanopore in a conducting fluid and application of a potential across it results in a small current due to the conduction of ions through the nanopore. The amount of current that flows is sensitive to the size of the nanopore. As a DNA molecule passes through the nanopore, each nucleotide in the DNA molecule blocks the nanopore to a different extent. Thus, changes in the current passing through the nanopore as the DNA molecule passes through it represent the reading of the DNA sequence.

[0074] Another example of a sequencing technique that can be used involves the use of chemically sensitive field effect transistor (chemFET) arrays to sequence DNA (e.g., as described in U.S. Patent Application Publication No. 2009 / 0026082). In one example of this technique, DNA molecules can be placed in a reaction chamber, and template molecules can be hybridized to a sequencing primer bound to a polymerase. Incorporation of one or more triphosphates into the new nucleic acid strand at the 3' end of the sequencing primer can be detected by a change in current flow via the chemFET. The array can have multiple chemFET sensors. In another example, a single nucleic acid can be attached to a bead, the nucleic acid can be amplified in the bead, and individual beads can be transferred to individual reaction chambers in a chemFET array, each chamber containing a chemFET sensor, allowing the nucleic acid to be sequenced.

[0075] Another example of a sequencing technique that can be used is described, for example, in Moudrianakis, This technique involves the use of an electron microscope, as described by E.N. and Beer M., "Base sequence determination in nucleic acids with the electron microscope," in III, Chemistry and microscopy of guanine-labeled DNA, PNAS 53:564-71 (1965). In one example of this technique, individual DNA molecules are labeled with metallic labels that are distinguishable using an electron microscope. These molecules are then spread out on a flat surface and imaged using an electron microscope to determine their sequence.

[0076] Prior to sequencing, adapter sequences can be attached to the nucleic acid molecules, allowing the nucleic acids to be enriched for particular sequences of interest. Sequence enrichment can occur before or after attachment of the adapter sequences.

[0077] The nucleic acid molecules or enriched nucleic acid molecules can be attached to any sequencing adapter suitable for use in any sequencing platform disclosed herein. For example, the sequence adapter can include a flow cell sequence, a sample barcode, or both. In another example, the sequence adapter can be a hairpin adapter, a Y-shaped adapter, a forked adapter, and / or can include a sample barcode. In some cases, the adapter does not include a sequencing primer region. In some cases, the DNA molecules to which the adapter is attached are amplified, and the amplification products are enriched for specific sequences as described herein. In some cases, the DNA molecules are enriched for specific sequences after preparing a sequencing library. The adapter can include a barcode sequence. The different barcodes can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or more nucleic acid bases (as described throughout this specification or any length), for example, 7 bases. The barcodes can be random, degenerate, semi-degenerate, or defined sequences. In some cases, there is sufficient diversity of barcodes such that substantially each nucleic acid molecule (e.g., at least 70%, at least 80%, at least 90%, or at least 99%) is tagged with a different barcode sequence. In some cases, there is sufficient diversity of barcodes such that substantially each nucleic acid molecule (e.g., at least 70%, at least 80%, at least 90%, or at least 99%) from a particular locus is tagged with a different barcode sequence.

[0078] The sequencing adaptor can include a sequence that can hybridize to one or more sequencing primers. The sequencing adaptor can further include a sequence that hybridizes to a solid support, such as a flow cell sequence. For example, the sequencing adaptor can be a flow cell adaptor. The sequencing adaptor can be attached to one or both ends of a polynucleotide fragment. In another example, the sequencing adaptor can be hairpin-shaped. For example, the hairpin-shaped adaptor can include a complementary double-stranded portion and a loop portion, and the double-stranded portion can be attached (e.g., ligated) to a double-stranded polynucleotide. The hairpin-shaped sequencing adaptor can be attached to both ends of a polynucleotide fragment to generate a circular molecule that can be sequenced multiple times.

[0079] In some cases, none of the library adapters contain a sample identification motif (or sample molecular barcode). Such a sample identification motif can be provided via a sequencing adapter. The sample identification motif can include at least 4, 5, 6, 7, 8, 9, 10, 20, 30, or 40 nucleotide bases of a sequencer, allowing the identification of polynucleotide molecules from a given sample from polynucleotide molecules from other samples. For example, this can allow polynucleotide molecules from two subjects to be sequenced in the same pool, and the sequence reads of the subjects can then be identified.

[0080] The sequencer motif comprises a nucleotide sequence(s) required for coupling the library adaptor to a sequencing system and sequencing the target polynucleotide coupled to the library adaptor. The sequencer motif can comprise a sequence complementary to the flow cell sequence and a sequence (sequencing initiation sequence) that can selectively hybridize to a primer (or priming sequence) for use in sequencing. For example, such a sequencing initiation sequence can be complementary to a primer used in sequencing by synthesis (e.g., Illumina). Such a primer can be included in the sequencing adaptor. The sequencing initiation sequence can be a primer hybridization site.

[0081] In some cases, none of the library adapters contain a complete sequencer motif. The library adapters may contain a partial sequencer motif or no sequencer motif. In some cases, the library adapters include a sequencing initiation sequence. The library adapters may include a sequencing initiation sequence but not a flow cell sequence. The sequencing initiation sequence may be complementary to a primer for sequencing. The primer may be a sequence-specific primer or a universal primer. Such a sequencing initiation sequence may be located in a single-stranded portion of the library adapter. Alternatively, such a sequencing initiation sequence may be a priming site (e.g., a kink or nick) that allows a polymerase to couple to the library adapter during sequencing.

[0082] Adapters can be attached to DNA molecules by ligation. In some cases, adapters can be ligated to double-stranded DNA molecules so that each adapter differentially tags complementary strands of the DNA molecule. In some cases, adapter sequences can be attached by PCR, where a first portion of the single-stranded DNA is complementary to the target sequence, and a second portion includes the adapter sequence.

[0083] Enrichment of specific sequences of interest can be achieved by sequence capture methods. Sequence capture can be achieved using immobilized probes that hybridize to the target of interest. Sequence capture can be achieved using probes attached to functional groups, such as biotin, that allow probes hybridized to specific sequences to be enriched from the sample by pull-down. In some cases, specific sequences, such as adapter sequences from library fragments, can be masked by annealing complementary, non-functionalized polynucleotide sequences to the fragments prior to hybridization to the functionalized probe to reduce non-specific or off-target binding. Sequence probes can target specific genes. Sequence capture probes can target specific loci or genes. Such genes can be cancer genes. Exemplary genes targeted by capture probes include those shown in FIG. 1. Exemplary genes with single mutations (SNVs) include AKT1, ATM, CCNE1, CTNNB1, FGFR1, GNAS, JAK3, MLH1, NPM1, PTPN11, RIT1, TERT, ALK, BRAF, CDH1, EGFR, FGFR2, HNF1A, and KIT Exemplary genes with copy number variations include, but are not limited to, AR, CCNE1, CDK6, ERBB2, FGFR2, KRAS, MYC, PIK3CA, BRAF, CDK4, EGFR, FGFR1, KIT, MET, PDGFRA, and RAF1. Exemplary genes with gene fusions include, but are not limited to, ALK, FGFR2, FGFR3, NTRK1, RET, and ROS1. Exemplary genes with indels include, but are not limited to, EGFR (e.g., in exons 19 and 20), ERBB2 (e.g., in exons 19 and 20), and MET (e.g., skipping exon 14). Exemplary targets can include CCND1 and CCND2. Sequence capture probes can be tiled across genes (e.g., probes can target overlapping regions). Sequence probes can target non-overlapping regions. Sequence probes can be optimized for length, melting temperature, and secondary structure.

[0084] A quantitative measure of guanine-cytosine (GC) content The guanine-cytosine content is the percentage of nitrogenous bases in a DNA molecule that are either guanine or cytosine. A quantitative measure related to the GC content of a locus can be the GC content of the entire locus. A quantitative measure related to the GC content of a locus can be the GC content of an exon region of a gene. A quantitative measure related to the GC content of a locus can be the GC content of a region covered by reads mapping to the locus. A quantitative measure related to the GC content can be the GC content of a sequence capture probe corresponding to the locus. A quantitative measure related to the GC content of a locus can be a measure related to the central tendency of the GC content of a sequence capture probe corresponding to the locus. The measure related to central tendency can be any measure of central tendency, such as the mean, median, or mode. The measure related to central tendency can be the median. The GC content of a given region can be measured by dividing the number of guanosine and cytosine bases across the region by the total number of bases.

[0085] Quantitative measures of sequencing read coverage A quantitative measure related to sequencing read coverage is a measure that indicates the number of reads derived from DNA molecules corresponding to a locus (e.g., a specific position, base, region, gene, or chromosome from a reference genome). To associate a read with a locus, the read can be mapped or aligned to the reference. Software for performing mapping or alignment (e.g., Bowtie, BWA, mrsFAST, BLAST, BLAT) can associate sequencing reads with loci. In the mapping process, certain parameters can be optimized. Non-limiting examples of optimizing the mapping process include masking repetitive regions; using a mapping quality (e.g., MAPQ) score cutoff; using different seed lengths to generate alignments; and limiting the edit distance between genomic positions.

[0086] Quantitative measures related to sequencing read coverage can include the count of the reads associated with loci. In some cases, the count is converted into a new metric to reduce the effect of different sequencing depths, library complexity or locus size. Exemplary metrics are Read Per Kilobase per Million, Fragments Per Kilobase per Million (FPKM), Trimmed Mean of M values ​​(TMM), variance stabilized raw counts, and logarithmically transformed raw counts. Other transformations that can be used for specific applications are known to those skilled in the art.

[0087] Quantitative measures can be determined using folded reads, with each folded read corresponding to an initial template DNA molecule. Methods for folding and quantifying read families can be found in PCT / US2013 / 058061 and PCT / US2014 / 000048, each of which is incorporated herein by reference in its entirety. In particular, folding methods that use sequence information from barcodes and sequencing reads can be used to fold reads into families, so that each family shares at least a portion of the barcode sequence and the sequencing read sequence. Each family is then derived from a single initial template DNA molecule for the majority of the family. The count derived from mapping sequences from a family can be referred to as "unique molecule count" (UMC). In some cases, determining a quantitative measure related to sequencing read coverage includes normalizing the UMC by a metric related to library size to provide a normalized UMC ("normalized UMC"). An exemplary method is to divide the UMC of a locus by the sum of all UMCs; divide the UMC of a locus by the sum of all autosomal UMCs.When comparing multiple sequencing read data sets, UMC can be normalized, for example, by the median UMC of the locus in two sequencing read data sets.In some cases, the quantitative measure related to sequencing read coverage can be the normalized UMC, which is further normalized as follows: (i) the normalized UMC is determined for the corresponding locus from the sequencing reads derived from training samples; (ii) for each locus, the normalized UMC of the sample is normalized by the median normalized UMC of the training samples at the corresponding locus, thereby providing the relative abundance (RA) of the locus.

[0088] A consensus sequence can be identified based on the sequence of a sequencing read, for example, by folding the sequencing read based on the identical sequence within the first 5, 10, 15, 20, or 25 bases. In some cases, folding allows for one difference, two differences, three differences, four differences, or five differences in otherwise identical reads. In some cases, folding uses the mapping position of the read, for example, the mapping position of the initial base of the sequencing read. In some cases, folding uses a barcode, and sequencing reads that share a barcode sequence are folded into a consensus sequence. In some cases, folding uses both the barcode and the sequence of the initial template molecule. For example, all reads that share a barcode and map to the same position in a reference genome can be folded. In another example, all reads that share a barcode and the sequence of the initial template molecule (or a percentage identity to the sequence of the initial template molecule) can be folded.

[0089] In some cases, a quantitative measure of sequencing read coverage is determined for a specific subregion of the genome. The region can be a region corresponding to a bin, a gene of interest, an exon, a sequence probe, a primer amplification product, or a primer binding site. In some cases, the subregion of the genome is a region corresponding to a sequence capture probe. If at least a portion of the read maps to at least a portion of the region corresponding to the sequence capture probe, the read can be mapped to the region corresponding to the sequence capture probe. If at least a portion of the read maps to most of the region corresponding to the sequence capture probe, the read can be mapped to the region corresponding to the sequence capture probe. If at least a portion of the read maps across the center point of the region corresponding to the sequence capture probe, the read can be mapped to the region corresponding to the sequence capture probe. In some cases, a quantitative measure related to the sequencing read coverage of a locus is the median RA of the probes corresponding to the genomic position within the locus. For example, if KRAS is covered by three probes with RAs of 2, 3, and 5, the RA of the locus will be 3.

[0090] “Saturation equilibrium” correction Generally, the methods described herein can be used to increase the specificity and sensitivity of variant calling (e.g., detecting copy number variants) in nucleic acid samples. For example, the methods can reduce the amount of noise or distortion in a data sample, thereby reducing the number of false-positive variants detected. As noise and / or distortion decreases, specificity and sensitivity increase. Noise can be considered to be an unwanted random addition to a signal. Distortion can be considered to be a change in the magnitude of a signal or part of a signal.

[0091] Noise can be introduced by errors in copying and / or reading polynucleotides. For example, in the sequencing process, a single polynucleotide can first be subjected to amplification. Amplification can introduce errors, such that a subset of amplified polynucleotides may contain a base at a particular locus that is not the same as the original base at that locus. Furthermore, during the reading process, a base at any particular locus may be read incorrectly. As a result, a collection of sequence reads may contain a certain percentage of base calls at that locus that are not the same as the original base. In typical sequencing techniques, this error rate can be in the single digits, for example, 2% to 3%. When a collection of molecules that are all presumed to have the same sequence is sequenced, this noise is small enough that the original base can be identified with high confidence.

[0092] However, when parent polynucleotide collection comprises a subset of polynucleotides with sequence variants at specific loci, noise can become a significant problem.For example, this can be the case when cell-free DNA not only comprises germline DNA, but also the DNA from other sources, such as fetal DNA or cancer cell-derived DNA.In this case, if the frequency of molecules with sequence variants is within the same range as the error frequency introduced by sequencing process, true sequence variants may be indistinguishable from noise.This can, for example, interfere with the detection of sequence variants in sample.

[0093] In sequencing process, distortion can be manifested as difference in signal intensity, for example, as the total number of sequence reads produced by molecules in parent populations at the same frequency.Distortion can be introduced by, for example, amplification bias, GC bias or sequencing bias.This can interfere with the detection of copy number variation in sample.GC bias causes unequal representation of GC-rich or GC-poor areas in sequence reading.

[0094] The methods disclosed herein include determining an initial set of loci for use in determining the baseline by removing loci from the dataset whose quantitative measures associated with sequencing read coverage or transformed quantitative measures associated with sequencing read coverage are most different from the predictive model (referred to herein as removing highly variable loci), thereby providing a first set of remaining loci. In some instances, removing these loci includes fitting a model that relates the quantitative measures associated with sequencing read coverage to a quantitative measure related to the GC content of the loci. For example, the predictive model can relate the RA of a locus to the GC content of the locus. In some cases, the predictive model is a regression model, including non-parametric regression models, such as LOESS and LOWESS regression models. In some cases, baselining is accomplished by removing 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, or 70% of the loci that deviate most from the predictive model. In some cases, baselining is accomplished by removing at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, or at least 70% of the loci that deviate most from the predictive model. In some cases, deviation is determined by measuring the remainder of the loci relative to the model. A precise cutoff can be selected to provide for the removal of a specific amount of variation from the remaining loci.

[0095] A method for determining a quantitative measure related to the probability that a DNA strand from a sample derived from a locus is represented in a sequencing read is disclosed in PCT / US2014 / 072383, the entire contents of which are incorporated herein by reference. Determining the quantitative measure can include estimating the number of initial template DNA molecules derived from the locus that were present in the sample. The probability that a double-stranded polynucleotide does not generate a sequence read can be determined based on the relative number of reads representing both strands of the initial template DNA molecule and reads representing only one strand of the initial template DNA molecule.

[0096] The number of undetected initial template DNA molecules in a sample can be estimated based on the relative number of reads representing both strands of the initial template DNA molecule and reads representing only a single strand of the initial template DNA molecule. As an example, counts are recorded for a particular locus, locus A, where 1,000 molecules are paired (e.g., both strands are detected) and 1,000 molecules are unpaired (e.g., only a single strand is detected). Note that the terms "paired" and "unpaired" are distinct from the terms sometimes applied to sequencing reads herein to indicate whether both ends of a molecule or a single end of a molecule are sequenced. Assuming a uniform probability p for each Watson-Crick strand to undergo the process following conversion, the proportion of molecules that fail to undergo the process (are not found) can be calculated as follows: R, the ratio of paired to unpaired molecules = 1,000 / 1,000 = 1, therefore, R = 1 = p 2 / (2p(1-p)). This means that p=2 / 3 and the amount of molecules lost is (1-p) 2= 1 / 9. Thus, in this example, approximately 11% of the converted molecules are lost and not detected. In addition to using the binomial distribution, other methods for estimating the number of undetected molecules include exponential, beta, gamma, or empirical distributions based on the redundancy of observed sequence reads. In the latter case, the distribution of read counts of paired and unpaired molecules can be derived from such redundancy to infer the underlying distribution of original polynucleotide molecules at a particular locus. This can often result in a better estimate of the number of undetected molecules. In some cases, p is a quantitative measure related to the probability that a strand of DNA from a sample derived from a locus is represented in the sequencing read. In some cases, p is derived in a similar way, but different models of read distribution are used (e.g., binomial, Poisson, beta, gamma, and negative binomial distributions).

[0097] The transformation for the quantitative measure related to sequencing read coverage can be determined by relating the quantitative measure from the set of loci with removed highly variable loci or the transformed sequencing read coverage to a quantitative measure related to GC content and a quantitative measure related to the probability that a strand of DNA from the locus is represented in a sequencing read. In some cases, the remaining loci are assumed to be diploid and / or present in the same copy number. In some instances, the transformation is determined by adapting the central tendency-related measure of the quantitative measure related to the sequencing read coverage of the remaining loci with a quantitative measure related to GC content and a quantitative measure related to the probability that a strand of DNA from the locus is represented in a sequencing read. For example, the transformation can be performed by (i) adapting the central tendency of the quantitative measure of the sequencing read coverage of the loci remaining after removing highly variable loci with both a quantitative measure related to GC content and a quantitative measure related to the probability that a strand of DNA from the locus is represented in a sequencing read. In some instances, the measure related to the central tendency of the quantitative measure of sequencing read coverage of the remaining loci is the central tendency of the UMC of the remaining loci. In some instances, a surface approximation is used to fit the surface of the UMC of the remaining loci or the central tendency of the UMC of the remaining loci with (i) a quantitative measure related to GC content and (ii) a quantitative measure related to the probability that a strand of DNA from the locus is represented in a sequencing read. For example, the surface approximation can be a two-dimensional quadratic polynomial surface fit of a measure related to the initial template DNA molecule (e.g., UMC) with the quantitative measure of GC content and p. In some cases, the transformed quantitative measure related to sequencing coverage is an expected value based on the transformation determined above, calculated from (i) a quantitative measure related to GC content and (ii) a quantitative measure related to the probability that a strand of DNA from the locus is represented in a sequencing read.In some cases, the transformed quantitative measure associated with sequencing coverage is the remainder for each locus (e.g., the difference or quotient of the expected quantitative measure associated with sequencing read coverage of the locus based on the surface approximation in the sample and the observed quantitative measure associated with sequencing read coverage of the locus). Optionally, after the transformed quantitative measure associated with sequencing coverage is determined, high variability loci can be removed again based on the new transformed quantitative measure of sequencing read coverage, as described above.

[0098] "Probe Efficiency" Correction Disclosed herein is a method for determining and removing the bias of loci using reference samples.In some cases, the reference sample is the sequencing reads from cell-free DNA from a subject without cancer.In some cases, the reference sample is the sequencing reads from cell-free DNA from a subject with cancer cells that are substantially lacking copy number variation at the locus of interest.In some cases, the reference sample is the sequencing reads from cell-free DNA from a subject with cancer, and in this case, the region suspected of having copy number variation is excluded from analysis.In some cases, the reference sample is the plasma sample from a subject without cancer.In some cases, the reference sample is the plasma sample from a subject with cancer.

[0099] Each locus of the reference sample can be processed as described in "Saturation Equilibrium Correction" above to provide a transformed quantitative measure of sequencing read coverage. In some cases, the transformed quantitative measure related to sequencing coverage is an expected value based on the transformation determined above, calculated from (i) a quantitative measure related to GC content and (ii) a quantitative measure related to the probability that a DNA strand derived from the locus from the reference locus is represented in the sequencing read. In some cases, the transformed quantitative measure related to sequencing coverage is the remainder of each reference locus (e.g., the difference or quotient of the expected quantitative measure related to the sequencing read coverage of the locus based on the surface approximation in the reference sample and the observed quantitative measure related to the sequencing read coverage of the locus). The transformed quantitative measure related to the sequencing read coverage of the locus in the reference sample can be considered the "efficiency" of the locus. For example, an inefficiently amplified locus will have a lower UMC than a highly efficiently amplified locus (present at the same copy number in the sample).

[0100] The transformed quantitative measure associated with the sequencing read coverage of the sample can be corrected based on the determined efficiency of the locus from the reference sample(s). This correction can reduce variation introduced into the sample by the process of producing sequencing reads from the sample, which may be related to ligation efficiency, pull-down efficiency, PCR efficiency, flow cell clustering loss, demultiplexing loss, folding loss, and alignment loss. In one embodiment, the correction involves dividing or subtracting the saturation-transformed quantitative measure of the sequencing coverage of the sample by the expected saturation-transformed quantitative measure associated with the sequencing coverage. In some instances, the expected saturation-transformed quantitative measure associated with the sequencing coverage of the locus is determined by fitting the relationship between the saturation-transformed quantitative measure associated with the sequencing coverage of the locus from the sample and the saturation-transformed quantitative measure associated with the sequencing read coverage of the reference. In some cases, the fitting includes performing a local regression (e.g., LOESS or LOWESS) or a robust linear regression of a saturation-transformed quantitative measure related to the sequencing coverage of the locus from the sample on a saturation-transformed quantitative measure related to the sequencing read coverage of the reference. In some cases, the fitting can be a linear regression, a nonlinear regression, or a nonparametric regression.

[0101] Optionally, the transformed quantitative measure from the probe efficiency correction can be input to a "saturation equilibrium correction" transformation to produce a third, further transformed quantitative measure related to sequencing read coverage with reduced variation. Generally, to further reduce variation within the transformed quantitative measure of sequencing read coverage, the transformed quantitative measure of sequencing coverage can be transformed using any of the methods disclosed herein additional times.

[0102] Gene-level overview The gene-level summary of inferred copy number can be determined based on the transformed quantitative measure of sequencing read coverage, determined as disclosed herein.Copy number can be inferred by comparing it with the baseline selected in the above-mentioned operation by discarding highly variable loci.For example, if the remaining loci are inferred to be diploid in the sample, the loci whose transformed quantitative measure related to sequencing coverage differs from the baseline are inferred to have undergone copy number alteration in tumor cells.In some instances, gene-level z-scores are calculated using the observed gene-level median of probe signal and the estimated standard deviation calculated using the observed probe-level standard deviation estimate for the gene and the whole genome normal diploid probe signal standard deviation.

[0103] Minor allele frequency baseline optimization Provided herein are methods for detecting and correcting errors in the gene-level copy number profile described herein using the minor allele frequency of variants in sequencing reads. Sequence variants present in 10% to 90%, 20% to 80%, 30% to 70%, 40% to 60%, or approximately 50% of sequencing reads derived from nucleic acids derived from acellular body fluids may be heterozygous variants present in the subject's germline sequence. In some instances, loci have been determined to have undergone amplification, as described above. The amount of variant is compared to the inferred copy number to determine whether the variant frequency is inconsistent with the inferred copy number. In one example, heterozygous loci can be tested at loci used to determine the baseline copy number (e.g., loci remaining after excluding high-variability loci). In some cases, multiple loci in a sample may be amplified, resulting in misidentification of this baseline. In such cases, heterozygosity may deviate from a 1:1 ratio, and inaccurate baselines may be detected and corrected. In a second example, a locus can be inferred to be present in triploid copy number based on a transformed quantitative measure related to sequencing read coverage. If a subject's germline genome has one chromosome with a first allele of a locus and a second chromosome with a second allele, then either the first or second allele may be duplicated in the cancer cell.

[0104] Langmuir-like saturation model Without being bound by theory, based on historical clinical data research and targeting experiments involving synthetic spike-in model systems, the Langmuir-like saturation model is disclosed herein, which is assumed to be the governing mechanism of bait-cfDNA interaction.Therefore, in the absence of interfering assay effects (for example, ligation efficiency, PCR amplification bias, sequencing artifacts, etc.), bait pull-down process can be expressed as follows: [ka]

[0105] K in this description is the bait efficiency, which depends on the bait sequence characteristics and its interaction with DNA fragments in the genomic vicinity of the targeted bait location. sat is a saturation parameter driven by the limited initial bait count in the pull-down reaction, which is a function of the total bait pool concentration and the replicate count. The replicate count, as used herein, refers to the relative or absolute amount of sequence capture probe present. For example, sequence capture arrays can provide different molar amounts of probe in the array to account for different probe efficiencies. Figure 8 shows the relationship between bait efficiency K and saturation parameter I. sat 1 illustrates a model relating true copy number and unique molecule counting based on

[0106] Bait efficiency K is largely driven by GC content, while I sat The variability is driven by a more complex bait depletion mechanism and RNA secondary structure interactions, which can be roughly investigated by studying the interplay between specific molecule counts and total read counts. Apart from the nonlinear pull-down reaction, the probe signal can be further modeled by a multiplicative model, which involves the following assumptions: under the naive model, cfDNA fragments are uniformly distributed by genomic location, and the stochastic sampling process is the dominant factor contributing to coverage variation. Copy number signals (e.g., UMC) can then be modeled by relating the observed UMC to the true molecule counts in the sample, taking into account the effects of the underlying positional cfDNA profile, ligation efficiency, pull-down efficiency, PCR efficiency, flow cell clustering loss, demultiplexing loss, folding loss, and alignment loss.

[0107] Apart from the nonlinear pull-down reaction, the probe signal can be further modeled by a simple multiplicative model, which involves the following assumptions: Under the naive model, cfDNA fragments are uniformly distributed by genomic location, and the stochastic sampling process is the dominant factor contributing to coverage variation. The copy number signal, i.e., the read count associated with a given probe, can then be modeled as follows:

[0108] Observed UMC = true UMC × underlying positional cfDNA profile (bait, cfDNA fragment) × ligation efficiency (position, size, cfDNA fragment) × pull-down efficiency (probe, cfDNA fragment) × PCR efficiency (DNA fragment) × flow cell clustering loss × demultiplexing loss and folding loss × alignment loss (cfDNA fragment sequence).

[0109] This model assumes the multiplicative nature of the model described above: The underlying bait-specific copy number signal can be inferred from the observed UMC (e.g., the UMC of a given sequence capture probe) relative to an established baseline by a series of steps, such as the baseline determination method disclosed herein.

[0110] The method disclosed herein provides an approach for estimating probe efficiency and bait saturation from sample and training sets. Alternatively, such parameters can be inferred by performing a set of bait titration experiments in which the effect of varying target sequence concentrations on UMC is observed for each probe. K, I sat When the UMC is known, it is possible to determine the UMC value or range corresponding to tumor cells that do not undergo copy number alterations. For example, assuming that the majority of loci do not undergo copy number alterations, the observed UMC will be mostly derived from diploid samples. Samples that undergo copy number alterations will have UMCs that are comparable to their corresponding K and I values. satIn some cases, for example, the UMC value or range may be determined by the K and I values ​​of each probe. sat For example, the UMC corresponding to a diploid copy number may differ between the two probes.

[0111] Computer Control System The present disclosure provides a computer control system programmed to implement the disclosed methods. Figure 12 shows a computer system 1201 programmed or otherwise configured to implement the disclosed methods. The computer system 1201 includes a central processing unit (CPU; collectively "processor" and "computer processor" herein) 1205, which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 1201 also includes memory or memory locations 1210 (e.g., random access memory, read-only memory, flash memory), electronic storage 1215 (e.g., hard disk), communication interface 1220 (e.g., network adapter) for communicating with one or more other systems, and peripherals 1225, such as cache, other memory, data storage, and / or electronic display adapters. The memory 1210, storage 1215, interface 1220, and peripherals 1225 communicate with the CPU 1205 via a communication bus (solid lines), such as a motherboard. The storage device 1215 may be a data storage device (or data repository) for storing data. The computer system 1201 may be operably coupled to a computer network ("network") 1230 utilizing a communication interface 1220. The network 1230 may be the Internet, an Internet and / or extranet, or an intranet and / or extranet in communication with the Internet. The network 1230, in some cases, is a telecommunications and / or data network. The network 1230 may include a local area network. The network 1230 may include one or more computer servers, which may enable distributed computing, such as cloud computing.Network 1230 may, in some cases, utilize computer system 1201 to implement a peer-to-peer network that may allow devices coupled to computer system 1201 to operate as clients or servers.

[0112] CPU 1205 can execute sequences of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as memory 1210. The instructions may be directed to CPU 1205, which may then program or otherwise configure CPU 1205 to implement methods of the present disclosure. Examples of operations performed by CPU 1205 may include fetch, decode, execute, and writeback.

[0113] The CPU 1205 may be part of a circuit, such as an integrated circuit. One or more other components of the system 1201 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0114] The storage device 1215 can store files such as drivers, libraries, and saved programs. The storage device 1215 can store user data, such as user preferences and user programs. The computer system 1201 can, in some cases, include one or more additional data storage devices that are external to the computer system 1201, such as located on a remote server that communicates with the computer system 1201 via an intranet or the Internet.

[0115] Computer system 1201 can communicate with one or more remote computer systems via network 1230. For example, computer system 1201 can communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone®, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access computer system 1201 via network 1230.

[0116] The methods described herein can be implemented as machine (e.g., a computer processor) executable code stored in an electronic storage location of computer system 1201, such as memory 1210 or electronic storage 1215. The machine-executable or machine-readable code can be provided in the form of software. In use, the code can be executed by processor 1205. In some cases, the code can be read from storage 1215 and stored in memory 1210 for easy access by processor 1205. In some cases, electronic storage 1215 can be eliminated, and machine-executable instructions are stored in memory 1210.

[0117] The code may be pre-compiled and configured for use by a machine having a processor adapted to execute the code, or may be compiled at run-time. The code may be supplied in a programming language that may be selected to allow the code to run in a pre-compiled or as-compiled manner.

[0118] Aspects of the systems and methods provided herein, such as computer system 1201, can be embodied in programming. Various aspects of the present technology can be considered "products" or "articles of manufacture," typically in the form of machine- (or processor-) executable code and / or associated data held or embodied in some type of machine-readable medium. The machine-executable code can be stored in electronic storage, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" type media can include any or all of the computer, processor, or other tangible memory or its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., that can provide non-transitory storage of software programming at any time. The software, in whole or in part, can sometimes be communicated via the Internet or various other telecommunications networks. Such communication can, for example, enable loading of the software from one computer or processor to another, e.g., from a management server or host computer to an application server computer platform. Thus, other types of media that may bear software elements include light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices via wired and optical landline networks and across various air-links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered media bearing software. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0119] Thus, machine-readable media, such as computer-executable code, can take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s), such as those that may be used to implement databases, etc., as shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media can take the form of electric or electromagnetic signals, or sound or light waves, such as those generated in radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example: a floppy disk, a flexible disk, a hard disk, magnetic tape, any other magnetic medium, a CD-ROM, a DVD or DVD-ROM, any other optical medium, punch cards, paper tape, any other physical storage medium with a pattern of holes, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, a cable or link transporting such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0120] The computer system 1201 may include or communicate with an electronic display 1235 that includes, for example, a user interface (UI) 1240 for providing reports. Examples of UIs include, without limitation, graphical user interfaces (GUIs) and web-based user interfaces.

[0121] The methods and systems of the present disclosure may be implemented by one or more algorithms, which may be implemented by software when executed by the central processing unit 1205. [Example]

[0122] Example 1 Examination of previously generated copy number variation spike-in data revealed significant probe-to-probe signal variation in both raw read counts and UMC, as well as probe / gene-level copy number signal response to underlying copy number alterations. See Figure 2. Figure 3 illustrates inferred copy numbers versus theoretical copy numbers for three genes (CCND1, CCND2, and ERBB2), demonstrating the nonlinear response of normalized coverage to the amount of bait in the sample. These results suggest bait depletion in the pulldown, which was confirmed by the subsequent bait titration effect on adjacent probes within the same gene with substantially larger differences in unique molecule counts (hence observing faster unique molecule count saturation for probes with higher initial UMC).

[0123] Figure 4A illustrates that the UMC associated with each probe has a nonlinear response with respect to probe p. Figure 4B illustrates that the UMC associated with each probe has a nonlinear response with respect to probe GC content.

[0124] Figure 5 illustrates the UMC of a probe without performing saturation correction or probe efficiency correction. Figure 6 shows the same sample after saturation correction. Figure 7 shows the same sample after probe efficiency correction. Variation within genomic locations is reduced at each stage, providing a clearer picture of the copy number alterations underlying the emerging tumor cells. Genes in Figure 7 with a probe efficiency-corrected signal for the median probe greater than 1.2 are called copy number altered. Different levels of probe efficiency-corrected signal may be due to tumor heterogeneity or secondary tumors.

[0125] FIG. 9 shows a typical progression of baselined locus probe signal-to-noise reduction after saturation correction and probe efficiency correction.

[0126] Figure 10A illustrates a plot of probe efficiency of reference sample(s) on the x-axis and saturation-corrected signal of a sample from a subject without copy number mutations in tumor cells. The relationship is approximately linear. Figure 10B illustrates a similar plot from a subject with copy number mutations in tumor cells. The response is not as linear as in Figure 10A. Correction by predicted efficiency, inferred by determining the relationship between probe efficiency from reference sample(s) and UMC after saturation correction of baselined loci (shown in black), would reduce mutations due to different probe efficiencies at loci presumed to have undergone copy number amplification in tumor cells (gray dots). Figure 11 illustrates an exemplary report of UMC after saturation and probe efficiency correction and copy number mutations from a patient sample based on MAF-optimized baselining. Asterisks indicate points indicated as belonging to loci with copy number mutations in the subject's tumor cells.

[0127] Example 2 Cell-free DNA is obtained from cancer patients, a barcoded sequencing library is prepared, and a panel of cancer genes is enriched by sequence capture using probe sets, and the barcoded sequencing library is sequenced.Sequencing reads are mapped to reference genome, and are collapsed into families based on their barcode sequences and mapping positions.For each genome coordinate corresponding to the midpoint of the probe from probe set, the number of read families that span this midpoint is counted to obtain the UMC per probe.The median UMC per probe is determined for each gene.To perform "saturation equilibrium correction", genes are grouped by their median GC content per probe.The genes whose median UMC per probe is significantly different from the genes with similar median GC content per probe are removed.

[0128] For each probe, determine p and GC content as described herein. Using the remaining genes from the previous step, perform a two-dimensional quadratic polynomial surface fit of the median gene-level UMC to probe p and GC content. Determine the expected UMC per probe using a function relating p and GC content to the expected UMC. Determine the remainder of the data set by dividing the observed UMC per probe by the expected UMC per probe. The remainder UMC for each probe is a transformed quantitative measure of sequencing coverage.

[0129] Genes are regrouped by their median GC content per probe, and the genes whose median GC content per probe residual UMC is significantly different from the genes with similar median GC content per probe are removed.Then, as described in the preceding paragraph, by obtaining the residual UMC of reference sample(s), " probe efficiency " correction is performed.Then, the residual UMC of each probe from sample is divided by the residual UMC of each corresponding probe from reference(s) to obtain the UMC after probe efficiency correction.

[0130] Similar to the saturation equilibrium correction described above, the remaining genes are used to perform a two-dimensional quadratic polynomial surface fit of the probe efficiency-corrected UMC to probe p and GC content. A function relating p and GC content to the expected probe efficiency-corrected UMC is used to determine the expected per-probe probe efficiency-corrected UMC. The remainder of the dataset is determined by dividing the observed per-probe probe efficiency-corrected UMC by the expected per-probe probe efficiency-corrected UMC. The probe efficiency-corrected UMC of the remainder of each probe is the post-probe GC-corrected signal.

[0131] The remaining genes are grouped by their median per-probe GC content, and genes whose median probe GC-corrected signal is significantly different from genes with similar median per-probe GC content are removed.

[0132] The process of this example is repeated with the probe GC corrected signal as the starting input instead of the initial UMC.

[0133] For each gene, the median probe GC-corrected signal is used to summarize each gene, and genes whose median probe GC-corrected signal is significantly different from other genes are considered candidates for gene amplification or deletion in tumor cells.

[0134] For each gene, the germline heterozygous allele is determined and the relative frequency of each allele is quantified. The loci used for baselining are found to have approximately a 1:1 ratio of alleles, validating the selection of baselining loci.

[0135] A Z-score is determined for each gene based on the median probe GC-corrected signal at the gene level and the estimated standard deviation from the whole-genome normal diploid probe signal. Genes with a Z-score higher than the cutoff are reported as undergoing gene amplification in tumor cells.

[0136] Example 3 The method described herein was validated by measuring ERBB2 copy number in the disclosed method versus a control method. The disclosed method produced a linear response of observed copy number (CN) versus theoretical copy number, with no false-positive CNV results observed in the normal (healthy) cohort. See Figure 13. Figure 13 shows the inferred gene copy number versus theoretical copy number. Theoretical copy number estimates are shown, with solid dots representing the observed copy number of approximately 2 (diploid samples), open dots representing detected amplification events, and the thick horizontal dashed line marking the average gene CN cutoff. See also Figure 14. Figure 14 depicts the data from Figure 13, with control data represented by squares. All CNVs followed the expected titration trend, dropping to 2.15 copies. Furthermore, the disclosed method reduced the observed "noise" in the data due to reduced variability, allowing CNVs to be more easily identified compared to the control method. See the far right side of Figure 15; triangles represent the disclosed method, while Xs represent the control method.

[0137] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided herein. While the present invention has been described with reference to the above specification, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. Furthermore, it is to be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It is to be understood that various alternatives to the embodiments of the present invention described herein can be employed in practicing the invention. It is therefore intended that the present invention cover any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of the claims and their equivalents be covered thereby. The present invention provides, for example, the following items. (Item 1) (a) obtaining sequencing reads for deoxyribonucleic acid (DNA) molecules of an acellular body fluid sample from a subject; (b) generating from the sequence reads a first dataset comprising a quantitative measure related to sequencing read coverage ("read coverage") for each locus in a plurality of loci; (c) correcting the first data set by performing a saturation equilibrium correction and a probe efficiency correction; (d) determining a baseline read coverage for the first dataset, the baseline read coverage being related to saturation equilibrium and probe efficiency; (e) determining a copy number state for each locus at the plurality of loci relative to the baseline read coverage; A method comprising: (Item 2) 2. The method of claim 1, wherein the first dataset comprises, for each locus in a plurality of loci, a quantitative measure related to the guanine-cytosine content ("GC content") of the locus. (Item 3) and prior to (c), removing loci from the first dataset that are highly variable loci, the removing step comprising: (i) fitting a model related to the quantitative measure related to guanine-cytosine content and the quantitative measure of sequencing read coverage of the locus; (ii) removing at least 10% of the loci from the first dataset, wherein at least 10% of the loci that are most different from the model are removed, thereby providing the first dataset of baselined loci; Item 3. The method according to item 2, comprising: (Item 4) 4. The method of claim 3, comprising removing at least 45% of the loci. (Item 5) performing a saturation equilibrium correction, (i) for each locus from the first dataset of baselined loci, determining a quantitative measure related to the probability that a strand of a DNA molecule from the sample that originates from that locus is represented in the sequencing read; (ii) determining a first transformation for the read coverage by relating the read coverage in the first dataset of baselined loci to both the GC content of the first dataset of baselined loci and the quantitative measure related to the probability that a strand of DNA from each locus in the first dataset of baselined loci is represented in the sequencing read; (iii) applying the first transformation to the read coverage for each locus from the first dataset of baselined loci to provide a saturation-corrected dataset comprising a first set of transformed read coverages for the first dataset of baselined loci. 4. The method of claim 3, comprising converting the first dataset of baselined loci into the saturation-corrected dataset by: (Item 6) 6. The method of claim 5, wherein determining the first transformation comprises: (i) determining a measure related to central tendency of the read coverage of the first dataset of baselined loci; (ii) determining a function that fits the measure related to central tendency of the read coverage of the first dataset of baselined loci based on the quantitative measure related to the GC content of the loci and the probability that a strand of DNA from the locus is represented in the sequencing reads; and (iii) determining, for each locus in the first dataset of baselined loci, the difference between the read coverage predicted by the function and the read coverage, wherein the difference is the transformed read coverage. (Item 7) 7. The method of claim 6, wherein the function is a surface approximation. (Item 8) 8. The method of claim 7, wherein the surface approximation is a two-dimensional quadratic polynomial. (Item 9) performing probe efficiency correction, (i) removing loci from the saturation-corrected dataset that are high variability loci with respect to the first set of transformed read coverages, thereby providing a second dataset of baselined loci; (ii) determining a second transformation for the first set of transformed read coverages related to the probe efficiency of the second dataset of baselined loci; (iii) transforming the first set of transformed read coverages of the second dataset of baselined loci with the second transformation, thereby providing a probe efficiency-corrected dataset comprising a second set of transformed read coverages of the second dataset of baselined loci. 6. The method of claim 5, further comprising converting the saturation-corrected dataset into the probe-efficiency-corrected dataset by: (Item 10) removing loci from the first dataset that are highly variable loci, (i) fitting a model relating to the first set of transformed read coverages of the GC content and saturation corrected dataset; (ii) removing at least 10% of the loci from the saturation-corrected dataset, the loci that are most different from the model, thereby providing the second dataset of baselined loci; Item 10. The method according to item 9, comprising: (Item 11) 11. The method of claim 10, comprising removing at least 45% of the loci. (Item 12) 10. The method of claim 9, wherein the probe efficiency is determined by performing the saturation equilibrium correction on one or more reference samples, and the probe efficiency is the converted read coverage obtained by performing the saturation equilibrium correction. (Item 13) 13. The method of claim 12, wherein the one or more reference samples are acellular body fluid samples from a subject without cancer. (Item 14) 13. The method of claim 12, wherein the one or more reference samples are acellular body fluid samples from a subject with cancer, and the corresponding loci do not undergo copy number alterations. (Item 15) 13. The method of claim 12, wherein determining the second transformation comprises: (i) fitting the probe efficiencies determined for the loci from the one or more reference samples to the first set of read coverages from the second dataset of baselined loci; and (ii) dividing the transformed read coverage for each locus in the second dataset of baselined loci by a predicted probe efficiency based on the fitting in (i). (Item 16) (g) determining a third transformation for the second set of transformed read coverages by relating the transformed read coverages of the second dataset of baselined loci to both the GC content of the second dataset of baselined loci and the quantitative measure related to the probability that a strand of DNA from each locus in the second dataset of baselined loci is represented in the sequencing reads; (h) applying the third transformation to the second set of transformed read coverages to provide a fourth dataset comprising a third set of transformed quantitative read coverages; Item 6. The method of item 5, further comprising: (Item 17) 2. The method of claim 1, wherein the DNA of the acellular body fluid sample is enriched for the set of loci using one or more oligonucleotide probes complementary to at least a portion of the loci from the set of loci. (Item 18) 18. The method of claim 17, wherein the GC content of each locus from the set of loci is a measure related to the central tendency of the guanine-cytosine content of the one or more oligonucleotide probes complementary to at least a portion of the locus from the set of loci. (Item 19) 18. The method of claim 17, wherein the read coverage of the locus is a measure related to the central tendency of the read coverage of the region of the locus corresponding to the one or more oligonucleotide probes. (Item 20) The steps of performing a saturation equilibrium correction and performing a probe efficiency correction include fitting a Langmuir model, wherein the Langmuir model is a function of the probe efficiency (K) and the saturation equilibrium constant (I sat Item 18. The method according to item 17, comprising: (Item 21) K and I sat 21. The method of claim 20, wherein is empirically determined for each oligonucleotide probe in the one or more oligonucleotide probes. (Item 22) 22. The method of claim 21, wherein performing saturation equilibrium correction and performing probe correction comprises fitting the read coverage of the locus to the Langmuir model assuming the locus is present in the same copy number state, thereby providing a baseline read coverage. (Item 23) 23. The method of item 22, wherein the same copy number state is diploid. (Item 24) 23. The method of claim 22, wherein the baseline read coverage is a function dependent on the probe efficiency and the saturation equilibrium. (Item 25) 23. The method of claim 22, wherein determining the copy number state comprises comparing the read coverage of the locus to the baseline read coverage. (Item 26) 10. The method of any one of the preceding items, wherein the acellular body fluid is selected from the group consisting of serum, plasma, urine, and cerebrospinal fluid. (Item 27) 10. The method of any one of the preceding items, wherein the read coverage is determined by mapping the sequencing reads to a reference genome. (Item 28) 10. The method of any one of the preceding items, wherein obtaining the sequencing reads comprises ligating adaptors to the DNA molecules from the acellular body fluid from the subject. (Item 29) 29. The method of claim 28, wherein the DNA molecule is a double-stranded DNA molecule, and the adapters are ligated to the double-stranded DNA molecule such that each adapter differently tags a complementary strand of the DNA molecule to provide a tagged strand. (Item 30) 30. The method of claim 29, wherein determining the quantitative measure related to the probability that a strand of DNA from the locus is represented in the sequencing reads comprises sorting sequencing reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to a sequence read generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in the set, and (ii) each unpaired read represents a first tagged strand that does not have a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule represented among the sequence reads in the set of sequence reads. (Item 31) 31. The method of claim 30, further comprising determining a quantitative measure of (i) the paired reads and (ii) the unpaired reads that map to each of one or more genetic loci, and determining a quantitative measure associated with total double-stranded DNA molecules in the sample that map to each of the one or more genetic loci based on the quantitative measure associated with the paired reads and unpaired reads that map to each locus. (Item 32) 29. The method of claim 28, wherein the adapter comprises a barcode sequence. (Item 33) 33. The method of claim 32, wherein determining the read coverage comprises collapsing the sequencing reads based on the mapping position of the sequencing reads to the reference genome and the barcode sequence. (Item 34) 10. The method of any one of the preceding items, wherein the locus comprises one or more cancer genes. (Item 35) 2. The method of any one of the preceding items, further comprising determining that at least a subset of the baselined loci have undergone copy number alterations in the tumor cells of the subject by determining the relative abundance of variants within the baselined loci at which the subject's germline genome is heterozygous. (Item 36) 36. The method of claim 35, wherein the relative amounts of the variants are not approximately equal. (Item 37) 37. The method of claim 36, wherein the baselined loci in which the relative amounts of the variants are not approximately equal are removed from the baselined loci, thereby providing allele frequency-corrected baselined loci. (Item 38) 38. The method of claim 37, wherein the allele frequency corrected baselined loci are used as the baselined loci in the method of any one of the preceding items. (Item 39) (a) receiving, in a memory, sequencing reads for deoxyribonucleic acid (DNA) molecules of an acellular body fluid sample of a subject; (b) executing the code using a computer processor to perform the following steps: (i) generating from the sequence reads a first dataset comprising a quantitative measure related to sequencing read coverage ("read coverage") for each locus in a plurality of loci; (ii) correcting the first data set by performing a saturation equilibrium correction and a probe efficiency correction; (iii) determining a baseline read coverage for the first dataset, wherein the baseline read coverage is related to saturation equilibrium and probe efficiency; (iv) determining a copy number state for each locus at the plurality of loci relative to the baseline read coverage; and A method comprising: (Item 40) (a) a network; (b) a database connected to the network, the database comprising a computer memory configured to store nucleic acid sequence data; (c) a bioinformatics computer connected to said network, said computer computer including a computer memory and one or more computer processors; A system comprising: The computer, when executed by the one or more computer processors, copies the nucleic acid sequence data stored in the database, writes the copied data to memory in the bioinformatics computer, and performs the following: (i) generating from the nucleic acid sequence data a first dataset comprising a quantitative measure related to sequencing read coverage ("read coverage") for each locus in a plurality of loci; (ii) correcting the first data set by performing a saturation equilibrium correction and a probe efficiency correction; (iii) determining a baseline read coverage for the first dataset, wherein the baseline read coverage is related to saturation equilibrium and probe efficiency; (iv) determining a copy number state for each locus at the plurality of loci relative to the baseline read coverage; The system further comprises machine-executable code for performing steps including: (Item 41) Item 41. The system of item 40, wherein the database is connected to a nucleic acid sequencer.

Claims

1. (a) obtaining sequencing reads of cell-free deoxyribonucleic acid (DNA) molecules of a sample of fluid in the spaces between cells, including serum, plasma, blood, saliva, urine, synovial fluid, whole blood, lymph, peritoneal fluid, interstitial or extracellular fluid, gingival crevicular fluid, bone marrow, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, or urine, from a subject; (b) generating a first dataset from the sequencing reads for each locus in a plurality of loci, the first dataset comprising (i) a quantitative measure related to sequencing read coverage (“read coverage”), and (ii) a quantitative measure related to guanine-cytosine content (“GC content”) of the locus; (c) removing loci from the first dataset that are highly variable loci, (i) fitting a model related to the quantitative measure related to guanine-cytosine content and the quantitative measure of sequencing read coverage of the locus; (ii) removing at least 10% of the loci from the first dataset, wherein removing the loci comprises removing at least 10% of the loci that are most different from the model, thereby providing the first dataset of baselined loci. removing the (d) correcting the first data set by performing a saturation equilibrium correction and a probe efficiency correction, performing a saturation equilibrium correction, (i) for each locus from the first dataset of baselined loci, determining a quantitative measure related to the probability that a strand of a DNA molecule from the sample that originates from that locus is represented in the sequencing read; (ii) determining a first transformation for the read coverage by relating the read coverage in the first dataset of baselined loci to both the GC content of the first dataset of baselined loci and the quantitative measure related to the probability that a strand of DNA from each locus in the first dataset of baselined loci is represented in the sequencing reads; wherein determining the first transformation comprises: (ii-1) determining a measure related to the central tendency of the read coverage of the first dataset of baselined loci; (ii-2) determining a function that fits the measure related to the central tendency of the read coverage of the first dataset of baselined loci based on the quantitative measure related to the GC content of the loci and the probability that a strand of DNA from the locus is represented in the sequencing reads; (ii-3) for each locus of the first dataset of baselined loci, determining the difference between the read coverage predicted by the function and the read coverage, wherein the difference is the transformed read coverage; and and (iii) applying the first transformation to the read coverage for each locus from the first dataset of baselined loci to provide a saturation-corrected dataset comprising a first set of transformed read coverages for the first dataset of baselined loci. transforming the first dataset of baselined loci into the saturation-corrected dataset by: performing probe efficiency correction, (i) removing loci from the saturation-corrected dataset that are high variability loci with respect to the first set of transformed read coverages, thereby providing a second dataset of baselined loci; (ii) determining a second transformation for the first set of transformed read coverages related to probe efficiencies of the second dataset of baselined loci, wherein the probe efficiencies are determined by performing the saturation equilibrium correction in one or more reference samples, and the probe efficiencies are the transformed read coverages obtained by performing the saturation equilibrium correction; wherein determining the second transformation comprises: (ii-1) matching the probe efficiencies determined for the loci from the one or more reference samples to the first set of read coverages from the second dataset of baselined loci; (ii-2) dividing the transformed read coverage for each locus in the second dataset of baselined loci by a predicted probe efficiency based on the fit in (ii-1); Including, (iii) transforming the first set of transformed read coverages of the second dataset of baselined loci with the second transformation, thereby providing a probe efficiency corrected dataset comprising a second set of transformed read coverages of the second dataset of baselined loci. converting the saturation corrected data set into the probe efficiency corrected data set by a correcting step; (e) determining a baseline read coverage for the first dataset, the baseline read coverage being related to saturation equilibrium and probe efficiency; (f) determining a copy number state for each locus at the plurality of loci relative to the baseline read coverage; A method comprising:

2. the method comprising removing at least 45% of the loci; The method of claim 1.

3. removing loci from the first dataset that are high variability loci in the probe efficiency correction, (i) fitting a model relating to the first set of transformed read coverages of the GC content and saturation corrected dataset; (ii) removing at least 10% of the loci from the saturation-corrected dataset, the loci that are most different from the model, thereby providing the second dataset of baselined loci; Including, The method of claim 1.

4. (a) the one or more reference samples are samples of fluid in the intercellular space, including serum, plasma, blood, saliva, urine, synovial fluid, whole blood, lymph, ascites, interstitial or extracellular fluid, gingival crevicular fluid, bone marrow, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, or urine, from a subject without cancer; or (b) the one or more reference samples are samples of fluid in the intercellular space, including serum, plasma, blood, saliva, urine, synovial fluid, whole blood, lymphatic fluid, ascites, interstitial or extracellular fluid, gingival crevicular fluid, bone marrow, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, or urine, from a subject with cancer, and the corresponding genetic loci are free of copy number alterations; The method of claim 1.

5. (iv) determining a third transformation for the second set of transformed read coverages that relates the transformed read coverages of the second dataset of baselined loci to both: (iv-1) the GC content of the second dataset of baselined loci; and (iv-2) the quantitative measure related to the probability that a strand of DNA from each locus in the second dataset of baselined loci is represented in the sequencing reads; (v) applying the third transformation to the second set of transformed read coverages to provide a fourth dataset comprising a third set of transformed quantitative read coverages; The method of claim 1 further comprising:

6. the DNA of the sample is enriched for the set of loci using one or more oligonucleotide probes complementary to at least a portion of the loci from the set of loci; (a) the GC content of each locus from the set of loci is a measure related to the central tendency of the guanine-cytosine content of the one or more oligonucleotide probes complementary to at least a portion of the locus from the set of loci; (b) the read coverage of the locus is a measure related to the central tendency of the read coverage of the region of the locus corresponding to the one or more oligonucleotide probes; or (c) the steps of performing a saturation equilibrium correction and performing a probe efficiency correction include fitting a Langmuir model, the Langmuir model fitting a probe efficiency (K) and a saturation equilibrium constant (I sat ), including K and I sat is empirically determined for each oligonucleotide probe in the one or more oligonucleotide probes, and performing a saturation equilibrium correction and performing a probe correction includes fitting the read coverage of the locus to the Langmuir model assuming the locus is present in the same copy number state, thereby providing a baseline read coverage. (i) the same copy number state is diploid; (ii) the baseline read coverage is a function that depends on the probe efficiency and the saturation equilibrium; or (iii) determining the copy number state comprises comparing the read coverage of the locus to the baseline read coverage; The method of claim 1.

7. (a) the sample is selected from the group consisting of serum, plasma, urine and cerebrospinal fluid; and / or (b) the read coverage is determined by mapping the sequencing reads to a reference genome; The method according to any one of claims 1 to 6.

8. obtaining the sequencing reads comprises ligating adaptors to the DNA molecules from the sample from the subject; (a) the DNA molecules are double-stranded DNA molecules, and the adapters are ligated to the double-stranded DNA molecules such that each adapter differently tags a complementary strand of the DNA molecule to provide tagged strands, and determining the quantitative measure related to the probability that a strand of DNA from the locus is represented in the sequencing reads comprises sorting sequencing reads into paired reads and unpaired reads, wherein (i) each paired read corresponds to a sequence read generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in the set, ii) determining a quantitative measure of (i) the paired reads and (ii) the unpaired reads, wherein each unpaired read represents a first tagged strand that does not have a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule represented among said sequence reads in said set of sequence reads, and further mapping to each of one or more genetic loci; and determining a quantitative measure associated with the total double-stranded DNA molecules in said sample that map to each of said one or more genetic loci based on the quantitative measure associated with paired reads and unpaired reads that map to each locus; or 8. The method of claim 7, wherein (b) the adapter comprises a barcode sequence, and wherein determining the read coverage comprises collapsing the sequencing reads based on the mapping position of the sequencing reads to the reference genome and the barcode sequence.

9. (a) the locus contains one or more cancer genes, and / or (b) determining that at least a subset of the baselined loci have undergone copy number alterations in the subject's tumor cells by determining the relative abundance of variants within the baselined loci at which the subject's germline genome is heterozygous, wherein the relative abundance of the variants is not approximately equal, and the baselined loci at which the relative abundance of the variants is not approximately equal are removed from the baselined loci, thereby providing allele frequency-corrected baselined loci, and wherein the allele frequency-corrected baselined loci are used as the baselined loci in the method of any one of claims 1 to 8. The method according to any one of claims 1 to 8.

10. (a) receiving, into a memory, sequencing reads of cell-free deoxyribonucleic acid (DNA) molecules of a sample of fluid in the spaces between cells, including serum, plasma, blood, saliva, urine, synovial fluid, whole blood, lymph, peritoneal fluid, interstitial or extracellular fluid, gingival crevicular fluid, bone marrow, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, or urine, from a subject; (b) executing the code with a computer processor to perform the following steps: (i) generating a first dataset from the sequencing reads for each locus in a plurality of loci, wherein the first dataset comprises: (a) a quantitative measure related to sequencing read coverage ("read coverage"); and (b) a quantitative measure related to guanine-cytosine content ("GC content") of the locus; (ii) removing loci from the first dataset that are highly variable loci; (a) fitting a model related to the quantitative measure related to guanine-cytosine content and the quantitative measure of sequencing read coverage of the locus; (b) removing at least 10% of the loci from the first dataset, wherein removing the loci comprises removing at least 10% of the loci that are most different from the model, thereby providing the first dataset of baselined loci. removing the (iii) correcting the first data set by performing a saturation equilibrium correction and a probe efficiency correction, performing a saturation equilibrium correction, (a) for each locus from the first dataset of baselined loci, determining a quantitative measure related to the probability that a strand of a DNA molecule from the sample that originates from that locus is represented in the sequencing read; (b) determining a first transformation for the read coverage by relating the read coverage in the first dataset of baselined loci to both the GC content of the first dataset of baselined loci and the quantitative measure related to the probability that a strand of DNA from each locus in the first dataset of baselined loci is represented in the sequencing reads; wherein determining the first transformation comprises: (b-1) determining a measure related to central tendency of the read coverage of the first dataset of baselined loci; (b-2) determining a function that fits the measure related to the central tendency of the read coverage of the first dataset of baselined loci based on the quantitative measure related to the GC content of the loci and the probability that a strand of DNA from the locus is represented in the sequencing reads; (b-3) for each locus of the first dataset of baselined loci, determining the difference between the read coverage predicted by the function and the read coverage, wherein the difference is the transformed read coverage; and and (c) applying the first transformation to the read coverage for each locus from the first dataset of baselined loci to provide a saturation-corrected dataset comprising a first set of transformed read coverages for the first dataset of baselined loci. transforming the first dataset of baselined loci into the saturation-corrected dataset by performing probe efficiency correction, (a) removing loci from the saturation-corrected dataset that are high variability loci with respect to the first set of transformed read coverages, thereby providing a second dataset of baselined loci; (b) determining a second transformation for the first set of transformed read coverages related to probe efficiencies of the second dataset of baselined loci, wherein the probe efficiencies are determined by performing the saturation equilibrium correction in one or more reference samples, and the probe efficiencies are the transformed read coverages obtained by performing the saturation equilibrium correction; wherein determining the second transformation comprises: (b-1) matching the probe efficiencies determined for the loci from the one or more reference samples to the first set of read coverages from the second dataset of baselined loci; (b-2) dividing the transformed read coverage for each locus in the second dataset of baselined loci by a predicted probe efficiency based on the fit of (b-1); Including, (c) transforming the first set of transformed read coverages of the second dataset of baselined loci with the second transformation, thereby providing a probe efficiency corrected dataset comprising a second set of transformed read coverages of the second dataset of baselined loci. converting the saturation corrected data set into the probe efficiency corrected data set by a correcting step; (iv) determining a baseline read coverage for the first dataset, wherein the baseline read coverage is related to saturation equilibrium and probe efficiency; (v) determining a copy number state for each locus at the plurality of loci relative to the baseline read coverage; and A method comprising:

11. (a) a network; (b) a database connected to the network, the database comprising a computer memory configured to store nucleic acid sequence data; (c) a bioinformatics computer connected to said network, said computer computer including a computer memory and one or more computer processors; A system comprising: The computer, when executed by the one or more computer processors, copies the nucleic acid sequence data stored in the database, writes the copied data to memory in the bioinformatics computer, and performs the following: (i) generating a first dataset from the nucleic acid sequence data for each locus in a plurality of loci, the first dataset comprising (a) a quantitative measure related to sequencing read coverage ("read coverage"), and (b) a quantitative measure related to guanine-cytosine content ("GC content") of the locus; (ii) removing loci from the first dataset that are highly variable loci; (a) fitting a model related to the quantitative measure related to guanine-cytosine content and the quantitative measure of sequencing read coverage of the locus; (b) removing at least 10% of the loci from the first dataset, wherein removing the loci comprises removing at least 10% of the loci that are most different from the model, thereby providing the first dataset of baselined loci. removing the (iii) correcting the first data set by performing a saturation equilibrium correction and a probe efficiency correction, performing a saturation equilibrium correction, (a) for each locus from the first dataset of baselined loci, determining a quantitative measure related to the probability that a strand of a DNA molecule from the sample that originates from that locus is represented in the sequencing read; (b) determining a first transformation for the read coverage by relating the read coverage in the first dataset of baselined loci to both the GC content of the first dataset of baselined loci and the quantitative measure related to the probability that a strand of DNA from each locus in the first dataset of baselined loci is represented in the sequencing reads; wherein determining the first transformation comprises: (b-1) determining a measure related to central tendency of the read coverage of the first dataset of baselined loci; (b-2) determining a function that fits the measure related to the central tendency of the read coverage of the first dataset of baselined loci based on the quantitative measure related to the GC content of the loci and the probability that a strand of DNA from the locus is represented in the sequencing reads; (b-3) for each locus of the first dataset of baselined loci, determining the difference between the read coverage predicted by the function and the read coverage, wherein the difference is the transformed read coverage; and Including, (c) applying the first transformation to the read coverage for each locus from the first dataset of baselined loci to provide a saturation-corrected dataset comprising a first set of transformed read coverages for the first dataset of baselined loci. transforming the first dataset of baselined loci into the saturation-corrected dataset by performing probe efficiency correction, (a) removing loci from the saturation-corrected dataset that are high variability loci with respect to the first set of transformed read coverages, thereby providing a second dataset of baselined loci; (b) determining a second transformation for the first set of transformed read coverages related to probe efficiencies of the second dataset of baselined loci, wherein the probe efficiencies are determined by performing the saturation equilibrium correction in one or more reference samples, and the probe efficiencies are the transformed read coverages obtained by performing the saturation equilibrium correction; wherein determining the second transformation comprises: (b-1) matching the probe efficiencies determined for the loci from the one or more reference samples to the first set of read coverages from the second dataset of baselined loci; (b-2) dividing the transformed read coverage for each locus in the second dataset of baselined loci by a predicted probe efficiency based on the fit of (b-1); Including, (c) transforming the first set of transformed read coverages of the second dataset of baselined loci with the second transformation, thereby providing a probe efficiency corrected dataset comprising a second set of transformed read coverages of the second dataset of baselined loci. converting the saturation corrected data set into the probe efficiency corrected data set by a correcting step; (iv) determining a baseline read coverage for the first dataset, wherein the baseline read coverage is related to saturation equilibrium and probe efficiency; (v) determining a copy number state for each locus at the plurality of loci relative to the baseline read coverage; further comprising machine-executable code to perform steps including: system.

Citation Information

Patent Citations

  • Systems and methods to detect rare mutations and copy number variation

    WO2014039556A1

  • Systems and methods to detect rare mutations and copy number variation

    WO2014149134A2

  • Methods and systems for detecting genetic variants

    WO2015100427A1