Probe design of EXON capture to identify EXON-skipping events in RNA-seq
A panel of oligonucleotide probes in RNA-Seq is designed to target specific exome regions, excluding certain genomic areas and focusing on 5'UTR and exon coverage to enhance sequencing efficiency and detect differential methylation patterns in cancer samples.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GUARDANT HEALTH INC
- Filing Date
- 2025-11-20
- Publication Date
- 2026-05-28
AI Technical Summary
Existing RNA-Seq methods face challenges in maximizing sequencing read budgets while minimizing reads consumed by exon-skipping events and repetitive elements, leading to inefficient use of sequencing resources.
A panel of oligonucleotide probes is designed to hybridize with cDNA derived from target exome regions, excluding regions like 3'UTR, SINES, and simple repeats, and focusing on 5'UTR and exon coverage to enrich for informative reads, including probes that span deletion or fusion breakpoints.
This approach maximizes informative sequencing reads by reducing reads consumed by exon-skipping events and repetitive elements, enhancing the detection of differential methylation patterns in cancerous samples, and optimizing sequencing budget allocation.
Smart Images

Figure US2025056462_28052026_PF_FP_ABST
Abstract
Description
PROBE DESIGN OF EXON CAPTURE TO IDENTIFY EXON-SKIPPING EVENTS IN RNA- SEQCROSS-REFERENCE TO OTHER APPLICATIONS
[0001] This application claims the benefit of the priority date of U.S. Provisional Patent Application Nos. 63 / 723,001 , filed on November 20, 2024, which is herein incorporated by reference in its entirety for all purposes.BACKGROUND
[0002] RNA-Seq is now the method of choice to study gene expression and identify novel RNA species. RNA-Seq offers less background noise and a greater dynamic range for detection Compared to DNA microarray-based methods. More importantly, RNA-Seq directly reveals sequence identity, crucial for analysis of unknown genes and novel transcript isoforms. cDNA library preparation from RNA is a required step for RNA-Seq, with each cDNA in an RNA-Seq library is composed of a cDNA insert of certain size flanked by adapter sequences used for amplification and sequencing on a specific platform.
[0003] The cDNA library preparation method varies depending on the RNA species under investigation, which can differ in size, sequence, structural features and abundance. Major considerations include (1 ) howto capture RNA molecules of interest; (2) howto convert RNA to double-stranded cDNAs with defined size ranges; and (3) how to place adapter sequences on the cDNA ends for amplification and sequencing. Sequencing relies on a single flowcell to deliver reads for somatic, epigenomic, and whole exome variant calling across multiple samples. As scale, a key challenge is spreading sequencing read budgets across all desired product features at desired loss of detection (LoD). There is a great need in the art to maximize read budget through probe design, including identification of exon-skipping events that might otherwise unnecessarily consume sequencing read budget in an informative manner.
[0004] Described herein is an RNA panel probe system that includes several probe design strategies to maximize informative sequencing read budget by eliminating uninformative sequence reads due to exon-skipping or repetitive elements, includingextend coverage towards 5’UTR, design probes to cover deletion breakpoints, exclude probes overlapping SINES and simple repeats, and exclude 3’UTR coverage.SUMMARY OF THE INVENTION
[0005] Described herein is a method for processing, comprising contacting converted complementary DNA (cDNA) molecules or amplification products thereof with a panel of different oligonucleotide probes configured to hybridize to cDNA derived from at least 500 target exome regions; and sequencing the enriched DNA or amplification products thereof, wherein each target exome region is associated with a target genomic regions, with a differential methylation pattern in cancerous samples when compared to normal sample. In other embodiments, the method includes enrichingfor probe-bound cDNA to produce enriched cDNA. In other embodiments, each target genomic regions is determined to be differentially methylated in cancer training samples based on criteria comprising a number of cancer samples that comprise a differential methylation In other embodiments, each of target genomic regions is determined to be differentially methylated based on criteria positively correlated with cancer and negatively correlated with non-cancer. In other embodiments, the method includes a mixture panel, wherein one or sub-panels is configured to detect a different biological analyte. In other embodiments, the mixture panel comprises one or more sub-panels, with different oligonucleotide probes. In other embodiments, the method includes one or more sub-panels, with different oligonucleotide probes is configured to detect cDNA, and one or more sub-panels, with different oligonucleotide probes is configured to detect cell-free DNA (cfDNA) molecules. In other embodiments, the cDNA, cfDNA molecules or amplification products thereof comprise adapter sequences at one or both ends. In other embodiments, the cfDNA molecules or amplification products thereof comprise a first adapter sequence at a first end and a second adapter sequence at a second end, wherein the first adapter sequence and the second adapter sequence are different. In other embodiments, the method includes one or more of plurality of the adapter sequences are configured to distinguish sequencing reads for different cfDNA molecules. In other embodiments, the adapter sequences comprise barcodes, unique molecular identifiers, or partition, sample indexes. In other embodiments, the methodincludes joining first adapter polynucleotides to a plurality of the cfDNA molecules. In other embodiments, the amplification products comprises amplifying the cDNA cfDNA molecules joined to said first adapter polynucleotides to produce first amplification products comprising a first adapter sequence. In other embodiments, preparing the amplification products further comprises joining second adapter polynucleotides comprising a second adapter sequence to the first amplification products; and further wherein (i) the second adapter polynucleotides are joined to an end distal to the first adapter sequence of the first amplification products; and (ii) the first adapter sequence and second adapter sequence are different. In other embodiments, preparingthe amplification products further comprises amplifyingthe first amplification products to produce second amplification products; and further wherein the second amplification products comprise (i) the first adapter sequence or a complement thereof, and (ii) the second adapter sequence or a complement thereof.
[0006] Described herein is a panel with a plurality of different oligonucleotides probes. In other embodiments, the panel comprises a mixture panel comprising one or more sub-panels, with different oligonucleotide probes. In other embodiments, a plurality of the one or more of the different oligonucleotide probes is configured to detect cDNA, a plurality of the one or more of the different oligonucleotide probes is configured to detect cell-free DNA (cfDNA) molecules and / or a plurality of the one or more of the different oligonucleotide probes is configured to detect RNA. In other embodiments, the panel comprises a plurality of hotspot region configured for variant calling at 2.5-5% maximum allele fraction (MAF), optionally comprising about 30-70 million sequence reads. In other embodiments, the panel comprises a plurality of hotspot region configured for variant calling at 1- 2.5% maximum allele fraction (MAF). In other embodiments, the cDNA is derived from exome sequences. In other embodiments, the panel is configured to detect 5-10% MAF within 20-80 million exome sequence reads. In other embodiments, the panel includes one or more oligonucleotide probes configured to span a breakpoint, optionally comprising a fusion breakpoint or deletion breakpoint. In other embodiments, the panel includes one or more oligonucleotide probes configured to bindingto an 5’UTR, or exon, optionally including the last 1 , 2, 3 or more exons of a target gene. In other embodiments, the panel comprises one or more oligonucleotide probes configured to avoid to a 3’UTR, shortinterspersed nuclear elements (SINEs) or short repeat element equal or less than 3 monomers. In other embodiments, a target gene of interest comprises a fusion, optionally SLC45A3, deletion, optionally EGFRv3, METex14, etc. and includes 1 , 2, 3 or more more probes spanning the fusion, deletion breakpoint. In other embodiments, the panel includes one or more oligonucleotide probes configured to bindingto an 5’UTR, optionally SLC45A3, or exon, optionally including the last 1 , 2, 3 or more exons of a target gene optionally SLC45A3. This may be further enhanced by including all 5’UTRs for SLC45A3. In other embodiments, the panel comprises one or more oligonucleotide probes configured to avoid to a 3’UTR, optionally androgen receptor v7 (ARV7), short interspersed nuclear elements (SINEs) or short repeat element equal or less than 3 monomers. In other embodiments, the panel includes one or more oligonucleotide probes configured to bindingto an 5’UTR, or exon, optionally including the last 1 , 2, 3 or more exons of a target gene such as KIF5B.
[0007] Described herein is a method for manufacturing a panel of different oligonucleotide probes configured to hybridize to cDNA derived from at least 500 target exome regions; and sequencing the enriched DNA or amplification products thereof, wherein each target exome region is associated with a target genomic regions, with a differential methylation pattern in cancerous samples when compared to normal sample. In other embodiments, the panel comprises a mixture panel comprising one or more sub-panels, with different oligonucleotide probes. In other embodiments, a plurality of the one or more of the different oligonucleotide probes is configured to detect cDNA, a plurality of the one or more of the different oligonucleotide probes is configured to detect cell-free DNA (cfDNA) molecules and / or a plurality of the one or more of the different oligonucleotide probes is configured to detect RNA. In other embodiments, the panel comprises a plurality of hotspot region configured for variant calling at 2.5-5% maximum allele fraction (MAF), optionally comprising about 30-70 million sequence reads. In other embodiments, the panel comprises a plurality of hotspot region configured for variant calling at 1 - 2.5% maximum allele fraction (MAF). In other embodiments, the cDNA is derived from exome sequences. In other embodiments, the panel is configured to detect 5-10% MAF within 20-80 million exome sequence reads. In other embodiments, the panel includes one or more oligonucleotide probes configured to span a breakpoint, optionally comprising a fusion breakpoint ordeletion breakpoint. In other embodiments, the panel includes one or more oligonucleotide probes configured to bindingto an 5’UTR, or exon, optionally including the last 1 , 2, 3 or more exons of a target gene. In other embodiments, the panel comprises one or more oligonucleotide probes configured to avoid to a 3’UTR, short interspersed nuclear elements (SINEs) or short repeat element equal or less than 3 monomers. In other embodiments, a target gene of interest comprises a fusion, optionally SLC45A3, deletion, optionally EGFRv3, METex14, etc. and includes 1 , 2, 3 or more more probes spanningthe fusion, deletion breakpoint. In other embodiments, the panel includes one or more oligonucleotide probes configured to bindingto an 5’UTR, optionally SLC45A3, or exon, optionally including the last 1 , 2, 3 or more exons of a target gene optionally SLC45A3. This may be further enhanced by including all 5’UTRs for SLC45A3. In other embodiments, the panel comprises one or more oligonucleotide probes configured to avoid to a 3’UTR, optionally androgen receptor v7 (ARV7), short interspersed nuclear elements (SINEs) or short repeat element equal or less than 3 monomers. In other embodiments, the panel includes one or more oligonucleotide probes configured to bindingto an 5’UTR, or exon, optionally including the last 1 , 2, 3 or more exons of a target gene such as KIF5B.BRIEF DESCRIPTION OF THE FIGURES
[0008] Figure 1 . SNV calling at 2.5% MAF at 95% LoD needs above 666 diversity. To estimate the feasibility of the 55M somatic read allocation for SNV detection. A VSCORE panel wide variant simulation was performed for clinical samples enriched on a somatic / epi blend to estimating the calling probability of MAFs within the sample. A sigmoidal curve fit for the relationship between the methylation hotspot diversity and SNV calling probability was fit for the varying subpanels / variant MAFs. The diversity estimates (red vertical dashed line) are where this fit sigmoid intersected 95% confidence (green horizontal dashed line)
[0009] Figure 2. 55M somatic reads are only sufficient for 5% panel-wide MAF.
[0010] Figure 3. 20M exome reads can detect 10% MAF at ~45% confidence.
[0011] Figure 4. Exome wide 10% MAF calling would take ~39M exome reads, 5% ~78M reads.
[0012] Figure 5. Calculations of blended ratio to fit within seq budget.
[0013] Figure 6. SLC45A3: 2 more probes spanning 240bp at fusion breakpoint. Request inclusion of all 5’UTRs for SLC45A3.
[0014] Figure 7. Androgen receptorv7 (ARV7): Design probes covering the 3’UTRs of AR down to 400 bp without any of the probes overlapping SINEs or simple repeats.
[0015] Figure 8. EGFRv3 fusion-designed probes. Exon-del-exon of EGFRv3. Proposed 3 fusion-specific probes and one 5’UTR to include spanning read support
[0016] Figure 9. METex14: probe design specific to detect deletion, with3 probes spanning ex14 deletion.
[0017] Figure 10. Repeat overlap: SINEs RNA probes can overlap with SINEs (>80bp); 75 of these overlap 100% ot its sequence with a SINE.
[0018] Figure 11 . Repeat overlap: Simple repeats monomer size <=3 covering >50% (60bp) due to Spurious hybridization.
[0019] Figure 12. KIF5B: Include design for last two exons (might include 120bp on 3’UTR) Include design for last two exons (might include 120bp on 3’UTR).
[0020] Figure 13. CD74: Design 2 probe to target the last exon.DETAILED DESCRIPTION
[0021] Exome profiling, including whole exosome profiling can play a powerful and complementary role in sequencing exercises related to cancer detection including DNA and forms such as cell free DNA (cfDNA). Given the limitations of sequencing bandwidth for sequencing breadth, and sequencing depth, along with bath samples, RNA exome profiling is highly informative and efficient in this regard. For example, RNA-seq provides a technique for analyzing RNAs of interest through a series of isolating, depletion / enrichment steps, and complementary DNA (cDNA) generation, which is fully compatible with the existing DNA sequencing platforms that therefore allow for complementary signal generation in parallel. RNA-seq can involve several steps, starting with RNA Isolation: RNA is isolated from tissue and mixed with Deoxyribonuclease (DNase). DNase reduces the amount of genomic DNA. The amount of RNA degradation is checked with gel and capillary electrophoresis and is used to assign an RNA integrity numberto the sample. This RNA quality and the total amount of starting RNA are taken into consideration duringthe subsequent library preparation, sequencing, and analysis steps.
[0022] Thereafter, RNA-seq involves RNA selection / depletion. Here, analysis is focused on signals of interest, where the isolated RNA can either be kept as is, enriched for RNA with 3' polyadenylated (poly(A)) tails to include only eukaryotic mRNA, depleted of ribosomal RNA (rRNA), and / or filtered for RNA that binds specific sequences. RNA molecules having 3' poly(A) tails in eukaryotes are mainly composed of mature, processed, coding sequences. Poly(A) selection is performed by mixing RNA with poly(T) oligomers covalently attached to a substrate. Poly(A) selection has important limitations in RNA biotype detection. Many RNA biotypes are not polyadenylated, including many noncoding RNA and histone-core protein transcripts, or are regulated via their poly(A) tail length (e.g., cytokines) and thus might not be detected after poly(A) selection.
[0023] In some instances, poly(A) selection may display increased 3' bias, especially with lower quality RNA. These limitations can be avoided with ribosomal depletion, removing rRNA that typically represents over 90% of the RNA in a cell. Both poly(A) enrichment and ribosomal depletion steps are labor intensive and could introduce biases, so more simple approaches have been developed to omit these steps. Small RNA targets, such as miRNA, can be further isolated through size selection with exclusion gels, magnetic beads, or commercial kits.
[0024] Finally, providing compatibility with DNA sequencing platform as complementary signal involves cDNA synthesis. Here, RNA is reverse transcribed to cDNA. Whereas amplification subsequent to reverse transcription results in loss of strandedness, which can be avoided with chemical labeling or single molecule sequencing. Fragmentation and size selection are performed to purify sequences that are the appropriate length for the sequencing machine. In some instances, one may opt for the RNA, cDNA, or both to be fragmented with enzymes, sonication, divalent ions, or nebulizers. Fragmentation of the RNA reduces 5' bias of randomly primed-reverse transcription and the influence of primer binding sites. On the other hand, one downside is that the 5' and 3' ends are converted to DNA less efficiently. Fragmentation is followed by size selection, where either small sequences are removed or a tight range of sequence lengths are selected. Because small RNAs like miRNAs are lost, these are analyzed independently. The cDNA for each experiment can be indexed with a hexameror octamer barcode, so that these experiments can be pooled into a single lane for multiplexed sequencing.Samples
[0025] A sample can be any biological sample isolated from a subject. A sample can be a bodily sample. Samples can include body tissues, such as known or suspected solid tumors, whole blood, platelets, serum, plasma, stool, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies, cerebrospinal fluid synovial fluid, lymphatic fluid, ascites fluid, interstitial or extracellular fluid, the fluid in spaces between cells, including gingival crevicular fluid, bone marrow, pleural effusions, cerebrospinal fluid, saliva, mucous, sputum, semen, sweat, urine. Samples are preferably body fluids, particularly blood and fractions thereof, and urine. A sample can be in the form originally isolated from a subject or can have been subjected to further processingto remove or add components, such as cells, or enrich for one component relative to another. Thus, a preferred body fluid for analysis is plasma or serum containing cell-free nucleic acids. A sample can be isolated or obtained from a subject and transported to a site of sample analysis. The sample may be preserved and shipped at a desirable temperature, e.g., room temperature, 4°C, -20°C, and / or -80°C. A sample can be isolated or obtained from a subject at the site of the sample analysis. The subject can be a human, a mammal, an animal, a companion animal, a service animal, or a pet. The subject may have a cancer. The subject may not have cancer or a detectable cancer symptom. The subject may have been treated with one or more cancer therapy, e.g., any one or more of chemotherapies, antibodies, vaccines or biologies. The subject may be in remission. The subject may or may not be diagnosed of being susceptible to cancer or any cancer-associated genetic mutations / disorders.
[0026] The volume of plasma can depend on the desired read depth for sequenced regions. Exemplary volumes are 0.4-40 ml, 5-20 ml, 10-20 ml. For examples, the volume can be 0.5 mL, 1 mL, 5 mL 10 mL, 20 mL, 30 mL, or 40 mL. A volume of sampled plasma may be 5 to 20 mL.
[0027] A sample can comprise various amount of nucleic acid that contains genome equivalents. For example, a sample of about 30 ng DNA can contain about 10,000 (104) haploid human genome equivalents and, in the case of cfDNA, about 200billion (2x1011 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents and, in the case of cfDNA, about 600 billion individual molecules.
[0028] A sample can comprise nucleic acids from different sources, e.g., from cells and cell-free of the same subject, from cells and cell-free of different subjects. A sample can comprise nucleic acids carrying mutations. For example, a sample can comprise DNA carrying germline mutations and / or somatic mutations. Germline mutations refer to mutations existing in germline DNA of a subject. Somatic mutations refer to mutations originating in somatic cells of a subject, e.g., cancer cells. A sample can comprise DNA carrying cancer-associated mutations (e.g., cancer-associated somatic mutations). A sample can comprise an epigenetic variant (i.e. a chemical or protein modification), wherein the epigenetic variant associated with the presence of a genetic variant such as a cancer-associated mutation. In some embodiments, the sample includes an epigenetic variant associated with the presence of a genetic variant, wherein the sample does not comprise the genetic variant.
[0029] Exemplary amounts of cell-free nucleic acids in a sample before amplification range from about 1 fg to about 1 pg, e.g., 1 pg to 200 ng, 1 ng to 100 ng, 10 ngto 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method can comprise obtaining 1 femtogram (fg) to 200 ng.
[0030] Cell-free nucleic acids are nucleic acids not contained within or otherwise bound to a cell or in other words nucleic acids remaining in a sample after removing intact cells. Cell-free nucleic acids include DNA, RNA, and hybrids thereof, including genomic DNA, mitochondrial DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long noncoding RNA (long ncRNA), or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or a hybrid thereof. A cell-free nucleic acid can bereleased into bodily fluid through secretion or cell death processes, e.g., cellular necrosis and apoptosis. Some cell-free nucleic acids are released into bodily fluid from cancer cells e.g., circulating tumor DNA, (ctDNA). Others are released from healthy cells. In some embodiments, cfDNA is cell-free fetal DNA (cffDNA) In some embodiments, cell free nucleic acids are produced by tumor cells. In some embodiments, cell free nucleic acids are produced by a mixture of tumor cells and nontumor cells.
[0031] Cell-free nucleic acids have an exemplary size distribution of about 100-500 nucleotides, with molecules of 110 to about 230 nucleotides representing about 90% of molecules, with a mode of about 168 nucleotides and a second minor peak in a range between 240 to 440 nucleotides. Cell-free nucleic acids can be isolated from bodily fluids through a fractionation or partitioning step in which cell-free nucleic acids, as found in solution, are separated from intact cells and other non-soluble components of the bodily fluid. Partitioning may include techniques such as centrifugation or filtration. Alternatively, cells in bodily fluids can be lysed and cell-free and cellular nucleic acids processed together. Generally, after addition of buffers and wash steps, nucleic acids can be precipitated with an alcohol. Further clean up steps may be used such as silica based columns to remove contaminants or salts. Non-specific bulk carrier nucleic acids, such as Cot-1 DNA, DNA or protein for bisulfite sequencing, hybridization, and / or ligation, may be added throughout the reaction to optimize certain aspects of the procedure such as yield.
[0032] After such processing, samples can include various forms of nucleic acid including double stranded DNA, single stranded DNA and single stranded RNA. In some embodiments, single stranded DNA and RNA can be converted to double stranded forms so they are included in subsequent processing and analysis steps.Analytes
[0033] Analytes can include nucleic acid analytes, and non-nucleic acid analytes. The disclosure provides for detecting genetic variations in biological samples from a subject. Biological samples may include polynucleotides from cancer cells. Polynucleotides may be DNA (e.g., genomic DNA, cDNA), RNA (e.g., mRNA, small RNAs), or any combination thereof. Biological samples may include tumor tissue, e.g.,from a biopsy. In some cases, biological samples may include blood or saliva. In particular cases, biological samples may comprise cell free DNA (“cfDNA”) or circulating tumor DNA (“ctDNA”). Cell free DNA can be present in, e.g., blood.
[0034] Examples of non-nucleic acid analytes include, but are not limited to, lipids, carbohydrates, peptides, proteins, glycoproteins (N-linked or O-linked), lipoproteins, phosphoproteins, specific phosphorylated or acetylated variants of proteins, amidation variants of proteins, hydroxylation variants of proteins, methylation variants of proteins, ubiquity lati on variants of proteins, sulfation variants of proteins, viral proteins (e.g., viral capsid, viral envelope, viral coat, viral accessory, viral glycoproteins, viral spike, etc.), extracellular and intracellular proteins, antibodies, and antigen binding fragments. This further includes receptor, an antigen, a surface protein, a transmembrane protein, a cluster of differentiation protein, a protein channel, a protein pump, a carrier protein, a phospholipid, a glycoprotein, a glycolipid, a cell-cell interaction protein complex, an antigen-presenting complex, a major histocompatibility complex, an engineered T-cell receptor, a T-cell receptor, a B-cell receptor, a chimeric antigen receptor, an extracellular matrix protein, a posttranslational modification (e.g., phosphorylation, glycosylation, ubiquitination, nitrosylation, methylation, acetylation or lipidation) state of a cell surface protein, a gap junction, and an adherens junction.
[0035] In general, the systems, apparatus, methods, and compositions can be used to analyze any number of analytes, further including both nucleic acid analytes and non-nucleic acid analytes. For example, the number of analytes that are analyzed can be at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11 , at least about 12, at least about 13, at least about 14, at least about 15, at least about 20, at least about 25, at least about 30, at least about 40, at least about 50, at least about 100, at least about 1 ,000, at least about 10,000, at least about 100,000 or more different analytes present in a region of the sample orwithin an individual feature of the substrate. Methods for performing multiplexed assays to analyze two or more different analytes will be discussed in a subsequent section of this disclosure.
[0036] One or more nucleic acid analytes and / or non-nucleic acid analytes constitute a set of molecular interactions in a biological system under study (e.g., cells), which may be regarded as “interactome” - the molecular interactions that occurbetween molecules belonging to different biochemical families (proteins, nucleic acids, lipids, carbohydrates, etc.) and also within a given family. In various embodiments, an interactome is a protein-DNA interactome (network formed by transcription factors (and DNA or chromatin regulatory proteins) and their target genes. In other embodiments, interactome refers to protein-protein interaction network(PPI), or protein interaction network (PIN). The methods described herein allow for study and analysis of the interactome. Techniques such as proteogenomics (whole genome sequencing, whole exome sequencing and RNA-seq, and mass spectrometry as examples) can support study of the interactome.Analysis
[0037] The present methods can be used to diagnose presence of conditions, particularly cancer, in a subject, to characterize conditions (e.g., staging cancer or determining heterogeneity of a cancer), monitor response to treatment of a condition, effect prognosis risk of developing a condition or subsequent course of a condition. The present disclosure can also be useful in determiningthe efficacy of a particular treatment option. Successful treatment options may increase the amount of copy number variation or rare mutations detected in subject's blood if the treatment is successful as more cancers may die and shed DNA. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of cancers over time. This correlation may be useful in selecting a therapy. Additionally, if a cancer is observed to be in remission after treatment, the present methods can be used to monitor residual disease or recurrence of disease.
[0038] The types and number of cancers that may be detected may include blood cancers, brain cancers, lung cancers, skin cancers, nose cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, bowel cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, mouth cancers, stomach cancers, solid state tumors, heterogeneous tumors, homogenous tumors and the like. Type and / or stage of cancer can be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure alterations, gene fusions,chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0039] Genetic and other analyte data can also be used for characterizing a specific form of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data may allow characterization of specific sub-types of cancer that may be important in the diagnosis or treatment of that specific sub-type. This information may also provide a subject or practitioner clues regarding the prognosis of a specific type of cancer and allow either a subject or practitioner to adapt treatment options in accord with the progress of the disease. Some cancers can progress to become more aggressive and genetically unstable. Other cancers may remain benign, inactive or dormant. The system and methods of this disclosure may be useful in determining disease progression.
[0040] The present analyses are also useful in determining the efficacy of a particular treatment option. Successful treatment options may increase the amount of copy number variation or rare mutations detected in subject's blood if the treatment is successful as more cancers may die and shed DNA. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of cancers over time. This correlation may be useful in selecting a therapy. Additionally, if a cancer is observed to be in remission after treatment, the present methods can be used to monitor residual disease or recurrence of disease.
[0041] The present methods can also be used for detecting genetic variations in conditions otherthan cancer. Immune cells, such as B cells, may undergo rapid clonal expansion upon the presence certain diseases. Clonal expansions may be monitored using copy number variation detection and certain immune states may be monitored. In this example, copy numbervariation analysis may be performed overtime to produce a profile of how a particular disease may be progressing. Copy number variation or even rare mutation detection may be used to determine how a population of pathogens is changing during the course of infection. This may be particularly important during chronic infections, such as HIV / AIDS or Hepatitis infections, whereby viruses may change life cycle state and / or mutate into more virulent forms during the course ofinfection. The present methods may be used to determine or profile rejection activities of the host body, as immune cells attempt to destroy transplanted tissue to monitor the status of transplanted tissue as well as altering the course of treatment or prevention of rejection.
[0042] Further, the methods of the disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods can include, e.g., generating a genetic profile of extracellular polynucleotides derived from the subject, wherein the genetic profile includes a plurality of data resulting from copy number variation and rare mutation analyses. In some embodiments, an abnormal condition is cancer. In some embodiments, the abnormal condition may be one resulting in a heterogeneous genomic population. In the example of cancer, some tumors are known to comprise tumor cells in different stages of the cancer. In other examples, heterogeneity may comprise multiple foci of disease. Again, in the example of cancer, there may be multiple tumor foci, perhaps where one or more foci are the result of metastases that have spread from a primary site.
[0043] The present methods can be used to generate or profile, fingerprint or set of data that is a summation of genetic information derived from different cells in a heterogeneous disease. This set of data may comprise copy number variation and mutation analyses alone or in combination.
[0044] The present methods can be used to diagnose, prognose, monitor or observe cancers, or other diseases. In some embodiments, the methods herein do not involve the diagnosing, prognosing or monitoring a fetus and as such are not directed to non-invasive prenatal testing. In other embodiments, these methodologies may be employed in a pregnant subject to diagnose, prognose, monitor or observe cancers or other diseases in an unborn subject whose DNA and other polynucleotides may cocirculate with maternal molecules.Captured Set
[0045] In some embodiments, a captured set of DNA (e.g., cDNA, cfDNA) is provided. With respect to the disclosed methods, the captured set of DNA may be provided, e.g., by performing a capturing step after RNA isolation, depletion and cDNA conversion, or a partitioning step. The captured set may comprise DNA correspondingto a sequence-variable target region set, an epigenetic target region set, or a combination thereof. In some embodiments the quantity of captured sequence-variable target region DNA is greater than the quantity of the captured epigenetic target region DNA, when normalized for the difference in the size of the targeted regions (footprint size).
[0046] Alternatively, first and second captured sets may be provided, comprising, respectively, DNA corresponding to a sequence-variable target region set and DNA corresponding to an epigenetic target region set. The first and second captured sets may be combined to provide a combined captured set.
[0047] In some embodiments in which a captured set comprising DNA corresponding to the sequence-variable target region set and the epigenetic target region set includes a combined captured set as discussed above, the DNA corresponding to the sequence-variable target region set may be present at a greater concentration than the DNA corresponding to the epigenetic target region set, e.g., a 1 .1 to 1 .2-fold greater concentration, a 1 .2- to 1 .4-fold greater concentration, a 1 .4- to 1 .6- fold greater concentration, a 1 .6- to 1 .8-fold greater concentration, a 1 .8- to 2.0-fold greater concentration, a 2.0- to 2.2-fold greater concentration, a 2.2- to 2.4-fold greater concentration a 2.4- to 2.6-fold greater concentration, a 2.6- to 2.8-fold greater concentration, a 2.8- to 3.0-fold greater concentration, a 3.0- to 3.5-fold greater concentration, a 3.5- to 4.0, a 4.0- to 4.5-fold greater concentration, a 4.5- to 5.0-fold greater concentration, a 5.0- to 5.5-fold greater concentration, a 5.5- to 6.0-fold greater concentration, a 6.0- to 6.5-fold greater concentration, a 6.5- to 7.0-fold greater, a 7.0- to 7.5-fold greater concentration, a 7.5- to 8.0-fold greater concentration, an 8.0- to 8.5- fold greater concentration, an 8.5- to 9.0-fold greater concentration, a 9.0- to 9.5-fold greater concentration, 9.5- to 10.0-fold greater concentration, a 10- to 11 -fold greater concentration, an 11 - to 12-fold greater concentration a 12- to 13-fold greater concentration, a 13- to 14-fold greater concentration, a 14- to 15-fold greater concentration, a 15- to 16-fold greater concentration, a 16- to 17-fold greater concentration, a 17- to 18-fold greater concentration, an 18- to 19-fold greater concentration, a 19- to 20-fold greater concentration, a 20- to 30-fold greater concentration, a 30- to 40-fold greater concentration, a 40- to 50-fold greater concentration, a 50- to 60-fold greater concentration, a 60- to 70-fold greaterconcentration, a 70- to 80-fold greater concentration, a 80- to 90-fold greater concentration, a 90- to 100-fold greater concentration, a 10- to 20-fold greater concentration, a 10- to 40-fold greater concentration, a 10- to 50-fold greater concentration, a 10- to 70-fold greater concentration, or a 10- to 100-fold greater concentration. The degree of difference in concentrations accounts for normalization for the footprint sizes of the target regions, as discussed in the definition section.
[0048] In some embodiments, RNA and / or associated DNA (e.g., cDNA, cfDNA) is obtained from a subject having a cancer. In some embodiments, the DNA (e.g., cDNA, cfDNA) is obtained from a subject suspected of having a cancer. In some embodiments, the DNA (e.g., cDNA, cfDNA) is obtained from a subject having a tumor. In some embodiments, the DNA (e.g., cDNA, cfDNA) is obtained from a subject suspected of having a tumor. In some embodiments, the DNA (e.g., cDNA, cfDNA) is obtained from a subject having neoplasia. In some embodiments, the DNA (e.g., cDNA, cfDNA) is obtained from a subject suspected of having neoplasia. In some embodiments, the DNA (e.g., cDNA, cfDNA) is obtained from a subject in remission from a tumor, cancer, or neoplasia (e.g., following chemotherapy, surgical resection, radiation, or a combination thereof). In any of the foregoing embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia may be of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the lung. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the colon or rectum. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the breast. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the prostate. In any of the foregoing embodiments, the subject may be a human subject.
[0049] In some embodiments, the sequence-variable target region probe set has a footprint of at least 0.5 kb, e.g., at least 1 kb, at least 2 kb, at least 5 kb, at least 10 kb, at least 20 kb, at least 30 kb, or at least 40 kb. In some embodiments, the epigenetic target region probe set has a footprint in the range of 0.5-100 kb, e.g., 0.5-2 kb, 2-10 kb, 10-20 kb, 20-30 kb, 30-40 kb, 40-50 kb, 50-60 kb, 60-70 kb, 70-80 kb, 80-90 kb, and 90- 100 kb.
[0050] In some embodiments, the probes specific for the sequence-variable target region set comprise probes specific for target regions from at least 10, 20, 30, or 35 cancer-related genes, such as AKT1 , ALK, BRAF, CCND1 , CDK2A, CTNNB1 , EGFR, ERBB2, ESR1 , FGFR1 , FGFR2, FGFR3, FOXL2, GATA3, GNA1 1 , GNAQ, GNAS, HRAS, IDH1 , IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11 , TP53, and U2AF1 .Sequencing panel
[0051] To improve the likelihood of detecting tumor indicating mutations, the region of DNA sequenced may comprise a panel of genes, exomes, or genomic regions. Selection of a limited region for sequencing (e.g., a limited panel) can reduce the total sequencing needed (e.g., a total amount of nucleotides sequenced. A sequencing panel can target a plurality of different genes or regions to detect a single cancer, a set of cancers, or all cancers.
[0052] In some aspects, a panel targets a plurality of different genes, exomes, or genomic regions is selected such that a determined proportion of subjects having a cancer exhibits a genetic variant or tumor marker in one or more different genes, exomes, or genomic regions in the panel. The panel may be selected to limit a region for sequencing to a fixed number of base pairs. The panel may be selected to sequence a desired amount of DNA. The panel may be further selected to achieve a desired sequence read depth. The panel may be selected to achieve a desired sequence read depth or sequence read coverage for an amount of sequenced base pairs. The panel may be selected to achieve a theoretical sensitivity, a theoretical specificity and / or a theoretical accuracy for detecting one or more genetic variants in a sample.
[0053] Probes for detecting the panel of regions can include those for detecting hotspots regions as well as nucleosome-aware probes (e.g., KRAS codons 12 and 13) and may be designed to optimize capture based on analysis of cfDNA coverage and fragment size variation impacted by nucleosome binding patterns and GC sequence composition. Regions used herein can also include non-hotspot regions optimized based on nucleosome positions and GC models. The panel can comprise a plurality of subpanels, including subpanels for identifying tissue of origin (e.g., use of published literature to define 50-100 baits representing genes with most diverse transcriptionprofile across tissues (not necessarily promoters)), whole genome scaffold (e.g., for identifying ultra-conservative genomic content and tiling sparsely across chromosomes with handful of probes for copy number base lining purposes), transcription start site (TSS) / CpG islands (e.g., for capturing differential methylated regions (e.g., Differentially Methylated Regions (DMRs)) in for example in promoters of tumor suppressor genes (e.g., SEPT9 / VIM in colorectal cancer)). In some embodiments, markers for a tissue of origin are tissue-specific epigenetic markers.
[0054] The one or more regions in the panel can comprise one or more loci from one or a plurality of genes. The plurality of genes may be selected for sequencing and tumor marker detection. Genes included in the region to be sequenced may be selected from genes known to be involved in cancer, or from genes not involved in cancer. For example, the plurality of genes in the panel may be oncogenes, tumor suppressors, growth factors, DNA repair genes, signaling genes, transcription factors, receptors or metabolic genes. Examples of genes that may be in the panel include, but are not limited to: APC, AR, ARID1A, BRAF, BRCA1 , BRCA2, CCND1 , CCND2, CCNE1 , CDK4, CDK6, CDKN2A, CDKN2B, EGFR, ERBB2, FGFR1 , FGFR2, HRAS, KIT, KRAS, MET, MYC, NF1 , NRAS, PDGFRA, PIK3CA, PTEN, RAF1 , TP53, AKT1 , ALK, ARAF, ATM, CDH1 , CTNNB1 ,ESR1 , EZH2, FBXW7, FGFR3, GATA3, GNA1 1 , GNAQ, GNAS, HNF1A, IDH1 , IDH2, JAK2, JAK3, MAP2K1 , MAP2K2, MLH1 , MPL, NFE2L2, NOTCH1 , NPM1 , NTRK1 , PTPN11 , RET, RHEB, RHOA, RIT1 , ROS1 , SMAD4, SMO, SRC, STK1 1 , TERT, VHL.
[0055] In some cases, the one or more regions in the panel can comprise one or more loci from one or a plurality of genes, including one or more of AKT1 , ALK, APC, ATM, BRAF, CTNNB1 , EGFR, ERBB2, ESR1 , FGFR2, GATA3, GNAS, IDH1 , IDH2, KIT, KRAS, MET, NRAS, PDGFRA, PIK3CA, PTEN, RB1 , SMAD4, STK11 , and TP53.
[0056] In some cases, the one or more regions in a panel for colorectal cancer can comprise one or more loci from one or a plurality of genes, including one of, two of, three of, four of, or five of TP53, APC, BRAF, KRAS, and NRAS. In some cases, the one or more regions in a panel for ovarian cancer can comprise one or more loci from one or a plurality of genes, including TP53. In some cases, the one or more regions in a panel for pancreatic cancer can comprise one or more loci from one or a plurality of genes, including one or both of TP53 and KRAS. In some cases, the one or more regions in a panel for lung adenocarcinoma can comprise one or more loci from one or a plurality ofgenes, including one of, two of, three of, four of, five of, six of, seven of, or eight of TP53, BRAF, KRAS, EGFR, ERBB2, MET, STK11 , and ALK. In some cases, the one or more regions in a panel for lung squamous cell carcinoma can comprise one or more loci from one or a plurality of genes, including one of, two of, three of, four of, or five of TP53, BRAF, KRAS, MET, and ALK. In some cases, the one or more regions in a panel for breast cancer can comprise one or more loci from one or a plurality of genes, including one of, two of, three of, or four of TP53, GATA3, PIK3CA, and ESR1 . In some cases, one or more regions in a panel can comprise one or more loci from a combination of any of the above genes, for example, to detect a combination of cancer types. In some cases, one or more regions in a panel can comprise one or more loci from each of the preceding genes, for example, in a pan-cancer panel.
[0057] In some cases, the one or more regions in a panel for lung cancer can comprise one or more loci from a plurality of genes, including one of, two of, three of, four of, five of, six of, seven of, eight of, nine of, 10 of, 11 of, 12 of, 13 of, 14 of, 15 of, 16 of, 17 of, 18 of, 19 of, or 20 of EGFR, KRAS, TP53, CDKN2A, STK11 , BRAF, PIK3CA, RB1 , ERBB2, PTEN, NFE2L2, MET, CTNNB1 , NRAS, MUC16, NF1 , BAI3, SMARCA4, ATM, NTRK3, and ERBB4. Such a panel also may include, or have substituted for any or all of the above, any or all of an EGFR Exon 19 deletion, EGFR L858R, EGFR C797S, EGFR T790M, EGFR S645C, ARAF S214C and S214F, ERBB2 S418T, MET exon 14 skipping, SNVs and indels. Many of these genes may be clinically actionable, such that an observed anomaly in MAF (e.g., significantly higher or lower than in normal control subjects) may be indicative of a clinical state relevant to lung cancer, such as diagnosis, prognosis, risk stratification, treatment selection, tumor resistance to treatment, tumor burden, etc. Such a lung cancer targeted panel may comprise a relatively small number of these lung cancer associated genes.
[0058] In some cases, the one or more regions in a panel for breast cancer can comprise one or more loci from a plurality of genes, including any one of, or any combination of, ACVRL1 , AFF2, AGMO, AGTR2, AHNAK, AHNAK2, AKAP9, AKT1 , AKT2, ALK, APC, ARID1 A, ARID1 B, ARID2, ARID5B, ASXL1 , ASXL2, ATR, BAP1 , BCAS3, BIRC6, BRAF, BRCA1 , BRCA2, BRIP1 , CACNA2D3, CASP8, CBFB, CCND3, CDH1 , CDKN1 B, CDKN2A, CHD1 , CHEK2, CLK3, CLRN2, COL12A1 , COL22A1 , COL6A3, CTCF, CTNNA1 , CTNNA3, DCAF4L2, DNAH11 , DNAH2, DNAH5, DTWD2, EGFR, EP300, ERBB2, ERBB3,ERBB4, FAM20C, FANCA, FANCD2, FBXW7, FLT3, F0X01 , F0X03, F0XP1 , FRMD3, GATA3, GH1 , GLDC, GPR124, GPR32, GPS2, HDAC9, HERC2, HIST1 H2BC, HRAS, JAK1 , KDM3A, KDM6A, KLRG1 , KMT2C, KRAS, L1 CAM, LAMA2, LAMB3, LARGE, LDLRAP1 , LIFR, LIPI, MAGEA8, MAP2K4, MAP3K1 , MAP3K10, MAP3K13, MBL2, MEN1 , LL2, MLLT4, MTAP, MUC16, MYH9, MY01A, MY03A, NC0A3, NC0R1 , NC0R2, NDFIP1 , NEK1 , NF1 , NF2, N0TCH1 , NPNT, NR2F1 , NR3C1 , NRAS, NRG3, NT5E, OR6A2, PALLD, PBRM1 , PDE4DIP, PIK3CA, PIK3R1 , PPP2CB, PPP2R2A, PRKACG, PRKCE, PRKCQ, PRKCZ, PRKG1 , PRPS2, PRR16, PTEN, PTPN22, PTPRD, PTPRM, RASGEF1 B, RB1 , R0S1 , RPGR, RUNX1 , RYR2, SBN01 , SETD1 A, SETD2, SETDB1 , SF3B1 , SGCD, SHANK2, SIAH1 , SIK1 , SIK2, SMAD2, SMAD4, SMARCB1 , SMARCC1 , SMARCC2, SMARCD1 , SPACA1 , STAB2, STK11 , STMN2, SYNE1 , TAF1 , TAF4B, TBL1 XR1 , TBX3, TG, THADA, THSD7A, TP53, TTYH1 , UBR5, USH2A, USP28, USP9X, UTRN, and ZFP36L1 . Many of these genes may be clinically actionable, such that an observed anomaly in MAF (e.g., significantly higher or lower than in normal control subjects) may be indicative of a clinical state relevant to breast cancer, such as diagnosis, prognosis, risk stratification, treatment selection, tumor resistance to treatment, tumor burden, etc. Such a breast cancer targeted panel may comprise a relatively small number of these breast cancer associated genes.
[0059] In some cases, the one or more regions in a panel for colorectal cancer can comprise one or more loci from a plurality of genes, including one of, two of, three of, four of, five of, or six of TP53, BRAF, KRAS, APC, TGFBR, and PIK3CA. Many of these genes may be clinically actionable, such that an observed anomaly in MAF (e.g., significantly higher or lower than in normal control subjects) may be indicative of a clinical state relevant to colorectal cancer, such as diagnosis, prognosis, risk stratification, treatment selection, tumor resistance to treatment, tumor burden, etc. Such a colorectal cancer targeted panel may comprise a relatively small number of these colorectal cancer associated genes.
[0060] In some embodiments, the one or more regions in the panel comprise one or more loci from one or a plurality of genes for detecting residual cancer after surgery. This detection can be earlier than is possible for existing methods of cancer detection. In some embodiments, the one or more regions in the panel comprise one or more loci from one or a plurality of genes for detecting cancer in a high-risk patient population. For example, smokers have much higher rates of lung cancer than the generalpopulation. Moreover, smokers can develop other lung conditions that make cancer detection more difficult, such as the development of irregular nodules in the lungs. In some embodiments, the methods described herein detect cancer in high risk patients earlierthan is possible for existing methods of cancer detection.
[0061] A region may be selected for inclusion in a sequencing panel based on a number of subjects with a cancerthat have a tumor marker in that gene or region. A region may be selected for inclusion in a sequencing panel based on prevalence of subjects with a cancer and a tumor marker present in that gene. Presence of a tumor marker in a region may be indicative of a subject having cancer.
[0062] In some instances, the panel may be selected using information from one or more databases. The information regarding a cancer may be derived from cancer tumor biopsies or cfDNA assays. A database may comprise information describing a population of sequenced tumor samples. A database may comprise information about mRNA expression in tumor samples. A databased may comprise information about regulatory elements in tumor samples. The information relatingto the sequenced tumor samples may include the frequency various genetic variants and describe the genes or regions in which the genetic variants occur. The genetic variants may be tumor markers. A non-limiting example of such a database is COSMIC. COSMIC is a catalogue of somatic mutations found in various cancers. For a particular cancer, COSMIC ranks genes based on frequency of mutation. A gene may be selected for inclusion in a panel by having a high frequency of mutation within a given gene. For instance, COSMIC indicates that 33% of a population of sequenced breast cancer samples have a mutation in TP53 and 22% of a population of sampled breast cancers have a mutation in KRAS. Other ranked genes, including APC, have mutations found only in about 4% of a population of sequenced breast cancer samples. TP53 and KRAS may be included in a sequencing panel based on having relatively high frequency among sampled breast cancers (compared to APC, for example, which occurs at a frequency of about 4%). COSMIC is provided as a non-limiting example, however, any database or set of information may be used that associates a cancer with tumor marker located in a gene or genetic region. In another example, as provided by COSMIC, of 1156 biliary tract cancer samples, 380 samples (33%) carried mutations in TP53. Several other genes, such as APC, have mutations in 4-8% of all samples. Thus, TP53 may be selected forinclusion in the panel based on a relatively high frequency in a population of biliary tract cancer samples.
[0063] A gene or region may be selected for a panel where the frequency of a tumor marker is significantly greater in sampled tumor tissue or circulating tumor DNA than found in a given background population. A combination of regions may be selected for inclusion of a panel such that at least a majority of subjects having a cancer will have a tumor marker present in at least one of the regions or genes in the panel. The combination of regions may be selected based on data indicatingthat, for a particular cancer or set of cancers, a majority of subjects have one or more tumor markers in one or more of the selected regions. For example, to detect cancer 1 , a panel comprising regions A, B, C, and / or D may be selected based on data indicating that 90% of subjects with cancer 1 have a tumor marker in regions A, B, C, and / or D of the panel. Alternately, tumor markers may be shown to occur independently in two or more regions in subjects having a cancer such that, combined, a tumor marker in the two or more regions is present in a majority of a population of subjects having a cancer. For example, to detect cancer 2, a panel comprising regions X, Y, and Z may be selected based on data indicating that 90% of subjects have a tumor marker in one or more regions, and in 30% of such subjects a tumor marker is detected only in region X, while tumor markers are detected only in regions Y and / or Z for the remainder of the subjects for whom a tumor marker was detected. Tumor markers present in one or more regions previously shown to be associated with one or more cancers may be indicative of or predictive of a subject having cancer if a tumor marker is detected in one or more of those regions 50% or more of the time. Computational approaches such as models employing conditional probabilities of detecting cancer given a known cancer frequency for a set of tumor markers within one or more regions may be used to predict which regions, alone or in combination, may be predictive of cancer. Other approaches for panel selection involve the use of databases describing information from studies employing comprehensive genomic profiling of tumors with large panels and / or whole genome sequencing (WGS, RNA-seq, Chip-seq, bisulfate sequencing, ATAC-seq, and others). Information gleaned from literature may also describe pathways commonly affected and mutated in certain cancers. Panel selection may be further informed by the use of ontologies describing genetic information.
[0064] Genes included in the panel for sequencing can include the fully transcribed region, the promoter region, enhancer regions, regulatory elements, and / or downstream sequence. To further increase the likelihood of detecting tumor indicating mutations only exons may be included in the panel. The panel can comprise all exons of a selected gene, or only one or more of the exons of a selected gene. The panel may comprise of exons from each of a plurality of different genes. The panel may comprise at least one exon from each of the plurality of different genes.
[0065] In some aspects, a panel of exons from each of a plurality of different genes is selected such that a determined proportion of subjects having a cancer exhibit a genetic variant in at least one exon in the panel of exons.
[0066] At least one full exon from each different gene in a panel of genes may be sequenced. The sequenced panel may comprise exons from a plurality of genes. The panel may comprise exons from 2 to 100 different genes, from 2 to 70 genes, from 2 to 50 genes, from 2 to 30 genes, from 2 to 15 genes, or from 2 to 10 genes.
[0067] A selected panel may comprise a varying number of exons. The panel may comprise from 2 to 3000 exons. The panel may comprise from 2 to 1000 exons. The panel may comprise from 2 to 500 exons. The panel may comprise from 2 to 100 exons. The panel may comprise from 2 to 50 exons. The panel may comprise no more than 300 exons. The panel may comprise no more than 200 exons. The panel may comprise no more than 100 exons. The panel may comprise no more than 50 exons. The panel may comprise no more than 40 exons. The panel may comprise no more than 30 exons. The panel may comprise no more than 25 exons. The panel may comprise no more than 20 exons. The panel may comprise no more than 15 exons. The panel may comprise no more than 10 exons. The panel may comprise no more than 9 exons. The panel may comprise no more than 8 exons. The panel may comprise no more than 7 exons.
[0068] The panel may comprise one or more exons from a plurality of different genes. The panel may comprise one or more exons from each of a proportion of the plurality of different genes. The panel may comprise at least two exons from each of at least 25%, 50%, 75% or 90% of the different genes. The panel may comprise at least three exons from each of at least 25%, 50%, 75% or 90% of the different genes. The panel may comprise at least four exons from each of at least 25%, 50%, 75% or 90% of the different genes.
[0069] The sizes of the sequencing panel may vary. A sequencing panel may be made larger or smaller (in terms of nucleotide size) depending on several factors including, for example, the total amount of nucleotides sequenced or a number of unique molecules sequenced for a particular region in the panel. The sequencing panel can be sized 5 kb to 50 kb. The sequencing panel can be 10 kb to 30 kb in size. The sequencing panel can be 12 kb to 20 kb in size. The sequencing panel can be 12 kb to 60 kb in size. The sequencing panel can be at least 10 kb, 12 kb, 15 kb, 20 kb, 25 kb, 30 kb, 35 kb, 40 kb, 45 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb in size. The sequencing panel may be less than 100 kb, 90 kb, 80 kb, 70 kb, 60 kb, or 50 kb in size.
[0070] The panel selected for sequencing can comprise at least 1 , 5, 10, 15, 20, 25, 30, 40, 50, 60, 80, or 100 regions. In some cases, the regions in the panel are selected that the size of the regions are relatively small. In some cases, the regions in the panel have a size of about 10 kb or less, about 8 kb or less, about 6 kb or less, about 5 kb or less, about 4 kb or less, about 3 kb or less, about 2.5 kb or less, about 2 kb or less, about 1 .5 kb or less, or about 1 kb or less or less. In some cases, the regions in the panel have a size from about 0.5 kb to about 10 kb, from about 0.5 kb to about 6 kb, from about 1 kb to about 11 kb, from about 1 kb to about 15 kb, from about 1 kb to about 20 kb, from about 0.1 kb to about 10 kb, or from about 0.2 kb to about 1 kb. For example, the regions in the panel can have a size from about 0.1 kb to about 5 kb.
[0071] The panel selected herein can allow for deep sequencing that is sufficient to detect low-frequency genetic variants (e.g., in cell-free nucleic acid molecules obtained from a sample). An amount of genetic variants in a sample may be referred to in terms of the minor allele frequency for a given genetic variant. The minor allele frequency may refer to the frequency at which minor alleles (e.g., not the most common allele) occurs in a given population of nucleic acids, such as a sample. Genetic variants at a low minor allele frequency may have a relatively low frequency of presence in a sample. In some cases, the panel allows for detection of genetic variants at a minor allele frequency of at least 0.0001 %, 0.001 %, 0.005%, 0.01 %, 0.05%, 0.1 %, or 0.5%. The panel can allow for detection of genetic variants at a minor allele frequency of 0.001 % or greater. The panel can allow for detection of genetic variants at a minor allele frequency of 0.01 % or greater. The panel can allow for detection of genetic variantpresent in a sample at a frequency of as low as 0.0001 %, 0.001%, 0.005%, 0.01 %, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1 .0%. The panel can allow for detection of tumor markers present in a sample at a frequency of at least 0.0001 %, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1 .0%. The panel can allow for detection of tumor markers at a frequency in a sample as low as 1 .0%. The panel can allow for detection of tumor markers at a frequency in a sample as low as 0.75%. The panel can allow for detection of tumor markers at a frequency in a sample as low as 0.5%. The panel can allow for detection of tumor markers at a frequency in a sample as low as 0.25%. The panel can allow for detection of tumor markers at a frequency in a sample as low as 0.1 %. The panel can allow for detection of tumor markers at a frequency in a sample as low as 0.075%. The panel can allow for detection of tumor markers at a frequency in a sample as low as 0.05%. The panel can allow for detection of tumor markers at a frequency in a sample as low as 0.025%. The panel can allow for detection of tumor markers at a frequency in a sample as low as 0.01 %. The panel can allow for detection of tumor markers at a frequency in a sample as low as 0.005%. The panel can allow for detection of tumor markers at a frequency in a sample as low as 0.001 %. The panel can allow for detection of tumor markers at a frequency in a sample as low as 0.0001 %. The panel can allow for detection of tumor markers in sequenced cfDNA at a frequency in a sample as low as 1 .0% to 0.0001 %. The panel can allow for detection of tumor markers in sequenced cfDNA at a frequency in a sample as low as 0.01 % to 0.0001 %.
[0072] A genetic variant can be exhibited in a percentage of a population of subjects who have a disease (e.g., cancer). In some cases, at least 1%, 2%, 3%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% of a population having the cancer exhibit one or more genetic variants in at least one of the regions in the panel. For example, at least 80% of a population having the cancer may exhibit one or more genetic variants in at least one of the regions in the panel.
[0073] The panel can comprise one or more regions from each of one or more genes. In some cases, the panel can comprise one or more regions from each of at least 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 genes. In some cases, the panel can comprise one or more regions from each of at most 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 genes. In some cases, the panel can comprise one or more regionsfrom each of from about 1 to about 80, from 1 to about 50, from about 3 to about 40, from 5 to about 30, from 10 to about 20 different genes.
[0074] The regions in the panel can be selected so that one or more epigenetically modified regions are detected. The one or more epigenetically modified regions can be acetylated, methylated, ubiquitylated, phosphorylated, sumoylated, ribosylated, and / or citrullinated. For example, the regions in the panel can be selected so that one or more methylated regions are detected.
[0075] The regions in the panel can be selected so that they comprise sequences differentially transcribed across one or more tissues. In some cases, the regions can comprise sequences transcribed in certain tissues at a higher level compared to other tissues. For example, the regions can comprise sequences transcribed in certain tissues but not in other tissues.
[0076] The regions in the panel can comprise coding and / or non-coding sequences. For example, the regions in the panel can comprise one or more sequences in exons, introns, promoters, 3’ untranslated regions, 5’ untranslated regions, regulatory elements, transcription start sites, and / or splice sites. In some cases, the regions in the panel can comprise other non-coding sequences, including pseudogenes, repeat sequences, transposons, viral elements, and telomeres. In some cases, the regions in the panel can comprise sequences in non-coding RNA, e.g., ribosomal RNA, transfer RNA, Piwi-interacting RNA, and microRNA.
[0077] The regions in the panel can be selected to detect (diagnose) a cancer with a desired level of sensitivity (e.g., through the detection of one or more genetic variants). For example, the regions in the panel can be selected to detect the cancer (e.g., through the detection of one or more genetic variants) with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The regions in the panel can be selected to detect the cancer with a sensitivity of 100%.
[0078] The regions in the panel can be selected to detect (diagnose) a cancer with a desired level of specificity (e.g., through the detection of one or more genetic variants). For example, the regions in the panel can be selected to detect cancer (e.g., through the detection of one or more genetic variants) with a specificity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or99.9%. The regions in the panel can be selected to detect the one or more genetic variant with a specificity of 100%.
[0079] The regions in the panel can be selected to detect (diagnose) a cancer with a desired positive predictive value. Positive predictive value can be increased by increasing sensitivity (e.g., chance of an actual positive being detected) and / or specificity (e.g., chance of not mistaking an actual negative for a positive). As a nonlimiting example, regions in the panel can be selected to detect the one or more genetic variant with a positive predictive value of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The regions in the panel can be selected to detect the one or more genetic variant with a positive predictive value of 100%.
[0080] The regions in the panel can be selected to detect (diagnose) a cancer with a desired accuracy. As used herein, the term “accuracy” may refer to the ability of a test to discriminate between a disease condition (e.g., cancer) and health. Accuracy may be can be quantified using measures such as sensitivity and specificity, predictive values, likelihood ratios, the area under the ROC curve, Youden’s index and / or diagnostic odds ratio.
[0081] Accuracy may presented as a percentage, which refers to a ratio between the number of tests giving a correct result and the total number of tests performed. The regions in the panel can be selected to detect cancer with an accuracy of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The regions in the panel can be selected to detect cancer with an accuracy of 100%.
[0082] A panel may be selected such that when one or more regions or genes in the panel are removed, specificity is appreciably decreased. Removal of one region from the panel may result in a decrease in specificity of at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more.
[0083] A panel may be selected such that the addition of one or more regions or genes to the panel does not appreciably increase the specificity of the panel, e.g., does not increase the specificity by more than 1 %, 2%, 5%, 10%, 15%, or 20%.
[0084] A panel may be of a size such that when one or more regions or genes in the panel are removed, this appreciably decreases sensitivity, e.g., sensitivity is decreased by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more.
[0085] A panel may be selected such that the addition of one or more regions or genes to the panel does not appreciably increase the sensitivity of the panel, e.g., does not increase the sensitivity by more than 1 %, 2%, 5%, 10%, 15%, or 20%.
[0086] A panel may be of a size such that when one or more regions or genes in the panel are removed, accuracy is appreciably decreased, e.g., accuracy is decreased by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more.
[0087] A panel may be selected such that the addition of one or more regions or genes to the panel does not appreciably increase the accuracy of the panel, e.g., does not increase the accuracy by more than 1%, 2%, 5%, 10%, 15%, or 20%.
[0088] A panel may be of a size such that when one or more regions or genes the panel are removed, positive predictive value is appreciably decreased, e.g., positive predictive value is decreased by at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more.
[0089] A panel may be selected such that the addition of one or more regions or genes to the panel does not appreciably increase the positive predictive value of the panel, e.g., does not increase the positive predictive value by more than 1%, 2%, 5%, 10%, 15%, or 20%
[0090] A panel may be selected to be highly sensitive and detect low frequency genetic variants. For instance, a panel may be selected such that a genetic variant or tumor marker present in a sample at a frequency as low as 0.01 %, 0.05%, or 0.001% may be detected at a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Regions in a panel may be selected to detect a tumor marker present at a frequency of 1 % or less in a sample with a sensitivity of 70% or greater. A panel may be selected to detect a tumor marker at a frequency in a sample as low as 0.1% with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. A panel may be selected to detect a tumor marker at a frequency in a sample as low as 0.01 % with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. A panel may be selected to detect a tumor marker at afrequency in a sample as low as O.OO1 % with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0091] A panel may be selected to be highly specific and detect low frequency genetic variants. For instance, a panel may be selected such that a genetic variant or tumor marker present in a sample at a frequency as low as 0.01 %, 0.05%, or 0.001% may be detected at a specificity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Regions in a panel may be selected to detect a tumor marker present at a frequency of 1 % or less in a sample with a specificity of 70% or greater. A panel may be selected to detect a tumor marker at a frequency in a sample as low as 0.1% with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. A panel may be selected to detect a tumor marker at a frequency in a sample as low as 0.01% with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. A panel may be selected to detect a tumor marker at a frequency in a sample as low as 0.001 % with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0092] A panel may be selected to be highly accurate and detect low frequency genetic variants. A panel may be selected such that a genetic variant ortumor marker present in a sample at a frequency as low as 0.01%, 0.05%, or 0.001 % may be detected at an accuracy of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Regions in a panel may be selected to detect a tumor marker present at a frequency of 1 % or less in a sample with an accuracy of 70% or greater. A panel may be selected to detect a tumor marker at a frequency in a sample as low as 0.1% with an accuracy of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. A panel may be selected to detect a tumor marker at a frequency in a sample as low as 0.01% with an accuracy of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. A panel may be selected to detect a tumor marker at a frequency in a sample as low as 0.001 % with an accuracy of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0093] A panel may be selected to be highly predictive and detect low frequency genetic variants. A panel may be selected such that a genetic variant ortumor marker present in a sample at a frequency as low as 0.01%, 0.05%, or 0.001 % may have apositive predictive value of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0094] The concentration of probes or baits used in the panel may be increased (2 to 6 ng / pL) to capture more nucleic acid molecule within a sample. The concentration of probes or baits used in the panel may be at least 2 ng / pL, 3 ng / pL, 4 ng / pL, 5 ng / pL, 6 ng / pL, or greater. The concentration of probes may be about 2 ng / pL to about 3 ng / pL, about 2 ng / pL to about 4 ng / pL, about 2 ng / pL to about 5 ng / pL, about 2 ng / pL to about 6 ng / pL. The concentration of probes or baits used in the panel may be 2 ng / pL or more to 6 ng / pL or less. In some instances, this may allow for more molecules within a biological to be analyzed thereby enabling lower frequency alleles to be detected.Sequencing depth
[0095] DNA enriched from a sample of cfDNA molecules may be sequenced at a variety of read depths to detect low frequency genetic variants in a sample. For a given position, read depth may refer to a number of all reads from all molecules from a sample that map to a position, including original molecules and molecules generated by amplifying original molecules. Thus, for example, a read depth of 50,000 reads can refer to the number of reads from 5,000 molecules, with 10 reads per molecule. Original molecules mappingto a position may be unique and non-redundant (e.g., nonamplified, sample cfDNA).
[0096] To assess read depth of sample molecules at a given position, sample molecules may be tracked. Molecular trackingtechniques may comprise various techniques for labeling DNA molecules, such as barcode tagging, to uniquely identify DNA molecules in a sample. For example, one or more unique barcode sequences may be attached to one or more ends of a sample cfDNA molecule. In determining read depth at a given position, the number of distinct barcode tagged cfDNA molecules which map to that position can be indicative of the read depth for that position. In another example, both ends of sample cfDNA molecules may be tagged with one of eight barcode sequences. The read depth at a given position may be determined by quantifyingthe number of original cfDNA molecules at a given position, for instance, by collapsing reads that are redundant from amplification and identifying unique molecules based on the barcode tags and endogenous sequence information.
[0097] The DNA may be sequenced to a read depth of at least 3,000 reads per base, at least 4,000 reads per base, at least 5,000 reads per base, at least 6,000 reads per base, at least 7,000 reads per base, at least 8,000 reads per base, at least 9,000 reads per base, at least 10,000 reads per base, at least 15,000 reads per base, at least 20,000 reads per base, at least 25,000 reads per base, at least 30,000 reads per base, at least 40,000 reads per base, at least 50,000 reads per base, at least 60,000 reads per base, at least 70,000 reads per base, at least 80,000 reads per base, at least 90,000 reads per base, at least 100,000 reads per base, at least 110,000 reads per base, at least 120,000 reads per base, at least 130,000 reads per base, at least 140,000 reads per base, at least 150,000 reads per base, at least 160,000 reads per base, at least 170,000 reads per base, at least 180,000 reads per base, at least 190,000 reads per base, at least 200,000 reads per base, at least 250,000 reads per base, at least 500,000 reads per base, at least 1 ,000,000 reads per base, or at least 2,000,000 reads per base. The DNA may be sequenced to a read depth of about 3,000 reads per base, about 4,000 reads per base, about 5,000 reads per base, about 6,000 reads per base, about 7,000 reads per base, about 8,000 reads per base, about 9,000 reads per base, about 10,000 reads per base, about 15,000 reads per base, about 20,000 reads per base, about 25,000 reads per base, about 30,000 reads per base, about 40,000 reads per base, about 50,000 reads per base, about 60,000 reads per base, about 70,000 reads per base, about 80,000 reads per base, about 90,000 reads per base, about 100,000 reads per base, about 110,000 reads per base, about 120,000 reads per base, about 130,000 reads per base, about 140,000 reads per base, about 150,000 reads per base, about 160,000 reads per base, about 170,000 reads per base, about 180,000 reads per base, about 190,000 reads per base, about 200,000 reads per base, about 250,000 reads per base, about 500,000 reads per base, about 1 ,000,000 reads per base, or about 2,000,000 reads per base. The DNA can be sequenced to a read depth from about 10,000 to about 30,000 reads per base, 10,000 to about 50,000 reads per base, 10,000 to about 5,000,000 reads per base, 50,000 to about 3,000,000 reads per base, 100,000 to about 2,000,000 reads per base, or about 500,000 to about 1 ,000,000 reads per base. In some embodiments, DNA can be sequenced to any of the above read depths on a panel size selected from: less than 70,000 bases, less than 65,000 bases, less than 60,000 bases, less than 55,000 bases, less than 50,000 bases, less than 45,000 bases,less than 40,000 bases, less than 35,000 bases, less than 30,000 bases, less than 25,000 bases, less than 20,000 bases, less than 15,000 bases, less than 10,000 bases, less than 5,000 bases, and less than 1 ,000 bases. For example, the total number of reads for a panel can be as low as 600,000 (3,000 reads per base for 1 ,000 bases) and as high as 1 .4 x 1011 (2,000,000 reads per base for 70,000 bases). In some embodiments, DNA can be sequenced to any of the above read depths on a panel size selected from: 5,000 bases to 70,000 bases, 5,000 bases to 60,000 bases, 10,000 bases to 70,000 bases, or 10,000 bases to 70,000 bases.
[0098] Read coverage can include reads from one or both strands of a nucleic acid molecule. For example, read coverage may include reads from both strands of at least 5,000, at least 10,000, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 35,000, at least 40,000, at least 45,000, or at least 50,000 DNA molecules from the sample mappingto each nucleotide in the of the panel.
[0099] A panel may be selected to optimize for a desired read depth given a fixed amount of base reads.Tagging[000100] [In some embodiments of the present disclosure, a nucleic acid library is prepared prior to sequencing. For example, individual polynucleotide fragments in a genomic nucleic acid sample (e.g., genomic DNA sample) can be uniquely identified by tagging with non-unique identifiers, e.g., non-uniquely taggingthe individual polynucleotide fragments. In some embodiments, nucleic acid molecules are non- uniquely tagged with respect to one another.[000101] Polynucleotides disclosed herein can be tagged. For example, doublestranded polynucleotides can be tagged with duplex tags, tags that differently label the complementary strands (i.e., the “Watson” and “Crick” strands) of a double-stranded molecule. In some cases, the duplex tags are polynucleotides having complementary and non-complementary portions.[000102] Tags can be any types of molecules attached to a polynucleotide, including, but not limited to, nucleic acids, chemical compounds, florescent probes, or radioactive probes. Tags can also be oligonucleotides (e.g., DNA or RNA). Tags can comprise known sequences, unknown sequences, or both. A tag can comprise randomsequences, pre-determined sequences, or both. A tag can be double-stranded or single-stranded. A double-stranded tag can be a duplex tag. A double-stranded tag can comprise two complementary strands. Alternatively, a double-stranded tag can comprise a hybridized portion and a non-hybridized portion. The double-stranded tag can be Y-shaped, e.g., the hybridized portion is at one end of the tag and the nonhybridized portion is at the opposite end of the tag. One such example is the “Y adapters” used in Illumina sequencing. Other examples include hairpin shaped adapters or bubble shaped adapters. Bubble shaped adapters have non- complementary sequences flanked on both sides by complementary sequences. In some embodiments, a Y-shaped adaptor comprises a barcode 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , or 32 nucleotides in length. In some combinations, this can be combined with blunt end repair and ligation.[000103] The number of different tags may be greater than an estimated or predetermined number of molecules in the sample. For example, for unique tagging, at least two times as many different tags may be used as the estimated or predetermined number of molecules in the sample.[000104] The number of different identifying tags used to tag molecules in a collection can range, for example, between any of 2, 3, 4, 5, 6, 7, 8,, 9, 10, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, or 49 at the low end of the range, and any of 50, 100, 500, 1000, 5000 and 10,000 and 100,000 at the high end of the range. The number of identifying tags used to tag molecules in a collection can be at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60 or more. So, for example, a collection of from 100 billion to 1 trillion molecules can be tagged with from 4 to 100 or 10 to 50,000 different identifying tags. A collection of from 100 billion to 1 trillion molecules may be tagged with from 8 to 10,000 different identifying tags. A collection of from 100 billion to 1 trillion molecules may be tagged with from 16 to 10,000 different identifying tags. A collection of from 100 billion to 1 trillion molecules may be tagged with from 16 to 5,000 different identifying tags. A collection of from 100 billion to 1 trillion molecules may be tagged with from 16 to 1 ,000 different identifying tags.[000105] A collection of molecules can be considered to be “non-uniquely tagged” if there are more molecules in the collection than tags (including tagging combinations). A collection of molecules can be considered to be non-uniquely tagged if each of at least 1 %, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, or at least or about 50% of the molecules in the collection bears an identifying tag that is shared by at least one other molecule in the collection (“non-unique tag” or “non-unique identifier”). An identifier can comprise a single barcode or a combination of two barcodes. The combination of two barcodes, e.g., one attached to each end of a molecule, function together to serve as an “identifier” or“tag”. A population of nucleic acid molecules can be non-uniquely tagged by tagging the nucleic acid molecules with fewer tags than the total number of nucleic acid molecules in the population. For a non-uniquely tagged population, no more than 1 %, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the molecules may be uniquely tagged. In some embodiments, nucleic acid molecules are identified by a combination of non-unique tags and the start and stop positions or sequences from sequence reads. In some embodiments, the number of nucleic acid molecules being sequenced is less than or equal to the number of combinations of identifiers and start and stop positions or sequences.[000106] In some instances, the tags herein comprise molecular barcodes. Such molecular barcodes can be used to differentiate polynucleotides in a sample. Molecular barcodes can be different from one another. For example, molecular barcodes can have a difference between them that can be characterized by a predetermined edit distance or a Hamming distance. In some instances, the molecular barcodes herein have a minimum edit distance of 1 , 2, 3, 4, 5, 6, 7, 8, 9, or 10. To further improve efficiency of conversion (e.g., tagging) of untagged molecular to tagged molecules, one utilizes short tags. For example, a library adapter tag can be up to 65, 60, 55, 50, 45, 40, or 35 nucleotide bases in length. A collection of such short library barcodes can include a number of different molecular barcodes, e.g., at least 2, 4, 6, 8, 10, 12, 14, 16, 18 or 20 different barcodes with a minimum edit distance of 1 , 2, 3 or more.[000107] Thus, a collection of molecules can include one or more tags. In some instances, some molecules in a collection can include an identifying tag (“identifier”)such as a molecular barcode that is not shared by any other molecule in the collection. For example, in some instances of a collection of molecules, 100% or at least 50%, 60%, 70%, 80%, 90%, 95%, 97%, 98%, or 99% of the molecules in the collection can include an identifier or molecular barcode that is not shared by any other molecule in the collection. As used herein, a collection of molecules is considered to be “uniquely tagged” if each of at least 95% of the molecules in the collection bears an identifier that is not shared by any other molecule in the collection (“unique tag” or“unique identifier”). In some embodiments, nucleic acid molecules are uniquely tagged with respect to one another. A collection of molecules is considered to be “non-uniquely tagged” if each of at least 1 %, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the molecules in the collection bears an identifying tag or molecular barcode that is shared by at least one other molecule in the collection (“non-unique tag” or “nonunique identifier”). In some embodiments, nucleic acid molecules are non-uniquely tagged with respect to one another. Accordingly, in a non-uniquely tagged population no more than 1 % of the molecules are uniquely tagged. For example, in a non-uniquely tagged population, no more than 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the molecules can be uniquely tagged.[000108] A number of different tags can be used based on the estimated number of molecules in a sample. In some tagging methods, the number of different tags can be at least the same as the estimated number of molecules in the sample. In other tagging methods, the number of different tags can be at least two, three, four, five, six, seven, eight, nine, ten, one hundred or one thousand times as many as the estimated number of molecules in the sample. In unique tagging, at least two times (or more) as many different tags can be used as the estimated number of molecules in the sample.[000109] The polynucleotides fragments (prior to tagging) can comprise sequences of any length. For example, polynucleotide fragments (prior to tagging) can comprise at least 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000 or more nucleotides in length. The polynucleotide fragment can be about the average length of cell-free DNA. For example, the polynucleotide fragments can comprise about 160bases in length. The polynucleotide fragment can also be fragmented from a larger fragment into smaller fragments about 160 bases in length.[000110] Improvements in sequencing can be achieved as long as at least some of the duplicate or cognate polynucleotides bear unique identifiers with respect to each other, that is, bear different tags. However, in certain embodiments, the number of tags used is selected so that there is at least a 95% chance that all duplicate molecules starting at any one position bear unique identifiers. For example, in a sample comprising about 10,000 haploid human genome equivalents of fragmented genomic DNA, e.g., cfDNA, z is expected to be between 2 and 8. Such a population can be tagged with between about 10 and 100 different identifiers, for example, about 2 identifiers, about 4 identifiers, about 9 identifiers, about 16 identifiers, about 25 identifiers, about 36 different identifiers, about 49 different identifiers, about 64 different identifiers, about 81 different identifiers, or about 100 different identifiers.[000111] Nucleic acid barcodes having identifiable sequences, including molecular barcodes, can be used for tagging. For example, a plurality of DNA barcodes can comprise various numbers of sequences of nucleotides. A plurality of DNA barcodes having , 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 100, 200, 500 or more identifiable sequences of nucleotides can be used. When attached to only one end of a polynucleotide, the plurality of DNA barcodes can produce 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 100, 200, 500 or more different identifiers. Alternatively, when attached to both ends of a polynucleotide, the plurality DNA barcodes can produce 4, 9, 16, 25, 36, 49, 64, 81 , 100, 121 , 144, 169, 196, 225, 256, 289, 324, 361 , 400, 2500, 10,000, 40,000, 250,000 or more different identifiers (which is the 2 of when the DNA barcode is attached to only 1 end of a polynucleotide). In one example, a plurality of DNA barcodes having 6, 7, 8, 9 or 10 identifiable sequences of nucleotides can be used. When attached to both ends of a polynucleotide, they produce 36, 49, 64, 81 or 100 possible different identifiers, respectively. In a particular example, the plurality of DNA barcodes can comprise 8 identifiable sequences of nucleotides. When attached to only one end of a polynucleotide, the plurality of DNA barcodes can produce 8 different identifiers. Alternatively, when attached to both ends of a polynucleotide, the plurality of DNAbarcodes can produce 64 different identifiers. Samples tagged in such a way can be those with a range of about 10 ng to any of about 200 ng, about 1 pg, about 10 pg of fragmented polynucleotides, e.g., genomic DNA, e.g., cfDNA.[000112] A polynucleotide can be uniquely identified in various ways. A polynucleotide can be uniquely identified by a unique barcode. For example, any two polynucleotides in a sample are attached two different barcodes. A barcode may be a DNA barcode or an RNA barcode. For example, a barcode may be a DNA barcode.[000113] Alternatively, a polynucleotide can be uniquely identified by the combination of a barcode and one or more endogenous sequences of the polynucleotide. The barcode may be a non-unique tag or a unique tag. In some cases, the barcode is a non-unique tag. For example, any two polynucleotides in a sample can be attached to barcodes comprising the same barcode, but the two polynucleotides can still be identified by different endogenous sequences. The two polynucleotides may be identified by information in the different endogenous sequences. Such information includes the sequence of the endogenous sequences or a portion thereof, the length of the endogenous sequences, the location of the endogenous sequences, one or more epigenetic modification of the endogenous sequences, or any other feature of the endogenous sequences. In some embodiments, polynucleotides can be identified by an identifier (comprising one barcode or comprisingtwo barcodes) in combination with start and stop sequences from the sequence read.[000114] Polynucleotides in a sample can be tagged with a sufficient number of different tags so that there is a high probability (e.g., at least 90%, at least 95%, at least 98%, at least 99%, at least 99.9% or at least 99.99%) that all polynucleotides mapping to a particular genomic region bear a different identifying tag (molecules within the region are substantially uniquely tagged). The genomic region to which the polynucleotides map can be, for example, (1 ) the entire panel of genes being sequenced, (2) some portion of that panel, such as mapping within a single gene, exon or intron, (3) a single nucleotide coordinate (e.g., at least one nucleotide in the polynucleotide maps to the coordinate, for example, the start position, stop position, mid-point or anywhere between) or (4) a particular pair of start / stop (begin / end) nucleotide coordinates. The number of different identifiers (tag counts) necessary to substantially uniquely tag polynucleotides is a function of how many originalpolynucleotide molecules in the sample that map to the region. This, in turn, is a function of several factors. One factor is the total number of haploid genome equivalents included in the assay. Another factor is the average size of the polynucleotide molecules. Another factor is the distribution of the molecules across the region. This, in turn, can be a function of the cleavage pattern - one may expect cleavage to occur primarily between nucleosomes so that more polynucleotides map across a nucleosome location than between nucleosomes. Another factor is the distribution of barcodes in the pool and the ligation efficiency of individual barcodes, potentially causing differences in affective concentration of one barcode versus another. Another factor is the size of the region within which the molecules to be uniquely tagged are confined (e.g., same start / stop or same exon).[000115] The identifier can be a single barcode attached to one end of a molecule, or two barcodes, each attached to different ends of the molecule. Attaching barcodes independently to both ends of a molecule increases by square the number of possible identifiers. In this case the number of different barcodes is selected such that the combination of barcodes on each end of a particular polynucleotide has a high probability of being unique with respect to other polynucleotides mapping to the same selected genomic region.[000116] In certain embodiments, the number of different identifiers or barcode combinations (tag count) used can be at least any of 64, 100, 400, 900, 1400, 2500, 5625, 10,000, 14,400, 22,500 or 40,000 and no more than any of 90,000, 40,000, 22,500, 14,400 or 10,000. For example, the number of identifiers or barcode combinations can be between 64 and 90,000, between 400 and 22,500, 400 and 14,400[000117] Types of cancer that a subject may have been diagnosed with include, but are not limited to: Acute lymphoblastic leukemia (ALL), Acute myeloid leukemia, Adrenocortical carcinoma, adult acute Myeloid leukemia, adult carcinoma of unknown primary site, adult malignant Mesothelioma, AIDS-related cancers, AIDS-related lymphoma, Anal cancer, Appendix cancer, Astrocytoma, childhood cerebellar or cerebral, Basal-cell carcinoma, Bile duct cancer, Bladder cancer, Bone tumor, osteosarcoma / malignant fibrous histiocytoma, Brain cancer, Brainstem glioma, Breast cancer, Bronchial adenomas / carcinoids, Burkitt Lymphoma, Carcinoid tumor, Carcinoma of unknown primary, Central nervous system lymphoma, cerebellarastrocytoma, cerebral astrocytoma / malignant glioma, Cervical cancer, childhood acute Myeloid leukemia, childhood cancer of unknown primary site, Childhood cancers, childhood cerebral astrocytoma, childhood Mesothelioma, Chondrosarcoma, Chronic lymphocytic leukemia, Chronic myelogenous leukemia, Chronic myeloproliferative disorders, Colon cancer, Cutaneous T-cell lymphoma, Desmoplastic small round cell tumor, Endometrial cancer, endometrial Uterine cancer, Ependymoma, Epitheliod Hemangioendothelioma (EHE), Esophageal cancer, Ewing family of tumors Sarcoma, Ewing’s sarcoma in the Ewing family of tumors, Extracranial germ cell tumor, Extragonadal germ cell tumor, Extrahepatic bile duct cancer, Eye cancer, intraocular melanoma, Gallbladder cancer, Gastric (stomach) cancer, Gastric carcinoid, Gastrointestinal carcinoid tumor, Gastrointestinal stromal tumor (GIST), Gestational trophoblastic tumor, Glioma of the brain stem, Glioma, Hairy cell leukemia, Head and neck cancer, Heart cancer, Hepatocellular (liver) cancer, Hodgkin lymphoma, Hypopharyngeal cancer, Hypothalamic and visual pathway glioma, Islet cell carcinoma (endocrine pancreas), Kaposi sarcoma, Kidney cancer (renal cell cancer), Laryngeal cancer, Leukaemia, acute lymphoblastic (also called acute lymphocytic leukaemia), Leukaemia, acute myeloid (also called acute myelogenous leukemia), Leukaemia, chronic lymphocytic (also called chronic lymphocytic leukemia), Leukaemias, Leukemia, chronic myelogenous (also called chronic myeloid leukemia), Leukemia, hairy cell, Lip and oral cavity cancer, Liposarcoma, Liver cancer (primary), Lung cancer, non-small cell, Lung cancer, small cell, Lymphoma (AIDS-related), Lymphomas, Macroglobulinemia, Waldenstrom, Male breast cancer, Malignant fibrous histiocytoma of bone / osteosarcoma, medulloblastoma, Melanoma, Merkel cell cancer, Metastatic squamous neck cancer with occult primary, Mouth cancer, Multiple endocrine neoplasia syndrome, childhood, multiple Myeloma (cancer of the bone-marrow), Multiple myeloma / plasma cell neoplasm, Mycosis fungoides, Myelodysplastic syndromes, Myelodysplastic / myeloproliferative diseases, Myelogenous leukemia, chronic, Myxoma, Nasal cavity and paranasal sinus cancer, Nasopharyngeal carcinoma, Neuroblastoma, Non-Hodgkin Lymphomas, Non-small cell lung cancer, Oligodendroglioma, Oral cancer, Oropharyngeal cancer, Osteosarcoma / malignant fibrous histiocytoma of bone, Ovarian cancer, Ovarian epithelial cancer (surface epithelial-stromal tumor), Ovarian germ cell tumor, Ovarian low malignant potentialtumor, Pancreatic cancer, Pancreatic cancer, islet cell, Paranasal sinus and nasal cavity cancer, Parathyroid cancer, Penile cancer, Pharyngeal cancer, Pheochromocytoma, Pineal astrocytoma, Pineal germinoma, Pineoblastoma and supratentorial primitive neuroectodermal tumors, Pituitary adenoma, Plasma cell neoplasia / Multiple myeloma, Pleuropulmonary blastoma, Primary central nervous system lymphoma, Prostate cancer, Rectal cancer, Renal cell carcinoma (kidney cancer), Renal pelvis and ureter transitional cell cancer, Retinoblastoma, Rhabdomyosarcoma, Salivary gland cancer, Sezary syndrome, Skin cancer (melanoma), Skin cancer (non-melanoma), Skin carcinoma, Merkel cell, Small cell lung cancer, Small intestine cancer, soft tissue Sarcoma, Squamous cell carcinoma , Squamous neck cancer with occult primary, metastatic, Stomach cancer, Supratentorial primitive neuroectodermal tumor, T-Cell lymphoma, cutaneous , Testicular cancer, Throat cancer, Thymoma and thymic carcinoma, Thymoma, Thyroid cancer, Transitional cell cancer of the renal pelvis and ureter, Ureter and renal pelvis, transitional cell cancer, Urethral cancer, Uterine sarcoma, Vaginal cancer, visual pathway and hypothalamic glioma, Visual pathway and hypothalamic glioma, childhood, Vulvar cancer, Waldenstrom macroglobulinemia, and Wilms tumor (kidney cancer).[000118] The endogenous sequence can be on an end of a polynucleotide. For example, the endogenous sequence can be adjacent (e.g., base in between) to the attached barcode. In some instances, the endogenous sequence can be at least 2, 4, 6, 8, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 bases in length. The endogenous sequence can be a terminal sequence of the fragment / polynucleotides to be analyzed. The endogenous sequence may be the length of the sequence. For example, a plurality of barcodes comprising 8 different barcodes can be attached to both ends of each polynucleotide in a sample. Each polynucleotide in the sample can be identified by the combination of the barcodes and about 10 base pair endogenous sequence on an end of the polynucleotide. Without being bound by theory, the endogenous sequence of a polynucleotide can also be the entire polynucleotide sequence.[000119] Also disclosed herein are compositions of tagged polynucleotides. The tagged polynucleotide can be single-stranded. Alternatively, the tagged polynucleotide can be double-stranded (e.g., duplex-tagged polynucleotides). Accordingly, this disclosure also provides compositions of duplex-tagged polynucleotides. Thepolynucleotides can comprise any types of nucleic acids (DNA and / or RNA). The polynucleotides comprise any types of DNA disclosed herein. For example, the polynucleotides can comprise DNA, e.g., fragmented DNA or cfDNA. A set of polynucleotides in the composition that map to a mappable base position in a genome can be non-uniquely tagged, that is, the number of different identifiers can be at least 2 and fewer than the number of polynucleotides that map to the mappable base position. The number of different identifiers can also be at least 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25 and fewer than the number of polynucleotides that map to the mappable base position.[000120] In some instances, as a composition goes from about 1 ngto about 10 pg or higher, a larger set of different molecular barcodes can be used. For example, between 5 and 100 different library adaptors can be used to tag polynucleotides in a cfDNA sample.[000121] The molecular barcodes can be assigned to any types of polynucleotides disclosed in this disclosure. For example, the molecular barcodes can be assigned to cell-free polynucleotides (e.g., cfDNA). Often, an identifier disclosed herein can be a barcode oligonucleotide that is used to tagthe polynucleotide. The barcode identifier may be a nucleic acid oligonucleotide (e.g., a DNA oligonucleotide). The barcode identifier can be single-stranded. Alternatively, the barcode identifier can be doublestranded. The barcode identifier can be attached to polynucleotides using any method disclosed herein. For example, the barcode identifier can be attached to the polynucleotide by ligation using an enzyme. The barcode identifier can also be incorporated into the polynucleotide through PCR. In other cases, the reaction may comprise addition of a metal isotope, either directly to the analyte or by a probe labeled with the isotope. Generally, assignment of unique or non-unique identifiers or molecular barcodes in reactions of this disclosure may follow methods and systems described by, for example, U.S. patent applications 2001 / 0053519, 2003 / 0152490, 2011 / 0160078 and U.S. Pat. No. 6,582,908, each of which is entirely incorporated herein by reference.[000122] Identifiers or molecular barcodes used herein may be completely endogenous whereby circular ligation of individual fragments may be performed followed by random shearing or targeted amplification. In this case, the combination ofa new start and stop point of the molecule and the original intramolecular ligation point can form a specific identifier.[000123] Identifiers or molecular barcodes used herein can comprise any types of oligonucleotides. In some cases, identifiers may be predetermined, random, or semirandom sequence oligonucleotides. Identifiers can be barcodes. For example, a plurality of barcodes may be used such that barcodes are not necessarily unique to one another in the plurality. Alternatively, a plurality of barcodes may be used such that each barcode is unique to any other barcode in the plurality. The barcodes can comprise specific sequences (e.g., predetermined sequences) that can be individually tracked. Further, barcodes may be attached (e.g., by ligation) to individual molecules such that the combination of the barcode and the sequence it may be ligated to creates a specific sequence that may be individually tracked. As described herein, detection of barcodes in combination with sequence data of beginning (start) and / or end (stop) portions of sequence reads can allow assignment of a unique identity to a particular molecule. The length or number of base pairs of an individual sequence read may also be used to assign a unique identity to such a molecule. As described herein, fragments from a single strand of nucleic acid having been assigned a unique identity, may thereby permit subsequent identification of fragments from the parent strand. In this way the polynucleotides in the sample can be uniquely or substantially uniquely tagged. A duplex tag can include a degenerate or semi-degenerate nucleotide sequence, e.g., a random degenerate sequence. The nucleotide sequence can comprise any number of nucleotides. For example, the nucleotide sequence can comprise 1 (if using a nonnatural nucleotide), 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleotides. In a particular example, the sequence can comprise 7 nucleotides. In another example, the sequence can comprise 8 nucleotides. The sequence can also comprise 9 nucleotides. The sequence can comprise 10 nucleotides.[000124] A barcode can comprise contiguous or non-contiguous sequences. A barcode that comprises at least 1 , 2, 3, 4, 5 or more nucleotides is a contiguous sequence or non-contiguous sequence, if the 4 nucleotides are uninterrupted by any other nucleotide. For example, if a barcode comprises the sequence TTGC, a barcode iscontiguous if the barcode is TTGC. On the other hand, a barcode is non-contiguous if the barcode is TTXGC, where X is a nucleic acid base.[000125] An identifier or molecular barcode can have an n-mer sequence which may be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleotides in length. A tag herein can comprise any range of nucleotides in length. For example, the sequence can be between 2 to 100, 10 to 90, 20 to 80, 30 to 70, 40 to 60, or about 50 nucleotides in length. A population of barcodes can comprise barcodes of the same length or of different lengths.[000126] The tag can comprise a double-stranded fixed reference sequence downstream of the identifier or molecular barcode. Alternatively, the tag can comprise a double-stranded fixed reference sequence upstream or downstream of the identifier or molecular barcode. Each strand of a double-stranded fixed reference sequence can be, for example, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50 nucleotides in length.[000127] Tagging disclosed herein can be performed using any method. A polynucleotide can be tagged with an adaptor by hybridization. For example, the adaptor can have a nucleotide sequence that is complementary to at least a portion of a sequence of the polynucleotide. As an alternative, a polynucleotide can be tagged with an adaptor by ligation.[000128] The barcodes or tags can be attached using a variety of techniques. Attachment can be performed by methods including, for example, ligation (blunt-end or sticky-end) or annealing-optimized molecular-inversion probes. For example, tagging can comprise using one or more enzymes. The enzyme can be a ligase. The ligase can be a DNA ligase. For example, the DNA ligase can be a T4 DNA ligase, E. coli DNA ligase, and / or mammalian ligase. The mammalian ligase can be DNA ligase I, DNA ligase III, or DNA ligase IV. The ligase can also be a thermostable ligase. Tags can be ligated to a blunt-end of a polynucleotide (blunt-end ligation). Alternatively, tags can be ligated to a sticky end of a polynucleotide (sticky-end ligation). Efficiency of ligation can be increased by optimizing various conditions. Efficiency of ligation can be increased by optimizing the reaction time of ligation. For example, the reaction time of ligation can beless than 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, or 20 hours. In a particular example, reaction time of ligation is less than 20 hours. Efficiency of ligation can be increased by optimizing the ligase concentration in the reaction. For example, the ligase concentration can be at least 10, 50, 100, 150, 200, 250, 300, 400, 500, or 600 units / microliter. Efficiency can also be optimized by adding or varying the concentration of an enzyme suitable for ligation, enzyme cofactors or other additives, and / or optimizing a temperature of a solution having the enzyme. Efficiency can also be optimized by varying the addition order of various components of the reaction. The end of tag sequence can comprise dinucleotide to increase ligation efficiency. When the tag comprises a non-complementary portion (e.g., Y-shaped adaptor), the sequence on the complementary portion of the tag adaptor can comprise one or more selected sequences that promote ligation efficiency. Such sequences are located at the terminal end of the tag. Such sequences can comprise 1 , 2, 3, 4, 5, or 6 terminal bases. Reaction solution with high viscosity (e.g., a low Reynolds number) can also be used to increase ligation efficiency. For example, solution can have a Reynolds number less than 3000, 2000, 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, 25, or 10. It is also contemplated that roughly unified distribution of fragments (e.g., tight standard deviation) can be used to increase ligation efficiency. For example, the variation in fragment sizes can vary by less than 20%, 15%, 10%, 5%, or 1 %. Tagging can also comprise primer extension, for example, by polymerase chain reaction (PCR). Tagging can also comprise any of ligation-based PCR, multiplex PCR, single strand ligation, or single strand circularization. Efficiency of tagging (e.g., by ligation) can be increased to an efficiency of tagging molecules (conversion efficiency) of at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 98%.[000129] A ligation reaction may be performed in which parent polynucleotides in a sample are admixed with a reaction mixture comprisingy different barcode oligonucleotides, wherein y = a square root of n. The ligation can result in the random attachment of barcode oligonucleotides to parent polynucleotides in the sample. The reaction mixture can then be incubated under ligation conditions sufficient to effect ligation of barcode oligonucleotides to parent polynucleotides of the sample. In some embodiments, random barcodes selected from the y different barcode oligonucleotidesare ligated to both ends of parent polynucleotides. Random ligation of the y barcodes to one or both ends of the parent polynucleotides can result in production of y2 unique identifiers. For example, a sample comprising about 10,000 haploid human genome equivalents of cfDNA can be tagged with about 36 unique identifiers. The unique identifiers can comprise six unique DNA barcodes. Ligation of 6 unique barcodes to both ends of a polynucleotide can result in 36 possible unique identifiers produced.[000130] In some embodiments, a sample comprising about 10,000 haploid human genome equivalents of DNA is tagged with a number of unique identifiers produced by ligation of a set of unique barcodes to both ends of parent polynucleotides. For example, 64 unique identifiers can be produced by ligation of 8 unique barcodes to both ends of parent polynucleotides. Likewise, 100 unique identifiers can be produced by ligation of 10 unique barcodes to both ends of parent polynucleotides, 225 unique identifiers can be produced by ligation of 15 unique barcodes to both ends of parent polynucleotides, 400 unique identifiers can be produced by ligation of 20 unique barcodes to both ends of parent polynucleotides, 625 unique identifiers can be produced by ligation of 25 unique barcodes to both ends of parent polynucleotides, 900 unique identifiers can be produced by ligation of 30 unique barcodes to both ends of parent polynucleotides, 1225 unique identifiers can be produced by ligation of 35 unique barcodes to both ends of parent polynucleotides, 1600 unique identifiers can be produced by ligation of 40 unique barcodes to both ends of parent polynucleotides, 2025 unique identifiers can be produced by ligation of 45 unique barcodes to both ends of parent polynucleotides, and 2500 unique identifiers can be produced by ligation of 50 unique barcodes to both ends of parent polynucleotides. The ligation efficiency of the reaction can be over 10%, over 20%, over 30%, over 40%, over 50%, over 60%, over 70%, over 80%, or over 90%. The ligation conditions can comprise use of bi-directional adaptors that can bind either end of the fragment and still be amplifiable. The ligation conditions can comprise sticky-end ligation adapters each having an overhang of at least one nucleotide base. In some instances, the ligation conditions can comprise adapters having different base overhangs to increase ligation efficiency. As a non-limiting example and described in more detail below, the ligation conditions may comprise adapters with single-base cytosine (C) overhangs (i.e., C-tailed adaptors), single-base thymine (T) overhangs (T-tailed adaptors), single-base adenine (A) overhangs (A-tailed adaptors), and / or singlebase guanine (G) overhangs (G-tailed adaptors). The ligation conditions can comprise blunt end ligation, as opposed to tailing. The ligation conditions can comprise careful titration of an amount of adapter and / or barcode oligonucleotides. The ligation conditions can comprise the use of over 2X, over 5X, over 10X, over 20X, over 40X, over 60X, over 80X, (e.g., ~100X) molar excess of adapter and / or barcode oligonucleotides as compared to an amount of parent polynucleotide fragments in the reaction mixture. The ligation conditions can comprise use of a T4 DNA ligase (e.g., NEBNExt Ultra Ligation Module). In an example, 18 microliters of ligase master mix is used with 90 microliter ligation (18 parts of the 90) and ligation enhancer. Accordingly, tagging parent polynucleotides with n unique identifiers can comprise use of a numbery different barcodes, wherein y= a square root of n. Samples tagged in such a way can be those with a range of about 10 ng to any of about 100 ng, about 200 ng, about 300 ng, about 400 ng, about 500 ng, about 1 pg, or about 10 pg of fragmented polynucleotides, e.g., genomic DNA, e.g. cfDNA. The numbery of barcodes used to identify parent polynucleotides in a sample can depend on the amount of nucleic acid in the sample.[000131] One method of increasing conversion efficiency involves using a ligase engineered for optimal reactivity on single-stranded DNA, such as a ThermoPhage single-stranded DNA (ssDNA) ligase derivative. Such ligases bypass traditional steps in library preparation of end-repair and A-tailing that can have poor efficiencies and / or accumulated losses due to intermediate cleanup steps, and allows for twice the probability that either the sense or anti-sense starting polynucleotide will be converted into an appropriately tagged polynucleotide. It also converts double-stranded polynucleotides that may possess overhangs that may not be sufficiently blunt-ended by the typical end-repair reaction. Optimal reactions conditions for this ssDNA reaction are: 1 x reaction buffer (50 millimolar (mM) MOPS (pH 7.5), 1 mM DTT, 5 mM MgCl2, 10 mM KOI). With 50 mM ATP, 25 mg / ml BSA, 2.5 mM MnCl2, 200 pmol 85 nt ssDNA oligomer and 5 U ssDNA ligase incubated at 65°C for 1 hour. Subsequent amplification using PCR can further convert the tagged single-stranded library to a double-stranded library and yield an overall conversion efficiency of well above 20%. Other methods of increasing conversion rate, e.g., to above 10%, include, for example, any of the following, alone or in combination: annealing-optimized molecular-inversion probes,blunt-end ligation with a well-controlled polynucleotide size range, selection of a high- efficiency polymerase, sticky-end ligation or an upfront multiplex amplification step with or without the use of fusion primers, optimization of end bases in a target sequence, optimization of reaction conditions (including reaction time), and the introduction of one or more steps to clean up a reaction (e.g., of unwanted nucleic acid fragments) during the ligation, and optimization of temperature of buffer conditions. Sticky end ligation may be performed using multiple-nucleotide overhangs. Sticky end ligation may be performed using single-nucleotide overhangs comprising an A, T, C, or G bases.[000132] The present disclosure also provides compositions of tagged polynucleotides. The polynucleotides can comprise fragmented DNA, e.g. cfDNA. A set of polynucleotides in the composition that map to a mappable base position in a genome can be non-uniquely tagged, that is, the number of different identifiers can be at least at least 2 and fewer than the number of polynucleotides that map to the mappable base position. A composition of between about 10 ngto about 10 pg (e.g., any of about 10 ng-1 pg, about 10 ng-100 ng, about 100 ng-10 pg, about 100 ng-1 pg, about 1 pg-10 pg) can bear between any of 2, 5, 10, 50 or 100 to any of 100, 1000, 10,000 or 100,000 different identifiers. For example, between 5 and 100 different identifiers can be used to tagthe polynucleotides in such a composition.Linking sample nucleic acid molecules to adapters[000133] Sample preparation for new generation sequencing platforms often follows a similar protocol. Samples typically contain double-stranded nucleic acid fragments with single-stranded overhangs. Such fragments can be blunt-ended and ligated to adapters directly. But such ligations also result in byproducts in which adapters or fragments form concatemers. Formation of such byproducts can be reduced by an alternative procedure in which blunt-ended fragments are A-tailed and ligated to T-tailed adapters. Commercial kits that perform end repair and tailing in a single tube are simple to use and fast and can be used with commercially available adaptors. (For example, NEBNext Ultra II (New England Biolabs, Ipswich, MA.). However, the kits are generally not optimized for A-tailing and can result in tailing withother nucleotides, such as G, T and C. The result of inefficient tailing is inefficient ligation of adapters and low complexity libraries.[000134] In certain aspects, the present disclosure provides improved methods of preparing double-stranded nucleic acids (preferably DNA) with single-stranded overhangs for amplification and subsequent analysis, particularly sequencing. These methods can be used in conjunction with early methods of cancer detection described herein or in other applications. It has been found that contacting blunt-ended doublestranded nucleic acids with Taq in the presence of all four standard nucleotide types results in non-templated directed addition of a single nucleotide to the 3’ ends of the nucleic acid such that A is added most frequently followed by G followed by C and T. Although inclusion of additional nucleic acid molecules increases the potential for off- target side reactions, it has been found that the proportion of single-G tailing is sufficiently high relative to single-A tailing that the efficiency of ligation of nucleic acid molecules in a sample to adapters can be significantly increased by including a customized mix of adapters tailed not only with T (as in prior methods) but also with C, which adapters anneal respectively to 3’ ends of DNA molecules tailed with A and G. The ligation efficiency can be increased even further by also including blunted-ended adapters (i.e., not tailed with any nucleotide) to ligate to blunt-ended nucleic acid molecules in the sample that have failed to undergo tailing with any nucleotide. Sequencing[000135] Tagged polynucleotides can be sequenced to generate sequence reads. For example, a tagged duplex polynucleotide can be sequenced. Sequence reads can be generated from only one strand of a tagged duplex polynucleotide. Alternatively, both strands of a tagged duplex polynucleotide can generate sequence reads. The two strands of the tagged duplex polynucleotide can comprise the same tags. Alternatively, the two strands of the tagged duplex polynucleotide can comprise different tags. When the two strands of the tagged duplex polynucleotide are differently tagged, sequence reads generated from one strand (e.g., a Watson strand) can be distinguished from sequence reads generated from the other strands (e.g., a Crick strand). Sequencing can involve generating multiple sequence reads for each molecule. This occurs, for example, as a result the amplification of individual polynucleotide strands during the sequencing process, e.g., by PCR.[000136] Methods disclosed herein can comprise amplifying of polynucleotides. Amplification can be performed before tagging, after tagging, or both. Polynucleotides amplification can result in the incorporation of nucleotides into a nucleic acid molecule or primer thereby forming a new nucleic acid molecule complementary to a template nucleic acid. The newly formed polynucleotide molecule and its template can be used as templates to synthesize additional polynucleotides. The polynucleotides being amplified can be any nucleic acids, for example, deoxyribonucleic acids, including genomic DNAs, cDNAs (complementary DNA), cfDNAs, and circulating tumor DNAs (ctDNAs). The polynucleotides being amplified can also be RNAs. As used herein, one amplification reaction may comprise many rounds of DNA replication. DNA amplification reactions can include, for example, polymerase chain reaction (PCR). One PCR reaction may comprise 2-100 “cycles” of denaturation, annealing, and synthesis of a DNA molecule. For example, 2-7, 5-10, 6-11 , 7-12, 8-13, 9-14, 10-15, 11 -16, 12-17, 13- 18, 14-19, or 15-20 cycles can be performed during the amplification step. The condition of the PCR can be optimized based on the GC content of the sequences, includingthe primers. Amplification primers can be chosen to select for a target sequence of interest. Primers can be designed to optimize or maximize conversion efficiency. In some embodiments, primers contain a short sequence between the primers so as to pull out a small region of interest. In some embodiments, primers target nucleosomal regions so that the primers hybridize to areas where nucleosomes are present, as opposed to areas between nucleosomes, because inter-nucleosomal areas are more highly cleaved and therefore less likely to be present as targets.[000137] In some embodiments, regions of the genome are targeted that are differentially protected by nucleosomes and other regulatory mechanisms in cancer cells, the tumor microenvironment, or immune system components (granulocytes, tumor infiltrating lymphocytes, etc.). In some embodiments, other regions are targeted that are stable and / or not differentially regulated in tumor cells. Within these regions, differences in coverage, cleavage sites, fragment length, sequence content, sequence content at fragment endpoints, or sequence content of the nearby genomic context can be used to inferthe presence or absence of a certain classification of cancer cells (e.g., EGFR mutant, KRAS mutant, ERBb2 amplified, or PD-1 expression cancers), or type of cancer (e.g., lung adenocarcinoma, breast, or colorectal cancer). Such targeting canalso enhance the sensitivity and / or specificity of the assay by enhancing coverage at certain sites or the probability of capture. These principles apply to methods of targeting including, but not limited to, ligation plus hybrid capture-based enrichment, amplification-based enrichment, rolling circle-based enrichment with sequence / genomic location specific initiation primers, and other methods. The regions that can be targeted with such methods and subsequent analysis include, but are not limited to, intronic regions, exonic regions, promoter regions, TSS regions, distant regulatory elements, enhancer regions, and super-enhancer regions and / or junctions of the preceding. These methods can also be used to inferthe tissue of origin of the tumor and / or a measure of tumor burden in combination with other techniques described herein for determining variants (e.g., germline or somatic variants) contained within the sample. For example, germline variants can determine predisposition for certain types of cancer, while somatic variants can correlate to certain types of cancer specifically based on the affected genes, pathways and percentages of the variants. This information can then be used in combination with epigenetic signatures relating to regulatory mechanisms and / or chemical modifications such as, for example, methylation, hydroxymethylation, acetylation, and / or RNA. The nucleic acid library can involve combined analysis of DNA, DNA modifications and RNA to enhance sensitivity and specificity to the detection of cancer, type of cancer, molecular pathways activated in the specific disease, tissue of origin as well as a measure that corresponds to tumor burden. Approaches for analyzing each of the above have been outlined elsewhere and can be combined for analysis of a single or multiple samples from the same patient, whereby the sample can be derived from various bodily specimens.[000138] Nucleic acid amplification techniques can be used with the assays described herein. Some amplification techniques are the PCR methodologies which can include, but are not limited to, solution PCR and in situ PCR. For example, amplification may comprise PCR-based amplification. Alternatively, amplification may comprise non-PCR-based amplification. Amplification of the template nucleic acid may comprise use of one or more polymerases. For example, the polymerase may be a DNA polymerase or an RNA polymerase. In some cases, high-fidelity amplification is performed such as with the use of high fidelity polymerase (e.g., Phusion RTM High- Fidelity DNA Polymerase) or PCR protocols. In some cases, the polymerase may be ahigh fidelity polymerase. For example, the polymerase may be KAPA HiFi DNA polymerase. The polymerase may also be Phusion DNA polymerase or an Ultra II polymerase. The polymerase may be used under reaction conditions that reduce or minimize amplification biases, e.g., due to fragment length and / or GC content.[000139] Amplification of a single strand of a polynucleotide by PCR will generate copies both of that strand and its complement. During sequencing, both the strand and its complement will generate sequence reads. However, sequence reads generated from the complement of, for example, the Watson strand, can be identified as such because they bear the complement of the portion of the duplex tag that tagged the original Watson strand. In contrast, a sequence read generated from a Crick strand or its amplification product will bearthe portion of the duplex tag that tagged the original Crick strand. In this way, a sequence read generated from an amplified product of a complement of the Watson strand can be distinguished from a complement sequence read generated from an amplification product of the Crick strand of the original molecule.[000140] Amplification, such as PCR amplification, is typically performed in rounds. Exemplary rounds of amplification include 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, or more rounds of amplification. Amplification conditions can be optimized, for example, for buffer conditions and polymerase type and conditions. The amplification also can be modified to reduce bias in the sample processing, for example, by reducing non-specific amplification bias, GC content bias, and size bias.[000141] In some embodiments, sequences can be enriched prior to sequencing. Enrichment can be performed for specific target regions or nonspecifically. In some embodiments, targeted genomic regions of interest may be enriched with capture probes (“baits”) selected for one or more bait set panels using a differential tiling and capture scheme. A differential tiling and capture scheme uses bait sets of different relative concentrations to differentially tile (e.g., at different “resolutions”) across genomic regions associated with baits, subject to a set of constraints (e.g., sequencer constraints such as sequencing load, utility of each bait, etc.), and capture them at a desired level for downstream sequencing. These targeted genomic regions of interest may include single-nucleotide variants (SNVs) and indels (i.e., insertions or deletions).The targeted genomic regions of interest may comprise backbone genomic regions of interest (“backbone regions”) or hot-spot genomic regions of interest (“hot-spot regions” or “hotspot regions” or “hot-spots” or “hotspots”). While “hotpots” can refer to particular loci associated with sequence variants, “backbone” regions can refer to larger genomic regions, each of which can have one or more potential sequence variants. For example, a backbone region can be a region containing one or more cancer-associated mutations, while a hotspot can be a locus with a particular mutation associated with recurring cancer or a locus with a particular recurring mutation associated with cancer. Both backbone and hot-spot genomic regions of interest may comprise tumor-relevant marker genes commonly included in liquid biopsy assays (e.g., BRAF, BRCA 1 / 2, EGFR, KRAS, PIK3CA, ROS1 , TP53, and others), for which one or more variants may be expected to be seen in subjects with cancer. In some embodiments, biotin-labeled beads with probes to one or more regions of interest can be used to capture target sequences, optionally followed by amplification of those regions, to enrich for the regions of interest.[000142] The amount of sequencing data that can be obtained from a sample is finite, and constrained by such factors as the quality of nucleic acid templates, number of target sequences, scarcity of specific sequences, limitations in sequencing techniques, and practical considerations such as time and expense. Thus, a “read budget” is a way to conceptualize the amount of genetic information that can be extracted from a sample. A per-sample read budget can be selected that identifies the total number of base reads to be allocated to a test sample comprising a predetermined amount of DNA in a sequencing experiment. The read budget can be based on total reads produced, e.g., including redundant reads produced through amplification. Alternatively, it can be based on number of unique molecules detected in the sample. In certain embodiments read budget can reflect the amount of double-stranded support for a call at a locus. That is, the percentage of loci for which reads from both strands of a DNA molecule are detected.[000143] Factors of a read budget include read depth and panel length. For example, a read budget of 3,000,000,000 reads can be allocated as 150,000 bases at an average read depth of 20,000 reads / base. Read depth can refer to number of molecules producing a read at a locus. In the present disclosure, the reads at each base can beallocated between bases in the backbone region of the panel, at a first average read depth and bases in the hotspot region of the panel, at a deeper read depth. In some embodiments, a sample is sequenced to a read depth determined by the amount of nucleic acid present in a sample. In some embodiments, a sample is sequenced to a set read depth, such that samples comprising different amounts of nucleic acid are sequenced to the same read depth. For example, a sample comprising 300 ng of nucleic acids can be sequenced to a read depth 1 / 10 that of a sample comprising 30 ng of nucleic acids. In some embodiments, nucleic acids from two or more different subjects can be added together at a ratio based on the amount of nucleic acids obtained from each of the subjects.[000144] By way of non-limiting example, if a read budget consists of 100,000 read counts for a given sample, those 100,000 read counts will be divided between reads of backbone regions and reads of hotspot regions. Allocating a large number of those reads (e.g., 90,000 reads) to backbone regions will result in a small number of reads (e.g., the remaining 10,000 reads) being allocated to hotspot regions. Conversely, allocating a large number of reads (e.g., 90,000 reads) to hotspot regions will result in a small number of reads (e.g., the remaining 10,000 reads) being allocated to backbone regions. Thus, a skilled worker can allocate a read budget to provide desired levels of sensitivity and specificity. In certain embodiments, the read budget can be between 100,000,000 reads and 100,000,000,000 reads, e.g., between 500,000,000 reads and 50,000,000,000 reads, or between about 1 ,000,000,000 reads and 5,000,000,000 reads across, for example, 20,000 bases to 100,000 bases.[000145] All polynucleotides (e.g., amplified polynucleotides) can be submitted to a sequencing device for sequencing. Alternatively, a sampling, or subset, of all of the amplified polynucleotides is submitted to a sequencing device for sequencing. With respect to any original double-stranded polynucleotide there can be three results with respect to sequencing. First, sequence reads can be generated from both complementary strands of the original molecule (that is, from both the Watson strand and from the Crick strand). Second, sequence reads can be generated from only one of the two complementary strands (that is, either from the Watson strand or from the Crick strand, but not both). Third, no sequence read may be generated from either of the two complementary strands. Consequently, counting unique sequence reads mappingto agenetic locus will underestimate the number of double-stranded polynucleotides in the original sample mappingto the locus. Described herein are methods of estimatingthe unseen and uncounted polynucleotides.[000146] The sequencing method can be massively parallel sequencing, that is, simultaneously (or in rapid succession) sequencing any of at least 100, 1000, 10,000, 100,000, 1 million, 10 million, 100 million, or 1 billion polynucleotide molecules.[000147] Sequencing methods may include, but are not limited to: high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), Next generation sequencing, Single Molecule Sequencing by Synthesis (SMSS) (Helicos), massively-parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Maxam-Gilbert or Sanger sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms and any other sequencing methods known in the art.[000148] The method can comprise sequencing at least 1 million, 10 million, 100 million, 500 million, 1 billion, 1.1 billion, 1.2 billion, 1.5 billion, 2 billion, 2.5 billion, 3 billion, 3.5 billion, 4 billion, 4.5 billion, 5 billion, 5.5 billion, 6 billion, 6.5 billion, 7 billion, 8 billion, 9 billion or 10 billion base pairs. In some cases, the methods can comprise sequencing from about 1 billion to about 7 billion, from about 1 .1 billion to about 6.8 billion, from about 1 .2 billion, to about 6.5 billion, from about 1.1 billion to about 6.4 billion, from about 1 .5 billion to about 7 billion, from about 2 billion to about 6 billion, from about 2.5 billion to about 5.5 billion, from about 3 billion to about 5 billion base pairs. For example, the methods can comprise sequencing from about 1 .2 billion, to about 6.5 billion base pairs.Tumor markers[000149] A tumor marker is a genetic variant associated with one or more cancers. Tumor markers may be determined using any of several resources or methods. Atumor marker may have been previously discovered or may be discovered de novo using experimental or epidemiologicaltechniques. Detection of a tumor marker may be indicative of cancer when the tumor marker is highly correlated a cancer. Detection of atumor marker may be indicative of cancer when a tumor marker in a region or gene occur with a frequency that is greater than a frequency for a given background population or dataset.[000150] Publicly available resources such as scientific literature and databases may describe in detail genetic variants found to be associated with cancer. Scientific literature may describe experiments or genome-wide association studies (GWAS) associating one or more genetic variants with cancer. Databases may aggregate information gleaned from sources such as scientific literature to provide a more comprehensive resource for determining one or more tumor markers. Non-limiting examples of databases include FANTOM, GTex, GEO, Body Atlas, INSiGHT, OMIM (Online Mendelian Inheritance in Man, omim.org), cBioPortal (cbioportal.org), CIViC (Clinical Interpretations of Variants in Cancer, civic.genome.wustl.edu), DOCM (Database of Curated Mutations, docm.genome.wustl.edu), and ICGC Data Portal (dcc.icgc.org). In a further example, the COSMIC (Catalogue of Somatic Mutations in Cancer) database allows for searching of tumor markers by cancer, gene, or mutation type. Tumor markers may also be determined de novo by conducting experiments such as case control or association (e.g., genome-wide association studies) studies.[000151] One or more tumor markers may be detected in the sequencing panel. A tumor marker may be one or more genetic variants associated with cancer. Tumor markers can be selected from single nucleotide variants (SNVs), copy number variants (CNVs), insertions or deletions (e.g., indels), gene fusions and inversions. Tumor markers may affect the level of a protein. Tumor markers may be in a promoter or enhancer, and may alter the transcription of a gene. The tumor markers may affect the transcription and / or translation efficacy of a gene. The tumor markers may affect the stability of a transcribed mRNA. The tumor marker may result in a change to the amino acid sequence of a translated protein. The tumor marker may affect splicing, may change the amino acid coded by a particular codon, may result in a frameshift, or may result in a premature stop codon. The tumor marker may result in a conservative substitution of an amino acid. One or more tumor markers may result in a conservative substitution of an amino acid. One or more tumor markers may result in a nonconservative substitution of an amino acid.[000152] One or more of the tumor markers may be a driver mutation. A driver mutation is a mutation that gives a selective advantage to a tumor cell in its microenvironment, through either increasing its survival or reproduction. None of the tumor markers may be a driver mutation. One or more of the tumor markers may be a passenger mutation. A passenger mutation is a mutation that has no effect on the fitness of a tumor cell but may be associated with a clonal expansion because it occurs in the same genome with a driver mutation.[000153] The frequency of a tumor marker may be as low as 0.001 %. The frequency of a tumor marker may be as low as 0.005%. The frequency of a tumor marker may be as low as 0.01%. The frequency of a tumor marker may be as low as 0.02%. The frequency of a tumor marker may be as low as 0.03%. The frequency of a tumor marker may be as low as 0.05%. The frequency of a tumor marker may be as low as 0.1%. The frequency of a tumor marker may be as low as 1 %.[000154] No single tumor marker may be present in more than 50%, of subjects having the cancer. No single tumor marker may be present in more than 40%, of subjects having the cancer. No single tumor marker may be present in more than 30%, of subjects having the cancer. No single tumor marker may be present in more than 20%, of subjects having the cancer. No single tumor marker may be present in more than 10%, of subjects having the cancer. No single tumor marker may be present in more than 5%, of subjects having the cancer. A single tumor marker may be present in 0.001% to 50% of subjects having cancer. A single tumor marker may be present in 0.01 % to 50% of subjects having cancer. A single tumor marker may be present in 0.01% to 30% of subjects having cancer. A single tumor marker may be present in 0.01 % to 20% of subjects having cancer. A single tumor marker may be present in 0.01 % to 10% of subjects having cancer. A single tumor marker may be present in 0.1 % to 10% of subjects having cancer. A single tumor marker may be present in 0.1% to 5% of subjects having cancer.[000155] Detection of a tumor marker may indicate the presence of one or more cancers. Detection may indicate presence of a cancer selected from the group comprising ovarian cancer, pancreatic cancer, breast cancer, colorectal cancer, nonsmall cell lung carcinoma (e.g., squamous cell carcinoma, or adenocarcinoma) or any other cancer. Detection may indicate the presence of any cancer selected from thegroup comprising ovarian cancer, pancreatic cancer, breast cancer, colorectal cancer, non-small cell lung carcinoma (squamous cell or adenocarcinoma) or any other cancer. Detection may indicate the presence of any of a plurality of cancers selected from the group comprising ovarian cancer, pancreatic cancer, breast cancer, colorectal cancer and non-small cell lung carcinoma (squamous cell or adenocarcinoma), or any other cancer. Detection may indicate presence of one or more of any of the cancers mentioned in this application.[000156] One or more cancers may exhibit a tumor marker in at least one exon in the panel. One or more cancers selected from the group comprising ovarian cancer, pancreatic cancer, breast cancer, colorectal cancer, non-small cell lung carcinoma (squamous cell or adenocarcinoma), or any other cancer, each exhibit a tumor marker in at least one exon in the panel. Each of at least 3 of the cancers may exhibit a tumor marker in at least one exon in the panel. Each of at least 4 of the cancers may exhibit a tumor marker in at least one exon in the panel. Each of at least 5 of the cancers may exhibit a tumor marker in at least one exon in the panel. Each of at least 8 of the cancers may exhibit a tumor marker in at least one exon in the panel. Each of at least 10 of the cancers may exhibit a tumor marker in at least one exon in the panel. All of the cancers may exhibit a tumor marker in at least one exon in the panel.[000157] If a subject has a cancer, the subject may exhibit a tumor marker in at least one exon or gene in the panel. At least 85% of subjects having a cancer may exhibit a tumor marker in at least one exon or gene in the panel. At least 90%, of subjects having a cancer may exhibit a tumor marker in at least one exon or gene in the panel. At least 92% of subjects having a cancer may exhibit a tumor marker in at least one exon or gene in the panel. At least 95% of subjects having a cancer may exhibit a tumor marker in at least one exon or gene in the panel. At least 96% of subjects having a cancer may exhibit a tumor marker in at least one exon or gene in the panel. At least 97% of subjects having a cancer may exhibit a tumor marker in at least one exon or gene in the panel. At least 98% of subjects having a cancer may exhibit a tumor marker in at least one exon or gene in the panel. At least 99% of subjects having a cancer may exhibit a tumor marker in at least one exon or gene in the panel. At least 99.5% of subjects having a cancer may exhibit a tumor marker in at least one exon or gene in the panel.[000158] If a subject has a cancer, the subject may exhibit a tumor marker in at least one region in the panel. At least 85% of subjects having a cancer may exhibit a tumor marker in at least one region in the panel. At least 90%, of subjects having a cancer may exhibit a tumor marker in at least one region in the panel. At least 92% of subjects having a cancer may exhibit a tumor marker in at least one region in the panel. At least 95% of subjects having a cancer may exhibit a tumor marker in at least one region in the panel. At least 96% of subjects having a cancer may exhibit a tumor marker in at least one region in the panel. At least 97% of subjects having a cancer may exhibit a tumor marker in at least one region in the panel. At least 98% of subjects having a cancer may exhibit a tumor marker in at least one region in the panel. At least 99% of subjects having a cancer may exhibit a tumor marker in at least one region in the panel. At least 99.5% of subjects having a cancer may exhibit a tumor marker in at least one region in the panel.[000159] Detection may be performed with a high sensitivity and / or a high specificity. Sensitivity can refer to a measure of the proportion of positives that are correctly identified as such. In some cases, sensitivity refers to the percentage of all existing tumor markers that are detected. In some cases, sensitivity refers to the percentage of sick people who are correctly identified as having certain disease. Specificity can refer to a measure of the proportion of negatives that are correctly identified as such. In some cases, specificity refers to the proportion of unaltered bases which are correctly identified. In some cases, specificity refers to the percentage of healthy people who are correctly identified as not having certain disease. The nonunique tagging method described previously significantly increases specificity of detection by reducing noise generated by amplification and sequencing errors, which reduces frequency of false positives. Detection may be performed with a sensitivity of at least 95%, 97%, 98%, 99%, 99.5%, or 99.9% and / or a specificity of at least 80%, 90%, 95%, 97%, 98% or 99%. Detection may be performed with a sensitivity of at least 90%, 95%, 97%, 98%, 99%, 99.5%, 99.6%, 99.98%, 99.9% or 99.95%. Detection may be performed with a specificity of at least 90%, 95%, 97%, 98%, 99%, 99.5%, 99.6%, 99.98%, 99.9% or 99.95%. Detection may be performed with a specificity of at least 70% and a sensitivity of at least 70%, a specificity of at least 75% and a sensitivity of at least 75%, a specificity of at least 80% and a sensitivity of at least 80%, a specificity ofat least 85% and a sensitivity of at least 85%, a specificity of at least 90% and a sensitivity of at least 90%, a specificity of at least 95% and a sensitivity of at least 95%, a specificity of at least 96% and a sensitivity of at least 96%, a specificity of at least 97% and a sensitivity of at least 97%, a specificity of at least 98% and a sensitivity of at least 98%, a specificity of at least 99% and a sensitivity of at least 99%, or a specificity of 100% a sensitivity of 100%. In some cases, the methods can detect a tumor marker at a sensitivity of sensitivity of about 80% or greater. In some cases, the methods can detect a tumor marker at a sensitivity of sensitivity of about 95% or greater. In some cases, the methods can detect a tumor marker at a sensitivity of sensitivity of about 80% or greater, and a sensitivity of sensitivity of about 95% or greater.[000160] Detection may be highly accurate. Accuracy may apply to the identification of tumor markers in cell free DNA, and / or to the diagnosis of cancer. Statistical tools, such as co-variate analysis described above, may be used to increase and / or measure accuracy. The methods can detect a tumor marker at an accuracy of at least 80%, 90%, 95%, 97%, 98% or 99%, 99.5%, 99.6%, 99.98%, 99.9%, or 99.95%. In some cases, the methods can detect a tumor marker at an accuracy of at least 95% or greater.[000161] In one implementation, using measurements from a plurality of samples collected substantially at once or over a plurality of time points, the diagnostic confidence indication for each variant can be adjusted to indicate a confidence of predicting the observation of the copy number variation (CNV) or mutation or tumor marker. The confidence can be increased by using measurements at a plurality of time points to determine whether cancer is advancing, in remission or stabilized. The diagnostic confidence indication can be assigned by any of a number of statistical methods and can be based, at least in part, on the frequency at which the measurements are observed over a period of time. For example, a statistical correlation of current and prior results can be done. Alternatively, for each diagnosis, a hidden Markov model can be built, such that a maximum likelihood or maximum a posteriori decision can be made based on the frequency of occurrence of a particular test event from a plurality of measurements or a time points. As part of this model, the probability of error and resultant diagnostic confidence indication for a particular decision can be output as well. In this manner, the measurements of a parameter, whether or not theyare in the noise range, may be provided with a confidence interval. Tested over time, one can increase the predictive confidence of whether a cancer is advancing, stabilized or in remission by comparing confidence intervals overtime. Two samplingtime points can be separated by at least about 1 microsecond, 1 millisecond, 1 second, 10 seconds, 30 seconds, 1 minute, 10 minutes, 30 minutes, 1 hour, 12 hours, 1 day, 1 week, 2 weeks, 3 weeks, one month, or one year. Two time points can be separated by about a month to about a year, about a year to about 5 years, or no more than about three months, two months, one month, three weeks, two weeks, one week, one day, or twelve hours. In some embodiments, two time points can be separated by a therapeutic event such as the administration of a treatment or the performance of a surgical procedure. When the two time points are separated by the therapeutic event, CNV or mutations detected can be compared before and after the event.[000162] After sequencing data of cell free polynucleotide sequences is collected, one or more bioinformatics processes may be applied to the sequence data to detect genetic features or variations such as cfDNA characteristics at regulatory elements, nucleosomal spacing / nucleosome binding patterns, chemical modifications of nucleic acids, copy number variation, and mutations or changes in epigenetic markers, including but not limited to methylation profiles, and genetic variants such as SNVs, CNVs, indels, and / or fusions. In some cases, in which copy number variation analysis is desired, sequence data may be: 1 ) aligned with a reference genome and mapped to individual molecules; 2) filtered; 4) partitioned into windows or bins of a sequence; 5) coverage reads and molecules counted for each window; 6) coverage molecules can then be normalized using a statistical modeling algorithm; and 7) an output file can be generated reflecting discrete copy number states at various positions in the genome. In some cases, the number of coverage reads / molecules or normalized coverage reads aligning to a particular locus of the reference genome is counted. In other cases, in which mutation analysis is desired, sequence data may be 1) aligned with a reference genome and mapped to individual molecules; 2) filtered; 4) frequency of variant bases calculated based on coverage reads for that specific base; 5) variant base frequency normalized using a stochastic, statistical or probabilistic modeling algorithm; and 6) an output file can be generated reflecting mutation states at various positions in the genome. In some cases, identifiers (such as those including barcodes) can be used togroup sequence reads during mutation analysis. In some cases, sequence reads are grouped into families, e.g., by using identifiers or a combination of identifiers and start / stop positions or sequences. In some cases, a base call can be made by comparing nucleotides in one or more families to a reference sequence and determining the frequency of a particular base 1 ) within each family, and 2) between the families and the reference sequences. A nucleotide base call can be made based on criteria such as the percentage of families having a base at a position. In some cases, a base call is reported if its frequency is greater than a noise threshold as determined by frequency in a plurality of reference sequences (e.g., sequences from healthy individuals). Temporal information from the current and prior analysis of the patient or subject is used to enhance the analysis and determination. In some embodiments, sequence information from the patient or subject is compared to sequence information obtained from a cohort of healthy individuals, a cohort of cancer patients, or germline DNA from the patient or subject. Germline DNA can be obtained, without limitation, from bodily fluid, whole blood, platelets, serum, plasma, stool, red blood cells, white blood cells or leukocytes, endothelial cells, tissue biopsies, synovial fluid, lymphatic fluid, ascites fluid, interstitial or extracellular fluid, the fluid in spaces between cells, including gingival crevicular fluid, bone marrow, cerebrospinal fluid, saliva, mucous, sputum, semen, sweat, urine, or any other bodily fluids. A cohort of cancer patients can have the same type of cancer as the patient or subject, the same stage of cancer as the patient or subject, both, or neither. In some embodiments, a cohort of cancer patients, a cohort of healthy individuals, or germline DNA from the subject is used to provide a baseline frequency of a base at a position, and the baseline frequency is used in making a base call in the subject. Without limitation, a frequency for a base at a position in a cohort of healthy individuals, or germline DNA from the subject can be compared to the frequency of a base detected among sequence reads from the subject.[000163] In some embodiments, the methods and systems of the present disclosure can be used to detect a minor allele frequency (MAF) of 0.025% or lower, 0.05% or lower, 0.075% or lower, or 0.1 % or lower. Copy number variation can be measured as a ratio of (1) unique molecule counts (UMCs) for a gene in a test sample to (2) UMCs for that gene in a reference sample (e.g., control sample). In some embodiments, the methods and systems of the present disclosure can be used todetect a copy number variation that is a copy number amplification (CNA). In some embodiments, the methods and systems of the present disclosure can be used to detect a CNA of at least 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, or more. In some embodiments, the methods and systems of the present disclosure can be used to detect a copy number variation that is a copy number loss (CNL). In some embodiments, the methods and systems of the present disclosure can be used to detect a CNL of less than 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1 , or 0.05.[000164] A variety of different reactions and / operations may occur within the systems and methods disclosed herein, including but not limited to: nucleic acid sequencing, nucleic acid quantification, sequencing optimization, detecting gene expression, quantifying gene expression, genomic profiling, cancer profiling, or analysis of expressed markers. Moreover, the systems and methods have numerous medical applications. For example, it may be used for the identification, detection, diagnosis, treatment, monitoring, staging of, or risk prediction of various genetic and non-genetic diseases and disorders including cancer. It may be used to assess subject response to different treatments of the genetic and non-genetic diseases, or provide information regarding disease progression and prognosis. Further information is found in U.S. Pat. No. US11795513, U.S. Pat. Pub. No. 20230313315, Inti. PCT Pub. No. 2016109452, 2016179049 and U.S. Prov. App. No. 63 / 786,007, each of which is fully incorporated by reference herein.EXAMPLESExample 1 - Sequencing read allocation in panels are demanding[000165] As shown in Figure 1 . SNV calling at 2.5% MAF at 95% LoD needs above 666 diversity. To estimate the feasibility of the 55M somatic read allocation for SNV detection. A VSCORE panel wide variant simulation was performed for clinical samples enriched on a somatic / epi blend to estimating the calling probability of MAFs within the sample. A sigmoidal curve fit for the relationship between the methylation hotspot diversity and SNV calling probability was fit for the varying subpanels / variant MAFs. Thediversity estimates (red vertical dashed line) are where this fit sigmoid intersected 95% confidence (green horizontal dashed line).[000166] What is apparent is that for the base ~3Mb methylation ‘hotspot’ somatic panel, to achieve the desired panel-wide 5% MAF SNV detection would take ~378 diversity, For a 229kb SNV range, to achieve the desired 2.5% MAF SNV would take ~666 diversity, For the 24 exon 3.8kb hotspot (current floor 2.5%), to achieve the desired 2.5% MAF SNV would take ~517 diversity.Example 2 - Increasing read allocation or reducing effective panel size achieve lower MAF[000167] As shown in Figure 2. 55M somatic reads are only sufficient for 5% panelwide MAF. The described simulation using clinical samples shows that hotspot diversities of 378 and 666 are needed for variant calling at 2.5-5% MAF panel-wide or a sub-set of panel genes. This translates to a read requirement of 38M (5% MAF) and 65M (2.5% MAF). In turn, this necessitates the need to boost reportable probes to reduce read requirements for lower detection sensitivity at 1 -2.5% MAF targetExample 3 - RNA exome panels are highly informative and efficient[000168] As shown in Figure 3. one can turn to exome coverage for germline variant calling to improve CNV sensitivity as 20M exome reads can detect 10% MAF at ~45% confidence. Therefore, a skewed molecule coverage for an exome panel was observed from the blending of somatic, epi, and exome panels to reduce the read allocation to the large epi (15Mb) and exome (34Mb) panels. For example, observed exome coverage rates (blacklisting 10% low-coverage regions) shows variant detection at 10% MAF of ~50% confidence with 20M reads, and 10% MAF dection at 95% confidence will take ~79M reads (excluding the 10% worst coverage probes). As futher shown in Figure 4., while exome wide 10% MAF calling would take ~39M exome reads, 5% MAF would require ~78M reads. Moreover, germline SNPs (i.e. ~50% MAF) require significantly fewer reads and can still offer significant value supporting enhanced CNV calling.Example 4 - Estimation of read / diversity depth needed by samples to achieve somatic SNV calling at various MAFs in an exome panel[000169] Germline SNVs are essentially 50% MAFs - and only need a diversity of 23 molecules to make a call. This implies our read count requirements for germline calling would be between 4.5M and 6.9M reads. One would need to take into account large, skewed variations in coverage across the panel, which will further drive up the required read depth to ensure panel-wide performance. Nevertheless, 10% MAF exome-wide detection should take ~39M+ exome reads andGermline exome-wide detection should take ~4.5-6.9M exome reads.Example 5 - Mixture of blended ratio to fit within sequence read budget[000170] As shown in Figure 5., one can calculate a mixture of blended ratio to fit within seq budget. Here, one enriches and blends the epi molecules after probe hybridization to avoid skewed coverage. Blended panel probe concentration on the order of ~0.15x-1x. Here, a read allocation to exome and epi panel remain high despite low per probe concentration. A low per probe concentration leads to observed skewed molecular coverage (potential dropouts) in exome panel.Example 6 - Probe design strategies for exon-skipping[000171] An area to maximize information efficiency from exome panel design is to account for repeat elements, untranslated regions, simple repeats that consume probe targeting, but are not informative, while also maximizing informative probes such as fusion breakpoint spanning. Two examples of fusion breakpoint spanning are shown in Figure 6. and Figure 7., with Figure 6. showing SLC45A3 including 2 more probes spanning 240bp at a fusion breakpoint. This may be further enhanced by including all 5’UTRs for SLC45A3. Similarly, in Figure 7. Androgen receptor v7 (ARV7) shows a design for probes covering the 3’UTRs of AR down to 400 bp without any of the probes overlapping SINEs or simple repeats. Exemplary related sequences include Genbank: FJ235916.1 ; Accession Nos. NM_001348063.1 / NM_001348064.1 / NM_001348061 .1 .[000172] As shown in Figure 8., EGFRv3 fusion-designed probes can include spanning exon-del-exon of EGFRv3 (Accession Nos. NM_001346941 .2) with proposed 3 fusion-specific probes and one 5’UTR to include spanning read support In Figure 9. METexI 4: probe design specific to detect deletions probes spanning ex14 deletion.Exemplary related sequences include Accession Nos. NM_001324402.2,NM_001127500.3)Example 7 - Probe design strategies enhancing informational efficiency avoiding SINES and simple repeats[000173] As shown in Figure 10., repeat overlap are notable as short interspersed nuclear elements (SINEs) RNA probes can overlap with SINEs (>80bp); 75 of these overlap 100% ot its sequence with a SINE, consuming sequence read budget in a largely non-informative manner. Thus, design strategies should look to avoid any SINE overlap from 3’UTRs. Similarly, as shown in Figure 11 . repeat overlap are also a source of concern as simple repeats of monomer size <=3 covering >50% (60bp) due to spurious hybridization., yet provide little useful information.Example 8 - Probe design strategies enhancing efficiency[000174] As shown in Figure 12., KIF5B include design for last two exons (might include 120bp on 3’UTR), similarly as shown in Figure 13. CD74: Design 2 probe to target the last exon
Claims
THE CLAIMS1 . A method for processing, comprising: contacting converted complementary DNA (cDNA) molecules or amplification products thereof with a panel of different oligonucleotide probes configured to hybridize to cDNA derived from at least 500 target exome regions; and sequencing the enriched DNA or amplification products thereof, wherein each target exome region is associated with a target genomic regions, with a differential methylation pattern in cancerous samples when compared to normal sample.
2. The method of claim 1 , comprising enriching for probe-bound cDNA to produce enriched cDNA.
3. The method of claim 1 , wherein each target genomic regions is determined to be differentially methylated in cancer training samples based on criteria comprising a number of cancer samples that comprise a differential methylation pattern.
4. The method of claim 3, wherein each of target genomic regions is determined to be differentially methylated based on criteria positively correlated with cancer and negatively correlated with non-cancer.
5. The method of claim 1 , comprising a mixture panel, wherein one or sub-panels is configured to detect a different biological analyte.
6. The method of claim 1 , wherein the mixture panel comprises one or more subpanels, with different oligonucleotide probes.
7. The method of claim 6, wherein one or more sub-panels, with different oligonucleotide probes is configured to detect cDNA, and one or more sub-panels, with different oligonucleotide probes is configured to detect cell-free DNA (cfDNA) molecules.
8. The method of claim 1 , wherein the cDNA, cfDNA molecules or amplification products thereof comprise adapter sequences at one or both ends.
9. The method of claim 13, wherein the cfDNA molecules or amplification products thereof comprise a first adapter sequence at a first end and a second adapter sequence at a second end, wherein the first adapter sequence and the second adapter sequence are different.
10. The method of claim 9, wherein one or more of plurality of the adapter sequences are configured to distinguish sequencing reads for different cfDNA molecules.11 . The method of claim 10, wherein the adapter sequences comprise barcodes, unique molecular identifiers, or partition, sample indexes.
12. A panel comprising a plurality of probes, wherein the panel comprises a mixture panel comprising one or more sub-panels, with different oligonucleotide probes.
13. The panel of claim 12, wherein one or more of the different oligonucleotide probes is configured to detect cDNA, one or more of the different oligonucleotide probes is configured to detect cell-free DNA (cfDNA) molecules and / or RNA.
14. The panel of claim 12, wherein the panel comprises a plurality of hotspot region configured for variant calling at 2.5-5% maximum allele fraction (MAF), optionally comprising about 30-70 million sequence reads.
15. The panel of claim 12, wherein the panel comprises a plurality of hotspot region configured forvariant calling at 1- 2.5% maximum allele fraction (MAF).
16. The panel of claim 12, wherein the cDNA is derived from exome sequences.
17. The panel of claim 12, wherein the panel is configured to detect 5-10% MAF within 20-80 million exome sequence reads.
18. The panel of claim 12, wherein the panel comprises one or more oligonucleotide probes configured to span a breakpoint, optionally comprising a fusion breakpoint or deletion breakpoint.
19. The panel of claim 12, wherein the panel comprises one or more oligonucleotide probes configured to bindingto an 5’UTR, or exon, optionally including the last 1 , 2, 3 or more exons of a target gene.
20. The panel of claim 12, wherein the panel comprises one or more oligonucleotide probes configured to avoid to a 3’UTR, short interspersed nuclear elements (SINEs) or short repeat element equal or less than 3 monomers.
Citation Information
Patent Citations
Oligonucleotides
US20010053519A1
Method and apparatus for imaging a sample on a device
US20030152490A1
Digital Counting of Individual Molecules by Stochastic Attachment of Diverse Labels
US20110160078A1
Lung cancer biomarker
US20160109452A1
Roll Deskewing Device for an Electrophotographic Image Forming Device
US20160179049A1