Method for rapid detection of aneuploidies

Amplicon-based sequencing with a single primer pair enhances the sensitivity and specificity of chromosomal abnormality detection, addressing limitations in current methods for cancer, prenatal, and birth defect testing, especially with low DNA input.

JP2026012754APending Publication Date: 2026-01-27JOHNS HOPKINS UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025173346
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-02-06
Filing Date
2025-10-15
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Current methods for detecting chromosomal abnormalities, such as aneuploidy, in cancer, non-invasive prenatal testing, and birth defect evaluation are limited in sensitivity and specificity, particularly when dealing with low input DNA samples.

Method used

A method utilizing amplicon-based sequencing with a single primer pair to amplify unique repeat elements across the genome, followed by mapping and quantifying features of amplicons to identify chromosomal abnormalities, enhancing detection sensitivity and specificity.

Benefits of technology

The method significantly improves the sensitivity and maintains specificity in detecting chromosomal abnormalities, even with low input DNA samples, enabling accurate cancer diagnosis, prenatal testing, and birth defect evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012754000052
    Figure 2026012754000052
  • Figure 2026012754000053
    Figure 2026012754000053
  • Figure 2026012754000054
    Figure 2026012754000054
Patent Text Reader

Abstract

To provide a method for identifying the presence of aneuploidy in the genome of a mammal.SOLUTION: Amplifying a plurality of chromosomal sequences in a cell-free DNA sample with a single primer pair complementary to the chromosomal sequences, wherein the primer pair amplifies a unique amplicon comprising short interspersed repeat sequences, determining at least a portion of the nucleic acid sequence of the plurality of amplicons to produce amplicon sequences, mapping the amplicon sequences to a reference genome, dividing the amplicon sequences into a plurality of genomic intervals, quantifying the number of reads for amplicon sequences mapped to the genomic intervals, and comparing the number of reads of amplicon sequences within a first genomic interval to: Comparing the number of reads of amplicon sequences in the one or more different genomic intervals, thereby identifying the presence of an aneuploidy in the genome of the mammal.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 62 / 849,662, filed May 17, 2019; U.S. Provisional Patent Application No. 62 / 905,327, filed September 24, 2019; and U.S. Provisional Patent Application No. 62 / 971,050, filed February 6, 2020. The disclosures of these prior applications are considered part of the disclosure of this application (and are incorporated herein by reference).

[0002] Government funding statement This invention was made with government support under grants CA230691 and CA230400 awarded by the National Institutes of Health. The United States Government has certain rights in this invention.

[0003] 1. Technical Field This document provides methods and materials for identifying chromosomal abnormalities that can be used in cancer diagnostics, non-invasive prenatal testing (NIPT), preimplantation genetic diagnosis, and birth defect evaluation. For example, this document provides methods and materials for evaluating sequencing data to identify a mammal as having a disease (e.g., cancer or birth defect) associated with one or more chromosomal abnormalities. Additionally or alternatively, this document provides methods and materials for evaluating sequencing data that can be used in cancer diagnostics, non-invasive prenatal testing (NIPT), preimplantation genetic diagnosis, and birth defect evaluation. [Background technology]

[0004] 2. Background information Aneuploidy is defined as an abnormality in the number of chromosomes. It was the first genomic abnormality identified in cancer (Boveri 2008 Journal of cell science 121(Supplement 1):1-84 (Non-Patent Document 1); and Nowell 1976 Science 194(4260):23-28 (Non-Patent Document 2)) and is estimated to be present in >90% of most histopathological types of cancer (Knouse et al. 2017 Annual Review of Cancer Biology 1:335-354 (Non-Patent Document 3)). Aneuploidy in cancer was first detected by karyotype testing, and later evaluated by microarray, Sanger sequencing, and more recently by massively parallel sequencing (Wang et al. 2002 Proceedings of the National Academy of Sciences 99(25):16156-16161 (Non-Patent Document 4)). Recent sequencing methods include those using circular binary segmentation, hidden Markov models, expectation maximization, and mean shift (as reviewed in (Zhao et al. 2013 BMC bioinformatics 14(11):S1 (Non-Patent Document 5))). In addition to applications in cancer genomics, such techniques provide the basis for non-invasive prenatal detection of fetuses with Down syndrome and other trisomies (Bianchi et al. 2015 JAMA 314(2):162-169 (Non-Patent Document 6); Zhao et al. 2015 Clinical chemistry 61(4):608-616 (Non-Patent Document 7)). [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Boveri 2008 Journal of cell science 121(Supplement 1):1-84

Non-Patent Document 2

Non-Patent Document 3

Non-Patent Document 4

Non-Patent Document 5

Non-Patent Document 6

Non-Patent Document 7

[0007] In one aspect, the present document provides a method for testing the existence of aneuploidy in mammalian genome.Method includes: Use primer pair complementary to chromosomal sequence to amplify a plurality of chromosomal sequences in DNA sample to form a plurality of amplicons; Determine at least a part of one or more nucleic acid sequences of the plurality of amplicons; Map the sequenced amplicons to reference genome; Divide DNA sample into a plurality of genome intervals; Quantify a plurality of features for the amplicons that are mapped to genome intervals; Compare the plurality of features of the amplicons in the first genome interval with the plurality of features of the amplicons in one or more different genome intervals; At least 100,000 amplicons are formed in the amplifying step (for example, the plurality of amplicons can include about 745,000 amplicons).

[0008] In some embodiments, the method is performed in vitro. In some embodiments, the plurality of amplicons is about 1,000,000 amplicons, e.g., about 1,000,000 to 10,000 amplicons; about 1,000,000 to 50,000 amplicons; about 1,000,000 to 100,000 amplicons; about 1,000,000 to 200,000 amplicons; about 1,000,000 to 300,000 amplicons; about 1,000,000 to 400,000 amplicons; about 1,000,000 to 500,000 amplicons; about 1,000,000 to 600,000 amplicons; about 1,000,000 to 700,000 amplicons; about 500,000 to 10,000 amplicons; about 400,000 to 10,000 amplicons; about 300,000 to 10,000 amplicons; about 200,000 to 10,000 amplicons; about 100,000 to 10,000 amplicons or about 50,000 to 10,000 amplicons.

[0009] In some embodiments, the plurality of amplicons is about 50,000 amplicons; about 100,000 amplicons; about 150,000 amplicons; about 200,000 amplicons; about 250,000 amplicons; about 300,000 amplicons; about 350,000 amplicons; about 400,000 amplicons; about 450,000 amplicons; about 500,000 amplicons; about 550,000 amplicons; about 600,000 amplicons; about 650,000 amplicons; about 700,000 amplicons; about 750,000 amplicons; about 800,000 amplicons; about 850,000 amplicons; about 900,000 amplicons; about 950,000 amplicons; or about 1,000,000 amplicons.

[0010] In some embodiments, the plurality of amplicons comprises about 750,000 amplicons.

[0011] In some embodiments, the plurality of amplicons comprises about 350,000 amplicons.

[0012] In some embodiments, the number of repeat elements (for example, amplicons) that can be amplified by a single primer pair disclosed herein is the function of the number of repeat elements present in sample and / or the length of the repeat elements present in sample.For example, in some samples, the number of repeat elements (for example, amplicons) that can be detected by a single primer pair is about 750,000 or less amplicons.In some embodiments, in other samples, the number of repeat elements (for example, amplicons) that can be detected by a single primer pair is about 350,000 or less amplicons.

[0013] In some embodiments, the DNA sample is a plurality of euploid DNA samples. In some embodiments, the DNA sample is a plurality of test DNA samples. In some embodiments, the DNA sample is a plurality of test DNA samples. In some embodiments, the DNA sample is derived from plasma. In some embodiments, the DNA sample is derived from serum. In some embodiments, the DNA sample comprises fetal cellular DNA. In some embodiments, the DNA sample comprises at least 3 picograms of DNA. In some embodiments, the mammal is a human. In some embodiments, the primer pair comprises a first primer comprising SEQ ID NO:1 and a second primer comprising SEQ ID NO:10. In some embodiments, the methods provided herein comprise one or more additional primer pairs. In some embodiments, the amplicon comprises a repetitive element (e.g., one or more types of repetitive element shown in Table 1). In some embodiments, the amplicon comprises a unique short interspersed nucleotide element (SINE). In some embodiments, the amplicon comprises a unique long interspersed nucleotide element (LINE).

[0014] In some embodiments, the average length of the amplicons is about 100 base pairs or less. In some embodiments, the average length of the amplicons is less than about 110 bp, for example, about 10 to 110 bp, about 10 to 105 bp, about 10 to 100 bp, about 10 to 99 bp, about 10 to 98 bp, about 10 to 97 bp, about 10 to 96 bp, about 10 to 95 bp, about 10 to 94 bp, about 10 to 93 bp, about 10 to 92 bp, about 10 to 91 bp, about 10 to 90 bp, about 10 to 89 bp, about 10 to 87 bp, about 10 to 86 bp, about 10 to 85 bp, about 10 to 84 bp, about 10 to 83 bp, about 10 to 82 bp, about 10 to 81 bp, about 10 to 8 ... 79bp, about 10-78bp, about 10-77bp, about 10-76bp, about 10-75bp, about 10-74bp, about 10-73bp, about 10-72bp, about 10-71bp, about 10-70bp, about 10-65bp, about 10-60bp, about 10-55bp, about 10-50 bp, about 10-40bp, about 10-30bp, about 10-20bp, about 15-110bp, about 20-110bp, about 25-110bp, about 30-110bp, about 35-110bp, about 40-110bp, about 45-110bp, about 50-110bp, about 55-110bp It is about 60 to 110 bp, about 65 to 110 bp, about 70 to 110 bp, about 75 to 110 bp, about 80 to 110 bp, about 85 to 110 bp, about 90 to 110 bp, about 95 to 110 bp, about 100 to 110 bp, or about 105 to 110 bp.

[0015] In some embodiments, the average length of the amplicons is about 10 bp; about 20 bp; about 30 bp; about 40 bp; about 45 bp; about 50 bp; about 60 bp; about 65 bp; about 70 bp; about 75 bp; about 80 bp; about 85 bp; about 90 bp; about 95 bp; about 100 bp; about 105 bp or about 110 bp.

[0016] In some embodiments, the amplicon comprises one or more long amplicons, the average length of which is 1000 base pairs or more. In some embodiments, the long amplicons comprise DNA from contaminating cells. In some embodiments, the contaminating cells are leukocytes. In some embodiments, the genomic interval comprises from about 100 nucleotides to about 125,000,000 nucleotides (e.g., the genomic interval can comprise about 500,000 nucleotides).

[0017] In another aspect, the disclosure provides a method of assessing a subject for the presence of or risk of developing any of a plurality of cancers, e.g., any of at least four cancers, in the subject, comprising: (i) obtaining, e.g., directly obtaining or indirectly obtaining, a value related to, e.g., detecting, the presence of one or more genetic biomarkers, e.g., one or more mutations (e.g., one or more driver gene mutations) in each of one or more genes (e.g., one or more driver genes, e.g., at least four driver genes), where optionally each gene, e.g., driver gene, is associated with the presence or risk of one of a plurality of cancers; (ii) obtaining, e.g., directly or indirectly obtaining, a value relating to, e.g., detecting, the level of each of a plurality of, e.g., at least four, protein biomarkers, wherein optionally, the level of each of the plurality of protein biomarkers is associated with the presence or risk of one of the plurality of cancers; or (iii) obtaining, e.g., directly or indirectly obtaining, a value relating to, e.g., detecting, aneuploidy, wherein the aneuploidy value is a function of the copy number or length of a genomic sequence located between at least two terminal repeat elements of a repetitive element family (RE family), wherein the RE family comprises: (a) RE families other than long interspersed nucleotide sequences (LINEs), (b) an RE family that, when amplified with a primer portion complementary to its own terminal repeat element, results in an amplicon having an average length of less than X nt, where X is 100, 105, or 110; (c) an RE family that is less than about 700 bp in length; or (d) RE families present in at least 100 copies / genome Includes; Optionally, the aneuploidy is associated with the presence or risk of one of multiple cancers; thereby assessing the subject for the presence of or risk of developing any of a plurality of cancers, e.g., any of at least four cancers. The present invention provides a method comprising:

[0018] In one embodiment, one of (i), (ii), and (iii) is obtained directly. In one embodiment, (i) and (ii) are obtained directly. In one embodiment, (i) and (iii) are obtained directly. In one embodiment, (ii) and (iii) are obtained directly. In one embodiment, all of (i), (ii), and (iii) are obtained directly.

[0019] In one embodiment, one of (i), (ii), and (iii) is obtained indirectly. In one embodiment, (i) and (ii) are obtained indirectly. In one embodiment, (i) and (iii) are obtained indirectly. In one embodiment, (ii) and (iii) are obtained indirectly. In one embodiment, all of (i), (ii), and (iii) are obtained indirectly.

[0020] In one embodiment, the method comprises sequencing one or more subgenomic intervals or amplicons comprising the genetic biomarkers. In one embodiment, the method comprises analyzing one or more genomic sequences for aneuploidy. In one embodiment, the method comprises contacting the protein biomarkers with a detection reagent. In one embodiment, the method comprises (1) sequencing one or more subgenomic intervals or amplicons comprising the genetic biomarkers, (2) analyzing one or more genomic sequences for aneuploidy, and / or (3) contacting the protein biomarkers with a detection reagent.

[0021] In one embodiment, the aneuploidy value is a function of the copy number of the genomic sequence located between at least two terminal repeat elements of the RE family. In one embodiment, the aneuploidy value is a function of the length of the genomic sequence located between at least two terminal repeat elements of the repetitive element family (RE family).

[0022] In some embodiments, the method is performed in vitro.

[0023] In one embodiment, a sample, e.g., a biological sample, collected from a subject is evaluated for one, two, or all of (i)-(iii). In one embodiment, the biological sample comprises a liquid sample, e.g., a blood sample. In one embodiment, the biological sample comprises a cell-free DNA sample, a plasma sample, or a serum sample. In one embodiment, the biological sample comprises cell-free DNA, e.g., circulating tumor DNA. In one embodiment, the biological sample comprises cells and / or tissues. In one embodiment, the biological sample comprises cells (e.g., normal cells or cancer cells) and cell-free DNA.

[0024] In one embodiment of any of the methods disclosed herein, the specificity of detecting one of the plurality of cancers involving (i), (ii), and (iii) is substantially the same as the specificity of detecting one of the plurality of cancers involving (i);(ii);(iii);(i) and (ii);(i) and (iii); or (ii) and (iii), e.g., not substantially less than the specificity of detecting one of the plurality of cancers involving (i);(ii);(iii);(i) and (ii);(i) and (iii); or (ii) and (iii).

[0025] In one embodiment of any of the methods disclosed herein, the sensitivity of detection of one of the plurality of cancers involving (i), (ii), and (iii) is greater than the sensitivity of detection of one of the plurality of cancers involving (i); (ii); (iii); (i) and (ii); (i) and (iii); or (ii) and (iii), e.g., about 1.1, 1.2, 1.3, 1.4, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 times greater. In one embodiment, an increase in detection sensitivity at a particular specificity, e.g., at a predetermined specificity, e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%, e.g., about a 1.1, 1.2, 1.3, 1.4, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5 or 10 fold increase in detection sensitivity.

[0026] In some embodiments, the plurality of amplicons is about 1,000,000 amplicons, e.g., about 1,000,000 to 10,000 amplicons; about 1,000,000 to 50,000 amplicons; about 1,000,000 to 100,000 amplicons; about 1,000,000 to 200,000 amplicons; about 1,000,000 to 300,000 amplicons; about 1,000,000 to 400,000 amplicons; about 1,000,000 to 500,000 amplicons; about 1,000,000 to 600,000 amplicons; about 1,000,000 to 700,000 amplicons; about 500,000 to 10,000 amplicons; about 400,000 to 10,000 amplicons; about 300,000 to 10,000 amplicons; about 200,000 to 10,000 amplicons; about 100,000 to 10,000 amplicons or about 50,000 to 10,000 amplicons.

[0027] In some embodiments, the plurality of amplicons is about 50,000 amplicons; about 100,000 amplicons; about 150,000 amplicons; about 200,000 amplicons; about 250,000 amplicons; about 300,000 amplicons; about 350,000 amplicons; about 400,000 amplicons; about 450,000 amplicons; about 500,000 amplicons; about 550,000 amplicons; about 600,000 amplicons; about 650,000 amplicons; about 700,000 amplicons; about 750,000 amplicons; about 800,000 amplicons; about 850,000 amplicons; about 900,000 amplicons; about 950,000 amplicons; or about 1,000,000 amplicons.

[0028] In some embodiments, the plurality of amplicons comprises about 750,000 amplicons.

[0029] In some embodiments, the plurality of amplicons comprises about 350,000 amplicons.

[0030] In some embodiments, the number of repeat elements (for example, amplicons) that can be amplified by a single primer pair disclosed herein is the function of the number of repeat elements present in sample and / or the length of the repeat elements present in sample.For example, in some samples, the number of repeat elements (for example, amplicons) that can be detected by a single primer pair is about 750,000 or less amplicons.In some embodiments, in other samples, the number of repeat elements (for example, amplicons) that can be detected by a single primer pair is about 350,000 or less amplicons.

[0031] In some embodiments, the average length of the amplicons is about 100 base pairs or less. In some embodiments, the average length of the amplicons is less than about 110 bp, for example, about 10 to 110 bp, about 10 to 105 bp, about 10 to 100 bp, about 10 to 99 bp, about 10 to 98 bp, about 10 to 97 bp, about 10 to 96 bp, about 10 to 95 bp, about 10 to 94 bp, about 10 to 93 bp, about 10 to 92 bp, about 10 to 91 bp, about 10 to 90 bp, about 10 to 89 bp, about 10 to 87 bp, about 10 to 86 bp, about 10 to 85 bp, about 10 to 84 bp, about 10 to 83 bp, about 10 to 82 bp, about 10 to 81 bp, about 10 to 8 ... 79bp, about 10-78bp, about 10-77bp, about 10-76bp, about 10-75bp, about 10-74bp, about 10-73bp, about 10-72bp, about 10-71bp, about 10-70bp, about 10-65bp, about 10-60bp, about 10-55bp, about 10-50 bp, about 10-40bp, about 10-30bp, about 10-20bp, about 15-110bp, about 20-110bp, about 25-110bp, about 30-110bp, about 35-110bp, about 40-110bp, about 45-110bp, about 50-110bp, about 55-110bp It is about 60 to 110 bp, about 65 to 110 bp, about 70 to 110 bp, about 75 to 110 bp, about 80 to 110 bp, about 85 to 110 bp, about 90 to 110 bp, about 95 to 110 bp, about 100 to 110 bp, or about 105 to 110 bp.

[0032] In some embodiments, the average length of the amplicons is about 10 bp; about 20 bp; about 30 bp; about 40 bp; about 45 bp; about 50 bp; about 60 bp; about 65 bp; about 70 bp; about 75 bp; about 80 bp; about 85 bp; about 90 bp; about 95 bp; about 100 bp; about 105 bp or about 110 bp.

[0033] In some embodiments, the method further comprises subjecting the subject to a radiation scan of an organ or internal body region, such as a PET-CT scan. In some embodiments, cancer is characterized by the radiation scan of the organ or internal body region. In some embodiments, the location of cancer is identified by the radiation scan of the organ or internal body region. In some embodiments, the radiation scan is a PET-CT scan. In some embodiments, the radiation scan is performed after the subject has been evaluated for the presence of each of multiple cancers.

[0034] In another aspect, the disclosure provides a method of testing for the presence of aneuploidy in a mammalian genome, the method comprising: a) amplifying a plurality of chromosomal sequences in a DNA sample to form a plurality of amplicons using primer portions, e.g., a primer or primer pair complementary to the chromosomal sequences, e.g., the primer portions amplify a sufficient number of sequences to allow detection of aneuploidy; b) determining at least a portion of the nucleic acid sequence of one or more of the plurality of amplicons; c) mapping the sequenced amplicons to a reference genome; d) dividing the DNA sample into a plurality of genomic intervals; e) quantifying a plurality of features for the amplicons mapped to the genomic interval; f) comparing a plurality of features of the amplicons in the first genomic interval with a plurality of features of the amplicons in one or more different genomic intervals. Includes; The amplifying step forms a sufficient number of amplicons to detect aneuploidy, for example, at least 10,000, 20,000, 50,000 or 100,000 amplicons.

[0035] In some embodiments, the method is performed in vitro.

[0036] In one embodiment of any of the methods disclosed herein, the specificity of the detection of one of the plurality of cancers is not affected, e.g., is not reduced or is not substantially reduced, by an increase in the sensitivity of detection of one of the plurality of cancers. In one embodiment, the specificity of the detection of one of the plurality of cancers plateaus, e.g., the specificity of detection does not change with the detection of additional biomarkers.

[0037] In another aspect, provided herein is a method of detecting aneuploidy in a sample containing low input DNA using any of the methods disclosed herein.

[0038] In some embodiments, the sample contains between about 0.01 picograms (pg) and 500 pg of DNA, or between about 0.01 and 500 pg, 0.05 and 400 pg, 0.1 and 300 pg, 0.5 and 200 pg, 1 and 100 pg, 10 and 90 pg, or 20 and 50 pg of DNA. In some embodiments, the sample contains at least 0.01 pg, at least 0.01 pg, at least 0.1 pg, at least 1 pg, at least 2 pg, at least 3 pg, at least 4 pg, at least 5 pg, at least 6 pg, at least 7 pg, at least 8 pg, at least 9 pg, at least 10 pg, at least 11 pg, at least 12 pg, at least 13 pg, at least 14 pg, at least 15 pg, at least 16 pg, at least 17 pg, at least 18 pg, at least 19 pg, at least 20 pg, at least 21 pg, at least 22 pg, at least 23 pg, at least 24 pg, at least 25 pg, at least 26 pg, at least 27 pg, at least 28 pg, at least 29 pg, at least 30 pg, at least 31 pg, at least 32 pg, at least 33 pg, at least 34 pg, at least 35 pg, at least 36 pg, at least 37 pg, at least 38 pg, at least 39 pg, at least 40 pg, at least 41 pg, at least 42 pg, at least 43 pg, at least 44 pg, at least 45 pg, at least 46 pg, at least 47 pg, at least 48 pg, at least 49 pg, at least 50 pg, at least 51 pg, at least 52 pg, at least 53 pg, at least 54 pg, at least 55 pg, at least 56 pg, at least 57 pg, at least 58 pg, at least 59 pg, at least 60 pg, at least 6 pg, at least 39 pg, at least 40 pg, at least 50 pg, at least 60 pg, at least 70 pg, at least 80 pg, at least 90 pg, at least 100 pg, at least 150 pg, at least 200 pg, at least 300 pg, at least 350 pg, at least 400 pg, at least 450 pg or at least 500 pg of DNA.

[0039] In some embodiments, the sample comprises 1 pg of DNA. In some embodiments, the sample comprises 2 pg of DNA. In some embodiments, the sample comprises 3 pg of DNA. In some embodiments, the sample comprises 4 pg of DNA. In some embodiments, the sample comprises 5 pg of DNA. In some embodiments, the sample comprises 10 pg of DNA.

[0040] In some embodiments, the sample is a biological sample from a subject. In one embodiment, the biological sample comprises a liquid sample, such as a blood sample. In one embodiment, the biological sample comprises a cell-free DNA sample, a plasma sample, or a serum sample. In one embodiment, the biological sample comprises cell-free DNA, such as circulating tumor DNA. In one embodiment, the biological sample comprises cells and / or tissues. In one embodiment, the biological sample comprises cells (e.g., normal cells or cancer cells) and cell-free DNA.

[0041] In some embodiments, the sample is a trisomy 21 sample. In some embodiments, the sample is a forensic sample. In some embodiments, the sample is derived from an embryo, for example, from a pre-implantation embryo.

[0042] In some embodiments, the sample is a biobank sample, for example, as described in Example 3.

[0043] In some embodiments, the methods are used for diagnostics, such as preimplantation diagnostics.

[0044] In some embodiments, the method is used for forensic purposes.

[0045] In some aspects, the method is an in vitro method.

[0046] In another aspect, provided herein is a method of identifying or distinguishing a sample using any of the methods disclosed herein.

[0047] In some embodiments, a sample, e.g., a first sample, from a subject (e.g., a first subject) is distinguished from a second sample from a second subject. In some embodiments, a sample, e.g., a first sample, is identified as being from the first subject based on a polymorphism (e.g., multiple polymorphisms, e.g., common polymorphisms). In some embodiments, a second sample is identified as being from the second subject based on a polymorphism (e.g., multiple polymorphisms, e.g., common polymorphisms). In some embodiments, the common polymorphisms are present within a repeat element, e.g., as described herein. In some embodiments, the method disclosed in Example 8 can be used to identify and / or distinguish samples.

[0048] In another aspect, provided herein is a reaction mixture comprising at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten detection reagents, wherein one detection reagent mediates a readout that is a value related to (i) one or more gene biomarkers referred to herein; (ii) one or more protein biomarkers referred to herein; and / or (iii) the copy number or length of a genomic sequence located between at least two terminal repeat elements of a repetitive element family (RE family) referred to herein, e.g., the level or presence of aneuploidy.

[0049] In yet another aspect, the present disclosure provides a kit comprising: (a) at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten detection reagents, wherein one detection reagent mediates a readout that is: (i) one or more gene biomarkers referred to herein; (ii) one or more protein biomarkers referred to herein; and / or (iii) a value related to the copy number or length of a genomic sequence located between at least two terminal repeat elements of a repetitive element family (RE family) referred to herein, e.g., the level or presence of aneuploidy; and (b) instructions for using the kit.

[0050] In some embodiments of any of the methods disclosed herein, quantifying the amplicons mapped to the genomic interval comprises identifying a plurality of genomic intervals that share one or more shared amplicon characteristics. In some embodiments, the shared amplicon characteristic is the number of mapped amplicons.

[0051] In some embodiments of any of the methods disclosed herein, the shared amplicon feature is the average length of the mapped amplicons. In some embodiments, a plurality of genomic intervals having the shared amplicon feature are grouped into clusters. In some embodiments, each cluster comprises about 200 genomic intervals. In some embodiments, the clusters comprise predefined clusters. In some embodiments, comparing the genomic intervals further comprises matching one or more genomic intervals from the test sample with the predefined clusters. In some embodiments, matching the genomic intervals from the test sample with the predefined clusters further comprises identifying one or more genomic intervals having the shared amplicon feature outside a predetermined significance threshold for the predefined clusters. In some embodiments, the method comprises supervised machine learning. In some embodiments, the supervised machine learning uses a support vector machine model.

[0052] In some embodiments of any of the methods disclosed herein, a single primer pair is used to amplify multiple amplicons from a DNA sample, the single primer including a first primer comprising a sequence at least 80% identical to SEQ ID NO:1 and a second primer comprising a sequence at least 80% identical to SEQ ID NO:10. In some embodiments, the sequence of the first primer is at least 90% identical to SEQ ID NO:1. In some embodiments, the sequence of the first primer is at least 95% identical to SEQ ID NO:1. In some embodiments, the sequence of the first primer is 100% identical to SEQ ID NO:1. In some embodiments, the sequence of the second primer is at least 90% identical to SEQ ID NO:10. In some embodiments, the sequence of the second primer is at least 95% identical to SEQ ID NO:10. In some embodiments, the sequence of the second primer is 100% identical to SEQ ID NO:10. In some embodiments, a kit comprising a primer pair is used to amplify multiple amplicons from a DNA sample, wherein a first primer of the primer pair comprises SEQ ID NO:1 or a sequence at least 80% identical thereto, and a second primer of the primer pair comprises SEQ ID NO:10 or a sequence at least 80% identical thereto.

[0053] In another aspect, the present disclosure provides a method for testing a mammal for the presence of cancer.The method includes: a) amplifying a plurality of chromosomal sequences in a DNA sample using a primer pair complementary to the chromosomal sequence to form a plurality of amplicons; b) determining at least a portion of one or more nucleic acid sequences of the plurality of amplicons; c) mapping the sequenced amplicons to a reference genome; d) dividing the DNA sample into a plurality of genomic intervals; e) quantifying a plurality of features for the amplicons mapped to the genomic intervals; f) comparing the plurality of features of the amplicons in a first genomic interval with the plurality of features of the amplicons in one or more different genomic intervals; and g) determining that the mammal has cancer if the plurality of features of the amplicons in the first genomic interval are different from the plurality of features of the amplicons in one or more different genomic intervals.In some embodiments, the method can include at least 100,000 amplicons formed in the amplifying step.In some embodiments, the cancer can be stage I cancer. In some aspects, the cancer may be liver cancer, ovarian cancer, esophageal cancer, gastric cancer, pancreatic cancer, colorectal cancer, lung cancer, breast cancer, or prostate cancer.

[0054] In some aspects, the method is an in vitro method.

[0055] In some embodiments of any of the methods, reaction mixtures, or kits disclosed herein, the plurality of amplicons is about 1,000,000 amplicons, e.g., about 1,000,000 to 10,000 amplicons; about 1,000,000 to 50,000 amplicons; about 1,000,000 to 100,000 amplicons; about 1,000,000 to 200,000 amplicons; about 1,000,000 to 300,000 amplicons; about 1,000,000 to 400,000 amplicons; about 1,000,000 to 500,000 amplicons; about 1,000,000 to 600,000 amplicons; about 1,000,000 to 700,000 amplicons. Approximately 1,000,000 to 800,000 amplicons; approximately 1,000,000 to 900,000 amplicons; approximately 900,000 to 10,000 amplicons; approximately 800,000 to 10,000 amplicons; approximately 700,000 to 10,000 amplicons; approximately 600,000 to 10,000 amplicons about 500,000 to 10,000 amplicons; about 400,000 to 10,000 amplicons; about 300,000 to 10,000 amplicons; about 200,000 to 10,000 amplicons; about 100,000 to 10,000 amplicons or about 50,000 to 10,000 amplicons.

[0056] In some embodiments, the plurality of amplicons is about 50,000 amplicons; about 100,000 amplicons; about 150,000 amplicons; about 200,000 amplicons; about 250,000 amplicons; about 300,000 amplicons; about 350,000 amplicons; about 400,000 amplicons; about 450,000 amplicons; about 500,000 amplicons; about 550,000 amplicons; about 600,000 amplicons; about 650,000 amplicons; about 700,000 amplicons; about 750,000 amplicons; about 800,000 amplicons; about 850,000 amplicons; about 900,000 amplicons; about 950,000 amplicons; or about 1,000,000 amplicons.

[0057] In some embodiments, the plurality of amplicons comprises about 750,000 amplicons.

[0058] In some embodiments, the plurality of amplicons comprises about 350,000 amplicons.

[0059] In some embodiments of any method disclosed herein, the number of repeat elements (for example, amplicons) that can be amplified by a single primer pair disclosed herein is the function of the number of repeat elements present in sample and / or the length of the repeat elements present in sample.For example, in some samples, the number of repeat elements (for example, amplicons) that can be detected by a single primer pair is about 750,000 or less amplicons.In some embodiments, in other samples, the number of repeat elements (for example, amplicons) that can be detected by a single primer pair is about 350,000 or less amplicons.

[0060] In some embodiments of any of the methods, reaction mixtures, or kits disclosed herein, the average length of the amplicons is about 100 base pairs or less. In some embodiments, the average length of the amplicons is less than about 110 bp, e.g., about 10-110 bp, about 10-105 bp, about 10-100 bp, about 10-99 bp, about 10-98 bp, about 10-97 bp, about 10-96 bp, about 10-95 bp, about 10-94 bp, about 10-93 bp, about 10-92 bp, about 10-91 bp, about 10-90 bp, about 10-89 bp, about 10-87 bp, about 10-86 bp, about 10-85 bp, about 10-84 bp, about 10-83 bp, about 10-82 bp, about 10-81 bp, about 10-80 bp, about 10- 79bp, about 10-78bp, about 10-77bp, about 10-76bp, about 10-75bp, about 10-74bp, about 10-73bp, about 10-72bp, about 10-71bp, about 10-70bp, about 10-65bp, about 10-60bp, about 10-55bp, about 10-50 bp, about 10-40bp, about 10-30bp, about 10-20bp, about 15-110bp, about 20-110bp, about 25-110bp, about 30-110bp, about 35-110bp, about 40-110bp, about 45-110bp, about 50-110bp, about 55-110bp It is about 60 to 110 bp, about 65 to 110 bp, about 70 to 110 bp, about 75 to 110 bp, about 80 to 110 bp, about 85 to 110 bp, about 90 to 110 bp, about 95 to 110 bp, about 100 to 110 bp, or about 105 to 110 bp.

[0061] In some embodiments, the average length of the amplicons is about 10 bp; about 20 bp; about 30 bp; about 40 bp; about 45 bp; about 50 bp; about 60 bp; about 65 bp; about 70 bp; about 75 bp; about 80 bp; about 85 bp; about 90 bp; about 95 bp; about 100 bp; about 105 bp or about 110 bp.

[0062] Further features of any of the methods disclosed herein include one or more of the following enumerated aspects.

[0063] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein which equivalents are intended to be encompassed by the following recited embodiments.

[0064] Enumeration E1. A method of assessing a subject for the presence or risk of developing any of a plurality of cancers, e.g., any of at least four cancers, in the subject, comprising the steps of: (i) obtaining, e.g., directly obtaining or indirectly obtaining, a value related to, e.g., detecting, the presence of one or more genetic biomarkers, e.g., one or more mutations (e.g., one or more driver gene mutations) in each of one or more genes (e.g., one or more driver genes, e.g., at least four driver genes), where optionally each gene, e.g., driver gene, is associated with the presence or risk of one of a plurality of cancers; (ii) obtaining, e.g., directly or indirectly obtaining, a value relating to, e.g., detecting, the level of each of a plurality of, e.g., at least four, protein biomarkers, wherein optionally, the level of each of the plurality of protein biomarkers is associated with the presence or risk of one of the plurality of cancers; or (iii) obtaining, e.g., directly or indirectly obtaining, a value relating to, e.g., detecting, aneuploidy, wherein the aneuploidy value is a function of the copy number or length of a genomic sequence located between at least two terminal repeat elements of a repetitive element family (RE family), wherein the RE family comprises: (a) RE families other than long interspersed nucleotide sequences (LINEs), (b) an RE family that, when amplified with a primer portion complementary to its terminal repeat element, results in a plurality of amplicons having an average length of less than X nt, where X is 100, 105, or 110; (c) an RE family that is less than about 700 bp in length; or (d) RE families present in at least 100 copies / genome Includes; Optionally, the aneuploidy is associated with the presence or risk of one of multiple cancers; thereby assessing the subject for the presence of or risk of developing any of a plurality of cancers, e.g., any of at least four cancers; Process. E2. (a) One of (i), (ii) and (iii) is obtained directly; (b)(i) and (ii) are obtained directly; (c) (i) and (iii) are obtained directly; (d)(ii) and (iii) are obtained directly; or (e) all of (i), (ii), and (iii) are obtained directly; The method of embodiment E1. E3. (a) One of (i), (ii) and (iii) is obtained indirectly; (b) (i) and (ii) are obtained indirectly; (c) (i) and (iii) are obtained indirectly; (d)(ii) and (iii) are obtained indirectly; or (e) all of (i), (ii), and (iii) are obtained indirectly; The method of embodiment E1. E4. (1) Sequencing one or more subgenomic intervals or amplicons containing genetic biomarkers; (2) analyzing one or more genomic sequences for aneuploidy; and / or (3) contacting the protein biomarker with a detection reagent The method of any one of aspects E1-E3, comprising: E5. The aneuploidy value is (a) the copy number of the genomic sequence located between at least two terminal repeat elements of the RE family; and / or (b) the length of the genomic sequence located between at least two terminal repeat elements of a repetitive element family (RE family); The method of any one of aspects E1-E4, wherein E6. The method of any one of embodiments E1-E5, wherein a biological sample obtained from the subject is assessed for one, two, or all of (i)-(iii). E7. The method of embodiment E6, wherein the biological sample comprises a liquid sample, eg, a blood sample. E8. The method of embodiment E6 or E7, wherein the biological sample comprises a cell-free DNA sample, a plasma sample, or a serum sample. E9. The method of any one of embodiments E6-E8, wherein the biological sample comprises cell-free DNA, eg, circulating tumor DNA. E10. (i) obtaining the sequence of a subgenomic interval of cell-free DNA from a sample; (ii) obtaining leukocyte parameters, e.g., the sequence of a subgenomic interval of leukocyte DNA from the sample; The method of any one of embodiments E1-E9, further comprising: E11. (i) Obtaining the sequence of a subgenomic interval of cell-free DNA from a sample for aneuploidy analysis; (ii) obtaining leukocyte parameters, e.g., the sequence of a subgenomic interval of leukocyte DNA from the sample for aneuploidy analysis; The method of any one of embodiments E1-E10, further comprising: E12. Comparing (i) to (ii) to assess genomic events, e.g., mutations, found in cell-free DNA subgenomic intervals or cell-free DNA aneuploidy analysis samples. The method of embodiment E10 or E11, further comprising: E13. Further classifying genomic events, e.g., mutations, in the cell-free DNA or subgenomic interval of the cell-free DNA of the aneuploidy analysis, e.g., assigning the mutations to a first class or a second class, the method of any one of aspects E10 to E12. E14. The method of any one of aspects E10-E13, further comprising classifying genomic events, e.g., mutations, in the cell-free DNA or subgenomic intervals of the cell-free DNA of the aneuploidy analysis as growth-dysregulating, e.g., cancerous. E15. The method of any one of aspects E10-E13, further comprising classifying genomic events, e.g., mutations, in the cell-free DNA or the subgenomic interval of the cell-free DNA of the aneuploidy analysis as other than proliferation deregulation, e.g., other than cancer. E16. Genomic events, e.g., mutations, in subgenomic intervals of cell-free DNA or cell-free DNA for aneuploidy analysis (a) the subgenomic interval is aneuploid in cell-free DNA and the subgenomic interval is not aneuploid in leukocytes; or (b) The genomic event is present in the subgenomic interval of cell-free DNA, and the genomic event is absent from the subgenomic interval of leukocytes. classified as cancerous if The method of any one of aspects E10-E14. E17. Genomic events, e.g., mutations, in subgenomic intervals of cell-free DNA or cell-free DNA for aneuploidy analysis (a) the subgenomic interval is aneuploid in cell-free DNA and the subgenomic interval is also aneuploid in leukocytes; or (b) Genomic events are present in subgenomic intervals of cell-free DNA, and genomic events are also present in subgenomic intervals of leukocytes. If the condition is classified as other than proliferation deregulation, The method of any one of aspects E10-E13 or E15. E18. The method of embodiment E17, wherein the genomic event is associated with clonal expansion of leukocytes, eg, age-associated clonal hematopoiesis, eg, clonal hematopoiesis of undetermined potential (CHIP). E19. The method of any one of aspects E1-E18, wherein the specificity for detecting one of the plurality of cancers involving (i), (ii), and (iii) is substantially the same as the specificity for detecting one of the plurality of cancers involving (i);(ii);(iii);(i) and (ii);(i) and (iii); or (ii) and (iii), e.g., not substantially less than the specificity for detecting one of the plurality of cancers involving (i);(ii);(iii);(i) and (ii);(i) and (iii); or (ii) and (iii). E20. The method of any one of aspects E1-E19, wherein the sensitivity of detection of one of the plurality of cancers involving (i), (ii), and (iii) is greater than the sensitivity of detection of one of the plurality of cancers involving (i); (ii); (iii); (i) and (ii); (i) and (iii); or (ii) and (iii), e.g., about 1.1, 1.2, 1.3, 1.4, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 times greater. E21. The method of any one of embodiments E1-E20, wherein (i), (ii), and (iii) result in an increase in detection sensitivity at a particular specificity, e.g., at a predetermined specificity, e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, e.g., about a 1.1, 1.2, 1.3, 1.4, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10-fold increase in detection sensitivity. E22. The method of any one of aspects E20-E21, wherein the increased sensitivity of detection of one of the plurality of cancers does not affect, e.g., does not decrease, or does not substantially decrease, the specificity of detection of one of the plurality of cancers. E23. The method of embodiment E22, wherein the specificity for detecting one of the cancers plateaus. E24. The method of any one of aspects E1-E23, wherein the RE family is other than LINE. E25. When the RE family contains a repeat element and is amplified using a primer for a terminal repeat element, the repeat element is less than about 110 bp, for example, about 10 to 110 bp, about 10 to 105 bp, about 10 to 100 bp, about 10 to 99 bp, about 10 to 98 bp, about 10 to 97 bp, about 10 to 96 bp, about 10 to 95 bp, about 10 to 94 bp, about 10 to 93 bp, about 10 to 92 bp, about 10 to 91 bp, about 10 to 90 bp, about 10 to 89 bp, about 10 to 87 bp, about 10 to 86 bp, about 10 to 85 bp, about 10 to 84 bp, about 10 to 83 bp, about 10 to 82 bp, about 10 to 81 bp, About 10-80bp, about 10-79bp, about 10-78bp, about 10-77bp, about 10-76bp, about 10-75bp, about 10-74bp, about 10-73bp, about 10-72bp, about 10-71bp, about 10-70bp, about 10-65bp, about 10-60bp, about 10-55bp , about 10-50bp, about 10-40bp, about 10-30bp, about 10-20bp, about 15-110bp, about 20-110bp, about 25-110bp, about 30-110bp, about 35-110bp, about 40-110bp, about 45-110bp, about 50-110bp, about 55-110bp The method of any one of aspects E1-E24, wherein the method results in a plurality of amplicons having an average length of about 60-110 bp, about 65-110 bp, about 70-110 bp, about 75-110 bp, about 80-110 bp, about 85-110 bp, about 90-110 bp, about 95-110 bp, about 100-110 bp, or about 105-110 bp. E26. The method of any one of embodiments E1-E25, wherein the RE family comprises one or more repeat elements set forth in Table 1. E27. The method of any one of embodiments E1-E26, wherein the RE family comprises a SINE or tandem repeat (e.g., microsatellite DNA, minisatellite DNA, satellite DNA, or DNA of a gene having multiple copies (e.g., DNA encoding ribosomal RNA)). E28. The method of embodiment E27, wherein the RE family is a SINE, e.g., an Alu family, MIR, or MIR3, or a SINE described in Vassetzky and Kramerov (2013) Nucleic Acids Res. 41:D83-89. E29. The method of any one of embodiments E1-E28, wherein the value of aneuploidy is further a function of the copy number or length of the genomic sequence located between the terminal repeat elements of the LINE repeat element. E30. The method of any one of embodiments E1-E29, wherein the value of aneuploidy is further a function of the copy number or length of a plurality of genomic sequences located between the terminal repeat elements of the repeat element family that, when amplified with primers complementary to their terminal repeat elements, result in an amplicon having an average length of greater than 100 bp. E31. Aneuploidy values ​​are further a) amplifying multiple chromosomal sequences in a DNA sample using primer pairs complementary to the chromosomal sequences to form multiple amplicons; b) determining at least a portion of the nucleic acid sequence of one or more of the plurality of amplicons; c) mapping the sequenced amplicons to the reference genome; d) dividing the DNA sample into multiple genomic intervals; e) quantifying multiple features for amplicons mapped to genomic intervals; f) comparing the plurality of features of the amplicons in the first genomic interval with the plurality of features of the amplicons in one or more different genomic intervals; and g) the amplifying step forms at least 100,000 amplicons The method of any one of aspects E1-E30, wherein the function is E32. The method of any one of embodiments E1-E31, comprising obtaining a value for aneuploidy, wherein the value is a function of the copy number of at least about 5, 10, 20, 30, 50, 100, 200, 500, or 1000 different genomic sequences disposed between the terminal repeat elements of the RE family. E33. The method of any one of aspects E1-E32, wherein the copy number is greater than 2 or less than 2. E34. At least about 100,000 amplicons, about 150,000 amplicons, about 200,000 amplicons; about 250,000 amplicons; about 300,000 amplicons; about 350,000 amplicons; about 400,000 amplicons; about 450,000 amplicons; about 500,000 amplicons; about 550,000 amplicons; about 600 The method of any one of embodiments E31-E33, wherein about 650,000 amplicons; about 700,000 amplicons; about 750,000 amplicons; about 800,000 amplicons; about 850,000 amplicons; about 900,000 amplicons; about 950,000 amplicons; or about 1,000,000 amplicons are formed. E35. Obtaining a value for aneuploidy, wherein the value is: (i) the copy number or length of a first genomic sequence located between terminal repeat elements of the RE family on a first segment of genomic DNA; and (ii) the copy number or length of the second genomic sequence located between the terminal repeat elements of the RE family (e.g., the same or different) on the second segment of genomic DNA; The method of any one of aspects E1-E34, wherein the function is E36. (i) A first segment of genomic DNA and a second segment of genomic DNA are present on different arms of the same chromosome, e.g., the first segment is present on the q arm and the second segment is present on the p arm of the same chromosome; or the first segment is present on the p arm and the second segment is present on the q arm of the same chromosome; (ii) the first segment of genomic DNA and the second segment of genomic DNA are present on the same arm of the same chromosome, e.g., the first segment and the second segment are both present on the p arm or the q arm of the chromosome; and / or (iii) the first segment of genomic DNA and the second segment of genomic DNA are present on different chromosomes, e.g., nonhomologous chromosomes; The method of embodiment E35. E37. The method of any one of embodiments E1-E36, comprising obtaining a value for aneuploidy, wherein the value is a function of the copy number or length of a third genomic sequence located between terminal repeat elements of the RE family on the third chromosome. E38. The method of any one of aspects E1 to E37, comprising obtaining a value for aneuploidy, wherein the value is a function of the copy number or length of the Nth genomic sequence located between the terminal repeat elements of the RE family on the Nth chromosome, where N is 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23. E39. The method of any one of embodiments E1-E38, comprising contacting the genomic nucleic acid of the subject with a primer portion that amplifies a sequence comprising a genomic sequence located between terminal repeat elements of an RE family. E40. The method of embodiment E39, wherein the primer portion is complementary to a terminal element of the RE family. E41. The method of embodiment E39 or E40, wherein the primer portion comprises a primer pair. E42. The method of any one of embodiments E39-E41, wherein the primer portion comprises a single primer, eg, for use in isothermal amplification. E43. The method of any one of aspects E1-E42, wherein the number of biomarkers detected (e.g., number of driver gene mutations) is sufficient such that the sensitivity of detection of the cancers in which each gene, e.g., driver gene, is implicated, among the plurality of cancers is not substantially increased by detection of one or more additional genetic biomarkers. E44. The method of any one of embodiments E1-E42, wherein detecting the genetic biomarker comprises obtaining a sequence (eg, a nucleotide sequence) of the genetic biomarker, eg, by sequencing. E45. The method of embodiment E44, wherein the number of genetic biomarker sequences obtained is sufficient such that the sensitivity of detection of cancers among the plurality of cancers with which each gene, e.g., driver gene, is implicated is not substantially increased by obtaining one or more sequences of additional genetic biomarkers. E46. The method of any one of embodiments E1-E42, wherein detecting the biomarkers comprises obtaining a sequence (eg, a nucleotide sequence) of one or more subgenomic intervals comprising the genetic biomarkers. E47. The method of embodiment E46, wherein the number of sequences of the subgenomic intervals obtained is sufficient such that the sensitivity of detection of cancers among the plurality of cancers to which each gene, e.g., driver gene, is implicated is not substantially increased by obtaining one or more sequences (e.g., nucleotide sequences) of additional subgenomic intervals. E48. The method of any one of embodiments E1-E42, wherein the step of detecting the genetic biomarker comprises obtaining the sequence of an amplicon comprising the genetic biomarker. E49. The method of embodiment E48, wherein the number of amplicon sequences obtained is sufficient such that the sensitivity of detection of cancers among the plurality of cancers in which each gene, e.g., driver gene, is implicated is not substantially increased by obtaining one or more sequences of additional amplicons. E50. The method of embodiment E46, wherein the number of sequences of the subgenomic intervals obtained is sufficient such that the specificity of detection of cancers among the plurality of cancers to which each gene, e.g., driver gene, is implicated is not substantially reduced by obtaining one or more sequences of additional subgenomic intervals. E51. The method of embodiment E48, wherein the number of amplicons obtained is sufficient such that the specificity of detection of cancers in which each gene of the plurality of genes, e.g., driver genes, is implicated, among the plurality of cancers, is not substantially reduced by obtaining one or more sequences of additional amplicons. E52. The method of any preceding embodiment, wherein the plurality of cancers comprises 4, 5, 6, 7, or 8 cancers. E53. The method of any of the preceding embodiments, wherein the multiple cancers are selected from solid tumors, e.g., mesothelioma (e.g., malignant pleural mesothelioma), lung cancer (e.g., non-small cell lung cancer, small cell lung cancer, squamous cell lung cancer, or large cell lung cancer), pancreatic cancer (e.g., pancreatic ductal adenocarcinoma), liver cancer (e.g., hepatocellular carcinoma or cholangiocarcinoma), esophageal cancer (e.g., esophageal adenocarcinoma or squamous cell carcinoma), head and neck cancer, ovarian cancer, colorectal cancer, bladder cancer, cervical cancer, uterine cancer (endometrial cancer), kidney cancer, breast cancer, prostate cancer, brain cancer (e.g., medulloblastoma or glioblastoma), or sarcoma (e.g., Ewing's sarcoma, osteosarcoma, rhabdomyosarcoma), or a combination thereof. E54. The method of any of the previous aspects, wherein the multiple cancers are selected from liver cancer, ovarian cancer, esophageal cancer, gastric cancer, pancreatic cancer, colorectal cancer, lung cancer, breast cancer, or prostate cancer, or a combination thereof. E55. The method of any of the previous embodiments, wherein one or more of the cancers is selected from liver cancer, ovarian cancer, esophageal cancer, gastric cancer, pancreatic cancer, colorectal cancer, lung cancer, or breast cancer. E56. The method of any of the preceding embodiments, wherein one or more of the cancers is a blood cancer. E57. One or more genes, e.g., one or more driver genes, e.g., genes listed in Tables 60 and 61 of US2019 / 0256924A1, e.g.,

[0039] The method of any preceding embodiment, wherein no more than 60, no more than 100, no more than 150, no more than 200, no more than 300, or no more than 400 subgenomic intervals or amplicons from TIFF2026012754000001.tif98161 are sequenced. E58. One or more genes, e.g., one or more driver genes, e.g., genes listed in Tables 60 and 61 of US2019 / 0256924A1, e.g.,

[0039] The method of any preceding embodiment, wherein at least 30, at least 40, at least 50, or at least 60 subgenomic intervals or amplicons from TIFF2026012754000002.tif99159 are sequenced. E59. One or more genes, e.g., one or more driver genes, e.g., one or more genes listed in Tables 60 and 61 of US2019 / 0256924A1, e.g.,

[0039] The method of any preceding embodiment, wherein at least 30 and no more than 400, at least 40 and no more than 300, at least 50 and no more than 200, at least 60 and no more than 150, or at least 60 and no more than 100 subgenomic intervals or amplicons from TIFF2026012754000003.tif99159 are sequenced. E60. The method of any of the preceding embodiments, wherein the number of subgenomic intervals or amplicons sequenced for genes does not exceed 125%, 150%, 200%, or 300% of the lowest number that plateaus in sensitivity for cancer detection. E61. Each subgenomic interval or amplicon of a genetic biomarker is between 6 and 800 bp, e.g., 6-750 bp, 6-700 bp, 6-650 bp, 6-600 bp, 6-550 bp, 6-500 bp, 6-450 bp, 6-400 bp, 6-350 bp, 6-300 bp, 6-250 bp, 6-200 bp, 6-150 bp, 6-100 bp, 10-800 bp, 15-800 bp, 20-800 bp, 25-800 bp, 30-800 bp, 35-800 bp, 40-800 bp, 45-800 bp, 50-800 bp, 55-800 bp, 60-800 bp, 65-800 bp,

[0023] The method of any of the preceding aspects, wherein the nucleic acid sequence comprises 00 bp, 70 to 800 bp, 75 to 800 bp, 80 to 800 bp, 85 to 800 bp, 90 to 800 bp, 95 to 800 bp, 100 to 800 bp, 200 to 800 bp, 300 to 800 bp, 400 to 800 bp, 500 to 800 bp, 600 to 800 bp, 700 to 800 bp, 10 to 700 bp, 20 to 600 bp, 30 to 500 bp, 40 to 400 bp, 50 to 300 bp, 60 to 200 bp, 61 to 150 bp, 62 to 140 bp, 63 to 130 bp, 64 to 120 bp, or 65 to 100 bp, for example, 66 to 80 bp. E62. The method of any preceding embodiment, wherein each subgenomic interval or amplicon of the genetic biomarker comprises about 35, about 40, about 45, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 100, or about 110 bp. E63. The method of any preceding embodiment, wherein each subgenomic interval or amplicon of the genetic biomarker comprises 50 bp or less, 55 bp or less, 60 bp or less, 65 bp or less, 70 bp or less, 75 bp or less, 80 bp or less, 85 bp or less, 90 bp or less, 95 bp or less, 100 bp or less, 200 bp or less, 300 bp or less, 400 bp or less, 500 bp or less, 600 bp or less, 700 bp or less, or 800 bp or less. E64. The method of any preceding embodiment, wherein each subgenomic interval or amplicon of the genetic biomarker comprises at least 6 bp, at least 10 bp, at least 15 bp, at least 20 bp, at least 25 bp, at least 30 bp, at least 35 bp, at least 40 bp, at least 45 bp, or at least 50 bp. E65. The method of any preceding embodiment, wherein each subgenomic interval or amplicon of the genetic biomarker comprises at least 6 pb and no more than 800 bp, at least 10 bp and no more than 700 bp, at least 15 bp and no more than 600 bp, at least 20 bp and no more than 600 bp, at least 25 bp and no more than 500 bp, at least 30 bp and no more than 400 bp, at least 35 bp and no more than 300 bp, at least 40 bp and no more than 200 bp, at least 45 bp and no more than 100 bp, at least 50 bp and no more than 95 bp, or at least 55 bp and no more than 90 bp. E66. The method of any of the previous embodiments, wherein each subgenomic interval or amplicon of the genetic biomarker comprises between 66 and 80 bp. E67. The method of any preceding embodiment, wherein the subgenomic interval or amplicon of the number of genetic biomarkers comprises 2000 bp or less, 2500 bp or less, 3000 bp or less, 3500 bp or less, 4000 bp or less, 5000 bp or less, 6000 bp or less, 7000 bp or less, 8000 bp or less, 9000 bp or less, 10,000 bp or less, 15,000 bp or less, or 20,000 bp or less. E68. The method of any preceding embodiment, wherein the subgenomic interval or amplicon of the number of genetic biomarkers comprises at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, at least 1000 bp, at least 1100 bp, at least 1200 bp, at least 1300 bp, at least 1400 bp, at least 1500 bp, at least 1600 bp, at least 1700 bp, at least 1800 bp, at least 1900 bp, or at least 2000 bp. E69. The method of any preceding embodiment, wherein the subgenomic interval or amplicon of the number of genetic biomarkers comprises at least 200 bp and no more than 20,000 bp, at least 300 bp and no more than 15,000 bp, at least 400 bp and no more than 10,000 bp, at least 500 bp and no more than 9000 bp, at least 600 bp and no more than 8000 bp, at least 700 bp and no more than 7000 bp, at least 800 bp and no more than 6000 bp, at least 900 bp and no more than 5000 bp, at least 1000 bp and no more than 4000 bp, at least 1100 bp and no more than 3500 bp, at least 1200 bp and no more than 3000 bp, at least 1300 bp and no more than 2500 bp, or at least 1500 bp and no more than 2000 bp. E70. Number of gene biomarkers in subgenomic intervals or amplicons: 200bp + 15%, 300bp + 15%, 400bp + 15%, 500bp + 15%, 600bp + 15%, 700bp + 15%, 800bp + 15%, 900bp + 15%, 1000bp + 15%, 1100bp + 15%, 1200bp + 15%, 1300bp + 15%, 1400bp + 15%, 1500bp + 15%, 1600bp + 15%, 1700bp + 15%, 1800bp + 15%, 1900bp + 15%, 2000bp + 15%, 2100bp + 15%, 2200bp + 15%, 2300bp + 15%, 2400bp + 15%, 2500bp + 15%, 2600bp + 15%, 2700bp + 15%, 2800bp + 15%, 2900bp + 15%, 3000bp + 15%, 3100bp + 15%, 3200bp + 15%, 3300bp + 15%, 3400bp + 15%, 3500bp + 15%, 3600bp + 15%, 3700bp + 15%, 3800bp + 15%, 3900bp + 15%, 4000bp + 15%, 4100bp + 15%, 4200bp + 15%, 4300bp + 15%, 4400bp + 15%, 4500bp + 15%, 4600bp + 15%, 4

[0030] The method of any preceding embodiment, comprising: 100bp + 15%, 1900bp + 15%, 2000bp + 15%, 2500bp + 15%, 3000bp + 15%, 3500bp + 15%, 4000bp + 15%, 5000bp + 15%, 6000bp + 15%, 7000bp + 15%, 8000bp + 15%, 9000bp + 15%, 10,000bp + 15%, 15,000bp + 15% or 20,000bp + 15%, e.g., 2000bp + 15%. E71. The method of any preceding embodiment, wherein the subgenomic interval or amplicon of the number of genetic biomarkers comprises 2000 bp. E72. The method of any preceding embodiment, wherein the average depth to which the subgenomic intervals or amplicons of the number of genetic biomarkers are sequenced is at least 5× the sequencing depth. E73. The method of any preceding embodiment, wherein the average depth to which the subgenomic intervals or amplicons of the number of genetic biomarkers are sequenced is a sequencing depth of 500× or less. E74. The method of any preceding embodiment, wherein the average depth to which the subgenomic intervals or amplicons of the number of genetic biomarkers are sequenced is between 5× and 500× the sequencing depth. E75. The method of any preceding embodiment, wherein said detecting step comprises sequencing each subgenomic interval to a depth of at least 50,000 reads / base. E76. The method of any preceding embodiment, wherein said detecting step comprises sequencing each subgenomic interval to a depth of no more than 150,000 reads / base. E77. The method of any preceding embodiment, wherein said detecting step comprises sequencing each subgenomic interval to a depth of between 50,000 reads / base and 150,000 reads / base. E78. The method of any preceding embodiment, wherein said detecting step comprises sequencing each subgenomic interval to a depth sufficient to detect mutations as low as 0.0005% in frequency within said region of interest. E79. 20 bp or less, 21 bp or less, 22 bp or less, 23 bp or less, 24 bp or less, 25 bp or less, 26 bp or less, 27 bp or less, 28 bp or less, 29 bp or less, 30 bp or less, 31 bp or less, 32 bp or less, 33 bp or less, 34 bp or less, 35 bp or less, 40 bp or less, 45 bp or less, 50 bp or less, 55 bp or less, 60 bp or less, 100 bp or less, 200 bp or less, or 300 bp or less is selected from the group consisting of: TIFF2026012754000004.tif99159. E80. At least 6 bp, at least 7 bp, at least 8 bp, at least 9 bp, at least 10 bp, at least 11 bp, at least 12 bp, at least 13 bp, at least 14 bp, at least 15 bp, at least 16 bp, at least 17 bp, at least 18 bp, at least 19 bp, or at least 20 bp are present in each biomarker, e.g., each gene, e.g., each driver gene, e.g., each gene disclosed in Table 60 or 61 of US2019 / 0256924A1, e.g., TIFF2026012754000005.tif99162. E81. At least 6 and no more than 300 bp, at least 7 and no more than 200 bp, at least 8 and no more than 100 bp, at least 9 and no more than 60 bp, at least 10 and no more than 55 bp, at least 11 and no more than 50 bp, at least 12 and no more than 45 bp, at least 13 and no more than 40 bp, at least 14 and no more than 35 bp, at least 15 and no more than 34 bp, at least 14 and no more than 33 bp, at least 15 and no more than 32 bp, at least 16 and no more than 31 bp, at least 17 and no more than 30 bp, at least 18 and no more than 29 bp, at least 19 and no more than 28 bp, at least 20 and no more than 27 bp, TIFF2026012754000006.tif99158. E82. Approximately 33 bp is selected from each biomarker, e.g., each gene, e.g., each driver gene, e.g., each gene disclosed in Table 60 or 61 of US2019 / 0256924A1, e.g., TIFF2026012754000007.tif99156. E83. The method of any preceding embodiment, wherein detecting a biomarker comprises obtaining the sequence of a subgenomic interval or amplicon of 20 bp or less, 21 bp or less, 22 bp or less, 23 bp or less, 24 bp or less, 25 bp or less, 26 bp or less, 27 bp or less, 28 bp or less, 29 bp or less, 30 bp or less, 31 bp or less, 32 bp or less, 33 bp or less, 34 bp or less, 35 bp or less, 40 bp or less, 45 bp or less, 50 bp or less, 55 bp or less, 60 bp or less, 100 bp or less, 200 bp or less, or 300 bp or less in length, wherein the subgenomic interval or amplicon comprises a biomarker, e.g., a driver gene comprising a driver mutation. E84. The method of any preceding embodiment, wherein detecting the biomarker comprises obtaining the sequence of a subgenomic interval or amplicon at least 6 bp, at least 7 bp, at least 8 bp, at least 9 bp, at least 10 bp, at least 11 bp, at least 12 bp, at least 13 bp, at least 14 bp, at least 15 bp, at least 16 bp, at least 17 bp, at least 18 bp, at least 19 bp, or at least 20 bp in length, wherein the subgenomic interval or amplicon comprises the biomarker, e.g., a driver gene comprising a driver mutation. E85. The method of any preceding embodiment, wherein detecting a biomarker comprises obtaining the sequence of a subgenomic interval or amplicon at least 6 and not more than 300 bp, at least 7 and not more than 200 bp, at least 8 bp and not more than 100 bp, at least 9 bp and not more than 60 bp, at least 10 bp and not more than 55 bp, at least 11 bp and not more than 50 bp, at least 12 bp and not more than 45 bp, at least 13 bp and not more than 40 bp, at least 14 bp and not more than 35 bp, at least 15 bp and not more than 34 bp, at least 14 bp and not more than 33 bp, at least 15 bp and not more than 32 bp, at least 16 bp and not more than 31 bp, at least 17 bp and not more than 30 bp, at least 18 bp and not more than 29 bp, at least 19 bp and not more than 28 bp, or at least 20 bp and not more than 27 bp in length, wherein the subgenomic interval or amplicon comprises a biomarker, e.g., a driver gene comprising a driver mutation. E86. The method of any of the preceding embodiments, wherein detecting a biomarker comprises obtaining the sequence of a subgenomic interval or amplicon between 6 bp and 300 bp, between 7 bp and 200 bp, or between 8 and 100 bp, between 9 bp and 60 bp, between 10 bp and 50 bp, between 15 bp and 40 bp, between 20 bp and 35 bp in length, wherein the subgenomic interval or amplicon comprises the biomarker, e.g., a driver gene comprising a driver mutation. E87. The method of any preceding embodiment, wherein detecting the biomarker comprises obtaining the sequence of a subgenomic interval or amplicon about 33 bp in length, wherein the subgenomic interval or amplicon comprises the biomarker, e.g., a driver gene comprising a driver mutation. E88. b) detecting in the biological sample the level of each of a plurality of, e.g., at least four, protein biomarkers, wherein the level of each of the plurality of protein biomarkers is associated with the presence of one of a plurality of cancers; (optionally) (c) comparing the detected level of each protein biomarker of the plurality of protein biomarkers to a reference level of the protein biomarker; and d) identifying the subject as having one of the plurality of cancers if the presence of one or more genetic biomarkers and the level of one of the plurality of protein biomarkers is detected. 4. The method of any preceding aspect, further comprising: E89. (i) The subject has not yet been determined to have cancer, e.g., a cancer selected from multiple cancers; (ii) the subject has not yet been determined to have cancer cells, e.g., cells of a cancer selected from a plurality of cancers; or (iii) the subject does not exhibit or has never exhibited symptoms associated with cancer, e.g., a cancer selected from a plurality of cancers; 10. The method of any of the preceding aspects. E90. The subject is (i) a pediatric subject or young adult; e.g., between 6 months and 21 years of age; or (ii) an adult, e.g., aged 18 years or older; 10. The method of any of the preceding aspects. E91. The method of any preceding embodiment, wherein the sample comprises a tumor sample, e.g., a biopsy sample (e.g., a liquid biopsy sample (e.g., a circulating tumor DNA sample or a cell-free DNA sample) or a solid tumor biopsy sample); a blood sample (e.g., a circulating tumor DNA sample or a cell-free DNA sample), an apheresis sample, a urine sample, a cyst fluid sample (e.g., a pancreatic cyst fluid sample), a Papanicolaou (Pap) sample, or a fixed tumor sample (e.g., a formalin-fixed or paraffin-embedded sample (FPPE)). E92. One or more, e.g., multiple, genes are selected from one, two, three, or four genes in Tables 60 and 61 of US2019 / 0256924A1, e.g., TIFF2026012754000008.tif99158. E93. The one or more, e.g., multiple, genes are 5, 6, 7, or 8 genes selected from Tables 60 and 61 of US2019 / 0256924A1, e.g., TIFF2026012754000009.tif99158. E94. The method of any preceding embodiment, wherein the one or more, e.g., multiple, genes are selected from NRAS, CTNNB1, PIK3CA, FBXW7, APC, EGFR, BRAF, CDKN2A, PTEN, FGFR2, HRAS, KRAS, AKT1, TP53, PPP2R1A, or GNAS. E95. The method of any preceding embodiment, wherein the one or more, e.g., multiple, biomarkers (e.g., one or more genes) are selected from KRAS, PIK3CA, HRAS, CDKN2A, TP53, AKT1, CTNNB1, APC, EGFR, GNAS, PPP2R1A, BRAF, FBXW7, PTEN, or FGFR2, or a combination thereof, and the cancer is selected from liver cancer, ovarian cancer, esophageal cancer, gastric cancer, pancreatic cancer, colorectal cancer, lung cancer, breast cancer, or prostate cancer. E96. The method of any of the preceding embodiments, wherein the one or more, e.g., multiple, biomarkers (e.g., one or more genes) are selected from KRAS, PIK3CA, HRAS, CDKN2A, TP53, TERT, ERBB2, FGFR3, MET, MLL, or VHL, or a combination thereof, and the cancer is selected from bladder cancer or upper tract urothelial carcinoma (UTUC). E97. The method of any of the preceding embodiments, wherein the one or more, e.g., multiple, biomarkers (e.g., one or more genes) are selected from KRAS, PIK3CA, CDKN2A, TP53, CTNNB1, PPP2R1A, BRAF, PTEN, CSMD3, FAT3, BRCA, or ARID1A, or a combination thereof, and the cancer is ovarian cancer or endometrial cancer. E98. The method of any of the preceding embodiments, wherein the one or more, e.g., multiple, biomarkers (e.g., one or more genes) are selected from KRAS, PIK3CA, CDKN2A, TP53, CTNNB1, GNAS, BRAF, NRAS, VHL, RNF43, or SMAD4, or a combination thereof, and the cancer is pancreatic cancer, e.g., pancreatic ductal adenocarcinoma (PDAC). E99. The method of any preceding embodiment, wherein the one or more, eg, multiple, biomarkers comprise 5, 6, 7, or 8 protein biomarkers. E100. The method of any preceding embodiment, wherein the one or more, e.g., multiple, biomarkers comprise a protein biomarker selected from CA19-9, CEA, HGF, OPN, CA125, prolactin (PRL), TIMP-1, CA15-3, AFP, or MPO. E101. Detecting the presence of one or more genetic biomarkers a. assigning a unique identifier (UID) to each of a plurality of template molecules present in a sample; b. amplifying each template molecule with a unique tag to generate a UID family; and c. Sequencing the amplification products in duplicate 4. The method of any preceding aspect, comprising: E102. The method of any of the preceding embodiments, further comprising detecting the presence of aneuploidy in the sample, e.g., detecting a gain or loss of one or more chromosomes, e.g., using the WALDO method as described in Example 6. E103. The method of embodiment 102, comprising: (i) assessing somatic mutational load; (ii) assessing carcinogen signatures; and / or (iii) detecting microsatellite instability (MSI). E104. The method of embodiment 102 or 103, which can be used to compare two samples, e.g., two unrelated samples, to assess genetic similarity between the samples, or to find somatic mutations within the samples, e.g., within LINE elements in the samples. E105. The method of embodiment 102 or 103, wherein the method provides increased specificity and / or sensitivity for detecting aneuploidy. E106. The method of embodiment 102, wherein the presence of aneuploidy is detected on one or more chromosome arms. E107. The method of any preceding embodiment, further comprising the step of attributing an origin or type of cancer to the cancer corresponding to the values ​​of the genetic markers, protein biomarkers, and / or aneuploidy status. E108. The method of any one of the preceding embodiments, comprising identifying the subject as having cancer or at risk of developing cancer corresponding to the values ​​of the genetic markers, protein biomarkers, and / or aneuploidy status. E109. The method of embodiment E108, further comprising administering to the subject a therapeutic agent to treat the cancer, or selecting a therapeutic agent to treat the cancer in the subject. E110. The method of embodiment E109, wherein the subject is administered the therapeutic agent in combination with one or more additional therapeutic agents. E111. A reaction mixture comprising at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 detection reagents, wherein one detection reagent is: (i) one or more genetic biomarkers referred to in this document; (ii) one or more protein biomarkers referred to in this document; and / or (iii) the copy number or length of the genomic sequence located between at least two terminal repeat elements of a repetitive element family (RE family) referred to in this document, e.g., aneuploidy The reaction mixture mediates a readout that is a value regarding the level or presence of E112. The reaction mixture of embodiment E111, comprising a plurality of detection reagents of (i). E113. The reaction mixture of any of embodiments E111-E112, comprising a plurality of detection reagents of (ii). E114. The reaction mixture of any of embodiments E111-E113, comprising a plurality of detection reagents of (iii). E115. The reaction mixture of any of embodiments E111-E114, comprising a sample derived from a subject, eg, a sample from the subject. E116. Kit, including: (a) at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten detection reagents, wherein one detection reagent is (i) one or more genetic biomarkers referred to in this document; (ii) one or more protein biomarkers referred to in this document; and / or (iii) the copy number or length of the genomic sequence located between at least two terminal repeat elements of a repetitive element family (RE family) referred to in this document, e.g., aneuploidy mediating a readout that is a value regarding the level or presence of detection reagents; and (b) Instructions for use of the kit. E117. The reaction mixture of embodiment E116, comprising a plurality of detection reagents of (i). E118. The reaction mixture of any of embodiments E116-E117, comprising a plurality of detection reagents of (ii). E119. The reaction mixture of any of embodiments E116-E118, comprising a plurality of detection reagents of (iii). E120. The method of any one of embodiments E1 to E110, wherein the aneuploidy status is assessed, eg, measured, using a first primer and a second primer. E121. The method of embodiment E120, wherein the first primer comprises a sequence that is at least 80%, 85%, 90%, 95%, 96%, 96%, 98%, 99%, or 100% identical to SEQ ID NO:1. E122. The method of embodiment E121, wherein the first primer comprises the sequence of SEQ ID NO:1. E123. The method of embodiment E120, wherein the second primer comprises a sequence that is at least 80%, 85%, 90%, 95%, 96%, 96%, 98%, 99%, or 100% identical to SEQ ID NO:10. E124. The method of embodiment E123, wherein the second primer comprises the sequence of SEQ ID NO:10. E125. The method of any one of embodiments E1-E110 or E120-E124, further comprising subjecting the subject to a radiation scan of an organ or internal body region, eg, a PET-CT scan. E126. The method of embodiment 125, wherein the cancer is characterized by radiological scanning of an organ or internal body region. E127. The method of embodiment 125, wherein the cancer is located by radiological scanning of an organ or internal region. E128. The method of any one of embodiments E125-E127, wherein the radiological scan is a PET-CT scan. E129. The method of any one of embodiments E125-E128, wherein a radiological scan is performed after the subject is evaluated for the presence of each of multiple cancers. E130. The method of any one of embodiments E1-E110 or E120-E129, comprising administering to the subject one or more therapeutic interventions (e.g., surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, immunotherapy, targeted therapy, and / or immune checkpoint inhibitors). E131. The method of any one of embodiments E1-E110 or E120-E130, wherein the evaluating comprises evaluating samples from the subject at a single time point or at different time points. E132. The method of any one of embodiments E1-E110 or E120-E131, comprising evaluating one or more samples, eg, multiple sampling, taken from the subject. E133. The method of E132, wherein one or more samples, e.g., multiple sampling samples, are taken annually, e.g., within one year of each other. E134. The method of any of embodiments E1-E110 or E120-E133, wherein the subject is simultaneously evaluated for the presence or absence of each of multiple cancers. E135. The method of any of embodiments E1-E110 or E120-E134, wherein the subject is co-evaluated for the presence or absence of each of multiple cancers. E136. The method of any of embodiments E1-E110 or E120-E135, comprising assessing the presence of each of a plurality of cancers in the subject at one or more time points within a predetermined period of time, e.g., at the same or substantially the same clinical stage, of at least one of the cancers in the subject. E137. The method of any of embodiments E1-E110 or E120-E136, comprising evaluating samples taken from the subject, eg, a single sample or multiple samples. E138. The method of any of embodiments E1-E110 or E120-E137, wherein the joint assessment is performed on a single sample, an aliquot of a single sample, or multiple samples taken within, e.g., 1 hour, 5 hours, 24 hours, or 48 hours of each other. E139. The method of any embodiment E1-E110 or E120-E138, wherein the subject is asymptomatic with respect to cancer. E140. The method of any of embodiments E1-E110 or E120-E139, wherein the subject is asymptomatic with respect to one of the cancers. E141. The method of any of aspects E1-E110 or E120-E140, wherein the subject is not known to have cancer cells and has not been determined to have cancer cells. E142. The method of any of aspects E1-E110 or E120-E141, wherein the subject has not been determined to have or diagnosed with cancer. E143. The method of any of embodiments E1-E110 or E120-E142, wherein the subject has early stage, eg, stage I or stage II, cancer. E144. The method of any of aspects E1-E110 or E120-E143, wherein the subject is in a pre-metastatic state. E145. The method of any of embodiments E1-E110 or E120-E144, wherein the subject does not have detectable metastases. E146. The method of any of aspects E1-E110 or E120-E145, wherein the subject does not exhibit symptoms associated with cancer. E147. The method of any of aspects E1-E110 or E120-E146, wherein the subject does not exhibit one, two or more symptoms clinically associated with cancer. E148. The method of any of embodiments E1-E110 or E120-E147, wherein if the aneuploidy status is positive, the subject has an early stage, e.g., stage I or stage II, cancer, e.g., as set forth in Table 3. E149. The method of any of embodiments E1-E110 or E120-E147, wherein if the aneuploidy status is negative, the subject has an early stage, e.g., stage I or stage II, cancer, e.g., as set forth in Table 3. E150. Methods for detecting aneuploidy in samples containing low input DNA. E151. The method of any of embodiments E1-E110 or E120-E150, wherein the sample comprises between about 0.01 picograms (pg) and 500 pg of DNA. E152. The method of embodiment E151, wherein the sample comprises about 0.01-500 pg, 0.05-400 pg, 0.1-300 pg, 0.5-200 pg, 1-100 pg, 10-90 pg, or 20-50 pg of DNA. E153. The sample contains at least 0.01 pg, at least 0.01 pg, at least 0.1 pg, at least 1 pg, at least 2 pg, at least 3 pg, at least 4 pg, at least 5 pg, at least 6 pg, at least 7 pg, at least 8 pg, at least 9 pg, at least 10 pg, at least 11 pg, at least 12 pg, at least 13 pg, at least 14 pg, at least 15 pg, at least 16 pg, at least 17 pg, at least 18 pg, at least 19 pg, at least 20 pg, at least 21 pg, at least 22 pg, at least 23 pg, at least 24 pg, at least 25 pg, at least 26 pg, at least 27 pg, at least 28 pg, at least 29 pg, at least 30 pg, at least 31 pg, at least 32 pg, at least 33 pg, at least 34 pg, at least 35 pg, at least 36 pg, at least 37 The method of embodiment E151, wherein the method comprises at least 100 pg, at least 200 pg, at least 300 pg, at least 350 pg, at least 400 pg, at least 450 pg, or at least 500 pg of DNA. E154. A method for identifying or distinguishing a sample, for example, using any of the methods disclosed herein. E155. The method of embodiment E154, wherein a sample from a subject, e.g., a first subject, e.g., a first sample, is distinguished from a second sample from a second subject. E156. The method of embodiment E154, wherein the sample is identified as being from the subject based on a polymorphism (eg, a plurality of polymorphisms, eg, common polymorphisms). E157. The method of embodiment E156, wherein the polymorphism, eg, a common polymorphism, is present within a repetitive element, eg, as described herein. E158. The method of embodiment E154, wherein the method disclosed in Example 8 is used to identify and / or differentiate the samples. E159. The method of any of embodiments E1-E110 or E120-E158, which is an in vitro method.

[0065] Unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. Although methods and materials similar or equivalent to those described herein can be used in carrying out embodiments of the present invention, suitable methods and materials are described below. All publications, patent applications, patents, and other references cited herein are incorporated herein by reference in their entirety. In case of conflict, the present specification, including definitions, will control. Additionally, the materials, methods, and examples are illustrative only and are not intended to be limiting.

[0066] [The present invention 1001] 1. A method for testing for the presence of aneuploidy in a mammalian genome, comprising the steps of: (a) amplifying a plurality of chromosomal sequences in a DNA sample to form a plurality of amplicons using primer pairs complementary to the chromosomal sequences, wherein the primer pairs amplify a sufficient number of sequences to allow detection of aneuploidy; (b) determining at least a portion of the nucleic acid sequence of one or more of said plurality of amplicons; (c) mapping the sequenced amplicons to a reference genome; (d) dividing the DNA sample into a plurality of genomic intervals; (e) quantifying a plurality of features for the amplicons mapped to the genomic interval; (f) comparing said plurality of features of amplicons in a first genomic interval with said plurality of features of amplicons in one or more different genomic intervals; and (g) said amplifying step produces a sufficient number of amplicons to detect aneuploidy, thereby testing for the presence of aneuploidy in the genome of said mammal. [The present invention 1002] 1001. The method of claim 1001, wherein the DNA sample comprises a plurality of euploid DNA samples. [The present invention 1003] 1001. The method of claim 1001, wherein the DNA sample comprises a plurality of test DNA samples. [The present invention 1004] The method of the present invention 1003, wherein the test DNA comprises DNA of unknown ploidy. [The present invention 1005] 1001. The method of claim 1001, wherein the DNA sample is derived from plasma. [The present invention 1006] 1001. The method of claim 1001, wherein the DNA sample is derived from serum. [The present invention 1007] 1001. The method of claim 1001, wherein the DNA sample comprises cell-free DNA. [The present invention 1008] 8. The method of any one of claims 1001 to 1007, wherein the DNA sample comprises at least 3 picograms of DNA. [The present invention 1009] The method of any one of claims 1001 to 1008, wherein the mammal is a human. [The present invention 1010] Any of the methods of claims 1001 to 1009, wherein the primer pair comprises a first primer and a second primer selected from Table 1, for example, a first primer comprising SEQ ID NO:1 and a second primer comprising SEQ ID NO:10, or a first primer having at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO:1 and a second primer having at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO:10. [The present invention 1011] The method of any of claims 1001 to 1010, wherein one or more additional primer pairs amplify one or more additional sets of a plurality of chromosomal sequences in the DNA sample of step (a). [The present invention 1012] 1012. The method of any of claims 1001 to 1011, wherein the amplicon comprises one or more repetitive elements as set forth in Table 1. [The present invention 1013] The method of claim 1012, wherein the amplicon comprises a unique short interspersed element (SINE). [The present invention 1014] The method of any one of claims 1001 to 1013, wherein the average length of the amplicons is 100 base pairs or less. [The present invention 1015] 15. The method of any one of claims 1001 to 1014, wherein the amplicon comprises one or more long amplicons, the average length of which is 1000 base pairs or more. [The present invention 1016] The method of any one of claims 1 to 5, wherein the long amplicon comprises DNA from contaminating cells. [The present invention 1017] The method of claim 1016, wherein the contaminating cells are leukocytes. [The present invention 1018] The method of any of claims 1001 to 1017, wherein the plurality of amplicons comprises sequences on multiple, e.g., two or more, different chromosomes. [The present invention 1019] 1018. The method of any of claims 1001 to 1018, wherein the genomic interval comprises from about 100 nucleotides to about 125,000,000 nucleotides. [The present invention 1020] 1020. The method of any of claims 1001 to 1019, wherein the step of quantifying amplicons mapped to a genomic interval comprises identifying a plurality of genomic intervals having one or more shared amplicon features. [The present invention 1021] The method of claim 1020, wherein the shared amplicon feature is the number of mapped amplicons. [The present invention 1022] The method of claim 1020, wherein the shared amplicon feature is the average length of the mapped amplicons. [The present invention 1023] The method of any of claims 1020 to 1022, wherein a plurality of genomic intervals having a shared amplicon characteristic are grouped into one or more clusters. [The present invention 1024] 1023. The method of claim 1023, wherein each cluster comprises about 200 genomic intervals. [The present invention 1025] The method of the present invention 1023, wherein the clusters include predefined clusters. [The present invention 1026] Any of the methods of claims 1001 to 1025, wherein comparing the genomic intervals further comprises matching one or more genomic intervals from the test sample to defined clusters. [The present invention 1027] The method of claim 1026, wherein matching genomic intervals from the test sample to the defined clusters further comprises identifying one or more genomic intervals having shared amplicon features outside a predetermined significance threshold for the defined clusters. [The present invention 1028] The method of any of claims 1001 to 1027, wherein the step of testing for the presence of aneuploidy comprises supervised machine learning. [The present invention 1029] The method of the present invention 1028, wherein the supervised machine learning uses a support vector machine model. [The present invention 1030] A primer pair for amplifying multiple amplicons from a DNA sample, comprising a first primer comprising a sequence at least 80% identical to SEQ ID NO:1 and a second primer comprising a sequence at least 80% identical to SEQ ID NO:10. [The present invention 1031] 1030. The primer pair of the present invention, wherein the sequence of the first primer is at least 90% identical to SEQ ID NO:1. [The present invention 1032] 1030. The primer pair of the present invention, wherein the sequence of the first primer is at least 95% identical to SEQ ID NO:1. [The present invention 1033] A primer pair of the present invention 1030, wherein the sequence of the first primer is or comprises a sequence that is 100% identical to SEQ ID NO:1, and / or the sequence of the second primer is or comprises a sequence that is 100% identical to SEQ ID NO:2. [The present invention 1034] The primer pair of any one of 1030 to 1032, wherein the sequence of the second primer is at least 90% identical to SEQ ID NO:10. [This invention 1035] The primer pair of any one of 1030 to 1032, wherein the sequence of the second primer is at least 95% identical to SEQ ID NO:10. [The present invention 1036] The primer pair of any one of claims 1030 to 1032, wherein the sequence of the second primer is 100% identical to SEQ ID NO:10 or comprises a sequence 100% identical to SEQ ID NO:10. [This invention 1037] A kit for amplifying multiple amplicons from a DNA sample, comprising a primer pair, wherein a first primer of the primer pair comprises SEQ ID NO:1 and a second primer of the primer pair comprises SEQ ID NO:10. [The present invention 1038] 1029. The method of any of claims 1001 to 1029, wherein at least 10,000 amplicons are formed in the amplifying step. [This invention 1039] The method of any of claims 1001 to 1037, wherein at least 20,000 amplicons are formed in the amplifying step. [The present invention 1040] The method of any of claims 1001 to 1037, wherein at least 50,000 amplicons are formed in the amplifying step. [This invention 1041] The method of any of claims 1001 to 1037, wherein at least 100,000 amplicons are formed in the amplifying step. [The present invention 1042] 1. A method for assessing a subject for the presence of, or risk of developing, each of a plurality of cancers in the subject, comprising: (i) obtaining a value for the presence of one or more mutations in each of one or more driver genes, each of which is associated with the presence or risk of one of the plurality of cancers; (ii) obtaining a value for the level of each of a plurality of protein biomarkers, wherein the level of each of the plurality of protein biomarkers is associated with the presence or risk of one of the plurality of cancers; (iii) obtaining a value for aneuploidy, the value for aneuploidy being a function of the copy number or length of the genomic sequence located between at least two terminal repeat elements of a repetitive element family (RE family), the RE family comprising: (a) RE families other than long interspersed nucleotide sequences (LINEs), (b) an RE family that, when amplified with a primer portion complementary to its own terminal repeat element, results in an amplicon having an average length of less than X nt, where X is 100, 105, or 110; (c) an RE family that is less than about 700 bp in length; or (d) RE families present in at least 100 copies / genome Includes; the aneuploidy is associated with the presence or risk of one of the cancers; thereby assessing the subject for the presence of, or risk of developing, any of the plurality of cancers. Process. [This invention 1043] The method of claim 1042, wherein one of (i), (ii) and (iii) is obtained directly. [This invention 1044] The method of the present invention 1042, wherein (i) and (ii) are obtained directly. [This invention 1045] The method of the present invention 1042, wherein (i) and (iii) are obtained directly. [The present invention 1046] The method of the present invention 1042, wherein (i) and (ii) are obtained directly. [This invention 1047] The method of claim 1042, wherein all of (i), (ii), and (iii) are obtained directly. [This invention 1048] The method of claim 1042, wherein one of (i), (ii) and (iii) is obtained indirectly. [This invention 1049] The method of claim 1042, wherein (i) and (ii) are obtained indirectly. [The present invention 1050] The method of claim 1042, wherein (i) and (iii) are obtained indirectly. [This invention 1051] The method of claim 1042, wherein (i) and (ii) are obtained indirectly. [This invention 1052] The method of claim 1042, wherein all of (i), (ii), and (iii) are obtained indirectly. [This invention 1053] (1) sequencing one or more subgenomic intervals or amplicons containing genetic biomarkers; (2) analyzing one or more genomic sequences for aneuploidy; and / or (3) contacting the protein biomarker with a detection reagent Any of the methods of the present invention 1042 to 1052, comprising: [This invention 1054] 1054. The method of any of claims 1042 to 1053, wherein the value of aneuploidy is a function of the copy number of the genomic sequence located between at least two terminal repeat elements of the RE family. [This invention 1055] 1054. The method of any of claims 1042 to 1054, wherein the value of aneuploidy is a function of the length of the genomic sequence located between at least two terminal repeat elements of a repetitive element family (RE family). [This invention 1056] (i) obtaining the sequence of a subgenomic interval of cell-free DNA from a sample; and (ii) obtaining leukocyte parameters of leukocyte DNA from the sample; Any of the methods of claims 1042 to 1055, further comprising: [This invention 1057] 1057. The method of any one of claims 1055 to 1056, wherein the white blood cell parameter comprises the sequence of said subgenomic interval. [This invention 1058] The method of any one of claims 1055 to 1056, further comprising the step of comparing (i) to (ii) to assess genomic events observed in the cell-free DNA subgenomic interval or cell-free DNA aneuploidy analysis sample. [This invention 1059] The method of any one of claims 10 to 58, wherein the genomic event comprises a mutation. [The present invention 1060] Any of the methods of claims 1042 to 1059, wherein the specificity of detecting one cancer among a plurality of cancers involving (i), (ii) and (iii) is substantially the same as the specificity of detecting said one cancer among said plurality of cancers involving (i); (ii); (iii); (i) and (ii); (i) and (iii); or (ii) and (iii). [This invention 1061] Any of the methods of claims 1042 to 1059, wherein the specificity of detecting one of the plurality of cancers involving (i), (ii) and (iii) is not substantially lower than the specificity of detecting the one of the plurality of cancers involving (i); (ii); (iii); (i) and (ii); (i) and (iii); or (ii) and (iii). [This invention 1062] Any of the methods of present inventions 1042 to 1061, wherein the detection sensitivity of one of multiple cancers involving (i), (ii), and (iii) is higher than the detection sensitivity of the one of the multiple cancers involving (i); (ii); (iii); (i) and (ii); (i) and (iii); or (ii) and (iii). [This invention 1063] The method of the present invention 1062, wherein the sensitivity of detecting one cancer among a plurality of cancers involving (i), (ii), and (iii) is about 1.1, 1.2, 1.3, 1.4, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5 or 10 times higher than the sensitivity of detecting said one cancer among said plurality of cancers involving (i); (ii); (iii); (i) and (ii); (i) and (iii); or (ii) and (iii). [This invention 1064] The method of any of claims 1042 to 1063, wherein (i), (ii), and (iii) result in increased detection sensitivity at a certain specificity. [This invention 1065] The method of the present invention 1064, wherein the increase in detection sensitivity is about a 1.1, 1.2, 1.3, 1.4, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5 or 10 fold increase in the specified specificity. [The present invention 1066] The method of claim 1064 or 1065, wherein the specificity is a predetermined specificity. [This invention 1067] The method of the present invention 1066, wherein the predetermined specificity is at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% specificity. [The present invention 1068] The method of any one of claims 1062 to 1067, wherein the specificity of the detection of one of the plurality of cancers is not affected by an increase in the sensitivity of the detection of the one of the plurality of cancers. [The present invention 1069] Any of the methods of claims 1062 to 1067, wherein the increase in the sensitivity of detection of one of the plurality of cancers does not decrease or substantially decrease the specificity of detection of said one of the plurality of cancers. [The present invention 1070] 1069. The method of claim 1068 or 1069, wherein the specificity for detecting one of the cancers plateaus. [This invention 1071] 1070. The method of any of claims 1042 to 1070, wherein the step of obtaining a value relating to the presence of one or more mutations comprises detecting one or more mutations in one or more driver genes. [This invention 1072] 1071. The method of claim 1071, wherein the one or more mutations comprise one or more driver gene mutations. [This invention 1073] 1072. The method of any of claims 1042 to 1072, wherein the one or more driver genes are selected from NRAS, CTNNB1, PIK3CA, FBXW7, APC, EGFR, BRAF, CDKN2A, PTEN, FGFR2, HRAS, KRAS, AKT1, TP53, PPP2R1A or GNAS. [This invention 1074] Any of the methods of claims 1042 to 1073, wherein the presence of one or more mutations is assessed in at least four driver genes selected from NRAS, CTNNB1, PIK3CA, FBXW7, APC, EGFR, BRAF, CDKN2A, PTEN, FGFR2, HRAS, KRAS, AKT1, TP53, PPP2R1A or GNAS. [This invention 1075] Any of the methods of claims 1042 to 1074, wherein the presence of one or more mutations is assessed in all of the following 16 driver genes: NRAS, CTNNB1, PIK3CA, FBXW7, APC, EGFR, BRAF, CDKN2A, PTEN, FGFR2, HRAS, KRAS, AKT1, TP53, PPP2R1A and GNAS. [This invention 1076] Any of the methods of claims 1042 to 1075, wherein the step of obtaining values ​​for each of a plurality of protein biomarkers comprises detecting each of the plurality of protein biomarkers selected from, for example, CA19-9, CEA, HGF, OPN, CA125, prolactin (PRL), TIMP-1, CA15-3, AFP or MPO. [This invention 1077] The method of any one of claims 1076 to 1076, wherein the plurality of protein biomarkers comprises at least four protein biomarkers. [This invention 1078] The method of any of claims 1042 to 1077, wherein the step of obtaining a value related to aneuploidy comprises detecting aneuploidy. [This invention 1079] The method of any of claims 1042 to 1078, wherein the plurality of cancers comprises at least four cancers. [The present invention 1080] The method of any of claims 1042 to 1079, further comprising the step of subjecting the subject to a radiation scan of an organ or internal body region, such as a PET-CT scan. [This invention 1081] The method of claim 1080, wherein the cancer is characterized by radiological scanning of an organ or internal body region. [This invention 1082] The method of the present invention 1080, wherein the cancer is located by radiological scanning of an organ or region within the body. [This invention 1083] 1083. The method of any one of claims 1080 to 1082, wherein the radiation scan is a PET-CT scan. [This invention 1084] The method of any of claims 1080 to 1083, wherein the radiological scanning is performed after the subject has been evaluated for the presence of each of a plurality of cancers. [This invention 1085] Any of the methods of claims 1042 to 1084, comprising the step of administering to the subject one or more therapeutic interventions (e.g., surgery, postoperative adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, immunotherapy, targeted therapy and / or immune checkpoint inhibitors). [The present invention 1086] The method of any of claims 1042 to 1085, wherein the subject is asymptomatic with respect to the cancer. [This invention 1087] The method of any of claims 1042 to 1085, wherein the subject is asymptomatic with respect to one of the cancers. [This invention 1088] The method of any one of claims 1042 to 1085, wherein the subject is not known to have cancer cells and has not been determined to have cancer cells. [This invention 1089] The method of any one of claims 1042 to 1085, wherein the subject has never been determined to have cancer or diagnosed with cancer. [The present invention 1090] The method of any of claims 1042 to 1085, wherein the subject has early stage, eg, stage I or stage II, cancer. [This invention 1091] The kit includes: (a) at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten detection reagents, wherein one detection reagent is (i) one or more genetic biomarkers referred to in this document; (ii) one or more protein biomarkers referred to in this document; and / or (iii) The copy number or length of the genomic sequence located between at least two terminal repeat elements of a repetitive element family (RE family) referred to in this document. mediating a readout that is a value regarding the level or presence of detection reagents; and (b) Instructions for using the kit. [This invention 1092] The kit of the present invention 1091, wherein the detection reagent mediates a readout that is a value related to the level or presence of aneuploidy in a genomic sequence. [This invention 1093] A method of testing a mammal for the presence of cancer, comprising the steps of: (a) amplifying a plurality of chromosomal sequences in a DNA sample using primer pairs complementary to the chromosomal sequences to form a plurality of amplicons; (b) determining at least a portion of the nucleic acid sequence of one or more of said plurality of amplicons; (c) mapping the sequenced amplicons to a reference genome; (d) dividing the DNA sample into a plurality of genomic intervals; (e) quantifying a plurality of features for the amplicons mapped to the genomic interval; (f) comparing said plurality of features of amplicons in a first genomic interval with said plurality of features of amplicons in one or more different genomic intervals; and (g) determining that cancer is present in the mammal if the plurality of features of amplicons in a first genomic interval differs from the plurality of features of amplicons in one or more different genomic intervals. [This invention 1094] The method of claim 1093, wherein the amplifying step forms at least 100,000 amplicons. [This invention 1095] The method of any one of claims 1093 to 1094, wherein the cancer is stage I cancer. [This invention 1096] 6. The method of any one of claims 1093 to 1095, wherein the cancer is liver cancer, ovarian cancer, esophageal cancer, gastric cancer, pancreatic cancer, colorectal cancer, lung cancer, breast cancer or prostate cancer. [This invention 1097] 1097. The method of any of claims 1093 to 1096, further comprising determining that aneuploidy is present if a plurality of features of the amplicon in the first genomic interval differs from said plurality of features of the amplicon in one or more different genomic intervals. [This invention 1098] A method of detecting aneuploidy in a sample containing low input DNA using any of the methods disclosed herein. [This invention 1099] 1098. The method of claim 1098, wherein the sample comprises between about 0.01 picograms (pg) and 500 pg of DNA. [The present invention 1100] The method of any one of claims 1098 to 1099, wherein the sample is a biological sample from a subject. [The present invention 1101] The method of any of claims 1098 to 1100, wherein the sample comprises a liquid sample, a blood sample, a cell-free DNA sample (e.g., a circulating tumor DNA sample), a plasma sample, a serum sample; or a tissue sample. [The present invention 1102] The method of any of claims 1098 to 1100, wherein the sample, e.g., a biological sample, comprises cells (e.g., normal cells or cancer cells) and cell-free DNA. [The present invention 1103] A method of identifying or distinguishing a sample using any of the methods disclosed herein. [The present invention 1104] The method of claim 1103, wherein a sample, e.g., a first sample, derived from a subject, e.g., a first subject, is distinguished from a second sample derived from a second subject. [This invention 1105] The method of claim 1103, wherein the sample is identified as being from the subject based on a polymorphism (e.g., a plurality of polymorphisms, e.g., common polymorphisms). [The present invention 1106] The method of claim 1105, wherein the polymorphism, eg, a common polymorphism, is present within a repetitive element, eg, as described herein. [This invention 1107] Any of the methods of 1001 to 1090 or 1093 to 1106 of the present invention, which is an in vitro method. The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the invention will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0067] [Figure 1A] Figure 1A shows the distribution of amplicon sizes when a single primer pair is used to amplify a repetitive element (see, e.g., Table 1 for a list of repetitive elements). The amplicon sizes shown in Figure 1A include the number of bases in the primer. [Figure 1B]Figure 1B shows the distribution of amplicon sizes when a single primer pair is used to amplify a repetitive element (see, e.g., Table 1 for a list of repetitive elements). The amplicon sizes shown in Figure 1B do not include the number of bases in the primer. [Figure 1C] Figure 1C shows the distribution of observed amplicon numbers in cell-free DNA from 2231 plasma samples. [Figure 2A] FIG. 2A. An exemplary overview of one embodiment of the workflow described herein. [Figure 2B] FIG. 2B is an exemplary overview of one embodiment of the Repetitive Element Aneuploidy Sequencing System (RealSeqS). [Figure 3] Figure 3 shows the sensitivity (at 99% specificity) of aneuploidy versus mutation in different cancer types. The percent of aneuploidy detected in each cancer type is shown on the Y-axis. [Figure 4] Figure 4 shows that aneuploidy demonstrates the sensitivity of aneuploidy compared to other cancer biomarkers. The percent of cancers detected (sensitivity) is shown on the Y-axis. [Figure 5] FIG. 5 shows pseudocode for creating a composite with multiple arm modifications. [Figure 6] Figure 6 shows an estimate of the relationship between reads and DNA concentration. [Figure 7A] Figure 7A shows a comparison of the sensitivity of different multi-analyte tests for cancer detection. Three different multi-analyte tests were evaluated for the detection sensitivity of the eight cancers indicated. The three tests were: (1) assessment of aneuploidy status, somatic mutation analysis, and protein biomarkers; (2) assessment of aneuploidy status and somatic mutation analysis; and (3) assessment of aneuploidy status and protein biomarkers. [Figure 7B]Figure 7B shows the sensitivity of a test incorporating aneuploidy, mutations, and abnormally high levels of eight proteins compared to the sensitivity of tests comparing aneuploidy + protein alone or mutations vs. protein alone. All sensitivities were calculated at an aggregate specificity of 99% (i.e., only 1% of plasma samples were positive for aneuploidy, mutation, or protein in a test incorporating aneuploidy, mutation, and protein using 10-fold cross-validation with 10 iterations). [Figure 8] Figure 8 shows the true positive rate (sensitivity) on the y-axis and the false positive rate for cancer detection using various tests. Tests include: (1) aneuploidy status; somatic mutations; and protein biomarkers; (2) aneuploidy status and protein biomarkers; (3) somatic mutations and protein biomarkers; (4) aneuploidy status and somatic mutations; (5) aneuploidy status; and (6) somatic mutations. True positive rates (sensitivity) were calculated using a threshold of 99% specificity. [Figure 9] Figure 9 shows the sensitivity of cancer detection with aneuploidy alone (at 98% or 99% specificity) compared to the sensitivity with aneuploidy and protein biomarkers (at 95% specificity) at different stages of cancer. [Figure 10] FIG. 10 shows aneuploidy (at 99% specificity) in different stages of cancer. [Figure 11] FIG. 11 shows aneuploidy (at 99% specificity) in different types of cancer. [Figure 12] FIG. 12 shows the sensitivity of combining aneuploidy (at 99% specificity) with the detection of protein biomarkers. [Figure 13] FIG. 13 shows pseudocode for generating in silico trisomy and monosomy samples used for comparison of whole genome sequencing, FAST-SeqS, and Real-SeqS. [Figure 14] Figure 14 shows the pseudocode for creating in silico simulation samples with multiple arm alterations used for the genome-wide aneuploidy SVM training set. [Figure 15A] Figures 15A-15C show the detection of aneuploidy using next-generation sequencing technology. Sensitivity was calculated at 99% specificity. Error bars represent 95% confidence intervals. Figure 15A: Comparison of sensitivity for monosomy and trisomy in non-acrocentric chromosome arms in all 39 cases at 5% cellularity. [Figure 15B] FIG. 15B Comparison of sensitivity for a 1.5 Mb DiGeorge deletion in 22q at 5% cellularity. [Figure 15C] FIG. 15C Comparison of sensitivity to 20 copies of ERBB2 focal amplification at 1% cell rate. [Figure 16] Figures 16A-16B show examples of plasma samples with focal deletions or amplifications. Figure 16A shows RealSeqS data for a plasma sample from a normal individual with a ~3 Mb deletion on chromosome 22 characteristic of DiGeorge syndrome. Note that many patients with microdeletions at this locus have mild signs and symptoms and are not clinically detectable. Figure 16B shows RealSeqS data for a typical plasma sample from a normal individual without a deletion at the DiGeorge locus. [Figure 17] Figures 17A-17B show examples of plasma samples with focal deletions or amplifications. Figure 17A shows RealSeqS data for a plasma sample from a patient with cancer showing a 2.5 MB focal amplification encompassing the ERBB2 locus on chromosome 17q. Figure 17B shows RealSeqS data for a typical plasma sample from a normal individual showing no amplification of the ERBB2 locus. [Figure 18] Figure 18 shows the sensitivity of RealSeqS in plasma samples with various amounts of tumor-derived DNA, estimated by the mutant allele frequency (MAF) of the driver mutation present in the plasma sample. [Figure 19]Figures 19A-19B show cancer detection in liquid biopsies from samples with eight different types of non-metastatic cancer. Sensitivity was calculated at 99% specificity during cross-validation. Error bars represent 95% confidence intervals. Figure 19A shows a comparison of aneuploidy status assessed by RealSeqS with somatic mutation status for tumor type. Figure 19B shows a comparison of aneuploidy status calculated by RealSeqS with somatic mutation status for cancer stage. DETAILED DESCRIPTION OF THE INVENTION

[0068] Detailed Description definition The term "driver gene mutation" or "driver mutation," as used herein, refers to a mutation that (i) occurs in a driver gene; and (ii) confers a growth advantage to the cell in which the mutation occurs. The growth advantage of a cell can include: a) an increased rate of cell division in cells carrying the driver gene mutation, for example compared to a reference cell, for example a similar cell, for example a similar cell adjacent to the cell carrying the mutation, for example compared to a cell of the same type that does not carry the driver gene mutation; b) an increased rate of clonal expansion in cells with a driver gene mutation, e.g., an increased rate of clonal expansion when compared to a reference cell, e.g., an originally similar cell, e.g., an originally similar cell adjacent to a cell with a mutation, e.g., an increased rate of clonal expansion when compared to a cell of the same type that does not have the driver mutation; c) an increase in the number of cells that are progeny, e.g., daughter cells, of a cell that has the driver gene mutation, e.g., an increase in the number of progeny cells compared to the number of progeny cells expected if the cell did not have the driver gene mutation; d) an increased ability to form tumors or promote tumor growth, e.g., tumor progression, e.g., compared to a reference cell, e.g., an otherwise similar cell that does not have the driver gene mutation; or e) the presence or occurrence in a second or subsequent site or location within the subject may be mentioned.

[0069] In one aspect, the driver gene mutation confers a growth advantage, e.g., an increase in the difference between cell birth and cell death, of 0.1 to 5%, e.g., 0.1 to 4.5%, 0.1 to 4%, 0.1 to 3.5%, 0.1 to 3%, 0.1 to 2.5%, 0.1 to 2%, 0.1 to 1.5%, 0.1 to 1%, 0.1 to 0.5%, 0.5 to 5%, 1 to 5%, 1.5 to 5%, 2 to 5%, 2.5 to 5%, 3 to 5%, 3.5 to 5%, 4 to 5%, 4.5 to 5%, 0.5 to 4.5%, 1 to 4%, 1.5 to 3.5%, or 2 to 3%. In one embodiment, the driver gene mutation confers a growth advantage, e.g., an increase in the difference between cell birth and cell death, of at least 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4% or 4.5%, e.g., about 0.4%. In one embodiment, the driver gene mutation confers proliferative potential to the cell in which the mutation occurs, e.g., enables expansion, e.g., clonal expansion, of the cell.

[0070] In some aspects, driver gene mutations may be causally linked to cancer progression.

[0071] In one embodiment, the driver gene mutation affects, e.g., alters, the regulation, expression, or function of a protein-coding gene. In one embodiment, the driver gene mutation affects, e.g., alters, the function of a non-coding region, e.g., a non-protein-coding region. In one embodiment, the driver gene mutation includes a translocation, a deletion (e.g., a homozygous deletion), an insertion (e.g., an intragenic insertion), a small insertion-deletion (indel), a single base substitution (e.g., a synonymous mutation, a non-synonymous mutation, a nonsense mutation, or a frameshift mutation), a copy number variation (CNV) (e.g., an amplification), or a single nucleotide variation (SNV) (e.g., a single nucleotide polymorphism (SNP)). Exemplary driver mutations are disclosed in Tables 60 and 61 of US2019 / 0256924A1.

[0072] In some embodiments, the presence of a driver gene mutation in a cell may alter (e.g., increase or decrease) the expression of a gene product in the cell. In some embodiments, the presence of a driver gene mutation in a cell may alter the function of a gene product. In some cases, the presence of a driver gene mutation in a cell may confer a growth advantage to the cell. For example, the presence of a driver gene mutation in a cell may cause an increased growth rate (e.g., compared to a reference cell). For example, the presence of a driver gene mutation in a cell may cause an increased clonal expansion rate of a cell having the driver gene mutation (e.g., compared to a reference cell). For example, the presence of a driver gene mutation in a cell may cause an increased number of progeny cells derived from a cell having the driver gene mutation (e.g., compared to a reference cell). For example, the presence of a driver gene mutation in a cell may cause an increased ability of the cell to form a tumor (e.g., compared to a reference cell). In some cases, the growth advantage may be measured as an increased difference between cell development (e.g., new cell formation) and cell death. For example, the presence of a driver gene mutation in a cell can confer a growth advantage to the cell of at least about 0.1% (e.g., about 0.2%, about 0.3%, about 0.4%, about 0.5%, about 0.6%, about 0.7%, about 0.8%, about 0.9%, about 1%, about 1.5%, about 2%, about 2.5%, about 3%, about 3.5%, about 4%, about 4.5% or more). For example, the presence of a driver gene mutation in a cell can confer a growth advantage to the cell of about 0.1% to about 5% (e.g., about 0.1 to about 5%, about 0.1 to about 4.5%, about 0.1 to about 4%, about 0.1 to about 3.5%, about 0.1 to about 3%, about 0.1 to about 2.5%, about 0.1 to about 2%, about 0.1 to about 1.5%, about 0.1 to about 1%, about 0.1 to about 0.5%, about 0.5 to about 5%, about 1 to about 5%, about 1.5 to about 5%, about 2 to about 5%, about 2.5 to about 5%, about 3 to about 5%, about 3.5 to about 5%, about 4 to about 5%, about 4.5 to about 5%, about 0.5 to about 4.5%, about 1 to about 4%, about 1.5 to about 3.5%, or about 2 to about 3%).

[0073] In some cases, a driver gene may contain more than one (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more) driver gene mutations. In some cases, a driver gene that contains one or more driver gene mutations may also contain one or more additional mutations (e.g., passenger gene mutations (somatic mutations that are not driver mutations)).

[0074] The term "driver gene," as used herein, refers to a gene that contains a driver gene mutation. In one embodiment, a driver gene is a gene in which one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more) acquired mutations, e.g., driver gene mutations, may be causally associated with cancer progression. In one embodiment, a driver gene modulates one or more cellular processes, e.g., cell fate determination, cell survival, and genome maintenance. A driver gene may be associated with (e.g., modulate) one or more signal transduction pathways. Examples of signal transduction pathways include, but are not limited to, the TGF-beta pathway, the MAPK pathway, the STAT pathway, the PI3K pathway, the RAS pathway, the cell cycle pathway, the apoptosis pathway, the NOTCH pathway, the Hedgehog (HH) pathway, the APC pathway, the chromatin modification pathway, the transcriptional regulation pathway, and the DNA damage control pathway. Examples of driver genes include, but are not limited to, TIFF2026012754000010.tif99160. Exemplary driver genes include oncogenes and tumor suppressors. In one embodiment, the driver gene has one or more driver gene mutations, e.g., as described herein. In one embodiment, the driver gene is a gene listed in Table 60 or 61 of US2019 / 0256924A1. In one embodiment, the driver gene is a gene that modulates one or more cellular processes, e.g., cell fate determination, cell survival, and genome maintenance, listed in Table 60 or 61 of US2019 / 0256924A1. In one embodiment, the driver gene is a gene that modulates one or more pathways listed in Table 60 or 61 of US2019 / 0256924A1. In one embodiment, the driver gene is a gene that modulates one or more signaling pathways listed in Table 62 of US2019 / 0256924A1.

[0075] In one embodiment, a driver gene comprises more than one driver mutation, and a first driver gene mutation confers a selective growth advantage to cells in which the mutation occurs. In one embodiment, subsequent mutations in the driver gene, e.g., second, third, fourth, fifth, or more mutations, e.g., driver mutations, confer growth potential to the cells in which the mutation occurs, e.g., allowing for cell expansion, e.g., clonal expansion. In one embodiment, a driver gene has one or more passenger gene mutations, e.g., somatic mutations that arise during cancer development but are not driver mutations. In one embodiment, a driver gene can be present, e.g., expressed, in any cell type, e.g., a cell type derived from any one of the three germ layers: ectoderm, endoderm, or mesoderm. In one embodiment, a driver gene is present, e.g., expressed, in somatic cells. In one embodiment, a driver gene is present, e.g., expressed, in embryonic cells. In one embodiment, a driver gene can be present in a majority of cancers, e.g., more than 5% of cancers. In one embodiment, a driver gene can be present in a minority of cancers, e.g., less than 5% of cancers. In one embodiment, driver gene has non-random and / or recurrent mutation pattern, that is, the location of driver mutation in driver gene is the same in different types of cancer.Exemplary recurrent driver gene mutations include the substrate binding site in IDH1 gene, for example, the mutation in codon 132, and the mutation in helical domain or kinase domain in PIK3CA gene, as shown in Vogelstein et al (2013) Science 339:1546-1558.

[0076] In one embodiment, the driver gene having driver gene mutation is a cancer gene.In one embodiment, the cancer gene is a gene having a cancer gene score of at least 20%, for example, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99% or at least 100%.In one embodiment, the cancer gene score is defined as the number of mutations, for example, clustered mutations (for example, missense mutations or identical in-frame insertions or deletions at the same amino acid), divided by the total number of mutations.In one embodiment, the driver gene having amplification, for example, as described herein, is a cancer gene.In one embodiment, the driver gene having driver gene mutation is a tumor suppressor gene (TSG). In one embodiment, the tumor suppressor gene is a gene that has a tumor suppressor gene score of at least 20%, for example, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99% or at least 100%. In one embodiment, the tumor suppressor gene score is defined as the number of inactivating mutations divided by the total number of mutations. In one embodiment, the driver gene with deletion, for example as described herein, is a tumor suppressor gene.

[0077] The phrase "repetitive element family" or "RE family," as used herein, refers to a family of repeat DNA elements (also known as repeated DNA elements or repeat units or DNA repeats) present in the genome of an organism. DNA repeat elements can be scattered throughout the genome of an organism or can be present on selected chromosomes. An RE family can contain one or more repeat DNA elements. Exemplary RE families in the human genome include interspersed repeats (e.g., long interspersed nucleotide sequences (LINEs); short interspersed nucleotide sequences (SINEs)); and tandem repeats (e.g., microsatellites, minisatellites, satellite DNA, or multicopy genes (e.g., ribosomal RNAs)). In some embodiments, an RE family contains one or more repeat elements, e.g., SINEs, listed in Table 1.

[0078] "Obtain" or "obtaining," as used herein, refers to obtaining possession of a physical entity or a value, e.g., a numerical value, by "directly obtaining" or "indirectly obtaining" the physical entity or value. "Directly obtaining," as used herein, refers to performing a process (e.g., performing a synthetic or analytical method) to obtain the physical entity or value. "Indirectly obtaining," as used herein, refers to receiving the physical entity or value from another party or source (e.g., a third-party laboratory that directly obtained the physical entity or value). Directly obtaining a physical entity includes performing a process that involves a physical change of a physical entity, e.g., a starting material. Obtaining a value directly includes performing a process that involves a physical change of a sample or another substance, such as performing an analytical process (sometimes referred to herein as "physical analysis") that involves a physical change of a substance, such as a sample, analyte, or reagent; performing an analytical method, such as a method that involves one or more of the following: separating or purifying a substance, such as an analyte, or a fragment or other derivative thereof, from another substance; combining an analyte, or a fragment or other derivative thereof, with another substance, such as a buffer, solvent, or reactant; or altering the structure of an analyte, or a fragment or other derivative thereof.

[0079] As used herein, a "biological sample," "sample," "patient sample," or "specimen" each refer to a sample taken from a subject or patient. The source of the sample can be biopsy tissue (e.g., liquid biopsy tissue), aspirate; blood or any blood component; or bodily fluid (e.g., cerebrospinal fluid, amniotic fluid, peritoneal fluid, or interstitial fluid). The sample can contain cells (e.g., any cells from the human body, e.g., normal cells and / or cancer cells) and / or cell-free DNA, e.g., circulating tumor DNA or circulating DNA from normal cells. In one embodiment, the sample, e.g., a tumor sample, contains tissue or cells from a surgical margin. In another embodiment, the sample, e.g., a tumor sample, contains one or more circulating tumor cells (CTCs) (e.g., CTCs obtained from a blood sample).

[0080] As used herein, the term "sensitivity" refers to the ability of a method to detect or identify the presence of a disease in a subject. For example, when used in reference to any of the various methods described herein that can detect the presence of cancer in a subject, high sensitivity means that the method correctly identifies the presence of cancer in a subject a high percentage of the time. For example, a method described herein is said to have a sensitivity of 95% if it correctly detects the presence of cancer in a subject 95% of the times the method is performed. In some embodiments, a method described herein that can detect the presence of cancer in a subject exhibits a sensitivity of at least 70% (e.g., about 70%, about 72%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 99.5%, or about 100%). In some embodiments, the methods provided herein that involve detecting the presence of one or more members of more than one class of biomarkers (e.g., genetic and / or protein biomarkers) exhibit greater sensitivity than methods that involve detecting the presence of one or more members of only one class of biomarkers.

[0081] In some embodiments, sensitivity refers to the measure of the ability of a method to detect sequence variants in a heterogeneous sequence population.Assuming that a sequence variant exists in at least F% of the sequences in a sample, and the method can detect the sequence S% of the times with C% confidence, the method has a sensitivity of S% for F% variants.As an example, assuming that a variant sequence exists in at least 5% of the sequences in a sample, and the method can detect the sequence 9 times out of 10 with 99% confidence, the method has a sensitivity of 90% for 5% variants (F=5%; C=99%; S=90%).Exemplary sensitivities include S=90%, 95%, 99%, and 99.9% for F=0.5%, 1%, 5%, 10%, 20%, 50%, and 100% sequence variants, with C=90%, 95%, 99%, and 99.9% confidence levels.

[0082] As discussed above, in various embodiments, sensitivity is the ability to attribute all first state samples to the same first state, or in other words, the ability of the test method to find or identify all first state samples. (Sensitivity does not address the tendency of the method to misattribute first state samples as second state samples.) In one embodiment, the first state is negative, and sensitivity is the ability to identify all negative samples. In one embodiment, the first state is positive, and sensitivity is the ability to identify all positive samples.

[0083] As used herein, the term "specificity" refers to the ability of a method to detect the presence of disease in a subject (e.g., the specificity of a method can be described as the ability to identify true positives in a subject compared to true negatives and / or the ability of a method to distinguish truly existing sequence variants from sequencing artifacts or other closely related sequences). For example, when used in connection with any of the various methods described herein that can detect the presence of cancer in a subject, high specificity means that the method correctly identifies the absence of cancer in a subject a high percentage of the time (e.g., the method does not incorrectly identify the presence of cancer in a subject a high percentage of the time). When applied to a sample set of N total sequences, where X true sequences are truly variants and X non-true sequences are truly non-variants, if the method can select at least X% of the non-true variants as non-variants, then the method has a specificity of X%. For example, when applied to a sample set of 1,000 sequences, where 500 sequences are truly variants and 500 are truly non-variants, if the method selects 90% of the 500 non-true variant sequences as non-variants, the method has a specificity of 90%. For example, a method described herein is said to have a specificity of 95% if it correctly detects the absence of cancer in a subject 95% of the times the method is performed. In some embodiments, a method described herein that can detect the absence of cancer in a subject exhibits a specificity of at least 80% (e.g., at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or higher). A method with high specificity minimizes or does not produce false positive results (e.g., when compared to other methods). False positive results can come from any source.For example, in various methods described herein that involve nucleic acid sequencing, false positives may be due to errors introduced into the sequence of interest during sample preparation, sequencing errors, and / or accidental sequencing of closely related sequences, such as pseudogenes or members of gene families. In some embodiments, the methods provided herein that involve detecting the presence of one or more members of two or more classes of biomarkers (e.g., gene biomarkers and / or protein biomarkers) exhibit higher specificity than methods that involve detecting the presence of one or more members of only one class of biomarkers.

[0084] As discussed above, in various embodiments, specificity is the ability of a testing method to truly assign a sample to a first state of identity. (Specificity does not address the ability of a method to find all true first-state samples, i.e., sensitivity.) In one embodiment, the first state is negative, and specificity is the ability to make true (rather than inaccurate) assignments to negatives (and not misassign second-state (e.g., positive) samples as first-state (negative) samples). In one embodiment, the first state is positive, and specificity is the ability to make true (rather than inaccurate) assignments to positives (and not misassign second-state (e.g., negative) samples as first-state (positive) samples).

[0085] As used herein, the phrase "subgenomic interval" refers to a portion of a genome sequence. A subgenomic interval can be of any suitable size (e.g., can include any suitable number of nucleotides). In some embodiments, a subgenomic interval can include a single nucleotide (e.g., a single nucleotide whose variants are associated (positively or negatively) with a tumor phenotype). In some embodiments, a subgenomic interval can include more than one nucleotide. For example, a subgenomic interval can include at least about 2 (e.g., about 5, about 10, about 50, about 100, about 150, about 250, or about 300) nucleotides. In some cases, a subgenomic interval can include an entire gene. In some cases, a subgenomic interval can include a portion of a gene (e.g., a coding region, e.g., an exon; a non-coding region, e.g., an intron; or a regulatory region, e.g., a promoter, an enhancer, a 5' untranslated region (5'UTR), or a 3' untranslated region (3'UTR)). In some cases, a subgenomic interval can include all or a portion of a naturally occurring (e.g., in a genome) nucleotide sequence. For example, a subgenomic interval may correspond to a fragment of genomic DNA that can be subjected to a sequencing reaction. In some cases, a subgenomic interval may be a contiguous nucleotide sequence derived from a genomic source. In some cases, a subgenomic interval may include nucleotide sequences that are not adjacent within a genome. For example, a subgenomic interval may include a nucleotide sequence that includes an exon-exon junction (e.g., within a cDNA reverse-transcribed from the subgenomic interval). In some cases, a subgenomic interval may include a mutation (e.g., an SNV, an SNP, a somatic mutation, a germline mutation, a point mutation, a rearrangement, a deletion mutation (e.g., an in-frame deletion, an intragenic deletion, or a whole-gene deletion), an insertion mutation (e.g., an intragenic insertion), an inversion mutation (e.g., an intrachromosomal inversion), an inverted duplication mutation, a tandem duplication (e.g., an intrachromosomal tandem duplication), a translocation (e.g., a chromosomal translocation or a non-reciprocal translocation), a change in gene copy number, or any combination thereof.

[0086] As used herein, the phrase "leukocyte parameter" refers to the sequence of nucleic acid, eg, chromosomal nucleic acid, of a leukocyte.

[0087] As used herein, the phrase "genomic event" refers to a difference in the sequence of a subgenomic interval from the sequence of a reference sequence. A genomic event can be, for example, a mutation, e.g., a point mutation, or a rearrangement, e.g., a translocation.

[0088] Aneuploidy detection This document provides methods and materials for identifying one or more chromosomal abnormalities (e.g., aneuploidy) in a sample. In some embodiments, the methods and materials described herein are used to identify one or more chromosomal abnormalities (e.g., aneuploidy) in an embryo. In some embodiments, the methods and materials described herein are used to identify one or more chromosomal abnormalities (e.g., aneuploidy) in a mammal (e.g., a young mammal or an adult mammal). For example, a mammal (e.g., a sample taken from a mammal) can be assessed for the presence or absence of one or more chromosomal abnormalities. In some cases, this document provides methods and materials for using amplicon-based sequencing data to identify a mammal as having a disease (e.g., cancer) associated with one or more chromosomal abnormalities. For example, the methods and materials described herein can be applied to a sample taken from a mammal, and the mammal can be identified as having one or more chromosomal abnormalities. For example, the methods and materials described herein can be applied to a sample taken from a mammal, and the mammal can be identified as having a disease (e.g., cancer) associated with one or more chromosomal abnormalities. This document also provides methods and materials for identifying and / or treating diseases or disorders associated with one or more chromosomal abnormalities (e.g., one or more chromosomal abnormalities identified as described herein). In some cases, one or more chromosomal abnormalities may be identified in DNA (e.g., genomic DNA) obtained from a sample collected from a mammal. For example, a prenatal mammal (e.g., a prenatal human) may be identified as having a disease or disorder based, at least in part, on the presence of one or more chromosomal abnormalities. In some embodiments, a mammalian embryo identified as having a disease or disorder based, at least in part, on one or more chromosomal abnormalities may be assessed for purposes of in vitro fertilization. In some embodiments, a mammal identified as having cancer based, at least in part, on the presence of one or more chromosomal abnormalities may be treated with one or more cancer treatments. In some embodiments, a mammal may be identified as having a congenital abnormality based, at least in part, on the presence of one or more chromosomal abnormalities.In some aspects, the methods and materials provided herein are used to test embryos (e.g., embryos obtained by in vitro fertilization) for chromosomal abnormalities before they are transferred to a uterus (e.g., a human uterus) for implantation.

[0089] Disclosed herein, inter alia, is a method for increasing the detection sensitivity of one or more cancers or multiple cancers without altering the specificity of the detection of the cancer or multiple cancers. In one embodiment, the sensitivity of cancer detection by assessing (i) a genetic biomarker, such as a somatic mutation; (ii) a protein biomarker; and (iii) an aneuploidy status is higher than the sensitivity of cancer detection by assessing (i) alone; (ii) alone; (iii) alone; (i) and (ii) alone; (i) and (iii) alone; or (ii) and (iii) alone, for example, about 1.1, 1.2, 1.3, 1.4, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 times. The increased sensitivity by the method comprising (i), (ii), and (iii) does not alter, for example, decrease, the specificity of the detection of the cancer or multiple cancers. An exemplary increase in sensitivity for cancer detection using the methods of the present disclosure is shown in Example 6 of the present disclosure.

[0090] Any suitable mammal can be assessed as described herein. The mammal can be a prenatal mammal (e.g., a prenatal human). The mammal can be a mammal suspected of having a disease (e.g., cancer or congenital abnormality) associated with one or more chromosomal abnormalities. In some cases, humans or other primates, such as monkeys, can be assessed for the presence of one or more chromosomal abnormalities as described herein. In some cases, dogs, cats, horses, cows, pigs, sheep, mice, and rats can be assessed for the presence of one or more chromosomal abnormalities as described herein. For example, humans can be assessed for the presence of one or more chromosomal abnormalities as described herein.

[0091] Any suitable sample from a mammal can be assessed (e.g., assessed for the presence of one or more chromosomal abnormalities) as described herein. The sample can include genomic DNA. In some cases, the sample can include circulating cell-free DNA (e.g., circulating fetal cell-free DNA). In some cases, the sample can include circulating tumor DNA (ctDNA). Examples of samples that can include DNA (e.g., ctDNA) include, but are not limited to, blood (e.g., whole blood, serum, or plasma), amniotic membrane, tissue, urine, cerebrospinal fluid, saliva, sputum, bronchoalveolar lavage fluid, bile, lymph, cyst fluid, stool, ascites, Pap smear, cerebrospinal fluid, internal cervical, endometrial, and fallopian tube samples. For example, the sample can be a plasma sample. For example, the sample can be a urine sample. For example, the sample can be a saliva sample. For example, the sample can be a cyst fluid sample. For example, the sample can be a sputum sample. In some cases, the sample can include a neoplastic cell rate (e.g., a low neoplastic cell rate).

[0092] In some embodiments, the sample may be processed to isolate and / or purify DNA from the sample. In some embodiments, DNA isolation and / or purification may include cell lysis (e.g., using detergents and / or surfactants). In some embodiments, further processing of the DNA (e.g., amplification reactions) is performed without purifying the DNA by cell lysis. In such cases, additional reagents, such as, but not limited to, protease inhibitors, are added to facilitate further processing. In some embodiments, DNA isolation and / or purification may include removing protein (e.g., using proteases). In some cases, DNA isolation and / or purification may include removing RNA (e.g., using RNases). In some embodiments, DNA isolation is performed using commercially available kits (e.g., but not limited to, Qiagen DNAeasy kits) or buffers known in the art (e.g., detergent-containing Tris-buffer).

[0093] In some embodiments, the amount of DNA ("input DNA") input into an isolation and / or purification reaction can vary depending on various factors, including, but not limited to, the average length of the DNA fragments, the overall DNA quality, and / or the type of DNA (e.g., gDNA, mitochondrial DNA, cfDNA). In some embodiments, any suitable amount of input DNA can be used in the methods described herein. In some embodiments, the amount of input DNA can be any amount between 1 picogram (pg) and 500 pg. In some embodiments, the amount of input DNA can be at least 0.01 pg, at least 0.01 pg, at least 0.1 pg, or at least 1 pg. In some embodiments, the amount of input DNA is at least 1 picogram (pg), at least 2 pg, at least 3 pg, at least 4 pg, at least 5 pg, at least 6 pg, at least 7 pg, at least 8 pg, at least 9 pg at least 10 pg, at least 11 pg, at least 12 pg, at least 13 pg, at least 14 pg, at least 15 pg, at least 16 pg, at least 17 pg, at least 18 pg, at least 19 pg, at least 20 pg, at least 21 pg, at least 22 pg, at least 23 pg, at least 24 pg, at least 25 pg, at least 26 pg, at least 27 pg, at least 28 pg, at least 29 pg, at least 30 pg, at least 31 pg, at least 32 pg, at least 33 pg, at least 34 pg, at least 35 pg, at least 36 pg, at least 37 pg, at least 38 pg, at least 39 pg, or at least 40 pg. In some embodiments, the amount of input DNA is 3 pg.

[0094] In some embodiments, the methods and materials for identifying one or more chromosomal abnormalities (e.g., aneuploidy) as described herein may include amplifying multiple amplicons. In some embodiments, multiple amplicons are amplified from multiple chromosomal sequences in a DNA sample. In some embodiments, multiple amplicons can be amplified from any variety of repetitive elements (e.g., see Table 1 for a list of repetitive elements). In some embodiments, multiple amplicons are amplified from multiple short interspersed nucleotide sequences (SINEs). In some embodiments, multiple amplicons are amplified from multiple long interspersed nucleotide sequences (LINEs). Methods for amplifying multiple amplicons include, but are not limited to, polymerase chain reaction (PCR) and isothermal amplification methods (e.g., rolling circle amplification or bridge amplification). In some embodiments, a second amplification step is performed. In some embodiments, DNA amplified in the first amplification reaction is used as a template in the second amplification reaction. In some embodiments, the amplified DNA is purified (e.g., PCR purification using methods known in the art) before the second amplification reaction.

[0095] In some embodiments, the amplification reaction comprises the use of a single primer pair, with a first primer having or comprising SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8 or SEQ ID NO:9. In some embodiments, the amplification reaction involves the use of a single primer pair, including a first primer having at least 80% (e.g., at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) sequence identity to SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9. In some embodiments, the amplification reaction comprises the use of a single primer pair with a second primer having or comprising SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, or SEQ ID NO:19.In some embodiments, the amplification reaction involves the use of a single primer pair, with a second primer having at least 80% (e.g., at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) sequence identity to SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, or SEQ ID NO:19.

[0096] In some embodiments, the first primer comprises: In some embodiments, the second primer has a sequence that is at least 80% identical (e.g., at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical) to TIFF2026012754000011.tif12157. TIFF2026012754000012.tif12158. In some embodiments, the amplification reaction comprises the use of a single primer pair comprising a first primer having SEQ ID NO:1 and a second primer having SEQ ID NO:10. In some embodiments, the amplification reaction comprises the use of a single primer pair comprising a first primer having at least 80% (e.g., at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) sequence identity to SEQ ID NO:1 and a second primer having at least 80% (e.g., at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) sequence identity to SEQ ID NO:10.

[0097] In some embodiments, the first primer comprises, from the 5' to the 3' end, a universal primer sequence (UPS), a unique identifier DNA sequence (UID) and an amplification sequence. In some embodiments, the first primer comprises, from the 5' to the 3' end, a UPS sequence and an amplification sequence. In some embodiments, the first primer comprises, from the 5' to the 3' end, an amplification sequence. In such cases where the first primer comprises at least an amplification sequence, any of a variety of library preparation methods known in the art can be used to prepare a next-generation sequencing library from the amplified amplicon.

[0098] In some embodiments, universal primer sequences (UPS) facilitate the generation of amplicon libraries in preparation for next-generation sequencing. For example, amplicons generated during an amplification reaction using a first primer (SEQ ID NO:1) and a second primer (SEQ ID NO:10) are used as templates in a second amplification reaction. In such cases, the second primer set, designed to bind to the UPS, contains a 5' grafting sequence necessary for hybridization to an Illumina flow cell.

[0099] In some embodiments, the UID comprises a sequence of 16-20 degenerate bases. In some embodiments, a degenerate sequence is a sequence that includes several possible bases at some positions in a nucleotide sequence. In some embodiments of any of the methods described herein, the degenerate sequence can be a degenerate nucleotide sequence that includes about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, or about 50, or at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 nucleotides. In some embodiments, the nucleotide sequence contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 0, 10, 15, 20, 25 or more degenerate positions within the nucleotide sequence. In some embodiments, the degenerate sequence is used as a unique identifier DNA sequence (UID). In some embodiments, the degenerate sequence is used to improve the amplification of amplicons. For example, the degenerate sequence may contain bases complementary to the chromosomal sequence to be amplified. In such cases, high complementarity may increase the affinity of the primer for the chromosomal sequence. In some embodiments, the UID (e.g., degenerate base) may be designed to increase the affinity of the primer for multiple chromosomal sequences.

[0100] In some embodiments, the amplification reaction includes one or more primer pairs (e.g., one or more primer pairs selected from Table 2). In some embodiments, the amplification reaction includes at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, or at least nine primer pairs. In some embodiments, when the amplification reaction includes more than one pair or primers, at least one primer pair includes a primer having SEQ ID NO:1 as a first primer and a primer having SEQ ID NO:10 as a second primer. In some embodiments, when more than one primer pair is included in the amplification reaction, at least one primer pair comprises a first primer having a sequence with at least 80% (e.g., at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) sequence identity to SEQ ID NO:1, and a second primer having a sequence with at least 80% (e.g., at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) sequence identity to SEQ ID NO:10.

[0101] In some embodiments, when one or more primer pairs are included in an amplification reaction, any of a variety of combinations of primers or primer pairs can be selected from Table 2. For example, an amplification reaction containing two primer pairs (e.g., four primers selected from Table 2) can include a first primer pair (e.g., first primer pair 1 in Table 2) comprising a first primer (e.g., a first primer having SEQ ID NO:1) and a second primer (e.g., a second primer having SEQ ID NO:10), and a second primer pair (e.g., second primer pair 2 in Table 2) comprising a third primer (e.g., a third primer having SEQ ID NO:2) and a fourth primer (e.g., a fourth primer having SEQ ID NO:11). Combining any of the forward primers listed in Table 2 (e.g., "FP" having SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9) with any of the reverse primers listed in Table 2 (e.g., "RP" having SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, or SEQ ID NO:19) generates amplicons derived from repetitive elements as described herein (e.g., see Table 1 for a list of exemplary repetitive elements).For example, an amplification reaction containing two primer pairs (e.g., four primers selected from Table 2) may include a first primer pair (e.g., first primer pair 1 in Table 2) comprising a first primer (e.g., a first primer having SEQ ID NO:1) and a second primer (e.g., a second primer having SEQ ID NO:10), and a second primer pair (e.g., a primer pair not listed in Table 2) comprising a third primer (e.g., a third primer having SEQ ID NO:2) and a fourth primer (e.g., a fourth primer having SEQ ID NO:12). In some embodiments, the amplification reaction includes one or more primer pairs in which the first primer is included in both primer pairs. For example, an amplification reaction may include a first primer pair (e.g., first primer pair 1 in Table 2) comprising a first primer (e.g., a first primer having SEQ ID NO:1) and a second primer (e.g., a second primer having SEQ ID NO:10), and a second primer pair comprising a third primer (e.g., a third primer having SEQ ID NO:1) and a fourth primer (e.g., a fourth primer having SEQ ID NO:11).

[0102] In some embodiments, a primer pair is complementary to multiple chromosomal sequences. As used herein, the term "complementary" or "complementarity" indicates that nucleic acid residues can undergo or participate in Watson-Crick or similar base-pairing interactions sufficient to support amplification. In some embodiments, the amplification sequence of a first primer can be designed to amplify one or more chromosomal sequences. In some embodiments, the one or more chromosomal sequences include any of a variety of repetitive elements as described herein (e.g., see Table 1 for a list of exemplary repetitive elements). In some embodiments, the chromosomal sequence is a SINE. In some embodiments, the chromosomal sequence is a LINE. In some embodiments, the chromosomal sequence is a mixture of different types of repetitive elements (e.g., SINEs, LINEs, and / or other exemplary repetitive elements listed in Table 1). In some embodiments, when two or more primer pairs are included in an amplification reaction, each primer pair amplifies a different type of repetitive element (e.g., see Table 1 for a list of exemplary repetitive elements). For example, a first primer pair may amplify a SINE, and a second primer pair may amplify a LINE. Optionally, a third, fourth, fifth, etc. primer pair may amplify a third, fourth, fifth, etc. type of repetitive element (see, e.g., Table 1 for a list of further exemplary repetitive elements). In some embodiments, when two or more primer pairs are included in an amplification reaction, each primer pair generates an amplicon derived from the same type of repetitive element (see, e.g., Table 1 for a list of exemplary repetitive elements). For example, a first primer pair may amplify a SINE, and a second primer pair may amplify a SINE. Optionally, a third, fourth, fifth, etc. primer pair may amplify a SINE. In some embodiments, when two or more primer pairs are included in an amplification reaction, each primer pair generates an amplicon derived from a mixture of different types of repetitive elements (see, e.g., Table 1 for a list of exemplary repetitive elements).

[0103] Table 1. List of exemplary repeat elements TIFF2026012754000013.tif120166TIFF2026012754000014.tif238166TIFF2026012754000015.tif51166

[0104] In some embodiments, one or both primers of a primer pair described herein comprise a primer modification, including, but not limited to, a spacer (e.g., C3 spacer, PC spacer, hexanediol, spacer 9, spacer 18, 1',2'-dideoxyribose (d-spacer)), phosphorylation, phosphorothioate linkage modification, modified nucleic acid, attachment chemistry, and / or linker modification. Examples of modified nucleic acids include, but are not limited to, 2-aminopurine, 2,6-diaminopurine (2-amino-dA), 5-bromo-dU, deoxyuridine, inverted dT, inverted dideoxy-T, dideoxy-C, 5-methyl-dC, deoxyinosine, Super T®, Super G®, locked nucleic acid (LNA), 5-nitroindole, 2'-O-methyl RNA bases, hydroxymethyl-dC, iso-dG, iso-dC, fluoro-C, fluoro-U, fluoro-A, fluoro-G, 2-methoxyethoxy-A, 2-methoxyethoxy-MeC, 2-methoxyethoxy-G, and / or 2-methoxyethoxy-T. Examples of attachment chemistries and linker modifications include, but are not limited to, Acrydite™, adenylation, azide (NHS ester), digoxigenin (NHS ester), cholesterol-TEG, I-linker, amino-modified (e.g., amino-C6, amino-C12, amino-C6 dT, amino-modified and / or Uni-Link™ amino-modified), alkyne (e.g., 5' hexynyl and / or 5-octadiynyl dU), biotinylation (e.g., biotin, biotin (azide), biotin dT, biotin-TEG, dual biotin, pC biotin, and / or desthiobiotin-TEG), and / or thiol-modified (e.g., thiol-C3 SS, dithiol and / or thiol-C6 SS). In some embodiments, any primer as described herein comprises a synthetic nucleic acid.

[0105] In some embodiments, one or both primers of a primer pair described herein contain a primer modification that improves processing of the amplified DNA. In some embodiments, any primer as described herein contains a primer modification that facilitates primer removal (e.g., removal of the primer after the amplification reaction). In some embodiments, the primer modification is transferred to the product of the amplification reaction (e.g., the amplification product contains a modified base). In such cases, the amplification product contains the modification and a property inherent to the modification (e.g., the ability to select for amplification products containing the modification).

[0106] In some embodiments, the method for identifying one or more chromosomal abnormalities as described herein comprises using the read of amplicon-based sequencing.In some embodiments, a plurality of amplicons (for example, the amplicons obtained from DNA samples) are sequenced.In some embodiments, each amplicon is sequenced at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20 times or more. In some embodiments, each amplicon can be sequenced about 1 to about 20 times (e.g., about 1 to about 15, about 1 to about 12, about 1 to about 10, about 1 to about 8, about 1 to about 5, about 5 to about 20, about 7 to about 20, about 10 to about 20, about 13 to about 20, about 3 to about 18, about 5 to about 16, or about 8 to about 12). In some cases, amplicon-based sequencing reads can include contiguous sequence reads. In some cases, the amplicon includes short interspersed nucleotide sequences (SINEs). In some cases, the amplicon-based sequencing reads are between about 100,000 and about 25 million (e.g., between about 100,000 and about 20 million, between about 100,000 and about 15 million, between about 100,000 and about 12 million, between about 100,000 and about 10 million, between about 100,000 and about 5 million, between about 100,000 and about 1 million, between about 100,000 and about 750,000, between about 100,000 and about 500,000, between about 100,000 and about 250,000). The sequence reads may include about 250,000 to about 25 million, about 500,000 to about 25 million, about 750,000 to about 25 million, about 1 million to about 25 million, about 5 million to about 25 million, about 10 million to about 25 million, about 15 million to about 25 million, about 200,000 to about 20 million, about 250,000 to about 15 million, about 500,000 to about 10 million, about 750,000 to about 5 million, or about 1 million to about 2 million).For example, sequencing multiple amplicons can include assigning a unique identifier (UID) to each template molecule (e.g., to each amplicon), amplifying each template molecule with a unique tag to create a UID family, and sequencing the amplification products in duplicate. For example, sequencing multiple amplicons can include assigning a unique identifier (UID) to each template molecule (e.g., to each amplicon), amplifying each template molecule with a unique tag to create a UID family, and sequencing the amplification products in duplicate. TIFF2026012754000016.tif12128 to calculate the Z-score of the variant in the selected chromosome arm, wherein w i is the UID depth for variant i, and Z i is the Z-score of variant i, and k is the number of variants observed on the chromosome arm. In some embodiments, the method of sequencing amplicon comprises a method known in the art (see, for example, US Patent Application Publication No. 2015 / 0051085; and Kinde et al. 2012 PloS ONE 7:e41162, which are incorporated herein by reference in their entirety). In some embodiments, amplicon is aligned to a reference genome (e.g., GRC37).

[0107] In some embodiments, the plurality of amplicons generated by the methods described herein may be from about 10,000 to about 1,000,000 (e.g., from about 15,000 to about 1,000,000, from about 25,000 to about 1,000,000, from about 35,000 to about 1,000,000, from about 50,000 to about 1,000,000, from about 75,000 to about 1,000,000, from about 100,000 to about 1,000,000, from about 125,000 to about 1,000,000, from about 160,000 to about 1,000,000, from about 180,000 to about 1,000,000, from about 200,000 to about 1,000,000, , about 300,000 to about 1,000,000, about 500,000 to about 1,000,000, about 750,000 to about 1,000,000, about 10,000 to about 800,000, about 10,000 to about 500,000, about 10,000 to about 250,000, about 10,000 to about 150,000, about 10,000 to about 100,000, about 10,000 to about 75,000, about 10,000 to about 50,000, about 10,000 to about 40,000, about 10,000 to about 30,000, or about 10,000 to about 20,000) amplicons (e.g., unique amplicons). As a non-limiting example, the plurality of amplicons can include about 745,000 amplicons (e.g., 745,000 unique amplicons). An amplicon in the plurality of amplicons can include about 50 to about 140 nucleotides (e.g., about 60 to about 140, about 76 to about 140, about 90 to about 140, about 100 to about 140, about 130 to about 140, about 50 to about 130, about 50 to about 120, about 50 to about 110, about 50 to about 100, about 50 to about 90, about 50 to about 80, about 60 to about 130, about 70 to about 125, about 80 to about 120, or about 90 to about 100). As a non-limiting example, an amplicon can include about 100 nucleotides.

[0108] In some embodiments, one or more amplicons in a plurality of amplicons produced by the methods described herein can be greater than 1000 base pairs (bp) in length ("long amplicons"). In some embodiments, the one or more long amplicons comprise at least 4.0% of the total amplicons across the plurality of amplicons. In some embodiments, the methods and materials described herein can detect long amplicons when they comprise at least 4.0% of the total amplicons across the plurality of amplicons. In some embodiments, the methods and materials described herein can detect long amplicons when they comprise 0.01% to 3.9% of the total amplicons across the plurality of amplicons.

[0109] In some embodiments, one or more amplicons having a length of >1000 bp are generated by amplification of DNA from cells that do not contain chromosomal abnormalities. In some embodiments, cells that do not contain chromosomal abnormalities are considered contaminating cells. In some embodiments, cells that do not contain chromosomal abnormalities are used as control cells or samples. In some embodiments, contaminating cells can be any of a variety of cells that may be found in a plasma sample and that may dilute the amplification of the intended target. In some embodiments, contaminating cells are white blood cells (e.g., leukocytes, granulocytes, eosinophils, basophils, B cells, T cells, or natural killer cells). For example, contaminating cells can be white blood cells.

[0110] In some embodiments, methods and materials for identifying one or more chromosomal abnormalities as described herein include grouping sequence reads (e.g., from multiple amplicons) into clusters of genomic intervals (e.g., unique clusters). In some embodiments, a given genomic interval is included in one or more clusters. In some embodiments, a given genomic interval may belong to about 100 to about 252 clusters (e.g., about 125 to about 252, about 150 to about 252, about 175 to about 252, about 200 to about 252, about 225 to about 252, about 100 to about 250, about 100 to about 225, about 100 to about 200, about 100 to about 175, about 100 to about 150, about 125 to about 225, about 150 to about 200, or about 160 to about 180). As a non-limiting example, a given genomic interval may belong to about 176 clusters. In some embodiments, each cluster comprises any suitable number of genomic intervals. In some embodiments, each cluster comprises the same number of genomic intervals. In some embodiments, different clusters comprise different numbers of genomic clusters. As a non-limiting example, each cluster may comprise about 200 genomic intervals.

[0111] In some embodiments, genome intervals are identified as having shared amplicon features.As used herein, the term " shared amplicon features " refers to one or more similar features that multiple amplicons have.In some embodiments, multiple genome intervals are grouped into clusters based on one or more shared amplicon features of the sequence reads that are mapped to genome intervals.In some embodiments, the shared amplicon features are the number of amplicons that are mapped to genome intervals (for example, the sum of the distribution of sequence reads in each genome interval).In some embodiments, the shared amplicon features are the average length of the mapped amplicons.

[0112] In some embodiments, a cluster of genomic intervals is between about 5000 and about 6000 (e.g., between about 5100 and about 6000, between about 5200 and about 6000, between about 5300 and about 6000, between about 5400 and about 6000, between about 5500 and about 6000, between about 5600 and about 6000, between about 5700 and about 6000, between about 5800 and about 6000, between about 5900 and about 6000, between about 5000 and about 5900, between about 5000 and about 5800, between about 5000 and about 5700, between about 5000 and about 5600, between about 5000 and about 5500, between about 5000 and about 5600, about 5400, about 5000 to about 5300, about 5000 to about 5200, about 5000 to about 5100, about 5100 to about 5800, about 5100 to about 5700, about 5100 to about 5600, about 5100 to about 5500, about 5100 to about 5400, about 5100 to about 5300, about 5100 to about 5200, about 5200 to about 5600, about 5200 to about 5500, about 5200 to about 5400, about 5200 to about 5300, about 5300 to about 5500, about 5300 to about 5400 or about 5400 to 5500 (about 5200 to about 5700 or about 5300 to about 5500). As a non-limiting example, a cluster of genomic intervals can comprise approximately 5344 genomic intervals.A genomic interval can be any suitable length.For example, a genomic interval can be the length of the amplicon sequenced as described herein.For example, a genomic interval can be the length of a chromosome arm.In some cases, a genomic interval may have a length of about 100 to about 125,000,000 (e.g., about 250 to about 125,000,000, about 500 to about 125,000,000, about 750 to about 125,000,000, about 1,000 to about 125,000,000, about 1,500 to about 125,000,000, about 2,000 to about 125,000,000, or about 3,000 to about 125,000,000). 5,000,000, about 5,000 to about 125,000,000, about 7,500 to about 125,000,000, about 10,000 to about 125,000,000, about 25,000 to about 125,000,000, about 50,000 to about 125,000,000, about 100,000 to about 125,000,000, about 250,000 to about 1 25,000,000, about 500,000 to about 125,000,000, about 100 to about 1,000,000, about 100 to about 750,000, about 100 to about 500,000, about 100 to about 250,000, about 100 to about 100,000, about 100 to about 50,000, about 100 to about 25,000, about 100 to about 10,000, about 1 A genomic interval may comprise approximately 500,000 nucleotides. In some embodiments, clusters of genomic intervals are formed using any suitable method known in the art. In some embodiments, clusters of genomic intervals are formed based on shared amplicon features of the genomic intervals (see, e.g., Douville et al. PNAS 201 115(8):1871-1876, which is incorporated herein by reference in its entirety).

[0113] In some embodiments, the methods and materials described herein for identifying one or more chromosomal abnormalities include assessing a genome (e.g., a mammalian genome) for the presence or absence of one or more chromosomal abnormalities (e.g., aneuploidy). The presence or absence of one or more chromosomal abnormalities in a mammalian genome can be determined, for example, by sequencing a plurality of amplicons obtained from a sample (e.g., a test sample) collected from a mammal to obtain sequence reads, and grouping the sequence reads into clusters of genomic intervals. In some cases, the number of reads of a genomic interval can be compared with the number of reads of other genomic intervals of the same sample. In some cases where the number of reads of a genomic interval is compared with the number of reads of other genomic intervals of the same sample, a second (e.g., control or reference) sample is not assayed. In some cases, the number of reads of a genomic interval can be compared with the number of reads of a genomic interval of another sample. For example, when using the methods and materials described herein to identify genetic relatedness, polymorphisms (e.g., somatic mutations) and / or microsatellite instability, the genomic interval can be compared with the number of reads of the genomic interval of a reference sample. The reference sample can be a synthetic sample. The reference sample can be from a database. In some cases, when the methods and materials described herein are used to identify abnormalities (for example, aneuploidy), the reference sample can be the normal sample collected from the same cancer patient (for example, the sample from the cancer patient that does not have cancer cells) or the normal sample from another source (for example, the patient that does not have cancer).In some cases, when the methods and materials described herein are used to identify abnormalities (for example, aneuploidy), the reference sample can be the normal sample collected from the same patient (for example, the sample from prenatal human that only contains maternal cells).

[0114] In some embodiments, the methods and materials described herein are used to detect aneuploidy in preimplantation embryos (e.g., embryos obtained by in vitro fertilization).In some embodiments, the presence or absence of one or more chromosomal abnormalities in preimplantation embryos is determined by sequencing a plurality of amplicons obtained from a sample (e.g., test sample, for example, but not limited to, one or more cells obtained from blastocysts) collected from preimplantation embryos to obtain sequence reads, and grouping the sequence reads into clusters of genomic intervals.In some cases, the read number of a genomic interval can be compared with the read number of other genomic intervals of the same sample.In some cases, the read number of a genomic interval can be compared with the read number of other genomic intervals of the same sample, and a second (e.g., control or reference) sample is not assayed.In some cases, the read number of a genomic interval can be compared with the read number of another sample's genomic interval (e.g., reference sample).In some embodiments, the reference sample is a sample collected from a reference mammal.In some embodiments, the reference sample is obtained from a database (e.g., the reference sample is an in silico sample whose sequence and / or ploidy of the genomic position of interest is known). Exemplary aneuploidies that can be detected in preimplantation embryos include trisomy 21 (e.g., resulting in Down's syndrome), trisomy 13, trisomy 18, Turner syndrome (e.g., females with only one X chromosome), and Klinefelter syndrome (e.g., males with more than one X chromosome).

[0115] In some embodiments, the methods and materials described herein are used to detect aneuploidy in mammalian genome.For example, a plurality of amplicons obtained from the sample collected from mammalian can be sequenced, sequence reads can be grouped into clusters of genome intervals, the sum of the distribution of sequence reads in each genome interval can be calculated, the Z-score of chromosome arm can be calculated, and the presence or absence of aneuploidy in mammalian genome can be identified.The sum of the distribution of sequence reads in each genome interval can be obtained.For example, the sum of the distribution of sequence reads in each genome interval can be calculated by the formula: TIFF2026012754000017.tif5128, where R i is the number of sequence reads, I is the number of clusters in a chromosome arm, and N is the parameter μ iおよび σ i 2 is a Gaussian distribution with μ i is the average number of sequence reads in each genome interval, and σ i 2 is the variance of sequence reads in each genomic interval. The Z-score of a chromosome arm can be calculated using any suitable method. For example, the Z-score of a chromosome arm can be calculated using the quantile function TIFF2026012754000018.tif5128 can be used to calculate the Z-score. The presence of aneuploidy in a mammalian genome can be identified if the Z-score is outside a predetermined significance threshold, and the absence of aneuploidy in a mammalian genome can be identified if the Z-score is within a predetermined significance threshold. The predetermined threshold can correspond to the reliability of the test and the acceptable number of false positives. For example, the significance threshold can be ±1.96, ±3, or ±5. In some embodiments, the methods and materials described herein use supervised machine learning. In some embodiments, supervised machine learning can detect small changes in one or more chromosome arms. For example, supervised machine learning can detect changes, such as the gain or loss of a chromosome arm, that are often present in diseases or disorders associated with chromosomal abnormalities, such as cancer or congenital anomalies. In some embodiments, supervised machine learning can detect changes, such as the gain or loss of a chromosome arm, that are present in preimplantation embryos (e.g., preimplantation embryos obtained by in vitro fertilization). In some cases, supervised machine learning can be used to classify samples according to aneuploidy status. For example, supervised machine learning can be used to make genome-wide aneuploidy calls. In some cases, a support vector machine model can include obtaining an SVM score. The SVM score can be obtained using any suitable method. In some cases, the SVM score can be obtained as described elsewhere (see, for example, Cortes 1995 Machine learning 20:273-297; and Meyer et al. 2015 R package version:1.6-3). At low read depths, samples typically have high raw SVM scores. Therefore, in some cases, the raw SVM probability can be calculated based on the read depth of the sample using the formula: TIFF2026012754000019.tif8128, where r is the ratio of the SVM score at a certain read depth to the lowest SVM score of a certain sample, assuming that the read depth is sufficient. A and B can be calculated as described in Example 1. For example, A=-7.076 *10^-7, x = number of unique template molecules in a given sample, and B = -1.946 * 10^-1.

[0116] Also provided herein is a new normalization method that reduces the amount of variation between samples.In some embodiments, principal component analysis (PCA) can be used for normalization.In some embodiments, PCA is performed on the sequencing data of a control.For example, PCA can reduce the number of 500 kb genomic intervals from n=5,344 to a number of more manageable sizes.Using the PCA coordinates of the control, a model can be created that predicts whether a certain 500 kb interval will be amplified with higher efficiency or lower efficiency in further samples based on its PCA coordinates. TIFF2026012754000020.tif11146

[0117] For example, for each test sample, the sample can be projected into PCA space, and a correction factor can be calculated for each 500 kb interval as a function of its PCA coordinate. After applying the correction factor to each 500 kb genomic interval, the test sample can be matched with one or more control samples based on the closest Euclidean distance of the 500 kb interval.

[0118] In some embodiments, samples are excluded to ensure data quality. In some embodiments, samples are excluded before, concurrently with, and / or after data analysis. In some embodiments, a series of coefficients may be applied to the data to exclude data that do not meet the criteria set forth in the series of coefficients. In some embodiments, the series of coefficients may be any reasonable numerical value. For example, a series of five coefficients may be used to exclude samples. Any combination of coefficients may be used to determine whether a sample should be excluded. In some embodiments, samples with fewer than 2.5M reads may be excluded. In some embodiments, samples with sufficient evidence of contamination may be excluded. For example, a sample may be considered contaminated if it has at least 10 significant allelic imbalanced chromosome arms (z score >= 2.5) and fewer than 10 significant gains or losses of chromosome arms (z >= 2.5 or z <= -2.5). In some embodiments, allelic imbalance may be determined by SNP, while gains or losses may be assessed by WALDO. In some embodiments, when examining the quality of plasma samples, samples in which more than 8.5% of the amplicons are larger than 94 bp (50 base pairs between the forward and reverse primers) can be excluded. Without wishing to be bound by theory, such samples may be contaminated by leukocyte DNA. In some embodiments, samples outside the dynamic range of the assay as defined by the following formula can be excluded. TIFF2026012754000021.tif14128

[0119] For example, the distribution of this metric has a long tail. Values ​​of >0.2450 and 0.2320 can be selected as the dynamic range within which the cutoff can be evaluated. In some embodiments, plasma samples with known aneuploidy in the white blood cells of the same patient can be excluded. For example, such patients may have clonal hematopoiesis with undetermined potential (CHIP) or a congenital disorder.

[0120] In some embodiments, provided herein are methods for detecting copy number variants (CNVs) of indeterminate length. In some embodiments, provided herein are methods for detecting copy number variations of approximately fixed length. In some embodiments, detecting copy number variations includes calculating values ​​of one or more variables. In some embodiments, a circular binary segmentation algorithm may be applied to examine copy number variants across each chromosome arm using the logarithms of the observed and WALDO predicted values ​​of test samples for each 500 kb interval on each chromosome arm. For example, copy number variants ≦5 Mb in size may be flagged. In some embodiments, flagged CNVs may be removed prior to, concurrently with, and / or after analysis. In some embodiments, small CNVs may be used to assess microdeletions or microamplifications. For example, microdeletions or microamplifications occur in DiGeorge syndrome (chromosome 22q11.2) or breast cancer (chromosome 17q12).

[0121] In some embodiments, methods for using synthetic aneuploidy samples are provided herein. In some embodiments, synthetic aneuploidy samples can be created by adding (or subtracting) reads from several chromosome arms to reads from such normal DNA samples. For example, reads from 1, 10, 15, or 20 chromosome arms can be added or subtracted from each sample. The addition and subtraction can be designed to result in a synthetic sample with a neoplastic cell rate in the range of 0.5% to 1.5% and containing exactly 10 million reads. Reads from each chromosome arm can be added or subtracted uniformly. In some embodiments, methods for creating synthetic aneuploidy samples are provided herein using exemplary pseudocode (FIG. 5). In some embodiments, those skilled in the art can create synthetic samples by applying known coding languages ​​and techniques to the exemplary pseudocode shown in FIG. 5.

[0122] Examples of chromosomal abnormalities that can be detected using the methods and materials described herein include, but are not limited to, numerical disorders, structural abnormalities, allelic imbalances, and microsatellite instability. Chromosomal abnormalities can include numerical disorders. For example, chromosomal abnormalities can include aneuploidy (e.g., abnormal chromosome number). In some cases, aneuploidy can involve an entire chromosome. In some cases, aneuploidy can involve a portion of a chromosome (e.g., gain of a chromosome arm or loss of a chromosome arm). Examples of aneuploidy include, but are not limited to, monosomy, trisomy, tetrasomy, and pentasomy. Chromosomal abnormalities can include structural abnormalities. Examples of structural abnormalities include, but are not limited to, deletions, duplications, translocations (e.g., reciprocal translocations and Robertsonian translocations), inversions, insertions, rings, and isochromosomes. Chromosomal abnormalities can occur in any chromosome pair (e.g., chromosome 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22) and / or one sex chromosome (e.g., X or Y). For example, aneuploidy can occur in , can occur in, but are not limited to, chromosome 13 (e.g., trisomy 13), chromosome 16 (e.g., trisomy 16), chromosome 18 (e.g., trisomy 18), chromosome 21 (e.g., trisomy 21), and / or sex chromosomes (e.g., X chromosome monosomy; sex chromosome trisomies, e.g., XXX, XXY, and XYY; sex chromosome tetrasomies, e.g., XXXX and XXYY; and sex chromosome pentasomy, e.g., XXXXX, XXXXY, and XYYYY).For example, structural abnormalities can occur in, but are not limited to, chromosome 4 (e.g., a partial deletion of the short arm of chromosome 4), chromosome 11 (e.g., a terminal deletion of 11q), chromosome 13 (e.g., a Robertsonian translocation in chromosome 13), chromosome 14 (e.g., a Robertsonian translocation in chromosome 14), chromosome 15 (e.g., a Robertsonian translocation in chromosome 15), chromosome 17 (e.g., a duplication of the gene encoding peripheral myelin protein 22), chromosome 21 (e.g., a Robertsonian translocation in chromosome 21), and chromosome 22 (e.g., a Robertsonian translocation in chromosome 22).

[0123] In some embodiments, the methods and materials as described herein are used to identify and / or treat diseases associated with one or more chromosomal abnormalities (e.g., one or more chromosomal abnormalities identified as described herein, such as, but not limited to, aneuploidy). In some cases, a DNA sample (e.g., a genomic DNA sample) collected from a mammal can be assessed for the presence or absence of one or more chromosomal abnormalities. For example, a mammal (e.g., a human) can be identified as having a disease based, at least in part, on the presence of one or more chromosomal abnormalities and treated with one or more cancer treatments. In some embodiments, a mammal identified as having cancer based, at least in part, on the presence of one or more chromosomal abnormalities is treated with one or more cancer treatments. In some embodiments, a mammal (e.g., a prenatal human) can be identified as having a disease or disorder based, at least in part, on the presence of one or more chromosomal abnormalities. In some embodiments, an embryo (e.g., an embryo obtained by in vitro fertilization) can be identified as unsuitable for transfer to a uterus (e.g., a human uterus) for implantation based, at least in part, on the presence of one or more chromosomal abnormalities. In some embodiments, an embryo (e.g., an embryo obtained by in vitro fertilization) may be identified as suitable for transfer to a uterus (e.g., a human uterus) for implantation based, at least in part, on the absence of one or more chromosomal abnormalities.

[0124] In some embodiments, a mammal that is identified as having one or more chromosomal abnormalities (for example, based at least in part on the existence of one or more chromosomal abnormalities, for example, but not limited to, aneuploidy) as described herein can have the diagnosis of the disease or disorder confirmed by any suitable method.The example of the method that can be used to confirm the existence of one or more chromosomal abnormalities includes, but is not limited to, karyotype analysis, fluorescence in situ hybridization (FISH), quantitative PCR of short tandem repeats, quantitative fluorescence PCR (QF-PCR), quantitative PCR quantity analysis, quantitative mass spectrometry of SNPs, comparative genomic hybridization (CGH), whole genome sequencing and exome sequencing.

[0125] Multi-analyte tests for cancer detection In some embodiments, detection of aneuploidy is used to identify a mammal as having cancer (e.g., any of the exemplary cancers described herein). In some embodiments, detection of one or more genetic biomarkers is used to confirm or identify a mammal as having cancer (e.g., any of the exemplary cancers described herein). In some embodiments, elevated levels of one or more peptide biomarkers are used to confirm or identify a mammal as having cancer (e.g., any of the exemplary cancers described herein). In some embodiments, a mammal identified as having cancer as described herein (e.g., based on detection of aneuploidy and / or based at least in part on the presence or absence of one or more genetic biomarkers (e.g., mutations) and / or elevated levels of one or more protein biomarkers (e.g., peptides)) may have a cancer diagnosis confirmed using any suitable method. Examples of methods that may be used to diagnose or confirm a cancer diagnosis include, but are not limited to, physical examination (e.g., pelvic exam), imaging tests (e.g., ultrasound or CT scan), cytology, and tissue examination (e.g., biopsy).

[0126] In some embodiments, the methods provided herein for identifying one or more chromosomal abnormalities (e.g., aneuploidy) are used to identify a mammal as having a specific stage of cancer. In some embodiments, the cancer of a mammal can be stage I cancer. In some embodiments, the cancer can be stage II cancer. In some embodiments, the cancer can be stage III cancer. In some embodiments, the cancer can be stage IV cancer. In some embodiments, the methods provided herein for identifying one or more chromosomal abnormalities (e.g., aneuploidy) are used to identify a mammal as having a stage of cancer that cannot be reliably detected by conventional cancer detection methods. For example, the methods provided herein for identifying one or more chromosomal abnormalities (e.g., aneuploidy) can be used to identify a mammal as having stage I cancer that cannot be reliably detected by conventional cancer detection methods. In some embodiments, the methods provided herein for identifying 1) one or more chromosomal abnormalities (e.g., aneuploidy) and 2) one or more genetic biomarkers (e.g., any of the genetic biomarkers provided herein) are used to identify a mammal as having a stage of cancer that cannot be reliably detected by conventional cancer detection methods. In some embodiments, the methods provided herein for identifying 1) one or more chromosomal abnormalities (e.g., aneuploidy) and 2) one or more protein biomarkers (e.g., any of the protein biomarkers provided herein) are used to identify a mammal as having a stage of cancer that cannot be reliably detected by conventional cancer detection methods. Non-limiting examples of cancers that can be identified as described herein (e.g., based on the detection of aneuploidy and / or based at least in part on the presence or absence of one or more genetic biomarkers (e.g., mutations) and / or elevated levels of one or more protein biomarkers (e.g., peptides)) include liver cancer, ovarian cancer, esophageal cancer, gastric cancer, pancreatic cancer, colorectal cancer, lung cancer, breast cancer, and prostate cancer.

[0127] In some embodiments, subjects who are detected to have one or more chromosomal abnormalities (e.g., aneuploidy) can be selected for further diagnostic testing. In some embodiments, the methods provided herein can be used to select subjects for further diagnostic testing before the period when the subject can be diagnosed as having early-stage cancer by conventional methods. For example, the methods provided herein for selecting subjects for further diagnostic testing can be used when the subject has not been diagnosed as having cancer by conventional methods and / or when it is unknown whether the subject has underlying cancer. In some embodiments, subjects who are selected for further diagnostic testing can be administered diagnostic testing (e.g., any diagnostic testing described herein) at a higher frequency than subjects who are not selected for further diagnostic testing. For example, subjects who are selected for further diagnostic testing can be administered diagnostic testing twice a day, daily, every other week, weekly, every other month, monthly, quarterly, twice a year, annually, or any frequency therebetween. In some embodiments, subjects who are selected for further diagnostic testing can be administered one or more further diagnostic testing compared to subjects who are not selected for further diagnostic testing. For example, subjects selected for further diagnostic testing may be administered two or more diagnostic tests, while subjects not selected for further diagnostic testing may be administered only one diagnostic test (or no diagnostic test). In some embodiments, the diagnostic testing method may determine that the same type of cancer as the initially detected cancer is present. Additionally or alternatively, the diagnostic testing method may determine that a different type of cancer than the initially detected cancer is present.

[0128] In some embodiments, the diagnostic examination method is a scan, hi some embodiments, the scan is a bone scan, computed tomography (CT), CT angiography (CTA), esophagography (barium swallow), barium enema, gallium scan, magnetic resonance imaging (MRI), mammography, monoclonal antibody scan (e.g., ProstaScint® scan for prostate cancer, OncoScint® scan for ovarian cancer, and CEA-Scan® for colon cancer), multi-gated acquisition (MUGA) scan, PET scan, PET / CT scan, thyroid scan, ultrasound (e.g., breast ultrasound, endobronchial ultrasound, endoscopic ultrasound, transvaginal ultrasound), X-ray, or DEXA scan.

[0129] In some embodiments, the diagnostic examination method is a physical examination, including but not limited to, anoscopy, biopsy, bronchoscopy (e.g., autofluorescence bronchoscopy, white light bronchoscopy, navigation bronchoscopy), breast digital tomosynthesis, digital rectal examination, endoscopy, including but not limited to, capsule endoscopy, virtual endoscopy, arthroscopy, bronchoscopy, colonoscopy, colposcopy, cystoscopy, esophagoscopy, gastroscopy, laparoscopy, laryngoscopy, neuroendscopy, proctoscopy, sigmoidoscopy, skin cancer examination, thoracoscopy, endoscopic retrograde cholangiopancreatography (ERCP), esophagogastroduodenoscopy, pelvic examination.

[0130] In some embodiments, the diagnostic test method is a biopsy (e.g., bone marrow aspiration, tissue biopsy). In some embodiments, the biopsy is performed by needle aspiration biopsy or surgical resection. In some embodiments, the diagnostic test method further comprises obtaining a biological sample (e.g., a tissue sample, a urine sample, a blood sample, a check swab, a saliva sample, a mucosal sample (e.g., sputum, bronchial secretions), a nipple aspirate, a secretion, or an excretion). In some embodiments, the diagnostic test method comprises measuring an exosome protein (e.g., an exosome surface protein (e.g., CD24, CD147, PCA-3)) (Soung et al. (2017) Cancers 9(1):pii:E8). In some embodiments, the diagnostic test method is the Oncotype DX® test (Baehner (2016) Ecancermedicalscience 10:675).

[0131] In some embodiments, the diagnostic test method is a test such as, but not limited to, an alpha-fetoprotein blood test, a bone marrow test, a fecal occult blood test, a human papillomavirus test, a low-dose helical computed tomography scan, a lumbar puncture, a prostate-specific antigen (PSA) test, a Pap smear, or a tumor marker test.

[0132] In some embodiments, the diagnostic testing method involves measuring the level of a known protein biomarker (e.g., CA-125 or prostate-specific antigen (PSA)). For example, elevated amounts of CA-125 can be found in the blood of subjects with ovarian, endometrial, fallopian tube, pancreatic, stomach, esophageal, colon, liver, breast, or lung cancer. The term "biomarker," as used herein, refers to "a biological molecule found in blood, other body fluids, or tissues that is indicative of normal or abnormal processes or a pathological state or disease," as defined, for example, by the National Cancer Institute. (See, e.g., URL www.cancer.gov / publications / dictionaries / cancer-terms?CdrID=45618). Biomarkers may include genetic biomarkers, such as, but not limited to, nucleic acids (e.g., DNA molecules, RNA molecules (e.g., microRNAs, long non-coding RNAs (lncRNAs) or other non-coding RNAs). Biomarkers may include protein biomarkers, such as, but not limited to, peptides, proteins or fragments thereof.

[0133] In some embodiments, the biomarkers are FLT3, NPM1, CEBPA, PRAM1, ALK, BRAF, KRAS, EGFR, Kit, NRAS, JAK2, KRAS, HPV virus, ERBB2, BCR-ABL, BRCA1, BRCA2, CEA, AFP and / or LDH. For example, Easton et al. (1995) Am. J. Hum. Genet. 56:265-271, Hall et al. (1990) Science 250:1684-1689, Lin et al. (2008) Ann. Intern. Med. 149:192-199, Allegra et al. (2009) (2009) J. Clin. Oncol. 27:2091-2096, Paik et al. (2004) N. Engl. J. Med. 351:2817-2826, Bang et al. (2010) Lancet 376:687-697, Piccart-Gebhart et al. (2005) N. Engl. J. Med. 353:1659-1672, Romond et al. (2005) N. Engl. J. Med. 353:1673-1684, Locker et al. (2006) J. Clin. Oncol. 24:5313-5327, Gilligan et al. (2010) J. Clin. Oncol. 28:3388-3404, Harris et al. (2007) J. Clin. Oncol. 25:5287-5312; Henry and Hayes (2012) Mol. Oncol. 6:140-146. In some embodiments, the biomarker is a biomarker for the detection of breast cancer in a subject, such as, but not limited to, MUC-1, CEA, p53, urokinase-type plasminogen activator, BRCA1, BRCA2, and / or HER2 (Gam (2012) World J. Exp. Med. 2(5):86-91).In some embodiments, the biomarker is a biomarker for detecting lung cancer in a subject, such as, but not limited to, KRAS, EGFR, ALK, MET, and / or ROS1 (Mao (2002) Oncogene 21:6960-6969; Korpanty et al. (2014) Front Oncol. 4:204). In some embodiments, the biomarker is a biomarker for detecting ovarian cancer in a subject, such as, but not limited to, HPV, CA-125, HE4, CEA, VCAM-1, KLK6 / 7, GST1, PRSS8, FOLR1, ALDH1 (Nolen and Lokshin (2012) Future Oncol. 8(1):55-71; Sarojini et al. (2012) J. Oncol. 2012:709049). In some embodiments, the biomarker is a biomarker for the detection of colorectal cancer in a subject, such as, but not limited to, MLH1, MSH2, MSH6, PMS2, KRAS, and BRAF (Gonzalez-Pons and Cruz-Correa (2015) Biomed. Res. Int. 2015:149014; Alvarez-Chaver et al. (2014) World J. Gastroenterol. 20(14):3804-3824). In some embodiments, diagnostic testing methods examine the presence and / or expression levels of nucleic acids (e.g., microRNAs (Sethi et al. (2011) J. Carcinog. Mutag. S1-005), RNA, SNPs (Hosein et al. (2013) Lab. Invest doi:10.1038 / labinvest.2013.54; Falzoi et al. (2010) Pharmacogenomics 11:559-571), methylation status (Castelo-Branco et al. (2013) Lancet Oncol 14:534-542), hotspot cancer mutations (Yousem et al. (2013) Chest 143:1679-1684)).Non-limiting examples of the method for detecting nucleic acid in sample include PCR, RT-PCR, sequencing (for example, next-generation sequencing, deep sequencing), DNA microarray, microRNA microarray, SNP microarray, fluorescent in situ hybridization (FISH), restriction fragment length polymorphism (RFLP), gel electrophoresis, Northern blot analysis, Southern blot analysis, colorimetric in situ hybridization (CISH), chromatin immunoprecipitation (ChIP), SNP genotyping and DNA methylation assay.For example, see Meldrum et al.(2011)Clin.Biochem.Rev.32(4):177-195;Sidranksy(1997)Science278(5340):1054-9.

[0134] In some embodiments, the diagnostic testing method involves determining the presence of a protein biomarker (e.g., a plasma biomarker (Mirus et al. (2015) Clin. Cancer Res. 21(7):1764-1771)) in a sample. Non-limiting examples of methods for determining the presence of protein biomarkers include Western blot analysis, immunohistochemistry (IHC), immunofluorescence, mass spectrometry (MS) (e.g., matrix-assisted laser desorption / ionization (MALDI)-MS, surface-enhanced laser desorption / ionization time-of-flight (SELDI-TOF)-MS), enzyme-linked immunosorbent assay (ELISA), flow cytometry, proximity assays (e.g., VeraTag proximity assay (Shi et al. (2009) Diagnostic molecular pathology: the American journal of surgical pathology, part B:18:11-21; Huang et al. (2010) AM. J. Clin. Pathol. 134:303-11)), protein microarrays (e.g., antibody microarrays (Ingvarsson et al. (2008) Proteomics 8:2211-9; Woodbury et al. (2002) J. Proteome Res. 1:233-237), IHC-based microarrays (Stromberg et al. (2007) Proteomics 7:2142-50), and microarray ELISA (Schroder et al. (2010) Mol. Cell. Proteomics 9:1271-80). In some embodiments, the method for determining the presence of a protein biomarker is a functional assay.In some embodiments, the functional assay is a kinase assay (Ghosh et al. (2010) Biosensors & Bioelectronics 26:424-31, Mizutani et al. (2010) Clin. Cancer Res. 16:3964-75, Lee et al. (2012) Biomed. Microdevices 14:247-57), a protease assay (Lowe et al. (2012) ACS nano. 6:851-7, Fujiwara et al. (2006) Breast cancer 13:272-8, Darragh et al. (2010) Cancer Res 70:1505-12). For example, see Powers and Palecek (2015) J. Heathc Eng. 3(4):503-534 for a review of protein analysis assays for diagnosing cancer patients.

[0135] In some embodiments, any suitable disease or condition associated with one or more chromosomal abnormalities as described herein (e.g., based at least in part on the presence of one or more chromosomal abnormalities, such as, but not limited to, aneuploidy) is identified as described herein. In some embodiments, the disease is cancer. Examples of cancers that may be associated with one or more chromosomal abnormalities include, but are not limited to, lung cancer (e.g., small cell lung cancer or non-small cell lung cancer), papillary thyroid cancer, medullary thyroid cancer, differentiated thyroid cancer, recurrent thyroid cancer, refractory differentiated thyroid cancer, lung adenocarcinoma, bronchiolopulmonary cell carcinoma, multiple endocrine neoplasia type 2A or type 2B (MEN2A or MEN2B, respectively), pheochromocytoma, parathyroid hyperplasia, breast cancer, colorectal cancer (e.g., metastatic colorectal cancer), papillary renal cell carcinoma, Gastrointestinal mucosal ganglioneuroma, inflammatory myofibroblastic tumor, or cervical cancer, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), adolescent cancer, adrenal cancer, adrenocortical carcinoma, anal cancer, appendix cancer, astrocytoma, atypical teratoma / rhabdoid tumor, basal cell carcinoma, bile duct cancer, bladder cancer, bone cancer, brain stem glioma, brain tumor, breast cancer, bronchial tumor, Burkitt lymphoma, carcinoid tumor, cancer of unknown primary site, cardiac tumor, cervical cancer, childhood cancer, chordoma, chronic lymphocytic Chronic leukemia (CLL), chronic myelogenous leukemia (CML), chronic myeloproliferative neoplasm, colon cancer, colorectal cancer, craniopharyngioma, cutaneous T-cell lymphoma, bile duct cancer, ductal carcinoma in situ, embryonal tumors, endometrial cancer, ependymoma, esophageal cancer, nasal neuroblastoma, Ewing's sarcoma, extracranial germ cell tumors, extragonadal germ cell tumors, extrahepatic bile duct cancer, eye cancer, fallopian tube cancer, fibrous histiocytoma of bone, gallbladder cancer, stomach cancer, carcinoid tumors of the gastrointestinal tract, gastrointestinal stromal tumors (GIST), germ cell alveolar tumor, gestational trophoblastic disease, glioma, hairy cell tumor, hairy cell leukemia, head and neck cancer, heart cancer, hepatocellular carcinoma, histiocytosis, Hodgkin's lymphoma, hypopharyngeal cancer, intraocular melanoma, pancreatic islet cell tumor, pancreatic neuroendocrine tumor, Kaposi's sarcoma, kidney cancer, Langerhans cell histiocytosis, laryngeal cancer, leukemia, lip and oral cavity cancer, liver cancer, lung cancer, lymphoma, macroglobulinemia, malignant fibrous histiocytoma of bone, bone cancer, melanoma, Merkel cell carcinoma, mesothelioma, metastatic cervical squamous cell carcinoma,Midline vein cancer, mouth cancer, multiple endocrine neoplasia syndrome, multiple myeloma, mycosis fungoides, myelodysplastic syndrome, myelodysplastic / myeloproliferative neoplasm, myelogenous leukemia, myeloid leukemia, multiple myeloma, myeloproliferative neoplasm, cancer of the nasal cavity and paranasal sinuses, nasopharyngeal carcinoma, neuroblastoma, non-Hodgkin's lymphoma, non-small cell lung cancer, oral cancer, oral cavity cancer, lip cancer, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, hepatobiliary cancer, upper urinary tract cancer, papillomatosis, paraganglioma, cancer of the paranasal sinuses and nasal cavity, parathyroid cancer, penile cancer, pharyngeal cancer, pheochromocytoma, pituitary cancer, plasma cell neoplasms, pleuropulmonary blastoma, breast cancer during pregnancy, primary central nervous system lymphoma, primary peritoneal cancer, prostate cancer, rectal cancer, renal cell carcinoma, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, sarcoma, Sézary syndrome, skin cancer, small cell lung cancer , small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, squamous cell carcinoma of the cervix, gastric cancer, T-cell lymphoma, testicular cancer, throat cancer, thymoma and thymic carcinoma, thyroid cancer, transitional cell carcinoma of the renal pelvis and ureter, cancer of unknown primary origin, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenström's macroglobulinemia, Wilms' tumor, 1p36 deletion syndrome, 1q21.1 deletion syndrome, 2q37 deletion syndrome, Wolf-Hirschhorn syndrome, Crissy-Cat syndrome (Cri du chat), 5q deletion syndrome, Williams syndrome, 8p monosomy, 8q monosomy, Alfie syndrome, Kleefstra syndrome, 10p monosomy, 10q monosomy, Jacobsen syndrome, Patau syndrome, Angelman syndrome, Prader-Willi syndrome, Miller-Dieker syndrome, Smith-Maginis syndrome, Edwards syndrome, Down syndrome, DiGeorge syndrome, Phelan-McDermid syndrome, 22q11.2 distal deletion syndrome, cat eye syndrome, XYY syndrome, triple X syndrome, Klinefelter syndrome, Wolf-Hirschhorn syndrome, Jacobsen syndrome, Charcot-Marie-Tooth disease type 1A, and Lynch syndrome.

[0136] Once identified as having a disease associated with one or more chromosomal abnormalities as described herein (e.g., based at least in part on the presence of one or more chromosomal abnormalities, such as, but not limited to, aneuploidy), the mammal (e.g., human) can be treated accordingly. For example, if a mammal is identified as having a cancer associated with one or more chromosomal abnormalities as described herein, the mammal can be treated with one or more cancer treatment methods. The one or more cancer treatment methods can include any suitable cancer treatment method. The cancer treatment method can include surgery. The cancer treatment method can include radiation therapy. The cancer treatment method can include the administration of drug therapy, such as chemotherapy, hormonal therapy, targeted therapy, and / or cytotoxic therapy. Examples of cancer treatments include, but are not limited to, platinum compounds (e.g., cisplatin or carboplatin), taxanes (e.g., paclitaxel or docetaxel), albumin-bound paclitaxel (nab-paclitaxel), altretamine, capecitabine, cyclophosphamide, etoposide (vp-16), gemcitabine, ifosfamide, irinotecan (cpt-11), liposomal doxorubicin, melphalan, pemetrexed, topotecan, vinorelbine, luteinizing hormone (luteinizing hormone), and luteinizing hormone (luteinizing hormone). These include hormone-releasing hormone (LHRH) agonists (e.g., goserelin and leuprolide), anti-estrogen therapy (e.g., tamoxifen), aromatase inhibitors (e.g., letrozole, anastrozole, and exemestane), angiogenesis inhibitors (e.g., bevacizumab), poly(ADP)-ribose polymerase (PARP) inhibitors (e.g., olaparib, rucaparib, and niraparib), external beam radiation therapy, brachytherapy, radioactive phosphorus, and any combination thereof.

[0137] Multi-analyte testing for increased detection sensitivity In some embodiments, the methods provided herein for detecting aneuploidy (e.g., using analysis of chromosomal sequences (e.g., see Table 1 for an exemplary list of repetitive elements that can be analyzed)) result in increased sensitivity for cancer detection compared to cancer detection using the presence of one or more genetic biomarkers as an indicator of cancer. In some embodiments, the methods provided herein for detecting aneuploidy (e.g., using analysis of chromosomal sequences (e.g., see Table 1 for an exemplary list of repetitive elements that can be analyzed)) result in increased sensitivity for cancer detection compared to cancer detection using the presence of one or more protein biomarkers as an indicator of cancer.

[0138] In some embodiments, the methods provided herein for detecting aneuploidy (e.g., using analysis of chromosomal sequences (e.g., see Table 1 for an exemplary list of repetitive elements that can be analyzed)) are combined with one or more methods for detecting the presence of one or more genetic biomarkers (e.g., mutations). In some embodiments, the combination of detecting aneuploidy and detecting genetic biomarkers increases the specificity and / or sensitivity of cancer detection. In some embodiments, the methods provided herein for detecting aneuploidy (e.g., using analysis of chromosomal sequences (e.g., see Table 1 for an exemplary list of repetitive elements that can be analyzed)) are combined with one or more methods for detecting the presence of one or more members of a panel of protein biomarkers (e.g., peptides). In some embodiments, the combination of detecting aneuploidy and detecting protein biomarkers increases the specificity and / or sensitivity of cancer detection. In some embodiments, the methods provided herein for detecting aneuploidy (e.g., using analysis of chromosomal sequences (e.g., see Table 1 for an exemplary list of repetitive elements that can be analyzed)) are combined with methods for detecting the presence of one or more genetic biomarkers (e.g., mutations) and / or methods for detecting the presence of one or more members of a panel of protein biomarkers (e.g., peptides). In some embodiments, the combination of detecting aneuploidy with detecting genetic and / or protein biomarkers increases the specificity and / or sensitivity of cancer detection.

[0139] In some embodiments, the method provided herein for detecting aneuploidy is used in combination with a method for detecting the presence of one or more gene biomarkers (such as mutations) in one or more genes selected from the group consisting of NRAS, PTEN, FGFR2, KRAS, POLE, AKT1, TP53, RNF43, PPP2R1A, MAPK1, CTNNB1, PIK3CA, FBXW7, PIK3R1, APC, EGFR, BRAF.In some embodiments, the method provided herein for detecting aneuploidy is used in combination with a method for detecting the presence of one or more gene biomarkers (such as mutations) in one or more genes selected from the group consisting of PTEN, TP53, PIK3CA, PIK3R1, CTNNB1, KRAS, FGFR2, POLE, APC, FBXW7, RNF43 and PPP2R1A.In some embodiments, assay is used in combination with a method for detecting the presence of one or more gene biomarkers (such as mutations) in one or more genes selected from the group consisting of PTEN, TP53, PIK3CA, PIK3R1, CTNNB1, KRAS, FGFR2, POLE, APC, FBXW7, RNF43 and PPP2R1A. TIFF2026012754000022.tif99157. In some embodiments, detecting aneuploidy in combination with detecting one or more genetic biomarkers (e.g., mutations) increases the specificity and / or sensitivity of cancer detection.

[0140] In some embodiments, detecting genetic biomarkers (e.g., one or more genetic biomarkers) comprises any of the various methods described in U.S. Patent No. 7,700,286, which is incorporated herein by reference in its entirety. Any of a variety of messenger RNA ("mRNA") isolation methods known in the art can be used to isolate RNA from a sample (e.g., Qiagen RNeasy Kit). Any of a variety of genomic DNA ("gDNA") isolation methods known in the art can be used to isolate gDNA from a sample (e.g., Qiagen DNeasy Kit). In some embodiments, detecting genetic biomarkers comprises a cancer detection assay. In some embodiments, the amount of gDNA and / or mRNA in a sample is measured for any of the genetic biomarkers disclosed herein. Changes in the amount of gDNA and / or mRNA can be indicative of cancer. For example, when measuring gDNA, gene amplification (e.g., an increase in the copy number of a chromosomal sequence (e.g., a coding region of a gene or non-coding DNA (see, e.g., Table 1 for an exemplary list of repetitive elements that can be measured)) can be indicative of cancer. For example, when measuring mRNA, an increase in the amount of RNA (e.g., increased expression of a genetic biomarker) can be indicative of cancer. In some cases, changes in DNA and changes in RNA can be correlated.

[0141] In some embodiments, the methods provided herein for detecting aneuploidy can be used in combination with methods for detecting the presence of one or more protein biomarkers (e.g., peptides) in one or more proteins selected from the group consisting of AFP, CA19-9, CEA, HGF, OPN, CA-125, CA15-3, MPO, prolactin (PRL) and / or TIMP-1 to examine the presence of cancer (e.g., ovarian or endometrial). In some embodiments, the protein biomarker can be any suitable peptide biomarker. In some embodiments, the peptide biomarker can be a peptide biomarker associated with cancer. For example, the peptide biomarker can be a peptide that has elevated levels in cancer (e.g., compared with a reference level of the peptide).

[0142] Exemplary, non-limiting threshold levels for some specific protein biomarkers include CA19-9 (>92 U / ml), CEA (>7,507 pg / ml), CA125 (>577 U / ml), AFP (>21,321 pg / ml), prolactin (>145,345 pg / ml), HGF (>899 pg / ml), OPN (>157,772 pg / ml), TIMP-1 (>176,989 pg / ml), follistatin (>1,970 pg / ml), and CA15-3 (>98 U / ml). In some embodiments, the threshold level of a protein biomarker can be higher (e.g., about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 100% or more) than the exemplary threshold levels described herein. In some embodiments, the threshold level of a protein biomarker can be lower (e.g., about 10%, about 20%, about 30%, about 40%, about 50% or more lower) than the exemplary threshold levels described herein.

[0143] In some embodiments, the threshold level of CA19-9 may be at least about 92 U / mL (e.g., about 92 U / mL). In some embodiments, the threshold level of CA19-9 may be 92 U / mL. In some embodiments, the threshold level of CEA may be at least about 7,507 pg / mL (e.g., about 7,507 pg / mL). In some embodiments, the threshold level of CEA may be 7.5 ng / mL. In some embodiments, the threshold level of HGF may be at least about 899 pg / mL (e.g., about 899 pg / mL). In some embodiments, the threshold level of HGF may be 0.92 ng / mL. In some embodiments, the threshold level of OPN may be at least about 157,772 pg / mL (e.g., about 157,772 pg / mL). In some embodiments, the threshold level of OPN may be 158 ng / mL. In some embodiments, the threshold level for CA125 may be at least about 577 U / ml (e.g., about 577 U / ml). In some embodiments, the threshold level for CA125 may be 577 U / mL. In some embodiments, the threshold level for AFP may be at least about 21,321 pg / ml (e.g., about 21,321 pg / ml). In some embodiments, the threshold level for AFP may be 21,321 pg / ml. In some embodiments, the threshold level for prolactin may be at least about 145,345 pg / ml (e.g., about 145,345 pg / ml). In some embodiments, the threshold level for prolactin may be 145,345 pg / ml. In some embodiments, the threshold level for TIMP-1 may be at least about 176,989 pg / ml (e.g., about 176,989 pg / ml). In some embodiments, the threshold level of TIMP-1 may be 176,989 pg / ml. In some embodiments, the threshold level of follistatin may be at least about 1,970 pg / ml (e.g., about 1,970 pg / ml). In some embodiments, the threshold level of CA15-3 may be at least about 98 U / ml (e.g., about 98 U / ml). In some embodiments, the threshold level of CA15-3 may be 98 U / ml.In some embodiments, the threshold levels of CA19-9, CEA, and / or OPN are 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100% or even higher than the threshold levels listed above (e.g., 92 U / mL for CA-19-9, 7,507 pg / ml for CEA, 899 pg / ml for HGF, 157,772 pg / ml for OPN, 577 U / ml for CA125, 21,321 pg / ml for AFP, 145,345 pg / ml for prolactin, 176,989 pg / ml for TIMP-1, 1,970 pg / ml for follistatin, 1,970 pg / ml for TIMP-1 ... pg / ml and / or for CA15-3 above the threshold level of 98 U / ml).

[0144] In some embodiments, the threshold level for a protein biomarker can be higher than the level typical for diagnostic or clinical testing, for example, the threshold level for CA19-9 can be greater than about 37 U / mL (e.g., greater than about 40 U / mL, greater than about 45 U / mL, greater than about 50 U / mL, greater than about 55 U / mL, greater than about 60 U / mL, greater than about 65 U / mL, greater than about 70 U / mL, greater than about 75 U / mL, greater than about 80 U / mL, greater than about 85 U / mL, greater than about 90 U / mL, greater than about 95 U / mL, or higher). Additionally or alternatively, the threshold level for CEA can be greater than about 2.5 ug / L (e.g., greater than about 3.0 ug / L, greater than about 3.5 ug / L, greater than about 4.0 ug / L, greater than about 4.5 ug / L, greater than about 5.0 ug / L, greater than about 5.5 ug / L, greater than about 6.0 ug / L, greater than about 6.5 ug / L, greater than about 7.0 ug / L, greater than about 7.5 ug / L or higher). Additionally or alternatively, the threshold level of CA125 can be greater than about 35 U / mL (e.g., greater than about 40 U / mL, greater than about 45 U / mL, greater than about 50 U / mL, greater than about 55 U / mL, greater than about 60 U / mL, greater than about 65 U / mL, greater than about 70 U / mL, greater than about 75 U / mL, greater than about 80 U / mL, greater than about 85 U / mL, greater than about 90 U / mL, greater than about 95 U / mL, greater than about 100 U / mL, greater than about 150 U / mL, greater than about 200 U / mL, greater than about 250 U / mL, greater than about 300 U / mL, greater than about 350 U / mL, greater than about 400 U / mL, greater than about 450 U / mL, greater than about 500 U / mL, greater than about 550 U / mL or higher). Additionally or alternatively, the threshold level for AFP can be greater than about 21 ng / mL (e.g., greater than about 25 ng / L, greater than about 30 ng / L, greater than about 40 ng / L, greater than about 50 ng / L, greater than about 60 ng / L, greater than about 70 ng / L, greater than about 80 ng / L, greater than about 90 ng / L, greater than about 100 ng / L, greater than about 150 ng / L, greater than about 200 ng / L, greater than about 250 ng / L, greater than about 300 ng / L, greater than about 350 ng / L, greater than about 400 ng / L or higher).Additionally or alternatively, the threshold level of TIMP-1 can be greater than about 2300 ng / mL (e.g., greater than about 2,500 ng / L, greater than about 3,000 ng / L, greater than about 4,000 ng / L, greater than about 5,000 ng / L, greater than about 6,000 ng / L, greater than about 7,000 ng / L, greater than about 8,000 ng / L, greater than about 9,000 ng / L, greater than about 10,000 ng / L, greater than about 15,000 ng / L, greater than about 20,000 ng / L, greater than about 25,000 ng / L, greater than about 30,000 ng / L, greater than about 35,000 ng / L, greater than about 40,000 ng / L or higher). Additionally or alternatively, the threshold level for follistatin can be greater than about 2 ug / mL (e.g., greater than about 2.5 ug / L, greater than about 3.0 ug / L, greater than about 3.5 ug / L, greater than about 4.0 ug / L, greater than about 4.5 ug / L, greater than about 5.0 ug / L, greater than about 5.5 ug / L, greater than about 6.0 ug / L, greater than about 6.5 ug / L, greater than about 7.0 ug / L, greater than about 7.5 ug / L or higher). Additionally or alternatively, the threshold level for CA15-3 can be greater than about 30 U / mL (e.g., greater than about 35 U / mL, greater than about 40 U / mL, greater than about 45 U / mL, greater than about 50 U / mL, greater than about 55 U / mL, greater than about 60 U / mL, greater than about 65 U / mL, greater than about 70 U / mL, greater than about 75 U / mL, greater than about 80 U / mL, greater than about 85 U / mL, greater than about 90 U / mL, greater than about 95 U / mL, or higher). In some embodiments, detecting one or more protein biomarkers at threshold levels higher than levels typical for testing during conventional diagnostic or clinical assays can improve the sensitivity of cancer detection.

[0145] Examples of peptide biomarkers include, but are not limited to, AFP, angiopoietin-2, AXL, CA125, CA 15-3, CA19-9, CD44, CEA, CYFRA 21-1, DKK1, endoglin, FGF2, follistatin, galectin-3, G-CSF, GDF15, HE4, HGF, IL-6, IL-8, kallikrein-6, leptin, LRG-1, mesothelin, midkine, myeloperoxidase, NSE, OPG, OPN, PAR, prolactin, sEGFR, sFas, SHBG, sHER2 / sEGFR2 / sErbB2, sPECAM-1, TGFα, thrombospondin-2, TIMP-1, TIMP-2, and vitronectin. For example, peptide biomarkers may include one or more of OPN, IL-6, CEA, CA125, HGF, myeloperoxidase, CA19-9, midkine, and / or TIMP- 1. In some embodiments, the detection of aneuploidy in combination with the detection of one or more protein biomarkers (e.g., peptides) increases the specificity and / or sensitivity of cancer detection.

[0146] In some embodiments, the presence of gene biomarkers and / or protein biomarkers can be detected in any of a variety of biological samples isolated or collected from a subject (e.g., a human subject), including, but not limited to, blood, plasma, serum, urine, cerebrospinal fluid, saliva, sputum, bronchoalveolar lavage fluid, bile, lymph, cyst fluid, stool, ascites, and combinations thereof. Any protein biomarker known in the art can be detected if a threshold value above that of a normal, healthy human subject but not a human subject with cancer is obtained. Any suitable method can be used to detect the level of one or more protein biomarkers as described herein. In some embodiments, the level of one or more protein biomarkers is compared to a predetermined threshold. In some embodiments, the predetermined threshold is a universal or global threshold. In some embodiments, the predetermined threshold is a threshold for a particular protein biomarker. In some embodiments, the level of one or more protein biomarkers is compared to the absolute amount of a reference protein biomarker. In some embodiments, the level of one or more protein biomarkers is relative to the amount of the reference protein biomarker. In some embodiments, the level of one or more protein biomarkers is an elevated level. In some embodiments, the level of one or more protein biomarkers is above a predetermined threshold. In some embodiments, the level of one or more protein biomarkers is within a predetermined range of thresholds. In some embodiments, the level of one or more protein biomarkers is at a predetermined threshold or is an approximation of a predetermined threshold. In some embodiments, the level of one or more protein biomarkers is below a predetermined threshold. In some embodiments, the level of one or more protein biomarkers from a biological sample is below a certain threshold. In some embodiments, the level of one or more protein biomarkers from a biological sample is reduced compared to a predetermined threshold.

[0147] In some embodiments, the methods and materials described herein can be used to detect one or more polymorphisms (e.g., somatic mutations) in the genome of a mammal. For example, a plurality of amplicons obtained from a sample taken from a first mammal (e.g., a test mammal or a mammal suspected of having one or more polymorphisms) can be sequenced, and a plurality of amplicons obtained from a sample taken from a second mammal (e.g., a reference mammal) can be sequenced, and the variant sequence reads from the sample taken from the first mammal can be grouped into clusters of genomic intervals, and the reference sequence reads from the sample taken from the second mammal can be grouped into clusters of genomic intervals. Chromosome arms can be selected in which the sum of variant sequence reads and reference sequence reads in both alleles is greater than about 3 (e.g., greater than about 4, greater than about 5, greater than about 6, greater than about 7, greater than about 8, greater than about 9, greater than about 10, greater than about 12, greater than about 15, greater than about 18, greater than about 20, greater than about 22, greater than about 25, or greater than about 30), and the variant-allele frequency (VAF) of the selected chromosome arm can be determined, and the presence or absence of one or more polymorphisms in the selected chromosome arm can be identified. The VAF of the selected chromosome arm can be determined using any suitable technique. For example, the VAF of the selected chromosome arm can be the numerical value of variant sequence reads / total number of sequence reads. The presence of one or more polymorphisms in a mammalian genome can be identified when the VAF is about 0.2 to about 0.8 (e.g., about 0.3 to about 0.8, about 0.4 to about 0.8, about 0.5 to about 0.8, about 0.6 to about 0.8, about 0.2 to about 0.7, about 0.2 to about 0.6, about 0.2 to about 0.5, or about 0.2 to about 0.4), and the absence of one or more polymorphisms in a mammalian genome can be identified when the VAF is within a predetermined significance threshold. For example, but not limited to, the presence of one or more polymorphisms in a mammalian genome can be identified when the VAF is about 0.4 to 0.6.

[0148] In some embodiments, the methods and materials described herein can be used for sample identification.The repeat elements amplified by the methods described herein contain common polymorphisms that can be used to establish sample identity between samples (for example, plasma, tumor and blood) or to prove the sample identity.For example, the genotype at each polymorphic location can be identified and compared between samples.The overall similarity between samples at polymorphic locations can be used to determine the sample identity.

[0149] In some cases, diseases associated with one or more chromosomal abnormalities (e.g., based at least in part on the presence of one or more chromosomal abnormalities, such as, but not limited to, aneuploidy) as described herein are also associated with a high mutation rate (e.g., a high mutation rate may be associated with the stage of the disease) when compared with a control (e.g., a non-disease sample).In such cases, the materials and methods described herein can be used to (a) identify the presence of one or more chromosomal abnormalities (e.g., aneuploidy), and (b) identify the stage of the disease (e.g., cancer stages I, II, III, and IV) based on the measurement of the mutation rate (e.g., number of mutations) compared with the control.

[0150] The invention is further described in the following examples, which do not limit the scope of the invention as defined in the claims. [Example]

[0151] Example 1: Detection of aneuploidy in patients with cancer This example describes a novel adaptation of amplicon-based aneuploidy detection. Using supervised machine learning to detect alterations in chromosome arms, an approach called WALDO (Within-Sample-Aneuploidy-Detection) has improved aneuploidy detection sensitivity compared to previous methods. Here, we demonstrate that using WALDO to analyze short interspersed nucleotide sequences (SINE) amplicons from DNA samples increases aneuploidy detection sensitivity. Furthermore, approximately 1,000,000 SINE amplicons, with an average length of approximately 100 bp, reduce the required cell-free DNA input while also increasing detection sensitivity.

[0152] material and method Primer To generate a list of candidate primers, we calculated the frequency of all possible hexamers (4^6 = 4096) within the hg19 RepeatMasker track. We then calculated the frequency of all possible tetramers (4^4 = 256) within 75 bp upstream or downstream of a hexamer. These hexamers and tetramers combined produced 2,097,152 candidate pairs. These pairs were selected for further assessment based on the number of unique genomic loci predicted from PCR-mediated amplification, the average size between a hexamer and its corresponding tetramer, and the distribution of such sizes, aiming for a unimodal distribution. This filtering criterion yielded 16 potential k-mer pairs, and 16 primer pairs were designed incorporating such k-mer pairs at their 3-terminus. A k-mer is understood in the art to refer to a subsequence of length k contained within a sequence.

[0153] In total, 16 primers were initially designed and tested (Table 2). One primer (SEQ ID NO:1) consistently produced low levels of primer dimers and was selected for use in testing the cohort. Primer pairs with SEQ ID NO:1 as one primer amplified 745,184 unique amplicons, with an average amplicon size of approximately 88 bp (Figure 1A). The amplicon size shown in Figure 1A includes the 45 bp primer. For example, without the primer, the amplicons have an average size of approximately 43 base pairs (Figure 1B).

[0154] (Table 2) TIFF2026012754000023.tif146170

[0155] Sequencing library preparation The first primer, bearing SEQ ID NO:1, contained, from the 5' to 3' end, a universal primer sequence (UPS), a unique identifier DNA sequence (UID), and an amplification sequence. Polymerase chain reactions (PCR) were performed in 25 uL reactions containing 7.25 uL of water, 0.125 uL of each primer, 12.5 uL of NEBNext Ultra II Q5 Master Mix (New England Biolabs catalog number M0544S), and 5 uL of DNA. The cycling conditions were 1 cycle at 98°C for 120 seconds, followed by 15 cycles at 98°C for 10 seconds, 57°C for 120 seconds, and 72°C for 120 seconds. For plasma experiments, 0.14 ng of DNA was used in 5 uL. A second PCR was then performed to add a dual index (barcode) to each PCR prior to sequencing. The forward and reverse primers used in the second PCR are listed in Table 2. The first amplification primers were not removed, and the amplified product from the first reaction was diluted 1:20, which was used directly in a second amplification using primers that annealed to the UPS site introduced by the first primer and also contained the 5' grafting sequence required for hybridization to the Illumina flow cell.

[0156] F indexes (e.g., sequences used to distinguish between samples) were introduced into each sample using a second reverse primer to enable subsequent multiplexed sequencing. Second-round PCR was performed in a 25-uL reaction containing 7.25 uL of water, 0.125 uL of each primer, 12.5 uL of NEBNext Ultra II Q5 Master Mix (New England Biolabs catalog number M0544S), and 5 uL of DNA containing 5% of the first-round PCR product. The cycling conditions were 1 cycle at 98°C for 120 seconds, followed by 15 cycles at 98°C for 10 seconds, 65°C for 15 seconds, and 72°C for 120 seconds. Amplification products were separated on an agarose gel to confirm amplification. Amplification products were purified using 1.2 volumes of AMPure XP beads and quantified by spectrophotometry, real-time PCR, or automated electrophoresis using an Agilent 2100 Bioanalyzer or Agilent TapeStation. All oligonucleotides were purchased from Integrated DNA Technology (Coralville, Iowa).

[0157] Sequencing and sequencing analysis Using Bowtie2, we aligned amplicon reads generated with each of the seven primer pairs to the human reference genome assembly GRC37 (Langmead et al. 2012). Primer pair 1 (primers with SEQ ID NO:1 and SEQ ID NO:10) uniquely aligned an average of 51.1% of all reads, with an average amplicon size of 88 bp (Figure 1A). The amplicon size shown in Figure 1A includes the 45-bp primers. For example, without the primers, amplicons have an average size of approximately 43 base pairs (Figure 1B). Primer pair 1 theoretically could amplify up to 745,184 uniquely aligned repeat elements, whereas the average sample contained an average of 350,000 repeat elements (see Figure 1C). While not wishing to be bound by theory, there were several potential reasons for the discrepancy between the potential and actual number of amplicons in plasma samples. (1) Polymorphisms within the sequence may have caused misalignments, resulting in "missing amplicons." (2) Polymorphisms within the primers may have prevented amplification. (3) Each amplicon may have different PCR efficiencies, resulting in low-efficiency amplicons being outcompeted during PCR. (4) Small DNA fragments may have been preferentially amplified, resulting in the absence of long amplicons (>100 bp). (5) The small size of DNA fragments within cell-free DNA may have prevented the presence of long amplicons within cell-free DNA. (6) The sequencing load used for these samples may not have been high enough to observe all amplicons, especially those with low PCR efficiencies. (7) Finally, some repetitive elements may not have been present in all individuals. 52,762 polymorphisms were identified within the amplicons generated by the primer pair SEQ ID NO:1 and SEQ ID NO:10. The average number of heterozygous sites in the test cohort of 1348 normal plasma samples and 883 plasma samples from cancer patients was 2200.These sites could be used to measure allelic imbalance, genetically identify samples, and determine whether the samples were accidentally mixed together. Using the same SNPs and synthesis experiments, we estimated that sample mixing could be detected when the amount of DNA from one sample was >4% of the amount of DNA from two samples in a given mixture.

[0158] statistical analysis Read-depth-based analysis methods are widely applied to whole-genome sequencing (WGS) protocols. Under the assumption that reads are uniformly and independently distributed, regions of normal copy number are predicted to follow a Poisson or normal distribution (Zhao et al., 2013 and Pirooznia et al., 2015). Amplicon-based protocols offer high coverage depth at relatively low cost, making them an attractive alternative to WGS. However, aligned reads from amplicon sequencing, such as those generated by the assays mentioned above, have different characteristics from those generated by WGS and WES. These reads are discontinuous because they are limited to a relatively small number of discrete loci. Furthermore, these reads are not randomly distributed, which makes it difficult to use statistical models of read depth coverage designed for WGS and WES. Within-sample aneuploidy detection (WALDO) is an algorithm specifically designed for amplicon-based aneuploidy detection (see, for example, Douville et al. PNAS 201 115(8):1871-1876). WALDO was applied to sequence reads (e.g., SINEs) mapped to the above genomic loci. A genome-wide aneuploidy score was used to identify whether a sample had aneuploidy.

[0159] The statistical principles underlying WALDO Unlike most conventional approaches to assessing copy number alterations, WALDO does not compare the normalized read counts from each chromosome arm in a test sample with the read rate of each chromosome arm in other samples. Such conventional comparisons are subject to batch effects and other artifacts associated with variables that are difficult to control. To evaluate whole-genome sequencing data, aneuploidy was detected by comparing the read counts within 5,344 genomic intervals, each containing 500 kb of sequence. The read counts within a 500 kb genomic interval within a sample were only compared with the read counts of other genomic intervals within the same sample (hence, denoted "Within-Sample" in WALDO). The previously described WALDO protocol was individually adjusted in this example, resulting in several analytical changes (see Figure 2). The modifications included a new normalization step, a new method for calling minor copy number changes of uncertain length, and an improved method for detecting genome-wide aneuploidy, as described below. These analytical improvements, combined with the increased genomic density of amplicons obtained using the primer pair SEQ ID NO:1 and SEQ ID NO:10, allowed for greater sensitivity and detection of local amplifications and deletions less than 1 Mb in size.

[0160] In euploid samples, the number of reads in each 500-kb genomic interval should be tracked together with the number of reads in certain other genomic regions. This is because amplicons within the interval are amplified to a similar extent in genomic intervals that are tracked together. Here, such genomic regions that are tracked together are referred to as "clusters." Clusters can be identified from sequencing data in euploid samples. In test samples, it is determined whether the number of reads within each genomic interval of each predefined cluster is within the predicted limits of other clusters in the same sample. If the reads within a genomic interval are outside the statistical predicted limits and there are many such outside the predicted limits on the same chromosome arm, the chromosome arm is classified as aneuploid. The statistical basis of this test has been described elsewhere (e.g., Douville et al. PNAS 201 115(8):1871-1876). Briefly, the number of reads is not randomly distributed across the genome, but the distribution of scaled reads within each cluster is approximately normal. A useful property of normal distributions is that the sum of multiple normal distributions is also normal, so the theoretical mean and variance of the total reads on each chromosome arm can be computed simply by summing the means and variances of all clusters represented on that chromosome arm.

[0161] WALDO also uses several other innovations that make it applicable to the analysis of PCR-generated amplicons from clinical samples. One example of such an innovation is the control of amplification bias due to the strong dependence of data on the size of the initial template. Another example is the use of machine learning algorithms (e.g., support vector machines (SVM)) that enable the detection of aneuploidy in samples containing low rates of neoplasia.

[0162] Normalization The improved WALDO method described in this example includes a new normalization method that reduces the amount of variation between samples.In this normalization, principal component analysis (PCA) is first performed on the control sequencing data.By PCA, the number of 500 kb genomic intervals is reduced from n=5,344 to a more manageable number.Using the control PCA coordinates, a model is created to predict whether a specific 500 kb interval will be amplified more efficiently or less efficiently in further samples based on its PCA coordinates. TIFF2026012754000024.tif11146

[0163] For each test sample, the sample was projected into PCA space and a correction factor was calculated for each 500 kb interval as a function of its PCA coordinate. After applying the correction factor to each 500 kb genomic interval, the test sample was matched to seven control samples based on the closest Euclidean distance of that 500 kb interval.

[0164] Creation of synthetic aneuploidy samples Data were selected from 84 likely euploid plasma samples, each containing at least 10 million reads and derived from normal WBC DNA. Synthetic aneuploid samples were created by adding (or subtracting) reads from several chromosome arms to (or from) reads from these normal DNA samples. Reads from 1, 10, 15, or 20 chromosome arms were added or subtracted from each sample. Additions and subtractions were designed to represent a neoplastic cell rate ranging from 0.5% to 1.5%, resulting in synthetic samples containing exactly 10 million reads. Reads from each chromosome arm were added or subtracted uniformly. For example, when modeling five missing chromosome arms, each was lost to the same degree, and we did not incorporate tumor heterogeneity into the model. Furthermore, we did not create synthetic samples containing more than three copies of any chromosome arm; for example, four copies of chromosome 3p. This simplistic approach did not comprehensively cover all biologically plausible aneuploidy events. However, by limiting the possible combinations of modified arms, sample generation became computationally trackable, and the resulting support vector machine performed well in practice. Synthetically generated samples, in which reads originating from only a single chromosome arm were added or subtracted, allowed us to estimate WALDO's performance when only a single chromosome arm of interest was gained or lost. The pseudocode for generating synthetic samples is shown in Figure 5.

[0165] Genome-wide aneuploidy measurement A two-class support vector machine (SVM) was trained to distinguish between euploid and aneuploid samples. The training set included 1,348 likely euploid negative-class plasma samples and 635 aneuploid samples from normal individuals with at least 2.5 million reads. The aneuploid class included a mixture of synthetic and true aneuploid samples. SVM training was performed in R using the e1071 package with a radial basis kernel and default parameters. Each sample had 39 Z-score features representing gains and losses of chromosome arms. During training, the positive class was randomly sampled to be 10% of the size of the negative class. The positive class was randomly sampled at a ratio of one synthetic sample to two true samples. This procedure was performed for 10 iterations. The final genome-wide aneuploidy score was the average of the raw SVM scores over the 10 iterations.

[0166] result The performance of this assay was assessed in a cohort of 1348 euploid plasma samples and 883 plasma samples from cancer patients (Table 3). Samples from cancer patients included breast, colorectal, esophageal, liver, lung, ovarian, pancreatic, and gastric cancers (Figure 3). Using a cutoff that yielded 99% specificity in the cohort of 1348 euploid samples, 49% of the plasma from cancer samples were found to be aneuploid.

[0167] Sample Exclusion Criteria To ensure that all samples included in the results section of this document were of high quality, several exclusion criteria were developed. First, samples with fewer than 2.5M reads were excluded. Second, samples with sufficient evidence of contamination were excluded. To be labeled as contaminated, samples had to have at least 10 significant allelically imbalanced chromosome arms (z score >= 2.5) and fewer than 10 significant chromosome arm gains or losses (z >= 2.5 or z <= -2.5). Allelic imbalance was determined by SNPs, while gains or losses were assessed by WALDO. Mixing experiments demonstrated that a relatively large number of allelically imbalanced chromosome arms, without significant gains or losses, indicated sample contamination with DNA from another individual. Third, for plasma analysis, samples with more than 8.5% of amplicons larger than 94 bp (50 base pairs between the forward and reverse primers) were excluded. Such samples were likely contaminated with leukocyte DNA. Fourth, samples outside the dynamic range of the assay as defined by the formula below were excluded. TIFF2026012754000025.tif14128

[0168] The distribution of this metric has a long tail. Values ​​>0.2450 and 0.2320 were selected as the dynamic range within which we could evaluate the cutoff. Fifth, plasma samples from the same patient with known aneuploidy in their leukocytes; we hypothesized that such patients had clonal hematopoiesis with undetermined potential (CHIP) or a congenital disorder.

[0169] Cancer detection using multi-analyte tests We investigated whether aneuploidy could be incorporated into the published framework as an additional biomarker, and compared the predictive ability of a logistic regression model with aneuploidy and protein markers with the original logistic regression model using somatic mutations and protein markers.

[0170] Here, 1348 plasma samples from healthy individuals and 883 cancer patients were analyzed. Of the 1348 healthy samples, only 248 overlapped with the original study. All 883 cancer samples were included in the original study. Sample demographic information is shown in Table 3.

[0171] A logistic regression model was trained using 812 original healthy samples (Cohen et al.) and 883 cancer samples, and performance was then assessed using ten 10-fold cross-validation runs. A complete list of samples and their biomarker values ​​is shown in Table 3. Because 564 original healthy samples were not analyzed for aneuploidy, the list of scores from 1348 normal samples was randomly sampled and an aneuploidy value was imputed for each missing sample. Ten runs of analysis were performed, and each new run randomly resampled the collection of scores from the 1348 normal samples and imputed a new score for the 564 samples.

[0172] To account for differences in the lower detection limit between different experiments, the 90th percentile feature value was used in the healthy training samples. Any feature values ​​below this threshold were set to the 90th percentile threshold. This transformation was performed for all training and test samples. This procedure was performed for aneuploidy scores, somatic mutation scores, and protein concentrations. The 90th percentile threshold and final feature coefficients from the logistic regression model are listed in Table 4.

[0173] Table 4. Logistic regression coefficients and thresholds TIFF2026012754000026.tif66128

[0174] Comparison of aneuploidy detection sensitivity with other cancer biomarkers Aneuploidy results were benchmarked against a driver gene mutation panel and a collection of seven protein markers (AFP, CA-125, CA15-3, CA19-9, CEA, HGF, OPN, and TIMP1) recently published as key biomarkers for cancer detection in plasma samples (Figure 4) (Cohen et al. 2018, Science 359(6378):926-930). Aneuploidy outperformed all protein markers. Furthermore, aneuploidy detected 42% of samples missed by mutations and 34% of samples missed by mutation panels and proteins. Given the high specificity of this aneuploidy assay and the availability of each additional cancer biomarker, it is conceivable that these components could be combined into a multianalyte test for cancer detection.

[0175] Example 2: Detection of aneuploidy in low-input DNA from trisomy 21 samples Preimplantation diagnostics and forensic applications require reliable detection of aneuploidies with only a few picograms (pg) of DNA. In preimplantation diagnostics, a few cells harvested from a blastocyst are used to assess copy number variation. For example, preimplantation diagnostics involves identifying a mammal as having an aneuploidy associated with Down syndrome. To test the detection limit for input DNA in the method featured in this disclosure, samples with aneuploidy associated with trisomy 21 were analyzed at input DNA concentrations ranging from 3 to 225 pg. The relationship between reads and DNA was relative to a negative control (water wells without DNA) and a euploid control of known concentration (Figure 6). Trisomy 21 aneuploidy was detected in each sample tested, even in those with 3 pg of input DNA, half the amount found in diploid cells. In the trisomy 21 samples, chromosome arms other than chromosome 21 were found to be non-aneuploid. In the euploid controls used in this experiment, the chromosome arm containing chromosome 21 was found to be non-aneuploid.

[0176] Example 3: Detection of aneuploidy in low-input DNA from biobank samples Low-input DNA samples from biobanks were assessed for either aneuploidy or identification purposes. The method described herein was applied to 793 plasma DNA samples stored in PCR plates for 10 years. The DNA volume in each well of the PCR plate was used in other experiments. Five microliters of water was added to the dry (empty) wells and then subjected to the method described herein. In 728 samples, more than 2.5 million aligned reads were sequenced, sufficient to reliably assess aneuploidy. In 768 of these samples, more than 1 million aligned reads were sequenced, sufficient to confirm the identity of the plasma DNA with other samples from the same donor.

[0177] Example 4: Detection of leukocyte DNA contamination in plasma samples Plasma cfDNA is often contaminated with DNA leaked from leukocytes during either blood collection or plasma preparation. This contaminating leukocyte DNA can reduce the sensitivity of aneuploidy testing in plasma samples because the leukocytes are not derived from fetal cells (in NIPT) or cancer cells (in liquid biopsy samples). Leukocyte genomic DNA (gDNA) has an average fragment size of >1000 bp, while plasma cell-free DNA has an average size of <160 bp. Given that smaller fragments are amplified more efficiently during PCR reactions, detection of contaminating leukocyte gDNA is challenging due to preferential amplification of shorter cfDNA fragments. Application of the method described herein enabled detection of contaminating leukocyte gDNA thanks to amplicons generated using primers SEQ ID NO:1 and SEQ ID NO:10. Using this method, we identified 1241 amplicons typically present in gDNA but absent in cfDNA. Sequencing of these amplicon reads demonstrated leukocyte contamination in plasma samples. By mixing leukocyte DNA with plasma cell-free DNA and using the methods described herein, we were able to detect samples containing >4% leukocyte DNA, as shown in Table 5.

[0178] Table 5. Prediction of gDNA contamination in plasma TIFF2026012754000027.tif83134

[0179] Example 5: Copy number analysis of uncertain length Copy number variants of uncertain length were detected. First, the logarithm of the ratio of the observed value in the test sample to the WALDO predicted value for each 500-kb interval on each chromosome arm was calculated. Using the logarithm of this ratio, a circular binary segmentation algorithm was applied to find copy number variants across each chromosome arm. Any copy number variants ≤5 Mb in size were flagged. These flagged CNVs were removed before calculating statistical significance on each chromosome arm. In general, small CNVs can be used to assess microdeletions or microamplifications, such as those occurring in DiGeorge syndrome (chromosome 22q11.2) or breast cancer (chromosome 17q12).

[0180] Example 6: Sensitivity of multi-analyte tests for cancer detection This example describes the sensitivity of different multi-analyte tests for cancer detection.

[0181] Three different multi-analyte tests were used to evaluate the sensitivity of detecting eight cancers in patient-derived plasma samples: breast, ovarian, liver, lung, pancreatic, esophageal, gastric, and colorectal cancers. The three tests were (1) a three-component test using assessment of aneuploidy status, somatic mutation analysis, and protein biomarkers; (2) a two-component test using assessment of aneuploidy status and somatic mutation analysis; and (3) a two-component test using assessment of aneuploidy status and protein biomarkers. The eight protein biomarkers and somatic mutations tested are as described in Cohen et al., Science 359, pp. 926-930, the entire contents of which are incorporated herein by reference.

[0182] As shown in Figures 7A-7B, the median sensitivity of the 3-component multi-analyte test for detecting ovarian, liver, lung, pancreatic, esophageal, gastric, and colorectal cancer was 80%, with a range of 77% to 97%. The sensitivity of the 3-component multi-analyte test for detecting breast cancer was 38%. Sensitivity was calculated using a threshold of 99% specificity.

[0183] Figure 8 further shows the true positive rates (a measure of sensitivity) for cancer detection using the following tests: (1) aneuploidy status; somatic mutations; and protein biomarkers; (2) aneuploidy status and protein biomarkers; (3) somatic mutations and protein biomarkers; (4) aneuploidy status and somatic mutations; (5) aneuploidy status; and (6) somatic mutations. Detection specificity remained at 99%.

[0184] As shown in Figure 8, the 3-component multi-analyte test (assessment of aneuploidy status, somatic mutation analysis, and protein biomarkers) detected cancer with a sensitivity of 73% at a specificity of 99%. The true positive rate (a measure of sensitivity) was highest for the 3-component multi-analyte test compared to the other tests.

[0185] As shown in Figure 9, multi-analyte testing (assessment of aneuploidy status and protein biomarkers) detected cancer with higher sensitivity than aneuploidy alone when looking at sample-based cancer staging.

[0186] Thus, the data disclosed in this example demonstrate that a three-component multi-analyte test involving assessment of aneuploidy status, somatic mutation analysis, and protein biomarkers can increase the sensitivity of cancer detection while maintaining high specificity for cancer detection.

[0187] Example 7: Determining somatic / germline status The materials and methods described herein may be used to identify somatic mutations in the sequence of a repetitive element amplified from a sample (e.g., a tumor sample or a non-tumor sample (i.e., a normal sample)). For example, when two types of samples, a non-tumor sample and a tumor sample, are available from the same patient, mutations present in one sample but absent in the other can be identified. For each sample, the number of somatic mutations can be counted, and the spectrum of single base substitutions (SBS) (e.g., A->T, A->C, etc.) can be measured. In addition, when a sample is analyzed by exome sequencing, the correlation between the number of SBS in the amplified repetitive element herein and the number of SBS in the exome can be examined. Therefore, somatic mutations in a sample can be identified using the materials and methods described herein.

[0188] Example 8: Sample Identification The materials and methods described herein may be used to identify and / or distinguish samples (e.g., distinguish a sample derived from one subject from a sample derived from a second subject). In such cases, the sample is identified based on the common polymorphisms present in the repeat elements amplified by the materials and methods described herein. The sample is then distinguished from other samples by comparing the sequences of the common polymorphisms between samples. Genotypes are assigned to samples by determining the genotype of each polymorphism in each amplicon. Genotypes can be compared between samples to identify the sample (e.g., distinguish a tumor sample from a non-tumor sample, or a sample derived from one subject from a sample derived from a different subject). Samples can be considered to be derived from different samples if the match rate (e.g., the percentage similarity between genotypes) is <0.99 and at least 5,000 amplicons have sufficient coverage.

[0189] Example 9: Detection of aneuploidy in different stages and types of cancer A set of experiments was conducted to assess the detection of aneuploidy in different stages and types of cancer. In this experiment, plasma from subjects with different stages of breast, colorectal, esophageal, liver, lung, ovarian, pancreatic, and gastric cancer was isolated according to the methods described herein. Figure 10 shows aneuploidy (at 99% specificity) in stage I (n=109), stage II (n=276), and stage III (173). Figure 11 shows aneuploidy (at 99% specificity) in the same cancers of Figure 7, but by cancer type (Figure 11) rather than cancer stage (Figure 10).

[0190] Using the Real-Seq method, aneuploidy was more commonly detected than mutations in plasma samples from cancer patients (49% and 34% of 883 samples, respectively; P<10-20, one-sided binomial test, Figure 19A). Regarding histology, aneuploidy was more commonly detected than mutations in samples from patients with esophageal, colorectal, pancreatic, lung, gastric, and breast cancer (all P<0.01), less commonly in ovarian cancer (P=0.048), and equally commonly in liver cancer (Figure 19A). Regarding stage, aneuploidy was more commonly detected than mutations in all stages, especially in stages I and II (Figure 19B, P<10-9).

[0191] Example 10: Detecting Cancer in Samples Using Aneuploidy and Protein Biomarkers A set of experiments was conducted to assess the sensitivity of cancer detection when aneuploidy detection is combined with protein biomarker detection as described herein. In these experiments, plasma from the same cohorts as in Example 8 (e.g., different stages of breast, colorectal, esophageal, liver, lung, ovarian, pancreatic, and gastric cancer) was assayed for aneuploidy and protein biomarkers. Figure 12 shows the detection sensitivity for different stages of cancer (stage I (n=109), stage II (n=276), and stage III (n=173)).

[0192] Example 11: Comparison of Real-Seq with other next-generation sequencing technologies A set of experiments was performed to assess the performance of Real-Seq relative to other next-generation sequencing technologies.

[0193] In the most common form of NIPT, the goal is to detect the gain or loss of a chromosome (e.g., chromosome 21 in Down syndrome). We assessed the performance of whole-genome sequencing (WGS), FAST-SeqS, and RealSeqS on samples with a DNA mixture typically found in noninvasive prenatal testing (NIPT), i.e., 5% fetal DNA. For this purpose, we used actual data from the three methods, but then added a predefined number of reads from various chromosomal regions of the same sample to simulate what would occur if aneuploidy were present in such regions. The pseudocode used to generate these in silico simulation samples is shown in Figures 13 and 14. Performance was calculated using the commonly used z-score, which compares the actual read rate for a particular chromosome arm to the mean read rate of a normal panel divided by the standard deviation of that normal panel. We report the total reads required for all three approaches, assuming single-end 100-bp reads and accounting for differences in alignment rates and typically used filtering criteria.

[0194] As shown in Figure 15A, RealSeqS consistently achieved high sensitivity at low sequencing loads. For example, RealSeqS had a sensitivity (at 99% specificity) of 99% for monosomy and trisomy at 5% cell rate, while WGS and FAST-SeqS had sensitivities of 94% and 81%, respectively (Figure 15A).

[0195] Another important aspect of copy number variation assays is the detection of relatively small deleted or amplified regions. For example, deletions in DiGeorge syndrome are often as small as 1.5 Mb. In data simulating a 5% deletion-containing cell rate, RealSeqS had a sensitivity of 75.0% (at 99% specificity) for the 1.5 Mb DiGeorge deletion, while WGS and FAST-SeqS had sensitivities of 19.0% and 29.0%, respectively (Figure 15B; and Figures 16A-16B).

[0196] Detection of amplifications, such as ERBB2 amplification in breast cancer, is important for determining whether a patient should be treated with trastuzumab or other targeted therapies. In this example, following the same protocol as described above, in silico simulation samples with a localized amplification (20 copies) of approximately 42 kb of the ERBB2 gene were generated for WGS, FAST-SeqS, and RealSeqS. RealSeqS detected the amplification in the in silico simulation samples with significantly fewer sequencing runs than WGS or Fast-SeqS. At a 1% cell rate, RealSeqS had a sensitivity of 91.0%, while WGS had a sensitivity of 50.0% (Figure 15C; and Figures 17A-17B).

[0197] This data demonstrates that the Real SEQ approach can detect small regions that have been amplified or deleted, and that the method has high sensitivity at low sequencing loads.

[0198] Example 12: Detection of aneuploidy in samples with low concentrations of tumor-derived DNA A set of experiments was performed to assess aneuploidy detection using the RealSEQ method in samples with varying concentrations of tumor-derived DNA. In an assessment of 302 samples in which the mutant allele rate had been measured by analysis of mutations present in plasma (Cohen et al., Science 359:926-930), aneuploidy was detected in 92% of 65 samples with a mutant allele rate of ≥2%, 71% of 65 samples with a mutant allele rate between 0.5% and 2%, and 49% of 172 samples with a mutant allele frequency ranging from 0.01% to 0.5% (Figure 18). The difference in aneuploidy between these three classes of samples was significant (P<10-3, one-sided binomial test).

[0199] This data indicates that the Real Seq method can detect aneuploidy even at low tumor DNA concentrations, and thus the sensitivity of aneuploidy detection is related to the circulating tumor DNA concentration in the sample.

[0200] Other Aspects While the present invention has been described in conjunction with a detailed description thereof, it will be understood that the foregoing description is intended to be illustrative and not to limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

[0201] (Table 3) TIFF2026012754000028.tif239165TIFF2026012754000029.tif240169TIFF2026012754000030.tif240169TIFF2026012754000031.tif239168TIFF2026012754000032.tif240169TIFF2026012754000033.tif240168TIFF2026012754000034.tif240169TIFF2026012754000035.tif240169TIFF2026012754000036.tif240169TIFF2026012754000037.tif240169TIFF2026012754000038.tif241168TIFF2026012754000039.tif241168TIFF2026012754000040.tif241169TIFF2026012754000041.tif239169TIFF2026012754000042.tif241169TIFF2026012754000043.tif240168TIFF2026012754000044.tif240169TIFF2026012754000045.tif240169TIFF2026012754000046.tif240169TIFF2026012754000047.tif241169TIFF2026012754000048.tif241169TIFF2026012754000049.tif241169TIFF2026012754000050.tif240169TIFF2026012754000051.tif241142

[0202] Array information SEQUENCE LISTING <110> The Johns Hopkins University <120> RAPID ANEUPLOIDY DETECTION <150> US 62 / 849,662 <151> 2019-05-17 <150> US 62 / 905,327 <151> 2019-09-24 <150> US 62 / 971,050 <151> 2020-02-06 <160> 19 <170> PatentIn version 3.5 <210> 1 <211> 57 <212> DNA <213> Artificial <220> <223> synthetic primer <220> <221> misc_feature <222> (22)..(37) <223> n is a, c, g, or t <400> 1 cgacgtaaaa cgacggccag tnnnnnnnnn nnnnnnnggt gaaaccccgt ctctaca 57 <210> 2 <211> 56 <212> DNA <213> Artificial <220> <223> synthetic primer <220> <221> misc_feature <222> (22)..(37) <223> n is a, c, g, or t <400> 2 cgacgtaaaa cgacggccag tnnnnnnnnn nnnnnnnggt gaaaccccgt ctctac 56 <210> 3 <211> 57 <212> DNA <213> Artificial <220> <223> synthetic primer <220> <221> misc_feature <222> (22)..(37) <223> n is a, c, g, or t <400> 3 cgacgtaaaa cgacggccag tnnnnnnnnn nnnnnnnggt gaaaccccgt ctctact 57 <210> 4 <211> 59 <212> DNA <213> Artificial <220> <223> synthetic primer <220> <221> misc_feature <222> (22)..(37) <223> n is a, c, g, or t <400> 4 cgacgtaaaa cgacggccag tnnnnnnnnn nnnnnnncat gcctgtagtc ccagctact 59 <210> 5 <211> 62 <212> DNA <213> Artificial <220> <223> synthetic primer <220> <221> misc_feature <222> (22)..(37) <223> n is a, c, g, or t <400> 5 cgacgtaaaa cgacggccag tnnnnnnnnn nnnnnnnata gtgaaacccc atctctacaa 60 aa 62 <210> 6 <211> 58 <212> DNA <213> Artificial <220> <223> synthetic primer <220> <221> misc_feature <222> (22)..(37) <223> n is a, c, g, or t <400> 6 cgacgtaaaa cgacggccag tnnnnnnnnn nnnnnnnggt gaaaccccat ctctacaa 58 <210> 7 <211> 61 <212> DNA <213> Artificial <220> <223> synthetic primer <220> <221> misc_feature <222> (22)..(37) <223> n is a, c, g, or t <400> 7 cgacgtaaaa cgacggccag tnnnnnnnnn nnnnnnnata gtgaaacccc atctctacaa 60 a 61 <210> 8 <211> 55 <212> DNA <213> Artificial <220> <223> synthetic primer <220> <221> misc_feature <222> (22)..(37) <223> n is a, c, g, or t <400> 8 cgacgtaaaa cgacggccag tnnnnnnnnnn nnnnnnngag gtgggaggat tgctt 55 <210> 9 <211> 55 <212> DNA <213> Artificial <220> <223> synthetic primer <220> <221> misc_feature <222> (22)..(37) <223> n is a, c, g, or t <400> 9 cgacgtaaaa cgacggccag tnnnnnnnnnn nnnnnnnacc agcctggggca acata 55 <210> 10 <211> 49 <212> DNA <213> Artificial <220> <223> synthetic primer <400> 10 cacacaggaa acagctatga ccatgcctcc taagtagctg ggactacag 49 <210> 11 <211> 49 <212> DNA <213> Artificial <220> <223> synthetic primer <400> 11 cacacaggaa acagctatga ccatgcctcc taagtagctg ggactacag 49 <210> 12 <211> 49 <212> DNA <213> Artificial <220> <223> synthetic primer <400> 12 cacacaggaa acagctatga ccatgcctcc taagtagctg ggactacag 49 <210> 13 <400> 13 000 <210> 14 <211> 60 <212> DNA <213> Artificial <220> <223> synthetic primer <400> 14 cacacaggaa acagctatga ccatgtgcag tggcacgatc atagctcact gcagccttga 60 <210> 15 <211> 44 <212> DNA <213> Artificial <220> <223> synthetic primer <400> 15 cacacaggaa acagctatga ccatgctccc gagtagctgg gact 44 <210> 16 <211> 46 <212> DNA <213> Artificial <220> <223> synthetic primer <400> 16 cacacaggaa acagctatga ccatgctccc gagtagctgg gactac 46 <210> 17 <211> 45 <212> DNA <213> Artificial <220> <223> synthetic primer <400> 17 cacacaggaa acagctatga ccatgcccga gtagctggga ctaca 45 <210> 18 <211> 42 <212> DNA <213> Artificial <220> <223> synthetic primer <400> 18 cacacaggaa acagctatga ccatgaggct ggagtgcagt gg 42 <210> 19 <211> 42 <212> DNA <213> Artificial <220> <223> synthetic primer <400> 19 cacacaggaa acagctatga ccatgccacc atgcctggct aa 42

Claims

1. 1. A method for identifying the presence of aneuploidy in a mammalian genome, comprising: (a) amplifying a plurality of chromosomal sequences in a cell-free DNA sample using a single primer pair complementary to the chromosomal sequences to form a plurality of amplicons, wherein the primer pair amplifies between 100,000 and 1,000,000 unique amplicons containing short interspersed element sequences (SINEs), and the plurality of amplicons comprises sequences from a plurality of different chromosomes; (b) determining at least a portion of the nucleic acid sequences of the plurality of amplicons to generate amplicon sequences; (c) mapping the amplicon sequences to a reference genome; (d) dividing the amplicon sequence into a plurality of genomic intervals; (e) quantifying the number of reads for the amplicon sequences mapped to the genomic interval; and (f) comparing the number of reads of amplicon sequences in a first genomic interval with the number of reads of amplicon sequences in one or more different genomic intervals; thereby identifying the presence of aneuploidy in the genome of the mammal; The method, wherein the presence of aneuploidy in the mammal's genome is identified when the number of reads of amplicon sequences in the first genomic interval is outside the predicted limits of the number of reads of amplicon sequences in one or more different genomic intervals.

2. The cell-free DNA sample contains multiple euploid DNA samples, or The cell-free DNA sample contains DNA of unknown ploidy.

10. The method of claim 1.

3. 2. The method of claim 1, wherein the cell-free DNA sample is derived from plasma or serum.

4. 4. The method of any one of claims 1 to 3, wherein the plurality of amplicons has an average size of 10 to 110 base pairs.

5. The method of any one of claims 1 to 4, wherein the mammal is a human.

6. 6. The method of any one of claims 1 to 5, wherein the primer pair amplifies between 200,000 and 1,000,000 unique amplicons.

7. 7. The method of any one of claims 1 to 6, wherein the primer pair amplifies between 300,000 and 1,000,000 unique amplicons.

8. 8. The method of any one of claims 1 to 7, wherein the plurality of amplicons comprises sequences on a plurality of different chromosomes.

9. 9. The method of any one of claims 1 to 8, wherein the genomic interval comprises between 100 nucleotides and 125,000,000 nucleotides.

10. 10. The method of any one of claims 1 to 9, wherein quantifying the number of reads for amplicon sequences mapped to a genomic interval comprises identifying a plurality of genomic intervals that have a shared amplicon feature, wherein the shared amplicon feature is the number of mapped amplicons or the average length of the mapped amplicons.

11. 11. The method of claim 10, wherein a plurality of genomic intervals having a shared amplicon feature are grouped into one or more clusters based on the shared amplicon feature.

12. 12. The method of claim 11, wherein each cluster comprises 200 genomic intervals.

Citation Information

Patent Citations

  • AM.2015