Methods and processes for non-invasive assessment of genetic variation

By optimizing genomic region partitioning and fetal fraction estimation through sequencing coverage analysis, the method enhances the accuracy of genetic variation and fetal DNA detection, addressing precision challenges in non-invasive genetic assessment.

JP7773301B2Active Publication Date: 2025-11-19SEQUENOM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2020187745
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2014-10-10
Filing Date
2020-11-11
Publication Date
2025-11-19
Estimated Expiration
2035-10-09

AI Technical Summary

Technical Problem

Current methods for non-invasive genetic variation assessment face challenges in accurately identifying genetic variations and fetal genetic information due to variability in sequencing coverage across the genome, which affects the precision of genetic diagnosis and prenatal testing.

Method used

A method involving partitioning genomic regions based on sequencing coverage variability, selecting initial portion lengths, and recalculating or repartitioning these regions to optimize the analysis, combined with fetal fraction estimation and nucleotide sequence read normalization, to enhance the detection of genetic variations and fetal DNA in maternal blood samples.

Benefits of technology

Improves the accuracy of genetic variation identification and fetal genetic analysis by optimizing genomic region partitioning and fetal fraction estimation, enabling precise detection of chromosomal abnormalities and genetic predispositions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007773301000022
    Figure 0007773301000022
  • Figure 0007773301000023
    Figure 0007773301000023
  • Figure 0007773301000024
    Figure 0007773301000024
Patent Text Reader

Abstract

To provide a method, process, and apparatus for non-invasive assessment of genetic variation.SOLUTION: Provided is a method for partitioning one or more genomic regions of a reference genome into a plurality of portion. The method comprises: (a) determining sequencing coverage variability across the reference genome; (b) selecting an initial portion length; (c) partitioning at least two genomic regions according to the initial portion length in (b); (d) comparing the sequencing coverage variability determined in (a) for each of the at least two genomic regions, thereby generating a comparison; (e) recalculating the number of portions for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized portion length; and (f) re-partitioning at least one of the genomic regions into a plurality of portions according to the optimized portion length in (e).SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Related patent applications) This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 062,748, filed October 10, 2014, entitled "METHODS AND PROCESSES FOR NON-INVASIVE ASSESSMENT OF GENETIC VARIATIONS," inventors Cosmin Deciu and Chen Zhao, and designated by attorney docket number SEQ-6081-PV, the entire contents of which are incorporated herein by reference for all purposes, including all text, tables, and figures.

[0002] (Field) The technology provided herein relates, in part, to methods, processes and devices for non-invasively assessing genetic variation. [Background technology]

[0003] (background) The genetic information of living organisms (e.g., animals, plants, and microorganisms) and other forms that replicate genetic information (e.g., viruses) is encoded in deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Genetic information is a sequence of nucleotides or modified nucleotides that represent the chemical or hypothetical primary structure of nucleic acids. In humans, the complete genome contains approximately 30,000 genes located on 24 chromosomes (see *The Human Genome*, T. Strachan, BIOS Scientific Publishers, 1992). Each gene encodes a specific protein, which, after expression via transcription and translation within living cells, performs a specific biochemical function. Many medical conditions are caused by one or more gene variations.Specific gene variations cause medical conditions, such as hemophilia, thalassemia, Duchenne muscular dystrophy (DMD), Huntington's disease (HD), Alzheimer's disease, and cystic fibrosis (CF) (Human Genome Mutations, DN Cooper and M. Krawczak, BIOS Publishers, 1993).Such genetic diseases can result from the addition, substitution, or deletion of a single nucleotide in the DNA of a specific gene.For example, certain congenital defects are caused by chromosomal abnormalities, also known as aneuploidy, such as trisomy 21 (Down syndrome), trisomy 13 (Patau syndrome), trisomy 18 (Edwards syndrome), trisomy 16, trisomy 22, monosomy X (Turner syndrome), and certain sex chromosome aneuploidies, such as Klinefelter syndrome (XXY). Another genetic variation is the sex of the fetus, which can often be determined based on the sex chromosomes X and Y. Some genetic variations can predispose an individual to or develop any of several diseases, such as diabetes, arteriosclerosis, obesity, various autoimmune diseases, and cancer (e.g., colorectal, breast, ovarian, lung). [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] The Human Genome, T. Strachan, BIOS Scientific Publishers, 1992 [Non-patent document 2] Human Genome Mutations, D.N. Cooper and M. Krawczak, BIOS Publishers, 1993 Summary of the Invention [Means for solving the problem]

[0005] Identifying one or more genetic variations or variances can lead to the diagnosis of a particular medical condition or the determination of a predisposition to such a condition. Identifying genetic variances can facilitate medical decisions and / or provide access to useful medical procedures. In certain embodiments, identifying one or more genetic variations or variances involves analyzing cell-free DNA. Cell-free DNA (CF-DNA) is composed of DNA fragments that arise from cell death and circulate in the peripheral blood. High concentrations of CF-DNA can be indicative of certain clinical conditions, such as cancer, trauma, burns, myocardial infarction, stroke, sepsis, infection, and other diseases. Furthermore, cell-free fetal DNA (CFF-DNA) can be detected in the maternal bloodstream and used for various non-invasive prenatal diagnostic methods.

[0006] (Abstract) In a particular aspect, a method is provided herein for partitioning one or more genomic regions of a reference genome into multiple parts, the method comprising the steps of: a) determining the variability of sequencing coverage across the reference genome; b) selecting an initial part length; c) partitioning at least two genomic regions according to the initial part lengths in (b); d) comparing the sequencing coverage variability determined in (a) for each of the at least two genomic regions, thereby generating a comparison; e) recalculating the number of parts for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized part length; and f) repartitioning at least one of the genomic regions into multiple parts according to the optimized part lengths in (e).

[0007] Also provided herein in certain embodiments is a method for partitioning one or more genomic regions of a reference genome into a plurality of portions, the method comprising: a) determining a variability in sequencing coverage across the reference genome; b) selecting an initial portion length; c) partitioning at least two genomic regions according to the lengths of the initial portions in (b); d) comparing the variability in sequencing coverage determined in (a) for each of the at least two genomic regions, thereby generating a comparison; and e) determining the number of portions for at least one of the genomic regions according to the comparison in (d). f) re-partitioning at least one of the genomic regions into a plurality of parts according to the optimized part lengths in (e), thereby generating a re-partitioned genomic region; g) estimating the fetal fraction for a test sample from a pregnant female having a fetus; h) determining the size of a minimal genomic region; and i) adjusting the number of parts to include at least two parts for each genomic region, thereby generating a refined re-partitioned genomic region.

[0008] Also provided herein in certain embodiments is a method for partitioning one or more genomic regions of a reference genome into a plurality of portions, the method comprising the steps of: a) determining a variability in sequencing coverage across the reference genome; b) selecting an initial portion length; c) partitioning at least two genomic regions according to the lengths of the initial portions in (b); d) comparing the variability in sequencing coverage determined in (a) for each of the at least two genomic regions, thereby generating a comparison; and e) recalculating the number of portions for at least one of the genomic regions according to the comparison in (d), thereby generating an optimized number of portions. f) repartitioning at least one of the genomic regions into a plurality of parts according to the optimized part lengths in (e), thereby generating a repartitioned genomic region; g) determining a region-specific fetal fraction for each genomic region according to the correlation between the nucleotide sequence read counts per part and a weighting coefficient; h) determining the size of a local minimum genomic region; and i) adjusting the number of parts to include at least two parts for each genomic region, thereby generating a refined repartitioned genomic region.

[0009] Also provided herein in certain embodiments is a method for partitioning one or more genomic regions of a reference genome into a plurality of portions, the method comprising the steps of: a) determining a variability in sequencing coverage across the reference genome; b) selecting an initial portion length; c) partitioning at least two genomic regions according to the initial portion lengths in (b); d) comparing the sequencing coverage variability determined in (a) for each of the at least two genomic regions, thereby generating a comparison; e) recalculating the number of portions for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized portion length; and f) partitioning the genomic regions. A method is also presented that includes the steps of: repartitioning at least one of the regions into multiple parts according to the optimized part lengths in (e), thereby generating a repartitioned genomic region; g) estimating the fetal fraction for a test sample derived from a pregnant female bearing a fetus; h) determining a region-specific fetal fraction for each genomic region according to the correlation between the counts of nucleotide sequence reads per part and a weighting coefficient; i) determining the size of a local minimum genomic region; and j) adjusting the number of parts for each genomic region to include at least two parts, thereby generating a refined repartitioned genomic region.

[0010] Also provided herein in certain aspects is a method for partitioning one or more genomic regions of a reference genome into multiple parts, the method comprising the steps of: a) determining the variability of sequencing coverage across the reference genome; b) selecting the length of the initial parts; c) partitioning at least two genomic regions according to the lengths of the initial parts in (b); d) determining a region-specific fetal fraction for each genomic region according to the correlation between the count of nucleotide sequence reads per part and a weighting coefficient; e) determining the size of a local minimum genomic region; and f) adjusting the number of parts for each genomic region to include at least two parts, thereby generating a re-partitioned genomic region.

[0011] Also provided herein in certain aspects is a method for partitioning a reference genome, or a part thereof, into a plurality of parts, the method comprising: a) generating a guanine and cytosine (GC) profile for the reference genome, or a part thereof; b) applying a segmentation process to the GC profile generated in (a), thereby resulting in individual segments; and c) partitioning the reference genome, or a part thereof, into a plurality of parts according to the individual segments provided in (b), thereby generating a GC-partitioned reference genome, or a part thereof.

[0012] Certain particular embodiments are further described in the following description, examples, claims, and drawings.

[0013] The drawings illustrate, but are not limiting, certain embodiments of the present technology. For clarity and ease of description, the drawings have not been made to scale and in some instances various aspects may be shown exaggerated or enlarged to facilitate an understanding of particular embodiments. The present invention provides, for example, the following items. (Item 1) 1. A method for partitioning one or more genomic regions of a reference genome into multiple portions, comprising: a) determining the variability of sequencing coverage across a reference genome; b) selecting the length of the initial portion; c) partitioning at least two genomic regions according to the length of the initial portions in (b); d) comparing the variability of the sequencing coverage determined in (a) for each of the at least two genomic regions, thereby generating a comparison; e) recalculating the number of segments for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized segment length; f) repartitioning at least one of the genomic regions into a plurality of portions according to the lengths of the optimized portions in (e); A method comprising: (Item 2) 1. A method for identifying the presence or absence of a genetic variation, comprising quantifying nucleotide sequence reads for a test sample, wherein the sequence reads comprise: a) determining the variability of sequencing coverage across the reference genome; b) selecting the length of the initial portion; c) partitioning at least two genomic regions according to the length of the initial portions in (b); d) comparing the variability of the sequencing coverage determined in (a) for each of the at least two genomic regions, thereby generating a comparison; e) recalculating the number of segments for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized segment length; f) repartitioning at least one of the genomic regions into a plurality of portions according to the lengths of the optimized portions in (e); The method comprises mapping the genomic region of a reference genome to one or more genomic regions of the reference genome, the genomic region being partitioned by a process comprising: (Item 3) 3. The method of claim 1 or 2, wherein the step of determining the variability of sequencing coverage in (a) comprises using a training set of nucleotide sequence reads mapped to portions of a reference genome, wherein the sequence reads are circulating cell-free nucleic acid reads from a plurality of samples from a pregnant female carrying a fetus. (Item 4) Item 4. The method according to item 3, wherein the length of the initial portion in (b) is selected for the training set according to sequencing depth. (Item 5) 5. The method according to item 3 or 4, wherein the length of the initial portion in (b) is selected for the training set according to an average fetal fraction. (Item 6) 6. The method of claim 5, wherein the average fetal fraction is determined using the training set. (Item 7) 7. The method according to any one of items 1 to 6, wherein the length of the initial portion is between about 1 kb and about 1000 kb. (Item 8) 8. The method of any one of items 1 to 7, wherein the length of the initial portion is about 30 kb. (Item 9) 8. The method of any one of items 1 to 7, wherein the length of the initial portion is about 40 kb. (Item 10) 8. The method of any one of items 1 to 7, wherein the length of the initial portion is about 50 kb. (Item 11) 8. The method of any one of items 1 to 7, wherein the length of the initial portion is not 50 kb. (Item 12) 8. The method of any one of items 1 to 7, wherein the length of the initial portion is about 60 kb. (Item 13) 8. The method of any one of items 1 to 7, wherein the length of the initial portion is about 70 kb. (Item 14) 14. The method of any one of items 1 to 13, wherein the total number of portions for the genome is determined according to the length of the initial portions in (b). (Item 15) 15. The method of any one of items 1 to 14, wherein the at least two genomic regions comprise a first genomic region and a second genomic region. (Item 16) Item 16. The method of item 15, wherein the first genomic region and the second genomic region are substantially similar in size. (Item 17) The step of comparing the variability of the sequencing coverage in (d) comprises calculating a proportionality coefficient (P) according to the following formula: P=(var1 / var2) 1 / 3 Formula A where var1 is the variability of the sequencing coverage of the first genomic region and var2 is the variability of the sequencing coverage of the second genomic region. 17. The method according to item 15 or 16, comprising calculating according to (Item 18) 18. The method of claim 17, wherein the variability of the sequencing coverage of the first genomic region is determined from the counts of nucleotide sequence reads for the first genomic region, or a derivative thereof, and the variability of the sequencing coverage of the second genomic region is determined from the counts of nucleotide sequence reads for the second genomic region, or a derivative thereof. (Item 19) 18. The method of claim 17, wherein the sequencing coverage variability of the first genomic region is determined from an average nucleotide sequence read count for the first genomic region, or a derivative thereof, and the sequencing coverage variability of the second genomic region is determined from an average nucleotide sequence read count for the second genomic region, or a derivative thereof. (Item 20) 20. The method of claim 19, wherein the average nucleotide sequencing read counts for each genomic region are determined using the training set. (Item 21) Item 19. The method of item 18, wherein the nucleotide sequence read counts are normalized nucleotide sequence read counts. (Item 22) 21. The method of claim 19 or 20, wherein the average nucleotide sequence read count is an average normalized nucleotide sequence read count. (Item 23) 23. The method of any one of items 17 to 22, wherein the step of recalculating the number of portions for at least one of the genomic regions in (e) is performed according to the proportionality coefficient and the total number of portions determined from the lengths of the initial portions in (b). (Item 24) 24. The method of any one of items 1 to 23, wherein the plurality of portions in (f) comprises portions of a uniform size. (Item 25) 24. The method of any one of items 1 to 23, wherein the plurality of portions in (f) comprises portions of varying sizes. (Item 26) 26. The method according to item 24 or 25, wherein the plurality of portions in (f) comprises portions between about 1 kb and about 1000 kb in length. (Item 27) 26. The method of item 24 or 25, wherein the multiple portions in (f) comprise portions of about 30 kb. (Item 28) 26. The method of item 24 or 25, wherein the multiple portions in (f) comprise a portion of about 40 kb. (Item 29) 26. The method of item 24 or 25, wherein the multiple portions in (f) comprise portions of about 50 kb. (Item 30) 26. The method of claim 24 or 25, wherein the plurality of portions in (f) does not include a 50 kb portion. (Item 31) 26. The method of item 24 or 25, wherein the multiple portions in (f) comprise a portion of about 60 kb. (Item 32) 26. The method of item 24 or 25, wherein the multiple portions in (f) comprise a portion of about 70 kb. (Item 33) 33. The method of any one of items 1 to 32, comprising sequencing nucleic acid from the test sample by a nucleotide sequencing process to generate nucleotide sequence reads. (Item 34) 34. The method of claim 33, wherein the nucleic acid is circulating cell-free nucleic acid from a pregnant female carrying a fetus. (Item 35) 35. The method of any one of items 1 to 34, comprising mapping nucleotide sequence reads from the test sample to portions of the repartitioned reference genome, thereby generating mapped nucleotide sequence reads. (Item 36) 36. The method of claim 35, comprising normalizing the counts of the mapped nucleotide sequence reads, thereby generating normalized counts. (Item 37) 37. The method of claim 36, wherein the normalizing step comprises LOESS normalization for guanine and cytosine (GC) bias (GC-LOESS normalization). (Item 38) 38. The method of claim 36 or 37, wherein the normalizing step comprises adjusting the counts of sequence reads according to a median count. (Item 39) 39. The method of claim 38, wherein the sequence read counts are adjusted according to the median fraction counts. (Item 40) 40. The method of any one of items 36 to 39, wherein the normalizing step comprises principal component normalization. (Item 41) 41. The method of any one of items 36 to 40, wherein the normalizing step comprises GC-LOESS normalization, followed by normalization according to fractional counts of the median, followed by principal component normalization. (Item 42) 42. The method of any one of items 36 to 41, comprising determining the presence or absence of a genetic variation for the test sample according to the normalized counts. (Item 43) 43. The method of any one of items 36 to 42, comprising determining a chromosome structure according to the normalized counts. (Item 44) 44. The method of any one of items 36 to 43, wherein the normalized counts represent chromosome dosages for the test sample. (Item 45) 45. The method of claim 44, wherein determining the presence or absence of a genetic variation is according to the chromosomal dosage. (Item 46) 46. ​​The method of any one of items 42 to 45, wherein determining the presence or absence of a genetic variation in the test sample comprises identifying the presence or absence of one copy of a chromosome, two copies of a chromosome, three copies of a chromosome, four copies of a chromosome, five copies of a chromosome, a deletion of one or more segments of a chromosome, or an insertion of one or more segments of a chromosome. (Item 47) 36. The method of claim 35, which does not include a step of normalizing the counts of the mapped nucleotide sequence reads. (Item 48) 48. The method of claim 47, comprising determining the presence or absence of a genetic variation for the test sample according to the raw counts of the mapped nucleotide sequence reads. (Item 49) 49. The method of claim 47 or 48, comprising determining chromosome structure according to the raw counts of the mapped nucleotide sequence reads. (Item 50) 50. The method of claim 48 or 49, wherein the raw counts represent chromosome dosages for the test sample. (Item 51) 51. The method of claim 50, wherein determining the presence or absence of a genetic variation is according to the chromosomal dosage. (Item 52) 52. The method of any one of items 48 to 51, wherein determining the presence or absence of a genetic variation in the test sample comprises identifying the presence or absence of one copy of a chromosome, two copies of a chromosome, three copies of a chromosome, four copies of a chromosome, five copies of a chromosome, a deletion of one or more segments of a chromosome, or an insertion of one or more segments of a chromosome. (Item 53) 1. A method for partitioning one or more genomic regions of a reference genome into multiple portions, comprising: a) determining the variability of sequencing coverage across a reference genome; b) selecting the length of the initial portion; c) partitioning at least two genomic regions according to the length of the initial portions in (b); d) comparing the variability of the sequencing coverage determined in (a) for each of the at least two genomic regions, thereby generating a comparison; e) recalculating the number of segments for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized segment length; f) repartitioning at least one of the genomic regions into a plurality of portions according to the lengths of the optimized portions in (e), thereby generating a repartitioned genomic region; g) estimating the fetal fraction for test samples derived from pregnant females bearing fetuses; h) determining the size of the smallest genomic region; i) adjusting the number of portions to include at least two portions for each genomic region, thereby generating refined, re-partitioned genomic regions; A method comprising: (Item 54) 1. A method for identifying the presence or absence of a genetic variation, comprising quantifying nucleotide sequence reads for a test sample, wherein the sequence reads comprise: a) determining the variability of sequencing coverage across the reference genome; b) selecting the length of the initial portion; c) partitioning at least two genomic regions according to the length of the initial portions in (b); d) comparing the variability of the sequencing coverage determined in (a) for each of the at least two genomic regions, thereby generating a comparison; e) recalculating the number of segments for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized segment length; f) repartitioning at least one of the genomic regions into a plurality of portions according to the lengths of the optimized portions in (e), thereby generating a repartitioned genomic region; g) estimating the fetal fraction for test samples derived from pregnant females bearing fetuses; h) determining the size of the smallest genomic region; and i) adjusting the number of segments to include at least two segments for each genomic region, thereby generating refined, re-partitioned genomic regions; The method comprises mapping the genomic region of a reference genome to one or more genomic regions of the reference genome, the genomic region being partitioned by a process comprising: (Item 55) 55. The method of claim 53 or 54, wherein the step of determining the variability of sequencing coverage in (a) comprises using a training set of nucleotide sequence reads mapped to portions of a reference genome, wherein the sequence reads are circulating cell-free nucleic acid reads from a plurality of samples from a pregnant female carrying a fetus. (Item 56) Item 56. The method of item 55, wherein the length of the initial portion in (b) is selected for the training set according to sequencing depth. (Item 57) 57. The method of claim 55 or 56, wherein the length of the initial portion in (b) is selected for the training set according to an average fetal fraction. (Item 58) 58. The method of claim 57, wherein the average fetal fraction is determined using the training set. (Item 59) 59. The method of any one of items 53 to 58, wherein the length of the initial portion is between about 1 kb and about 1000 kb. (Item 60) 60. The method of any one of items 53 to 59, wherein the length of the initial portion is about 30 kb. (Item 61) 60. The method of any one of items 53 to 59, wherein the length of the initial portion is about 40 kb. (Item 62) 60. The method of any one of items 53 to 59, wherein the length of the initial portion is about 50 kb. (Item 63) 60. The method of any one of items 53 to 59, wherein the length of the initial portion is not 50 kb. (Item 64) 60. The method of any one of items 53 to 59, wherein the length of the initial portion is about 60 kb. (Item 65) 60. The method of any one of items 53 to 59, wherein the length of the initial portion is about 70 kb. (Item 66) 66. The method of any one of items 53 to 65, wherein the total number of portions for the genome is determined according to the length of the initial portions in (b). (Item 67) 67. The method of any one of items 53 to 66, wherein the at least two genomic regions comprise a first genomic region and a second genomic region. (Item 68) 68. The method of claim 67, wherein the first genomic region and the second genomic region are substantially similar in size. (Item 69) The step of comparing the variability of the sequencing coverage in (d) comprises calculating a proportionality coefficient (P) according to the following formula: P=(var1 / var2) 1 / 3 Formula A where var1 is the variability of the sequencing coverage of the first genomic region and var2 is the variability of the sequencing coverage of the second genomic region. 69. The method according to item 67 or 68, comprising calculating according to (Item 70) 70. The method of claim 69, wherein the variability of the sequencing coverage of the first genomic region is determined from the counts of nucleotide sequence reads for the first genomic region, or a derivative thereof, and the variability of the sequencing coverage of the second genomic region is determined from the counts of nucleotide sequence reads for the second genomic region, or a derivative thereof. (Item 71) 71. The method of claim 70, wherein the sequencing coverage variability of the first genomic region is determined from an average nucleotide sequence read count for the first genomic region, or a derivative thereof, and the sequencing coverage variability of the second genomic region is determined from an average nucleotide sequence read count for the second genomic region, or a derivative thereof. (Item 72) 72. The method of claim 71, wherein the average nucleotide sequencing read counts for each genomic region are determined using the training set. (Item 73) 71. The method of claim 70, wherein the nucleotide sequence read counts are normalized nucleotide sequence read counts. (Item 74) 73. The method of claim 71 or 72, wherein the average nucleotide sequence read count is an average normalized nucleotide sequence read count. (Item 75) 75. The method of any one of items 69 to 74, wherein the step of recalculating the number of portions for at least one of the genomic regions in (e) is performed according to the proportionality coefficient and the total number of portions determined from the lengths of the initial portions in (b). (Item 76) 76. The method of any one of items 53 to 75, wherein the plurality of portions in (f) comprises portions of a uniform size. (Item 77) 76. The method of any one of items 53 to 75, wherein the plurality of portions in (f) comprises portions of varying sizes. (Item 78) 78. The method according to item 76 or 77, wherein the plurality of portions in (f) comprises portions between about 1 kb and about 1000 kb in length. (Item 79) 78. The method of item 76 or 77, wherein the plurality of portions in (f) comprises a portion of about 30 kb. (Item 80) 78. The method of item 76 or 77, wherein the plurality of portions in (f) comprises a portion of about 40 kb. (Item 81) 78. The method of item 76 or 77, wherein the plurality of portions in (f) comprises a portion of about 50 kb. (Item 82) 78. The method of claim 76 or 77, wherein the plurality of portions in (f) does not include a 50 kb portion. (Item 83) 78. The method of item 76 or 77, wherein the plurality of portions in (f) comprises a portion of about 60 kb. (Item 84) 78. The method of item 76 or 77, wherein the plurality of portions in (f) comprises a portion of about 70 kb. (Item 85) 85. The method of any one of items 53 to 84, wherein the step of estimating the fetal fraction in (g) comprises determining an error value. (Item 86) 86. The method of any one of items 53 to 85, wherein the step of determining the size of the smallest genomic region in (h) comprises determining the size of the smallest genomic region detectable for a sample having the estimated fetal fraction in (g). (Item 87) 87. The method of item 86, wherein the size of the smallest genomic region is determined according to the upper 95% confidence interval of the fetal fraction. (Item 88) (j) re-estimating the fetal fraction from the refined, re-partitioned genomic regions. (Item 89) 89. The method of claim 88, comprising comparing the estimated fetal fraction in (g) with the re-estimated fetal fraction in (j). (Item 90) 90. The method of claim 89, comprising repeating parts (g), (h), and (i) if the estimated fetal fraction in (g) differs from the re-estimated fetal fraction in (j) by a predetermined tolerance value. (Item 91) Item 91. The method according to item 90, wherein the predetermined tolerance value is between about 1% and about 25%. (Item 92) 92. The method of any one of items 53 to 91, comprising sequencing nucleic acid from the test sample by a nucleotide sequencing process to generate nucleotide sequence reads. (Item 93) 93. The method of claim 92, wherein the nucleic acid is circulating cell-free nucleic acid from a pregnant female carrying a fetus. (Item 94) 94. The method of any one of items 53 to 93, comprising mapping nucleotide sequence reads from the test sample to portions of the refined, repartitioned reference genome, thereby generating mapped nucleotide sequence reads. (Item 95) 95. The method of claim 94, comprising normalizing the counts of the mapped nucleotide sequence reads, thereby generating normalized counts. (Item 96) Item 96. The method of item 95, wherein the normalizing step comprises LOESS normalization for guanine and cytosine (GC) bias (GC-LOESS normalization). (Item 97) 97. The method of claim 95 or 96, wherein the normalizing step comprises adjusting the counts of sequence reads according to a median count. (Item 98) 98. The method of item 97, wherein the counts of the sequence reads are adjusted according to the median fractional counts. (Item 99) 99. The method of any one of items 95 to 98, wherein the normalizing step comprises principal component normalization. (Item 100) 100. The method of any one of items 95 to 99, wherein the normalizing step comprises GC-LOESS normalization, followed by normalization according to fractional counts of the median, followed by principal component normalization. (Item 101) 101. The method of any one of items 95 to 100, comprising determining the presence or absence of a genetic variation for the test sample according to the normalized counts. (Item 102) 102. The method of any one of items 95 to 101, comprising determining a chromosome structure according to the normalized counts. (Item 103) 103. The method of any one of paragraphs 95 to 102, wherein the normalized counts represent chromosome dosages for the test sample. (Item 104) 104. The method of claim 103, wherein determining the presence or absence of genetic variation is according to the chromosomal dosage. (Item 105) 105. The method of any one of items 101 to 104, wherein determining the presence or absence of a genetic variation in the test sample comprises identifying the presence or absence of one copy of a chromosome, two copies of a chromosome, three copies of a chromosome, four copies of a chromosome, five copies of a chromosome, a deletion of one or more segments of a chromosome, or an insertion of one or more segments of a chromosome. (Item 106) 95. The method of item 94, which does not include a step of normalizing the counts of the mapped nucleotide sequence reads. (Item 107) 107. The method of claim 106, comprising determining the presence or absence of a genetic variation for the test sample according to the raw counts of the mapped nucleotide sequence reads. (Item 108) 108. The method of claim 106 or 107, comprising determining chromosome structure according to raw counts of the mapped nucleotide sequence reads. (Item 109) 109. The method of claim 107 or 108, wherein the raw counts represent chromosome doses for the test sample. (Item 110) 110. The method of claim 109, wherein determining the presence or absence of a genetic variation is according to the chromosomal dosage. (Item 111) 111. The method of any one of items 107 to 110, wherein determining the presence or absence of a genetic variation in the test sample comprises identifying the presence or absence of one copy of a chromosome, two copies of a chromosome, three copies of a chromosome, four copies of a chromosome, five copies of a chromosome, a deletion of one or more segments of a chromosome, or an insertion of one or more segments of a chromosome. (Item 112) 1. A method for partitioning one or more genomic regions of a reference genome into multiple portions, comprising: a) determining the variability of sequencing coverage across a reference genome; b) selecting the length of the initial portion; c) partitioning at least two genomic regions according to the length of the initial portions in (b); d) comparing the variability of the sequencing coverage determined in (a) for each of the at least two genomic regions, thereby generating a comparison; e) recalculating the number of segments for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized segment length; f) repartitioning at least one of the genomic regions into a plurality of portions according to the lengths of the optimized portions in (e), thereby generating a repartitioned genomic region; g) determining a region-specific fetal fraction for each genomic region according to the correlation between the count number of nucleotide sequence reads per portion and a weighting coefficient; h) determining the size of a local minimal genomic region; i) adjusting the number of portions to include at least two portions for each genomic region, thereby generating refined, re-partitioned genomic regions; A method comprising: (Item 113) 1. A method for identifying the presence or absence of a genetic variation, comprising quantifying nucleotide sequence reads for a test sample, wherein the sequence reads comprise: a) determining the variability of sequencing coverage across the reference genome; b) selecting the length of the initial portion; c) partitioning at least two genomic regions according to the length of the initial portions in (b); d) comparing the variability of the sequencing coverage determined in (a) for each of the at least two genomic regions, thereby generating a comparison; e) recalculating the number of segments for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized segment length; f) repartitioning at least one of the genomic regions into a plurality of portions according to the lengths of the optimized portions in (e), thereby generating a repartitioned genomic region; g) determining the region-specific fetal fraction for each genomic region according to the correlation between the count number of nucleotide sequence reads per portion and the weighting coefficient; h) determining the size of the local smallest genomic region; and i) adjusting the number of segments to include at least two segments for each genomic region, thereby generating refined, re-partitioned genomic regions; The method comprises mapping the genomic region of a reference genome to one or more genomic regions of the reference genome, the genomic region being partitioned by a process comprising: (Item 114) 114. The method of claim 112 or 113, wherein the step of determining the variability of sequencing coverage in (a) comprises using a training set of nucleotide sequence reads mapped to portions of a reference genome, wherein the sequence reads are circulating cell-free nucleic acid reads from a plurality of samples from a pregnant female carrying a fetus. (Item 115) Item 115. The method of item 114, wherein the length of the initial portion in (b) is selected for the training set according to sequencing depth. (Item 116) 116. The method of claim 114 or 115, wherein the length of the initial portion in (b) is selected for the training set according to an average fetal fraction. (Item 117) 117. The method of claim 116, wherein the average fetal fraction is determined using the training set. (Item 118) 118. The method of any one of items 112 to 117, wherein the length of the initial portion is between about 1 kb and about 1000 kb. (Item 119) 119. The method of any one of items 112 to 118, wherein the length of the initial portion is about 30 kb. (Item 120) 119. The method of any one of items 112 to 118, wherein the length of the initial portion is about 40 kb. (Item 121) 119. The method of any one of items 112 to 118, wherein the length of the initial portion is about 50 kb. (Item 122) 119. The method of any one of items 112 to 118, wherein the length of the initial portion is not 50 kb. (Item 123) 119. The method of any one of items 112 to 118, wherein the length of the initial portion is about 60 kb. (Item 124) 119. The method of any one of items 112 to 118, wherein the length of the initial portion is about 70 kb. (Item 125) 125. The method of any one of items 112 to 124, wherein the total number of portions for the genome is determined according to the length of the initial portions in (b). (Item 126) 126. The method of any one of items 112 to 125, wherein the at least two genomic regions comprise a first genomic region and a second genomic region. (Item 127) 127. The method of claim 126, wherein the first genomic region and the second genomic region are substantially similar in size. (Item 128) The step of comparing the variability of the sequencing coverage in (d) comprises calculating a proportionality coefficient (P) according to the following formula: P=(var1 / var2) 1 / 3 Formula A where var1 is the variability of the sequencing coverage of the first genomic region and var2 is the variability of the sequencing coverage of the second genomic region. Item 128. The method according to item 126 or 127, comprising calculating according to (Item 129) 129. The method of claim 128, wherein the variability of the sequencing coverage of the first genomic region is determined from the counts of nucleotide sequence reads for the first genomic region, or a derivative thereof, and the variability of the sequencing coverage of the second genomic region is determined from the counts of nucleotide sequence reads for the second genomic region, or a derivative thereof. (Item 130) 129. The method of claim 128, wherein the sequencing coverage variability of the first genomic region is determined from an average nucleotide sequence read count for the first genomic region, or a derivative thereof, and the sequencing coverage variability of the second genomic region is determined from an average nucleotide sequence read count for the second genomic region, or a derivative thereof. (Item 131) 131. The method of claim 130, wherein the average nucleotide sequencing read counts for each genomic region are determined using the training set. (Item 132) 130. The method of claim 129, wherein the nucleotide sequence read counts are normalized nucleotide sequence read counts. (Item 133) 132. The method of claim 130 or 131, wherein the average nucleotide sequence read count is an average normalized nucleotide sequence read count. (Item 134) 134. The method of any one of items 128 to 133, wherein the step of recalculating the number of portions for at least one of the genomic regions in (e) is performed according to the proportionality coefficient and the total number of portions determined from the lengths of the initial portions in (b). (Item 135) 135. The method of any one of items 112 to 134, wherein the plurality of portions in (f) comprises portions of a uniform size. (Item 136) 135. The method of any one of items 112 to 134, wherein the plurality of portions in (f) comprises portions of varying sizes. (Item 137) 137. The method according to item 135 or 136, wherein the plurality of portions in (f) comprises portions between about 1 kb and about 1000 kb in length. (Item 138) 137. The method of item 135 or 136, wherein the plurality of portions in (f) comprises a portion of about 30 kb. (Item 139) 137. The method of item 135 or 136, wherein the plurality of portions in (f) comprises a portion of about 40 kb. (Item 140) 137. The method of item 135 or 136, wherein the plurality of portions in (f) comprises a portion of about 50 kb. (Item 141) 137. The method of claim 135 or 136, wherein the plurality of portions in (f) does not include the 50 kb portion. (Item 142) 137. The method of item 135 or 136, wherein the plurality of portions in (f) comprises a portion of about 60 kb. (Item 143) 137. The method of item 135 or 136, wherein the plurality of portions in (f) comprises a portion of about 70 kb. (Item 144) 144. The method of any one of items 112 to 143, wherein the step of determining the size of the local smallest genomic region in (h) comprises determining the size of the local genomic region detectable for a sample having an average fetal fraction. (Item 145) 145. The method of any one of items 112 to 144, comprising sequencing nucleic acid from the test sample by a nucleotide sequencing process to generate nucleotide sequence reads. (Item 146) 146. The method of claim 145, wherein the nucleic acid is circulating cell-free nucleic acid from a pregnant female carrying a fetus. (Item 147) 147. The method of any one of paragraphs 112 to 146, comprising mapping nucleotide sequence reads from the test sample to portions of the refined and repartitioned reference genome, thereby generating mapped nucleotide sequence reads. (Item 148) 148. The method of claim 147, comprising normalizing the counts of the mapped nucleotide sequence reads, thereby generating normalized counts. (Item 149) Item 149. The method of item 148, wherein the normalizing step comprises LOESS normalization for guanine and cytosine (GC) bias (GC-LOESS normalization). (Item 150) 150. The method of claim 148 or 149, wherein the normalizing step comprises adjusting the counts of sequence reads according to a median count. (Item 151) 151. The method of claim 150, wherein the counts of the sequence reads are adjusted according to the median fractional counts. (Item 152) 152. The method of any one of items 148 to 151, wherein the normalizing step comprises principal component normalization. (Item 153) 153. The method of any one of items 148 to 152, wherein the normalizing step comprises GC-LOESS normalization, followed by normalization according to fractional counts of the median, followed by normalization by principal components. (Item 154) 154. The method of any one of items 148 to 153, comprising determining the presence or absence of a genetic variation for the test sample according to the normalized counts. (Item 155) 155. The method of any one of items 148 to 154, comprising determining a chromosome structure according to the normalized counts. (Item 156) 156. The method of any one of items 148 to 155, wherein the normalized counts represent chromosome doses for the test sample. (Item 157) 157. The method of claim 156, wherein determining the presence or absence of genetic variation is according to the chromosomal dosage. (Item 158) 158. The method of any one of items 154 to 157, wherein determining the presence or absence of a genetic variation in the test sample comprises identifying the presence or absence of one copy of a chromosome, two copies of a chromosome, three copies of a chromosome, four copies of a chromosome, five copies of a chromosome, a deletion of one or more segments of a chromosome, or an insertion of one or more segments of a chromosome. (Item 159) 148. The method of claim 147, which does not include a step of normalizing the counts of the mapped nucleotide sequence reads. (Item 160) 160. The method of claim 159, comprising determining the presence or absence of a genetic variation for the test sample according to the raw counts of the mapped nucleotide sequence reads. (Item 161) 161. The method of claim 159 or 160, comprising determining the chromosome structure according to the raw counts of the mapped nucleotide sequence reads. (Item 162) 162. The method of claim 160 or 161, wherein the raw counts represent chromosome doses for the test sample. (Item 163) 163. The method of claim 162, wherein determining the presence or absence of a genetic variation is according to the chromosomal dosage. (Item 164) 164. The method of any one of items 160 to 163, wherein determining the presence or absence of a genetic variation in the test sample comprises identifying the presence or absence of one copy of a chromosome, two copies of a chromosome, three copies of a chromosome, four copies of a chromosome, five copies of a chromosome, a deletion of one or more segments of a chromosome, or an insertion of one or more segments of a chromosome. (Item 165) 1. A method for partitioning one or more genomic regions of a reference genome into multiple portions, comprising: a) determining the variability of sequencing coverage across a reference genome; b) selecting the length of the initial portion; c) partitioning at least two genomic regions according to the length of the initial portions in (b); d) comparing the variability of the sequencing coverage determined in (a) for each of the at least two genomic regions, thereby generating a comparison; e) recalculating the number of segments for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized segment length; f) repartitioning at least one of the genomic regions into a plurality of portions according to the lengths of the optimized portions in (e), thereby generating a repartitioned genomic region; g) estimating the fetal fraction for test samples derived from pregnant females bearing fetuses; h) determining a region-specific fetal fraction for each genomic region according to the correlation between the count number of nucleotide sequence reads per portion and a weighting coefficient; i) determining the size of a local minimal genomic region; j) adjusting the number of portions to include at least two portions for each genomic region, thereby generating refined, re-partitioned genomic regions; A method comprising: (Item 166) 1. A method for identifying the presence or absence of a genetic variation, comprising quantifying nucleotide sequence reads for a test sample, wherein the sequence reads comprise: a) determining the variability of sequencing coverage across the reference genome; b) selecting the length of the initial portion; c) partitioning at least two genomic regions according to the length of the initial portions in (b); d) comparing the variability of the sequencing coverage determined in (a) for each of the at least two genomic regions, thereby generating a comparison; e) recalculating the number of segments for at least one of the genomic regions according to the comparison in (d), thereby determining an optimized segment length; f) repartitioning at least one of the genomic regions into a plurality of portions according to the lengths of the optimized portions in (e), thereby generating a repartitioned genomic region; g) estimating the fetal fraction for test samples derived from pregnant females bearing fetuses; h) determining a region-specific fetal fraction for each genomic region according to the correlation between the count number of nucleotide sequence reads per region and a weighting coefficient; i) determining the size of the local smallest genomic region; j) adjusting the number of segments to include at least two segments for each genomic region, thereby generating refined, re-partitioned genomic regions; The method comprises mapping the genomic region of a reference genome to one or more genomic regions of the reference genome, the genomic region being partitioned by a process comprising: (Item 167) 167. The method of claim 165 or 166, wherein the step of determining the variability of sequencing coverage in (a) comprises using a training set of nucleotide sequence reads mapped to portions of a reference genome, wherein the sequence reads are circulating cell-free nucleic acid reads from a plurality of samples from a pregnant female carrying a fetus. (Item 168) Item 168. The method of item 167, wherein the length of the initial portion in (b) is selected for the training set according to sequencing depth. (Item 169) 169. The method of claim 167 or 168, wherein the length of the initial portion in (b) is selected for the training set according to the average fetal fraction. (Item 170) 169. The method of claim 168, wherein the average fetal fraction is determined using the training set. (Item 171) 171. The method of any one of items 165 to 170, wherein the length of the initial portion is between about 1 kb and about 1000 kb. (Item 172) 172. The method of any one of items 165 to 171, wherein the length of the initial portion is about 30 kb. (Item 173) 172. The method of any one of items 165 to 171, wherein the length of the initial portion is about 40 kb. (Item 174) 172. The method of any one of items 165 to 171, wherein the length of the initial portion is about 50 kb. (Item 175) 172. The method of any one of items 165 to 171, wherein the length of the initial portion is not 50 kb. (Item 176) 172. The method of any one of items 165 to 171, wherein the length of the initial portion is about 60 kb. (Item 177) 172. The method of any one of items 165 to 171, wherein the length of the initial portion is about 70 kb. (Item 178) 178. The method of any one of items 165 to 177, wherein the total number of portions for the genome is determined according to the length of the initial portions in (b). (Item 179) 179. The method of any one of items 165 to 178, wherein the at least two genomic regions comprise a first genomic region and a second genomic region. (Item 180) 180. The method of claim 179, wherein the first genomic region and the second genomic region are substantially similar in size. (Item 181) The step of comparing the variability of the sequencing coverage in (d) comprises calculating a proportionality coefficient (P) according to the following formula: P=(var1 / var2) 1 / 3 Formula A where var1 is the variability of the sequencing coverage of the first genomic region and var2 is the variability of the sequencing coverage of the second genomic region. 181. The method according to item 179 or 180, comprising calculating according to (Item 182) 182. The method of claim 181, wherein the sequencing coverage variability of the first genomic region is determined from a count of nucleotide sequence reads for the first genomic region, or a derivative thereof, and the sequencing coverage variability of the second genomic region is determined from a count of nucleotide sequence reads for the second genomic region, or a derivative thereof. (Item 183) 182. The method of claim 181, wherein the sequencing coverage variability of the first genomic region is determined from an average nucleotide sequence read count for the first genomic region, or a derivative thereof, and the sequencing coverage variability of the second genomic region is determined from an average nucleotide sequence read count for the second genomic region, or a derivative thereof. (Item 184) 184. The method of claim 183, wherein the average nucleotide sequencing read counts for each genomic region are determined using the training set. (Item 185) 183. The method of claim 182, wherein the nucleotide sequence read counts are normalized nucleotide sequence read counts. (Item 186) 185. The method of claim 183 or 184, wherein the average nucleotide sequence read count is an average normalized nucleotide sequence read count. (Item 187) 187. The method of any one of items 181 to 186, wherein the step of recalculating the number of portions for at least one of the genomic regions in (e) is performed according to the proportionality coefficient and the total number of portions determined from the lengths of the initial portions in (b). (Item 188) 188. The method of any one of items 165 to 187, wherein the plurality of portions in (f) comprises portions of a uniform size. (Item 189) 188. The method of any one of items 165 to 187, wherein the plurality of portions in (f) comprises portions of varying sizes. (Item 190) 189. The method according to item 188 or 189, wherein the plurality of portions in (f) comprises portions between about 1 kb and about 1000 kb in length. (Item 191) 189. The method of claim 188, wherein the plurality of portions in (f) comprises a portion of about 30 kb. (Item 192) 189. The method of claim 188, wherein the plurality of portions in (f) comprises a portion of about 40 kb. (Item 193) 189. The method of claim 188, wherein the plurality of portions in (f) comprises a portion of about 50 kb. (Item 194) 189. The method of claim 188, wherein the plurality of portions in (f) does not include the 50 kb portion. (Item 195) 189. The method of claim 188, wherein the plurality of portions in (f) comprises a portion of about 60 kb. (Item 196) 189. The method of claim 188, wherein the plurality of portions in (f) comprises a portion of about 70 kb. (Item 197) 197. The method of any one of items 165 to 196, wherein the step of estimating the fetal fraction in (g) comprises determining an error value. (Item 198) 198. The method of any one of items 165 to 197, wherein the step of determining the size of the smallest local genomic region in (i) comprises determining the size of the smallest local genomic region detectable for the sample having the estimated fetal fraction in (g). (Item 199) 199. The method of item 198, wherein the size of the local minimal genomic region is determined according to the upper 95% confidence interval for the fetal fraction. (Item 200) (k) re-estimating the fetal fraction from the refined, re-partitioned genomic regions. (Item 201) 201. The method of claim 200, comprising comparing the estimated fetal fraction in (g) with the re-estimated fetal fraction in (k). (Item 202) 202. The method of claim 201, comprising repeating parts (g), (h), (i), and (j) if the estimated fetal fraction in (g) differs from the re-estimated fetal fraction in (k) by a predetermined tolerance value. (Item 203) Item 203. The method according to item 202, wherein the predetermined tolerance value is between about 1% and about 25%. (Item 204) 204. The method of any one of items 165 to 203, comprising sequencing nucleic acid from the test sample by a nucleotide sequencing process to generate nucleotide sequence reads. (Item 205) 205. The method of claim 204, wherein the nucleic acid is circulating cell-free nucleic acid from a pregnant female carrying a fetus. (Item 206) 206. The method of any one of items 165 to 205, comprising mapping nucleotide sequence reads from the test sample to portions of the refined, repartitioned reference genome, thereby generating mapped nucleotide sequence reads. (Item 207) 207. The method of claim 206, comprising normalizing the counts of the mapped nucleotide sequence reads, thereby generating normalized counts. (Item 208) 208. The method of claim 207, wherein the normalizing step comprises LOESS normalization for guanine and cytosine (GC) bias (GC-LOESS normalization). (Item 209) 209. The method of claim 207 or 208, wherein the normalizing step comprises adjusting the counts of sequence reads according to a median count. (Item 210) 209. The method of claim 208, wherein the counts of the sequence reads are adjusted according to the median fractional counts. (Item 211) 211. The method of any one of items 207 to 210, wherein the normalizing step comprises principal component normalization. (Item 212) 212. The method of any one of items 207 to 211, wherein the normalizing step comprises GC-LOESS normalization, followed by normalization according to fractional counts of the median, followed by principal component normalization. (Item 213) 213. The method of any one of items 207 to 212, comprising determining the presence or absence of a genetic variation for the test sample according to the normalized counts. (Item 214) 214. The method of any one of items 207 to 213, comprising determining chromosome structure according to the normalized counts. (Item 215) 215. The method of any one of paragraphs 207 to 214, wherein the normalized counts represent chromosome dosages for the test sample. (Item 216) 216. The method of claim 215, wherein determining the presence or absence of genetic variation is according to the chromosomal dosage. (Item 217) 217. The method of any one of paragraphs 213 to 216, wherein determining the presence or absence of a genetic variation in the test sample comprises identifying the presence or absence of one copy of a chromosome, two copies of a chromosome, three copies of a chromosome, four copies of a chromosome, five copies of a chromosome, a deletion of one or more segments of a chromosome, or an insertion of one or more segments of a chromosome. (Item 218) 207. The method of claim 206, which does not include a step of normalizing the counts of the mapped nucleotide sequence reads. (Item 219) 219. The method of claim 218, comprising determining the presence or absence of a genetic variation for the test sample according to the raw counts of the mapped nucleotide sequence reads. (Item 220) 210. The method of claim 218 or 219, comprising determining the chromosome structure according to the raw counts of the mapped nucleotide sequence reads. (Item 221) 221. The method of claim 219 or 220, wherein the raw counts represent chromosome doses for the test sample. (Item 222) 222. The method of claim 221, wherein determining the presence or absence of genetic variation is according to the chromosomal dosage. (Item 223) 223. The method of any one of paragraphs 219 to 222, wherein determining the presence or absence of a genetic variation in the test sample comprises identifying the presence or absence of one copy of a chromosome, two copies of a chromosome, three copies of a chromosome, four copies of a chromosome, five copies of a chromosome, a deletion of one or more segments of a chromosome, or an insertion of one or more segments of a chromosome. (Item 224) 1. A method for partitioning one or more genomic regions of a reference genome into multiple portions, comprising: a) determining the variability of sequencing coverage across a reference genome; b) selecting the length of the initial portion; c) partitioning at least two genomic regions according to the length of the initial portions in (b); d) determining a region-specific fetal fraction for each genomic region according to the correlation between the count number of nucleotide sequence reads per portion and a weighting coefficient; e) determining the size of a local minimal genomic region; f) adjusting the number of portions to include at least two portions for each genomic region, thereby generating repartitioned genomic regions; A method comprising: (Item 225) 1. A method for identifying the presence or absence of a genetic variation, comprising quantifying nucleotide sequence reads for a test sample, wherein the sequence reads comprise: a) determining the variability of sequencing coverage across the reference genome; b) selecting the length of the initial portion; c) partitioning at least two genomic regions according to the length of the initial portions in (b); d) determining the region-specific fetal fraction for each genomic region according to the correlation between the count number of nucleotide sequence reads per region and the weighting coefficient; e) determining the size of the local smallest genomic region; and f) adjusting the number of segments to include at least two segments for each genomic region, thereby generating re-partitioned genomic regions; The method comprises mapping the genomic region of a reference genome to one or more genomic regions of the reference genome, the genomic region being partitioned by a process comprising: (Item 226) 226. The method of claim 224 or 225, wherein the step of determining the variability of sequencing coverage in (a) comprises using a training set of nucleotide sequence reads mapped to portions of a reference genome, wherein the sequence reads are circulating cell-free nucleic acid reads from a plurality of samples from a pregnant female carrying a fetus. (Item 227) 227. The method according to item 224 or 226, wherein the length of the initial portion in (b) is selected according to the depth of sequencing. (Item 228) 228. The method of claim 224, 226, or 227, wherein the length of the initial portion in (b) is selected according to the average fetal fraction. (Item 229) 229. The method of claim 228, wherein the average fetal fraction is determined using the training set. (Item 230) 229. The method of any one of items 224 to 229, wherein the length of the initial portion is between about 1 kb and about 1000 kb. (Item 231) 231. The method of any one of items 224 to 230, wherein the length of the initial portion is about 30 kb. (Item 232) 231. The method of any one of items 224 to 230, wherein the length of the initial portion is about 40 kb. (Item 233) 231. The method of any one of items 224 to 230, wherein the length of the initial portion is about 50 kb. (Item 234) 231. The method of any one of items 224 to 230, wherein the length of the initial portion is not 50 kb. (Item 235) 231. The method of any one of items 224 to 230, wherein the length of the initial portion is about 60 kb. (Item 236) 231. The method of any one of items 224 to 230, wherein the length of the initial portion is about 70 kb. (Item 237) 237. The method of any one of items 224 to 236, wherein the total number of portions for the genome is determined according to the length of the initial portions in (b). (Item 238) 238. The method of any one of items 224 to 237, wherein the at least two genomic regions comprise a first genomic region and a second genomic region. (Item 239) 239. The method of claim 238, wherein the first genomic region and the second genomic region are substantially similar in size. (Item 240) 239. The method of any one of items 224 to 239, wherein the repartitioned genomic region in (f) comprises a portion of a certain size. (Item 241) 239. The method of any one of items 224 to 239, wherein the repartitioned genomic region in (f) comprises portions of varying size. (Item 242) 242. The method of item 240 or 241, wherein the repartitioned genomic region in (f) comprises a portion having a size between about 1 kb and about 1000 kb. (Item 243) 242. The method of claim 240 or 241, wherein the repartitioned genomic region in (f) comprises a portion having a size of about 30 kb. (Item 244) 242. The method of claim 240 or 241, wherein the repartitioned genomic region in (f) comprises a portion having a size of about 40 kb. (Item 245) 242. The method of claim 240 or 241, wherein the repartitioned genomic region in (f) comprises a portion having a size of about 50 kb. (Item 246) 242. The method of item 240 or 241, wherein the repartitioned genomic region does not comprise a portion of 50 kb. (Item 247) 242. The method of claim 240 or 241, wherein the repartitioned genomic region in (f) comprises a portion having a size of about 60 kb. (Item 248) 242. The method of claim 240 or 241, wherein the repartitioned genomic region in (f) comprises a portion having a size of about 70 kb. (Item 249) 249. The method of any one of items 224 to 248, wherein the step of determining the size of the local smallest genomic region in (e) comprises identifying the size of the local genomic region detectable for a sample having an average fetal fraction. (Item 250) (g) re-estimating the fetal fraction from the re-partitioned genomic region. (Item 251) 251. The method of claim 250, comprising comparing the region-specific fetal fraction in (d) with the re-estimated fetal fraction in (g). (Item 252) 252. The method of claim 251, comprising repeating parts (d), (e), and (f) if the region-specific fetal fraction in (d) differs from the re-estimated fetal fraction in (g) by a predetermined tolerance value. (Item 253) Item 253. The method of item 252, wherein the predetermined tolerance value is between about 1% and about 25%. (Item 254) 254. The method of any one of items 224 to 253, comprising sequencing nucleic acid from the test sample by a nucleotide sequencing process to generate nucleotide sequence reads. (Item 255) 255. The method of claim 254, wherein the nucleic acid is circulating cell-free nucleic acid from a pregnant female carrying a fetus. (Item 256) 256. The method of any one of items 224 to 255, comprising mapping nucleotide sequence reads from the test sample to portions of the repartitioned reference genome, thereby generating mapped nucleotide sequence reads. (Item 257) 257. The method of claim 256, comprising normalizing the counts of the mapped nucleotide sequence reads, thereby generating normalized counts. (Item 258) 258. The method of claim 257, wherein the normalizing step comprises LOESS normalization for guanine and cytosine (GC) bias (GC-LOESS normalization). (Item 259) 259. The method of claim 257 or 258, wherein the normalizing step comprises adjusting the counts of sequence reads according to a median count. (Item 260) 259. The method of claim 259, wherein the sequence read counts are adjusted according to the median fraction counts. (Item 261) 261. The method of any one of items 257 to 260, wherein the normalizing step comprises principal component normalization. (Item 262) 262. The method of any one of items 257 to 261, wherein the normalizing step comprises GC-LOESS normalization, followed by normalization according to median fraction counts, followed by principal component normalization. (Item 263) 263. The method of any one of items 257 to 262, comprising determining the presence or absence of a genetic variation for the test sample according to the normalized counts. (Item 264) 264. The method of any one of items 257 to 263, comprising determining chromosome structure according to the normalized counts. (Item 265) 265. The method of any one of paragraphs 257 to 264, wherein the normalized counts represent chromosome dosages for the test sample. (Item 266) 266. The method of claim 265, wherein determining the presence or absence of genetic variation is according to the chromosomal dosage. (Item 267) 267. The method of any one of paragraphs 263 to 266, wherein determining the presence or absence of a genetic variation in the test sample comprises identifying the presence or absence of one copy of a chromosome, two copies of a chromosome, three copies of a chromosome, four copies of a chromosome, five copies of a chromosome, a deletion of one or more segments of a chromosome, or an insertion of one or more segments of a chromosome. (Item 268) 257. The method of claim 256, which does not include a step of normalizing the counts of the mapped nucleotide sequence reads. (Item 269) 269. The method of claim 268, comprising determining the presence or absence of a genetic variation for the test sample according to the raw counts of the mapped nucleotide sequence reads. (Item 270) 269. The method of claim 268, comprising determining the chromosome structure according to the raw counts of the mapped nucleotide sequence reads. (Item 271) 271. The method of claim 269 or 270, wherein the raw counts represent chromosome doses for the test sample. (Item 272) 272. The method of claim 271, wherein determining the presence or absence of genetic variation is according to the chromosomal dosage. (Item 273) 273. The method of any one of items 269 to 272, wherein determining the presence or absence of a genetic variation in the test sample comprises identifying the presence or absence of one copy of a chromosome, two copies of a chromosome, three copies of a chromosome, four copies of a chromosome, five copies of a chromosome, a deletion of one or more segments of a chromosome, or an insertion of one or more segments of a chromosome. (Item 274) 1. A method for partitioning a reference genome, or a part thereof, into a plurality of portions, comprising: a) generating a guanine and cytosine (GC) profile for a reference genome, or a part thereof; b) applying a segmentation process to the GC profile generated in (a), thereby providing individual segments; c) partitioning the reference genome, or part thereof, into a plurality of parts according to the individual segments provided in (b), thereby generating a GC-partitioned reference genome, or part thereof; A method comprising: (Item 275) 1. A method for identifying the presence or absence of a genetic variation, comprising quantifying nucleotide sequence reads for a test sample, wherein the sequence reads comprise: a) generating a guanine and cytosine (GC) profile for a reference genome, or part thereof; b) applying a segmentation process to the GC profile generated in (a), thereby providing individual segments; c) partitioning said reference genome, or part thereof, into a plurality of parts according to said individual segments provided in (b), thereby generating a GC-partitioned reference genome, or part thereof; A method for mapping to a reference genome, or part thereof, partitioned by a process including: (Item 276) 276. The method of claim 274 or 275, comprising partitioning a chromosome, or a segment of a chromosome, from the reference genome, thereby generating a GC-partitioned chromosome or a GC-partitioned chromosome segment. (Item 277) 277. The method of claim 274, 275, or 276, wherein the GC profile in (a) comprises a GC content level determined for each 1 kb of nucleotide sequence in the reference genome. (Item 278) Item 278. The method of item 277, wherein the segmentation process in (b) is performed on the GC content level. (Item 279) 279. The method of claim 278, wherein 1 kb of nucleotide sequences having similar GC content levels are combined into the individual segments. (Item 280) 279. The method of any one of items 274 to 279, wherein the segmentation process in (b) generates a decomposed rendering including the individual segments. (Item 281) The segmentation process in (b) determines the length of the chromosome (L chr ) and the length of the smallest part (L min 281. The method according to any one of items 274 to 280, wherein the method is carried out according to the level of decomposition based on (Item 282) 282. The method of any one of items 274 to 281, wherein the segmentation process in (b) comprises Haar wavelet segmentation. (Item 283) 283. The method of any one of items 274 to 282, wherein the plurality of portions comprises portions of varying sizes. (Item 284) Item 284. The method of Item 283, wherein the plurality of portions comprises portions having a size between about 30 kb and about 300 kb. (Item 285) 284. The method of claim 283, wherein the plurality of portions comprises a portion of about 32 kb. (Item 286) 284. The method of claim 283, wherein the plurality of portions comprises a portion of about 64 kb. (Item 287) 284. The method of claim 283, wherein the plurality of portions comprises a portion of about 128 kb. (Item 288) 284. The method of claim 283, wherein the plurality of portions comprises a portion of about 256 kb. (Item 289) 284. The method of claim 283, wherein the plurality of portions does not include a 50 kb portion. (Item 290) 289. The method of any one of items 274 to 289, comprising determining the GC content for said individual segments in (b). (Item 291) 291. The method of any one of items 274 to 290, comprising sequencing nucleic acid from the test sample by a nucleotide sequencing process to generate nucleotide sequence reads. (Item 292) 292. The method of claim 291, wherein the nucleic acid is circulating cell-free nucleic acid from a pregnant female carrying a fetus. (Item 293) 293. The method of any one of items 274 to 292, comprising mapping nucleotide sequence reads from the test sample to portions of the GC-partitioned reference genome, thereby generating mapped nucleotide sequence reads. (Item 294) 294. The method of claim 293, comprising normalizing the counts of the mapped nucleotide sequence reads, thereby generating normalized counts. (Item 295) 295. The method of claim 294, wherein the normalizing step comprises LOESS normalization for guanine and cytosine (GC) bias (GC-LOESS normalization). (Item 296) 295. The method of claim 293 or 294, wherein the normalizing step comprises adjusting the counts of sequence reads according to a median count. (Item 297) 297. The method of claim 296, wherein the sequence read counts are adjusted according to the median fractional counts. (Item 298) 298. The method of any one of items 293 to 297, wherein the normalizing step comprises principal component normalization. (Item 299) 299. The method of any one of items 293 to 298, wherein the normalizing step comprises GC-LOESS normalization, followed by normalization according to fractional counts of the median, followed by principal component normalization. (Item 300) 300. The method of any one of items 293 to 299, comprising determining the presence or absence of a genetic variation for the test sample according to the normalized counts. (Item 301) 301. The method of any one of items 294 to 300, comprising determining chromosome structure according to said normalized counts. (Item 302) 302. The method of any one of paragraphs 294 to 301, wherein the normalized counts represent chromosome doses for the test sample. (Item 303) 303. The method of claim 302, wherein determining the presence or absence of genetic variation is according to the chromosomal dosage. (Item 304) 304. The method of any one of paragraphs 300 to 303, wherein determining the presence or absence of a genetic variation in the test sample comprises identifying the presence or absence of one copy of a chromosome, two copies of a chromosome, three copies of a chromosome, four copies of a chromosome, five copies of a chromosome, a deletion of one or more segments of a chromosome, or an insertion of one or more segments of a chromosome. (Item 305) 294. The method of claim 293, which does not include a step of normalizing the counts of the mapped nucleotide sequence reads. (Item 306) 306. The method of claim 305, comprising determining the presence or absence of a genetic variation for the test sample according to the raw counts of the mapped nucleotide sequence reads. (Item 307) 307. The method of claim 305 or 306, comprising determining chromosome structure according to raw counts of the mapped nucleotide sequence reads. (Item 308) 308. The method of claim 306 or 307, wherein the raw counts represent chromosome doses for the test sample. (Item 309) 309. The method of claim 308, wherein determining the presence or absence of genetic variation is according to the chromosomal dosage. (Item 310) 309. The method of any one of paragraphs 306 to 309, wherein determining the presence or absence of a genetic variation in the test sample comprises identifying the presence or absence of one copy of a chromosome, two copies of a chromosome, three copies of a chromosome, four copies of a chromosome, five copies of a chromosome, a deletion of one or more segments of a chromosome, or an insertion of one or more segments of a chromosome. [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 1 shows a diagram illustrating genomic regions partitioned into segments with similar GC content using wavelet binning methods.

[0015] [Figure 2] FIG. 2 shows an example of a bin size distribution using the wavelet binning method.

[0016] [Figure 3] Figure 3 shows classification results for chromosomes 21, 18, and 13 using wavelet binning and 50 kb binning methods in the LDTv4CE2 study. Accuracy is identical for the two methods.

[0017] [Figure 4] FIG. 4 shows the truth table for chromosomes 21, 18, and 13 for the LDTv4CE2 study.

[0018] [Figure 5] FIG. 5 is a diagram illustrating an example of a workflow using the wavelet binning method.

[0019] [Figure 6] Figure 6 shows the distribution of z-scores for euploid events (left peak) and trisomy events (right peak). The shaded band indicates the false negative (FN) rate α.

[0020] [Figure 7] FIG. 7 shows the minimum fetal fraction for the detection of microdeletions and / or microduplications that achieves a population-level false negative (FN) rate of 1%.

[0021] [Figure 8] FIG. 8 presents a list of certain examples of microdeletions and microduplications.

[0022] [Figure 9] Figure 9 shows a LOESS regression plot of normalized counts versus enet (Elastic Net) bin coefficients (excluding bins with coefficients of 0) for three samples. Sample 1 had a small fetal fraction (approximately 5%), Sample 2 had a medium fetal fraction (approximately 10%), and Sample 3 had a large fetal fraction (approximately 20%).

[0023] [Figure 10] FIG. 10 illustrates certain steps that may be used in the optimal discretization method.

[0024] [Figure 11]FIG. 11 illustrates an example workflow for the optimal discretization method using certain steps illustrated in FIG.

[0025] [Figure 12] FIG. 12 illustrates an exemplary embodiment of a system in which certain embodiments of the present technology may be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0026] (Detailed explanation) Provided herein are methods for determining the presence or absence of genetic variations (e.g., chromosomal aneuploidies, microduplications, or microdeletions), where the determination is made partially and / or completely according to nucleic acid sequences. Also provided herein are methods for partitioning one or more genomic regions of a reference genome into multiple portions according to variability in sequencing coverage and / or sequence content (e.g., guanine and cytosine (GC) content). In some embodiments, nucleic acid sequences are obtained from a sample obtained from a pregnant female (e.g., the blood of a pregnant female). Also provided herein are improved data manipulation methods, as well as systems, devices, and modules, which, in some embodiments, perform the methods described herein. In some embodiments, identifying genetic variations using the methods described herein can lead to the diagnosis of a particular medical condition or determine a predisposition to a particular medical condition. Identifying genetic variances can facilitate medical decisions and / or provide access to useful medical procedures.

[0027] sample This paper provides a method and composition for analyzing nucleic acid.In some embodiments, the nucleic acid fragment in a mixture of nucleic acid fragments is analyzed.The mixture of nucleic acid can comprise two or more nucleic acid fragment species with different nucleotide sequences, different fragment lengths, different origins (for example, genomic origin, fetal origin vs. maternal origin, cell origin or tissue origin, sample origin, subject origin, etc.), or combinations thereof.

[0028] Often, nucleic acids or nucleic acid mixtures utilized in the methods and devices described herein are isolated from samples obtained from subjects. The subject can be any living or non-living organism, including, but not limited to, humans, non-human animals, plants, bacteria, fungi, or protists. Any human or non-human animal can be selected, including, but not limited to, mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovines (e.g., cows), equines (e.g., horses), caprines and ovines (e.g., sheep, goats), swine (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursidae (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. The subject can be male or female (e.g., female, pregnant woman). The subject can be of any age (eg, embryo, fetus, infant, child, adult).

[0029] Nucleic acids can be isolated from any type of suitable biological specimen or sample (e.g., a test sample). A sample or test sample can be any specimen isolated or obtained from a subject or part thereof (e.g., a human subject, a pregnant female, a fetus). Non-limiting examples of specimens include bodily fluids or tissues obtained from a subject, including, but not limited to, blood or blood products (e.g., serum, plasma, etc.), umbilical cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, spinal fluid, lavage fluid (e.g., from the bronchoalveolar, stomach, peritoneal cavity, duct, ear, arthroscopy), biopsy samples (e.g., samples obtained from preimplantation embryo biopsies), peritoneal aspirate samples, cells (blood cells, placental cells, embryonic or fetal cells, fetal nucleated cells or fetal cell remnants) or parts thereof (e.g., mitochondria, nuclei, extracts, etc.), female reproductive tract washings, urine, feces, sputum, saliva, nasal mucus, prostatic fluid, lavage fluid, semen, lymph, bile, tears, sweat, milk, mammary fluid, etc., or combinations thereof. In some embodiments, the biological sample is a cervical swab obtained from a subject. In some embodiments, the biological sample may be blood, and sometimes may be plasma or serum. The term "blood," as used herein, refers to a sample or preparation of blood obtained from a pregnant woman or a woman being tested for possible pregnancy. This term encompasses whole blood, blood products, or any fraction of blood, such as serum, plasma, buffy coat, etc., according to conventional definitions. Blood or fractions thereof often contain nucleosomes (e.g., maternal and / or fetal nucleosomes). Nucleosomes contain nucleic acids and are sometimes acellular or intracellular nucleosomes. Blood also includes buffy coats, which are sometimes isolated using a Ficoll gradient. Buffy coats can contain white blood cells (e.g., leukocytes, T cells, B cells, platelets, etc.). In certain embodiments, buffy coats contain maternal and / or fetal nucleic acids. Plasma refers to the fraction of whole blood obtained by centrifugation of blood treated with an anticoagulant. Serum refers to the aqueous liquid portion remaining after a blood sample has clotted. Body fluid or tissue samples are often collected according to standard protocols commonly followed by hospitals or outpatient clinics.In the case of blood, an appropriate volume of peripheral blood (e.g., 3-40 milliliters) is often collected and can be stored according to standard procedures before or after preparation. The body fluid or tissue sample from which nucleic acid is extracted may be free of cells (e.g., acellular). In some embodiments, the body fluid or tissue sample may contain cellular elements or cellular remnants. In some embodiments, fetal cells or cancerous cells may be included in the sample.

[0030] Often, sample is heterogeneous, which means that more than one type of nucleic acid species exists in sample.For example, heterogeneous nucleic acid can include, but is not limited to, (i) fetal nucleic acid and maternal nucleic acid, (ii) cancerous nucleic acid and non-cancerous nucleic acid, (iii) pathogenic nucleic acid and host nucleic acid, and more generally, (iv) mutated nucleic acid and wild-type nucleic acid.Sample can be heterogeneous because there are more than one cell type, for example, fetal cell and maternal cell, cancerous cell and non-cancerous cell, or pathogenic cell and host cell.In some embodiments, there are minor nucleic acid species and major nucleic acid species.

[0031] When the technology described herein is applied prenatally, body fluid or tissue samples can be collected from females at the appropriate gestational age for testing, or from females that are being tested for potential pregnancy.The appropriate gestational age can vary depending on the prenatal test being performed.In certain embodiments, the pregnant female subject is sometimes in the first trimester, sometimes in the second trimester, or sometimes in the third trimester. In certain embodiments, bodily fluids or tissues are collected from pregnant females at about 1 to about 45 weeks of gestation (e.g., 1 to 4, 4 to 8, 8 to 12, 12 to 16, 16 to 20, 20 to 24, 24 to 28, 28 to 32, 32 to 36, 36 to 40, or 40 to 44 weeks of gestation), and sometimes at about 5 to about 28 weeks of gestation (e.g., 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 weeks of gestation). In certain embodiments, bodily fluid or tissue samples are collected from pregnant females during or shortly after (e.g., 0 to 72 hours after) delivery (e.g., vaginal or non-vaginal delivery (e.g., surgical delivery)).

[0032] Obtaining blood samples and extracting DNA The methods herein often involve the isolation, enrichment, and analysis of fetal DNA found in maternal blood as a non-invasive means to detect the presence or absence of maternal and / or fetal genetic variations and / or to monitor the health of the fetus and / or pregnant female during, and sometimes after, pregnancy. Thus, the first step in carrying out certain methods herein often involves obtaining a blood sample from a pregnant woman and extracting DNA from the sample.

[0033] Obtaining blood samples Using the methods of the present technology, blood samples can be obtained from pregnant women at an appropriate gestational age for testing. The appropriate gestational age can vary depending on the disorder being tested, as discussed below. Collection of blood from women is often performed according to standard protocols commonly followed by hospitals or outpatient clinics. An appropriate amount of peripheral blood, typically 5-50 ml, is often collected and can be stored according to standard procedures before further processing. Blood samples can be collected, stored, or transported in a manner that minimizes degradation of the quality of the nucleic acids present in the sample.

[0034] Blood sample preparation Analysis of fetal DNA found in maternal blood can be performed using, for example, whole blood, serum, or plasma. Methods for preparing serum or plasma from maternal blood are known. For example, a pregnant woman's blood can be placed in a tube containing EDTA or a specialized commercial product, such as Vacutainer SST (Becton Dickinson, Franklin Lakes, NJ), to prevent blood clotting, and plasma can then be obtained from the whole blood by centrifugation. Serum can be obtained with or without centrifugation after blood clotting. When centrifugation is used, it is typically, but not necessarily, performed at an appropriate speed, e.g., 1,500 to 3,000 times g. The plasma or serum may be subjected to an additional centrifugation step before being transferred to a new tube for DNA extraction.

[0035] In addition to the cell-free portion of whole blood, DNA can also be recovered from the cellular fraction and concentrated in the buffy coat portion, which can be obtained by centrifuging a whole blood sample obtained from a woman and removing the plasma.

[0036] DNA extraction There are many known methods for extracting DNA from biological samples, including blood.General methods for DNA preparation (for example, as described by Sambrook and Russell, Molecular Cloning: A Laboratory Manual, 3rd Edition, 2001) can be followed, and various commercially available reagents or kits, such as Qiagen's QIAamp Circulating Nucleic Acid Kit, QiaAmp DNA Mini Kit, or QiaAmp DNA Blood Mini Kit (Qiagen, Hilden, Germany), GenomicPrep™ Blood DNA Isolation Kit (Promega, Madison, Wis.), and GFX™ Genomic Blood DNA Purification Kit (Amersham, Piscataway, NJ) can be used to obtain DNA from blood samples obtained from pregnant women.Also, a combination of more than one of these methods can be used.

[0037] In some embodiments, sample can be enriched or enriched to some extent for fetal nucleic acid by one or more methods first.For example, the composition and process of the present technology can be used alone or in combination with other distinguishing factors to distinguish between fetal DNA and maternal DNA.Examples of these factors include but are not limited to the single nucleotide difference between X chromosome and Y chromosome, the sequence specific to Y chromosome, polymorphisms located in other places in genome, the size difference between fetal DNA and maternal DNA, and the difference in methylation pattern between maternal tissue and fetal tissue.

[0038] Other methods for enriching a sample for particular species of nucleic acid are described in PCT Patent Application No. PCT / US07 / 69991, filed May 30, 2007, PCT Patent Application No. PCT / US2007 / 071232, filed June 15, 2007, and U.S. Provisional Applications Nos. 60 / 968,876 and 60 / 968,878 (assigned to the present applicant) (PCT Patent Application No. PCT / EP05 / 012707, filed November 28, 2005), all of which are incorporated herein by reference. In certain embodiments, parental nucleic acids are selectively removed (partially, substantially, almost completely, or completely) from the sample.

[0039] The terms "nucleic acid" and "nucleic acid molecule" can be used interchangeably throughout this disclosure. These terms refer to nucleic acids of any composition derived from DNA (e.g., complementary DNA (cDNA), genomic DNA (gDNA), etc.), RNA (e.g., messenger RNA (mRNA), small interfering RNA (siRNA), ribosomal RNA (rRNA), tRNA, microRNA, RNA highly expressed by the fetus or placenta, etc.), and / or DNA or RNA analogs (e.g., those containing base analogs, sugar analogs, and / or exogenously added backbones, etc.), RNA / DNA hybrids, polyamide nucleic acids (PNAs), etc., all of which can be in single-stranded or double-stranded form and, unless otherwise limited, can include known analogs of natural nucleotides that can function in a manner similar to naturally occurring nucleotides. In certain embodiments, the nucleic acid may be or may be derived from a plasmid, phage, autonomously replicating sequence (ARS), centromere, artificial chromosome, chromosome, or other nucleic acid capable of replicating or being replicated in vitro or in a host cell, cell, cell nucleus, or cell cytoplasm. In some embodiments, the template nucleic acid may be derived from a single chromosome (e.g., a nucleic acid sample may be derived from one chromosome of a sample obtained from a diploid organism). Unless otherwise specified, the term encompasses nucleic acids containing known analogs of natural nucleotides that have similar binding properties to the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise specified, a particular nucleic acid sequence implicitly encompasses not only the sequence explicitly indicated, but also its conservatively modified variants (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs), and complementary sequences. Specifically, degenerate codon substitutions can be obtained by generating sequences in which the third position of one or more selected (or all) codons is substituted with a mixed-base residue and / or a deoxyinosine residue. The term nucleic acid is used interchangeably with locus, gene, cDNA, and mRNA encoded by a gene.The term can also include, as equivalents, RNA or DNA derivatives, variants, and analogs synthesized from nucleotide analogs, single-stranded ("sense" or "antisense" strand, "plus" or "minus" strand, "forward" or "reverse" reading frame), and double-stranded polynucleotides. The term "gene" means a segment of DNA involved in producing a polypeptide chain, including regions preceding and following the coding region (leader and trailer) that are involved in the transcription / translation of the gene product and the regulation of transcription / translation, as well as intervening sequences (introns) between individual coding segments (exons).

[0040] Deoxyribonucleotides include deoxyadenosine, deoxycytidine, deoxyguanosine, and deoxythymidine. In the case of RNA, the base cytosine is replaced with uracil. Template nucleic acids can be prepared using nucleic acids obtained from a subject as templates.

[0041] Nucleic Acid Isolation and Processing Nucleic acids can be obtained from one or more sources (e.g., cells, serum, plasma, buffy coat, lymph, skin, soil, etc.) by methods known in the art. Any suitable method can be used to isolate, extract, and / or purify DNA from a biological sample (e.g., blood or blood products), including, but not limited to, methods for preparing DNA (e.g., as described by Sambrook and Russell, Molecular Cloning: A Laboratory Manual, 3rd Edition, 2001), various commercially available reagents or kits, such as Qiagen's QIAamp Circulating Nucleic Acid Kit, QiaAmp DNA Mini Kit, or QiaAmp DNA Blood Mini Kit (Qiagen, Hilden, Germany), GenomicPrep™ Blood DNA Isolation Kit (Promega, Madison, Wis.), and GFX™ Genomic Blood DNA Purification Kit (Amersham, Piscataway, NJ), or a combination thereof.

[0042] Cell lysis procedures and reagents are known in the art and can generally be performed by chemical methods (e.g., detergents, hypotonic solutions, enzymatic procedures, etc., or a combination thereof), physical methods (e.g., French press, sonication, etc.), or electrolyte lysis methods. Any suitable lysis procedure can be used. For example, chemical methods generally utilize a lysing agent to disrupt cells and extract nucleic acids from the cells, followed by treatment with a chaotropic salt. Physical methods, such as freeze / thaw followed by trituration; use of a cell press, etc., are also useful. High salt lysis procedures are also commonly used. For example, alkaline lysis procedures can be used. The latter procedures traditionally incorporate the use of phenol-chloroform solutions, although alternative phenol-chloroform-free procedures involving three solutions are also available. For the latter procedure, one solution may contain 15 mM Tris, pH 8.0; 10 mM EDTA, and 100 μg / ml ribonuclease A; a second solution may contain 0.2 N NaOH and 1% SDS; and a third solution may contain 3 M KOAc, pH 5.5. These procedures may be found in Current Protocols in Molecular Biology, John Wiley & Sons, NY, vols. 6.3.1-6.3.6 (1989), which are incorporated herein in their entirety.

[0043] When comparing a nucleic acid with another nucleic acid, it can be isolated at different times, and each of the samples can be from the same or different sources. For example, the nucleic acid can be derived from a nucleic acid library, such as a cDNA library or an RNA library. The nucleic acid can be the result of nucleic acid purification or isolation and / or amplification of nucleic acid molecules obtained from a sample. The nucleic acids provided for the processing described herein can contain nucleic acids from one sample or from two or more samples (e.g., one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, sixteen or more, seventeen or more, eighteen or more, nineteen or more, or twenty or more samples).

[0044] In certain embodiments, nucleic acids can include extracellular nucleic acids. As used herein, the term "extracellular nucleic acid" can refer to nucleic acids isolated from a source that is substantially cell-free, and is also referred to as "cell-free" nucleic acid and / or "cell-free circulating" nucleic acid. Extracellular nucleic acids are present in and can be obtained from blood (e.g., the blood of a pregnant female). Extracellular nucleic acids often do not contain detectable cells and may contain cellular elements or cellular remnants. Non-limiting examples of cell-free sources for obtaining extracellular nucleic acids are blood, plasma, serum, and urine. As used herein, the term "obtaining cell-free circulating sample nucleic acids" includes obtaining a sample directly (e.g., collecting a sample, e.g., a test sample) or obtaining a sample from another person who has collected the sample. Without being limited by theory, extracellular nucleic acids can be products of cellular apoptosis and cell degradation, which often result in extracellular nucleic acids with a range of lengths spanning a spectrum (e.g., a "ladder").

[0045] In certain embodiments, extracellular nucleic acids can contain different nucleic acid species and are therefore referred to herein as "heterogeneous." For example, serum or plasma obtained from a person with cancer may contain nucleic acids derived from cancerous cells and nucleic acids derived from non-cancerous cells. In another example, serum or plasma obtained from a pregnant female may contain maternal nucleic acids and fetal nucleic acids. In some cases, fetal nucleic acids sometimes represent about 5% to about 50% of the total nucleic acids (e.g., about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, or 49% of all nucleic acids are fetal nucleic acids). In some embodiments, the majority of the fetal nucleic acids in the nucleic acid are about 500 base pairs or less, about 250 base pairs or less, about 200 base pairs or less, about 150 base pairs or less, about 100 base pairs or less, about 50 base pairs or less, or about 25 base pairs or less in length. In some embodiments, the majority of the fetal nucleic acids in the nucleic acid are about 500 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acids are about 500 base pairs or less in length). In some embodiments, the majority of the fetal nucleic acids in the nucleic acid are about 250 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acids are about 250 base pairs or less in length). In some embodiments, the majority of the fetal nucleic acids in the nucleic acid are about 200 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acids are about 200 base pairs or less in length). In some embodiments, the majority of the fetal nucleic acids in the nucleic acid are about 150 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acids are about 150 base pairs or less in length).In some embodiments, the majority of the fetal nucleic acids in the nucleic acid are about 100 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acids are about 100 base pairs or less in length), In some embodiments, the majority of the fetal nucleic acids in the nucleic acid are about 50 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acids are about 50 base pairs or less in length). In some embodiments, the majority of the fetal nucleic acids in the nucleic acid are about 25 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acids are about 25 base pairs or less in length).

[0046] In some embodiments, nucleic acid fragments of a certain length, length range, or length below or above a certain threshold or cutoff are analyzed. In some embodiments, fragments with lengths below a certain threshold or cutoff (e.g., 500bp, 400bp, 300bp, 200bp, 150bp, 100bp) are referred to as "short" fragments, and fragments with lengths above a certain threshold or cutoff (e.g., 500bp, 400bp, 300bp, 200bp, 150bp, 100bp) are referred to as "long" fragments. For example, fragments with lengths below 200bp are referred to as "short" fragments, and fragments with lengths equal to or greater than 200bp are referred to as "long" fragments. In some embodiments, fragments of a certain length, length range, or length below or above a certain threshold or cutoff are analyzed, but fragments of different lengths, length ranges, or lengths below or above a threshold or cutoff are not analyzed. In some embodiments, fragments of a particular length, range of lengths, or lengths below a particular threshold or cutoff are analyzed separately from fragments of a different length, range of lengths, or lengths above a particular threshold or cutoff.

[0047] In some embodiments, fragments less than about 500 bp are analyzed. In some embodiments, fragments less than about 400 bp are analyzed. In some embodiments, fragments less than about 300 bp are analyzed. In some embodiments, fragments less than about 200 bp are analyzed. In some embodiments, fragments less than about 150 bp are analyzed. For example, fragments less than about 200 bp, 190 bp, 180 bp, 170 bp, 160 bp, 150 bp, 140 bp, 130 bp, 120 bp, 110 bp, or 100 bp are analyzed. In some embodiments, fragments between about 100 bp and about 200 bp are analyzed. For example, fragments of about 190 bp, 180 bp, 170 bp, 160 bp, 150 bp, 140 bp, 130 bp, 120 bp, or 110 bp are analyzed. In some embodiments, fragments in the range of about 100 bp to about 200 bp are analyzed. For example, fragments in the range of about 110 bp to about 190 bp, 130 bp to about 180 bp, 140 bp to about 170 bp, 140 bp to about 150 bp, 150 bp to about 160 bp, 145 bp to about 155 bp, or 130 bp to 140 bp are analyzed. In some embodiments, fragments of about 135 bp are analyzed. In some embodiments, fragments of about 200 bp or more are analyzed. In some embodiments, fragments of about 200 bp are analyzed. In some embodiments, fragments of a particular length or length range are analyzed that are about 10 bp to about 30 bp shorter than other fragments. In some embodiments, fragments of a particular length or length range are analyzed that are about 10 bp to about 20 bp shorter than other fragments. In some embodiments, fragments of a particular length or length range are analyzed that are about 10 bp to about 15 bp shorter than other fragments.

[0048] In certain embodiments, a sample containing nucleic acids can be provided without processing before performing the methods described herein. In some embodiments, a sample containing nucleic acids is processed before providing the nucleic acids and performing the methods described herein. For example, nucleic acids can be extracted, isolated, purified, partially purified, or amplified from a sample. The term "isolated," as used herein, refers to removing a nucleic acid from its original environment (e.g., the natural environment if naturally occurring, or a host cell if exogenously expressed); thus, the nucleic acid is altered in that it has been removed from its original environment by human intervention (e.g., "by the hand of man"). The term "isolated nucleic acid," as used herein, can refer to a nucleic acid removed from a subject (e.g., a human subject). An isolated nucleic acid can be provided with less non-nucleic acid components (e.g., proteins, lipids) than the amount of those components present in the source sample. A composition containing isolated nucleic acids can be about 50% to greater than 99% free of non-nucleic acid components. A composition comprising an isolated nucleic acid may be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of non-nucleic acid components. The term "purified," as used herein, can refer to providing a nucleic acid that contains fewer non-nucleic acid components (e.g., proteins, lipids, carbohydrates) than were present before the nucleic acid was subjected to a purification procedure. A composition comprising a purified nucleic acid may be about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of other non-nucleic acid components. The term "purified," as used herein, can refer to providing a nucleic acid that contains fewer nucleic acid species than in the sample source from which the nucleic acid was derived. A composition containing purified nucleic acid can be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of other nucleic acid species. For example, fetal nucleic acid can be purified from a mixture containing maternal and fetal nucleic acids.In certain instances, nucleosomes containing small fragments of fetal nucleic acid can be purified from a mixture of larger nucleosome complexes containing larger fragments of maternal nucleic acid.

[0049] In some embodiments, nucleic acids are fragmented or cleaved before, during, or after the methods described herein. The fragmented or cleaved nucleic acids can have a nominal, average, or mean length of about 5 to about 10,000 base pairs, about 100 to about 1,000 base pairs, about 100 to about 500 base pairs, or about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, or 9000 base pairs. Fragments can be generated by any suitable method known in the art, and the average, mean or nominal length of the nucleic acid fragments can be controlled by selecting an appropriate fragment generation procedure.

[0050] Nucleic acid fragments can contain overlapping nucleotide sequences, and such overlapping sequences can facilitate the construction of corresponding unfragmented nucleic acid nucleotide sequences, or segments thereof. For example, one fragment can have subsequences x and y, and another fragment can have subsequences y and z, where x, y, and z are nucleotide sequences that can be 5 nucleotides or more in length. In certain embodiments, overlapping sequence y can be utilized to facilitate the construction of the nucleotide sequence xyz in nucleic acids derived from a sample. In certain embodiments, nucleic acids can be partially fragmented (e.g., from an incomplete or aborted specific cleavage reaction) or completely fragmented.

[0051] In some embodiments, nucleic acids are fragmented or cleaved by suitable methods, non-limiting examples of which include physical methods (e.g., shearing, e.g., sonication, French press, heating, UV irradiation, etc.), enzymatic treatment (e.g., enzymatic cleavage agents (e.g., suitable nucleases, suitable restriction enzymes, suitable methylation-sensitive restriction enzymes)), chemical methods (e.g., alkylation, DMS, piperidine, acid hydrolysis, base hydrolysis, heating, etc., or combinations thereof), treatments described in U.S. Patent Application Publication No. 20050112590, etc., or combinations thereof.

[0052] As used herein, "fragmentation" or "cleavage" refers to a procedure or condition that can separate a nucleic acid molecule, e.g., a nucleic acid template gene molecule or its amplification product, into two or more smaller nucleic acid molecules. Such fragmentation or cleavage can be sequence-specific, base-specific, or non-specific, and can be achieved by any of a variety of methods, reagents, or conditions, including, for example, chemical, enzymatic, or physical fragmentation.

[0053] As used herein, "fragments," "cleavage products," "cleaved products," or grammatical variations thereof, refer to nucleic acid molecules resulting from fragmentation or cleavage of a nucleic acid template gene molecule, or amplification products thereof. While such fragments or cleaved products may refer to all nucleic acid molecules resulting from a cleavage reaction, typically such fragments or cleaved products refer only to nucleic acid molecules resulting from fragmentation or cleavage of a nucleic acid template gene molecule, or amplification product segments thereof, that contain the corresponding nucleotide sequence of the nucleic acid template gene molecule. The term "amplification," as used herein, refers to subjecting a target nucleic acid in a sample to a process that linearly or exponentially produces amplicon nucleic acids having the same or substantially the same nucleotide sequence as the target nucleic acid or a segment thereof. In certain embodiments, the term "amplification" refers to methods including polymerase chain reaction (PCR). For example, an amplification product can contain one or more more nucleotides than the amplified nucleotide region of the nucleic acid template sequence (e.g., a primer can contain "extra" nucleotides, such as a transcription initiation sequence, in addition to nucleotides complementary to the nucleic acid template gene molecule, resulting in an amplification product containing "extra" nucleotides, or nucleotides that do not correspond to the amplified nucleotide region of the nucleic acid template gene molecule). Thus, fragments can include fragments resulting from segments or parts of the amplified nucleic acid molecule that contain, at least in part, nucleotide sequence information obtained from or based on the indicated nucleic acid template molecule.

[0054] As used herein, the term "complementary cleavage reaction" refers to cleavage reactions carried out on the same nucleic acid using different cleavage reagents or by changing the cleavage specificity of the same cleavage reagent, thus generating alternative cleavage patterns of the same target or reference nucleic acid or protein.In certain embodiments, nucleic acid can be treated with one or more specific cleavage agents (for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more specific cleavage agents) in one or more reaction vessels (for example, nucleic acid is treated with each specific cleavage agent in a separate vessel).As used herein, the term "specific cleavage agent" refers to an agent, sometimes a chemical or enzyme, that can cleave nucleic acid at one or more specific sites.

[0055] Furthermore, before providing nucleic acid to the method described herein, nucleic acid can be exposed to a treatment that modifies specific nucleotides in nucleic acid.For example, nucleic acid can be subjected to a treatment that selectively modifies nucleic acid based on the methylation status of nucleotides therein.In addition, conditions such as high temperature, ultraviolet radiation, X-rays, etc. can cause changes in the sequence of nucleic acid molecules.Nucleic acid can be provided in any suitable form that is useful for performing appropriate sequence analysis.

[0056] The nucleic acid may be single-stranded or double-stranded. For example, single-stranded DNA can be generated by denaturing double-stranded DNA, for example, by heating or treating with alkali. In certain embodiments, the nucleic acid adopts a D-loop structure formed by intercalating an oligonucleotide into a strand of a double-stranded DNA molecule, or is a DNA-like molecule, for example, a peptide nucleic acid (PNA). The formation of a D-loop can be promoted by adding E. coli RecA protein and / or by changing the salt concentration, for example, using methods known in the art.

[0057] Determination of fetal nucleic acid content In some embodiments, the amount (e.g., concentration, relative amount, absolute amount, copy number, etc.) of fetal nucleic acid in the nucleic acid is determined. In certain embodiments, the amount of fetal nucleic acid in a sample is referred to as the "fetal fraction." In some embodiments, "fetal fraction" refers to the fraction of fetal nucleic acid in the circulating cell-free nucleic acid in a sample (e.g., a blood sample, a serum sample, a plasma sample) obtained from a pregnant female. In certain embodiments, the amount of fetal nucleic acid is determined according to markers specific to male fetuses (e.g., Y chromosome STR markers (e.g., DYS19, DYS385, DYS392 markers); RhD markers in RhD-negative females), the allelic ratio of polymorphic sequences, or according to one or more markers specific to fetal nucleic acid but not maternal nucleic acid (e.g., differences in epigenetic biomarkers between the mother and fetus (e.g., methylation; described in more detail below), or fetal RNA markers in maternal plasma (see, e.g., Lo, 2005, Journal of Histochemistry and Cytochemistry, 53(3):293-296)).

[0058] Determination of fetal nucleic acid content (e.g., fetal fraction) is sometimes performed using a fetal quantifier assay (FQA), for example, as described in U.S. Patent Application Publication No. 2010 / 0105049, which is incorporated herein by reference. This type of assay allows fetal nucleic acid in a maternal sample to be detected and quantified based on the methylation status of the nucleic acid in the sample. In certain embodiments, the amount of fetal nucleic acid from the maternal sample can be determined relative to the total amount of nucleic acid present, thereby providing the percentage of fetal nucleic acid in the sample. In certain embodiments, the copy number of fetal nucleic acid in the maternal sample can be determined. In certain embodiments, the amount of fetal nucleic acid can be determined in a sequence-specific (or segment-specific) manner, sometimes with sufficient sensitivity to allow accurate chromosome dosage analysis (e.g., to detect the presence or absence of fetal aneuploidy, microduplication, or microdeletion).

[0059] A fetal quantification assay (FQA) can be performed in conjunction with any of the methods described herein. Such an assay can be performed by any method known in the art and / or described in U.S. Patent Application Publication No. 2010 / 0105049, such as a method that can distinguish between maternal and fetal DNA based on differences in methylation status and quantify (i.e., determine the amount of) fetal DNA. Methods for differentiating nucleic acids based on methylation status include, but are not limited to, methylation-sensitive capture, e.g., using the MBD2-Fc fragment (in which the methyl-binding domain of MBD2 is fused to the Fc fragment of an antibody (MBD-FC)) (Gebhard et al. (2006) Cancer Res. 66(12):6118-28); methylation-specific antibodies; bisulfite conversion methods, e.g., MSP (methylation-sensitive PCR), COBRA, methylation-sensitive single nucleotide primer extension (Ms-SNuPE), or Sequenom MassCLEAVE™ technology; and the use of methylation-sensitive restriction enzymes (e.g., digesting maternal DNA in a maternal sample with one or more methylation-sensitive restriction enzymes, thereby enriching for fetal DNA). Methyl-sensitive enzymes can also be used to differentiate nucleic acids based on methylation status; these enzymes can preferentially or substantially cleave or digest at their DNA recognition sequences, e.g., when the latter is unmethylated. Thus, unmethylated DNA samples are sheared into smaller fragments than methylated DNA samples, and hypermethylated DNA samples are not sheared. Unless otherwise specified, any method for differentiating nucleic acids based on methylation status can be used with the compositions and methods of the technology herein. The amount of fetal DNA can be determined, for example, by introducing one or more competitors at known concentrations during the amplification reaction. The amount of fetal DNA can also be determined, for example, by RT-PCR, primer extension, sequencing, and / or counting. In certain cases, the amount of nucleic acid can be determined using BEAMing technology as described in U.S. Patent Application Publication No. 2007 / 0065823.In certain embodiments, the restriction efficiency can be determined and the ratio of efficiencies used to further determine the amount of DNA in the fetus.

[0060] In certain embodiments, a fetal quantification assay (FQA) can be used to determine the concentration of fetal DNA in a maternal sample, for example, by the following methods: a) determining the total amount of DNA present in the maternal sample; b) selectively digesting the maternal DNA in the maternal sample using one or more methylation-sensitive restriction enzymes, thereby enriching for fetal DNA; c) determining the amount of fetal DNA obtained from step b); d) comparing the amount of fetal DNA obtained from step c) with the total amount of DNA obtained from step a), thereby determining the concentration of fetal DNA in the maternal sample. In certain embodiments, the absolute copy number of fetal nucleic acids in the maternal sample can be determined, for example, using a system that uses mass spectrometry and / or a competitive PCR approach to measure absolute copy number. See, for example, Ding and Cantor (2003) Proc. Natl. Acad. Sci. USA 100:3059-3064, and US Patent Application Publication No. 2004 / 0081993, both of which are incorporated herein by reference.

[0061] In certain embodiments, the fetal fraction can be determined based on the ratio of alleles of a polymorphic sequence (e.g., a single nucleotide polymorphism (SNP)), for example, using a method such as that described in U.S. Patent Application Publication No. 2011 / 0224087, which is incorporated herein by reference. In such methods, nucleotide sequence reads are obtained for a maternal sample, and the fetal fraction is determined by comparing the total number of nucleotide sequence reads mapped to a first allele and the total number of nucleotide sequence reads mapped to a second allele at a reference polymorphic site (e.g., an SNP) in a reference genome. In certain embodiments, for example, the fetal allele is distinguished by the maternal nucleic acid's large contribution to the mixture of fetal and maternal nucleic acids in the sample, compared with the relatively small contribution of the fetal allele. Therefore, the relative abundance of fetal nucleic acid in a maternal sample can be determined as a parameter of the total number of unique sequence reads mapped to the target nucleic acid sequence in the reference genome for each of the two alleles at the polymorphic site.

[0062] In some embodiments, the fetal fraction can be determined using methods that incorporate information derived from maternal chromosomal abnormalities, such as those described in International Patent Application Publication No. WO 2014 / 055774, which is incorporated herein by reference. In some embodiments, the fetal fraction can be determined using methods that incorporate information derived from sex chromosomes.

[0063] In some embodiments, the fetal fraction can be determined using methods that incorporate fragment length information (e.g., fragment length ratio (FLR) analysis, fetal ratio statistic (FRS) analysis, as described in International Application Publication No. WO 2013 / 177086, incorporated herein by reference). Fragments of cell-free fetal nucleic acid are generally shorter than fragments of maternal nucleic acid (see, e.g., Chan et al. (2004) Clin. Chem. 50:88-92; Lo et al. (2010) Sci. Transl. Med. 2:61ra91). Thus, in some embodiments, the fetal fraction can be determined by counting fragments below a certain length threshold and comparing these counts to, for example, the counts obtained from fragments above a certain length threshold and / or the amount of total nucleic acid in the sample. Methods for counting nucleic acid fragments of a particular length are described in further detail in International Application Publication No. WO2013 / 177086.

[0064] In some embodiments, fetal fractions can be determined according to the estimated value of the portion-specific fetal fraction. The portion-specific fetal fraction may also be referred to as binned fetal fraction (BFF), elastic net bin coefficient, and sequence-based fetal fraction (SeqFF). Without being limited by theory, the amount of reads obtained from fetal CCF fragments (e.g., fragments of a specific length or length range) is often mapped using a frequency range for the portion (e.g., within the same sample, e.g., within the same sequencing run). Also, without being limited by theory, when compared between multiple samples, a specific portion shows a similar representation of reads obtained from fetal CCF fragments (e.g., fragments of a specific length or length range), and this representation tends to correlate with the portion-specific fetal fraction (e.g., the relative amount, percentage, or ratio of CCF fragments originating from the fetus).

[0065] In some embodiments, estimates of the part-specific fetal fraction are determined, in part, based on part-specific parameters and their relationship to the fetal fraction. The part-specific parameters can be any suitable parameter that reflects (e.g., correlates with) the amount or proportion of reads obtained from CCF fragment lengths of a particular size (e.g., size range) in the part. The part-specific parameters can be the average, mean, or median of part-specific parameters determined for multiple samples. Any suitable part-specific parameter can be used. Non-limiting examples of part-specific parameters include FLR (e.g., FRS), the amount of reads with a length less than a selected fragment length, genome coverage (i.e., coverage), mappability, counts (e.g., counts of sequence reads mapped to the part, e.g., normalized counts, ChAI-normalized counts), DNase I sensitivity, methylation status, acetylation, histone distribution, guanine-cytosine (GC) content, chromatin structure, etc., or combinations thereof. The part-specific parameter can be any suitable parameter that correlates with FLR and / or FRS in a part-specific manner. In some embodiments, some or all of the part-specific parameters are direct or indirect indications of the FLR for the part. In some embodiments, the part-specific parameter is not guanine-cytosine (GC) content.

[0066] In some embodiments, the portion-specific parameter is any suitable value that indicates, correlates with, or is proportional to the amount of reads obtained from a CCF fragment, where the reads mapped to the portion have a length less than the selected fragment length. In certain embodiments, the portion-specific parameter is an indication of the amount of reads obtained from a relatively short CCF fragment (e.g., about 200 base pairs or less) mapped to the portion. CCF fragments having a length less than the selected fragment length are often relatively short CCF fragments, and sometimes the selected fragment length is about 200 base pairs or less (e.g., CCF fragments that are about 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, or 80 bases in length). The length of the CCF fragment or the reads obtained from the CCF fragment can be determined (e.g., estimated or inferred) by any suitable method (e.g., sequencing, hybridization approach). In some embodiments, the length of the CCF fragment is determined (e.g., estimated or inferred) from reads obtained from paired-end sequencing. In certain embodiments, the length of the CCF fragment template is determined directly from the length of reads obtained from the CCF fragment (e.g., reads from a single end).

[0067] The part-specific parameters can be weighted or adjusted by one or more weighting factors. In some embodiments, the weighted or adjusted part-specific parameters can provide an estimate of the part-specific fetal fraction for a sample (e.g., a test sample). In some embodiments, the weighting or adjustment generally converts part counts (e.g., reads mapped to parts) or another part-specific parameter into an estimate of the part-specific fetal fraction, and such conversion is sometimes considered a conversion.

[0068] In some embodiments, the weighting coefficients are coefficients or constants that describe and / or define, in part, the relationship between the fetal fractions (e.g., fetal fractions determined from a plurality of samples) and the part-specific parameters for a plurality of samples (e.g., a training set). In some embodiments, the weighting coefficients are determined according to the relationships between the fetal fraction determinations and the part-specific parameters. A relationship can be defined by one or more weighting coefficients, and one or more weighting coefficients can be determined from a relationship. In some embodiments, the weighting coefficients (e.g., one or more weighting coefficients) are determined from a relationship for the parts that is fitted according to (i) the fraction of fetal nucleic acid determined for each of the plurality of samples and (ii) the part-specific parameters for the plurality of samples.

[0069] The weighting coefficients can be any suitable coefficients, estimated coefficients, or constants obtained from a suitable relationship (e.g., a suitable mathematical relationship, algebraic relationship, fitted relationship, regression, regression analysis, regression model). The weighting coefficients can be determined according to, derived from, or estimated from a suitable relationship. In some embodiments, the weighting coefficients are estimated coefficients from the fitted relationship. Fitting a relationship for a plurality of samples is sometimes referred to as training a model. Any suitable model and / or method for fitting a relationship (e.g., training a model to obtain a training set) can be used. Non-limiting examples of suitable models that can be used include regression models, linear regression models, simple regression models, ordinary least squares regression models, multiple regression models, general multiple regression models, polynomial regression models, general linear models, generalized linear models, discrete choice regression models, logistic regression models, multinomial logit models, mixed logit models, probit models, multinomial probit models, ordered logit models, ordered probit models, Poisson models, multivariate response regression models, multilevel models, fixed effects models, random effects models, mixed models, nonlinear regression models, nonparametric models, semiparametric models, robust models, quantile models, isotonic models, principal component models, least angle models, local models, segmented models, and errors in variables models. In some embodiments, the fitted relationship is not a regression model. In some embodiments, the fitted relationship is selected from a decision tree model, a support vector machine model, and a neural network model. The result of training a model (e.g., a regression model, a relationship) is often a relationship that can be mathematically described, and the relationship includes one or more coefficients (e.g., weighting coefficients). More complex multivariate models can determine one, two, three, or more weighting coefficients. In some embodiments, a model is trained according to fetal fractions obtained from multiple samples and two or more part-specific parameters (e.g., coefficients) (e.g., fitting relationships to multiple samples, e.g., by a matrix).

[0070] The weighting coefficients may be obtained from any suitable relationship (e.g., any suitable mathematical relationship, algebraic relationship, fitted relationship, regression, regression analysis, regression model) by any suitable method. In some embodiments, the fitted relationship is fitted by estimation, non-limiting examples of which include least squares, ordinary least squares, linear regression, partial regression, full regression, generalized regression, weighted regression, nonlinear regression, iteratively weighted regression, ridge regression, least absolute deviation, Bayes, Bayesian multivariate, reduced rank, LASSO, Weighted Rank Selection Criteria (WRSC), Rank Selection Criteria (RSC), elastic net estimation methods (e.g., elastic net regression), and combinations thereof.

[0071] Weighting coefficients can be determined for or associated with any suitable portion of a genome. Weighting coefficients can be determined for or associated with any suitable portion of any suitable chromosome. In some embodiments, weighting coefficients are determined for or associated with some or all portions in a genome. In some embodiments, weighting coefficients are determined for or associated with portions of some or all chromosomes in a genome. Weighting coefficients are sometimes determined for or associated with selected chromosome portions. Weighting coefficients can be determined for or associated with portions of one or more autosomes. Weighting coefficients can be determined for or associated with portions of a plurality of portions, including portions of autosomes or a subset thereof. In some embodiments, weighting coefficients are determined for or associated with portions of sex chromosomes (e.g., ChrX and / or ChrY). Weighting coefficients can be determined for or associated with portions of one or more autosomes and one or more sex chromosomes. In certain embodiments, weighting coefficients are determined for or associated with the portions of the plurality of portions in all autosomes and chromosomes X and Y. Weighting coefficients can be determined for or associated with the portions of the plurality of portions that do not include the portions in chromosomes X and / or Y. In certain embodiments, weighting coefficients are determined for or associated with the portions of a chromosome, and this chromosome includes aneuploidy (e.g., whole chromosome aneuploidy). In certain embodiments, weighting coefficients are determined for or associated only with the portions of a chromosome, and this chromosome is not aneuploid (e.g., is a euploid chromosome). Weighting coefficients can be determined for or associated with the portions of the plurality of portions that do not include the portions in chromosomes 13, 18, and / or 21.

[0072] In some embodiments, weighting coefficients are determined for portions according to one or more samples (e.g., samples from a training set). The weighting coefficients are often portion-specific. In some embodiments, one or more weighting coefficients are independently assigned to portions. In some embodiments, weighting coefficients are determined according to a relationship between fetal fraction determinations for multiple samples (e.g., sample-specific fetal fraction determinations) and portion-specific parameters determined according to multiple samples. Weighting coefficients are often determined from multiple samples, for example, from about 20 to about 100,000 or more, from about 100 to about 100,000 or more, from about 500 to about 100,000 or more, from about 1,000 to about 100,000 or more, or from about 10,000 to about 100,000 or more samples. The weighting factor can be determined from a sample that is euploid (e.g., a sample obtained from a subject with a euploid fetus, e.g., a sample in which no aneuploid chromosomes are present). In some embodiments, the weighting factor is obtained from a sample that contains aneuploid chromosomes (e.g., a sample obtained from a subject with a euploid fetus). In some embodiments, the weighting factor is determined from multiple samples obtained from a subject with a euploid fetus and a subject with a trisomic fetus. The weighting factor can be obtained from multiple samples, and these samples are obtained from subjects with male and / or female fetuses.

[0073] The fetal fraction is often determined for one or more samples of the training set, from which the weighting coefficients are derived. The fetal fraction for determining the weighting coefficients is sometimes the result of determining a sample-specific fetal fraction. The fetal fraction for determining the weighting coefficients can be determined by any suitable method described herein or known in the art. In some embodiments, the fetal nucleic acid content (e.g., fetal fraction) is determined using a suitable fetal quantification assay (FQA) described herein or known in the art, including, but not limited to, determination according to markers specific to male fetuses, determination based on the allelic ratio of polymorphic sequences, determination according to one or more markers specific to fetal nucleic acids but not maternal nucleic acids, determination using methylation-based DNA discrimination (e.g., A. Nygren et al. (2010) Clinical Chemistry 56(10):1627-1635), determination by mass spectrometry methods and / or systems using a competitive PCR approach, determination by the methods described in U.S. Patent Application Publication No. 2010 / 0105049, which is incorporated herein by reference, or a combination thereof. Often, the fetal fraction is determined, in part, at the Y chromosome level (e.g., at the level of one or more genomic segments; at the level of a profile). In some embodiments, the fetal fraction is determined according to a suitable assay of the Y chromosome (e.g., using quantitative real-time PCR to compare the amount of a fetal-specific locus (e.g., the SRY locus on the Y chromosome in the case of a male fetus) with the amount of a locus on any autosome common to both the mother and the fetus (e.g., Lo YM et al. (1998) Am J Hum Genet 62:768-775)).

[0074] The part-specific parameters (e.g., for the test sample) can be weighted or adjusted by one or more weighting factors (e.g., weighting factors derived from a training set). For example, weighting factors can be derived for parts according to the relationship between the part-specific parameters and the fetal fraction determination results for a training set of multiple samples. The part-specific parameters of the test sample can then be adjusted and / or weighted according to the weighting factors derived from the training set. In some embodiments, the part-specific parameters from which the weighting factors are derived are the same as the part-specific parameters (e.g., for the test sample) that are adjusted or weighted (e.g., both parameters are FLR). In certain embodiments, the part-specific parameters from which the weighting factors are derived are different from the part-specific parameters (e.g., for the test sample) that are adjusted or weighted. For example, weighting factors can be determined from the relationship between coverage (i.e., part-specific parameters) and fetal fraction for the samples in the training set, and the FLR (i.e., another part-specific parameter) for the parts of the test sample can be adjusted according to the weighting factor derived from coverage. Without being limited by theory, the part-specific parameters (e.g., for a test sample) can sometimes be adjusted and / or weighted by weighting factors derived from different part-specific parameters (e.g., of a training set) due to the relationship and / or correlation between each part-specific parameter and the common part-specific FLR.

[0075] An estimate of the part-specific fetal fraction can be determined for a sample (e.g., a test sample) by weighting the part-specific parameters by the weighting coefficients determined for that part. Weighting can include adjusting, converting, and / or transforming the part-specific parameters by the weighting coefficients by applying any suitable mathematical operation, non-limiting examples of which include multiplication, division, addition, subtraction, integration, symbolic calculation, algebraic calculation, algorithm, trigonometric or geometric function, transformation (e.g., Fourier transform), etc., or combinations thereof. Weighting can include adjusting, converting, and / or transforming the part-specific parameters by the weighting coefficients by a suitable mathematical model.

[0076] In some embodiments, the fetal fraction is determined for a sample according to one or more portion-specific fetal fraction estimates. In some embodiments, the fetal fraction is determined (e.g., estimated) for a sample (e.g., a test sample) according to weighting or adjustment of portion-specific parameters for one or more portions. In certain embodiments, the fraction of fetal nucleic acid for a test sample is estimated based on adjusted counts or adjusted subset counts. In certain embodiments, the fraction of fetal nucleic acid for a test sample is estimated based on adjusted FLR, adjusted FRS, adjusted coverage, and / or adjusted mappability for the portions. In some embodiments, weighting or adjustment of about 1 to about 500,000, about 100 to about 300,000, about 500 to about 200,000, about 1,000 to about 200,000, about 1,500 to about 200,000, or about 1,500 to about 50,000 portion-specific parameters is performed.

[0077] The fetal fraction (e.g., for a test sample) is determined according to multiple part-specific fetal fraction estimates (e.g., for the same test sample) by any suitable method. In some embodiments, a method for improving the accuracy of estimating the fraction of fetal nucleic acid in a test sample obtained from a pregnant female includes determining one or more part-specific fetal fraction estimates, and the fetal fraction estimate for the sample is determined according to the one or more part-specific fetal fraction estimates. In some embodiments, estimating or determining the fraction of fetal nucleic acid for a sample (e.g., a test sample) includes a substep of summing one or more part-specific fetal fraction estimates. The summing substep can include determining the mean, average, median, AUC, or integral value according to the multiple part-specific fetal fraction estimates.

[0078] In some embodiments, a method for improving the accuracy of estimating the fraction of fetal nucleic acid in a test sample obtained from a pregnant female includes obtaining counts of sequence reads mapped to portions of a reference genome, where these sequence reads are reads of circulating cell-free nucleic acid obtained from a test sample from a pregnant female, and at least a subset of the obtained counts are obtained from a region of the genome, where the region provides a higher number of counts obtained from fetal nucleic acid compared to the total number of counts from this region than the total number of counts obtained from fetal nucleic acid compared to another region of the genome. In some embodiments, the estimate of the fraction of fetal nucleic acid is determined according to a subset of the portions, where the subset of portions is selected according to a portion to which the number of counts obtained from fetal nucleic acid compared to non-fetal nucleic acid is mapped to a higher number than the number of counts obtained from fetal nucleic acid compared to non-fetal nucleic acid in another portion. In some embodiments, the subset of portions is selected according to a portion to which the number of counts obtained from fetal nucleic acid compared to non-fetal nucleic acid is mapped to a higher number than the number of counts obtained from fetal nucleic acid compared to non-fetal nucleic acid in another portion. The counts mapped to all or a subset of the portions can be weighted to obtain weighted counts. The weighted counts can be used to estimate the fraction of fetal nucleic acid, and the counts can be weighted according to the portion to which the counts obtained from fetal nucleic acid are mapped that are greater than the counts of fetal nucleic acid in another portion. In some embodiments, the counts are weighted according to the portion to which the counts obtained from fetal nucleic acid relative to non-fetal nucleic acid are mapped that are greater than the counts of fetal nucleic acid relative to non-fetal nucleic acid in another portion.

[0079] The fetal fraction can be determined for a sample (e.g., a test sample) according to a plurality of portion-specific fetal fraction estimates for the sample, where the portion-specific estimates are obtained from portions of any suitable region or segment of the genome. The portion-specific fetal fraction estimates can be determined for one or more portions of suitable chromosomes (e.g., one or more selected chromosomes, one or more autosomes, sex chromosomes (e.g., ChrX and / or ChrY), aneuploid chromosomes, euploid chromosomes, etc., or combinations thereof).

[0080] In some embodiments, determining the fetal fraction includes the substeps of (a) obtaining counts of sequence reads mapped to portions of the reference genome (these sequence reads are reads of circulating cell-free nucleic acids obtained from a test sample derived from a pregnant female); (b) using, for example, a microprocessor, weighting (i) the counts of sequence reads mapped to each portion, or (ii) other portion-specific parameters, for the portion-specific fraction of fetal nucleic acid according to a weighting coefficient independently associated with each portion, thereby obtaining an estimate of the portion-specific fetal fraction according to the weighting coefficient (for a plurality of samples, each of the weighting coefficients has been determined from a relationship fitted for each portion between (i) the fraction of fetal nucleic acid for each of the plurality of samples and (ii) the counts of sequence reads mapped to each portion or other portion-specific parameter); and (c) estimating the fraction of fetal nucleic acid for the test sample based on the estimate of the portion-specific fetal fraction.

[0081] The amount of fetal nucleic acid in extracellular nucleic acid can be quantified and used in conjunction with the methods provided herein. Thus, in certain embodiments, the methods of the technology described herein include an additional step of determining the amount of fetal nucleic acid. The amount of fetal nucleic acid in a nucleic acid sample obtained from a subject can be determined before or after processing to prepare the sample nucleic acid. In certain embodiments, after processing and preparing the sample nucleic acid, the amount of fetal nucleic acid in the sample is determined, and this amount is used for further evaluation. In some embodiments, the outcome includes factoring the fraction of fetal nucleic acid in the sample nucleic acid (e.g., adjusting the count, removing the sample, making a call, or not making a call).

[0082] The determining step can be performed before, during, at any point in the methods described herein, or after a particular method (e.g., detecting aneuploidy, detecting microduplications or microdeletions, determining fetal sex) described herein. For example, to perform a fetal sex or aneuploidy, microduplication or microdeletion determination method with a given sensitivity or specificity, a method of quantifying fetal nucleic acid can be performed before, during, or after determining fetal sex or aneuploidy, microduplication or microdeletion to identify samples with more than about 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25% or more fetal nucleic acid. In some embodiments, for example, the sample that is determined to have a certain threshold amount of fetal nucleic acid (for example, about 15% or more fetal nucleic acid; about 4% or more fetal nucleic acid) is further analyzed to determine fetal sex or aneuploidy, microduplication or microdeletion, or for the presence or absence of aneuploidy or genetic variation.In certain embodiments, only when a sample has a certain threshold amount of fetal nucleic acid (for example, about 15% or more fetal nucleic acid; about 4% or more fetal nucleic acid), for example, the determination of fetal sex or the presence or absence of aneuploidy, microduplication or microdeletion is selected (for example, selected and communicated to patient).

[0083] In some embodiments, determining the presence or absence of chromosomal aneuploidy, microduplication or microdeletion does not require or require determining the fetal fraction or the amount of fetal nucleic acid.In some embodiments, determining the presence or absence of chromosomal aneuploidy, microduplication or microdeletion does not require differentiation of the sequence between fetal DNA and maternal DNA.In certain embodiments, this is because the combined contribution of both maternal and fetal sequences in a specific chromosome, chromosomal portion or segment is analyzed.In some embodiments, determining the presence or absence of chromosomal aneuploidy, microduplication or microdeletion does not depend on a priori sequence information that would distinguish between fetal DNA and maternal DNA.

[0084] Nucleic acid concentration In some embodiments, nucleic acids (e.g., extracellular nucleic acids) are enriched or relatively enriched to obtain a subpopulation or species of nucleic acids. Subpopulations of nucleic acids can include, for example, fetal nucleic acids, maternal nucleic acids, nucleic acids comprising fragments of a specific length or range of lengths, or nucleic acids derived from a specific genomic region (e.g., a single chromosome, a set of chromosomes, and / or a specific chromosomal region). Such enriched samples can be used in conjunction with the methods provided herein. Thus, in certain embodiments, the methods of the present technology include an additional step of enriching for a subpopulation of nucleic acids in a sample, such as fetal nucleic acids. In certain embodiments, the above-described methods for determining the fetal fraction can also be used to enrich for fetal nucleic acids. In certain embodiments, maternal nucleic acids are selectively (partially, substantially, almost completely, or completely) removed from the sample. In certain embodiments, enrichment for specific low-copy-number species of nucleic acids (e.g., fetal nucleic acids) can improve quantitative sensitivity. Methods for enriching a sample for a particular species of nucleic acid are described, for example, in U.S. Pat. No. 6,927,028, International Patent Application Publication No. WO2007 / 140417, International Patent Application Publication No. WO2007 / 147063, International Patent Application Publication No. WO2009 / 032779, International Patent Application Publication No. WO2009 / 032781, International Patent Application Publication No. WO2010 / 033639, International Patent Application Publication No. WO2011 / 034631, International Patent Application Publication No. WO2006 / 056480, and International Patent Application Publication No. WO2011 / 143659, all of which are incorporated herein by reference.

[0085] In some embodiments, nucleic acids are enriched to obtain specific target and / or reference fragment species. In certain embodiments, nucleic acids are enriched to obtain specific nucleic acid fragment lengths or ranges of fragment lengths using one or more length-based separation methods described below. In certain embodiments, nucleic acids are enriched to obtain fragments derived from selected genomic regions (e.g., chromosomes) using one or more sequence-based separation methods described herein and / or known in the art. Specific methods for enriching for subpopulations of nucleic acids (e.g., fetal nucleic acids) in a sample are described in detail below.

[0086] Some methods for enriching for a subpopulation of nucleic acids (e.g., fetal nucleic acids) that can be used with the methods described herein include methods that exploit epigenetic differences between maternal and fetal nucleic acids. For example, fetal nucleic acids can be differentiated from and separated from maternal nucleic acids based on differences in methylation. Methylation-based methods for enriching for fetal nucleic acids are described in U.S. Patent Application Publication No. 2010 / 0105049, incorporated herein by reference. Such methods sometimes include binding sample nucleic acids to a methylation-specific binder (e.g., methyl-CpG binding protein (MBD), methylation-specific antibody, etc.) and separating bound nucleic acids from unbound nucleic acids based on differences in methylation status. Such methods can also include the use of methylation-sensitive restriction enzymes (described above; e.g., HhaI and HpaII), which allow for the enrichment of a region of fetal nucleic acid in a maternal sample by selectively digesting nucleic acids from the maternal sample with an enzyme that selectively and completely or substantially digests maternal nucleic acids, enriching the sample for at least one region of fetal nucleic acid.

[0087] Another method for enriching for a subpopulation of nucleic acids (e.g., fetal nucleic acids) that can be used with the methods described herein is the approach of enhancing polymorphic sequences with restriction endonucleases, such as the method described in U.S. Patent Application Publication No. 2009 / 0317818, incorporated herein by reference. Such a method includes cleaving nucleic acids containing non-target alleles with a restriction endonuclease that recognizes nucleic acids containing non-target alleles but not target alleles, and amplifying the uncleaved nucleic acids without amplifying the cleaved nucleic acids, where the uncleaved, amplified nucleic acids are target nucleic acids (e.g., fetal nucleic acids) enriched relative to non-target nucleic acids (e.g., maternal nucleic acids). In certain embodiments, for example, nucleic acids can be selected to contain alleles with polymorphic sites that are susceptible to selective digestion by a cleavage agent.

[0088] Some methods for enriching for a subpopulation of nucleic acids (e.g., fetal nucleic acids) that can be used with the methods described herein involve selective enzymatic degradation approaches. Such methods include protecting target sequences from exonuclease digestion, thereby facilitating the elimination of unwanted sequences (e.g., maternal DNA) in the sample. For example, one approach involves denaturing sample nucleic acids to generate single-stranded nucleic acids, contacting the single-stranded nucleic acids with at least one target-specific primer pair under appropriate annealing conditions, extending the annealed primers by nucleotide polymerization to generate double-stranded target sequences, and digesting the single-stranded nucleic acids using a nuclease that digests single-stranded (i.e., non-target) nucleic acids. In certain embodiments, this method can be repeated for at least one additional cycle. In certain embodiments, the same target-specific primer pair is used for primer extension in each of the first and second cycles, and in certain embodiments, different target-specific primer pairs are used for the first and second cycles.

[0089] Some methods for enriching for a subpopulation of nucleic acids (e.g., fetal nucleic acids) that can be used with the methods described herein include massively parallel signature sequencing (MPSS) approaches. MPSS is typically a solid-phase method that uses adapter (i.e., tag) ligation, followed by adapter decoding and aliquoting and reading of nucleic acid sequences. Typically, tagged PCR products are amplified, resulting in a PCR product with a unique tag for each nucleic acid. Often, tags are used to link the PCR products to microbeads. After several rounds of ligation-based sequencing, for example, a sequence signature can be identified from each bead. Each signature sequence (MPSS tag) in the MPSS dataset is analyzed and compared to all other signatures, and all identical signatures are counted.

[0090] In certain embodiments, certain enrichment methods (e.g., certain MPS- and / or MPSS-based enrichment methods) can include amplification (e.g., PCR)-based approaches. In certain embodiments, locus-specific amplification methods can be used (e.g., using locus-specific amplification primers). In certain embodiments, a multiplex SNP allele PCR approach can be used. In certain embodiments, a multiplex SNP allele PCR approach can be used in combination with uniplex sequencing. For example, such an approach can include the use of multiplex PCR (e.g., a MASSARRAY system) and incorporation of capture probe sequences into amplicons, followed by sequencing using, for example, an Illumina MPSS system. In certain embodiments, a multiplex SNP allele PCR approach can be used in combination with a three-primer system and index sequencing. For example, such an approach can involve using multiplex PCR (e.g., a MASSARRAY system) with a primer having a first capture probe incorporated into a particular locus-specific forward PCR primer and an adapter sequence incorporated into a locus-specific reverse PCR primer for sequencing using, for example, an Illumina MPSS system, thereby generating an amplicon, followed by a second PCR to incorporate the reverse capture sequence and molecular index barcode. In certain embodiments, a multiplex SNP allele PCR approach can be used in combination with a four-primer system and index sequencing.For example, such an approach can include using multiplex PCR (e.g., a MASSARRAY system) with primers having adapter sequences incorporated into both the locus-specific forward PCR primer and the locus-specific reverse PCR primer for sequencing using, for example, an Illumina MPSS system, followed by a second PCR to incorporate both the forward and reverse capture sequences and the molecular index barcode. In certain embodiments, a microfluidic approach can be used. In certain embodiments, an array-based microfluidic approach can be used. For example, such an approach can include using a microfluidic array (e.g., Fluidigm) to perform low-plex amplification and incorporate index and capture probes, followed by sequencing. In certain embodiments, an emulsion microfluidic approach, such as digital droplet PCR, can be used.

[0091] In certain embodiments, universal amplification methods can be used (e.g., using universal primers or non-locus-specific amplification primers). In certain embodiments, universal amplification methods can be used in combination with pull-down approaches. In certain embodiments, methods can include biotinylated ultramer pull-down from a universally amplified sequencing library (e.g., biotinylated pull-down assays from Agilent or IDT). For example, such an approach can include standard library preparation, enrichment for selected regions by pull-down assay, and a second universal amplification step. In certain embodiments, pull-down approaches can be used in combination with ligation-based methods. In certain embodiments, methods can include biotinylated ultramer pull-down using sequence-specific adapter ligation (e.g., HALOPLEX PCR, Halo Genomics). For example, such an approach can include the use of selector probes to capture restriction enzyme digest fragments, followed by ligation of the captured products to adapters and universal amplification, followed by sequencing. In certain embodiments, pull-down approaches can be used in combination with extension and ligation-based methods. In certain embodiments, the method can include molecular inversion probe (MIP) extension and ligation. For example, such an approach can include the use of molecular inversion probes in combination with sequence adapters, followed by universal amplification and sequencing. In certain embodiments, complementary DNA can be synthesized and sequenced without amplification.

[0092] In certain embodiments, the extension and ligation approach can be performed without the pull-down component. In certain embodiments, the method can include hybridization, extension, and ligation using a forward primer and a reverse primer specific to the locus. Such methods can further include universal amplification or amplification-free complementary DNA synthesis followed by sequencing. In certain embodiments, such methods can reduce or eliminate background sequences during analysis.

[0093] In certain embodiments, the pull-down approach can be used with or without an optional amplification component. In certain embodiments, the method can include a modified pull-down assay and ligation that fully incorporates a capture probe and does not involve universal amplification. For example, such an approach can include the use of a modified selector probe to capture restriction enzyme digestion fragments, followed by ligation of the captured product to an adapter, optional amplification, and sequencing. In certain embodiments, the method can include a biotinylated pull-down assay with extension and ligation of an adapter sequence combined with circular single-stranded ligation. For example, such an approach can include the use of a selector probe against a capture region of interest (i.e., a target sequence), probe extension, adapter ligation, single-stranded circular ligation, optional amplification, and sequencing. In certain embodiments, analysis of the sequencing results can separate the target sequence from background.

[0094] In some embodiments, nucleic acids are enriched to obtain fragments derived from selected genomic regions (e.g., chromosomes) using one or more sequence-based separation methods described herein. Sequence-based separation is generally based on the presence of nucleotide sequences in fragments of interest (e.g., target and / or reference fragments) and that are substantially absent or present in only trace amounts (e.g., 5% or less) in other fragments of the sample. In some embodiments, sequence-based separation can separate target fragments and / or reference fragments. The separated target fragments and / or separated reference fragments are often isolated and removed from the remaining fragments in the nucleic acid sample. In certain embodiments, the separated target fragments and the separated reference fragments are also isolated and removed from each other (e.g., as separate assay compartments). In certain embodiments, the separated target fragments and the separated reference fragments are isolated together (e.g., as the same assay compartment). In some embodiments, unbound fragments can be differentially removed, degraded, or digested.

[0095] In some embodiments, selective nucleic acid capture process is used to separate and extract target fragment and / or reference fragment from nucleic acid sample.Commercially available nucleic acid capture systems include, for example, Nimblegen sequence capture system (Roche NimbleGen, Madison, WI); Illumina BEADARRAY platform (Illumina, San Diego, CA); Affymetrix GENECHIP platform (Affymetrix, Santa Clara, CA); Agilent SureSelect Target Enrichment System (Agilent Technologies, Santa Clara, CA); and related platforms.This method typically includes hybridization of capture oligonucleotide with the nucleotide sequence of a segment or the entirety of target fragment or reference fragment, and can include the use of solid phase (for example, solid phase array) and / or solution-based platform. Capture oligonucleotides (sometimes called "baits") are selected or designed to preferentially hybridize to nucleic acid fragments derived from a selected genomic region or locus (e.g., one of chromosomes 21, 18, 13, X, or Y, or a reference chromosome). In certain embodiments, hybridization-based methods (e.g., using oligonucleotide arrays) are used to enrich for and obtain nucleic acid sequences derived from specific chromosomes (e.g., potentially aneuploid chromosomes, reference chromosomes, or other chromosomes of interest), or segments thereof of interest.

[0096] In some embodiments, one or more length-based separation methods are used to enrich nucleic acids for a particular nucleic acid fragment length, a range of lengths, or lengths below or above a particular threshold or cutoff. Nucleic acid fragment length typically refers to the number of nucleotides in the fragment. Nucleic acid fragment length is also sometimes referred to as nucleic acid fragment size. In some embodiments, length-based separation methods are performed without measuring the length of individual fragments. In some embodiments, length-based separation methods are performed in conjunction with methods for determining the length of individual fragments. In some embodiments, length-based separation refers to a size fractionation procedure, and all or a portion of the fractionated pool can be isolated (e.g., retained) and / or analyzed. Size fractionation procedures are known in the art (e.g., separation on an array, separation by molecular sieving, separation by gel electrophoresis, separation by column chromatography (e.g., molecular sieving column), and microfluidic technology-based approaches). In certain embodiments, length-based separation approaches can include, for example, circularization of fragments, chemical treatment (e.g., formaldehyde, polyethylene glycol (PEG)), mass spectrometry, and / or size-specific nucleic acid amplification.

[0097] In some embodiments, nucleic acid fragments of a certain length, length range, or length below or above a certain threshold or cutoff are separated from a sample. In some embodiments, fragments having lengths below a certain threshold or cutoff (e.g., 500 bp, 400 bp, 300 bp, 200 bp, 150 bp, 100 bp) are referred to as "short" fragments, and fragments having lengths above a certain threshold or cutoff (e.g., 500 bp, 400 bp, 300 bp, 200 bp, 150 bp, 100 bp) are referred to as "long" fragments. In some embodiments, fragments of a certain length, length range, or length below or above a certain threshold or cutoff are retained for analysis, while fragments of a different length, length range, or length below or above a threshold or cutoff are not retained for analysis. In some embodiments, fragments less than about 500 bp are retained. In some embodiments, fragments less than about 400 bp are retained. In some embodiments, fragments less than about 300 bp are retained. In some embodiments, fragments of less than about 200 bp are retained. In some embodiments, fragments of less than about 150 bp are retained. For example, fragments of less than about 190 bp, 180 bp, 170 bp, 160 bp, 150 bp, 140 bp, 130 bp, 120 bp, 110 bp, or 100 bp are retained. In some embodiments, fragments of about 100 bp to about 200 bp are retained. For example, fragments of about 190 bp, 180 bp, 170 bp, 160 bp, 150 bp, 140 bp, 130 bp, 120 bp, or 110 bp are retained. In some embodiments, fragments in the range of about 100 bp to about 200 bp are retained. For example, fragments in the range of about 110 bp to about 190 bp, 130 bp to about 180 bp, 140 bp to about 170 bp, 140 bp to about 150 bp, 150 bp to about 160 bp, or 145 bp to about 155 bp are retained. In some embodiments, fragments that are about 10 bp to about 30 bp shorter than other fragments of a particular length or length range are retained. In some embodiments, fragments that are about 10 bp to about 20 bp shorter than other fragments of a particular length or length range are retained. In some embodiments, fragments that are about 10 bp to about 15 bp shorter than other fragments of a particular length or length range are retained.

[0098] In some embodiments, one or more bioinformatics-based (e.g., in silico) methods are used to enrich nucleic acids for specific nucleic acid fragments of a certain length, a length range, or lengths below or above a certain threshold or cutoff.For example, suitable nucleotide sequencing processes can be used to obtain nucleotide sequence reads for nucleic acid fragments.In some cases, such as when using a sequencing method that reads from both ends, the length of a specific fragment can be determined based on the position of the mapped sequence reads obtained from each end of the fragment.As described herein, the sequence reads used for a specific analysis (e.g., an analysis that determines the presence or absence of genetic variation) can be enriched or filtered according to one or more selected fragment lengths or fragment length thresholds for the corresponding fragments.

[0099] A method for specific length-based separation that can be used with the methods described herein utilizes, for example, a selective sequence tagging approach. The term "sequence tagging" refers to incorporating a recognizable and distinct sequence into a nucleic acid or a population of nucleic acids. As used herein, the term "sequence tagging" has a different meaning from the term "sequence tag" described later in this specification. In such a sequence tagging method, nucleic acids of a certain fragment size species (e.g., short fragments) are subjected to selective sequence tagging in a sample containing long and short nucleic acids. Such a method typically includes a step of performing a nucleic acid amplification reaction using a set of nested primers including an inner primer and an outer primer. In certain embodiments, one or both of the inner primers can be tagged, thereby introducing the tag onto the target amplification product. The outer primer generally does not anneal to the short fragment carrying the (inner) target sequence. The inner primer can anneal to the short fragment and generate an amplification product carrying the tag and the target sequence. Typically, tagging of long fragments is inhibited through a combination of mechanisms, including, for example, blocking extension of inner primers by prior annealing and extension of outer primers. Enrichment for tagged fragments can be achieved by any of a variety of methods, including, for example, exonuclease digestion of single-stranded nucleic acids and amplification of tagged fragments using at least one tag-specific amplification primer.

[0100] Another length-based separation method that can be used with the methods described herein involves subjecting a nucleic acid sample to polyethylene glycol (PEG) precipitation. Exemplary methods include those described in International Patent Application Publication Nos. WO2007 / 140417 and WO2010 / 115016. This method generally involves contacting a nucleic acid sample with PEG in the presence of one or more monovalent salts under conditions sufficient to substantially precipitate large nucleic acids without substantially precipitating small (e.g., less than 300 nucleotides) nucleic acids.

[0101] Another size-based enrichment method that can be used with the methods described herein includes circularization by ligation, for example, by ligation using circligase. Short nucleic acid fragments can typically be circularized more efficiently than long fragments. Non-circularized sequences can be separated from circularized sequences, and the enriched short fragments can be used for further analysis.

[0102] Nucleic Acid Library In some embodiments, a nucleic acid library is a plurality of polynucleotide molecules (e.g., a sample of nucleic acids) that are prepared, collected, and / or modified for a particular process (non-limiting examples of which include immobilization on a solid phase (e.g., a solid support, e.g., a flow cell, beads), enrichment, amplification, cloning, detection) and / or for nucleic acid sequencing. In certain embodiments, the nucleic acid library is prepared before or during the sequencing process. Nucleic acid libraries (e.g., sequencing libraries) can be prepared by suitable methods known in the art. Nucleic acid libraries can be prepared by targeted or non-targeted preparation processes.

[0103] In some embodiments, the library of nucleic acids is modified to include chemical moieties (e.g., functional groups) configured for immobilization of the nucleic acids to a solid support. In some embodiments, the library of nucleic acids is modified to include biological molecules (e.g., functional groups) and / or members of binding pairs configured for immobilization of the library to a solid support, non-limiting examples of which include thyroxine-binding globulin, steroid-binding proteins, antibodies, antigens, haptens, enzymes, lectins, nucleic acids, repressors, protein A, protein G, avidin, streptavidin, biotin, complement component C1q, nucleic acid-binding proteins, receptors, carbohydrates, oligonucleotides, polynucleotides, complementary nucleic acid sequences, etc., and combinations thereof. Some examples of specific binding pairs include, but are not limited to, an avidin moiety and a biotin moiety; an antigenic epitope and an antibody or immunologically reactive fragment thereof; an antibody and a hapten; a digoxigen moiety and an anti-digoxigen antibody; a fluorescein moiety and an anti-fluorescein antibody; an operator and a repressor; a nuclease and a nucleotide; a lectin and a polysaccharide; a steroid and a steroid binding protein; an active compound and a receptor for the active compound; a hormone and a hormone receptor; an enzyme and a substrate; an immunoglobulin and Protein A; an oligonucleotide or polynucleotide and its corresponding complement, and the like, or combinations thereof.

[0104] In some embodiments, a library of nucleic acids is modified to include one or more polynucleotides of known composition, including, but not limited to, identifiers (e.g., tags, index tags), capture sequences, labels, adapters, restriction enzyme sites, promoters, enhancers, origins of replication, stem-loops, complementary sequences (e.g., primer binding sites, annealing sites), suitable integration sites (e.g., transposons, viral integration sites), modified nucleotides, etc., or combinations thereof. Polynucleotides of known sequence can be added to any suitable position, such as the 5' end, 3' end, or internal position of the nucleic acid sequence. Polynucleotides of known sequence can be the same or different sequences. In some embodiments, polynucleotides of known sequence are configured to hybridize to one or more oligonucleotides immobilized on a surface (e.g., a surface in a flow cell). For example, a nucleic acid molecule containing a 5' known sequence can be hybridized to a first plurality of oligonucleotides, while the 3' known sequence of the molecule can be hybridized to a second plurality of oligonucleotides. In some embodiments, the nucleic acid library can include chromosome-specific tags, capture sequences, labels, and / or adapters. In some embodiments, the nucleic acid library includes one or more detectable labels. In some embodiments, one or more detectable labels can be incorporated into the nucleic acid library at the 5' end, the 3' end, and / or at any nucleotide position within the nucleic acids in the library. In some embodiments, the nucleic acid library includes hybridized oligonucleotides. In certain embodiments, the hybridized oligonucleotides are labeled probes. In some embodiments, the nucleic acid library includes hybridized oligonucleotide probes prior to immobilization on a solid phase.

[0105] In some embodiments, the polynucleotide of known sequence comprises a universal sequence. A universal sequence is a specific nucleotide sequence that is incorporated into two or more nucleic acid molecules, or two or more subsets of nucleic acid molecules, and the universal sequence is the same for all molecules in the molecules or subsets into which it is incorporated. Universal sequences are often designed to hybridize to and / or amplify multiple different sequences using a single universal primer that is complementary to the universal sequence. In some embodiments, two (e.g., pairs) or more universal sequences and / or universal primers are used. Universal primers often comprise a universal sequence. In some embodiments, an adapter (e.g., a universal adapter) comprises a universal sequence. In some embodiments, one or more universal sequences are used to capture, identify, and / or detect multiple species or subsets of nucleic acids.

[0106] In certain embodiments of nucleic acid library preparation (e.g., in the case of specific sequencing by synthesis procedures), nucleic acids are size-selected and / or fragmented to lengths of a few hundred base pairs or less (e.g., in preparation for library generation). In some embodiments, library preparation is performed without fragmentation (e.g., when using ccfDNA).

[0107] In certain embodiments, ligation-based library preparation methods are used (e.g., ILLUMINA TRUSEQ, Illumina, San Diego, CA). Ligation-based library preparation methods often utilize adapter (e.g., methylated adapter) designs, which can incorporate index sequences in the initial ligation step and can often be used to prepare samples for single-end sequencing, double-end sequencing, and multiplex sequencing. End repair of nucleic acids (e.g., fragmented nucleic acids or ccfDNA) is sometimes performed, for example, by a fill-in reaction, an exonuclease reaction, or a combination thereof. In some embodiments, the resulting blunt-end repaired nucleic acid can then be extended with a single nucleotide that is complementary to the single-nucleotide overhang on the 3' end of the adapter / primer. Any nucleotide can be used for the extension / overhang nucleotide. In some embodiments, nucleic acid library preparation includes ligation of an adapter oligonucleotide. Adapter oligonucleotides often exhibit complementarity to flow cell anchors and are sometimes utilized, for example, to immobilize nucleic acid libraries to a solid support, such as the inner surface of a flow cell. In some embodiments, the adapter oligonucleotide comprises an identifier, one or more sequencing primer hybridization sites (e.g., a sequence exhibiting complementarity to a universal sequencing primer, a single-end sequencing primer, a double-end sequencing primer, a multiplex sequencing primer, etc.), or a combination thereof (e.g., adapter / sequencing, adapter / identifier, adapter / identifier / sequencing).

[0108] An identifier is a suitable detectable label incorporated into or tethered to a nucleic acid (e.g., a polynucleotide), allowing the detection and / or identification of the nucleic acid containing it. In some embodiments, the identifier is incorporated into or tethered to a nucleic acid (e.g., by a polymerase) during a sequencing method. Non-limiting examples of identifiers include nucleic acid tags, nucleic acid indexes or barcodes, radiolabels (e.g., isotopes), metal labels, fluorescent labels, chemiluminescent labels, phosphorescent labels, fluorophore quenchers, dyes, proteins (e.g., enzymes, antibodies or parts thereof, linkers, members of binding pairs), etc., or combinations thereof. In some embodiments, the identifier (e.g., nucleic acid index or barcode) is a unique, known, and / or identifiable sequence of nucleotides or nucleotide analogs. In some embodiments, the identifier is six or more adjacent nucleotides. Numerous fluorophores are available with a variety of different excitation and emission spectra. Any suitable type and / or number of fluorophores can be used as identifiers. In some embodiments, one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, or fifty or more different identifiers are utilized in the methods described herein (e.g., nucleic acid detection and / or sequencing methods). In some embodiments, one or two types of identifiers (e.g., fluorescent labels) are linked to each nucleic acid in the library.Detection and / or quantification of the identifiers may be performed by any suitable method or device, non-limiting examples of which include flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, a luminometer, a fluorometer, a spectrophotometer, suitable gene chip or microarray analysis, Western blot, mass spectrometry, chromatography, cytofluorimetric analysis, fluorescence microscopy, suitable fluorescence or digital imaging methods, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, electric field suspension, suitable nucleic acid sequencing methods and / or nucleic acid sequencing devices, and the like, and combinations thereof.

[0109] In some embodiments, transposon-based library preparation methods are used (e.g., EPICENTRE NEXTERA, Epicentre, Madison WI). Transposon-based methods typically use in vitro transposition to simultaneously fragment and tag DNA in a single-tube reaction (often allowing the incorporation of platform-specific tags and optional barcodes) to prepare sequencing-compatible libraries.

[0110] In some embodiments, the nucleic acid library, or portions thereof, are amplified (e.g., amplified by a PCR-based method). In some embodiments, sequencing methods involve amplification of the nucleic acid library. The nucleic acid library can be amplified before or after immobilization on a solid support (e.g., a solid support in a flow cell). Nucleic acid amplification involves a process that amplifies or increases the number of nucleic acid templates and / or their complements present (e.g., in a nucleic acid library) by generating one or more copies of the templates and / or their complements. Amplification can be carried out by any suitable method. The nucleic acid library can be amplified by thermocycling or isothermal amplification. In some embodiments, rolling circle amplification is used. In some embodiments, amplification occurs on a solid support (e.g., inside a flow cell) to which the nucleic acid library, or portions thereof, are immobilized. In certain sequencing methods, the nucleic acid library is added to a flow cell and immobilized by hybridization to anchors under appropriate conditions. This type of nucleic acid amplification is often referred to as solid-phase amplification. In some embodiments of solid-phase amplification, all or part of the amplification products are synthesized by extension initiated from immobilized primers. Solid-phase amplification reactions are similar to standard solution-phase amplification, except that at least one of the amplification oligonucleotides (eg, primers) is immobilized on a solid support.

[0111] In some embodiments, solid-phase amplification includes nucleic acid amplification reactions that include only one type of oligonucleotide primer immobilized on a surface. In certain embodiments, solid-phase amplification includes multiple different immobilized oligonucleotide primer species. In some embodiments, solid-phase amplification can include nucleic acid amplification reactions that include one type of oligonucleotide primer immobilized on a solid surface and a second, different oligonucleotide primer species in solution. Multiple different species of immobilized or solution-based primers can be used. Non-limiting examples of solid-phase nucleic acid amplification reactions include interface amplification, bridge amplification, emulsion PCR, WildFire amplification (e.g., U.S. Patent Publication No. US20130012399), etc., or combinations thereof.

[0112] Sequencing In some embodiments, nucleic acids (e.g., nucleic acid fragments, sample nucleic acids, cell-free nucleic acids) are sequenced. In certain embodiments, complete or substantially complete sequences are obtained, and sometimes partial sequences are obtained.

[0113] In some embodiments, the fragment length is determined using a sequencing method. In some embodiments, the fragment length is determined using a sequencing platform that reads from both ends. Such platforms involve sequencing both ends of a nucleic acid fragment. Generally, sequences corresponding to both ends of a fragment can be mapped to a reference genome (e.g., a reference human genome). In certain embodiments, both ends are sequenced separately for each fragment end, with a read length sufficient to map to a reference genome. Examples of sequence read lengths read from both ends are described below. In certain embodiments, all or part of the sequence reads can be mapped to a reference genome without mismatches. In some embodiments, each read is mapped independently. In some embodiments, information from both sequence reads (i.e., from each end) is incorporated into the mapping process. The fragment length can be determined, for example, by calculating the difference between the genomic coordinates assigned to each of the mapped reads read from both ends.

[0114] In some embodiments, the length of a fragment can be determined using a sequencing process that yields the complete, or substantially complete, nucleotide sequence for the fragment, including platforms that generate relatively long read lengths (e.g., Roche 454, Ion Torrent, single molecule (Pacific Biosciences) platforms, real-time SMRT technology, etc.).

[0115] In some embodiments, some or all of the nucleic acids in a sample are enriched and / or amplified (e.g., non-specifically, e.g., by PCR-based methods) before or during sequencing. In certain embodiments, a specific portion or subset of nucleic acids in a sample are enriched and / or amplified before or during sequencing. In some embodiments, a portion or subset of a preselected pool of nucleic acids is sequenced randomly. In some embodiments, nucleic acids in a sample are not enriched and / or amplified before or during sequencing.

[0116] As used herein, a "read" (i.e., a read, a sequence read) is a short nucleotide sequence generated by any sequencing process described herein or known in the art. Reads can be generated from one end of a nucleic acid fragment (a "single-end read"), and sometimes are generated from both ends of a nucleic acid (e.g., a double-end read, a two-end read).

[0117] The length of a sequence read is often associated with a particular sequencing technology. For example, high-throughput methods provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). For example, nanopore sequencing can provide sequence reads that can vary in size from tens to hundreds or thousands of base pairs. In some embodiments, the mean, median, average, or absolute length of the sequence reads is about 15 bp to about 900 bp long. In certain embodiments, the mean, median, average, or absolute length of the sequence reads is about 1000 bp or greater.

[0118] In some embodiments, the nominal, average, mean, or absolute length of a read from a single end is sometimes about 15 to about 50 or more contiguous nucleotides, about 15 to about 40 or more contiguous nucleotides, and sometimes about 15 contiguous nucleotides, or about 36 or more contiguous nucleotides. In certain embodiments, the nominal, average, mean, or absolute length of a read from a single end is about 20 to about 30 bases long, or about 24 to about 28 bases long. In certain embodiments, the nominal, average, mean or absolute length of a read from a single end is about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, or about 29 bases in length or more.

[0119] In certain embodiments, the nominal, average, mean, or absolute length of the reads read from both ends is, optionally, from about 10 contiguous nucleotides to about 25 contiguous nucleotides or more (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length or more), from about 15 contiguous nucleotides to about 20 contiguous nucleotides or more, and optionally about 17 contiguous nucleotides, about 18 contiguous nucleotides, about 20 contiguous nucleotides, about 25 contiguous nucleotides, about 36 contiguous nucleotides, or about 45 contiguous nucleotides.

[0120] A read is generally a physical nucleic acid representation of a nucleotide sequence. For example, in a read containing a sequence depicted as ATGC, "A" represents an adenine nucleotide, "T" represents a thymine nucleotide, "G" represents a guanine nucleotide, and "C" represents a cytosine nucleotide as a physical nucleic acid. A sequence read obtained from the blood of a pregnant female may be a read derived from a mixture of fetal and maternal nucleic acids. A mixture of relatively short reads can be converted into a representation of genomic nucleic acids present in the pregnant female and / or fetus by the processes described herein. A mixture of relatively short reads can be converted into a representation of, for example, copy number variations (e.g., maternal and / or fetal copy number variations), genetic variations, or aneuploidies, microduplications, or microdeletions. A read of a mixture of maternal and fetal nucleic acids can be converted into a representation of a composite chromosome or a segment thereof that includes features of one or both maternal and fetal chromosomes. In certain embodiments, "obtaining" a nucleic acid sequence read of a sample obtained from a subject and / or "obtaining" a nucleic acid sequence read of a biological specimen obtained from one or more reference individuals can include directly sequencing the nucleic acid to obtain the sequence information. In some embodiments, "obtaining" can include receiving sequence information obtained directly from the nucleic acid by another person.

[0121] In some embodiments, a representative fraction of the genome is sequenced, sometimes referred to as "coverage" or "coverage factor." For example, 1x coverage indicates that approximately 100% of the nucleotide sequence of the genome is represented by the reads. In some embodiments, "coverage factor" is a term that refers to and compares a previous sequencing run as a reference. For example, a second sequencing run may have half the coverage of the first sequencing run. In some embodiments, genomes are sequenced with redundancy, where a given region of the genome can be covered by two or more reads or overlapping reads (e.g., a "coverage factor" greater than 1, e.g., 2x coverage).

[0122] In some embodiments, the nucleic acid sample obtained from one individual is sequenced.In certain embodiments, the nucleic acid obtained from each of two or more samples is sequenced, where the sample is obtained from one individual or from different individuals.In certain embodiments, the nucleic acid samples obtained from two or more biological samples are pooled, where each biological sample is obtained from one individual or two or more individuals, and the pooled sample is sequenced.In the latter embodiment, the nucleic acid sample obtained from each biological sample is often identified by one or more unique identifiers.

[0123] In some embodiments, sequencing method utilizes identifiers that allow multiplexing of sequencing reaction in sequencing process.The more unique identifiers there are, for example, the more samples and / or chromosomes that can be detected in sequencing process can be multiplexed.Can use any suitable number of unique identifiers (for example, 4, 8, 12, 24, 48, 96 or more) to perform sequencing process.

[0124] The sequencing process sometimes uses a solid phase, sometimes including a flow cell, onto which nucleic acids from a library can be tethered and through which reagents can flow and contact the tethered nucleic acids. Flow cells sometimes include flow cell lanes, and the use of identifiers can facilitate the analysis of several samples in each lane. Flow cells are often solid supports that can be configured to hold bound analytes and / or allow reagent solutions to pass in an orderly fashion over the bound analytes. Flow cells are often planar, optically transparent, generally millimeter or submillimeter scale, and often contain channels or lanes within which analyte-reagent interactions occur. In some embodiments, the number of samples analyzed in a given lane of a flow cell depends on the number of unique identifiers utilized during library preparation and / or probe design. Lanes of a single flow cell. For example, multiplexing using 12 identifiers allows for the simultaneous analysis of 96 samples in an 8-lane flow cell (e.g., equivalent to the number of wells in a 96-well microwell plate). Similarly, multiplexing using, for example, 48 identifiers allows for the simultaneous analysis of 384 samples in an 8-lane flow cell (e.g., equivalent to the number of wells in a 384-well microwell plate). Non-limiting examples of commercially available multiplex sequencing kits include Illumina's Multiplexed Sample Preparation Oligonucleotide Kit, and Multiplexed Sequencing Primer and PhiX Control Kit (e.g., Illumina catalog numbers PE-400-1001 and PE-400-1002, respectively).

[0125] Any suitable method for sequencing nucleic acids can be used, including, but not limited to, Maxim & Gilbert, chain termination, sequencing by synthesis, sequencing by ligation, mass spectrometry sequencing, microscopy-based techniques, and the like, or a combination thereof. In some embodiments, the methods provided herein can use first-generation techniques, such as Sanger sequencing (including automated Sanger sequencing, including microfluidic Sanger sequencing). In some embodiments, sequencing techniques can be used that involve the use of nucleic acid imaging techniques (e.g., transmission electron microscopy (TEM) and atomic force microscopy (AFM)). In some embodiments, high-throughput sequencing methods are used. High-throughput sequencing methods generally involve clonal amplification of DNA templates or single DNA molecules, and sequencing of these templates or molecules in a massively parallel manner, sometimes inside a flow cell. Next-generation (e.g., second- and third-generation) sequencing techniques capable of massively parallel DNA sequencing can be used for the methods described herein, and are collectively referred to herein as "massively parallel sequencing" (MPS). In some embodiments, MPS sequencing methods utilize a targeted approach, in which specific chromosomes, genes, or regions of interest are sequenced. In certain embodiments, a non-targeted approach is used, in which most or all nucleic acids in a sample are randomly sequenced, amplified, and / or captured.

[0126] In some embodiments, targeted approaches for enrichment, amplification, and / or sequencing are used. Targeting approaches often involve isolating, selecting, and / or enriching a subset of nucleic acids in a sample for further processing using sequence-specific oligonucleotides. In some embodiments, a library of sequence-specific oligonucleotides is utilized to target (e.g., hybridize to) one or more sets of nucleic acids in a sample. Often, the sequence-specific oligonucleotides and / or primers are selective for specific sequences (e.g., unique nucleic acid sequences) present in one or more chromosomes, genes, exons, introns, and / or regulatory regions of interest. Any suitable method or combination of methods can be used to enrich, amplify, and / or sequence one or more subsets of targeted nucleic acids. In some embodiments, the targeted sequences are isolated and / or enriched by capturing them on a solid phase (e.g., a flow cell, beads) using one or more sequence-specific anchors. In some embodiments, targeted sequences are enriched and / or amplified by polymerase-based methods (e.g., PCR-based methods with any suitable polymerase-based extension) using sequence-specific primers and / or primer sets. Sequence-specific anchors can often be used as sequence-specific primers.

[0127] MPS sequencing sometimes uses sequencing-by-synthesis and specific visualization processes. Nucleic acid sequencing technologies that can be used in the methods described herein include sequencing-by-synthesis and reversible chain-terminating nucleotide-based sequencing (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ2000; HISEQ2500 (Illumina, San Diego, CA)). This technology allows parallel sequencing of millions of nucleic acid (e.g., DNA) fragments. One example of this type of sequencing technology uses a flow cell containing an optically transparent slide with eight individual lanes, on whose surface oligonucleotide anchors (e.g., adapter primers) are attached. Flow cells are often solid supports that can be configured to hold bound analytes and / or allow reagent solutions to pass in an orderly fashion over the bound analytes. Flow cells are often planar, optically transparent, generally millimeter or submillimeter scale, and often contain channels or lanes within which analyte-reagent interactions occur.

[0128] In some embodiments, sequencing by synthesis involves the iterative addition (e.g., by covalent addition) of nucleotides to a primer or an existing nucleic acid strand in a template-guided manner. After each iterative nucleotide addition, detection is performed, and this process is repeated multiple times until the sequence of the nucleic acid strand is obtained. The length of the resulting sequence depends, in part, on the number of addition and detection steps performed. In some sequencing by synthesis embodiments, one, two, three, or more nucleotides of the same type (e.g., A, G, C, or T) are added and detected in a single nucleotide addition. Nucleotides can be added by any suitable (e.g., enzymatic or chemical) method. For example, in some embodiments, a polymerase or ligase adds nucleotides to a primer or an existing nucleic acid strand in a template-guided manner. Some sequencing by synthesis embodiments use different types of nucleotides, nucleotide analogs, and / or identifiers. In some embodiments, reversible chain-terminating nucleotides and / or removable (e.g., cleavable) identifiers are used. In some embodiments, fluorescently labeled nucleotides and / or nucleotide analogs are used. In certain embodiments, sequencing by synthesis includes cleavage (e.g., cleavage and removal of the identifier) ​​and / or a washing step. In some embodiments, the addition of one or more nucleotides is detected by a suitable method described herein or known in the art, non-limiting examples of which include any suitable imaging device, suitable camera, digital camera, CCD (charge-coupled device)-based imaging device (e.g., CCD camera), CMOS (complementary metal oxide silicon)-based imaging device (e.g., CMOS camera), photodiode (e.g., photomultiplier tube), electron microscopy, field-effect transistor (e.g., DNA field-effect transistor), ISFET ion sensor (e.g., CHEMFET sensor), etc., or combinations thereof. Other sequencing methods that can be used to practice the methods herein include digital PCR and sequencing by hybridization.

[0129] Other sequencing methods that can be used to practice the methods herein include digital PCR and hybridization sequencing. Digital polymerase chain reaction (digital PCR or dPCR) can be used to directly identify and quantify nucleic acids in a sample. In some embodiments, digital PCR can be performed in an emulsion. For example, individual nucleic acids can be separated, for example, in a microfluidic chamber device, and each nucleic acid can be individually amplified by PCR. Nucleic acids can be separated so that only one nucleic acid is present per well. In some embodiments, different probes can be used to distinguish between various alleles (e.g., fetal alleles and maternal alleles). Alleles can be enumerated to determine copy number.

[0130] In certain embodiments, sequencing by hybridization can be used. The method includes contacting a plurality of polynucleotide sequences with a plurality of polynucleotide probes, each of which can optionally be tethered to a substrate. In some embodiments, the substrate can be a flat surface having a large number of known nucleotide sequences. The pattern of hybridization to the array can be used to determine the polynucleotide sequences present in the sample. In some embodiments, each probe is tethered to a bead, such as a magnetic bead. Hybridization to the bead can be identified and used to identify multiple polynucleotide sequences within the sample.

[0131] In some embodiments, the methods described herein can use nanopore sequencing, a single-molecule sequencing technique that directly determines the sequence of a single nucleic acid molecule (e.g., DNA) as it passes through a nanopore.

[0132] Nucleic acid sequencing reads can be obtained using any MPS method, system or technology platform suitable for the practice described herein. Non-limiting examples of MPS platforms include Illumina / Solex / HiSeq (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ2000; HISEQ), SOLiD, Roche / 454, PACBIO and / or SMRT, Helicos True Single Molecule Sequencing, Ion Torrent and ion semiconductor-based sequencing (e.g., developed by Life Technologies), WildFire, 5500, 5500xl W and / or 5500xl W Genetic Analyzer-based technology (e.g., developed and sold by Life Technologies, U.S. Patent Publication No. US20130012399); polony sequencing, pyrosequencing, massively parallel signature sequencing (MPSS), RNA polymerase (RNAP) sequencing, LaserGen systems and methods, nanopore-based platforms, chemically sensitive field effect transistor (CHEMFET) arrays, electron microscopy-based sequencing (e.g., ZS Genetics, Halcyon developed by NASA Molecular), and nanoball sequencing.

[0133] In some embodiments, chromosome-specific sequencing is performed. In some embodiments, chromosome-specific sequencing is performed using DANSR (digital analysis of selected regions). cfDNA-dependent catenation of two locus-specific oligonucleotides via an intervening "bridge" oligonucleotide to form a PCR template allows for the simultaneous quantification of hundreds of loci by digital analysis of selected regions. In some embodiments, chromosome-specific sequencing is performed by generating a library enriched for chromosome-specific sequences. In some embodiments, sequence readings are obtained only for a selected set of chromosomes. In some embodiments, sequence readings are obtained only for chromosomes 21, 18, and 13.

[0134] Read Mapping Sequence reads can be mapped, and the number of reads that map to a particular nucleic acid region (e.g., a chromosome, a portion thereof, or a segment) is referred to as a count. Any suitable mapping method (e.g., a process, an algorithm, a program, software, a module, etc., or a combination thereof) can be used. Specific aspects of the mapping process are described below.

[0135] Mapping of nucleotide sequence reads (i.e., sequence information obtained from fragments whose physical location in the genome is unknown) can be performed in several ways, and often involves aligning the obtained sequence reads with matching sequences in a reference genome. In such alignment, the sequence reads are generally aligned to a reference sequence, and the aligned reads are referred to as "mapped" or "mapped sequence reads." In certain embodiments, mapped sequence reads are referred to as "hits" or "counts." In some embodiments, mapped sequence reads are grouped together and assigned to specific portions according to various parameters, which are discussed in more detail below.

[0136] As used herein, the term "aligned," "alignment," or "aligning" refers to two or more nucleic acid sequences that can be identified as identical (e.g., 100% identical) or partially identical. Alignment can be performed manually or by computer (e.g., software, program, module, or algorithm), non-limiting examples of which include the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysis pipeline. The alignment of sequence reads can be 100% sequence identical. In some cases, the alignment is less than 100% sequence identical (i.e., incomplete match, partial match, partial alignment). In some embodiments, the alignment is about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, or 75% identical. In some embodiments, the alignment includes mismatches. In some embodiments, the alignment includes 1, 2, 3, 4, or 5 mismatches. Two or more sequences can be aligned using either strand. In certain embodiments, a nucleic acid sequence is aligned with the reverse complement of another nucleic acid sequence.

[0137] Various computational methods can be used to map each sequence read to a part.Non-limiting examples of computer algorithms that can be used to align sequences include, but are not limited to, BLAST, BLITZ, FASTA, BOWTIE1, BOWTIE2, ELAND, MAQ, PROBEMATCH, SOAP or SEQMAP, or their modifications or combinations.In some embodiments, sequence reads can be aligned with the sequence in a reference genome.In some embodiments, sequence reads can be found in and / or aligned with the sequences in nucleic acid databases known in the art, including, for example, GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Laboratory) and DDBJ (DNA Databank of Japan).BLAST or similar tools can be used to search identified sequences against sequence databases.Then, for example, search hits (as described below) can be used to sort identified sequences into appropriate parts.

[0138] In some embodiments, the mapped sequence reads and / or information associated with the mapped sequence reads are stored on and / or accessed from a non-transitory computer-readable storage medium in a suitable computer-readable format. As used herein, "computer-readable format" is sometimes loosely referred to as a format. In some embodiments, the mapped sequence reads are stored and / or accessed in a suitable binary format, text format, etc., or a combination thereof. The binary format is sometimes a BAM format. The text format is sometimes a Sequence Alignment / Map (SAM) format. Non-limiting examples of binary and / or text formats include BAM, SAM, SRF, FASTQ, Gzip, etc., or a combination thereof. In some embodiments, the mapped sequence reads are stored in and / or converted to a format that requires less storage space (e.g., fewer bytes) than a conventional format (e.g., a SAM format or a BAM format). In some embodiments, the mapped sequence reads in a first format are compressed into a second format that requires less storage space than the first format. The term "compressed," as used herein, refers to a process of data compression, source coding, and / or bitrate reduction that reduces the size of a computer-readable data file. In some embodiments, the mapped sequence reads are compressed from a binary SAM format. When compressing a file, some data is sometimes lost. Sometimes, no data is lost in the compression process. In some file compression embodiments, some data is replaced with an index and / or reference to another data file containing information about the mapped sequence reads. In some embodiments, the mapped sequence reads are stored in a binary format that includes or consists of a read count, a chromosome identifier (e.g., identifying the chromosome to which the read is mapped), and a chromosome location identifier (e.g., identifying the location on the chromosome to which the read is mapped).In some embodiments, the binary format includes a 20-byte sequence, a 16-byte sequence, an 8-byte sequence, a 4-byte sequence, or a 2-byte sequence. In some embodiments, the mapped read information is stored in a 10-byte format, a 9-byte format, an 8-byte format, a 7-byte format, a 6-byte format, a 5-byte format, a 4-byte format, a 3-byte format, or a 2-byte format. Sometimes, the mapped read data is stored in a 4-byte sequence, including a 5-byte format. In some embodiments, the binary format includes a 5-byte format including a 1-byte chromosome ordinal number and a 4-byte chromosome position. In some embodiments, the mapped reads are stored in a compressed binary format that is about 1 / 100, about 1 / 90, about 1 / 80, about 1 / 70, about 1 / 60, about 1 / 55, about 1 / 50, about 1 / 45, about 1 / 40, or about 1 / 30 of the sequence alignment / map (SAM) format. In some embodiments, the mapped reads are stored in a compressed binary format that is about 1 / 2 to about 1 / 50 (e.g., about 1 / 30, 1 / 25, 1 / 20, 1 / 19, 1 / 18, 1 / 17, 1 / 16, 1 / 15, 1 / 14, 1 / 13, 1 / 12, 1 / 11, 1 / 10, 1 / 9, 1 / 8, 1 / 7, 1 / 6, or about 1 / 5) smaller than the GZip format.

[0139] In some embodiments, the system includes a compression module. In some embodiments, the compression module compresses the mapped sequence read information stored on a non-transitory computer-readable storage medium in a computer-readable format. The compression module sometimes converts the mapped sequence reads to or from an appropriate format. In some embodiments, the compression module can receive the mapped sequence reads in a first format, convert them to a compressed format (e.g., binary format, 5), and transfer the compressed reads to another module (e.g., bias density module, 6). The compression module often provides the sequence reads in a binary format, 5 (e.g., BReads format). Non-limiting examples of compression modules include GZIP, BGZF, BAM, etc., or modifications thereof. Below is an example of converting an integer to a 4-byte array using Java. public static final byte[ ] convertToByteArray(int value) { return new byte[ ] { (byte)(value >>> 24), (byte)(value >>> 16), (byte)(value >>> 8), (byte)value}; }

[0140] In some embodiments, reads can be uniquely or non-uniquely mapped to portions in a reference genome. A read is considered "uniquely mapped" if it aligns with a single sequence in the reference genome. A read is considered "non-uniquely mapped" if it aligns with two or more sequences in the reference genome. In some embodiments, non-uniquely mapped reads are excluded from further analysis (e.g., quantification). In certain embodiments, a specific, low degree of mismatch (0-1) may be accounted for as a single nucleotide polymorphism that may exist between the reference genome and the reads obtained from the individual samples being mapped. In some embodiments, no degree of mismatch is allowed for reads mapped to the reference sequence.

[0141] As used herein, the term "reference genome" can refer to any particular known sequenced or characterized genome of any organism or virus, whether a partial sequence or a complete sequence, that can be used to reference identified sequences from a subject. For example, reference genomes for human subjects and many other organisms can be found at the National Center for Biotechnology Information at the World Wide Web URL ncbi.nlm.nih.gov. "Genome" refers to the complete genetic information of an organism or virus, expressed as a nucleic acid sequence. As used herein, a reference sequence or reference genome is often a compiled or partially compiled genome sequence obtained from one or more individuals. In some embodiments, a reference genome is a compiled or partially compiled genome sequence obtained from one or more human individuals. In some embodiments, a reference genome includes sequences assigned to chromosomes.

[0142] In certain embodiments, when the sample nucleic acid is derived from a pregnant female, the reference sequence is sometimes not derived from the fetus, the fetus's mother, or the fetus's father, and is referred to herein as an "external reference." In some embodiments, a maternal reference can be prepared and used. When preparing a reference from a pregnant female based on an external reference (a "maternal reference sequence"), reads obtained from the pregnant female's DNA, which does not substantially contain fetal DNA, are often mapped and aggregated against the external reference sequence. In certain embodiments, the external reference is derived from the DNA of an individual of substantially the same ethnicity as the pregnant female. The maternal reference sequence may not completely cover the maternal genomic DNA (e.g., it may cover about 50%, 60%, 70%, 80%, 90% or more of the maternal genomic DNA), and the maternal reference may not completely match the maternal genomic DNA sequence (e.g., the maternal reference sequence may contain multiple mismatches).

[0143] In certain embodiments, mappability is evaluated for a genome region (for example, a portion, a genome portion, a portion).Mappability is the ability to unambiguously align a nucleotide sequence read to a portion of a reference genome, typically with only a certain number of mismatches, including, for example, 0, 1, 2 or more mismatches.For a given genome region, a sliding window approach of preset read length can be used, and the obtained read-level mappability values ​​can be averaged to estimate expected mappability.Genome regions that contain stretches of unique nucleotide sequences sometimes have high mappability values.

[0144] portion In some embodiments, mapped sequence reads (i.e., sequence tags) are grouped together and assigned to specific portions (e.g., portions of a reference genome) according to various parameters. Often, individual mapped sequence reads can be used to identify a portion (e.g., the presence, absence, or amount of a portion) present in a sample. In some embodiments, the amount of a portion indicates the amount of a larger sequence (e.g., a chromosome) in the sample. The term "portion" may also be referred to herein as a "genome section," "bin," "region," "section," "reference genome portion," "portion of a chromosome," or "genomic portion." In some embodiments, a portion is an entire chromosome, a segment of a chromosome, a segment of a reference genome, a segment spanning multiple chromosomes, a segment of multiple chromosomes, and / or a combination thereof. In some embodiments, a portion is predefined based on certain parameters. In some embodiments, a portion is arbitrarily defined based on partitioning of the genome (e.g., partitioning by size, GC content, variability in sequencing coverage, contiguous regions, contiguous regions of arbitrarily defined size, etc.). Methods for partitioning a genome (e.g., a reference genome, or parts thereof) are presented herein and described in further detail below.

[0145] In some embodiments, the portions are delineated based on one or more parameters, including, for example, sequence length or one or more specific characteristics. Portions can be selected, filtered, and / or removed from consideration using any suitable criteria known in the art or described herein. In some embodiments, the portions are based on a specific length of the genomic sequence. In some embodiments, the method can include analyzing multiple reads of a sequence mapped to multiple portions. The portions may be approximately the same length, or the portions may be different lengths. In some embodiments, the portions are approximately equal lengths. In some embodiments, the portions are not equal lengths. In some embodiments, the portions are a first equal length in a particular genomic region of interest and a second equal length in a different genomic region of interest. For example, a portion can be 30 kb long in genomic region A and 70 kb long in genomic region B. Methods for optimizing the length of portions in genomic regions of interest are presented herein and described in more detail below. In some embodiments, portions of different lengths are adjusted or weighted. In some embodiments, the genome is partitioned according to the length of an initial portion, and then repartitioned according to the length of one or more optimal portions. In some embodiments, the portions are from about 1 kilobase (kb) to about 1000 kb, from about 1 kb to about 500 kb, from about 10 kb to about 300 kb, from about 10 kb to about 100 kb, from about 20 kb to about 80 kb, from about 30 kb to about 70 kb, from about 40 kb to about 60 kb, and optionally about 50 kb. In some embodiments, the portions are less than 50 kb. In some embodiments, the portions are from about 10 kb to about 20 kb. In some embodiments, the portions are about 30 kb. In some embodiments, the portions are about 10 kb. In some embodiments, the portions are about 20 kb. In some embodiments, the portions are about 30 kb. In some embodiments, the portions are about 40 kb. In some embodiments, the portions are about 50 kb. In some embodiments, the portions are about 60 kb. In some embodiments, the portion is about 70 kb, hi some embodiments, the portion is about 80 kb.In some embodiments, the portion is about 90 kb. In some embodiments, the portion is about 100 kb. In some embodiments, the portion is about 30 kb to about 300 kb. In some embodiments, the portion is about 32 kb. In some embodiments, the portion is about 64 kb. In some embodiments, the portion is about 128 kb. In some embodiments, the portion is about 256 kb. A portion is not limited to a contiguous run of sequence. Thus, a portion can be composed of contiguous and / or discontinuous sequences. A portion is not limited to a single chromosome. In some embodiments, a portion includes all or part of one chromosome, or all or part of two or more chromosomes. In some embodiments, a portion can span one, two, or more chromosomes. Additionally, a portion can span connected or interspersed regions of multiple chromosomes.

[0146] In some embodiments, the portion can be a specific chromosomal segment in a chromosome of interest, such as a chromosome for which genetic variation (e.g., aneuploidy of chromosomes 13, 18, and / or 21, or sex chromosomes) is being evaluated. The portion can also be the genome of a pathogen (e.g., a bacterium, fungus, or virus), or a fragment thereof. The portion can be a gene, a fragment of a gene, a regulatory sequence, an intron, an exon, etc.

[0147] A "segment" of a chromosome is generally a part of a chromosome, typically a different part of a chromosome from the segment. A chromosomal segment is sometimes in a different region of the chromosome than the segment, sometimes does not share polynucleotides with the segment, and sometimes contains polynucleotides that are in the segment. A chromosomal segment often contains a larger number of nucleotides than the segment (e.g., a segment sometimes contains a segment), and a chromosomal segment sometimes contains a smaller number of nucleotides than the segment (e.g., a segment sometimes is within a segment). As used herein, a "genomic region" often contains a larger number of nucleotides than the segment (e.g., a genomic region sometimes contains one or more segments).

[0148] Genome compartmentalization In some embodiments, a genome (e.g., a human genome, a reference genome, a part of a reference genome, a genomic region, one or more chromosomes, a segment of a chromosome) is partitioned into portions based on the information content of specific regions and / or other parameters. Partitioning a genome is sometimes referred to as discretization, binning, segmentation, segmentation, subdivision, division, grouping, aggregation, and aggregation. In some embodiments, a genome is partitioned according to guanine and cytosine (GC) content. In some embodiments, a genome is partitioned according to variability in sequencing coverage. In some embodiments, partitioning a genome can eliminate or reduce bias associated with the information content of specific regions and / or other parameters. In some embodiments, partitioning a genome can establish fine grids (i.e., small portions) for certain regions and coarse grids (i.e., large portions) for other regions. In some embodiments, the genome can be partitioned to eliminate similar regions (e.g., identical or homologous regions or sequences) across the genome and retain only unique regions. The regions excluded during partitioning can be within a single chromosome or across multiple chromosomes. In some embodiments, the partitioned genome is trimmed and optimized for rapid alignment, often allowing for a focus on uniquely identifiable sequences.

[0149] In some embodiments, the partitioning of a genome into regions beyond chromosomal boundaries can be based on the information gain obtained in a classification context. For example, information content can be quantified using a p-value profile, which measures the significance of a particular genomic location for distinguishing between confirmed normal and confirmed abnormal subjects (e.g., euploid and trisomic subjects, respectively). In some embodiments, the partitioning of a genome into regions beyond chromosomal boundaries can be based on any other criteria, such as the speed / convenience of aligning tags, GC content (e.g., high or low GC content), uniformity of GC content, other measures of sequence content (e.g., percentage of individual nucleotides, percentage of pyrimidines or purines, percentage of natural to non-natural nucleic acids, percentage of methylated nucleotides, and CpG content), variability of sequencing coverage, methylation status, duplex melting temperature, amenability to sequencing or PCR, uncertainty values ​​assigned to individual portions of the reference genome, and / or search results targeting specific features. For example, a method for partitioning a genome according to GC content is presented herein. Also presented herein are methods for partitioning genomes, for example, according to variability in sequencing coverage.

[0150] GC Compartmentalization In some embodiments, the genome is partitioned according to the content of guanine and cytosine (GC). GC partitioning is sometimes referred to herein as "wavelet binning." Each chromosome, or each part of a chromosome, is often partitioned separately (i.e., one at a time) from other chromosomes in the reference genome. Although the method described below is generally applied to a single chromosome, one or more chromosomes or all chromosomes in the reference genome can also be partitioned according to the following method.

[0151] In some embodiments, partitioning a genome according to GC content involves generating a GC profile for a chromosome or a segment of a chromosome. The GC profile can be generated by quantifying the GC content (i.e., the number of guanine and cytosine bases) for a given length of genomic sequence (i.e., a window) through a chromosome in a reference genome, or a portion thereof. A window is typically a relatively short length of genomic sequence (e.g., 100 bases to 10 kilobases (kb)). The GC content is typically determined for consecutive windows through a chromosome or a segment thereof. In certain cases, the window is 1 kb. Thus, for example, a GC profile can be generated by quantifying the GC content per consecutive 1 kb window through a chromosome in a reference genome.

[0152] In some embodiments, partitioning a genome according to GC content comprises segmenting. In some embodiments, segmenting modifies and / or transforms a profile (e.g., a GC profile), thereby resulting in one or more decomposed renderings of the profile. The profile subjected to the segmentation process is often a profile of GC content within a reference genome or a portion thereof (e.g., autosomes and sex chromosomes). The decomposed rendering of a profile is often a transformation of the profile. The decomposed rendering of a profile is sometimes a transformation of the profile into a representation of a genome, chromosome, or a segment thereof.

[0153] In certain embodiments, the segmentation process utilized for segmentation locates and identifies one or more GC content levels in the profile that are different (e.g., substantially or significantly different) from one or more other GC content levels in the profile. A GC content level identified in the profile following the segmentation process that is different from other GC content levels in the profile and has edges that are different from other GC content levels in the profile is referred to herein as a wavelet, or more generally, as a GC content level for an individual segment. The segmentation process can generate a decomposition rendering from a profile of GC content or GC content levels that can identify one or more individual segments or wavelets. Individual segments are generally shorter than the segmented entity (e.g., a chromosome, chromosomes, autosomes).

[0154] In some embodiments, segmentation locates and identifies edges of individual segments and wavelets in a profile. In certain embodiments, one or both of the edges of one or more individual segments and the edges of one or more wavelets are identified. For example, the segmentation process can identify the location (e.g., genomic coordinates, e.g., location of a portion) of the right edge and / or the left edge of an individual segment or wavelet in a profile. Individual segments or wavelets often include two edges. For example, an individual segment or wavelet may include a left edge and a right edge. In some embodiments, depending on the display or illustration, the left edge may be the 5'-edge and the right edge may be the 3'-edge of a nucleic acid segment in the profile. In some embodiments, the left edge may be the 3'-edge and the right edge may be the 5'-edge of a nucleic acid segment in the profile. The edges of a profile are often known prior to segmentation; thus, in some embodiments, the edges of a profile determine which edges of a level are 5'-edges and which edges are 3'-edges. In some embodiments, one or both of the edges of the profile and / or the individual segments (eg, wavelets) are edges of chromosomes.

[0155] In some embodiments, the edges of individual segments or wavelets are determined according to a decomposition rendering generated for a reference sample (e.g., a reference profile). In some embodiments, the distribution of null edge heights is determined according to a decomposition rendering of a reference profile (e.g., a profile of a chromosome or a segment thereof). In certain embodiments, the edges of individual segments or wavelets in a profile are identified when the level of the individual segment or wavelet is outside the distribution of null edge heights. In some embodiments, the edges of individual segments or wavelets in a profile are identified according to a Z-score calculated according to a decomposition rendering for the reference profile.

[0156] In some embodiments, segmentation produces two or more individual segments or wavelets in the profile (e.g., two or more fragmentation levels, two or more fragmented segments). In some embodiments, the decomposition rendering derived from the segmentation process is over-segmented or fragmented and includes multiple individual segments or wavelets. In some embodiments, the individual segments or wavelets produced by segmentation are substantially different; in some embodiments, the individual segments or wavelets produced by segmentation are substantially similar. Substantially similar individual segments or wavelets (e.g., substantially similar levels) often refer to two or more adjacent individual segments or wavelets in the segmented profile, each having a GC content level that differs by less than a predetermined level of uncertainty. In some embodiments, substantially similar individual segments or wavelets are adjacent to each other and are not separated by an intervening segment or wavelet. In some embodiments, substantially similar individual segments or wavelets are separated by one or more smaller segments or wavelets. In some embodiments, substantially different individual segments or wavelets are not adjacent. The GC content levels of substantially different individual segments or wavelets are generally substantially different.

[0157] In some embodiments, the segmentation process includes determining (e.g., calculating) a GC-content level (e.g., a quantitative value, e.g., an average or median level), a level of uncertainty (e.g., an uncertainty value), a Z-score, a Z-value, a p-value, etc., or a combination thereof, for one or more individual segments or wavelets (e.g., GC-content levels) within the profile or within that segment. In some embodiments, a GC-content level (e.g., a quantitative value, e.g., an average or median level), a level of uncertainty (e.g., an uncertainty value), a Z-score, a Z-value, a p-value, etc., or a combination thereof is determined (e.g., calculated) for an individual segment or wavelet.

[0158] In some embodiments, segmentation is achieved by one or more sub-processes, non-limiting examples of which include a decomposition generation process (e.g., a wavelet decomposition generation process), thresholding, leveling, smoothing, etc., or combinations thereof. Thresholding, leveling, smoothing, etc., can be performed in conjunction with the decomposition generation process and are described herein below in reference to a wavelet decomposition rendering process.

[0159] In some embodiments, the segmentation is performed according to a wavelet decomposition generation process. In some embodiments, the segmentation is performed according to two or more wavelet decomposition generation processes. In some embodiments, the wavelet decomposition generation processes identify one or more wavelets in the profile and present a decomposed rendering of the profile.

[0160] Segmentation may be performed, in whole or in part, by any suitable wavelet decomposition generating process described herein or known in the art. Non-limiting examples of wavelet decomposition generation processes include Haar wavelet segmentation (Haar, Alfred (1910), "Zur Theorie der orthogonalen Funktionensysteme", Mathematische Annalen, Vol. 69(3): pp. 331-371; Nason, G.P. (2008), "Wavelet methods in Statistics", R. Springer, New York) (e.g., WaveThresh), a suitable binary recursive segmentation process, circular binary segmentation (CBS) (Olshen, A.B., Venkatraman, E.S., Lucito, R., Wigler, M. (2004), "Circular binary segmentation for the analysis of array-based DNA copy number data", Biostatistics, Vol. 5, No. 4: pp. 557-72; Venkatraman, E.S., Olshen, A.B. (2007), "A faster "Circular binary segmentation algorithm for the analysis of array CGH data," Bioinformatics, Vol. 23, No. 6: pp. 657-63), MODWT (Maximal Overlap Discrete Wavelet Transform) (L. Hsu, S. Self, D. Grove, T. Randolph, K. Wang, J. Delrow, L. Loo, and P. Porter, "Denoising array-based comparative genomic hybridization data using wavelets," Biostatistics (Oxford, England), Vol. 6, No. 2, pp. 211-226, 2005), and Stationary Wavelet (SWT) (Y. Wang and S.Wang, "A novel stationary wavelet denoising algorithm for array-based DNA copy number data," International Journal of Bioinformatics Research and Applications, Vol. 3, No. 2, pp. 206-222, 2007), the dual-tree complex wavelet transform (DTCWT) (Nha, N., H. Heng, S. Oraintara, and W. Yuhang (2007), "Denoising of Array-Based DNA Copy Number Data Using the Dual-tree Complex Wavelet Transform," pp. 137-144), convolution with edge detection kernels, Jensen-Shannon divergence, Kullback-Leibler divergence, binary recursive segmentation, Fourier transform, etc., or combinations of these.

[0161] The wavelet decomposition generation process may be represented or implemented by suitable software, modules, and / or code written in any suitable language (e.g., a computer programming language known in the art) and / or operating system, non-limiting examples of which include UNIX, Linux, Oracle, Windows, Ubuntu, ActionScript, C, C++, C#, Haskell, Java, JavaScript, Objective-C, Perl, Python, Ruby, Smalltalk, SQL, Visual Basic, COBOL, Fortran, UML, HTML (e.g., with PHP), PGP, G, R, S, etc., or combinations thereof. In some embodiments, a suitable wavelet decomposition generation process is represented in S code or R code or package (e.g., an R package). R, R source code, R programs, R packages, and R documentation for wavelet decomposition generation processes are available for download from CRAN or a CRAN mirror site (e.g., Comprehensive R Archive Network (CRAN); internet URL: cran.us.r-project.org). CRAN is a network of ftp and web servers that store identical, up-to-date versions of code and documentation for R worldwide. For example, WaveThresh (WaveThresh: Wavelets statistics and transforms; internet URL: cran.r-project.org / web / packages / wavethresh / index.html) and a detailed description of WaveThresh ("WaveThresh" package; internet URL: cran.r-project.org / web / package / wavethresh / wavethresh.pdf) may be available for download.Examples of R code for the CBS method can be downloaded (e.g., DNAcopy; internet URL: bioconductor.org / packages / 2.12 / bioc / html / DNAcopy.html or the "DNAcopy" package; internet URL: bioconductor.org / packages / release / bioc / manuals / DNAcopy / man / DNAcopy.pdf).

[0162] In some embodiments, the wavelet decomposition generation process (e.g., Haar wavelet segmentation, e.g., WaveThresh) includes thresholding. In some embodiments, thresholding distinguishes signal from noise. In certain embodiments, thresholding determines which wavelet coefficients (e.g., nodes) indicate signal and should be retained, and which wavelet coefficients indicate a reflection of noise and should be excluded. In some embodiments, thresholding includes one or more variable parameters, where a user defines the values ​​of the parameters. In some embodiments, thresholding parameters (e.g., thresholding parameters, policy parameters) can describe or prescribe the amount of segmentation utilized in the wavelet decomposition generation process. Any suitable parameter value can be used. In some embodiments, thresholding parameters are used. In some embodiments, the thresholding parameter values ​​are soft thresholding. In certain embodiments, soft thresholding is utilized to exclude small and insignificant coefficients. In certain embodiments, hard thresholding is utilized. In certain embodiments, thresholding includes policy parameters. Any suitable policy value can be used. In some embodiments, the policy used is a "universal" policy, and in some embodiments, the policy used is a "sure" policy.

[0163] In some embodiments, the wavelet decomposition generation process (e.g., Haar wavelet segmentation, e.g., WaveThresh) includes leveling. In some embodiments, after thresholding, several high-level coefficients remain. These coefficients represent steep changes or large spikes in the original signal and, in certain embodiments, are filtered out by leveling. In some embodiments, leveling includes assigning a value to a parameter known as the decomposition level, c. In certain embodiments, the optimal decomposition level is determined according to one or more determined values, such as the length of the chromosome (e.g., the length of the profile), the desired wavelet length, etc., to detect fetal fraction, sequence coverage (e.g., plex level), and noise level of the normalized profile. For a given length (L) of a genome, chromosome, or segment of a profile, the optimal decomposition level is determined according to one or more determined values, such as the length of the chromosome (e.g., the length of the profile), the desired wavelet length, etc., to detect fetal fraction, sequence coverage (e.g., plex level), and noise level of the normalized profile. chr ), the wavelet decomposition level c is sometimes expressed as min =L chr / 2 c+1 According to the minimum wavelet length or minimum part length L min In some embodiments, the decomposition level c is related to the following formula: c=log(L chr / L min );c=log2(L chr / L min )+1;c=log2(L chr / L min In some embodiments, the decomposition level c is about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some embodiments, L min is the given L min In some embodiments, L min is predetermined according to the limit of detection (LoD) analysis described in Example 1. In some embodiments, the amount of sequence coverage (e.g., plex level) and fetal fraction are determined by L minis inversely proportional to. For example, as the amount of fetal fraction in a sample increases, the desired minimum wavelet length (e.g., minimum segment length) decreases (i.e., resolution increases). In some embodiments, as sequencing coverage increases, the desired minimum wavelet length (e.g., minimum segment length) decreases (i.e., resolution increases). In some embodiments, thresholding is performed before leveling, and in some cases, thresholding is performed after leveling.

[0164] In some embodiments, the decomposed rendering is polished, resulting in a polished decomposed rendering. In some embodiments, the decomposed rendering is polished two or more times. In some embodiments, the decomposed rendering is polished before and / or after one or more steps of the segmentation process. In some embodiments, the segmentation of the genome includes two or more segmentation processes, each segmentation process including one or more polishing processes. A decomposed rendering may refer to a polished or unpolished decomposed rendering.

[0165] Thus, in some embodiments, the segmentation process includes a refinement process. In some embodiments, the refinement process identifies two or more substantially similar individual segments or wavelets (e.g., in a decomposition rendering) and integrates them into a single individual segment or wavelet. In some embodiments, the refinement process identifies two or more adjacent segments or wavelets that are substantially similar and integrates them into a single level, segment, or wavelet. Thus, in some embodiments, the refinement process includes an integration process. In certain embodiments, adjacent fragmented individual segments or wavelets are integrated according to their GC content level. In some embodiments, the integration of two or more adjacent individual segments or wavelets includes calculating a median level for the two or more adjacent individual segments or wavelets that are ultimately integrated. In some embodiments, two or more adjacent individual segments or wavelets that are substantially similar are integrated, thereby resulting in a single segment, wavelet, or GC content level as a result of the refinement process. In certain embodiments, two or more adjacent individual segments or wavelets are integrated by a process described by Willenbrock and Fridly (Willenbrock H, Fridlyand J, A comparison study: applying segmentation to array CGH data for downstream analyses, Bioinformatics (2005, November 15), 21(22):4084-91). In some embodiments, two or more adjacent individual segments or wavelets are integrated by a process known as GLAD and described in Hupe, P. et al. (2004), "Analysis of array CGH data: from signal ratio to gain and loss of DNA regions," Bioinformatics, 20, 3413-3422.

[0166] In some embodiments, the segmentation process comprises a "sliding edge" process or a "sliding window" process. A suitable "sliding edge" process can be used directly to validate individual segments during decomposition rendering, or can be adapted for this purpose. In some embodiments, the "sliding edge" process comprises segmenting an individual segment into multiple subsets of parts. In some embodiments, an individual segment is a set of parts for an entire chromosome or a segment of a chromosome.

[0167] In certain embodiments, the "sliding edge" process involves segmenting the identified individual segments into multiple subsets of portions, where each subset of portions represents individual segments that are similar but have different edges. In some embodiments, the original identified individual segments are incorporated into the analysis. For example, the original identified individual segments are incorporated as one of multiple subsets of portions. The subset of portions can be determined by altering one or both edges of the original identified individual segments in any suitable manner. In some embodiments, the left edge can be altered, thereby generating individual segments with different left edges. In some embodiments, the right edge can be altered, thereby generating individual segments with different right edges. In some embodiments, both the right edge and the left edge can be altered. In some embodiments, the edge is altered by moving the edge to the left or right of the original edge by one or more adjacent reference genome portions.

[0168] In some embodiments, one or both edges are altered by 5 to 30 reference genome portions. In some embodiments, the edges are moved in either direction by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 reference genome portions. In some embodiments, regardless of portion size, for one or both edges, the edges are altered to generate edges in the range of about 100,000 to about 2,000,000 base pairs, 250,000 to about 1,500,000 base pairs, or about 500,000 to about 1,000,000 base pairs. In some embodiments, regardless of the size of the portion, for one or both edges, the edges are varied to generate edges in the range of about 500,000, 600,000, 700,000, 750,000, 800,000, 900,000, or about 1,000,000 base pairs.

[0169] In some embodiments, the identified individual segments include a first end and a second end, and segmentation includes (i) recursively excluding one or more moieties from the first end of the set of moieties, thereby subjecting each recursively excluding subset of moieties; (ii) terminating the recursive excluding of (i) after n iterations, thereby resulting in n+1 subsets of moieties, where the set of moieties is a subset, each subset including a different number of moieties, the end of the first subset, and the end of the second subset; (iii) excluding one or more moieties from the end of the second subset of each of the n+1 subsets of moieties resulting from the recursive excluding of (ii); and (iv) terminating the recursive excluding of (iii) after n iterations, thereby resulting in a plurality of subsets of moieties. In some embodiments, the plurality of subsets equals (n+1) subsets. In some embodiments, n equals an integer between 5 and 30. In some embodiments, n is equal to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30.

[0170] In certain embodiments of the sliding edge method, a level of significance (e.g., Z-score, p-value) is determined for each subset of portions of the reference genome, and an average, mean, or median level of significance is determined according to the level of significance determined for all of the subsets. In some embodiments, the level of significance is a Z-score or p-value. In some embodiments, the Z-score is calculated using the following formula: Z i =(E i -Med.E (n) ) / MAD

[0171] [In the formula, E i is the quantitative determination of the level of individual segment i, Med.E (n)is the median level for all individual segments generated by the sliding edge process, and MAD is Med.E (n) is the median absolute deviation for Z i is the resulting Z-score for individual segment i. In some embodiments, MAD can be replaced by any suitable measure of uncertainty. In some embodiments, E i is any suitable measure of level, non-limiting examples of which include median level, mean level, average level, sum, etc. of the number of counts for the portion.

[0172] In some embodiments, the segmentation process described above is applied to the GC profile to produce individual segments. The segmentation process can be performed on GC content levels, as described herein. In certain cases, windows with similar GC content levels are merged into individual segments during the segmentation process. In certain embodiments, the segmentation process produces a decomposed rendering that includes the individual segments. In certain embodiments, the chromosome is partitioned into multiple parts according to the individual segments. Thus, in certain embodiments, the location and length of the individual segments are the same as or similar to the location and length of the parts in the chromosome partitioned by GC.

[0173] Segmentation by variability in sequencing coverage In some embodiments, the genome is partitioned according to variability in sequencing coverage. In certain cases, sequencing of nucleic acids containing a mixture of maternal and fetal genetic material (e.g., ccfDNA) can be characterized by variations in sequencing coverage as a function of location in the genome. Without being limited by theory, certain genomic regions may exhibit abundant sequencing data, while other regions may exhibit sparse sequencing data. Optimizing the length of segments for certain regions in terms of variability in sequencing coverage can allow for the use of fine grids (i.e., small segments) for certain regions and coarse grids (i.e., large segments) for other regions. A fine grid can be useful, for example, for detecting small genetic variations (e.g., small microdeletions or microduplications). A coarse grid can be useful, for example, for capturing sequence reads that may be filtered out using a small or standard grid rather than a coarse grid.

[0174] In some embodiments, partitioning the genome involves determining the variability of sequencing coverage across the reference genome. In some embodiments, a training set of nucleotide sequence reads mapped to portions of the reference genome is used to determine the variability of sequencing coverage. The training set may include, for example, nucleotide sequence reads from multiple samples including a mixture of maternal nucleic acid and fetal ccf nucleic acid (e.g., ccfDNA). The variability of sequencing coverage across the reference genome can be determined by quantifying the sequence reads for the training set. Quantifying the sequence reads may include quantifying raw sequence reads and / or normalized sequence reads for one or more portions or regions, as described herein. In certain cases, an average sequence read count is determined. In certain cases, an average normalized sequence read count is determined.

[0175] In some embodiments, partitioning the genome includes selecting the length of the initial portion. The length of the initial portion can be selected, for example, according to certain characteristics of the training set. For example, the length of the initial portion can be selected according to the sequencing depth for the training set. In certain cases, the length of the initial portion can be selected for the training set according to the average fetal fraction. In some embodiments, the length of the initial portion is selected according to the sequencing depth and the average fetal fraction for the training set. The average fetal fraction for the training set can be determined using any suitable method for determining fetal fraction, known in the art or described herein (e.g., determining a portion-specific fetal fraction). In some embodiments, the average fetal fraction for the training set is known or can be calculated based on sample records. The length of the initial portion can be between about 1 kb and about 1000 kb. In some embodiments, the length of the initial portion is about 10 kb. In some embodiments, the length of the initial portion is about 20 kb. In some embodiments, the length of the initial portion is about 30 kb. In some embodiments, the length of the initial portion is about 40 kb. In some embodiments, the length of the initial portion is about 50 kb. In some embodiments, the length of the initial portion is about 60 kb. In some embodiments, the length of the initial portion is about 70 kb. In some embodiments, the length of the initial portion is about 80 kb. In some embodiments, the length of the initial portion is about 90 kb. In some embodiments, the length of the initial portion is about 100 kb. In some embodiments, the length of the initial portion is less than 50 kb. Generally, a larger initial portion length (e.g., greater than 50 kb) can be selected for training sets with a small average fetal fraction (e.g., less than 10%) and / or a small sequencing depth, and a smaller initial portion length (e.g., less than 50 kb) can be selected for training sets with a large average fetal fraction (e.g., 10%-20%) and / or a large sequencing depth.In some embodiments, the total number of portions for a reference genome can be determined according to the length of the initial portions and the total genome size.

[0176] In some embodiments, the partitioning of the genome comprises partitioning at least two genomic regions according to the size of the initial portion. The genomic regions can be selected according to one or more known genetic variations (e.g., any form of copy number variation) that may exist in the region (e.g., microdeletion, microduplication, aneuploidy), or can be selected randomly. The genomic regions can be chromosomes or chromosomal segments. Generally, pairs of genomic regions are selected for comparison of variability in sequencing coverage, as described below. If necessary, additional pairs of genomic regions can be selected. The first pair of genomic regions can include a first genomic region and a second genomic region. The first genomic region and the second genomic region are often substantially similar or equal in size (i.e., length). For example, the first genomic region and the second genomic region can differ in length by about 1 kb or less.

[0177] In some embodiments, partitioning the genome comprises comparing the variability of sequencing coverage for each pair of genomic regions. Comparing the variability of sequencing coverage is performed by calculating a proportionality coefficient (P) according to the following formula: P=(var1 / var2) 1 / 3 Formula A

[0178] where var1 is the variability of sequencing coverage of the first genomic region and var2 is the variability of sequencing coverage of the second genomic region. In some embodiments, the variability of the sequencing coverage of the first genomic region is determined from the nucleotide sequence read counts for the first genomic region, or a derivative thereof, and the variability of the sequencing coverage of the second genomic region is determined from the nucleotide sequence read counts for the second genomic region, or a derivative thereof. As used herein, a derivative of the sequence read counts may be a processed sequence read count (e.g., a filtered, adjusted, and / or normalized sequence read count as described herein). In some embodiments, the variability of the sequencing coverage of the first genomic region is determined from the average nucleotide sequence read counts for the first genomic region, or a derivative thereof, and the variability of the sequencing coverage of the second genomic region is determined from the average nucleotide sequence read counts for the second genomic region, or a derivative thereof. In some embodiments, the average nucleotide sequencing read counts for each genomic region are determined using a training set. In some embodiments, the nucleotide sequence read counts are normalized nucleotide sequence read counts, hi some embodiments, the average nucleotide sequence read counts are average normalized nucleotide sequence read counts.

[0179] In some embodiments, partitioning the genome includes recalculating the number of parts for the genomic region. Typically, recalculating the number of parts for the genomic region is performed according to a proportionality factor (e.g., the proportionality factor described above). In some embodiments, recalculating the number of parts for the genomic region is performed according to the proportionality factor and the total number of parts (e.g., for the reference genome) determined from the initial part size as described above. For example, let N1 be the number of parts for genomic region 1, N2 be the number of parts for genomic region 2, and N3 be the number of parts for genomic region 3. Let N be the total number of regions, and the ratio between these numbers (derived from formula A above) be calculated as follows: N1 / N2=P1 N1 / N3=P2 Assuming that N1+N2+N3=N Given that N2=N1 / P1 and N3=N1 / P2, N1=N×P1×P2 / (P1×P2+P1+P2), N2=N×P2 / (P1×P2+P1+P2), and

[0180] N3=N×P1 / (P1×P2+P1+P2) This becomes:

[0181] In some embodiments, partitioning the genome includes determining the length of the optimized portion according to the number of recalculated portions for the genomic region. The length of the optimized portion can be between about 1 kilobase (kb) and about 1000 kb. In some embodiments, the length of the optimized portion is between about 1 kb and about 500 kb, between about 10 kb and about 300 kb, between about 10 kb and about 100 kb, between about 20 kb and about 80 kb, between about 30 kb and about 70 kb, or between about 40 kb and about 60 kb. In some embodiments, the length of the optimized portion is less than 50 kb. In some embodiments, the length of the optimized portion is between about 10 kb and about 20 kb. In some embodiments, the length of the optimized portion is about 30 kb. In some embodiments, the length of the optimized portion is about 10 kb. In some embodiments, the length of the optimized portion is about 20 kb. In some embodiments, the optimized portion is about 30 kb in length. In some embodiments, the optimized portion is about 40 kb in length. In some embodiments, the optimized portion is about 50 kb in length. In some embodiments, the optimized portion is about 60 kb in length. In some embodiments, the optimized portion is about 70 kb in length. In some embodiments, the optimized portion is about 80 kb in length. In some embodiments, the optimized portion is about 90 kb in length. In some embodiments, the optimized portion is about 100 kb in length.

[0182] In some embodiments, partitioning a genome comprises repartitioning a genomic region into a plurality of portions according to optimized portion sizes. In some embodiments, the plurality of portions comprises portions of constant (i.e., equal or substantially equal) length. In some embodiments, the plurality of portions comprises portions of varying size. In certain cases, a genome partitioning method may also include an additional genome partitioning method (e.g., GC partitioning, as described herein), which can result in portions of varying size. In some embodiments, partitioning a genome comprises repartitioning one or more additional genomic regions of the reference genome using the methods described herein. In some embodiments, partitioning a genome comprises repartitioning all or substantially all of the genomic regions of the reference genome using the methods described herein.

[0183] In some embodiments, partitioning the genome includes estimating the fetal fraction for the test sample. The fetal fraction can be estimated using any suitable method for estimating fetal fraction known in the art or described herein (e.g., part-specific fetal fraction estimation, fetal quantification assay, SNP-based fetal fraction estimation, Y chromosome fetal fraction estimation). Estimating the fetal fraction optionally includes determining an error value. The error value can be expressed (or displayed), for example, as an uncertainty value, calculated variance, standard deviation, Z-score, p-value, mean absolute deviation, average absolute deviation, median absolute deviation, etc. In some embodiments, the error value defines a range above and below the estimated fetal fraction. In some embodiments, the error is expressed as a range of values ​​(e.g., a confidence interval). In some embodiments, the region-specific fetal fraction is determined for a genomic region according to the correlation between the nucleotide sequence read counts per portion (e.g., raw sequence read counts, normalized sequence read counts) and a weighting factor (e.g., determining the portion-specific fetal fraction described herein).

[0184] In some embodiments, partitioning the genome comprises determining the size (i.e., length) of a minimum genomic region. In some embodiments, partitioning the genome comprises determining the size of a minimum genomic region detectable for a sample having a given fetal fraction (e.g., estimated according to the methods described above). In certain cases, the size of the minimum genomic region is determined according to the limit of detection (LoD) analysis for a particular genetic abnormality described in Example 1 and presented in FIG. 7. In certain cases, the size of the minimum genomic region is determined according to a particular confidence interval for the fetal fraction. For example, the size of the minimum genomic region can be determined according to the upper 80% confidence interval for the fetal fraction. In certain embodiments, the size of the minimum genomic region can be determined according to the upper 90% confidence interval for the fetal fraction. In certain embodiments, the size of the minimum genomic region can be determined according to the upper 95% confidence interval for the fetal fraction. In certain embodiments, the size of the minimum genomic region can be determined according to the upper 99% confidence interval for the fetal fraction.

[0185] In some embodiments, partitioning the genome involves determining the size (i.e., length) of a smallest local genomic region. "Local" refers to within a particular repartitioned genomic region. In some embodiments, partitioning the genome involves determining the size of a local genomic region detectable for a sample with an average fetal fraction. The average fetal fraction can be between about 5% and about 20%. For example, the average fetal fraction can be about 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, 10%, 10.5%, 11%, 11.5%, 12%, 12.5%, 13%, 13.5%, 14%, 14.5%, 15%, 15.5%, 16%, 16.5%, 17%, 17.5%, 18%, 18.5%, 19%, or 19.5%. In certain cases, the size of the local minimum genomic region is determined according to the limit of detection (LoD) analysis described in Example 1 and presented in Figure 7 for a particular genetic abnormality.

[0186] In some embodiments, partitioning the genome comprises determining the size (i.e., length) of a local minimum genomic region. In some embodiments, partitioning the genome comprises determining the size of a local minimum genomic region detectable for a sample with a given fetal fraction (e.g., estimated according to the methods described above). In certain cases, the size of the local minimum genomic region is determined according to the limit of detection (LoD) analysis for a particular genetic abnormality described in Example 1 and presented in FIG. 7. In certain cases, the size of the local minimum genomic region is determined according to a particular confidence interval for the fetal fraction. For example, the size of the local minimum genomic region can be determined according to the upper 80% confidence interval for the fetal fraction. In certain embodiments, the size of the local minimum genomic region can be determined according to the upper 90% confidence interval for the fetal fraction. In certain embodiments, the size of the local minimum genomic region can be determined according to the upper 95% confidence interval for the fetal fraction. In certain embodiments, the size of the local minimal genomic region can be determined according to the upper 99% confidence interval for the fetal fraction.

[0187] In certain cases, the size of the smallest genomic region or the size of a local smallest genomic region may span a single portion. To address this possibility, partitioning the genome may further include adjusting the number of portions for each genomic region so that each region includes at least two portions. Adjusting the number of portions for each genomic region may generate a refined grid (i.e., a refined, re-partitioned genome).

[0188] In some embodiments, partitioning the genome includes re-estimating the fetal fraction from the refined, re-partitioned genomic region. In some embodiments, the re-estimated fetal fraction is compared to an initial fetal fraction estimate for the sample. In some embodiments, the re-estimated fetal fraction is compared to a region-specific fetal fraction estimate. In some embodiments, certain method components are repeated if the initial estimated fetal fraction or the region-specific fetal fraction differs from the re-estimated fetal fraction by a predetermined tolerance value. The predetermined tolerance value may be between about 1% and about 25%. For example, the predetermined tolerance value may be about 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, or 24%.

[0189] Count number In some embodiments, sequence reads that are mapped or partitioned based on selected features or variables can be quantified to determine the number of reads that map to one or more portions (e.g., reference genome portions). In certain embodiments, the quantity of sequence reads that map to a portion is referred to as a count (e.g., 1 count). Often, the count is associated with the portion. In certain embodiments, the counts for two or more portions (e.g., a series of portions) are mathematically manipulated (e.g., averaged, added, normalized, etc., or a combination thereof). In some embodiments, the count is determined from some or all of the sequence reads that map to (i.e., are associated with) the portion. In certain embodiments, the count is determined from a predefined subset of mapped sequence reads. Any suitable feature or variable can be used to define or select the predefined subset of mapped sequence reads. In some embodiments, the predefined subset of mapped sequence reads can include 1 to n sequence reads, where n is a number equal to the sum of all sequence reads generated from the test subject or reference subject sample.

[0190] In certain embodiments, the counts are derived from sequence reads that are processed or manipulated by suitable methods, operations, or mathematical processes known in the art. Counts (e.g., number counts) can be determined by suitable methods, operations, or mathematical processes. In certain embodiments, the counts are derived from sequence reads associated with a moiety, where some or all of the sequence reads are weighted, removed, filtered, normalized, adjusted, averaged, derived as an average, added, subtracted, or a combination thereof. In some embodiments, the counts are derived from raw sequence reads and / or filtered sequence reads. In certain embodiments, the count value is determined by mathematical processes. In certain embodiments, the count value is the average, mean, or sum of the sequence reads mapped to a moiety. Often, the count is the average number of counts. In some embodiments, the count is associated with an uncertainty value.

[0191] In some embodiments, the counts can be manipulated or transformed (e.g., normalized, combined, added, filtered, selected, averaged, derived as an average, etc., or combinations thereof). In some embodiments, the counts can be transformed to obtain normalized counts. The counts can be processed (e.g., normalized) by methods known in the art and / or described herein (e.g., fractional normalization, median count (median bin count, median fractional count) normalization, GC content normalization, linear least squares regression and nonlinear least squares regression, LOESS (e.g., GC LOESS), LOWESS, ChAI, principal component normalization, RM, GCRM, cQn, and / or combinations thereof). In certain embodiments, the counts can be processed (e.g., normalized) by one or more of LOESS, median count (median bin count, median fractional count) normalization, and principal component normalization. In certain embodiments, the counts can be processed (e.g., normalized) by LOESS followed by normalization by median counts (median bin counts, median fractional counts). In certain embodiments, the counts can be processed (e.g., normalized) by LOESS followed by normalization by median counts (median bin counts, median fractional counts), followed by normalization by principal components.

[0192] The counts (e.g., raw, filtered, and / or normalized counts) can be processed and normalized to one or more levels. Levels and profiles are described in more detail below. In certain embodiments, the counts can be processed and / or normalized to a reference level. Reference levels are described later in this specification. The processed counts (e.g., processed counts) according to level can be associated with an uncertainty value (e.g., a calculated variance, error, standard deviation, Z-score, p-value, mean absolute deviation, etc.). In some embodiments, the uncertainty value defines a range above and below a certain level. A value for deviation can be used in place of the uncertainty value; non-limiting examples of measures of deviation include standard deviation, average absolute deviation, median absolute deviation, standard score (e.g., Z-score, Z-score, normal score, standardized variable), etc.

[0193] Counts are often obtained from a nucleic acid sample from a pregnant female with a fetus.Counts of the nucleic acid sequence readings that are mapped to one or more parts are often the counts that represent both the fetus and the mother of the fetus (for example, pregnant female subjects).In certain embodiments, some of the counts that are mapped to a certain part are derived from the genome of the fetus, and some of the counts that are mapped to the same part are derived from the genome of the mother.

[0194] Data Processing and Normalization Herein, the mapped sequence reads that have been counted are referred to as raw data because they represent unmanipulated counts (e.g., raw counts). In some embodiments, the sequence read data in a dataset can be further processed (e.g., mathematically and / or statistically manipulated) and / or displayed to facilitate an outcome. In certain embodiments, datasets, including larger datasets, may benefit from preprocessing to facilitate further analysis. Preprocessing of a dataset sometimes includes removing duplicated and / or uninformative portions or portions of the reference genome (e.g., portions of the reference genome with uninformative data, duplicated mapped reads, portions with a median count of zero, over- or under-represented sequences). Without being limited by theory, data processing and / or preprocessing can (i) remove noisy data, (ii) remove uninformative data, (iii) remove redundant data, (iv) reduce the complexity of a larger data set, and / or (v) facilitate the conversion of data from one form to one or more other forms. As used herein, the terms "preprocessing" and "processing" are collectively referred to as "processing" when used in reference to data or data sets. Processing can make data more suitable for further analysis and, in some embodiments, can produce an outcome. In some embodiments, one or more or all of the processing methods (e.g., normalization methods, filtering portions, mapping, validation, etc., or a combination thereof) are performed by a processor, microprocessor, computer in conjunction with memory, and / or by a microprocessor-controlled device.

[0195] The term "noisy data," as used herein, refers to (a) data that, when analyzed or plotted, shows significant variance between data points, (b) data with significant standard deviations (e.g., greater than 3 standard deviations), (c) data with significant standard errors of the mean, and the like, as well as combinations of the above. Noisy data sometimes arises due to the quantity and / or quality of the starting material (e.g., nucleic acid sample), and sometimes arises from part of the process for preparing or replicating the DNA used to obtain the sequence reads. In certain embodiments, noise results from overrepresentation of certain sequences when prepared using PCR-based methods. The methods described herein can reduce or eliminate the contribution of noisy data, thus reducing the effect of noisy data on the resulting outcome.

[0196] The terms "non-informative data," "non-informative portion of the reference genome," and "non-informative portion," as used herein, refer to portions or data derived therefrom having a numerical value significantly different from a predetermined threshold value or a numerical value that falls outside a predefined limit range of values. The terms "threshold" and "threshold value," as used herein, refer to any number calculated using a qualified dataset and serving as a limit for diagnosing a genetic variation (e.g., copy number variation, aneuploidy, microduplication, microdeletion, chromosomal abnormality, etc.). In certain embodiments, the results obtained by the methods described herein exceed the threshold value, and the subject is diagnosed with a genetic variation (e.g., trisomy 21). In some embodiments, the threshold value or range of values ​​is often calculated by mathematically and / or statistically manipulating sequence read data (e.g., obtained from the reference and / or subject); in certain embodiments, the sequence read data manipulated to obtain the threshold value or range of values ​​is the sequence read data (e.g., obtained from the reference and / or subject). In some embodiments, an uncertainty value is determined. The uncertainty value is generally a measure of variance or error and may be any suitable measure of variance or error. In some embodiments, the uncertainty value is a standard deviation, a standard error, a calculated variance, a p-value, or a mean absolute deviation (MAD). In some embodiments, the uncertainty value may be calculated according to the formulas described herein.

[0197] Any suitable procedure can be used to process the datasets described herein.Non-limiting examples of suitable procedures for processing datasets include filtering, normalizing, weighting, monitoring peak heights, monitoring peak areas, monitoring peak edges, determining area ratios, mathematically processing data, statistically processing data, applying statistical algorithms, analyzing with a fixed variable, analyzing with an optimized variable, plotting data, identifying patterns or trends, and further processing, etc., and combinations thereof.In some embodiments, datasets are processed based on various features (e.g., GC content, overlapping, mapped reads, centromeric regions, telomeric regions, etc., and combinations thereof) and / or variables (e.g., fetal sex, maternal age, maternal ploidy, percent contribution of fetal nucleic acids, etc., or combinations thereof).In certain embodiments, processing datasets as described herein can reduce the complexity and / or dimensionality of large and / or complex datasets. Non-limiting examples of complex datasets include sequence read data generated from one or more test subjects and multiple reference subjects of different age and ethnic backgrounds. In some embodiments, the dataset can include thousands to millions of sequence reads for each test subject and / or reference subject.

[0198] In certain embodiments, data processing can be performed in any number of steps. For example, in some embodiments, data can be processed using only a single processing procedure, and in certain embodiments, data can be processed using one or more, five or more, ten or more, or twenty or more processing steps (e.g., one or more processing steps, two or more processing steps, three or more processing steps, four or more processing steps, five or more processing steps, six or more processing steps, seven or more processing steps, eight or more processing steps, nine or more processing steps, ten or more processing steps, eleven or more processing steps, twelve or more processing steps, thirteen or more processing steps, fourteen or more processing steps, fifteen or more processing steps, sixteen or more processing steps, seventeen or more processing steps, eighteen or more processing steps, nineteen or more processing steps, or twenty or more processing steps). In some embodiments, the processing step can be the same step repeated two or more times (e.g., filtering two or more times, normalizing two or more times), while in certain embodiments, the processing step can be two or more different processing steps performed simultaneously or sequentially (e.g., filtering, normalizing; normalizing, monitoring peak height and edges; filtering, normalizing, normalizing to a reference, statistically manipulating, determining p-values, etc.). In some embodiments, any suitable number and / or combination of the same or different processing steps can be utilized to process sequence read data to facilitate obtaining an outcome. In certain embodiments, processing a dataset according to the criteria described herein can reduce the complexity and / or dimensionality of the dataset.

[0199] In some embodiments, one or more processing steps can include one or more filtering steps. The term "filtering," as used herein, refers to removing a portion or portions of a reference genome from consideration. Portions of a reference genome can be selected for removal based on any appropriate criteria, including, but not limited to, redundant data (e.g., duplicated or overlapping mapped reads), uninformative data (e.g., portions of a reference genome with a median count of zero), portions of a reference genome with over- or under-represented sequences, noisy data, etc., or combinations of the above. The filtering process often involves removing one or more portions of a reference genome from consideration and subtracting the counts in one or more portions of the reference genome selected for removal from the counts tallied or summed for the reference genome, one or more chromosomes, or portion of the genome under consideration. In some embodiments, portions of a reference genome can be removed sequentially (e.g., removed one by one, allowing evaluation of the effect of removing each individual portion), and in certain embodiments, all portions of the reference genome marked for removal can be removed simultaneously. In some embodiments, portions of the reference genome characterized by variance above or below a certain level are removed, sometimes referred to herein as filtering "noisy" portions of the reference genome. In certain embodiments, the filtering process comprises obtaining data points from the dataset that deviate from the mean profile level of the portion, chromosome, or chromosome segment by a predetermined multiple of the profile variance, and in certain embodiments, the filtering process comprises removing data points from the dataset that do not deviate from the mean profile level of the portion, chromosome, or chromosome segment by a predetermined multiple of the profile variance. In some embodiments, the filtering process is used to reduce the number of candidate portions of the reference genome to be analyzed for the presence or absence of genetic variation.Reducing the number of candidate portions of the reference genome that are analyzed for the presence or absence of genetic variations (e.g., microdeletions, microduplications) often reduces the complexity and / or dimensionality of the dataset, sometimes increasing the speed of searching for and / or identifying genetic variations and / or abnormalities by two orders of magnitude or more.

[0200] In some embodiments, one or more processing steps may include one or more normalization steps. Normalization may be performed by any suitable method described herein or known in the art. In certain embodiments, normalization involves adjusting values ​​measured on different scales to a conceptually common scale. In certain embodiments, normalization involves sophisticated mathematical adjustments to bring the probability distributions of the adjusted values ​​into alignment. In some embodiments, normalization involves fitting the distributions to a normal distribution. In certain embodiments, normalization involves mathematical adjustments that allow for comparison of corresponding normalized values ​​for different data sets in a manner that eliminates the effects of certain global influences (e.g., errors and anomalies). In certain embodiments, normalization involves scaling. Normalization sometimes involves division of one or more data sets by a predetermined variable or formula. Normalization sometimes involves division of one or more data sets by a predetermined variable or formula. Non-limiting examples of normalization methods include sectional normalization, GC content normalization, median count (median bin count, median sectional count) normalization, linear least squares regression and non-linear least squares regression, LOESS, GC LOESS, LOWESS (locally weighted scatterplot flattening), ChAI, principal component normalization, repeat masking (RM), GC normalized repeat masking (GCRM), cQn, and / or combinations thereof. In some embodiments, the determination of the presence or absence of genetic variations (e.g., aneuploidy, microduplication, microdeletion) utilizes normalization methods (e.g., section normalization, GC content normalization, median counts (median bin counts, median section counts) normalization, linear least squares regression and non-linear least squares regression, LOESS, GC LOESS, LOWESS (locally weighted scatterplot flattening), ChAI, principal component normalization, repeat masking (RM), GC-normalized repeat masking (GCRM), cQn, normalization methods known in the art, and / or combinations thereof).In some embodiments, determining the presence or absence of genetic variations (e.g., aneuploidy, microduplication, microdeletion) utilizes one or more of LOESS, median count (median bin count, median partial count) normalization, and principal component normalization. In some embodiments, determining the presence or absence of genetic variations utilizes LOESS followed by median count (median bin count, median partial count) normalization. In some embodiments, determining the presence or absence of genetic variations utilizes LOESS followed by median count (median bin count, median partial count) normalization, and principal component normalization.

[0201] Any suitable number of normalizations can be used. In some embodiments, a dataset can be normalized one or more times, five or more times, ten or more times, or even twenty or more times. A dataset can be normalized to a value (e.g., normalization value) that represents any suitable feature or variable (e.g., sample data, reference data, or both). Non-limiting examples of the types of data normalization that can be used include: normalizing the raw count data for one or more selected test or reference portions to the total number of counts that are mapped to the chromosome or whole genome to which the selected portion or section is mapped; normalizing the raw count data for one or more selected portions to the median of the reference counts for one or more portions or chromosomes to which the selected portion or section is mapped; normalizing the raw count data to pre-normalized data or their derivatives; and normalizing the pre-normalized data to one or more other predetermined normalization variables. Normalizing a dataset sometimes has the effect of isolating statistical errors, depending on the feature or characteristic selected as the predetermined normalization variable. Also, normalizing a dataset sometimes allows for comparison of data features of data with different scales by giving the data a common scale (e.g., a predetermined normalization variable). In some embodiments, one or more normalizations to a statistically derived value can be used to minimize data differences and reduce the importance of outlier data. Normalizing a portion or a portion of a reference genome with respect to a normalization value is sometimes referred to as "partial normalization."

[0202] In certain embodiments, processing steps involving normalization include normalizing to a stationary window; in some embodiments, processing steps involving normalization include normalizing to a moving or sliding window. The term "window," as used herein, refers to one or more portions selected for analysis and is sometimes used as a reference for comparison (e.g., for normalization and / or other mathematical or statistical operations). The term "normalizing to a stationary window," as used herein, refers to a normalization process that uses one or more selected portions to compare a test dataset with a reference dataset. In some embodiments, the selected portions are used to generate a profile. A stationary window generally includes a predetermined set of portions that do not change during manipulation and / or analysis. The terms "normalizing to a moving window" and "normalizing to a sliding window," as used herein, refer to normalization performed to portions localized to a genomic region of the selected test portion (e.g., adjacent portions or sections immediately surrounding a gene, etc.), where one or more selected test portions are normalized to portions immediately surrounding the selected test portion. In certain embodiments, the selected portions are used to generate a profile. Sliding window or moving window normalization often involves iteratively moving or sliding toward adjacent test portions and normalizing the newly selected test portion to portions immediately surrounding or adjacent to the newly selected test portion, where the adjacent windows have one or more portions in common. In certain embodiments, multiple selected test portions and / or chromosomes can be analyzed using sliding window processing.

[0203] In some embodiments, one or more values ​​can be generated by normalizing over a sliding or moving window, where each value represents the result of normalization over a different set of reference portions selected from different regions (e.g., chromosomes) of the genome. In certain embodiments, the generated value or values ​​are cumulative sums (e.g., a numerical estimate of the integral of the normalized count profile over the selected portion, domain (e.g., part of a chromosome), or chromosome). The values ​​generated by the sliding or moving window process can be used to generate profiles and facilitate arriving at outcomes. In some embodiments, the cumulative sum of one or more portions can be displayed as a function of genomic position. Sometimes, moving or sliding window analysis is used to analyze a genome for the presence or absence of microdeletions and / or microinsertions. In certain embodiments, displaying the cumulative sum of one or more portions is used to identify the presence or absence of regions of genetic variation (e.g., microdeletions, microduplications). In some embodiments, moving or sliding window analysis is used to identify genomic regions containing microdeletions, and in certain embodiments, moving or sliding window analysis is used to identify genomic regions containing microduplications.

[0204] Specific examples of normalization processes that can be utilized are described in more detail below, such as the LOESS, ChAI, and principal component normalization methods.

[0205] In some embodiments, the processing step includes weighting. The terms "weighted," "weighting," or "weighting function," or grammatical derivatives or equivalents thereof, as used herein, refer to a mathematical manipulation of part or all of a dataset that may be utilized to vary the influence of a particular dataset feature or variable relative to other dataset features or variables (e.g., to increase or decrease the significance and / or contribution of data contained in one or more portions or portions of a reference genome based on the quality or usefulness of the data in the selected portion or portions of the reference genome). In some embodiments, a weighting function may be used to increase the influence of data with a relatively small measurement variance and / or decrease the influence of data with a relatively large measurement variance. For example, portions of a reference genome with underrepresented or low-quality sequence data may be "weighted down" to minimize their influence on the dataset, while selected portions of a reference genome may be "weighted up" to increase their influence on the dataset. A non-limiting example of a weighting function is [1 / (standard deviation) 2 ]. The weighting step is sometimes performed in a manner substantially similar to the normalization step. In some embodiments, the dataset is divided by a predetermined variable (e.g., a weighting variable). Often, the predetermined variable (e.g., a minimization objective function, Phi) is selected to weight different parts of the dataset differently (e.g., to increase the influence of certain types of data while decreasing the influence of other types of data).

[0206] In certain embodiments, the processing step may include one or more mathematical and / or statistical operations. Any suitable mathematical and / or statistical operations, alone or in combination, may be used to analyze and / or manipulate the datasets described herein. Any suitable number of mathematical and / or statistical operations may be used. In some embodiments, the dataset may be mathematically and / or statistically manipulated one or more times, five or more times, ten or more times, or twenty or more times. Non-limiting examples of mathematical and statistical operations that may be used include addition, subtraction, multiplication, division, algebraic functions, least squares estimators, curve fitting, differential equations, rational polynomials, double polynomials, orthogonal polynomials, z-scores, p-values, chi values, phi values, analyzing peak levels, determining the location of peak edges, calculating peak area ratios, analyzing chromosome-level medians, calculating the mean absolute deviation, sum of squared residuals, mean, standard deviation, standard error, etc., or combinations thereof. Mathematical and / or statistical manipulations can be performed on all or a portion of the sequence read data or their processed products. Non-limiting examples of variables or features of a dataset that can be statistically manipulated include raw counts, filtered counts, normalized counts, peak height, peak width, peak area, peak edge, lateral tolerance, P-value, median level, mean level, distribution of counts within a genomic region, relative representation of nucleic acid species, etc., or combinations thereof.

[0207] In some embodiments, the processing step can include the use of one or more statistical algorithms. Any suitable statistical algorithm can be used alone or in combination to analyze and / or manipulate the datasets described herein. Any suitable number of statistical algorithms can be used. In some embodiments, one or more, five or more, ten or more, or twenty or more statistical algorithms can be used to analyze the dataset. Non-limiting examples of statistical algorithms suitable for use with the methods described herein include decision trees, alternative hypothesis, multiple comparisons, omnibus tests, Behrens-Fisher tests, bootstrapping, Fisher's method for combining independence tests of significance, null hypothesis, type I error, type II error, exact test, one-sample Z-test, two-sample Z-test, one-sample t-test, paired t-test, two-sample pooled t-test with equal variances, two-sample unpooled t-test with unequal variances, one proportion z-test, two proportion z-test pooled, two proportion z-test unpooled, one-sample chi-square test, two-sample F-test for homogeneity of variances, confidence interval, credible interval, significance, meta-analysis, simple regression, robust linear regression, etc., or a combination of the above. Non-limiting examples of variables or features of a dataset that can be analyzed using statistical algorithms include raw counts, filtered counts, normalized counts, peak height, peak width, peak edge, lateral tolerance, P-value, median level, mean level, distribution of counts within a genomic region, relative representation of nucleic acid species, and the like, or combinations thereof.

[0208] In certain embodiments, a dataset can be analyzed by utilizing multiple (e.g., two or more) statistical algorithms (e.g., least squares regression, principal component analysis, linear discriminant analysis, quadratic discriminant analysis, bagging, neural networks, support vector machine models, random forests, classification tree models, K-nearest neighbors, logistic regression, and / or loss smoothing) and / or mathematical and / or statistical operations (e.g., referred to herein as operations). In some embodiments, the use of multiple operations can generate an N-dimensional space that can be used to generate an outcome. In certain embodiments, analyzing a dataset by utilizing multiple operations can reduce the complexity and / or dimensionality of the dataset. For example, multiple operations can be used on a reference dataset to generate an N-dimensional space (e.g., a probability plot) that can be used to display the presence or absence of genetic variations depending on the genetic status of the reference sample (e.g., positive or negative for selected genetic variations). Analysis of test samples using a substantially similar set of operations can be used to generate N-dimensional points for each of the test samples. The complexity and / or dimensionality of the test data set is sometimes simplified to a single value or N-dimensional point that can be easily compared with the N-dimensional space generated from the reference data.The test sample data that belongs to the N-dimensional space where the reference data exists shows a genetic status that is substantially similar to the genetic status of the reference data.The test sample data that exists outside the N-dimensional space where the reference data exists shows a genetic status that is substantially dissimilar to the genetic status of the reference data.In some embodiments, the reference is euploid or otherwise does not have genetic variations or medical conditions.

[0209] In some embodiments, after the datasets have been counted and optionally filtered and normalized, these processed datasets can be further manipulated by one or more filtering and / or normalizing procedures. In certain embodiments, datasets that have been further manipulated by one or more filtering and / or normalizing procedures can be used to generate profiles. In some embodiments, sometimes the complexity and / or dimensionality of a dataset can be reduced by one or more filtering and / or normalizing procedures. An outcome can be provided based on the dataset of reduced complexity and / or dimensionality.

[0210] In some embodiments, portions can be filtered according to an error measure (e.g., standard deviation, standard error, calculated variance, p-value, mean absolute error (MAE), mean absolute deviation, and / or mean absolute deviation (MAD)). In certain embodiments, the error measure refers to the variability of the counts. In some embodiments, portions are filtered according to the variability of the counts. In certain embodiments, the variability of the counts is a measure of error determined for counts mapped to a portion (i.e., portion) of the reference genome for multiple samples (e.g., multiple samples obtained from multiple subjects, e.g., 50 or more, 100 or more, 500 or more, 1000 or more, 5000 or more, or 10,000 or more subjects). In some embodiments, portions with variability of counts above a predetermined upper range are filtered (e.g., eliminated from consideration). In some embodiments, the predetermined upper range is a MAD value equal to or greater than about 50, equal to or greater than about 52, equal to or greater than about 54, equal to or greater than about 56, equal to or greater than about 58, equal to or greater than about 60, equal to or greater than about 62, equal to or greater than about 64, equal to or greater than about 66, equal to or greater than about 68, equal to or greater than about 70, equal to or greater than about 72, equal to or greater than about 74, or equal to or greater than about 76. In some embodiments, the portion of the count variability that falls below the predetermined lower range is filtered (e.g., eliminated from consideration). In some embodiments, the predetermined lower range is a MAD value equal to or less than about 40, equal to or less than about 35, equal to or less than about 30, equal to or less than about 25, equal to or less than about 20, equal to or less than about 15, equal to or less than about 10, equal to or less than about 5, equal to or less than about 1, or equal to or less than about 0. In some embodiments, the portion of the counts that has variability outside the predetermined range is filtered (e.g., eliminated from consideration).In some embodiments, the predetermined range is a MAD value from greater than zero to less than about 76, less than about 74, less than about 73, less than about 72, less than about 71, less than about 70, less than about 69, less than about 68, less than about 67, less than about 66, less than about 65, less than about 64, less than about 62, less than about 60, less than about 58, less than about 56, less than about 54, less than about 52, or less than about 50. In some embodiments, the predetermined range is a MAD value from greater than zero to less than about 67.7. In some embodiments, a variability portion of the counts within the predetermined range is selected (e.g., for use in determining the presence or absence of a genetic variation).

[0211] In some embodiments, the variability of the counts of the portions exhibits a distribution (e.g., a normal distribution). In some embodiments, the portions are selected within quantiles of the distribution. In some embodiments, the variability of the counts of the portions is equal to or less than about 99.9%, equal to or less than about 99.8%, equal to or less than about 99.7%, equal to or less than about 99.6%, equal to or less than about 99.5%, equal to or less than about 99.4%, equal to or less than about 99.3%, equal to or less than about 99.2%, equal to or less than about 99.1%, equal to or less than about 99.0%, equal to or less than about 98.9%, equal to or less than about 98.8%, equal to or less than about 98.7%, equal to or less than about 98.6%, equal to or less than about 98.5%, or equal to or less than about 99.6% of the distribution. The following quantiles are selected: 8.4% or less, about 98.3% or less, about 98.2% or less, about 98.1% or less, about 98.0% or less, about 97% or less, about 96% or less, about 95% or less, about 94% or less, about 93% or less, about 92% or less, about 91% or less, about 90%, about 85%, about 80%, or about 75%. In some embodiments, the 99% quantile of the variability distribution of counts is selected. In some embodiments, the 99% quantile of the MAD>0 and the MAD<67.725 is selected, thereby identifying a series of stable portions of the reference genome.

[0212] Portions can be filtered based on an error measure, or a portion of an error measure. In certain embodiments, an error measure including the absolute value of deviation, such as an R factor, can be used to remove or weight portions. In some embodiments, the R factor is defined as the sum of the absolute deviations of the count values ​​predicted from the actual measurements divided by the count values ​​predicted from the actual measurements (e.g., Formula II herein). An error measure including the absolute value of deviation can be used, although any suitable error measure can be used instead. In certain embodiments, an error measure that does not include the absolute value of deviation, such as a square-based variability, can be used. In some embodiments, portions are filtered or weighted according to a mappability measure (e.g., mappability score). Sometimes, portions are filtered or weighted according to a relatively low number of sequence reads mapped to the portion (e.g., 0, 1, 2, 3, 4, or 5 reads mapped to the portion). Portions can be filtered or weighted according to the type of analysis being performed. For example, for analysis of aneuploidies of chromosomes 13, 18 and / or 21, the sex chromosomes can be filtered out and only autosomes or a subset of autosomes can be analyzed.

[0213] In certain embodiments, the following filtering process can be used: The same series of segments (e.g., a portion of the reference genome) within a given chromosome (e.g., chromosome 21) is selected, and the number of reads is compared between affected and non-affected samples. Gaps are used to link trisomy 21 samples with euploid samples, including a series of segments covering most of chromosome 21. These series of segments are the same between euploid and T21 samples. Because the segments can be defined, the distinction between a series and a single section is less important. The same genomic region is compared in different patients. This process can be used for trisomy analysis, for example, T13 or T18, in addition to or instead of T21.

[0214] In some embodiments, after the data sets have been counted and optionally filtered and normalized, these processed data sets can be manipulated by weighting. In certain embodiments, one or more portions can be selected and weighted to reduce the influence of data contained in the selected portions (e.g., noisy data, uninformative data), and in some embodiments, one or more portions can be selected and weighted to enhance or increase the influence of data contained in the selected portions (e.g., data with small measured variances). In some embodiments, a single weighting function is utilized to weight the data sets, reducing the influence of data with large variances and increasing the influence of data with small variances. Sometimes, a weighting function is used to reduce the influence of data with large variances and increase the influence of data with small variances (e.g., [1 / (standard deviation) 2 ]). In some embodiments, weighting is used to generate a profile plot of the processed data that is further manipulated to facilitate classification and / or providing an outcome. An outcome can be provided based on the weighted data profile plot.

[0215] Filtering or weighting of part can be carried out at one or more suitable points in analysis.For example, before or after mapping sequence readings to the part of reference genome, part can be filtered or weighted.In some embodiments, before or after determining the experimental bias of each genome part, part can be filtered or weighted.In certain embodiments, before or after calculating the level of genome division, part can be filtered or weighted.

[0216] In some embodiments, after the datasets have been counted, optionally filtered, normalized, and optionally weighted, these processed datasets can be manipulated by one or more mathematical and / or statistical (e.g., statistical functions or algorithms) operations. In certain embodiments, the processed datasets can be further manipulated by calculating Z-scores for one or more selected portions, chromosomes, or portions of chromosomes. In some embodiments, the processed datasets can be further manipulated by calculating P-values. In certain embodiments, the mathematical and / or statistical operations include one or more assumptions regarding ploidy and / or fetal fraction. In some embodiments, a profile plot of the processed data further manipulated by one or more statistical and / or mathematical operations is generated to facilitate classification and / or providing an outcome. An outcome can be provided based on the plot of the profile of the statistically and / or mathematically manipulated data. Outcomes provided based on the plot of the profile of the statistically and / or mathematically manipulated data often include one or more assumptions regarding ploidy and / or fetal fraction.

[0217] In certain embodiments, after the dataset is counted and optionally filtered and normalized, multiple operations are performed on the processed dataset to generate an N-dimensional space and / or N-dimensional points, and an outcome can be provided based on a plot of the profile of the dataset analyzed in the N dimensions.

[0218] In some embodiments, the dataset is processed using one or more of peak-level analysis, peak width analysis, peak edge location analysis, peak lateral tolerance, etc., derivatives thereof, or combinations thereof, as part of or after processing and / or manipulation of the dataset. In some embodiments, a profile plot of the processed data using one or more of peak-level analysis, peak width analysis, peak edge location analysis, peak lateral tolerance, etc., derivatives thereof, or combinations thereof is generated to facilitate classification and / or providing an outcome. An outcome can be provided based on a profile plot of the processed data using one or more of peak-level analysis, peak width analysis, peak edge location analysis, peak lateral tolerance, etc., derivatives thereof, or combinations thereof.

[0219] In some embodiments, one or more reference samples that do not substantially contain the genetic variation in questio...

Claims

1. 1. A processor and / or computer implemented method for partitioning a reference genome, or part thereof, into multiple portions, comprising: a) generating a guanine and cytosine (GC) profile for a reference genome, or part thereof; b) applying a segmentation process to the GC profile generated in (a) based on GC content, thereby providing individual segments, wherein the segmentation process includes determining GC content levels for candidate wavelets, the candidate wavelets being determined using a wavelet decomposition generation process, the wavelet decomposition generation process being based on chromosome length (L chr ) and the predetermined minimum length (L min ) is performed on the GC profile according to a decomposition level C based on L chr is the length of the reference genome or a part thereof, or the length of the GC profile for the reference genome or a part thereof, and is calculated from the average number of counts per bin T A , the length of the genome L Genome , and the total number of counts T C , as follows: L min =T A ×L Genome / T C , and C is calculated as follows: C=log2(L chr / L min ), C=log2(L chr / L min ) + 1 and C = log2(L chr / L min )-1; and c) partitioning the reference genome, or part thereof, into a plurality of parts according to the individual segments provided in (b), thereby generating a GC-partitioned reference genome, or part thereof; A method comprising:

2. 1. A processor and / or computer implemented method for identifying the presence or absence of a genetic variation in a test sample, said method comprising: a) generating a guanine and cytosine (GC) profile for a reference genome, or part thereof; b) applying a segmentation process to the GC profile generated in (a) based on GC content, thereby providing individual segments, wherein the segmentation process includes determining GC content levels for candidate wavelets, the candidate wavelets being determined using a wavelet decomposition generation process, the wavelet decomposition generation process being based on chromosome length (L chr ) and the predetermined minimum length (L min ) is performed on the GC profile according to a decomposition level C based on L chr is the length of the reference genome or a part thereof, or the length of the GC profile for the reference genome or a part thereof, and is calculated from the average number of counts per bin T A , the length of the genome L Genome , and the total number of counts T C , as follows: L min =T A ×L Genome / T C , and C is calculated as follows: C=log2(L chr / L min ), C=log2(L chr / L min ) + 1 and C = log2(L chr / L min )-1; and c) partitioning the reference genome, or part thereof, into a plurality of parts according to the individual segments provided in (b), thereby generating a GC-partitioned reference genome, or part thereof; d) sequencing nucleic acid from said test sample by a nucleotide sequencing process to generate nucleotide sequence reads, wherein said nucleic acid is circulating cell-free nucleic acid from a pregnant female carrying a fetus; e) mapping nucleotide sequence reads from the test sample to the GC-partitioned reference genome, or portions thereof, thereby generating mapped nucleotide sequence reads; f) normalizing the counts of the mapped nucleotide sequence reads, thereby generating normalized counts; g) determining the presence or absence of said genetic variation for said test sample according to said normalized counts; A method comprising:

3. 3. The method of claim 1 or 2, further comprising partitioning chromosomes, or segments of chromosomes, from the reference genome, thereby generating GC-partitioned chromosomes or GC-partitioned chromosome segments.

4. The method of claim 1 or 2, wherein the GC profile in (a) comprises a GC content level determined for each 1 kb of nucleotide sequence in the reference genome.

5. 5. The method of claim 4, wherein the segmentation process in (b) is performed on the GC content level determined for each 1 kb of nucleotide sequence.

6. The method of claim 5, wherein 1 kb of nucleotide sequences with similar GC content levels are combined into the individual segments.

7. The method of claim 1 or 2, wherein the segmentation process in (b) produces a decomposed rendering that includes the individual segments.

8. 3. The method of claim 1, further comprising determining GC content for the individual segments in (b).

9. 3. The method of claim 2, wherein the normalizing step comprises LOESS normalization for guanine and cytosine (GC) bias (GC-LOESS normalization).

10. 3. The method of claim 2, wherein the normalizing step comprises adjusting the counts of the sequence reads according to a median count.

11. 11. The method of claim 10, wherein the counts of the sequence reads are adjusted according to the median fractional counts.

12. The method of claim 2 , wherein the normalizing step comprises principal component normalization.

13. 3. The method of claim 2, wherein the normalizing step comprises GC-LOESS normalization, followed by normalization according to median fractional counts, followed by principal component normalization.

14. The method of claim 2 , further comprising determining a chromosome structure according to the normalized counts.

15. 15. The method of claim 14, wherein the normalized count represents a chromosome dose for the test sample.

16. 16. The method of claim 15, wherein determining the presence or absence of the genetic variation is according to the chromosomal dosage.

17. 17. The method of claim 16, wherein determining the presence or absence of the genetic variation in the test sample comprises identifying the presence or absence of one copy of a chromosome, two copies of a chromosome, three copies of a chromosome, four copies of a chromosome, five copies of a chromosome, a deletion of one or more segments of a chromosome, or an insertion of one or more segments of a chromosome.

18. 1. A processor and / or computer implemented method for identifying the presence or absence of a genetic variation in a test sample, said method comprising: a) generating a guanine and cytosine (GC) profile for a reference genome, or part thereof; b) applying a segmentation process to the GC profile generated in (a) based on GC content, thereby providing individual segments, wherein the segmentation process includes determining GC content levels for candidate wavelets, the candidate wavelets being determined using a wavelet decomposition generation process, the wavelet decomposition generation process being based on chromosome length (L chr ) and the predetermined minimum length (L min ) is performed on the GC profile according to a decomposition level C based on L chr is the length of the reference genome or a part thereof, or the length of the GC profile for the reference genome or a part thereof, and is calculated from the average number of counts per bin T A , the length of the genome L Genome , and the total number of counts T C , as follows: L min =T A ×L Genome / T C , and C is calculated as follows: C=log2(L chr / L min ), C=log2(L chr / L min ) + 1 and C = log2(L chr / L min )-1; and c) partitioning the reference genome, or part thereof, into a plurality of parts according to the individual segments provided in (b), thereby generating a GC-partitioned reference genome, or part thereof; d) sequencing nucleic acid from said test sample by a nucleotide sequencing process to generate nucleotide sequence reads, wherein said nucleic acid is circulating cell-free nucleic acid from a pregnant female carrying a fetus; e) mapping nucleotide sequence reads from the test sample to the GC-partitioned reference genome, or portions thereof, thereby generating mapped nucleotide sequence reads; f) generating raw counts of the mapped nucleotide sequence reads; g) determining the presence or absence of said genetic variation for said test sample according to said raw counts; A method comprising: