Genetic analysis method and system

Microhaplotypes based on single base pair substitutions address the challenge of detecting and quantifying DNA contamination in complex samples, enhancing forensic and disease analysis by identifying and distinguishing multiple DNA sources without relying on minor allele frequency.

JP7722929B2Active Publication Date: 2025-08-13PERSONAL GENOME DIAGNOSTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2021562794
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-04-22
Filing Date
2020-04-21
Publication Date
2025-08-13
Estimated Expiration
2040-04-21

AI Technical Summary

Technical Problem

Existing genetic analysis methods struggle to accurately detect DNA contamination and quantify its amount due to overlapping minor allele frequencies (MAFs) in complex DNA mixtures, particularly in samples with low yield or low-quality DNA, leading to missed contamination levels.

Method used

Utilizing microhaplotypes associated with single base pair substitutions (SBSs) to identify and quantify the frequency of haplotypes within SNP sets, allowing for the detection of DNA contamination and genetic markers in forensic and disease analysis, without relying on minor allele frequency (MAF).

Benefits of technology

Enhances the accuracy of DNA contamination detection and disease analysis by leveraging microhaplotypes, providing a method to distinguish between multiple DNA sources and identify genetic markers without prior knowledge of the sample components, improving forensic analysis and disease detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007722929000047
    Figure 0007722929000047
  • Figure 0007722929000048
    Figure 0007722929000048
  • Figure 0007722929000049
    Figure 0007722929000049
Patent Text Reader

Abstract

The present disclosure provides computational methods for genetic analysis and systems for implementing such analysis. The present disclosure provides methods of genetic analysis that utilize microhaplotypes associated with SNPs that are single base pair substitutions (SBSs) rather than insertion or deletion SNPs. Such microhaplotype analysis is useful, among other things, in forensic genetic applications, sample contamination analysis, and disease analysis. TIFF2022530393000048.tif86154
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority under 35 U.S.C. §119(e) of U.S. Patent Application No. 62 / 837,034, filed April 22, 2019, the entire contents of which are incorporated herein by reference.

[0002] FIELD OF THE INVENTION The present invention relates generally to genetic analysis, and more particularly to methods and systems for performing microhaplotype analysis to determine genetic identity in complex DNA mixtures. [Background technology]

[0003] Background information Sequence variation in the human genome is a cornerstone of human identification and forensic applications. Genetic fingerprinting is a forensic technique used to identify individuals by the characteristics of their genetic information (e.g., RNA, DNA). A genetic fingerprint is a small set of one or more nucleic acid variations that are likely to be different in all unrelated individuals and, thus, are unique to an individual, similar to a fingerprint.

[0004] Sequence variation is useful in genetic analysis for many applications, such as detection of contamination in biological samples, forensic analysis, disease detection, and population genetics. Single nucleotide polymorphisms (SNPs) have long been used in genetic analysis for such applications.

[0005] DNA contamination in biological samples is a widespread problem. Contamination can occur at almost every stage of sample collection / processing. For example, slides can be contaminated during cutting, liquids can be inadvertently transferred between tubes, libraries can be mixed, and sample barcodes can be impure or have low-quality sequences. Contamination is likely to be more pronounced in samples with low yield and / or low-quality DNA.

[0006] SNPCheck™ is a batch testing tool for the presence of SNPs and can be used to confirm the presence of DNA contamination in samples. For "well-behaved" DNA, such as normal tissue or cfDNA, SNPCheck™ can provide reasonable results because the minor allele frequency (MAF) is almost always around 0 or 0.5. However, extremely high contamination levels can be missed because the MAF can be very high, approaching 0.5. Tumor DNA is not "well-behaved," with extreme copy number variation resulting in a MAF ranging from 0.02 to 0.98. This means that the MAFs of contamination and actual variants can overlap significantly.

[0007] To be able to detect DNA contamination and to accurately quantify the amount of contamination, a detection method that is independent or largely independent of MAF is required. Summary of the Invention

[0008] The present disclosure provides methods of genetic analysis that utilize microhaplotypes associated with SNPs that are single base pair substitutions (SBSs) rather than insertion or deletion SNPs. Such microhaplotype analysis is particularly useful in forensic genetic applications, sample contamination analysis, and disease analysis.

[0009] In one embodiment, the present disclosure provides a method of genetic analysis comprising: a) identifying a SNP set having at least three microhaplotypes in a sample; and b) quantifying the frequency of haplotypes within the SNP set having more than two microhaplotypes.

[0010] In another embodiment, the present disclosure provides a method of genetic analysis comprising: a) identifying a SNP set having at least three microhaplotypes in a sample; and b) quantifying the frequency of haplotypes within the SNP set having more than two microhaplotypes to determine the presence or absence of DNA contamination in the sample.

[0011] In yet another embodiment, the present disclosure provides a method of genetic analysis comprising: a) identifying a SNP set having at least three microhaplotypes in a sample; and b) quantifying the frequency of haplotypes within the SNP set having more than two microhaplotypes to determine the presence or absence of a genetic marker indicative of a disease or disorder.

[0012] In yet another embodiment, the present disclosure provides a method for identifying microhaplotypes in a genome, the method comprising: a) identifying a region of interest in the genome; b) detecting SBSs within the region of interest, thereby generating a plurality of sets of sequence variants; and c) To identify candidate microhaplotypes, Each variant set was analyzed for linkage disequilibrium vinegar and d) identifying candidate microhaplotypes.

[0013] In another embodiment, the present disclosure provides a method for detecting a SNP set having at least three microhaplotypes derived from multiple subjects present in a sample. The method includes: a) identifying microhaplotypes in the genome of the sample; b) determining the number of SNP sets in the sample having at least three microhaplotypes; and c) quantifying the frequency of haplotypes within the SNP set with more than two microhaplotypes to determine the presence of DNA derived from the multiple subjects in the sample, thereby detecting DNA derived from the multiple subjects in the sample. In one embodiment, the identifying includes: i) identifying a region of interest in the genome; ii) detecting SBSs within the region of interest, thereby generating multiple sequence variant sets; and iii) analyzing each variant set for LD to identify microhaplotypes.

[0014] In an embodiment, the present disclosure provides a method for detecting a SNP set having at least two microhaplotypes derived from multiple subjects present in a sample, the method comprising: a) determining the presence or absence of a SNP set having more than two microhaplotypes in the sample, wherein the SNP set comprises multiple single base pair substitutions and corresponds to a genomic region set forth in Tables 5, 6, and 7; and b) quantifying the frequency of haplotypes within the SNP set to determine the presence of DNA derived from the multiple subjects in the sample, thereby detecting a SNP set having more than two microhaplotypes derived from the multiple subjects in the sample.

[0015] In one embodiment, the disclosure provides an oligonucleotide panel comprising oligonucleotides for amplifying or hybrid capture regions of a genome corresponding to one or more genomic regions listed in Tables 5, 6, and 7.

[0016] In another embodiment, the present disclosure provides a method of genetic analysis comprising: a) amplifying a region of a genome present in a sample, wherein the region corresponds to a genomic region set forth in Table 5, Table 6, and Table 7, and generating an amplicon upon amplification; and b) sequencing the amplicon to determine the nucleic acid sequence of the amplicon.

[0017] In a further embodiment, the present disclosure provides a method for detecting a disease or disorder in a subject. The method includes: a) obtaining a sample from the subject; b) identifying microhaplotypes in DNA molecules present in the sample; c) determining the presence or absence of a SNP set having more than two microhaplotypes in the sample; and d) quantifying the frequency of haplotypes within the SNP set to determine the presence or absence of a genetic marker indicative of the disease or disorder, thereby detecting the disease or disorder. In one embodiment, the identifying includes: i) identifying a region of interest, where the region of interest is associated with the disease or disorder; ii) detecting SBS within the region of interest, thereby generating a plurality of sequence variant sets; and iii) analyzing each variant set for LD to identify microhaplotypes.

[0018] In an embodiment, the present disclosure provides a genetic analysis system including: a) at least one processor operatively connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in a DNA sample; and c) an analysis component executed by the at least one processor configured to: i) identify microhaplotypes in the sample based on the presence of single base pair substitutions; ii) confirm the presence of a number of SNP sets for the microhaplotypes in the DNA sample; and iii) quantify the frequency of genotypes within the SNP sets with more than two microhaplotypes in the DNA sample.

[0019] In related embodiments, the present disclosure provides a genetic analysis system configured to perform the methods of the present disclosure, the system including: a) at least one processor operably connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in a DNA sample; and c) an analysis component executed by the at least one processor, the analysis component configured to perform the methods of the present disclosure.

[0020] In yet another embodiment, the present invention provides a non-transitory computer-readable storage medium encoded with a computer program, the program including instructions that, when executed by one or more processors, cause the one or more processors to perform operations to perform the method of the present disclosure.

[0021] In yet another embodiment, the present invention provides a computing system including a memory and one or more processors coupled to the memory, the one or more processors configured to perform operations implementing the method of the present disclosure. [The present invention 1001] 1. A method for identifying microhaplotypes in a genome, comprising: a) Identifying target regions of the genome; b) detecting single base pair substitutions (SBS) within the target region to generate a plurality of sets of sequence variants; c) To identify candidate microhaplotypes, Each variant set was analyzed for linkage disequilibrium vinegar and; d) identifying candidate microhaplotypes; and The method comprising: [The present invention 1002] 1001. The method of claim 1001, further comprising detecting SBS in regions flanking said region of interest. [The present invention 1003] The method of claim 1002, wherein the flanking regions of the region of interest comprise less than about 50, less than about 100, less than about 150, less than about 180, or less than about 200 nucleotide base pairs sequenceable by a short-read sequencer. [The present invention 1004] 1002. The method of claim 1002, wherein the flanking regions of the region of interest comprise less than about 10,000 nucleotide base pairs sequenceable by a long-read sequencer. [The present invention 1005] 1001. The method of claim 10, wherein the region of interest in a) has an SBS frequency of between about 10 and 90%. [The present invention 1006] 1002. The method of claim 1002, wherein said flanking regions of said region of interest have SBS at a frequency of about 5-95%. [The present invention 1007] 1001. The method of claim 1001, further comprising calibrating a cutoff value for the candidate microhaplotype to assess contamination of the sample. [The present invention 1008] The method of claim 1006, wherein only DNA sequence reads that overlap with said candidate microhaplotype are used to calculate a threshold for contamination detection and a degree of contamination. [The present invention 1009] 1008. The method of claim 10, wherein the DNA sequences used to calibrate the threshold for contamination detection and the degree of contamination are mixed pairwise in silico, using each DNA sequence alternately as the primary sample and the contaminant. [The present invention 1010] The method of claim 1008 or 1009, wherein the number and genotype of SNP sets with one and / or two microhaplotypes are compared between different individuals to assess identity or contamination. [The present invention 1011] The method of the present invention 1007 further comprises assessing sample contamination using a cutoff value determined for the frequency of a candidate microhaplotype having a single nucleotide polymorphism (SNP) set involving at least three microhaplotypes. [The present invention 1012] The method of claim 1011, further comprising assessing sample contamination using a cutoff value determined for the frequency of candidate microhaplotypes having a SNP set with at least four or more microhaplotypes. [The present invention 1013] 1001. The method of claim 1001, wherein said candidate microhaplotypes correspond to one or more genomic regions selected from those set forth in Table 5, Table 6, or Table 7. [The present invention 1014] 1007. The method of claim 10, wherein said sample comprises DNA derived from a tumor or a liquid biopsy. [The present invention 1015] 1007. The method of claim 1007, wherein said sample comprises DNA extracted from a formalin-fixed, paraffin-embedded block, slide, or curl. [The present invention 1016] The method of claim 1014, wherein the liquid biopsy is derived from amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, chyme, endolymph, perilymph, stool, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal mucus and sputum), pericardial fluid, ascites, pleural fluid, pus, mucosal secretions, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomit, prostatic fluid, nipple aspirate, tears, sweat, buccal specimen collection, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluid. [The present invention 1017] 1015. The method of claim 1014, wherein said sample is derived from circulating tumor cells. [The present invention 1018] 1007. The method of claim 10, wherein said calibration comprises analysis of candidate microhaplotypes in a plurality of samples obtained from humans of different ethnicities. [The present invention 1019] 1001. The method of claim 1001, wherein said candidate microhaplotype comprises a set of SNPs having at least three, four or more sets of SNP sequence variants. [The present invention 1020] 1002. The method of claim 1001, wherein said region of interest is within a gene, an intron, and / or an exon, or between genes. [The present invention 1021] 1001. The method of claim 1001, wherein the region of interest is within the exome. [The present invention 1022] 1001. The method of claim 1001, further comprising isolating DNA comprising said candidate microhaplotype. [The present invention 1023] 1001. The method of claim 1001, wherein the genome is derived from a human. [The present invention 1024] The method of claim 1001, further comprising assessing sample contamination by analyzing the median, mean, or other measure of microhaplotype frequency of haplotypes within a SNP set with at least three or four microhaplotypes. [The present invention 1025] Any of the methods of the invention further comprising determining the source of contamination of the sample by identifying microhaplotypes common to or specific to the microhaplotypes of the sample and the contaminant. [The present invention 1026] The method of the present invention 1025, wherein the microhaplotype information is stored in a database and compared with newly / concurrently sequenced individuals to identify whether the DNA samples are from the same or different individuals. [The present invention 1027] The method of the present invention 1025, wherein the microhaplotype information is stored in a database and compared with newly / concurrently sequenced individuals to identify whether a particular DNA sample is contaminating other samples. [The present invention 1028] The method of claim 1026 or 1027, wherein the number and genotype of SNP sets with one and / or two microhaplotypes are compared between different individuals to assess identity or contamination. [The present invention 1029] Any of the aforementioned methods of the present invention further comprising determining the ethnicity of said sample and said contaminant. [The present invention 1030] 1001. The method of claim 1001, wherein the microhaplotype frequencies are calculated using only common genotypes found in the population used in said method. [The present invention 1031] The method of claim 1030, wherein said common genotype is present at more than 1% in 1000 Genomes™ or other database. [The present invention 1032] Use of the method of the present invention 1001 to assess the quality of a sample from a particular source, from a vendor, or from a technician preparing or sequencing the sample. [The present invention 1033] 1. A method for detecting a set of single nucleotide polymorphisms (SNPs) having at least three microhaplotypes from a plurality of subjects present in a sample, the method comprising: a) i) identifying a region of interest in the genome; ii) detecting single base pair substitutions (SBS) within the target region, thereby generating a plurality of sets of sequence variants; and iii) Analyzing each variant set for linkage disequilibrium to identify microhaplotypes identifying microhaplotypes in the genome in the sample; b) determining the number of SNP sets having at least three microhaplotypes in the sample; c) quantifying the frequency of SNP sets with more than two microhaplotypes to detect the presence of DNA from multiple subjects in the sample, thereby detecting DNA from multiple subjects in the sample; The method comprising: [The present invention 1034] The method of claim 1033, further comprising isolating DNA from said sample comprising said microhaplotype. [This invention 1035] The method of claim 1033, further comprising detecting SBS in genomic regions flanking said region of interest. [The present invention 1036] The method of claim 1035, wherein the flanking regions of the target region comprise less than about 50, less than about 100, less than about 150, less than about 180, or less than about 200 nucleotide base pairs sequenceable by a short-read sequencer. [This invention 1037] 1035. The method of claim 1035, wherein the flanking regions of the region of interest comprise less than about 10,000 nucleotide base pairs sequenceable by a long-read sequencer. [The present invention 1038] The method of the present invention 1033, wherein the target region of i) has an SBS with a genotype at a frequency of about 10 to 90%. [This invention 1039] 1035. The method of claim 1035, wherein said flanking regions of said region of interest have an SBS with a genotype at a frequency of about 5-95%. [The present invention 1040] The method of claim 1033, wherein a cutoff value for a set of SNPs involving two, three, four, or more microhaplotypes is calibrated to assess the presence of DNA from multiple subjects in the sample. [The present invention 1041] 1034. The method of claim 1033, wherein said sample comprises DNA derived from a tumor or a liquid biopsy. [The present invention 1042] The method of claim 1041, wherein the liquid biopsy is derived from amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, chyme, endolymph, perilymph, stool, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal mucus and sputum), pericardial fluid, ascites, pleural fluid, pus, mucosal secretions, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomit, prostatic fluid, nipple aspirate, tears, sweat, buccal specimen collection, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluid. [This invention 1043] 1042. The method of claim 1041, wherein said sample is derived from circulating tumor cells. [This invention 1044] The method of claim 1033, wherein a set of SNPs with more than two microhaplotypes from two or more subjects is detected. [This invention 1045] The method of claim 1033, wherein said sample comprises maternal DNA and fetal DNA. [The present invention 1046] The method of claim 1045, further comprising distinguishing said fetal DNA from said maternal DNA. [This invention 1047] The method of claim 1046, further comprising assessing the presence of DNA other than said maternal DNA and said fetal DNA. [This invention 1048] 1034. The method of claim 1033, wherein the subject is a human. [This invention 1049] 1. A method for detecting a set of single nucleotide polymorphisms (SNPs) having at least three microhaplotypes from a plurality of subjects present in a sample, the method comprising: a) determining the presence or absence of a set of SNPs having more than two microhaplotypes in a sample, wherein the set of SNPs comprises a plurality of single base pair substitutions and corresponds to a genomic region selected from the regions set forth in Table 5, Table 6, and Table 7; b) quantifying the frequency of the set of SNPs to determine the presence of DNA from multiple subjects in the sample, thereby detecting a set of SNPs having at least three microhaplotypes from multiple subjects in the sample; The method comprising: [The present invention 1050] An oligonucleotide panel comprising oligonucleotides for amplifying or hybrid capturing a region of a genome corresponding to one or more genomic regions containing an SBS set identified in any of claims 1001 to 1006. [This invention 1051] An oligonucleotide panel comprising oligonucleotides for amplifying or hybrid capture regions of the genome corresponding to one or more genomic regions selected from the regions set forth in Table 5, Table 6, and Table 7. [This invention 1052] a) amplifying a region of a genome present in a sample, said region corresponding to a genomic region selected from the regions set forth in invention 1050, Table 5, Table 6, or Table 7, and generating an amplicon by amplification; b) sequencing the amplicon to determine the nucleic acid sequence of the amplicon; A method comprising: [This invention 1053] The method of claim 1052, further comprising quantifying the number of SNP sets with more than two microhaplotypes present in said sample. [This invention 1054] The method of claim 1053, further comprising quantifying the number of SNP sets with more than three microhaplotypes present in said sample. [This invention 1055] The method of claim 1054, further comprising quantifying the number of SNP sets with more than four microhaplotypes present in said sample. [This invention 1056] 1. A method for detecting a disease or disorder in a subject, comprising: a) obtaining a sample from a subject; b) i) identifying a region of interest, said region of interest being associated with a disease or disorder; ii) detecting single base pair substitutions (SBS) within the target region, thereby generating a plurality of sets of sequence variants; and iii) Analyzing each variant set for linkage disequilibrium to identify microhaplotypes identifying microhaplotypes in DNA molecules present in the sample; c) determining the presence or absence of a set of single nucleotide polymorphisms (SNPs) having greater than two microhaplotypes in the sample; d) quantifying the frequency of the set of SNPs to determine the presence or absence of a genetic marker indicative of a disease or disorder, thereby detecting the disease or disorder; The method comprising: [This invention 1057] 1056. The method of claim 1056, wherein said disease or disorder is trisomy 13, 18, or 21. [This invention 1058] 1056. The method of claim 1056, wherein the disease or disorder is a gene copy number variation. [This invention 1059] The method of claim 1056, wherein said disease or disorder is a fetal disorder. [The present invention 1060] The method of any of claims 1056 to 1059, wherein the frequency of the third microhaplotype at the particular chromosome or chromosomal region is compared to the frequency of the third microhaplotype elsewhere in the genome. [The present invention 1061] a) at least one processor operatively connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in the DNA sample; c) i) identifying microhaplotypes in the sample based on the presence of single base pair substitutions; ii) confirming the presence of a number of SNP sets for the microhaplotype in the DNA sample; and iii) quantifying the frequency of genotypes within the set of SNPs with more than two microhaplotypes in the DNA sample; an analysis component executed by at least one processor; A genetic analysis system comprising: [The present invention 1062] 1061. The system of claim 1061, wherein said analytical component is further configured to determine the possible presence of DNA contaminants in said sample. [This invention 1063] The system of claim 1061, wherein said analysis component is further configured to determine the presence or absence of a genetic mutation. [This invention 1064] The system of claim 1063, wherein the genetic mutation is associated with a disease or disorder. [This invention 1065] The system of the present invention 1064, wherein the disease or disorder is associated with gene copy number variation. [The present invention 1066] The system of the present invention 1065, wherein the disease or disorder is trisomy 13, 18, or 21. [This invention 1067] a) at least one processor operatively connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in the DNA sample; c) an analysis component executed by said at least one processor and configured to perform (a)-(d) of the present invention 1001; A genetic analysis system comprising: [The present invention 1068] a) at least one processor operatively connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in the DNA sample; c) an analysis component configured to be executed by the at least one processor and to perform (a) to (c) of the present invention 1033; A genetic analysis system comprising: [The present invention 1069] a) at least one processor operatively connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in the DNA sample; c) an analysis component configured to be executed by said at least one processor and to perform the method of the invention 1049 or 1052; A genetic analysis system comprising: [The present invention 1070] a) at least one processor operatively connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in the DNA sample; c) an analysis component configured to be executed by said at least one processor and to perform (b)-(d) of the present invention 1056; A genetic analysis system comprising: [This invention 1071] a) identifying a set of single nucleotide polymorphisms (SNPs) having at least three microhaplotypes in a sample; b) quantifying the frequency of haplotypes within a set of SNPs with more than two microhaplotypes to determine the presence or absence of DNA contamination in the sample; A method comprising: [This invention 1072] The method of claim 1071, further comprising quantifying the frequency of haplotypes within a SNP set having at least three or four microhaplotypes in said sample to determine the amount of DNA contamination in said sample. [This invention 1073] 1072. The method of claim 1071, wherein said sample comprises DNA derived from a tumor or a liquid biopsy. [This invention 1074] 1073. The method of claim 1073, wherein said liquid biopsy is derived from amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, chyme, endolymph, perilymph, stool, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal mucus and sputum), pericardial fluid, ascites, pleural fluid, pus, mucosal secretions, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomit, prostatic fluid, nipple aspirate, tears, sweat, buccal specimen collection, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluid. [This invention 1075] 1072. The method of claim 1071, wherein said sample is derived from circulating tumor cells. [This invention 1076] 1071. The method of claim 1071, wherein said set of SNPs comprises sequence variants having single base pair substitutions. [This invention 1077] a) identifying a set of single nucleotide polymorphisms (SNPs) having at least three microhaplotypes in a sample; b) quantifying the frequency of haplotypes within a set of SNPs with more than two microhaplotypes to determine the presence or absence of a genetic marker indicative of a disease or disorder; A method comprising: [This invention 1078] The method of claim 1077, further comprising quantifying the frequency of haplotypes within a set of SNPs having at least three or four microhaplotypes in said sample. [This invention 1079] 1077. The method of claim 1077, wherein the disease or disorder is a gene copy number variation. [The present invention 1080] 1079. The method of claim 1079, wherein said disease or disorder is trisomy 13, 18, or 21. [This invention 1081] The method of claim 1077, wherein said disease or disorder is a fetal disorder. [This invention 1082] 1077-1081, wherein the method of any of claims 1077-1081 increases the number of SNP sets on a particular chromosome, thereby enhancing the identification of trisomies. [This invention 1083] 1082. The method of claim 1082, wherein said specific chromosome is one or more of chromosomes 13, 18, and / or 21. [This invention 1084] 1084. Any of the methods of claims 1077 to 1083, which is carried out earlier in the woman's pregnancy compared to the use of conventional methods. [This invention 1085] The method of any of claims 1077 to 1084, wherein the method has improved specificity due to reduced sensitivity to the effects of errors due to maternal copy number. [The present invention 1086] a) identifying a set of single nucleotide polymorphisms (SNPs) having at least three microhaplotypes in a sample; b) quantifying the frequency of haplotypes within a set of SNPs with more than two microhaplotypes to determine the proportion of fetal DNA in maternal DNA sources; A method comprising: [This invention 1087] 1086. The method of claim 1086, wherein said source of maternal DNA is derived from a biological fluid. [This invention 1088] The method of claim 1086, wherein the source of maternal DNA is derived from amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, chyme, endolymph, perilymph, stool, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal mucus and sputum), pericardial fluid, ascites, pleural fluid, pus, mucosal secretions, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomit, prostatic fluid, nipple aspirate, tears, sweat, buccal specimen collection, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluid. [This invention 1089] A non-transitory computer-readable storage medium encoded with a computer program, the program including instructions that, when executed by one or more processors, cause the one or more processors to perform operations that implement any of the methods of the present invention 1001-1031, 1033-1049, 1052-1060, or 1077-1088. [The present invention 1090] A computing system comprising: a memory; and one or more processors coupled to the memory, wherein the one or more processors are configured to perform operations that implement any of the methods of inventions 1001-1031, 1033-1049, 1052-1060, or 1077-1088. [Brief explanation of the drawings]

[0022] [Figure 1] FIG. 1 is a graph illustrating data generated using the method of the present disclosure in one embodiment of the present invention. [Figure 2] FIG. 2 is a graph illustrating data generated using the method of the present disclosure in one embodiment of the present invention. [Figure 3] FIG. 3 is an image showing microhaplotype frequencies in the presence of contamination in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0023] Detailed Description of the Invention The present invention is based on innovative methods and systems for genetic analysis of microhaplotypes. Before describing the compositions and methods of the present invention, it is to be understood that the present invention is not limited to the particular methods and experimental conditions described, as such compositions, methods, and conditions may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, as the scope of the present invention will be limited only by the appended claims.

[0024] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, reference to "the method" includes one or more methods, and / or steps, of the type described herein that would become apparent to those skilled in the art upon reading this disclosure.

[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred methods and materials are described below.

[0026] The present disclosure provides innovative methods and systems for genetic analysis using microhaplotypes. The methods utilize SBS SNPs, and in some embodiments, SBS variations in low-error genomic regions. This can improve accuracy in DNA contamination detection, disease detection, and forensic analysis. The methods disclosed herein use SBS and not STR or insertion / deletion SNPs, because the latter have unacceptably high error rates that impact the detection of low levels of contamination in samples. All of the disclosed methods focus on SNP variants that are close genetic distances from each other, ideally within a single sequence read. With long-read technology, the distance can be even greater, as long as the SNP variants are located within a single read. While even greater distances can be used, the use of paired reads results in higher error rates, and the greater the distance between variants, the lower the coverage. Furthermore, certain methods of the present disclosure advantageously utilize a two-stage analysis: first detecting contamination and then quantifying it. Detection of DNA contamination through the methods disclosed herein relies on the number of microhaplotypes and / or the frequency of tertiary / quaternary haplotypes for each SNP set, and not on the MAF of individual SNPs.

[0027] Previous studies have illustrated the utility of markers based on multiple, tightly linked SNPs in anthropology due to their ability to provide plausible explanations for population relationships and recent patterns of human diversity. In addition, multi-allelic SNPs are increasingly being used as markers for addressing forensic problems, such as family / ancestry cohorts, ancestry inference, and individual identification. The Kidd laboratory has proposed a new type of genetic marker called microhaplotypes (e.g., "microhaps" or MH) to complement current DNA typing tools for forensic and population genetics. These are short segments of DNA (<300 nucleotides, hence the "micro") characterized by the presence of two or more tightly linked SNPs representing three or more allelic combinations (i.e., "haplotypes") within a population. The short distance between SNPs means that the recombination rate between them is extremely low. The level of heterozygosity in a microhaplotype depends on a variety of factors, including the historical accumulation of allelic variants at different positions within the region of interest, the occurrence of rare crossover events, the expression of uncontrolled genetic drift, and / or selection. Because microhaplotypes are multi-SNP haplotypes, they can provide more information per locus than a single SNP marker.

[0028] Furthermore, when variants are close to each other on the genome, they tend to be correlated. A set of different SNPs on a single chromosomal allele is called a haplotype (a set of linked SNP alleles that tend to always be expressed together (i.e., statistically associated)). Because each individual has two copies of their genome, each person has two haplotypes in the autosomal region. These haplotypes can be different (heterozygous) or identical (homozygous). As mentioned above, microhaplotypes are short haplotypes of approximately 300 nucleotides or less, or even longer distances in the case of long reads. For the purposes of the methods described herein, microhaplotypes are short enough that variants are on the same sequencing read, allowing them to be phased without ambiguity. Most microhaplotypes are not very useful in genetic analysis, because only two, and only two, microhaplotypes have been found in a given population. However, the methods of the present invention allow for the identification of microhaplotypes that can provide statistically useful information, such as microhaplotypes in which three, four, five, or more different haplotypes are found among different individuals (but no more than two haplotypes are found in any one individual).

[0029] As used herein, a "SNP" is a single nucleotide substitution in which one base (e.g., cytosine, thymine, uracil, adenine, or guanine) is substituted for another base at a particular location, i.e., locus, in the genome, and which substitution is present to an appreciable extent in a population (e.g., greater than 1% of the population).

[0030] In certain embodiments, the methods of the present disclosure relate to determining and quantifying the presence of DNA contamination in a DNA sample.

[0031] In related embodiments, the methods of the present disclosure relate to determining whether a sample contains a complex mixture of DNA from multiple individuals, which may be related or unrelated individuals, as well as mothers and offspring.

[0032] Traditional forensic analysis uniquely identifies individual DNA samples through extraction of short tandem repeats (STRs) and / or determination of mitochondrial DNA (mtDNA) sequences. Capillary electrophoresis is often used to quantify STR length and mtDNA sequence. This methodology has proven accurate for individual profile identification.

[0033] Importantly for the methods according to the present disclosure, the ability of these methods to deconvolute complex DNA mixtures into component profiles does not require any prior knowledge of the components.For example, the methods described herein are effective for deconvoluting complex DNA mixtures into component profiles, and do not require knowledge of the genetic markers or DNA sequences belonging to any individual or component that contributes to any one of the complex DNA mixtures.Thus, one of the outstanding features of the methods of the present disclosure is that they do not require any prior knowledge or data about the individual profiles, contributors, or components of the complex DNA mixture.

[0034] In some embodiments, the techniques described herein can be used to determine the ethnicity of an individual associated with DNA present in a biological sample.

[0035] In embodiments, the present disclosure provides methods for identifying microhaplotypes in a genome. Microhaplotypes are useful for use in any of the methods disclosed herein, for example, in detecting sample contamination, disease analysis, and / or deconvolution of complex samples.

[0036] Thus, the present disclosure provides a method for identifying microhaplotypes in a genome, the method including: a) identifying a region of interest in the genome; b) detecting SBS within the region of interest, thereby generating a plurality of sets of sequence variants; c) analyzing each set of variants for LD to identify candidate microhaplotypes; and d) identifying the candidate microhaplotypes.

[0037] Also provided is a method comprising: a) identifying a SNP set with at least three microhaplotypes in a sample; and b) quantifying the frequency of haplotypes within the SNP set with more than two microhaplotypes.

[0038] Additionally, the present disclosure also provides a method comprising: a) identifying a SNP set with at least three microhaplotypes in a sample; and b) quantifying the frequency of haplotypes within the SNP set with more than two microhaplotypes to determine the presence or absence of DNA contamination in the sample.

[0039] Also provided is a method of genetic analysis comprising: a) identifying a SNP set with at least three microhaplotypes in a sample; and b) quantifying the frequency of haplotypes within the SNP set with more than two microhaplotypes to determine the presence or absence of a genetic marker indicative of a disease or disorder.

[0040] In various embodiments, the methodology of the present disclosure may further include quantifying the frequency of SNP sets with at least 3, 4, 5, 6, or more microhaplotypes in the sample. This may be done to determine the amount of DNA contamination in the sample. In embodiments, as discussed in Example 1, the method further includes calibrating a cutoff value for candidate microhaplotypes. Sample contamination can be assessed using a cutoff value determined for the frequency of candidate microhaplotypes with SNP sets involving at least 3, 4, 5, 6, 7, 8, or more microhaplotypes.

[0041] The microhaplotypes of the present invention can use different sets of SNPs, but the principles for their selection are the same. As discussed here, these principles include: using databases such as gnomAD™ (for exons, approximately 52% European, 7% East Asian, and 6% African) to select candidate SNPs, and the 1000 Genomes™ database (approximately 20% European, 20% East Asian, and 26% African) to assess LD; and using 1000 Genomes™ databases for 3rd / 4th haplotypes to equalize variation between ancestries. Selecting the final set of SNPs based on Genomes frequencies (or a similar database) (using the gnomAD database allows for slightly greater variation among Europeans); variants must be close enough to be on the same sequence read; avoiding repetitive sequences / indels and using single base substitutions to minimize error rates; avoiding homopolymers and low-confidence sequence regions; selecting SNPs in low LD to increase the frequency of tertiary / quaternary haplotypes; maximizing the distance between SNP sets to ensure information independence; and testing the candidate SNP sets against actual samples to ensure high coverage, diverse genotypes, and a low rate of tertiary / quaternary haplotypes in pure samples.

[0042] The methodology of the present disclosure may include identifying a set of candidate variants for analysis, as discussed in Example 1.

[0043] This may involve identifying a region of interest in the genome and determining the nucleotide sequence of that region for use in the analysis. The region of interest is examined for the presence of SBS. In embodiments, the SBS frequency is typically between about 5-95%, which may be determined using a suitable genome database, such as the gnomAD™ database (gnomad.broadinstitute.org / ).

[0044] In embodiments, the utilized region of interest optionally includes flanking regions that are also tested for the presence of SBSs and whose frequency is determined to be between about 5-95%. In various embodiments, the flanking regions of the region of interest comprise less than about 50, less than about 100, less than about 150, less than about 180, or less than about 200 nucleotide base pairs. In various embodiments, the total length of the region of interest, including optional flanking regions, is less than about 500, less than about 450, less than about 400, less than about 350, less than about 300, less than about 250, less than about 200, less than about 150, less than about 100, less than about 90, less than about 80, less than about 70, less than about 60, less than about 50, less than about 40, less than about 30, less than about 20, or less than about 10 base pairs.

[0045] In embodiments, the identified candidate variant pairs are then examined for LD, which may be performed using the 1000 Genomes™ database (ldlink.nci.nih.gov / ?tab=ldhap).

[0046] Pairs, triplets, quartets, and the like, having at least three haplotypes, with the third and further haplotypes having a combined frequency greater than 1%, are then considered for use. In various embodiments, the microhaplotype variant set is selected to avoid insertions / deletions, because such variants have a higher inherent sequencing error rate and are more likely to generate noise. In some embodiments, the variant may not be present in the 1000 Genomes™ database, and therefore cannot be easily assessed for LD. However, such variants may be utilized if the MAF found in the gnomAD™ database suggests that they are suitable.

[0047] It will be understood that the region of interest may be within a gene, intron, and / or exon, or between genes. Alternatively, the region of interest may be within an exome. In embodiments, the region of interest may include a genetic marker associated with a disease. In embodiments, the region of interest may include a genetic marker associated with a particular ethnicity.

[0048] This approach may be utilized to generate oligonucleotide panels for amplifying or hybrid capture specific regions containing microhaplotypes identified using the methods of the present disclosure. In one embodiment, the oligonucleotide panel comprises oligonucleotides for amplifying or hybrid capture regions of the genome corresponding to one or more genomic regions listed in Table 5. In another embodiment, the oligonucleotide panel comprises oligonucleotides for amplifying or hybrid capture regions of the genome corresponding to one or more genomic regions listed in Tables 6 or 7.

[0049] Thus, the present disclosure also provides methods of genetic analysis that include: a) amplifying a region of a genome present in a sample, where the region corresponds to a genomic region set forth in Table 5, Table 6, and Table 7, to generate an amplicon by amplification; and b) sequencing the amplicon to determine the nucleic acid sequence of the amplicon.

[0050] As discussed herein, microhaplotypes identified by the methods of the present disclosure may be utilized in a variety of applications, including, but not limited to, DNA contamination detection, disease analysis, and sample deconvolution (i.e., detection of DNA from multiple subjects or cell types in a single sample).

[0051] In one embodiment, the present disclosure provides a method for detecting SNP sets with at least three microhaplotypes derived from multiple subjects present in a sample. The method includes: a) identifying microhaplotypes in the genome of the sample; b) determining the number of SNP sets with at least three microhaplotypes in the sample; and c) quantifying the frequency of SNP sets with more than two microhaplotypes to determine the presence of DNA derived from the multiple subjects in the sample, thereby detecting DNA derived from the multiple subjects in the sample. In one embodiment, the identifying includes: i) identifying a region of interest in the genome; ii) detecting SBSs within the region of interest, thereby generating multiple sequence variant sets; and iii) analyzing each variant set for LD to identify microhaplotypes.

[0052] In another embodiment, the present disclosure provides a method for detecting a SNP set having at least three microhaplotypes derived from a plurality of subjects present in a sample, the method comprising: a) determining the presence or absence of a SNP set having at least three microhaplotypes in the sample, wherein the SNP set comprises a plurality of single base pair substitutions and corresponds to a genomic region set forth in Tables 5, 6, and 7; and b) quantifying the frequency of the SNP set to determine the presence of DNA derived from the plurality of subjects in the sample, thereby detecting a SNP set having at least three microhaplotypes derived from the plurality of subjects in the sample.

[0053] Therefore, the disclosed method of deconvoluting or decomposing components from a complex DNA mixture can be carried out by analyzing a single complex DNA mixture.In certain embodiments of the disclosed method of deconvoluting or decomposing components from a complex DNA mixture, the method can analyze two or more complex DNA mixtures.The resolution of DNA profile using these methods increases as the number of SNP loci in the panel used increases.As used herein, the term complex DNA mixture refers to a DNA mixture that contains DNA from two or more contributors.Preferably, the complex DNA mixture of the method described herein contains DNA from at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more contributors.

[0054] The disclosed method is superior to existing methods for deconvoluting DNA profiles. It is noteworthy that the application of the method described herein is not limited to forensic analysis or DNA contamination detection. For example, the disclosed method may be used for medical diagnosis and / or prognosis. To detect disease, a region of interest may be selected to contain genetic markers associated with a disease or condition, such as cancer or fetal disorders. In this method, the region of interest may be located on chromosome 21, for example, to enable the diagnosis of trisomy 21, also known as Down syndrome. If samples are determined to be derived from the mother and fetus and the frequency of a third microhaplotype differs on chromosome 21 compared to other chromosomes, this indicates a mutation in a gene copy, such as trisomy 21. Other trisomies, including chr13 trisomy and chl18 trisomy, can also be detected.

[0055] Thus, the methods described herein may be used in a variety of ways to predict, diagnose, and / or monitor diseases such as cancer, fetal disorders, etc. Additionally, the methods may be utilized to distinguish various cell types from one another.

[0056] In the field of cancer, biopsy samples often contain many cell types, only a small portion of which may form any part of the tumor. As a result, DNA obtained from a tumor biopsy is another form of complex DNA mixture and may contain somatic variants occurring on specific DNA molecules. In the case of somatic diversity, restrictions on SBS can be relaxed because somatic diversity may be indels or other modifications that might otherwise be avoided. Furthermore, within a tumor, many cells may be molecularly distinct, for example, with respect to the expression of factors indicative of or promoting angiogenesis and / or metastasis. DNA mixtures obtained from tumor samples may also form the complex DNA mixtures of the present disclosure. In both of these non-limiting examples, the methods of the present disclosure may be used to construct individual profiles for each cell or cell type contributing to the complex DNA mixture. Furthermore, the methods of the present disclosure may be used to deconvolute contributors to a complex DNA mixture. As an illustrative example, a complex DNA mixture obtained from a breast cancer tumor biopsy may be used to construct individual profiles of malignant cells. For the same patient, brain cancer tumor biopsy, this individual profile may be used to deconvolute the contributors to the complex DNA mixture obtained from the brain cancer tumor biopsy, to determine, for example, whether malignant breast cancer cells from that subject have metastasized to the brain and formed a secondary tumor. This method may address the question of whether tumors arise independently, or alternatively, whether these tumors are related.

[0057] Thus, the present disclosure provides a method for detecting a disease or disorder in a subject. The method includes: a) obtaining a sample from the subject; b) identifying microhaplotypes in DNA molecules present in the sample; c) determining the presence or absence of a SNP set having more than two microhaplotypes in the sample; and d) quantifying the frequency of haplotypes within the SNP set to determine the presence or absence of a genetic marker indicative of the disease or disorder, thereby detecting the disease or disorder. In one embodiment, the identifying includes: i) identifying a region of interest, where the region of interest is associated with the disease or disorder; ii) detecting SBS within the region of interest, thereby generating a plurality of sequence variant sets; and iii) analyzing each variant set for LD to identify microhaplotypes.

[0058] In various embodiments, the genome is present in a biological sample obtained from a subject. The biological sample can be virtually any type of biological sample, particularly a sample containing DNA. The biological sample can be a germline cell line, stem cells, reprogrammed cells, cultured cells, or a tissue sample containing 1,000 to about 10,000,000 cells, or a fluid containing circulating DNA. In embodiments, the sample comprises DNA from a tumor or liquid biopsy, such as, but not limited to, amniotic fluid, aqueous humor, vitreous fluid, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, chyme, endolymph, perilymph, stool, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal mucus and sputum), pericardial fluid, peritoneal fluid, pleural fluid, pus, mucosal secretions, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomit, prostatic fluid, nipple aspirate, tears, sweat, buccal specimen collection, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluid. In one embodiment, the sample comprises DNA from circulating tumor cells. In embodiments utilizing amplification protocols such as PCR, it is possible to obtain a sample containing a large number of cells, even a single cell. The sample need not contain intact cells, as long as it contains sufficient biological material (e.g., DNA) to perform genetic analysis of one or more regions of the genome.

[0059] In some embodiments, biological or tissue samples can be obtained from any tissue containing cells with DNA or from fluids with circulating DNA. Biological or tissue samples can be obtained by surgery, biopsy, mucosal specimen collection, stool, or other collection methods. In some embodiments, the sample is derived from blood, plasma, serum, lymph, nerve cell-containing tissue, cerebrospinal fluid, biopsy material, tumor tissue, bone marrow, neural tissue, skin, hair, tears, urine, fetal material, amniocentesis material, uterine tissue, saliva, stool, or sperm. Methods for isolating PBLs from whole blood are well known in the art.

[0060] As disclosed above, the biological sample can be a blood sample. Blood samples can be obtained using methods known in the art, such as finger prick or phlebotomy. Preferably, the blood sample is about 0.1 to 20 ml, or about 1 to 15 ml, with a blood volume of about 10 ml. Small amounts can be used, as can free DNA circulating in the blood. Microsampling, needle biopsy, catheter sampling, and sampling by excretion or production of DNA-containing bodily fluids are also potential sources of biological samples.

[0061] In the present invention, the subject is typically a human, but can also be any species, including, but not limited to, a dog, cat, rabbit, cow, bird, rat, horse, pig, or monkey.

[0062] The disclosed methods utilize nucleic acid sequence information and, therefore, can include any method for performing nucleic acid sequencing, including nucleic acid amplification, polymerase chain reaction (PCR), nanopore sequencing, 454 sequencing, and insertion-tagged sequencing. In embodiments, the disclosed methodology utilizes systems from Illumina, Inc. (including, but not limited to, HiSeq™ X10, HiSeq™ 1000, HiSeq™ 2000, HiSeq™ 2500, Genome Analyzers™, MiSeq™, NextSeq, and NovaSeq systems), Applied Biosystems Life Technologies (SOLiD™ System, Ion PGM™ Sequencer, ion Proton™ Sequencer), Genapsys, BGI MGI, and others. Nucleic acid analysis can also be performed by systems from Oxford Nanopore Technologies (GridiON™, MinION™) or Pacific Biosciences (Pacbio™ RS II or Sequel I or II). Importantly, in embodiments, sequencing can be performed using any of the methods described herein. When long-read technologies such as PacBio™ or Oxford Nanopore™ are used, the length limitations imposed on the DNA are relaxed, allowing SNPs to be further apart, consistent with longer read lengths.

[0063] The present invention, including systems that perform the steps of the disclosed methods, is described, in part, in terms of functional components and various process steps. Such functional components and process steps may be realized by any number of components, operations, and techniques configured to perform designated functions and achieve various results. For example, the present invention may employ various biological samples, biomarkers, elements, materials, computers, data sources, storage systems and media, information collection techniques and processes, data processing standards, statistical analysis, regression analysis, and the like, which may perform various functions.

[0064] Genetic analysis methods according to various aspects of the present invention may be implemented in any suitable manner, for example, using a computer program running on a computer system. Exemplary genetic analysis systems according to various aspects of the present invention may be implemented in conjunction with a computer system, such as a conventional computer system including a processor and random access memory, such as a remotely accessible application server, network server, personal computer, or workstation. The computer system may also preferably include additional memory or information storage systems, such as a mass storage system, and a user interface, such as a conventional monitor, keyboard, and tracking device. However, the computer system may include any suitable computer system and associated equipment and may be configured in any suitable manner. In one embodiment, the computer system comprises a standalone system. In another embodiment, the computer system is part of a network of computers, including a server and a database.

[0065] The software necessary to receive, process, and analyze genetic information may be implemented on a single device or multiple devices. The software may be accessible over a network, allowing information storage and processing to occur remotely from the user. Genetic analysis systems according to various aspects of the present invention, and their various components, provide functions and operations that facilitate genetic analysis, such as data collection, processing, analysis, reporting, and / or diagnosis. For example, in this embodiment, a computer system may execute a computer program that receives, stores, retrieves, analyzes, and reports information related to the human genome or regions thereof. The computer program may include multiple modules that perform various functions or operations, such as a processing module that processes raw data to generate supplemental data and an analysis module that analyzes the raw data and supplemental data to quantitatively assess contamination or disease models and / or generate diagnostic information.

[0066] The procedures performed by the genetic analysis system may include any suitable steps that facilitate genetic analysis and / or disease diagnosis. In one embodiment, the genetic analysis system is configured to establish a disease model and / or determine a patient's disease state. Determining or identifying a disease state may include generating any useful information regarding the patient's condition related to the disease, such as making a diagnosis, providing information useful for the diagnosis, assessing the stage or progression of the disease, identifying conditions that may indicate susceptibility to the disease and whether further testing is recommended, predicting and / or assessing the effectiveness of one or more treatment programs, or otherwise assessing the disease state, likelihood of disease, or other aspects of the patient's health.

[0067] The genetic analysis system preferably generates a disease model and / or provides a diagnosis for a patient based on the genetic data and / or additional subject data related to the subject. The genetic data may be obtained from any suitable biological sample, as well as from a database storing genetic information.

[0068] The following examples are provided to further illustrate the advantages and features of the present invention, but are not intended to limit the scope of the invention. The examples are typical of those that might be used, but other procedures, methodologies, or techniques known to those skilled in the art may be used instead. [Example]

[0069] Example 1 Sample contamination detection In this example, the methodology of the present disclosure was utilized to detect sample contamination. The following provides a detailed discussion of the methods and steps used for detection.

[0070] Identification of candidate variant sets For each region of interest, the region to be sequenced, along with any additional border regions (up to 100 bp), was examined for SNP frequencies between 10% and 90% according to the gnomAD™ database (gnomad.broadinstitute.org / ). Once a variant was found that was not in a low-confidence region, the adjacent 180-bp regions in both directions were examined for SNP frequencies between 5% and 95%. These cutoffs may vary depending on the type of sample being analyzed for various panels and the number of SNP sets required. Such variant pairs were then examined for LD using the 1000 genomes data (ldlink.nci.nih.gov / ?tab=ldhap). Pairs, triplets, etc., with at least three haplotypes, with the third and / or further haplotypes having a combined frequency greater than 1%, were considered for use. These cutoffs could potentially be expanded to include additional variant sets, if necessary, or restricted to retain only the most informative variant sets to minimize noise. For example, the variant set was selected to avoid insertions / deletions because such variants have a higher intrinsic sequencing error rate and are more likely to introduce noise. Similarly, other sequence contexts may be advantageous based on error rate. Furthermore, some variants were not found in the 1000 Genomes™ database and therefore could not be assessed for LD, but were advanced for candidate testing if the MAF observed in gnomAD™ suggested they might be suitable. While SNPs could theoretically be as far apart as their paired read counterparts, to simplify analysis, we selected SNPs that are located closer to each other and covered by a single read.

[0071] Characterization of the candidate variant set The candidate variant sets were further evaluated in actual samples to ensure there were enough reads to harbor both / all variants on the reads to generate phased haplotypes. A cutoff of 100x median coverage for each SBS was used to ensure that all or nearly all SNP sets were included in each comparison. High coverage is necessary to maximize analytical sensitivity. For other panels, the correct set of SBSs to use will vary depending on the panel being examined. Furthermore, some sequence contexts have higher error rates than others, and their use may result in additional artifactual microhaplotypes. Variant sets that tended to overrepresent third / fourth microhaplotypes in supposedly pure samples were excluded from use because they could introduce high levels of noise into the signal.

[0072] A set of 106 variants was selected for use with the 507-gene panel (Table 5) based on high coverage and low background noise levels. Whenever possible, the distance between the SBS sets was maximized to minimize redundant information. The MAFs listed for the SBSs in this table were obtained from the "All Populations" 1000 Genomes™ database and differ from the original MAFs obtained from gnomAD™.

[0073] Estimating contamination levels Because any sample could theoretically be contaminated, it was necessary to characterize the samples before using them for calibration so that the process could begin with a pure sample. Furthermore, because variant and microhaplotype frequencies can vary significantly by ethnicity, it is useful to characterize samples with different ethnicities to ensure that a given set of SBSs works well across all samples and contaminants. For this dataset, five African, five Asian, and six European (all self-identified) individuals were selected based on coverage of at least 10 / 10 variant sets and fewer than two variant sets with more than two microhaplotypes. These samples and their characteristics are listed in Table 1. The European samples have a non-significantly reduced number of single-microhaplotype SBSs.

[0074] (Table 1) Samples used for calibration TIFF0007722929000001.tif113128

[0075] To mimic in silico contamination, unfiltered fastQ™ reads from a pure sample were computationally mixed with other samples to generate artificially "contaminated" samples. To target X% contamination, 100-X% of the reads from the nominal sample were mixed with X% of the reads from the "contaminant." These mixed samples were then run through the pipeline and aligned and called using our standard methods. For each sample, the number of haplotypes and their frequencies in each SBS set were counted and tabulated. The frequencies of third haplotypes, if any, were then examined within each sample for each SBS set, and the minimum, maximum, median, and mean for each set of third haplotype frequencies were calculated. The mixtures were then examined to see how well these parameters predicted contamination.

[0076] Before examining the results in detail, we considered how multiple technical and biological confounding factors might affect the results. Even in "pure" samples, technical noise exists, resulting in a low number of third / fourth haplotypes. To avoid these factors interfering with contaminant detection, we set a minimum number of third / fourth haplotypes. Since the desired level of contaminant detection is 1–2%, we chose the minimum number of third / fourth haplotypes to be in the range of 5–10. This avoids the issue of misidentifying low-level technical noise as contaminant.

[0077] (Table 2) Number of SBS sets with more than two microhaplotypes (n = 70 each) TIFF0007722929000002.tif22128

[0078] The percentage of SNPs with more than two microhaplotypes determines whether a sample is contaminated, but this percentage is relatively insensitive to the degree of contamination. Because the percentage of more than two microhaplotypes rapidly reaches a maximum, contaminations of 2%, 5%, and 20% appear very similar when looking at this parameter alone. To circumvent this issue, we quantified contamination levels using the MAF for the third haplotype. This value can be misleading at low contamination levels due to technical artifacts. This value can appear abnormally high due to the possibility that contaminating DNA contributes two copies of the third haplotype, making contamination appear two-fold higher than it actually is (Figure 3). Extreme copy number variation, commonly found in tumor samples, also affects apparent contamination in either direction, depending on which haplotype is in excess. While this is not typically an issue with normal DNA, it can be a serious problem with tumor DNA. To avoid these challenges, we use the median MAF for the third haplotype to minimize the effect of either unusually high or low MAFs. There is additional information found in the allele frequencies for the second and fourth microhaplotypes, but this data was not used in the calculations. More complex analyses of haplotype frequencies can be used if there is a sufficient set to examine.

[0079] For samples with more than the set number of third / fourth haplotypes, various factors may interfere with accurate frequency determination. In this calibration series, one technical challenge is whether the nominal contamination level is actually accurate. Although the number of spiked reads can be precisely controlled, each sample has different characteristics in terms of DNA quality, which may affect the functional contamination level. Samples with a wide range of DNA lengths will have different functional contamination levels due to differences in DNA quality or the proportion of on-target reads due to different capture efficiencies. This is because the frequency of a set of SNPs appearing in the same read is length-dependent. This means that 1% spiked reads are functionally equivalent to 0.5%, 2%, or any length in between. For this reason, each sample and its contaminants were swapped in parallel as sample and contaminant. This normalizes quality differences to some extent and provides a better estimate of the functional contamination level. When applying these methods to real samples, functional contamination rather than stoichiometric contamination becomes more important due to the possibility of incorrect variant calls.

[0080] There are also biological reasons for the quantitative challenge. A pure sample may have one or two microhaplotypes in each SBS set, and one or two microhaplotypes from the contaminant may match one, two, or neither of the microhaplotypes in the primary sample. At low contamination levels, when signal is just beginning to emerge, new third haplotypes will predominantly consist of double contributions that do not match the sample's microhaplotypes, whereas at higher contamination levels, there will be a mix of single and double contributions. Therefore, it is advisable not to expect a simple linear relationship between contamination levels and the frequencies of various haplotypes. Compounding this difficulty, extensive copy number variation occurs between tumor samples, which can also significantly affect haplotype frequencies. Because of these caveats, we used empirical estimates of contamination because simply looking at the frequency of third haplotypes will overestimate low contamination levels and underestimate high contamination levels. With a larger set of variants at very high coverage levels, it may be possible to combine frequency data to obtain even better estimates of functional contamination. As shown in Table 3, using this SNP set and coverage criteria, the region where overcounting and undercounting are balanced to yield a relatively accurate contamination estimate is approximately 2%. This is approximately the level at which we want to set sensitivity, so we use the median value of the third haplotype as an approximation of the contamination level, and accuracy may be challenged if the range is too far from 2%. Accurate estimation of other contamination levels will require examining more admixtures, as was done with other SBS sets.

[0081] Table 3. Median frequency of the third haplotype by ethnicity. TIFF0007722929000003.tif39128

[0082] Application to actual samples Samples used in the in silico contaminant mixture were selected based on their high quality. Unfortunately, there is much greater variability in real-world samples, making it necessary to set standards for which samples can be analyzed and how the analysis should be performed. Ideally, every sample would have greater than 100× coverage across all 106 SBS sets, but this is often not the case. Missing SBS sets can lead to inconsistent comparisons, and low coverage at a particular SBS can result in a significant overestimation or absence of the third haplotype frequency. Therefore, 1000 samples were run through a standard pipeline and microhaplotype data were examined. Of these 1000 samples, 151 failed standard quality control metrics, leaving 849 samples for microhaplotype analysis. A minimum coverage of 20 is required for SBS counting. The majority of samples (709) have data for all 106 SBS sets. However, some samples have significantly fewer SBS sets that meet the minimum criteria. The point at which more samples fail than pass other quality control measures is 100 SBS calls. Therefore, the following analysis uses only the 825 samples that passed more than 100 SBS calls. Of these 825 samples, 24 failed SNPCheck™, which was previously used to monitor sample contamination.

[0083] Table 4 shows the effect of varying the cutoff on contamination detection for these 825 samples. Samples are considered passing if either the SBS set with more than two microhaplotypes is less than the cutoff number or the median MAF of the third microhaplotype is below the set threshold. Based on the in silico experiments described above, the number of SBS sets with more than two microhaplotypes should be in the range of 5 to 10 with these microhaplotypes. In addition, samples with a median third haplotype frequency below 1.5% are also considered passing, even if more microhaplotypes than the cutoff number are present. Using these cutoffs, 804 to 811 samples are passing, including 18 to 19 samples that failed SNPCheck™. If the third haplotype frequency is between 2 and 4%, samples are optionally checked to determine whether the level of contamination is potentially problematic based on the observed somatic mutation frequency. Of these 11-18 samples, 4-5 failed SNPCheck™. Samples with a third microhaplotype frequency greater than 4% failed. In each case, this resulted in three samples, one of which failed SNPCheck™. In addition to the 825 passing runs described above, SNPCheck™ was also run on samples that failed other QC metrics or had too few SBSs called in the disclosed microhaplotyping method. Of the four samples that failed QC and SNPCheck™, three failed the microhaplotyping method, with contamination greater than 10%. Of the seven samples that failed SNPCheck™, which would not normally be assessed by microhaplotypes with fewer than 101 SBSs called, four also failed the microhaplotyping method regardless of cutoff, while another failed at some cutoff value.

[0084] Table 4. Comparison of Microhaplotypes and SNPCheck™ TIFF0007722929000004.tif60166

[0085] Perfect agreement between the method of the present invention and SNPCheck™ was not expected. SNPCheck™ generates false positives by calling pure samples as contaminated, thereby failing some tumor samples with very high copy number diversity. False negatives are also known to occur when the level of contamination is so high that the diversity is mistaken for germline variation.

[0086] Contamination detection in the exome Many of the SBSs used in the 507-gene panel are in non-coding regions, making them useless in exome analysis. Therefore, a new set of SBSs was selected to interrogate the exome. Because exome coverage is low per ROI, capturing variants with as much coverage as possible is more important. Therefore, the SBS set was selected to ensure closer spacing between variants and closer to exons than the 507-gene panel. Also, because the number of ROIs is so large, we attempted to include more informative SBSs, which were selected in ROIs with higher-than-average coverage. These were then examined in the exome data set, and SBSs with a median coverage greater than 80 and diverse haplotypes were selected for use in the panel. These SBS sets are listed in Table 6. Using a method similar to that described above, two exomes suspected of contamination were examined and found to be more than 15% contaminated using this SBS set.

[0087] Using the initial set of microhaplotypes used in the 507-gene panel, we observed differences in sensitivity between different ancestry groups. This challenge was likely due to both bias in the database used to select the microhaplotype set and differences in heterozygosity across different ancestries. To correct for this, we used population haplotype frequencies obtained from the 1000 genomes project to balance the third / fourth haplotype frequencies so that they were approximately equal across all ancestries. We summed the frequencies of third / fourth haplotypes across SNP sets and dropped SNP sets that contributed excessive frequencies in overrepresented ancestries. This allowed us to generate a microhaplotype set with the same expected average number of third / fourth haplotypes among people of East Asian, African, and European ancestry. However, it was not possible to simultaneously generate identical frequencies for the other two 1000 genome ancestries: mixed-race Americans and South Asians. Both of these ancestors had higher frequencies of the third / fourth microhaplotype than the other three ancestors, so contamination should be easily detectable using the same thresholds as the other ancestors.

[0088] To further improve performance characteristics, we attempted to select only microhaplotype sets with high coverage and low noise among pure samples. We increased the minimum average coverage of the SNP set from 100 to 250. However, high coverage is a double-edged sword. While high coverage improves sensitivity and accuracy, it can also generate artifactual third haplotypes due to inherent sequencing errors, typically at a level of 0.1%. To minimize the impact of such technical errors, low-frequency haplotypes can be excluded from consideration. The level at which this should be set can be optimized based on coverage and sequencing quality. For our experiments, we set the threshold at 0.2%, where haplotypes with frequencies below 0.2% were considered impractical. Other thresholds can be used depending on sequence quality and other factors.

[0089] Additionally, a larger set of SNPs was used to enhance the signal and allow for more accurate contamination estimates. Based on these considerations, a set of 164 SNPs was selected for the second microhaplotype panel that met all of these criteria. Fifty-one of these SNP sets were also present in the first panel. Both sets are shown in Table 7, along with the region, dbSNP number, and 1000 genome frequency of the third and fourth haplotypes.

[0090] As discussed above, generating samples with precise levels of contamination is extremely challenging. Combining samples in silico can yield mixed samples with precise levels of contamination, but the functional impact is not necessarily precise. Because microhaplotype detection depends on the length of the sequenced molecule, samples with identical partial components but different DNA qualities will have different effects on microhaplotype frequencies. To minimize this effect, samples were analyzed in pairs, swapping the "sample" and "contaminant" and then averaging the results within each pair. Fifteen such pairs for each category (African, East Asian, European, and mixed-race individuals) were then analyzed for the number of third / fourth microhaplotypes as a function of contamination level. As shown in Figure 1, the third / fourth MH counts for East Asian and European-ancestry individuals were nearly overlapping. The third / fourth MH counts for individuals of African-American ancestry and those of mixed ancestry were higher than those for East Asians / Europeans, but were similar to each other. The discrepancy in African Americans is likely due to the composition of the 1000 genomes African panel, which includes five subgroups from Africa and two subgroups from African Americans. These two groups are somewhat admixed and therefore produce higher values than the other groups. The more even frequency of the third / fourth microhaplotypes, combined with the larger set of microhaplotypes tested, allows for more reliable identification of contaminated samples.

[0091] Although the number of third / fourth microhaplotypes varies slightly between different ancestries, the median frequency of third microhaplotypes as a function of contamination level is nearly identical across ancestries, including mixed samples from different ancestries (Figure 2). This relationship is linear from approximately 1%. Contamination levels below 1% are not only significantly affected by sequencing artifacts, but also the possibility of additional DNA being present that may be contaminating beyond the intended level. Above 1%, the median observed frequency is roughly half the contamination level. This is expected based on the way third MHs are generated, as shown in Figure 3. At higher contamination levels, this value begins to decline, but this is due to multiple factors, including the possibility that third microhaplotypes are actually derived from the sample rather than from the contaminant.

[0092] Using the relationship of contamination level = 2 × median third-level microhaplotype level, the detection results for different levels of contamination are shown in Table 8 for each ancestor. The patterns are similar: when the expected contamination level is twice the third-level microhaplotype, the percentage of samples detected at higher contamination levels decreases. This table provides guidance on where the threshold must be set to achieve near-100% detection of contamination at a given level. For example, if one wants to detect almost all samples contaminated at 2%, setting the third-level microhaplotype cutoff = 0.75% will detect 97% of samples contaminated at 2%, while 82% of samples contaminated at 1.5%, only 15% of samples contaminated at 1%, and none at 0.5%. The choice of threshold can be based on the relative levels of false positives and false negatives.

[0093] Example 2 Use of microhaplotypes for NIPT detection of chromosomal abnormalities Noninvasive prenatal testing (NIPT) for detecting chromosomal abnormalities is performed by collecting a maternal blood sample and assessing circulating fetal DNA in the presence of a large background proportion of maternal DNA. Typically, sequence reads are simply aligned and the number aligning to each chromosome is counted. A positive diagnosis is indicated by an excess of reads aligning to trisomy-susceptible chromosomes (usually chr13, chr18, and chr21). The test is typically performed after the 10th week of pregnancy, when the amount of fetal DNA in the maternal blood is sufficient for accuracy. The use of microhaplotypes allows for earlier testing because more accurate quantification is possible at lower DNA concentrations, resulting in more accurate results, independent of pre-existing benign maternal copy number variations that can lead to misinterpretation.

[0094] The behavior of NIPT samples will be more straightforward than that of tumor samples for two reasons. First, the complexity of extensive copy number variation will be less of a challenge. Second, one of the fetal haplotypes will already be present in the mother, and the third haplotype coming from the paternal side will only be a single copy and therefore will not be overcounted at low levels. Thus, a more predictable increase in frequency is expected.

[0095] In most cases of trisomy 21, the extra chromosome originates maternal, reducing the contribution of the new paternal haplotype on that chromosome. Therefore, the paternal haplotype frequency on unaffected chromosomes can be determined and compared to the paternal haplotype frequency on potentially affected chromosomes. Because many SBS sets are available, a list of well-behaved SBSs can be generated directly. These SBSs can be enriched by target capture or PCR amplification, allowing for earlier detection than currently possible. Unbiased PCR amplification of DNA for typical NIPT is challenging because slight nonlinearities affect quantitation. Microhaplotyping methods look at the ratio of microhaplotypes rather than simply counting the number of reads, reducing sensitivity to amplification bias. Accuracy can be further improved by selecting SBS sets that are less prone to sequencing errors or by selecting multiple SBS sets that result in two or more sequence changes going from the maternal microhaplotype to the paternal microhaplotype. Additionally, fetal DNA fraction can be easily determined by examining the genotype frequencies in SNP sets with three microhaplotypes. The fetal fraction is twice the frequency of the third microhaplotype. Knowledge of the fetal fraction and its variability allows for a more accurate determination of whether a test result is valid or indeterminate.

[0096] To determine trisomy or other DNA copy number abnormalities, the frequencies of the third microhaplotypes in different regions are compared. If the frequency of the third microhaplotype from any large genomic region (part or entire chromosome) differs from the frequency of other genomic regions, this indicates trisomy or other amplification (increased frequency of the third microhaplotype) or deletion (absence of the third microhaplotype). Supplementary table

[0097] Table 5. SBS set of 507 gene panel TIFF0007722929000005.tif229106TIFF0007722929000006.tif229132TIFF000 7722929000007.tif229122TIFF0007722929000008.tif229112TIFF00077229290 00009.tif229122TIFF0007722929000010.tif229122TIFF0007722929000011.t if229122TIFF0007722929000012.tif229122TIFF0007722929000013.tif229132

[0098] (Table 6) SBS set for exome analysis TIFF0007722929000014.tif195110TIFF0007722929000015.tif190129TIFF000 7722929000016.tif190129TIFF0007722929000017.tif190128TIFF00077229290 00018.tif190128TIFF0007722929000019.tif190128TIFF0007722929000020.t if190128TIFF0007722929000021.tif190128TIFF0007722929000022.tif190128

[0099] Table 7: SNP set TIFF0007722929000023.tif24078TIFF0007722929000024.tif235128TIFF0007722929000025.ti f235137TIFF0007722929000026.tif235128TIFF0007722929000027.tif234125TIFF00077229290 00028.tif235137TIFF0007722929000029.tif234137TIFF0007722929000030.tif234125TIFF000 7722929000031.tif234125TIFF0007722929000032.tif234125TIFF0007722929000033.tif234125 TIFF0007722929000034.tif234125TIFF0007722929000035.tif234125TIFF0007722929000036.t if234125TIFF0007722929000037.tif234117TIFF0007722929000038.tif234125TIFF0007722929 000039.tif234125TIFF0007722929000040.tif235129TIFF0007722929000041.tif235138TIFF00 07722929000042.tif237129TIFF0007722929000043.tif235129TIFF0007722929000044.tif23569

[0100] Table 8. Observed frequency of third MH (x2) TIFF0007722929000045.tif213124TIFF0007722929000046.tif59128

[0101] Although the present invention has been described with reference to illustrative embodiments, it will be understood that modifications, changes, and variations are encompassed within the spirit and scope of the invention. Accordingly, the invention is limited only by the appended claims.

Claims

1. 1. A method for identifying microhaplotypes in a genome, comprising: a) identifying a region of interest in the genome; b) determining additional border regions for the region of interest, the border regions comprising flanking regions of the region of interest, the flanking regions of the region of interest comprising less than 200 nucleotide base pairs; c) detecting single base pair substitutions (SBS) within the target region and the additional border regions to generate a plurality of sequence variant sets; d) filtering the set of variants based on a frequency threshold for detected SBS; e) analyzing each remaining variant set for linkage disequilibrium to identify candidate microhaplotypes, wherein the linkage disequilibrium is assessed using one or more population-wide genomic databases, and the candidate microhaplotypes are identified based on the presence of at least three microhaplotypes, wherein the third and more microhaplotypes have a combined frequency of greater than 1% when compared to data from the one or more population-wide genomic databases; f) providing candidate microhaplotypes; The method comprising:

2. 2. The method of claim 1, wherein the flanking regions of the region of interest comprise fewer than 50, fewer than 100, fewer than 150, or fewer than 180 nucleotide base pairs sequenceable by a short-read sequencer.

3. 2. The method of claim 1, wherein SBSs with a frequency of less than 10 or greater than 90% are filtered.

4. 2. The method of claim 1, wherein SBSs originating from the flanking regions of the region of interest and having a frequency of less than 5 or greater than 95% are filtered.

5. 10. The method of claim 1, further comprising calibrating a cutoff value for the candidate microhaplotype to assess contamination of the sample.

6. 5. The method of claim 4, wherein only DNA sequence reads that overlap with the candidate microhaplotype are used to calculate a contamination detection threshold and a degree of contamination.

7. 7. The method of claim 6, wherein the DNA sequences used to calibrate the contamination detection threshold and the degree of contamination are mixed pairwise in silico, using each DNA sequence alternately as the primary sample and the contaminant.

8. 8. The method of claim 6 or 7, wherein the number and genotype of SNP sets with one and / or two microhaplotypes are compared between different individuals to assess identity or contamination.

9. 6. The method of claim 5, further comprising assessing sample contamination using a cutoff value determined for the frequency of a candidate microhaplotype having a single nucleotide polymorphism (SNP) set involving at least three microhaplotypes.

10. 10. The method of claim 9, further comprising assessing sample contamination using a cutoff value determined for the frequency of a candidate microhaplotype having a SNP set with at least four or more microhaplotypes.

11. 10. The method of claim 1, wherein the candidate microhaplotypes correspond to one or more genomic regions selected from the genomic regions listed in Table 5, Table 6, or Table 7. 1) Genomic regions listed in Table 5: chr1:120057158-120057246, chr1:156846120-156846233, chr1:226589833-226589958, chr1:23885498-23885599, chr10:104386934-104387019, chr10:43615505-43615633, chr10:70332580-70332672, chr11:534197-534242, chr11:8246326-8246343, chr12:121416622-121416650, chr12:121431272-121431300, chr12:121435427-121435475, chr12:121437114-121437221, chr12:133208886-133208979, chr12:133226159-133226196, chr12:133253995-133254083, chr12:18656174-18656225, chr12:56494991-56494998, chr13:21562832-21562948, chr14:102568296-102568367, chr14:104165753-104165927, chr14:105239146-105239192, chr14:105258892-105258893, chr14:35872792-35872926, chr15:40998305-40998342, chr15:41857216-41857303, chr15:41860411-41860490, chr15:67457335-67457485, chr16:2138269-2138398, chr16:2138398-2138422, chr16:68857289-68857441, chr16:81819768-81819820, chr16:89806343-89806347, chr16:89849583-89849629, chr16:89858505-89858525, chr17:1782952-1782957, chr17:78599562-78599655, chr17:78820329-78820374, chr17:78865546-78865630,<h2 style=";text-align:left;direction:ltr">chr17:78897547-78897561, chr17:78921117-78921211, chr19:10267011-10267077, chr19:17937758-17937786, chr19:17955001-17955021, chr19:2226676-2226772, chr19:3119184-3119239, chr19:50919797-50919828, chr19:5210622-5210782, chr19:5210762-5210782, chr19:5212380-5212482, chr19:7166376-7166388, chr2:112754828-112754880, chr2:112754943-112755001, chr2:141259283-141259376, chr2:29416366-29416481, chr2:29416481-29416615, chr2:29446184-29446202, chr2:48010488-48010558, chr20:40714307-40714479, chr20:40714539-40714540, chr20:57478807-57478939, chr20:9543622-9543681, chr21:42845374-42845383, chr22:21337266-21337325, chr22:21348914-21349037, chr22:24158895-24158899, chr3:178922222-178922274, chr3:183211906-183212026, chr4:106196829-106196951, chr4:143043340-143043404, chr4:143324036-143324094, chr4:187534362-187534375, chr4:187629497-187629538, chr5:149456772-149456811, chr5:149495287-149495395, chr5:176517326-176517461, chr5:176523562-176523597, chr5:176721198-176721272, chr5:180046209-180046344,<h2 style=";text-align:left;direction:ltr">chr5:180051003-180051118, chr5:180057231-180057293, chr5:231111-231143, chr5:35861068-35861159, chr5:35871190-35871273, chr5:57754808-57754851, chr5:67522722-67522851, chr6:117725448-117725578, chr6:117730673-117730819, chr6:152382311-152382325, chr6:26056549-26056708, chr6:30865115-30865204, chr6:32188603-32188642, chr7:100410597-100410657, chr7:6026775-6026942, chr7:78119109-78119199, chr8:30999122-30999123, chr8:31024638-31024654, chr8:90958422-90958530, chr9:139403268-139403280, chr9:139405093-139405261, chr9:139410424-139410589, chr9:139411714-139411880, chr9:21968159-21968199, chr9:93639846-93639973, chr9:93641175-93641199, chr9:98238358-98238379, 2) Genomic regions listed in Table 6: chr1:3743319-3743391, chr1:10431132-10431158, chr1:32672908-32672932, chr1:94544234-94544276, chr1:154832290-154832304, chr1:159409857-159409884, chr1:171168545-171168584, chr1:183616884-183616926, chr11:4928841-4928866, chr11:5345128-5345170, chr11:5566030-5566051, chr11:63883985-63884027, chr11:85436303-85436352, chr11:116703640-116703671, chr12:6030405-6030437, chr12:40834918-40834955, chr12:113348849-113348870, chr12:121600180-121600253, chr12:132688115-132688137, chr13:25367282-25367301, chr14:23549285-23549319, chr14:65263300-65263347, chr14:96136775-96136794, chr15:41819283-41819322, chr15:79310256-79310288, chr15:89398330-89398407, chr15:94945704-94945719, chr16:2812890-2812939, chr16:87678144-87678165, chr17:1782952-1782957, chr17:3101578-3101590, chr17:3352294-3352309, chr17:6331803-6331836, chr17:10223697-10223714, chr17:33772658-33772689, chr17:42989063-42989088, chr17:45695832-45695914, chr17:80887206-80887244, chr18:56204747-56204768, chr19:4510530-4510560,<h2 style=";text-align:left;direction:ltr">chr19:8148301-8148314, chr19:9362297-9362343, chr19:11227554-11227602, chr19:36237227-36237245, chr19:44352639-44352666, chr19:58131576-58131623, chr19:58213952-58213969, chr19:58572959-58572979, chr2:33623720-33623734, chr2:37579937-37579971, CHR2:71058184-71058226, CHR2:231775094-231775144, CHR2:239184569-239184581, CHR20:744382-744415, CHR20:5904028-5904040, CHR20:52645534-52645541, CHR20:62597666-62597694, CHR21:43557698-43557736, CHR21:46321659-46321677, CHR22:17589209-17589246, chr22:19951207-19951271, chr22:21377301-21377334, chr22:33253280-33253292, chr22:35817553-35817597, chr22:44322922-44322970, chr3:122003757-122003769, chr3:129155451-129155463, chr3:136574501-136574521, chr3:142277536-142277575, chr3:178968634-178968660, chr4:156289900-156289917, chr5:147024476-147024509, chr5:148206440-148206473, chr5:150666933-150666962, chr5:150901613-150901630, chr5:174870150-174870196, chr6:4069133-4069166, chr6:29913201-29913266, chr6:30080231-30080274, chr6:30993533-30993590,<h2 style=";text-align:left;direction:ltr">chr6:31170514-31170528, chr6:31930441-31930462, chr6:33141253-33141280, chr6:36291985-36292007, chr6:167754702-167754721, chr7:4213975-4214023, chr7:21640361-21640405, chr7:27196069-27196113, chr7:30795288-30795331, chr7:55220177-55220202, chr7:100677455-100677523, chr8:142490120-142490166, chr8:145639681-145639726, chr9:117166206-117166246, chr9:125315542-125315557, chr9:134385435-134385436, chr9:136412255-136412296, chrX:23019317-23019346, 3) Genomic regions listed in Table 7: chr1:10431132-10431158, chr1:120057158-120057246, chr1:154832290-154832304, chr1:156846120-156846233, chr1:159409857-159409884, chr1:171168545-171168584, chr1:183616884-183616926, chr1:226573364-226573402, chr1:226589833-226589958, chr1:23885498-23885599, chr1:32672908-32672932, chr1:3743319-3743391, chr1:94544234-94544276, chr10:104386934-104387019, chr10:123194558-123194609, chr10:123199092-123199095, chr10:123275662-123275666, chr10:123335839-123335866, chr10:123346116-123346190, chr10:123396728-123396806, chr10:123406645-123406663, chr10:43611708-43611865, chr10:43615505-43615633, chr10:70332580-70332672, chr11:116703640-116703671, chr11:4928841-4928866, chr11:534197-534242, chr11:5345128-5345170, chr11:5566030-5566051, chr11:63883985-63884027, chr11:69412090-69412124, chr11:8246326-8246343, chr11:85436303-85436352, chr12:113348849-113348870, chr12:12009741-12009874, chr12:12013572-12013612, chr12:12016008-12016089, chr12:12020114-12020170, chr12:12035649-12035664,chr12:121416622-121416650, chr12:121431272-121431300, chr12:121435427-121435475, chr12:121437114-121437221, chr12:121600180-121600253, chr12:132688115-132688137, chr12:133208886-133208979, chr12:133226159-133226196, chr12:133253995-133254083, chr12:18656174-18656225, chr12:40834918-40834955, chr12:4346169-4346177, chr12:4351884-4352027, chr12:4376089-4376091, chr12:4399036-4399087, chr12:4399917-4399970, chr12:4411639-4411683, chr12:4417127-4417232, chr12:56494991-56494998, chr12:6030405-6030437, chr12:69169222-69169316, chr12:69265196-69265278, chr12:69277127-69277165, chr13:21562832-21562948, chr13:25367282-25367301, chr13:32986219-32986340, chr14:102568296-102568367, chr14:104165753-104165927, chr14:105239146-105239192, chr14:105258892-105258893, chr14:23549285-23549319, chr14:35872792-35872926, chr14:65263300-65263347, chr14:96136775-96136794, chr15:40998305-40998342, chr15:41819283-41819322, chr15:41857216-41857303, chr15:41860411-41860490, chr15:67457335-67457485,<h2 style=";text-align:left;direction:ltr">chr15:79310256-79310288, chr15:88488326-88488428, chr15:88549118-88549151, chr15:88646922-88647038, chr15:88667852-88667948, chr15:89398330-89398407, chr15:94945704-94945719, chr16:2138269-2138398, chr16:2138398-2138422, chr16:2812890-2812939, chr16:68857289-68857441, chr16:81819768-81819820, chr16:87678144-87678165, chr16:89806343-89806347, chr16:89849480-89849629, chr16:89858505-89858525, chr17:1782952-1782957, chr17:3101578-3101590, chr17:33772658-33772689, chr17:37832279-37832315, chr17:37834715-37834808, chr17:41616392-41616456, chr17:42989063-42989088, chr17:45695832-45695914, chr17:6331803-6331836, chr17:78599562-78599655, chr17:78820329-78820374, chr17:78865546-78865630, chr17:78896488-78896529, chr17:78897547-78897561, chr17:78921117-78921211, chr17:80887206-80887244, chr18:56204747-56204768, chr19:10267011-10267077, chr19:11227554-11227602, chr19:17937758-17937786, chr19:17955001-17955021, chr19:2226676-2226772, chr19:30253901-30253998, chr19:30255068-30255090,<h2 style=";text-align:left;direction:ltr">chr19:30290349-30290357, chr19:30340381-30340412, chr19:30361995-30362112, chr19:3119184-3119239, chr19:36237227-36237245, chr19:41724820-41724885, chr19:41781493-41781579, chr19:44352639-44352666, chr19:4510530-4510560, chr19:50919797-50919828, chr19:5210622-5210782, chr19:5210762-5210782, chr19:5212380-5212482, chr19:58131576-58131623, chr19:58213952-58213969, chr19:58572959-58572979, chr19:7163154-7163230, chr19:7166376-7166388, chr19:8148301-8148314, chr19:9362297-9362343, chr2:112754828-112754880, chr2:112754943-112755001, chr2:113983937-113984033, chr2:113984503-113984594, chr2:113989236-113989267, chr2:141259283-141259376, chr2:16042003-16042051, chr2:16073257-16073263, chr2:16112814-16112828, chr2:16113594-16113723, chr2:202122956-202122995, CHR2:231775094-231775144, CHR2:239184569-239184581, chr2:29416366-29416481, chr2:29416481-29416615, chr2:29446184-29446202, chr2:29446701-29446721, chr2:29447108-29447253, chr2:33623720-33623734, chr2:37579937-37579971,<h2 style=";text-align:left;direction:ltr">chr2:47800577-47800603, chr2:47852559-47852643, chr2:48010488-48010558, chr2:71058184-71058226, chr20:30729488-30729523, chr20:40714307-40714479, chr20:40714479-40714540, chr20:40714539-40714540, chr20:52645534-52645541, chr20:57478807-57478939, chr20:5904028-5904040, chr20:62597666-62597694, chr20:744382-744415, chr20:9543622-9543681, chr21:42845374-42845383, chr21:42876400-42876447, chr21:43557698-43557736, chr21:46321659-46321677, chr22:17589209-17589246, chr22:17640022-17640045, chr22:19951207-19951271, chr22:21337266-21337325, chr22:21348914-21349037, chr22:21377301-21377334, chr22:24158895-24158899, chr22:29690246-29690345, chr22:33253280-33253292, chr22:35817553-35817597, chr22:44322922-44322970, chr3:122003757-122003769, chr3:12649857-12649937, chr3:129155451-129155463, chr3:136574501-136574521, chr3:138327951-138328016, chr3:142277536-142277575, chr3:178922222-178922274, chr3:178968634-178968660, chr3:178984575-178984679, chr3:178986121-178986203, chr3:178990402-178990462,chr3:183211906-183212026, chr3:36986932-36986992, chr3:71247257-71247304, chr4:106196829-106196951, chr4:143043340-143043404, chr4:143324036-143324094, chr4:156289900-156289917, chr4:1745492-1745500, chr4:1750487-1750584, chr4:1788994-1789044, chr4:1796629-1796636, chr4:1797741-1797852, chr4:187534362-187534375, chr4:187629497-187629538, chr4:54269096-54269173, chr4:54657737-54657790, chr4:55208737-55208788, chr4:55501109-55501195, chr4:55582037-55582068, chr4:55619846-55619859, chr4:55982752-55982784, chr4:56026865-56026914, chr5:147024476-147024509, chr5:148206440-148206473, chr5:149456772-149456811, chr5:149495287-149495395, chr5:150666933-150666962, chr5:150901613-150901630, chr5:174870150-174870196, chr5:176517326-176517461, chr5:176523562-176523597, chr5:176531772-176531857, chr5:176721198-176721272, chr5:180046209-180046344, chr5:180051003-180051118, chr5:180057231-180057293, chr5:231111-231143, chr5:35861068-35861159, chr5:35871190-35871273, chr5:56178111-56178217,<h2 style=";text-align:left;direction:ltr">chr5:57754808-57754851, chr5:67477132-67477234, chr5:67492589-67492652, chr5:67517563-67517646, chr5:67522722-67522851, chr5:67534039-67534057, chr5:67553771-67553827, chr6:117725448-117725578, chr6:117730673-117730819, chr6:152382311-152382325, chr6:167754702-167754721, chr6:26056549-26056708, chr6:29913201-29913266, chr6:30080231-30080274, chr6:30865115-30865204, chr6:30993533-30993590, chr6:31170514-31170528, chr6:31930441-31930462, chr6:32188603-32188642, chr6:32190390-32190484, chr6:33141253-33141280, chr6:36291985-36292007, chr6:4069133-4069166, chr6:41924853-41924931, chr6:42013020-42013049, chr6:42039487-42039542, chr6:42039551-42039666, chr6:42052577-42052667, chr7:100410597-100410657, chr7:100416139-100416250, chr7:100677455-100677523, chr7:116336880-116336947, chr7:116471122-116471227, chr7:21640361-21640405, chr7:27196069-27196113, chr7:30795288-30795331, chr7:4213975-4214023, chr7:55220177-55220202, chr7:55251541-55251648, chr7:6026775-6026942, chr7:6026942-6026988,<h2 style=";text-align:left;direction:ltr">chr7:78119109-78119199, chr8:128700175-128700233, chr8:128713221-128713364, chr8:128889285-128889371, CHR8:142490120-142490166, CHR8:145639681-145639726, chr8:145737636-145737816, chr8:30999122-30999123, chr8:31024638-31024654, chr8:38299624-38299715, chr8:38310910-38311001, chr8:38350292-38350315, chr8:38361379-38361430, chr8:90958422-90958530, chr9:117166206-117166246, chr9:125315542-125315557, chr9:134385435-134385436, chr9:136412255-136412296, chr9:139401504-139401577, chr9:139403268-139403280, chr9:139405093-139405261, chr9:139410424-139410589, chr9:139411714-139411880, chr9:21968159-21968199, chr9:5408242-5408358, chr9:5415025-5415111, chr9:5420254-5420266, chr9:5458035-5458095, chr9:5484100-5484203, chr9:87478135-87478172, chr9:93639846-93639973, chr9:93641175-93641199, chr9:98238358-98238379, chrX:23019317-23019346。,

12. 6. The method of claim 5, wherein the sample comprises DNA from a tumor or liquid biopsy.

13. 6. The method of claim 5, wherein the sample comprises DNA extracted from a formalin-fixed, paraffin-embedded block, slide, or curl.

14. 13. The method of claim 12, wherein the liquid biopsy is derived from amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, chyme, endolymph, perilymph, stool, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal mucus and sputum), pericardial fluid, ascites, pleural fluid, pus, mucosal secretions, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomit, prostatic fluid, nipple aspirate, tears, sweat, buccal specimen collection, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluid.

15. 13. The method of claim 12, wherein the sample is derived from circulating tumor cells.

16. 6. The method of claim 5, wherein the calibration comprises analysis of candidate microhaplotypes in multiple samples obtained from humans of different ethnicities.

17. 10. The method of claim 1, wherein the candidate microhaplotype comprises a SNP set having at least three, four, or more sets of SNP sequence variants.

18. 2. The method of claim 1, wherein the region of interest is within a gene, an intron, and / or an exon, or between genes.

19. 10. The method of claim 1, wherein the region of interest is within an exome.

20. The method of claim 1 , further comprising isolating DNA containing the candidate microhaplotype.

21. The method of claim 1 , wherein the genome is of human origin.

22. 10. The method of claim 1, further comprising assessing sample contamination by analyzing the median, mean, or other measure of microhaplotype frequency of haplotypes within a SNP set with at least three or four microhaplotypes.

23. 2. The method of claim 1, wherein the microhaplotype frequencies are calculated using only common genotypes found in the population used in the method.

24. 24. The method of claim 23, wherein the common genotype is present at greater than 1% in 1000 Genomes™ or other database.

25. 10. Use of the method of claim 1 to assess the quality of a sample from a particular source, from a vendor, or from a technician preparing or sequencing the sample.

26. 1. A method for detecting a set of single nucleotide polymorphisms (SNPs) having at least three microhaplotypes from a plurality of subjects present in a sample, the method comprising: a) i) identifying regions of interest in the genome; ii) determining additional border regions for the region of interest, the border regions comprising flanking regions of the region of interest, the flanking regions of the region of interest comprising less than 200 nucleotide base pairs; iii) detecting single base pair substitutions (SBS) within the target region and the additional border regions, thereby generating a plurality of sets of sequence variants; iv) filtering the set of variants based on a frequency threshold for detected SBS; and v) analyzing each remaining set of variants for linkage disequilibrium to identify candidate microhaplotypes, wherein the linkage disequilibrium is assessed using one or more population-wide genomic databases, and the candidate microhaplotypes are identified based on the presence of at least three microhaplotypes, wherein the third and more microhaplotypes have a combined frequency of greater than 1% when compared to data from one or more population-wide genomic databases; identifying microhaplotypes in the genome in the sample, comprising: b) determining the number of SNP sets in the sample that have at least three microhaplotypes; c) quantifying the frequency of SNP sets with more than two microhaplotypes to detect the presence of DNA from multiple subjects in the sample, thereby detecting DNA from multiple subjects in the sample; The method comprising:

27. 27. The method of claim 26, further comprising isolating DNA containing the microhaplotype from the sample.

28. 27. The method of claim 26, wherein the flanking regions of the region of interest comprise fewer than 50, fewer than 100, fewer than 150, or fewer than 180 nucleotide base pairs sequenceable by a short-read sequencer.

29. 27. The method of claim 26, wherein the sample comprises DNA from a tumor or liquid biopsy.

30. 30. The method of claim 29, wherein the liquid biopsy is derived from amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, chyme, endolymph, perilymph, stool, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal mucus and sputum), pericardial fluid, peritoneal fluid, pleural fluid, pus, mucosal secretions, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomit, prostatic fluid, nipple aspirate, tears, sweat, buccal specimen collection, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluid.

31. 30. The method of claim 29, wherein the sample is derived from circulating tumor cells.

32. 27. The method of claim 26, wherein a SNP set with more than two microhaplotypes from two or more subjects is detected.

33. 27. The method of claim 26, wherein the sample comprises maternal DNA and fetal DNA.

34. 34. The method of claim 33, further comprising distinguishing the fetal DNA from the maternal DNA.

35. 35. The method of claim 34, further comprising assessing the presence of DNA other than the maternal DNA and the fetal DNA.

36. 27. The method of claim 26, wherein the subject is a human.