Methods and systems for genetic analysis

JP2025111575A5Pending Publication Date: 2025-11-07PERSONAL GENOME DIAGNOSTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025068897
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-04-22
Filing Date
2025-04-18
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing genetic analysis methods struggle to accurately detect DNA contamination and quantify its presence or absence, particularly in samples with low-quality or low-yield DNA, due to the overlap of minor allele frequencies (MAF) with contamination levels, especially in complex DNA mixtures.

Method used

The method utilizes microhaplotypes associated with single nucleotide polymorphisms (SNPs) that are single base pair substitutions (SBS) to identify and quantify the frequencies of multiple haplotypes, enabling detection of DNA contamination and genetic markers in samples, independent of minor allele frequency.

Benefits of technology

Enhances the accuracy of DNA contamination detection and quantification, as well as the identification of genetic markers for diseases, by using microhaplotypes, which are less prone to errors and can deconvolute complex DNA mixtures without prior knowledge of individual profiles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide methods of genetic analysis which utilize microhaplotypes that are associated with SNPs that are single base pair substitutions (SBSs) in preference to insertion or deletion SNPs.SOLUTION: Provided is a method of identifying microhaplotypes in a genome comprising: a) identifying a target region of the genome; b) detecting single base pair substitutions (SBSs) within the target region thereby generating multiple sequence variant sets; c) analyzing each variant set for linkage disequilibrium to identify candidate microhaplotypes; and d) identifying candidate microhaplotypes.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. Patent Application No. 62 / 837,034, filed Apr. 22, 2019, the entire content of which is incorporated herein by reference.

[0002] Field of the Invention The present invention generally relates to genetic analysis, and more particularly, to methods and systems for performing microhaplotype analysis to determine genetic identity in complex DNA mixtures.

Background Art

[0003] Background Information Sequence diversity in the human genome forms the basis in human identification and forensic applications. Genetic fingerprinting is a forensic technique used to identify an individual based on the characteristics of the individual's genetic information (e.g., RNA, DNA). A genetic fingerprint is a small set of one or more nucleic acid diversities, which is likely to be different in all individuals who are not related by blood, and thus is unique to an individual, similar to a fingerprint.

[0004] Sequence diversity is useful in genetic analysis for many applications such as detection of contamination in biological samples, forensic analysis, disease detection, and population genetics. Single - nucleotide polymorphisms (SNPs) have long been used for genetic analysis for such applications.

[0005] DNA contamination in biological samples is a widespread problem. Contamination can occur at almost every stage of sample collection / processing. For example, slides can be contaminated during cutting, liquids can be inadvertently transferred between tubes, libraries can be mixed, and sample barcodes can be impure or have low - quality sequences. Contamination is even more likely to be prominent in samples with low - yield and / or low - quality DNA.

[0006] SNPCheck (trademark) is a tool for batch checking the presence of SNPs and can be used to confirm the presence of DNA contamination in a sample. In "well-behaved" DNA such as normal tissue or cfDNA, since the minor allele frequency (MAF) is almost all 0 or around 0.5, SNPCheck (trademark) can provide reasonable results. However, extremely high contamination levels can be missed because the MAF can be very high and approach 0.5. Tumor DNA, due to extreme copy number diversity, can have an MAF in the range of 0.02 to 0.98 and thus does not "behave well". This means that the MAF of contamination and the actual variant can overlap significantly.

[0007] In order to be able to detect DNA contamination and further accurately quantify the amount of contamination, a detection method that is independent of or almost independent of MAF is required. SUMMARY OF THE INVENTION

[0008] The present disclosure provides a method of genetic analysis that utilizes microhaplotypes associated with single nucleotide polymorphisms (SNPs) that are single base pair substitutions (SBS) rather than insertion SNPs or deletion SNPs. Analysis of such microhaplotypes is particularly useful in forensic genetic applications, sample contamination analysis, and disease analysis.

[0009] In one embodiment, the present disclosure provides a method of genetic analysis comprising: a) identifying a set of SNPs having at least three microhaplotypes in a sample; and b) quantifying the frequencies of the haplotypes within the set of SNPs having more than two microhaplotypes.

[0010] In another embodiment, the present disclosure provides a method of genetic analysis comprising: a) identifying a set of SNPs having at least three microhaplotypes in a sample; and b) quantifying the frequencies of the haplotypes within the set of SNPs having more than two microhaplotypes to determine the presence or absence of DNA contamination in the sample.

[0011] In yet another embodiment, the present disclosure provides a method of genetic analysis comprising: a) identifying a SNP set having at least three microhaplotypes in a sample; and b) quantifying the frequencies of the haplotypes within the SNP set having more than two microhaplotypes to determine the presence or absence of a genetic marker indicative of a disease or disorder.

[0012] In yet another embodiment, the present disclosure provides a method of identifying microhaplotypes in a genome. The method comprises: a) identifying a target region of the genome; b) detecting SBS within the target region, thereby generating a plurality of sets of sequence variants; c) analyzing each set of variants for linkage disequilibrium to identify candidate microhaplotypes; and d) identifying the candidate microhaplotypes.

[0013] In another embodiment, the present disclosure provides a method for detecting a SNP set having at least three microhaplotypes derived from a plurality of subjects present in a sample. The method comprises: a) identifying microhaplotypes in the genome in the sample; b) determining the number of SNP sets having at least three microhaplotypes in the sample; and c) quantifying the frequencies of the haplotypes within the SNP sets with more than two microhaplotypes to determine the presence of DNA derived from the plurality of subjects in the sample, thereby detecting DNA derived from the plurality of subjects in the sample. In one embodiment, identifying comprises: i) identifying a target region of the genome; ii) detecting SBS within the target region, thereby generating a plurality of sets of sequence variants; and iii) analyzing each set of variants for LD to identify the microhaplotypes.

[0014] In one embodiment, the present disclosure provides a method for detecting a set of SNPs having at least two microhaplotypes derived from a plurality of analytes present in a sample. The method includes: a) determining the presence or absence of a set of SNPs having more than two microhaplotypes in the sample, wherein the set of SNPs includes a plurality of single nucleotide substitutions and corresponds to the genomic regions described in Tables 5, 6, and 7; and b) quantifying the frequencies of the haplotypes within the set of SNPs to determine the presence of DNA derived from a plurality of analytes in the sample, thereby detecting a set of SNPs having more than two microhaplotypes derived from a plurality of analytes in the sample.

[0015] In one embodiment, the present disclosure provides an oligonucleotide panel. The panel includes oligonucleotides for amplifying or hybrid-capturing genomic regions corresponding to one or more of the genomic regions described in Tables 5, 6, and 7.

[0016] In another embodiment, the present disclosure provides a method for genetic analysis including: a) amplifying a genomic region present in a sample, wherein the region corresponds to the genomic regions described in Tables 5, 6, and 7, and generating an amplicon by the amplification; and b) sequencing the amplicon to determine the nucleic acid sequence of the amplicon.

[0017] In a further embodiment, the present disclosure provides a method for detecting a disease or disorder in a subject. The method includes: a) obtaining a sample from the subject; b) identifying microhaplotypes in DNA molecules present in the sample; c) determining the presence or absence of a set of SNPs having more than two microhaplotypes in the sample; and d) quantifying the frequencies of the haplotypes within the set of SNPs to determine the presence or absence of a genetic marker indicative of the disease or disorder, thereby detecting the disease or disorder. In one embodiment, the identifying includes: i) identifying a target region, wherein the target region is associated with the disease or disorder; ii) detecting SBS within the target region, thereby generating a plurality of sets of sequence variants; and iii) analyzing each set of variants for LD to identify microhaplotypes.

[0018] In certain embodiments, the present disclosure provides a genetic analysis system. The system includes: a) at least one processor operably connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in a DNA sample; and c) an analysis component executed by the at least one processor, the analysis component being configured to: i) identify microhaplotypes in the sample based on the presence of single nucleotide substitutions; ii) verify the presence of the number of sets of SNPs for the microhaplotypes in the DNA sample; and iii) quantify the frequencies of the genotypes within the set of SNPs accompanied by more than two microhaplotypes in the DNA sample.

[0019] In related embodiments, the present disclosure provides a genetic analysis system configured to execute the methods of the present disclosure. The system includes: a) at least one processor operably connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in a DNA sample; and c) an analysis component executed by the at least one processor and configured to execute the methods of the present disclosure.

[0020] In yet another embodiment, the present invention provides a non-transitory computer-readable storage medium encoded with a computer program. The program includes instructions that, when executed by one or more processors, cause the one or more processors to perform operations for executing the methods of the present disclosure.

[0021] In yet another embodiment, the present invention provides a computing system. The system includes a memory and one or more processors coupled to the memory, and the one or more processors are configured to perform operations for implementing the methods of the present disclosure. [Invention 1001] A method for identifying microhaplotypes in a genome, comprising: a) identifying a target region of the genome; b) detecting single nucleotide substitutions (SBS) within the target region to generate a plurality of sets of sequence variants; c) analyzing each set of variants for linkage disequilibrium to identify candidate microhaplotypes; d) identifying candidate microhaplotypes. The method as described above. [Invention 1002] The method of Invention 1001, further comprising detecting SBS in the flanking regions of the target region. [Invention 1003] The method of the present invention 1002, wherein the flanking region of the target region comprises less than about 50, less than about 100, less than about 150, less than about 180, or less than about 200 nucleotide base pairs that can be sequenced by a short-read sequencer. [The present invention 1004] The method of the present invention 1002, wherein the flanking region of the target region comprises less than about 10,000 nucleotide base pairs that can be sequenced by a long-read sequencer. [The present invention 1005] The method of the present invention 1001, wherein the target region of a) has SBS at a frequency between about 10% and 90%. [The present invention 1006] The method of the present invention 1002, wherein the flanking region of the target region has SBS at a frequency of about 5% to 95%. [The present invention 1007] The method of the present invention 1001, further comprising calibrating a cut-off value for a candidate microhaplotype to evaluate sample contamination. [The present invention 1008] The method of the present invention 1006, wherein only DNA sequence reads overlapping with the candidate microhaplotype are used, and a threshold for contamination detection and the degree of contamination are calculated. [The present invention 1009] The method of the present invention 1008, wherein the DNA sequences used to calibrate the threshold for contamination detection and the degree of contamination are mixed in silico in pairs by alternately using each DNA sequence as a primary sample and a contaminant. [The present invention 1010] The method of the present invention 1008 or 1009, wherein the number and genotype of SNP sets with one and / or two microhaplotypes are compared between different individuals to evaluate identity or contamination. [The present invention 1011] The method of the present invention 1007, further comprising evaluating sample contamination using a cut-off value determined for the frequency of a candidate microhaplotype having a single nucleotide polymorphism (SNP) set with at least three microhaplotypes. [The present invention 1012] The method of the present invention 1011, further comprising evaluating sample contamination using a cut-off value determined for the frequency of candidate microhaplotypes having a SNP set with at least four or more microhaplotypes. [The present invention 1013] The method of the present invention 1001, wherein the candidate microhaplotype corresponds to one or more genomic regions selected from those described in Table 5, Table 6, or Table 7. [The present invention 1014] The method of the present invention 1007, wherein the sample contains DNA derived from a tumor or a liquid biopsy. [The present invention 1015] The method of the present invention 1007, wherein the sample contains DNA extracted from a formalin-fixed paraffin-embedded block, slide, or curl. [The present invention 1016] The method of the present invention 1014, wherein the liquid biopsy is derived from amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, chyme, endolymph, peripheral lymph, feces, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal mucus and sputum), pericardial fluid, ascites, pleural effusion, pus, mucosal secretions, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomit, prostatic fluid, nipple aspirate fluid, lacrimal fluid, sweat, oral mucosal swab, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluids. [The present invention 1017] The method of the present invention 1014, wherein the sample is derived from circulating tumor cells. [The present invention 1018] The method of the present invention 1007, wherein the calibration includes analysis of candidate microhaplotypes in a plurality of samples obtained from humans of different ethnicities. [The present invention 1019] The method of the present invention 1001, wherein the candidate microhaplotype includes a SNP set having at least three, four, or more sets of SNP sequence variants. [The present invention 1020] The method of the present invention 1001, wherein the target region is within a gene, within an intron, and / or within an exon, or between genes. [The present invention 1021] The method of the present invention 1001, wherein the target region is within an exome. [The present invention 1022] The method of the present invention 1001, further comprising separating DNA containing the candidate microhaplotype. [The present invention 1023] The method of the present invention 1001, wherein the genome is of human origin. [The present invention 1024] The method of the present invention 1001, further comprising evaluating sample contamination by analyzing the median, mean, or other measure of the microhaplotype frequency of haplotypes within an SNP set with at least three or four microhaplotypes. [The present invention 1025] Any of the methods of the present invention, further comprising determining the source of sample contamination by identifying microhaplotypes common or specific to the microhaplotypes of the sample and the contaminant. [The present invention 1026] The method of the present invention 1025, wherein microhaplotype information is stored in a database and it is identified whether a DNA sample is from the same individual or a different individual as compared to a newly / simultaneously sequenced individual. [The present invention 1027] The method of the present invention 1025, wherein microhaplotype information is stored in a database and it is identified whether a particular DNA sample is contaminating other samples as compared to a newly / simultaneously sequenced individual. [The present invention 1028] The method of the present invention 1026 or 1027, wherein the number and genotype of SNP sets with one and / or two microhaplotypes are compared between different individuals and identity or contamination is evaluated. [The present invention 1029] Any of the methods of the present invention, further comprising determining the ethnicity of the sample and the contaminant. [The present invention 1030] The method of the present invention 1001, wherein the frequency of the microhaplotype is calculated using only the common genotypes found in the population used in the method. [The present invention 1031] The method of the present invention 1030, wherein the common genotype is present in more than 1% in 1000 Genomes (trademark) or other databases. [The present invention 1032] Use of the method of the present invention 1001 for assessing the quality of samples from a particular source, from a vendor, or from a technician preparing or sequencing the samples. [The present invention 1033] A method for detecting a single nucleotide polymorphism (SNP) set having at least three microhaplotypes derived from a plurality of subjects present in a sample, comprising: a) i) Identifying a target region of the genome, ii) Detecting single base substitutions (SBS) within the target region, thereby generating a plurality of sets of sequence variants, and iii) Analyzing each set of variants for linkage disequilibrium to identify microhaplotypes to identify microhaplotypes in the genome in the sample; b) Determining the number of SNP sets having at least three microhaplotypes in the sample; c) Quantifying the frequency of SNP sets with more than two microhaplotypes to detect the presence of DNA derived from a plurality of subjects in the sample, thereby detecting DNA derived from a plurality of subjects in the sample, and the method. [The present invention 1034] The method of the present invention 1033, further comprising separating DNA containing the microhaplotype from the sample. [The present invention 1035] The method of the present invention 1033, further comprising detecting SBS in the flanking genomic regions of the target region. [The present invention 1036] The method of the present invention 1035, wherein the flanking region of the target region contains less than about 50, less than about 100, less than about 150, less than about 180, or less than about 200 nucleotide base pairs that can be sequenced by a short-read sequencer. [The present invention 1037] The method of the present invention 1035, wherein the flanking region of the target region contains less than about 10,000 nucleotide base pairs that can be sequenced by a long-read sequencer. [The present invention 1038] The method of the present invention 1033, wherein the target region of i) has SBSs associated with genotypes at a frequency of about 10 - 90%. [The present invention 1039] The method of the present invention 1035, wherein the flanking region of the target region has SBSs associated with genotypes at a frequency of about 5 - 95%. [The present invention 1040] The method of the present invention 1033, wherein the cut-off values of SNP sets associated with two, three, four, or more microhaplotypes are calibrated to evaluate the presence of DNA derived from multiple subjects in the sample. [The present invention 1041] The method of the present invention 1033, wherein the sample contains DNA derived from a tumor or a liquid biopsy. [The present invention 1042] The method of the present invention 1041, wherein the liquid biopsy is derived from amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, chyme, endolymph, peripheral lymph, feces, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal discharge and sputum), pericardial fluid, ascites, pleural effusion, pus, mucosal secretion, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomitus, prostatic fluid, nipple aspirate fluid, lacrimal fluid, sweat, oral mucosal swab, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluids. [The present invention 1043] The method of the present invention 1041, wherein the sample is derived from circulating tumor cells. [The present invention 1044] The method of the present invention 1033, in which an SNP set with more than two microhaplotypes derived from two or more subjects is detected. [The present invention 1045] The method of the present invention 1033, in which the sample contains maternal DNA and fetal DNA. [The present invention 1046] The method of the present invention 1045, further comprising identifying the fetal DNA from the maternal DNA. [The present invention 1047] The method of the present invention 1046, further comprising evaluating the presence of DNA other than the maternal DNA and the fetal DNA. [The present invention 1048] The method of the present invention 1033, in which the subject is human. [The present invention 1049] A method for detecting a single nucleotide polymorphism (SNP) set having at least three microhaplotypes derived from a plurality of subjects present in a sample, comprising: a) determining the presence or absence of an SNP set having more than two microhaplotypes in the sample, wherein the SNP set comprises a plurality of single nucleotide substitutions and corresponds to a genomic region selected from the regions described in Tables 5, 6, and 7; b) quantifying the frequency of the SNP set to determine the presence of DNA derived from a plurality of subjects in the sample, thereby detecting an SNP set having at least three microhaplotypes derived from a plurality of subjects in the sample; [[ID=B]] The method comprising the above. [The present invention 1050] An oligonucleotide panel comprising an oligonucleotide for amplifying or hybrid-capturing a genomic region corresponding to one or more genomic regions containing the SBS set identified in any of the present inventions 1001 to 1006. [The present invention 10B1] I An oligonucleotide panel comprising oligonucleotides for amplifying or hybridizing to capture genomic regions corresponding to one or more genomic regions selected from the regions described in Tables 5, 6, and 7. [Inventive concept 1052] a) Amplifying a genomic region present in a sample, wherein the region corresponds to a genomic region selected from the regions described in Inventive concept 1050, Table 5, Table 6, or Table 7, and generating an amplicon by amplification; b) Sequencing the amplicon to determine the nucleic acid sequence of the amplicon; A method comprising the steps of. [Inventive concept 1053] The method of Inventive concept 1052, further comprising quantifying the number of SNP sets having more than two microhaplotypes present in the sample. [Inventive concept 1054] The method of Inventive concept 1053, further comprising quantifying the number of SNP sets having more than three microhaplotypes present in the sample. [Inventive concept 1055] The method of Inventive concept 1054, further comprising quantifying the number of SNP sets having more than four microhaplotypes present in the sample. [Inventive concept 1056] A method for detecting a disease or disorder in a subject, comprising: a) Obtaining a sample from the subject; b) i) Identifying a target region, wherein the target region is related to a disease or disorder; ii) Detecting single base substitutions (SBS) within the target region, thereby generating a plurality of sequence variant sets; and iii) Analyzing each variant set for linkage disequilibrium to identify microhaplotypes; Identifying microhaplotypes in DNA molecules present in the sample; c) Determining the presence or absence of single nucleotide polymorphism (SNP) sets having more than two microhaplotypes in the sample; d) Quantifying the frequency of the SNP set to determine the presence or absence of a genetic marker indicative of a disease or disorder, thereby detecting the disease or disorder; The method comprising the above. [Inventive Concept 1057] The method of Inventive Concept 1056, wherein the disease or disorder is trisomy 13, 18, or 21. [Inventive Concept 1058] The method of Inventive Concept 1056, wherein the disease or disorder is a gene copy number variation. [Inventive Concept 1059] The method of Inventive Concept 1056, wherein the disease or disorder is a fetal disorder. [Inventive Concept 1060] The method according to any one of Inventive Concepts 1056 to 1059, comparing the frequency of a third microhaplotype in a specific chromosome or chromosomal region with the frequency of the third microhaplotype at other locations in the genome. [Inventive Concept 1061] a) at least one processor operably connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in a DNA sample; c) i) identifying a microhaplotype in a sample based on the presence of a single nucleotide substitution, ii) confirming the presence of the number of SNP sets for the microhaplotype in the DNA sample, and iii) quantifying the frequency of genotypes within SNP sets with more than two microhaplotypes in the DNA sample, an analysis component executed by at least one processor, and A gene analysis system comprising the above. [Inventive Concept 1062] The system of Inventive Concept 1061, wherein the analysis component is further configured to determine the likelihood of the presence of DNA contaminants in the sample. [Inventive Concept 1063] The system of the present invention 1061, wherein the analysis component is further configured to determine the presence or absence of a gene mutation. [The present invention 1064] The system of the present invention 1063, wherein the gene mutation is associated with a disease or disorder. [The present invention 1065] The system of the present invention 1064, wherein the disease or disorder is associated with a gene copy number variation. [The present invention 1066] The system of the present invention 1065, wherein the disease or disorder is trisomy 13, 18, or 21. [The present invention 1067] a) at least one processor operably connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in a DNA sample; c) an analysis component executed by the at least one processor and configured to perform (a)-(d) of the present invention 1001, A gene analysis system comprising. [The present invention 1068] a) at least one processor operably connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in a DNA sample; c) an analysis component executed by the at least one processor and configured to perform (a)-(c) of the present invention 1033, A gene analysis system comprising. [The present invention 1069] a) at least one processor operably connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in a DNA sample; c) an analysis component that is executed by the at least one processor and configured to execute the method of the present invention 1049 or 1052; A gene analysis system comprising: [The present invention 1070] a) at least one processor operably connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in a DNA sample; c) an analysis component that is executed by the at least one processor and configured to execute (b) to (d) of the present invention 1056; A gene analysis system comprising: [The present invention 1071] a) identifying a set of single nucleotide polymorphisms (SNPs) having at least three microhaplotypes in a sample; b) quantifying the frequencies of haplotypes within the set of SNPs with more than two microhaplotypes to determine the presence or absence of DNA contamination in the sample; A method comprising: [The present invention 1072] The method of the present invention 1071, further comprising quantifying the frequencies of haplotypes within a set of SNPs having at least three or four microhaplotypes in the sample to determine the amount of DNA contamination in the sample. [The present invention 1073] The method of the present invention 1071, wherein the sample comprises DNA derived from a tumor or a liquid biopsy. [The present invention 1074] The method of the present invention 1073, wherein the liquid biopsy is derived from amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, chyme, endolymph, peripheral lymph, feces, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal mucus and sputum), pericardial fluid, ascites, pleural effusion, pus, mucosal secretion, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomitus, prostatic fluid, nipple aspirate fluid, lacrimal fluid, sweat, oral mucosal swab, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluid. [The present invention 1075] The method of the present invention 1071, wherein the sample is derived from circulating tumor cells. [The present invention 1076] The method of the present invention 1071, wherein the SNP set includes sequence variants having single nucleotide substitutions. [The present invention 1077] a) Identifying a set of single nucleotide polymorphisms (SNPs) having at least three microhaplotypes in a sample; b) Quantifying the frequencies of haplotypes within the SNP set with more than two microhaplotypes to determine the presence or absence of a genetic marker indicative of a disease or disorder; A method comprising: [The present invention 1078] The method of the present invention 1077, further comprising quantifying the frequencies of haplotypes within an SNP set having at least three or four microhaplotypes in the sample. [The present invention 1079] The method of the present invention 1077, wherein the disease or disorder is a gene copy number variation. [The present invention 1080] The method of the present invention 1079, wherein the disease or disorder is trisomy 13, 18, or 21. [The present invention 1081] The method of the present invention 1077, wherein the disease or disorder is a fetal disorder. [The present invention 1082] Any method of the present inventions 1077 - 1081, increasing the number of SNP sets on a specific chromosome, thereby enhancing the identification of trisomy. [The present invention 1083] The method of the present invention 1082, wherein the specific chromosome is one or more of chromosomes 13, 18, and / or 21. [The present invention 1084] Any method of the present inventions 1077 - 1083, performed early in a woman's pregnancy as compared to the use of conventional methods. [The present invention 1085] A method according to any one of the present inventions 1077 to 1084, which has improved specificity due to low sensitivity to the influence of errors caused by the copy number of the parent body. [The present invention 1086] a) Identifying a set of single nucleotide polymorphisms (SNPs) having at least three microhaplotypes in a sample; b) Quantifying the frequencies of haplotypes within the set of SNPs with more than two microhaplotypes to determine the fetal DNA ratio in the maternal DNA source; A method comprising. [The present invention 1087] The method of the present invention 1086, wherein the maternal DNA source is derived from a biological fluid. [The present invention 1088] The method of the present invention 1086, wherein the maternal DNA source is derived from amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, chyme, endolymph, peripheral lymph, feces, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal mucus and sputum), pericardial fluid, ascites, pleural effusion, pus, mucosal secretions, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomitus, prostatic fluid, nipple aspirate fluid, lacrimal fluid, sweat, oral mucosal biopsy specimen, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluid. [The present invention 1089] A non-transitory computer-readable storage medium encoded with a computer program, wherein when the program is executed by one or more processors, the program includes instructions for causing the one or more processors to perform operations for executing any one of the methods of the present inventions 1001 to 1031, 1033 to 1049, 1052 to 1060, or 1077 to 1088. [The present invention 1090] A computing system including a memory and one or more processors coupled to the memory, wherein the one or more processors are configured to perform operations for executing any one of the methods of the present inventions 1001 to 1031, 1033 to 1049, 1052 to 1060, or 1077 to 1088.

Brief Description of the Drawings

[0022]

Figure 1

Figure 2

Figure 3

Best Mode for Carrying Out the Invention

[0023] Detailed Description of the Invention The present invention is based on innovative methods and systems for genetic analysis of microhaplotypes. Before explaining the constitution and method of the present invention, it should be understood that the present invention is not limited to the specific methods and experimental conditions described, because such constitutions, methods, and conditions may change. Also, it should be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to be limiting, because the scope of the present invention will be limited only in the appended claims.

[0024] As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to "the method" includes one or more methods and / or steps of the kind described herein, which will be apparent to those skilled in the art upon reading the present disclosure and the like.

[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred methods and materials are described below.

[0026] The present disclosure provides innovative methods and systems for genetic analysis utilizing microhaplotypes. This method utilizes SBS SNPs and, in some embodiments, SBS variations in low-error genomic regions. Thereby, it is possible to enhance the accuracy not only in the detection of DNA contamination and diseases but also in forensic analysis. The methods disclosed herein use SBS and do not use STR or insertion / deletion SNPs because the latter have an unacceptably high error rate that affects the detection of low levels of contamination in a sample. Any method of the present disclosure focuses on SNP variants with short genetic distances to each other, and ideally, those variants can be present on a single sequence read. In long-read technology, the distance can be even longer as long as the SNP variant is on a single read. Longer distances can also be used, but using paired reads results in a higher error rate and lower coverage as the distance between the variants increases. Further, certain methods of the present disclosure advantageously utilize a two-step analysis that first detects contamination and then quantifies it. The detection of DNA contamination through the methods disclosed herein depends on the number of microhaplotypes for each SNP set and / or the frequency of the third / fourth haplotypes and does not depend on the MAF of individual SNPs.

[0027] In previous research, the utility of markers based on multiple tightly linked SNPs has been exemplified in anthropology because of their ability to provide plausible explanations for group relationships and recent patterns of human diversity. In addition, multi-allelic SNPs are gaining status as markers suitable for addressing forensic issues such as family / ancestry common groups, pedigree estimation, and individual identification. The Kidd laboratory proposed a new type of genetic marker called microhaplotypes (e.g., "microhaps" or MH) with the aim of complementing current DNA typing tools for forensic medicine and population genetics. These are short segments of DNA (less than 300 nucleotides, thus "micro"), characterized by the presence of two or more tightly linked SNPs representing combinations of three or more alleles (i.e., "haplotypes") within a population. The short distance between SNPs means that the recombination rate between them is extremely low. The level of microhaplotype heterozygosity depends on various factors, including the historical accumulation of allelic variants at different positions within the region of interest, the occurrence of rare crossover events, the manifestation of random genetic drift, and / or selection. Since microhaplotypes are multi-SNP haplotypes, they can provide more information per locus than single SNP markers alone.

[0028] Furthermore, when variants are close to each other on the genome, they tend to be correlated. A set of different SNPs on a single chromosomal allele is called a haplotype (a set of linked SNP alleles that tend to be expressed together (i.e., are statistically related)). Since each individual has two copies of their genome, each person has two haplotypes in the autosomal region. These haplotypes can be different (heterozygous) or identical (homozygous). As described above, a microhaplotype is a short haplotype of about 300 nucleotides or less, or a longer distance in the case of long reads. For the purposes of the methods described herein, microhaplotypes are short enough such that the variants are on the same sequencing read and can thus be unambiguously phased. Most microhaplotypes are not very useful in genetic analysis because only two and only two microhaplotypes have been found to date in a given population. However, the methods of the present invention enable the identification of microhaplotypes that can provide statistically useful information, such as microhaplotypes in which 3, 4, 5, or more different haplotypes are found between different individuals (although no individual will have more than two haplotypes).

[0029] As used herein, a "SNP" is a single nucleotide substitution at a specific location in the genome, i.e., a locus, where one base (e.g., cytosine, thymine, uracil, adenine, or guanine) is replaced by another base, and this substitution is present to an extent that is evaluable within a population (e.g., more than 1% of the population).

[0030] In certain embodiments, the methods of the disclosure relate to determining and quantifying the presence of DNA contamination in a DNA sample.

[0031] In related embodiments, the methods of the present disclosure relate to determining whether a sample contains a complex mixture of DNA from multiple individuals. Such individuals may include not only mothers and descendants, but also related or unrelated individuals.

[0032] In conventional forensic analysis, individual DNA samples are uniquely identified through the extraction of short tandem repeats (STRs) and / or determination of mitochondrial DNA (mtDNA) sequences. Capillary electrophoresis is often used to quantify the length of STRs and the sequence of mtDNA. This methodology has proven to be accurate for individual profile identification.

[0033] What is important for the methods according to the present disclosure is that the ability of these methods to deconvolute complex DNA mixtures into component profiles does not require any prior knowledge about the components. For example, the methods described herein are effective in deconvoluting complex DNA mixtures into component profiles and do not require knowledge of any individual or genetic markers or DNA sequences belonging to any component contributing to any one of the complex DNA mixtures. Thus, one of the excellent characteristics of the methods of the present disclosure is that the method does not require any prior knowledge or data about the individual profiles, donors, or components of the complex DNA mixture.

[0034] In some aspects, the techniques described herein can be used to determine the ethnicity of an individual associated with DNA present in a biological sample.

[0035] In embodiments, the present disclosure provides methods for identifying microhaplotypes in the genome. Microhaplotypes are useful in any of the methods disclosed herein, for example, in detecting sample contamination, disease analysis, and / or deconvolution of complex samples.

[0036] Accordingly, the present disclosure provides a method for identifying microhaplotypes in a genome. The method includes: a) identifying a target region of the genome; b) detecting SBS within the target region, thereby generating a plurality of sets of sequence variants; c) analyzing each set of variants for LD to identify candidate microhaplotypes; and d) identifying candidate microhaplotypes.

[0037] Also provided is a method including: a) identifying a set of SNPs having at least three microhaplotypes in a sample; and b) quantifying the frequencies of haplotypes within the set of SNPs with more than two microhaplotypes.

[0038] In addition, the present disclosure also provides a method including: a) identifying a set of SNPs having at least three microhaplotypes in a sample; and b) quantifying the frequencies of haplotypes within the set of SNPs with more than two microhaplotypes to determine the presence or absence of DNA contamination in the sample.

[0039] Also provided is a method of genetic analysis including: a) identifying a set of SNPs having at least three microhaplotypes in a sample; and b) quantifying the frequencies of haplotypes within the set of SNPs with more than two microhaplotypes to determine the presence or absence of a genetic marker indicative of a disease or disorder.

[0040] In various embodiments, the methodology of the present disclosure may further include quantifying the frequencies of sets of SNPs having at least 3, 4, 5, 6, or more microhaplotypes in a sample. This may be performed to determine the amount of DNA contamination in the sample. In an embodiment, as discussed in Example 1, the method further includes calibrating a cut-off value for candidate microhaplotypes. Sample contamination can be evaluated using the cut-off value determined for the frequency of candidate microhaplotypes having a set of SNPs with at least 3, 4, 5, 6, 7, 8, or more microhaplotypes.

[0041] The microhaplotypes of the present invention can use different SNP sets, but the principle for selecting them is the same. As considered herein, this principle is: to select candidate SNPs, use databases such as gnomAD™ (for exons, about 52% Europeans, 7% East Asians, 6% Africans), and to evaluate LD, use the 1000 Genomes™ database (about 20% Europeans, 20% East Asians, 26% Africans), etc.; to select the final set of SNPs based on the 1000 Genomes frequency (or a similar database) of the third / fourth haplotypes to equalize variation among ancestries (using the gnomAD database results in slightly greater variation among Europeans); variants must be close enough to be on the same sequence read; avoid repetitive sequences / indels, use single nucleotide substitutions, and minimize error rates; avoid homopolymers and low-confidence sequence regions; select SNPs in low LD such that the frequency of the third / fourth haplotypes is high; maximize the distance between SNP sets so that the information is independent; and examine the candidate SNP set against actual samples to ensure high coverage, diverse genotypes, and a low rate of the third / fourth haplotypes in pure samples.

[0042] The methodology of the present disclosure may include, as considered in Example 1, the identification of a candidate variant set for analysis.

[0043] This may include identifying a target region of the genome and determining the nucleotide sequence of that region for use in the analysis. The target region is examined for the presence of SBSs. In embodiments, the SBS frequency is typically between about 5-95%, which may be determined using a suitable genomic database, such as the gnomAD™ database (gnomad.broadinstitute.org / ).

[0044] In embodiments, the target region utilized optionally includes a flanking region, and this flanking region is also examined for the presence of SBS, and its frequency is determined to be between about 5% and 95%. In various embodiments, the flanking region of the target region includes less than about 50, less than about 100, less than about 150, less than about 180, or less than about 200 nucleotide base pairs. In various embodiments, the total length of the target region, optionally including the flanking region, is less than about 500, less than about 450, less than about 400, less than about 350, less than about 300, less than about 250, less than about 200, less than about 150, less than about 100, less than about 90, less than about 80, less than about 70, less than about 60, less than about 50, less than about 40, less than about 30, less than about 20, or less than about 10 base pairs.

[0045] In embodiments, the identified candidate variant pairs are then examined for LD. This may be performed using the 1000 Genomes™ database (ldlink.nci.nih.gov / ?tab=ldhap).

[0046] Pairs, triplets, quartets, and the like that have at least three haplotypes and in which the third and any additional haplotypes have a combined frequency greater than 1% are then considered as candidate uses. In various embodiments, the set of microhaplotype variants was selected to avoid insertions / deletions because the inherent sequencing error rate in such variants increases and the likelihood of generating noise increases. In some embodiments, it may not be possible to easily evaluate LD because the variant may not be present in the 1000 Genomes™ database. However, such variants may be used if it is suggested by the MAF found in the gnomAD™ database that it is appropriate.

[0047] It will be understood that the target region may be within a gene, intron, and / or exon, or between genes. Alternatively, the target region may be within the exome. In an embodiment, the target region may include a gene marker associated with a disease. In an embodiment, the target region may include a gene marker associated with a specific ethnicity.

[0048] Using this approach, an oligonucleotide panel may be generated to amplify or hybrid capture a specific region containing the microhaplotype identified using the methods of the present disclosure. In one embodiment, the oligonucleotide panel includes oligonucleotides for amplifying or hybrid capturing a genomic region corresponding to one or more of the genomic regions set forth in Table 5. In another embodiment, the oligonucleotide panel includes oligonucleotides for amplifying or hybrid capturing a genomic region corresponding to one or more of the genomic regions set forth in Table 6 or 7.

[0049] Thus, the present disclosure also provides a method for gene analysis, including: a) amplifying a genomic region present in a sample, wherein the region corresponds to the genomic regions set forth in Table 5, Table 6, and Table 7, and generating an amplicon by amplification; and b) sequencing the amplicon to determine the nucleic acid sequence of the amplicon.

[0050] As discussed herein, the microhaplotypes identified by the methods of the present disclosure may be utilized in various applications, including, but not limited to, DNA contamination detection, disease analysis, and sample deconvolution (i.e., detection of DNA derived from multiple subjects or cell types in a single sample).

[0051] In one embodiment, the present disclosure provides a method for detecting an SNP set having at least three microhaplotypes derived from a plurality of analytes present in a sample. The method includes: a) identifying microhaplotypes in the genome of the sample; b) determining the number of SNP sets having at least three microhaplotypes in the sample; and c) quantifying the frequency of SNP sets accompanied by more than two microhaplotypes to determine the presence of DNA derived from a plurality of analytes in the sample, thereby detecting DNA derived from a plurality of analytes in the sample. In one embodiment, the identifying includes: i) identifying a target region of the genome; ii) detecting SBSs within the target region, thereby generating a plurality of sequence variant sets; and iii) analyzing each variant set for LD to identify microhaplotypes.

[0052] In another embodiment, the present disclosure provides a method for detecting an SNP set having at least three microhaplotypes derived from a plurality of analytes present in a sample. The method includes: a) determining the presence or absence of an SNP set having at least three microhaplotypes in the sample, the SNP set including a plurality of single nucleotide substitutions and corresponding to the genomic regions described in Tables 5, 6, and 7; and b) quantifying the frequency of the SNP set to determine the presence of DNA derived from a plurality of analytes in the sample, thereby detecting an SNP set having at least three microhaplotypes derived from a plurality of analytes in the sample.

[0053] Thus, the methods of the present disclosure for deconvoluting or resolving components from a complex DNA mixture may be carried out by analyzing a single complex DNA mixture. In certain embodiments of the methods of the present disclosure for deconvoluting or resolving components from a complex DNA mixture, the method may analyze two or more complex DNA mixtures. The resolution of DNA profiles using these methods increases as the number of SNP loci increases in the panel used. As used herein, the term complex DNA mixture refers to a DNA mixture that contains DNA from two or more donors. Preferably, the complex DNA mixtures of the methods described herein contain DNA from at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more donors.

[0054] The methods of the present disclosure are superior to existing methods for deconvoluting DNA profiles. It should be noted that the applications of the methods described herein are not limited to forensic analysis or DNA contamination detection scenarios. For example, the methods of the present disclosure may be used in medical diagnosis and / or prognosis. To detect a disease, the target region may be selected to include genetic markers associated with a disease or condition such as cancer or fetal disorder. In this method, the target region may be, for example, on chromosome 21 that enables the diagnosis of trisomy 21, also known as Down syndrome. When it is determined that the sample is from the mother and the fetus and the frequency of a third microhaplotype is different on chromosome 21 compared to other chromosomes, this indicates a variation in gene copy, such as trisomy 21. Other trisomies, including trisomy 13 and trisomy 18, can be detected in a similar manner.

[0055] Thus, the methods described herein may be used in various ways to predict, diagnose, and / or monitor diseases such as cancer and fetal disorders. Further, the method may be utilized to distinguish between various cell types.

[0056] In the field of cancer, biopsy samples often contain many cell types, and only a very small fraction of them may form any part of the tumor. As a result, the DNA obtained from a tumor biopsy is another form of a complex DNA mixture and may contain somatic variants that occur on specific DNA molecules. In the case of somatic diversity, the restrictions on SBS can be relaxed because somatic diversity can be an indel or another modification that may otherwise be avoided. Furthermore, within a tumor, a large number of cells may be molecularly different with respect to, for example, the expression of factors that indicate or promote angiogenesis and / or metastasis. The DNA mixture obtained from a tumor sample may also form the complex DNA mixture of the present disclosure. In both of these non-limiting examples, the methods of the present disclosure may be used to construct an individual profile for each cell or cell type contributing to the complex DNA mixture. Furthermore, the methods of the present disclosure may be used to deconvolute the contributors to the complex DNA mixture. As an example, an individual profile of malignant cells may be constructed using the complex DNA mixture obtained from a breast cancer tumor biopsy. In the same patient, in a brain cancer tumor biopsy, this individual profile may be used to deconvolute the contributors to the complex DNA mixture obtained from the brain cancer tumor biopsy to determine, for example, whether malignant breast cancer cells from that subject have metastasized to the brain and formed a secondary tumor. This method has the potential to resolve questions as to whether the tumors arose independently or, on the other hand, whether these tumors are related.

[0057] Accordingly, the present disclosure provides a method for detecting a disease or disorder in a subject. The method includes: a) obtaining a sample from the subject; b) identifying microhaplotypes in DNA molecules present in the sample; c) determining the presence or absence of a set of SNPs having more than two microhaplotypes in the sample; and d) quantifying the frequencies of the haplotypes within the set of SNPs to determine the presence or absence of a genetic marker indicative of the disease or disorder, thereby detecting the disease or disorder. In one embodiment, identifying includes: i) identifying a target region, wherein the target region is associated with the disease or disorder; ii) detecting SBSs within the target region, thereby generating a plurality of sets of sequence variants; and iii) analyzing each set of variants for LD to identify microhaplotypes.

[0058] In various embodiments, the genome is present in a biological sample collected from a subject. The biological sample can be virtually any type of biological sample, particularly a sample containing DNA. The biological sample can be a germline, stem cell, reprogrammed cell, cultured cell, or a tissue sample containing from 1000 to about 10,000,000 cells, or a fluid containing circulating DNA. In embodiments, the sample is a tumor or a liquid biopsy, such as amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, chyme, endolymph, peripheral lymph, feces, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal mucus and sputum), pericardial fluid, ascites, pleural effusion, pus, mucosal secretions, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomit, prostatic fluid, nipple aspirate fluid, lacrimal fluid, sweat, oral mucosal swab, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluids, etc., but not limited thereto, and contains DNA derived therefrom. In one embodiment, the sample contains DNA derived from circulating tumor cells. In embodiments utilizing an amplification protocol such as PCR, it is possible to obtain a sample containing a large number of cells, even if it is a single cell. This sample need not contain intact cells as long as it contains sufficient biological material (e.g., DNA) to perform genetic analysis of one or more regions of the genome.

[0059] In some embodiments, the biological sample or tissue sample can be collected from any tissue containing cells with DNA or from a fluid with circulating DNA. The biological sample or tissue sample may be obtained by surgery, biopsy, mucosal swab, feces, or other collection methods. In some embodiments, the sample is derived from blood, plasma, serum, lymph, nerve cell-containing tissue, cerebrospinal fluid, biopsy material, tumor tissue, bone marrow, nerve tissue, skin, hair, tears, urine, fetal material, amniocentesis material, uterine tissue, saliva, feces, or sperm. Methods for separating PBLs from whole blood are well known in the art.

[0060] As disclosed above, the biological sample can be a blood sample. The blood sample can be obtained using methods known in the art such as finger prick or venipuncture. Preferably, the blood sample is about 0.1 to 20 ml, or about 1 to 15 ml, and the volume of the blood is about 10 ml. Similar to the circulating free DNA in blood, a small amount can be used. Micro-sampling, and needle biopsy, catheter, sampling by excretion or production of body fluids containing DNA are also potential sources of biological samples.

[0061] In the present invention, the subject can be typically a human, but can also be any species including but not limited to dogs, cats, rabbits, cows, birds, rats, horses, pigs, or monkeys.

[0062] The methods of the present disclosure utilize nucleic acid sequence information and thus can include any method of performing nucleic acid sequencing, including nucleic acid amplification, polymerase chain reaction (PCR), nanopore sequencing, 454 sequencing, and tag - based sequencing. In embodiments, the methodology of the present disclosure utilizes systems from Illumina, Inc. (including, but not limited to, HiSeq™ X10, HiSeq™ 1000, HiSeq™ 2000, HiSeq™ 2500, Genome Analyzers™, MiSeqTM, NextSeq, NovaSeq systems), Applied Biosystems Life Technologies (SOLiD™ System, Ion PGM™ Sequencer, ion Proton™ Sequencer), or Genapsys, or BGI MGI, and other systems. Further, nucleic acid analysis can be performed by systems provided by Oxford Nanopore Technologies (GridiON™, MiniON™) or Pacific Biosciences (Pacbio™ RS II or Sequel I or II). Importantly, in embodiments, sequencing can be performed using any of the methods described herein. When long - read technologies such as PacBio™ or Oxford Nanopore™ are used, the length restrictions imposed on DNA are relaxed, and SNPs can be made further apart to match the longer read lengths.

[0063] The present invention includes a system for performing the steps of the disclosed method and is described, in part, from the perspective of functional components and various processing steps. Such functional components and processing steps may be implemented by any number of components, operations, and techniques configured to perform the specified functions and achieve various results. For example, the present invention may employ various biological samples, biomarkers, elements, materials, computers, data sources, storage systems and media, information collection techniques and processes, data processing criteria, statistical analysis, regression analysis, and the like, which may perform various functions.

[0064] The method of genetic analysis according to various aspects of the present invention may be implemented in any suitable manner, for example, using a computer program operating on a computer system. An exemplary genetic analysis system according to various aspects of the present invention may be implemented in combination with a computer system, such as a conventional computer system including a processor and a random access memory, such as a remotely accessible application server, a network server, a personal computer, or a workstation. The computer system may preferably also include additional memory devices or information storage systems, such as a mass storage system and a user interface, such as a conventional monitor, keyboard, and tracking device. However, the computer system may include any suitable computer system and related equipment and may be configured in any suitable manner. In one embodiment, the computer system includes a stand-alone system. In another embodiment, the computer system is part of a network of computers including a server and a database.

[0065] The software necessary for receiving, processing, and analyzing genetic information may be implemented in a single device or in multiple devices. The software may be accessible via a network such that information storage and processing are performed remotely with respect to the user. A genetic analysis system according to various aspects of the present invention, and its various components, provide functions and operations that facilitate genetic analysis, such as data collection, processing, analysis, reporting, and / or diagnosis. For example, in the present embodiment, a computer system may execute a computer program that receives, stores, retrieves, analyzes, and reports information related to the human genome or a region thereof. The computer program may include a plurality of modules that perform various functions or operations, such as a processing module that processes raw data to generate supplementary data, and an analysis module that analyzes the raw data and the supplementary data to generate a quantitative assessment of contamination or a disease state model, and / or diagnostic information.

[0066] The procedures executed by the genetic analysis system may include any suitable steps that facilitate genetic analysis and / or disease diagnosis. In one embodiment, the genetic analysis system is configured to establish a disease state model and / or to determine the disease state of a patient. Determining or identifying a disease state may involve generating any useful information regarding the state of a patient related to a disease, such as making a diagnosis, providing information useful for diagnosis, assessing the stage or progression of a disease, identifying a condition that may indicate susceptibility to the disease, identifying whether further tests are recommended, predicting and / or evaluating the effectiveness of one or more treatment programs, or otherwise assessing the disease state, the likelihood of a disease, or other aspects of the patient's health.

[0067] The gene analysis system is preferably configured to generate a disease model and / or provide a diagnosis for a patient based on gene data and / or additional subject data related to the subject. The gene data may be obtained from any suitable biological sample, not only from a database storing genetic information.

[0068] The following examples are provided to further illustrate the advantages and features of the present invention, but are not intended to limit the scope of the present invention. These examples are typical of those that may be used, but other procedures, methodologies, or techniques known to those skilled in the art may be used instead.

Example

[0069] Example 1 Detection of Sample Contamination In this example, the methodology of the present disclosure was utilized to detect sample contamination. The following provides a detailed consideration of the methods and steps used for detection.

[0070] Identification of Candidate Variant Sets For each target region, the regions to be sequenced, together with additional boundary regions (up to 100 bp), were examined for SBSs with frequencies of 10 - 90% according to the gnomAD™ database (gnomad.broadinstitute.org / ). Once a variant was found in a region that was not a low-confidence region, the 180-bp region adjacent in both directions was examined for SBSs with frequencies of 5 - 95%. These cutoffs may vary depending on the type of sample to be analyzed for various panels and the number of SNP sets required. Such variant pairs were then examined for LD using 1000 genomes data (ldlink.nci.nih.gov / ?tab=ldhap). Pairs, triplets, etc., that had at least three haplotypes and where the third and higher haplotypes had a combined frequency greater than 1% were considered as candidate uses. These cutoffs may potentially be extended to include additional variant sets as needed, or restricted to retain only the most informative variant sets to minimize noise. For example, the variant set was selected to avoid insertions / deletions because such variants have an inherently increased sequencing error rate and a higher likelihood of generating noise. Similarly, other sequence contexts may be advantageous based on the error rate. Additionally, although some variants were not found in the 1000 Genomes™ database and could not be evaluated for LD, they were advanced to candidate testing if their appropriateness was suggested by the MAF observed in gnomAD™. SNPs may theoretically be present as far apart as their paired reads, but for simplicity of analysis, SNPs that were closer together and covered by a single read were selected.

[0071] Characteristic evaluation of candidate variant sets By further evaluating the candidate variant sets in actual samples, sufficient reads with both / all variants on the reads were ensured to generate phased haplotypes. For each SBS, a cutoff of 100 times the median coverage was used so that all or almost all SNP sets could be included in each comparison. High coverage is required to maximize the sensitivity of the analysis. For other panels, the correct set of SBSs will vary depending on the panel to be examined. Additionally, some sequence contexts have higher error rates than others, and using those variants can result in additional artifactual microhaplotypes. Variant sets that tended to have too many third / fourth microhaplotypes in samples considered to be pure were excluded from use because they could generate a high level of noise relative to the signal.

[0072] Based on high coverage and low background noise levels, 106 variant sets were selected for use with the 507-gene panel (Table 5). The distance between SBS sets was maximized to minimize redundant information as much as possible. The MAFs listed for the SBSs in this table were obtained from "All Populations" in the 1000 Genomes™ database and differ from the original MAFs obtained from gnomAD™.

[0073] Estimation of contamination level Since any sample could potentially be contaminated theoretically, it was necessary to characterize the samples before using them for calibration so that the process could be started with pure samples. Furthermore, since the frequencies of variants and microhaplotypes can vary widely depending on ethnicity, it is useful to characterize samples with different ethnicities to ensure that a given set of SBSs will work well with any sample and contaminant. For this dataset, five Africans, five Asians, and six Europeans (all self-identified) were selected based on covering at least 105 / 106 variant sets and having less than two variant sets with more than two microhaplotypes. These samples and their characteristics are shown in Table 1. The European samples have a non-significantly decreased number of single-microhaplotype SBSs.

[0074] (Table 1) Samples used for calibration TIFF2025111575000002.tif113128

[0075] To mimic in-silico contamination, unfiltered fastQ™ reads taken from pure samples were computationally mixed with other samples for the purpose of artificially generating "contaminated" samples. To target X% contamination, in principle, 100 - X% of the reads taken from the original sample were mixed with X% of the reads taken from the "contaminant". These mixed samples were then run through the pipeline and aligned and called using our standard methods. For each sample, the number and frequency of haplotypes in each SBS set were counted and tabulated. Then, for each SBS set, the frequency of the third haplotype was examined within each sample, if any, and the minimum, maximum, median, and mean values were calculated for each set of third haplotype frequencies. The mixtures were then examined to see how well contamination could be predicted by these parameters.

[0076] Prior to examining the results in detail, we considered how multiple technical and biological confounding factors could potentially affect the results. As observed even in "pure" samples, there is technical noise that results in a decrease in the number of the 3rd / 4th haplotypes. To avoid these interfering with contamination detection, we set a minimum number of the 3rd / 4th haplotypes. Since the desired level of contamination detection is in the range of 1 - 2%, we selected the minimum number of the 3rd / 4th haplotypes to be in the range of 5 - 10. This avoids the problem of misaligning low levels of technical noise as contamination.

[0077] (Table 2) Number of SBS sets with more than two microhaplotypes (n = 70 each) TIFF2025111575000003.tif22128

[0078] The percentage of SNPs with more than two microhaplotypes determines whether a sample is contaminated, but this percentage is relatively insensitive to the degree of contamination. Since the % value of more than two microhaplotypes reaches its maximum rapidly, looking at this parameter alone, contaminations of 2% vs 5% vs 20% appear very similar. To avoid this problem, we quantified the contamination level using the MAF for the third haplotype. This value can be misleading at low contamination due to technical artifacts. This value can appear abnormally high due to the possibility that the DNA causing the contamination gives two copies of the third haplotype, so the contamination may appear to be twice as high as it actually is (Figure 3). Also, the extreme copy number diversity often seen in tumor samples also affects apparent contamination in either direction, depending on which haplotype is in excess. This is usually not a problem with normal DNA but can be a serious problem with tumor DNA. To avoid these problems, we used the median MAF for the third haplotype to minimize the influence of either an abnormally high or low MAF. There is additional information found in the allele frequencies for the second and fourth microhaplotypes, but this data was not used in the calculations. If there are sufficient sets available to examine, more complex analyses of haplotype frequencies can be used.

[0079] For samples with a third / fourth haplotype that exceeds a set number, various factors may interfere with the determination of the exact frequency. In this calibration series, one technical challenge is whether the nominal contamination level is actually accurate. The number of reads added can be precisely controlled, but each sample has different characteristics in terms of DNA quality, and this characteristic may affect the contamination level in terms of functionality. In samples with a wide range of DNA lengths, due to differences in DNA quality or capture efficiency resulting in different proportions of targeted reads, the contamination level in terms of functionality will be different because the frequency of the SNP set appearing on the same read depends on the length. This means that 1% of the added reads is equivalent in terms of length and functionality to 0.5% or 2%, or any value in between. For this reason, each sample and its contaminant were swapped simultaneously and in parallel as samples and contaminants. Thus, this normalizes the quality difference to some extent and provides a better estimated result of the contamination level in terms of functionality. When applying these methods to actual samples, considering the possibility of incorrect variant calls, contamination in terms of functionality becomes more important than stoichiometric contamination.

[0080] Regarding the issue of quantification, there are also biological reasons. A pure sample may have one or two microhaplotypes in each SBS set, and one or two microhaplotypes of the contaminating contaminants may or may not match one or two of the microhaplotypes of the primary sample. When the contamination is low and the signal is just starting to appear, the new third haplotype is preferentially composed of double contributions that do not match the microhaplotypes of the sample, while at higher contamination levels, there will be a mixture of single / double contributions. Therefore, it is desirable not to expect a simple linear relationship between the contamination level and the frequencies of various haplotypes. Adding to this difficulty, extensive copy number diversity occurs between tumor samples, which can also greatly affect haplotype frequencies. For these reasons, empirical estimates of contamination were used because simply looking at the frequency of the third haplotype would overestimate low contamination levels and underestimate high contamination levels. If there are more variant sets at a very high coverage level, it may be possible to incorporate frequency data to better estimate contamination in terms of function. As shown in Table 3, using this SNP set and coverage conditions, the area where an appropriate balance of overcounts and undercounts can be achieved and relatively accurate contamination estimates can be obtained is approximately 2%. Since this is approximately the same level as where we want to set the sensitivity, the median of the third haplotype will be used as an approximation of the contamination level, and there may be problems with accuracy if it deviates greatly from 2%. To accurately estimate other contamination levels, more mixtures would need to be examined as was done with other SBS sets.

[0081] (Table 3) Median frequency of the third haplotype by ethnicity TIFF2025111575000004.tif39128

[0082] Application to actual samples Samples used for in-silico contaminant mixtures were selected based on their high quality. Unfortunately, due to much greater variability in actual samples, it is necessary to set criteria for which samples can be analyzed and how their analysis should be performed. Ideally, every sample could potentially have coverage greater than 100× across all 106 SBS sets, but this is often not the case in practice. Absence of an SBS set results in inconsistent comparison results, and low coverage at a particular SBS can lead to significant overestimation or omission of the frequency of a third haplotype. Thus, 1000 samples were run through a standard pipeline to examine microhaplotype data. Of these 1000 samples, 151 samples did not meet standard quality control metrics, leaving 849 samples for microhaplotype analysis. A coverage of at least 20 is required to count an SBS. The majority of samples (709) have data for all 106 SBS sets. However, there are also samples with a significantly low number of SBS sets meeting the minimum criteria. The point at which samples fail more often than those passing other quality control metrics is 100 SBS calls. Thus, for the following analysis, only 825 samples that passed with more than 100 SBS calls were used. Of these 825 samples, 24 failed SNPCheck™, which was previously used to monitor sample contamination.

[0083] Table 4 shows the effect on contamination detection when changing the cut-off for these 825 samples. Samples pass if either more than two microhaplotype SBS sets are less than the cut-off number or the median MAF of the third microhaplotype is below the set threshold. Based on the in-silico experiments above, the number of SBS sets with more than two microhaplotypes should be in the range of 5 - 10 with these microhaplotypes. Additionally, even if there are microhaplotypes more than the cut-off number, samples with a median third haplotype frequency of less than 1.5% are also judged to pass. Using these cut-offs, 804 - 811 samples, including 18 - 19 samples that failed SNPCheck™, pass. When the frequency of the third haplotype is 2 - 4%, samples are randomly checked to confirm whether the contamination level may cause problems based on the observed somatic mutation frequency. 4 - 5 out of these 11 - 18 samples failed SNPCheck™. Samples with a frequency of the third microhaplotype more than 4% failed. In any case, this was 3 samples, one of which failed SNPCheck™. In addition to the above 825 passing runs, SNPCheck™ was also performed on samples that failed other QC metrics or samples with too few SBSs called by the microhaplotype method of the present disclosure. Of the 4 samples that failed both QC and SNPCheck™, 3 failed by the microhaplotype method and the contamination was higher than 10%. 4 out of 7 samples that failed SNPCheck™, which would not be normally evaluated by microhaplotypes with less than 101 called SBSs, also failed by the microhaplotype method regardless of the cut-off, while another one failed at some cut-off values.

[0084] (Table 4) Comparison of Microhaplotype and SNPCheck™ TIFF2025111575000005.tif60166

[0085] A complete match between the method of the present invention and SNPCheck™ was not expected. SNPCheck™ calls pure samples as contaminated, which results in false positives by failing some tumor samples with very high copy number variations. False negatives are also known to occur when the contamination level is very high and its diversity is misidentified as germline diversity.

[0086] Contamination Detection in Exome Since many of the SBSs used in the 507-gene panel are in non-coding regions, they are worthless for exome analysis. Therefore, a new set of SBSs was selected to examine the exome. Since the coverage of the exome is low for each ROI, it is more important to capture variants with as much coverage as possible. Therefore, the SBS set was selected such that the interval between variants is shorter than that of the 507-gene panel and is confined closer to exons. Also, since the number of ROIs is very large, an attempt was made to include more informative SBSs, and the attempt was selected in ROIs having a coverage higher than the average. These were then examined in a set of exome data, and SBSs having a median coverage greater than 80 and diverse haplotypes were selected for use in the panel. These SBS sets are listed in Table 6. Using the same method as above, two exomes suspected of contamination were examined, and using this SBS set, it was found that they were contaminated by more than 15%.

[0087] Using the initial set of microhaplotypes used in the 507 gene panel, differences in sensitivity were observed among different ancestral groups. This issue is likely due to both bias in the databases used to select the microhaplotype sets and differences in heterozygosity among different ancestries. To correct for this, population haplotype frequencies obtained from the 1000 Genomes Project were used to balance the third / fourth haplotype frequencies to be approximately equal across all ancestries. The frequencies of the third / fourth haplotypes among SNP sets were summed, and SNP sets that contributed to excessive frequencies in overrepresented ancestries were dropped. This enabled the generation of a set of microhaplotypes such that the expected average number of third / fourth haplotypes was the same among people with East Asian, African, and European ancestries. However, it was not possible to generate the same frequencies simultaneously for the other two 1000 Genome ancestries, namely admixed Americans and South Asians. Since both of these ancestries had higher frequencies of third / fourth microhaplotypes than the other three ancestries, contamination should be easily detectable using the same threshold as the other ancestries.

[0088] To further improve performance characteristics, an attempt was made to select only microhaplotype sets with high coverage and low noise in pure samples. The minimum average coverage of the SNP sets was increased from 100 to 250. However, high coverage is a double-edged sword. While high coverage improves sensitivity and accuracy, it can also generate artifactual third haplotypes due to inherent sequencing errors, typically at a level of about 0.1%. To minimize the impact of such technical errors, low-frequency haplotypes can be excluded from consideration. The level to set this can be optimized based on coverage and sequencing quality. For this experiment, the threshold was set at 0.2%, where haplotypes with frequencies below 0.2% were considered non-real. Other thresholds can be used depending on sequence quality and other factors.

[0089] In addition, more SNP sets were used to enhance the signal to enable improved accuracy of contamination estimation. Based on these considerations, 164 SNP sets were selected for a second microhaplotype panel that met all of these criteria. Fifty-one of these SNP sets were also present in the first panel, and both sets are shown in Table 7 along with the region, dbSNP number, and 1000 genome frequency of the third / fourth haplotypes.

[0090] As discussed above, generating samples with the correct level of contamination is very difficult. When samples are combined in silico, mixed samples with the correct level of contamination are obtained, but the functional effects are not necessarily accurate. Since microhaplotype detection depends on the length of the sequenced molecule, samples with different DNA qualities but the same partial components will have different effects on microhaplotype frequencies. To minimize this effect, samples were analyzed in pairs, with the "sample" and "contaminant" swapped, and then the results were averaged within each pair. Then, 15 such pairs for each category (African, East Asian, European, admixed) were analyzed for the number of third / fourth microhaplotypes as a function of the contamination level. As shown in Figure 1, the number of third / fourth MHs for individuals of East Asian and European ancestry almost overlapped. The number of third / fourth MHs for individuals of African American ancestry and admixed ancestry was higher than that of East Asians / Europeans but similar to each other. The discrepancy in African Americans is likely due to the composition of the 1000 genomes African panel, which includes five subgroups from Africa and two subgroups from African Americans. These two groups are somewhat admixed and thus generate higher values than the other groups. Further, the combination of more even third / fourth microhaplotype frequencies and a larger number of microhaplotype sets being examined will enable more reliable identification of contaminated samples.

[0091] The number of the 3rd / 4th microhaplotypes varies slightly among different ancestors. Nevertheless, the median frequency of the 3rd microhaplotype as a function of the contamination level is nearly the same among those ancestors, including the mixed samples from different ancestors (Figure 2). This relationship is linear from about 1%. Contamination levels below 1% not only greatly affect sequencing artifacts but also, inadvertently, the possibility of additional DNA causing contamination. Above 1%, the median of the observed frequencies is roughly half of the contamination level. This is as expected based on how the 3rd MH is generated, as shown in Figure 3. At even higher contamination levels, this value begins to decline, which is due to multiple factors, including the possibility that the 3rd microhaplotype actually originates from the sample rather than the contaminant.

[0092] Using the relationship that the contamination level = 2 × the median of the 3rd microhaplotype level, the detection results of the contamination levels at different levels are shown in Table 8 for each ancestor. Their patterns are similar, and when the predicted contamination level is twice the 3rd microhaplotype level, the proportion of samples detected at even higher contamination levels decreases. This table gives guidance on where to set the threshold to achieve nearly 100% detection of contamination at a given level. For example, if you want to detect almost all samples contaminated at 2%, setting a cut-off of the 3rd microhaplotype = 0.75% will detect 97% of the samples contaminated at 2%, while including 82% of the samples contaminated at 1.5%, only 15% of the samples contaminated at 1%, and none of those contaminated at 0.5%. The choice of the threshold can be made based on the relative levels of false positives and false negatives.

[0093] Example 2 Use of Microhaplotypes for NIPT Detection of Chromosomal Abnormalities Non-invasive prenatal testing (NIPT) for detecting chromosomal abnormalities is performed by collecting a blood sample from the mother and evaluating circulating fetal DNA in the presence of a large background proportion of maternal DNA. Typically, sequence reads are simply aligned and the number aligned to each chromosome is counted. If there are excessive reads aligned to chromosomes that are highly sensitive to trisomy (usually chr13, chr18, chr21), a positive diagnosis is made. This test is typically performed after the 10th week when the amount of fetal DNA contained in the mother's blood is sufficient for test accuracy. By using microhaplotypes, the test can be performed earlier because more accurate quantification is possible at even lower DNA concentrations, resulting in more accurate results, which is due to being independent of benign copy number variations already present in the mother that can lead to misinterpretation.

[0094] The behavior of NIPT samples is even more straightforward than that of tumor samples for two reasons. First, the complexity of extensive copy number variations is less likely to be a problem. Second, one of the fetal haplotypes will already be present in the mother, and the third haplotype coming from the father will be only a single copy, so it will not be overcounted at low levels. Thus, a more predictable increase in frequency is expected.

[0095] In most cases of trisomy 21, the extra chromosome is maternally derived, reducing the contribution of the new paternal haplotype on that chromosome. Thus, the paternal haplotype frequency on the unaffected chromosomes can be determined and potentially compared to the frequency of the paternal haplotype on the potentially affected chromosome. Since many SBS sets can be made available, a list of well-behaved SBSs will be directly generated. These SBSs can be enriched by target capture or PCR amplification and detected earlier than current possible detection. Non-biased PCR amplification of DNA for typical NIPT is difficult because slight non-linearity affects quantification. In the microhaplotype method, instead of simply counting the number of reads, the ratio of microhaplotypes is looked at, reducing the sensitivity to amplification bias. By selecting SBS sets less prone to sequencing errors or by selecting multi-SBS sets that result in two or more sequence changes going from the maternal microhaplotype to the paternal microhaplotype, the accuracy can be further enhanced. Additionally, the fetal DNA fraction can be easily determined by examining the genotype frequencies in SNP sets with three microhaplotypes. The fetal fraction is twice the frequency of the third microhaplotype. Knowledge of the fetal fraction and its diversity allows for a more accurate determination of whether the test results are valid or uncertain.

[0096] To determine trisomy or other DNA copy number abnormalities, the frequencies of the third microhaplotype in different regions are compared. If the frequency of the third microhaplotype from any large genomic region (part or whole of a chromosome) differs from the frequency of other genomic regions, it means trisomy or other amplification (increase in the frequency of the third microhaplotype) or deletion (absence of the third microhaplotype). Supplementary table

[0097] (Table 5) SBS sets of 507 gene panel TIFF2025111575000006.tif229106TIFF2025111575000007.tif229132TIFF2025111575000008.tif229122TIFF2025111575000009.tif229112TIFF2025111575000010.tif229122TIFF2025111575000011.tif229122TIFF2025111575000012.tif229122TIFF2025111575000013.tif229122TIFF2025111575000014.tif229132

[0098] (Table 6) SBS Set for Exome Analysis TIFF2025111575000015.tif195110TIFF2025111575000016.tif190129TIFF2025111575000017.tif190129TIFF2025111575000018.tif190128TIFF2025111575000019.tif190128TIFF2025111575000020.tif190128TIFF2025111575000021.tif190128TIFF2025111575000022.tif190128TIFF2025111575000023.tif190128

[0099] (Table 7) SNP Set TIFF2025111575000024.tif24078TIFF2025111575000025.tif235128TIFF2025111575000026.tif235137TIFF2025111575000027.tif235128TIFF2025111575000028.tif234125TIFF2025111575000029.tif235137TIFF2025111575000030.tif234137TIFF2025111575000031.tif234125TIFF2025111575000032.tif234125TIFF2025111575000033.tif234125TIFF2025111575000034.tif234125TIFF2025111575000035.tif234125TIFF2025111575000036.tif234125TIFF2025111575000037.tif234125TIFF2025111575000038.tif234117TIFF2025111575000039.tif234125TIFF2025111575000040.tif234125TIFF2025111575000041.tif235129TIFF2025111575000042.tif235138TIFF2025111575000043.tif237129TIFF2025111575000044.tif235129TIFF2025111575000045.tif23569

[0100] (Table 8) Observed frequency of the third MH (×2) TIFF2025111575000046.tif213124TIFF2025111575000047.tif59128

[0101] As described above, the present invention has been described with reference to the embodiments, but it will be understood that modifications and variations are included within the spirit and scope of the present invention. Therefore, the present invention is limited only by the appended claims.

Claims

1. 1. A method for detecting a set of single nucleotide polymorphisms (SNPs) having at least three microhaplotypes from a plurality of subjects present in a sample, the method comprising: a) determining the presence or absence of a set of SNPs having more than two microhaplotypes in a sample, the set of SNPs comprising a plurality of single base pair substitutions and corresponding to a genomic region selected from a set of predetermined regions; b) quantifying the frequency of the SNP set to determine the presence of DNA from multiple subjects in the sample, thereby detecting a SNP set having at least three microhaplotypes from multiple subjects in the sample; The method comprising:

2. The method of claim 1, wherein the set of predetermined regions comprises one or more sets of genomic regions described below: 1) A set of genomic regions listed in Table 5: chr1:120057158-120057246, chr1:156846120-156846233, chr1:226589833-226589958, chr1:23885498-23885599, chr10:104386934-104387019, chr10:43615505-43615633, chr10:70332580-70332672, chr11:534197-534242, chr11:8246326-8246343, chr12:121416622-121416650, chr12:121431272-121431300, chr12:121435427-121435475, chr12:121437114-121437221, chr12:133208886-133208979, chr12:133226159-133226196, chr12:133253995-133254083, chr12:18656174-18656225, chr12:56494991-56494998, chr13:21562832-21562948, chr14:102568296-102568367, chr14:104165753-104165927, chr14:105239146-105239192, chr14:105258892-105258893, chr14:35872792-35872926, chr15:40998305-40998342, chr15:41857216-41857303, chr15:41860411-41860490, chr15:67457335-67457485, chr16:2138269-2138398, chr16:2138398-2138422, chr16:68857289-68857441, chr16:81819768-81819820, chr16:89806343-89806347, chr16:89849583-89849629, chr16:89858505-89858525, chr17:1782952-1782957, chr17:78599562-78599655, chr17:78820329-78820374, chr17:78865546-78865630,chr17:78897547-78897561, chr17:78921117-78921211, chr19:10267011-10267077, chr19:17937758-17937786, chr19:17955001-17955021, chr19:2226676-2226772, chr19:3119184-3119239, chr19:50919797-50919828, chr19:5210622-5210782, chr19:5210762-5210782, chr19:5212380-5212482, chr19:7166376-7166388, chr2:112754828-112754880, chr2:112754943-112755001, chr2:141259283-141259376, chr2:29416366-29416481, chr2:29416481-29416615, chr2:29446184-29446202, chr2:48010488-48010558, chr20:40714307-40714479, chr20:40714539-40714540, chr20:57478807-57478939, chr20:9543622-9543681, chr21:42845374-42845383, chr22:21337266-21337325, chr22:21348914-21349037, chr22:24158895-24158899, chr3:178922222-178922274, chr3:183211906-183212026, chr4:106196829-106196951, chr4:143043340-143043404 chr4:143324036-143324094, chr4:187534362-187534375, chr4:187629497-187629538, chr5:149456772-149456811, chr5:149495287-149495395, chr5:176517326-176517461, chr5:176523562-176523597, chr5:176721198-176721272, chr5:180046209-180046344,chr5:180051003-180051118, chr5:180057231-180057293, chr5:231111-231143, chr5:35861068-35861159, chr5:35871190-35871273, chr5:57754808-57754851, chr5:67522722-67522851, chr6:117725448-117725578, chr6:117730673-117730819, chr6:152382311-152382325, chr6:26056549-26056708, chr6:30865115-30865204, chr6:32188603-32188642, chr7:100410597-100410657, chr7:6026775-6026942, chr7:78119109-78119199, chr8:30999122-30999123, chr8:31024638-31024654, chr8:90958422-90958530, chr9:139403268-139403280, chr9:139405093-139405261, chr9:139410424-139410589, chr9:139411714-139411880, chr9:21968159-21968199, chr9:93639846-93639973, chr9:93641175-93641199, chr9:98238358-98238379, 2) A set of genomic regions listed in Table 6: chr1:3743319-3743391, chr1:10431132-10431158, chr1:32672908-32672932, chr1:94544234-94544276, chr1:154832290-154832304, chr1:159409857-159409884, chr1:171168545-171168584, chr1:183616884-183616926, chr11:4928841-4928866, chr11:5345128-5345170, chr11:5566030-5566051, chr11:63883985-63884027, chr11:85436303-85436352, chr11:116703640-116703671, chr12:6030405-6030437, chr12:40834918-40834955, chr12:113348849-113348870, chr12:121600180-121600253, chr12:132688115-132688137, chr13:25367282-25367301, chr14:23549285-23549319, chr14:65263300-65263347, chr14:96136775-96136794, chr15:41819283-41819322, chr15:79310256-79310288, chr15:89398330-89398407, chr15:94945704-94945719, chr16:2812890-2812939, chr16:87678144-87678165, chr17:1782952-1782957, chr17:3101578-3101590, chr17:3352294-3352309, chr17:6331803-6331836, chr17:10223697-10223714, chr17:33772658-33772689, chr17:42989063-42989088, chr17:45695832-45695914, chr17:80887206-80887244, chr18:56204747-56204768, chr19:4510530-4510560,chr19:8148301-8148314, chr19:9362297-9362343, chr19:11227554-11227602, chr19:36237227-36237245, chr19:44352639-44352666, chr19:58131576-58131623, chr19:58213952-58213969, chr19:58572959-58572979, CHR2:33623720-33623734, CHR2:37579937-37579971, CHR2:71058184-71058226, CHR2:231775094-231775144, CHR2:239184569-239184581, chr20:744382-744415, chr20:5904028-5904040, chr20:52645534-52645541, chr20:62597666-62597694, chr21:43557698-43557736, chr21:46321659-46321677, chr22:17589209-17589246, chr22:19951207-19951271 chr22:21377301-21377334, chr22:33253280-33253292, chr22:35817553-35817597, chr22:44322922-44322970, chr3:122003757-122003769, chr3:129155451-129155463, chr3:136574501-136574521, chr3:142277536-142277575, chr3:178968634-178968660, chr4:156289900-156289917 chr5:147024476-147024509, chr5:148206440-148206473, chr5:150666933-150666962, chr5:150901613-150901630, chr5:174870150-174870196, chr6:4069133-4069166, chr6:29913201-29913266, chr6:30080231-30080274, chr6:30993533-30993590,chr6:31170514-31170528, chr6:31930441-31930462, chr6:33141253-33141280, chr6:36291985-36292007, chr6:167754702-167754721, chr7:4213975-4214023, chr7:21640361-21640405, chr7:27196069-27196113, chr7:30795288-30795331, chr7:55220177-55220202, chr7:100677455-100677523, CHR8:142490120-142490166, CHR8:145639681-145639726, chr9:117166206-117166246, chr9:125315542-125315557, chr9:134385435-134385436, chr9:136412255-136412296, chrX:23019317-23019346, 3) A set of genomic regions listed in Table 7: chr1:10431132-10431158, chr1:120057158-120057246, chr1:154832290-154832304, chr1:156846120-156846233, chr1:159409857-159409884, chr1:171168545-171168584, chr1:183616884-183616926, chr1:226573364-226573402, chr1:226589833-226589958, chr1:23885498-23885599, chr1:32672908-32672932, chr1:3743319-3743391, chr1:94544234-94544276, chr10:104386934-104387019, chr10:123194558-123194609, chr10:123199092-123199095, chr10:123275662-123275666, chr10:123335839-123335866, chr10:123346116-123346190, chr10:123396728-123396806, chr10:123406645-123406663, chr10:43611708-43611865, chr10:43615505-43615633, chr10:70332580-70332672, chr11:116703640-116703671, chr11:4928841-4928866, chr11:534197-534242, chr11:5345128-5345170, chr11:5566030-5566051, chr11:63883985-63884027, chr11:69412090-69412124, chr11:8246326-8246343, chr11:85436303-85436352, chr12:113348849-113348870, chr12:12009741-12009874, chr12:12013572-12013612, chr12:12016008-12016089, chr12:12020114-12020170, chr12:12035649-12035664,chr12:121416622-121416650, chr12:121431272-121431300, chr12:121435427-121435475, chr12:121437114-121437221, chr12:121600180-121600253, chr12:132688115-132688137, chr12:133208886-133208979, chr12:133226159-133226196, chr12:133253995-133254083, chr12:18656174-18656225, chr12:40834918-40834955, chr12:4346169-4346177, chr12:4351884-4352027, chr12:4376089-4376091, chr12:4399036-4399087, chr12:4399917-4399970, chr12:4411639-4411683, chr12:4417127-4417232, chr12:56494991-56494998, chr12:6030405-6030437, chr12:69169222-69169316, chr12:69265196-69265278, chr12:69277127-69277165, chr13:21562832-21562948, chr13:25367282-25367301, chr13:32986219-32986340, chr14:102568296-102568367, chr14:104165753-104165927, chr14:105239146-105239192, chr14:105258892-105258893, chr14:23549285-23549319, chr14:35872792-35872926, chr14:65263300-65263347, chr14:96136775-96136794, chr15:40998305-40998342, chr15:41819283-41819322, chr15:41857216-41857303, chr15:41860411-41860490, chr15:67457335-67457485,chr15:79310256-79310288, chr15:88488326-88488428, chr15:88549118-88549151, chr15:88646922-88647038, chr15:88667852-88667948, chr15:89398330-89398407, chr15:94945704-94945719, chr16:2138269-2138398, chr16:2138398-2138422, chr16:2812890-2812939, chr16:68857289-68857441, chr16:81819768-81819820, chr16:87678144-87678165, chr16:89806343-89806347, chr16:89849480-89849629, chr16:89858505-89858525, chr17:1782952-1782957, chr17:3101578-3101590, chr17:33772658-33772689, chr17:37832279-37832315, chr17:37834715-37834808, chr17:41616392-41616456, chr17:42989063-42989088, chr17:45695832-45695914, chr17:6331803-6331836, chr17:78599562-78599655, chr17:78820329-78820374, chr17:78865546-78865630, chr17:78896488-78896529, chr17:78897547-78897561, chr17:78921117-78921211, chr17:80887206-80887244, chr18:56204747-56204768, chr19:10267011-10267077, chr19:11227554-11227602, chr19:17937758-17937786, chr19:17955001-17955021, chr19:2226676-2226772, chr19:30253901-30253998, chr19:30255068-30255090,chr19:30290349-30290357, chr19:30340381-30340412, chr19:30361995-30362112, chr19:3119184-3119239, chr19:36237227-36237245, chr19:41724820-41724885, chr19:41781493-41781579, chr19:44352639-44352666, chr19:4510530-4510560, chr19:50919797-50919828, chr19:5210622-5210782, chr19:5210762-5210782, chr19:5212380-5212482, chr19:58131576-58131623, chr19:58213952-58213969, chr19:58572959-58572979, chr19:7163154-7163230, chr19:7166376-7166388, chr19:8148301-8148314, chr19:9362297-9362343, chr2:112754828-112754880, chr2:112754943-112755001, chr2:113983937-113984033, chr2:113984503-113984594, chr2:113989236-113989267, chr2:141259283-141259376, chr2:16042003-16042051, chr2:16073257-16073263, chr2:16112814-16112828, chr2:16113594-16113723, chr2:202122956-202122995, CHR2:231775094-231775144, CHR2:239184569-239184581, chr2:29416366-29416481, chr2:29416481-29416615, chr2:29446184-29446202, chr2:29446701-29446721, chr2:29447108-29447253, chr2:33623720-33623734, chr2:37579937-37579971,chr2:47800577-47800603, chr2:47852559-47852643, chr2:48010488-48010558, CHR2:71058184-71058226, chr20:30729488-30729523, chr20:40714307-40714479, chr20:40714479-40714540, chr20:40714539-40714540, chr20:52645534-52645541, chr20:57478807-57478939 chr20:5904028-5904040, chr20:62597666-62597694, chr20:744382-744415, chr20:9543622-9543681, chr21:42845374-42845383, chr21:42876400-42876447, chr21:43557698-43557736, chr21:46321659-46321677, chr22:17589209-17589246, chr22:17640022-17640045, chr22:19951207-19951271 chr22:21337266-21337325, chr22:21348914-21349037, chr22:21377301-21377334, chr22:24158895-24158899, chr22:29690246-29690345, chr22:33253280-33253292, chr22:35817553-35817597, chr22:44322922-44322970, chr3:122003757-122003769, chr3:12649857-12649937, chr3:129155451-129155463, chr3:136574501-136574521, chr3:138327951-138328016, chr3:142277536-142277575, chr3:178922222-178922274, chr3:178968634-178968660, chr3:178984575-178984679, chr3:178986121-178986203, chr3:178990402-178990462,chr3:183211906-183212026, chr3:36986932-36986992, chr3:71247257-71247304, chr4:106196829-106196951, chr4:143043340-143043404, chr4:143324036-143324094, chr4:156289900-156289917, chr4:1745492-1745500, chr4:1750487-1750584, chr4:1788994-1789044, chr4:1796629-1796636, chr4:1797741-1797852, chr4:187534362-187534375, chr4:187629497-187629538, chr4:54269096-54269173, chr4:54657737-54657790, chr4:55208737-55208788, chr4:55501109-55501195, chr4:55582037-55582068, chr4:55619846-55619859, chr4:55982752-55982784, chr4:56026865-56026914, chr5:147024476-147024509, chr5:148206440-148206473, chr5:149456772-149456811, chr5:149495287-149495395, chr5:150666933-150666962, chr5:150901613-150901630, chr5:174870150-174870196, chr5:176517326-176517461, chr5:176523562-176523597, chr5:176531772-176531857, chr5:176721198-176721272, chr5:180046209-180046344, chr5:180051003-180051118, chr5:180057231-180057293, chr5:231111-231143, chr5:35861068-35861159, chr5:35871190-35871273, chr5:56178111-56178217,chr5:57754808-57754851, chr5:67477132-67477234, chr5:67492589-67492652, chr5:67517563-67517646, chr5:67522722-67522851, chr5:67534039-67534057, chr5:67553771-67553827, chr6:117725448-117725578, chr6:117730673-117730819, chr6:152382311-152382325, chr6:167754702-167754721, chr6:26056549-26056708, chr6:29913201-29913266, chr6:30080231-30080274, chr6:30865115-30865204, chr6:30993533-30993590, chr6:31170514-31170528, chr6:31930441-31930462, chr6:32188603-32188642, chr6:32190390-32190484, chr6:33141253-33141280, chr6:36291985-36292007, chr6:4069133-4069166, chr6:41924853-41924931, chr6:42013020-42013049, chr6:42039487-42039542, chr6:42039551-42039666, chr6:42052577-42052667, chr7:100410597-100410657, chr7:100416139-100416250, chr7:100677455-100677523 chr7:116336880-116336947, chr7:116471122-116471227, chr7:21640361-21640405, chr7:27196069-27196113, chr7:30795288-30795331, chr7:4213975-4214023, chr7:55220177-55220202, chr7:55251541-55251648, chr7:6026775-6026942, chr7:6026942-6026988,chr7:78119109-78119199, chr8:128700175-128700233, chr8:128713221-128713364, chr8:128889285-128889371, CHR8:142490120-142490166, CHR8:145639681-145639726, chr8:145737636-145737816, chr8:30999122-30999123, chr8:31024638-31024654, chr8:38299624-38299715, chr8:38310910-38311001, chr8:38350292-38350315, chr8:38361379-38361430, chr8:90958422-90958530, chr9:117166206-117166246, chr9:125315542-125315557, chr9:134385435-134385436, chr9:136412255-136412296, chr9:139401504-139401577, chr9:139403268-139403280, chr9:139405093-139405261, chr9:139410424-139410589, chr9:139411714-139411880, chr9:21968159-21968199, chr9:5408242-5408358, chr9:5415025-5415111, chr9:5420254-5420266, chr9:5458035-5458095, chr9:5484100-5484203, chr9:87478135-87478172, chr9:93639846-93639973 chr9:93641175-93641199, chr9:98238358-98238379, chrX:23019317-23019346。,

3. An oligonucleotide panel comprising oligonucleotides for amplifying or hybrid capturing a region of a genome corresponding to one or more genomic regions selected from the set of genomic regions described in claim 2.

4. a) amplifying a region of a genome present in a sample, said region corresponding to a genomic region selected from the set of genomic regions of claim 2, and generating an amplicon by amplification; b) sequencing the amplicon to determine the nucleic acid sequence of the amplicon; A method comprising:

5. 5. The method of claim 4, further comprising quantifying the number of SNP sets with more than two microhaplotypes, more than three microhaplotypes, or more than four microhaplotypes present in the sample.

6. 1. A method for detecting a disease or disorder in a subject, comprising: a) i) identifying a region of interest, said region of interest being associated with a disease or disorder; ii) detecting single base pair substitutions (SBS) within the target region, thereby generating a plurality of sets of sequence variants; and iii) analyzing each variant set for linkage disequilibrium to identify microhaplotypes; identifying microhaplotypes in DNA molecules present in a sample obtained from the subject; b) determining the presence or absence of a set of single nucleotide polymorphisms (SNPs) having greater than two microhaplotypes in said sample; c) quantifying the frequency of the set of SNPs to determine the presence or absence of a genetic marker indicative of a disease or disorder, thereby detecting the disease or disorder; The method comprising:

7. 7. The method of claim 6, wherein the disease or disorder is (i) trisomy 13, 18, or 21, (ii) a gene copy number variation, or (iii) a fetal disorder.

8. 7. The method of claim 6, wherein the frequency of the third microhaplotype at a particular chromosome or chromosomal region is compared to the frequency of the third microhaplotype elsewhere in the genome.

9. a) at least one processor operatively connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in the DNA sample; c) i) identifying microhaplotypes in the sample based on the presence of single base pair substitutions; ii) confirming the presence of a number of SNP sets for the microhaplotype in the DNA sample; and iii) quantifying the frequency of genotypes within a set of SNPs with more than two microhaplotypes in the DNA sample; an analysis component executed by at least one processor, configured to: A genetic analysis system comprising:

10. 10. The system of claim 9, wherein the analysis component is further configured to (i) determine the possible presence of a DNA contaminant in the sample or (ii) determine the presence or absence of a genetic mutation.

11. a) at least one processor operatively connected to a memory; b) a receiver component configured to receive DNA analysis information including microhaplotype sequence information generated from PCR amplification of DNA in the DNA sample; c) an analysis component executed by said at least one processor and configured to perform the method of any one of claims 1 to 2 and 4 to 8; A genetic analysis system comprising:

12. 10. A non-transitory computer-readable storage medium encoded with a computer program, the program comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations for performing the method of any one of claims 1-2 and 4-8.