Method and system for kinship assessment for missing persons and disaster / conflict victims
The method addresses the challenge of confirming family relationships using small, low-quality DNA samples by employing multiplex PCR and sequencing to generate accurate DNA profiles, overcoming the limitations of existing technologies and privacy concerns.
Patent Information
- Application Number
- JP2025508470
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-14
- Filing Date
- 2023-08-15
- Publication Date
- 2025-09-17
AI Technical Summary
Current methods for generating DNA profiles for family relationship confirmation in forensic cases require large, high-quality DNA samples and are not suitable for small, low-quality samples, and publicly accessible genetic databases are often not used due to privacy concerns in cases of missing persons or disaster victims.
A method involving multiplex PCR reactions with primers targeting 2,000 to 50,000 SNPs for DNA amplification, followed by sequencing and genotyping to generate DNA profiles, which can be used to calculate relatedness without relying on publicly accessible databases.
Enables confirmation of family relationships using small, low-quality DNA samples by generating accurate DNA profiles and calculating relatedness, suitable for missing persons and disaster victims.
Smart Images

Figure 2025530659000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 398,512, entitled "METHODS AND SYSTEMS FOR KINSHIP EVALUATION FOR MISSING PERSONS AND DISASTER / CONFLICT VICTIMS," filed August 16, 2022, and U.S. Provisional Application No. 63 / 445,541, entitled "METHODS AND SYSTEMS FOR KINSHIP EVALUATION FOR MISSING PERSONS AND DISASTER / CONFLICT VICTIMS," filed February 14, 2023, the contents of which are incorporated by reference in their entireties. Field
[0002] The present disclosure relates, in some embodiments, to methods and systems for DNA-based kinship assessment of persons of interest, such as missing persons and victims of conflict and disaster. [Background technology]
[0003] background Current methods for generating DNA profiles for comparison in genetic databases include genotyping using high-density SNP microarrays and whole-genome sequencing (WGS), followed by linking evidentiary samples to distant relatives in the database. These require large, high-quality DNA samples and are not designed for family search or forensic purposes. Forensic casework samples are generally small and low-quality and typically require querying publicly accessible genetic databases to confirm family relationships. However, in the context of missing persons or victims of disasters or conflicts, family members may be hesitant to provide genetic samples that would help identify such individuals if their genetic samples were to be uploaded to publicly accessible databases. Therefore, there is a need for new and improved methods for generating DNA-based profile analyses that enable confirmation of family relationships of individuals of interest, such as missing persons or victims of disasters or conflicts, without requiring the use of publicly accessible genetic databases. Summary of the Invention [Means for solving the problem]
[0004] Abstract Provided herein is a method for performing DNA-based kinship analysis, the method comprising: providing a nucleic acid sample from a person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences collectively comprising a plurality of at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplified products, wherein the amplification is performed in one or more multiplex PCR reactions; generating a nucleic acid library from the amplified products; sequencing the nucleic acid library generated from the amplified products; analyzing the sequences of the amplified products; genotyping the plurality of SNPs, thereby generating a DNA profile; and calculating a degree of relatedness between the DNA profile and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included in a reference set of DNA profiles that includes one or more reference DNA profiles from relatives of the person of interest.
[0005] Also provided herein is a method for performing DNA-based kinship analysis, comprising: providing a nucleic acid sample from a person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences collectively comprising a plurality of at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplified products, wherein the amplification is performed in one or more multiplex PCR reactions; generating a nucleic acid library from the amplified products; sequencing the nucleic acid library generated from the amplified products; genotyping the plurality of SNPs, thereby generating a DNA profile; and calculating a degree of relatedness between the DNA profile and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included in a reference set of DNA profiles that includes one or more reference DNA profiles from relatives of the person of interest.
[0006] In some of any such embodiments, the sequencing is performed using massively parallel sequencing (MPS). In some of any such embodiments, the sequencing does not include whole genome sequencing (WGS).
[0007] In some of any such embodiments, the method further includes generating a family tree that includes DNA profiles related to the one or more DNA profiles.
[0008] Also provided herein are methods for constructing a nucleic acid library for a person of interest, comprising: providing a nucleic acid sample from the person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise at least or about 2,000 to 50,000 single nucleotide polymorphisms (SNPs), thereby generating a nucleic acid library comprising amplified products, wherein the amplification is performed in one or more multiplex PCR reactions. In some embodiments, the method further comprises sequencing the amplified products to produce a DNA profile of the person of interest.
[0009] Also provided herein is a method of constructing a nucleic acid library for a reference DNA sample, the method comprising: providing a nucleic acid sample from a relative of a person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of between at least or about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating a nucleic acid library comprising the amplified products, wherein the amplification is performed in one or more multiplex PCR reactions.
[0010] In some embodiments, the relative is a first-, second-, third-, fourth-, or fifth-degree relative of the person of interest. In some of any of such embodiments, the relative is a first-, second-, or third-degree relative of the person of interest.
[0011] In some of such embodiments, the nucleic acid sample comprises genomic DNA. In some of such embodiments, the nucleic acid sample comprises one or more enzyme inhibitors. In some of such embodiments, the one or more enzyme inhibitors comprise one or more inhibitors selected from the group consisting of hematin, heme, humic acid, indigo, tannic acid, collagen, calcium, and hydroxyapatite. In some of such embodiments, the nucleic acid sample comprises low-quality and / or low-abundance nucleic acid molecules. In some of such embodiments, the low-quality nucleic acid molecules are degraded and / or fragmented genomic DNA. In some of any of such embodiments, the low quality nucleic acid molecules are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200 or a Degradation Index (DI) of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200. In some of any of such embodiments, low-quality nucleic acid molecules have a DI of at least 1 and up to 158.3 or less. In some of any of such embodiments, the nucleic acid sample comprises high-quality nucleic acid molecules. In some of such embodiments, high-quality nucleic acid molecules have a DI of less than 1.
[0012] In some of such embodiments, the person of interest is a missing person. In some of such embodiments, the person of interest is a victim of a disaster or conflict.
[0013] In some of any such embodiments, the nucleic acid sample is derived from saliva, blood, semen, hair, teeth, bone, or skin. In some of any such embodiments, the nucleic acid sample is derived from saliva, blood, or semen. In some of such embodiments, the nucleic acid sample is derived from bone or hair. In some of such embodiments, the nucleic acid sample is derived from a buccal swab, paper, fabric, or other substrate or object impregnated with saliva, blood, semen, or other bodily fluid, or containing hair or skin cells.
[0014] In some of such embodiments, the nucleic acid sample comprises between or about 3 pg and 100 ng of genomic DNA. In some of such embodiments, the nucleic acid sample comprises between or about 100 pg and 5 ng of genomic DNA, between or about 50 pg and 5 ng of genomic DNA, or between or about 3 pg and 5 ng of genomic DNA. In some of such embodiments, the nucleic acid sample comprises at or about 1 ng of genomic DNA.
[0015] In some of any such embodiments, the plurality of SNPs comprises kinship SNPs (kiSNPs). In some of any such embodiments, the plurality of SNPs comprises Y chromosome SNPs (Y-SNPs). In some of such embodiments, the plurality of SNPs comprises kiSNPs and Y-SNPs. In some of such embodiments, the plurality of SNPs comprises kiSNPs, biogeographic ancestry SNPs (aiSNPs), identity SNPs (iiSNPs), phenotypic SNPs (piSNPs), X chromosome SNPs (X-SNPs), and Y chromosome SNPs (Y-SNPs). In some of such embodiments, the plurality of SNPs comprises SNPs selected from one or more of the group consisting of kiSNPs, aiSNPs, iiSNPs, piSNPs, X-SNPs, and Y-SNPs. In some of any of such embodiments, at least or at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the plurality of SNPs are related SNPs.
[0016] In some of any of such embodiments, the reference set of DNA profiles includes up to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 reference DNA profiles. In some of any of such embodiments, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest. In some of any of such embodiments, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest, and at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are each first-, second-, third-, fourth-, or fifth-degree relatives. In some of any such embodiments, at least 50% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest.
[0017] In some of such embodiments, each relative of the person of interest in the reference set of DNA profiles is a first-, second-, third-, fourth-, or fifth-degree relative of the person of interest, respectively. In some of such embodiments, each relative of the person of interest in the reference set of DNA profiles is a first-, second-, or third-degree relative of the person of interest, respectively. In some of such embodiments, the identity of each relative of the person of interest in the reference set of DNA profiles is known. In some of such embodiments, the identity of each of the one or more reference DNA profiles in the reference set of DNA profiles is known. In some of such embodiments, the reference set of DNA profiles is in a database. In some embodiments, the database is not publicly accessible.
[0018] In some of such embodiments, the sequencing comprises a sequencing plexity of up to 40-plex. In some of such embodiments, the sequencing comprises a sequencing plexity of up to 32-plex. In some of such embodiments, the sequencing comprises a sequencing plexity of 12-plex to 32-plex. In some of such embodiments, the sequencing comprises a sequencing plexity of 24-plex to 32-plex. In some of any of such embodiments, the sequencing is performed at or about 10-plex, 11-plex, 12-plex, 13-plex, 14-plex, 15-plex, 16-plex, 17-plex 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex, or 35-plex. In some such embodiments, the sequencing comprises a sequencing plex of 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex, or 35-plex. In some such embodiments, the sequencing comprises a sequencing plex of 8-16-plex or about 8-16-plex for post-mortem samples, and / or the sequencing comprises a sequencing plex of 24-40-plex or about 24-40-plex for ante-mortem samples. In some such embodiments, the sequencing comprises a sequencing plex of 12-plex or about 12-plex for post-mortem samples, and / or the sequencing comprises a sequencing plex of 32-plex or about 32-plex for ante-mortem samples.In some of any of such embodiments, the sequencing comprises a sequencing complexity of 30-plex, 31-plex, or 32-plex, or about 30-plex, 31-plex, or 32-plex.
[0019] In some of any such embodiments, the method further includes identifying the person of interest.
[0020] Also provided herein is a method for calculating degree of relatedness, comprising obtaining a DNA profile comprising genotypes for at least or between about 2,000 and 50,000 SNPs, wherein the DNA profile is from a person of interest; and calculating a degree of relatedness between the DNA profile and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included in a reference set of DNA profiles that includes one or more reference DNA profiles from relatives of the person of interest.
[0021] Also provided herein is a method for calculating degree of relatedness, comprising generating a DNA profile comprising genotypes for at least or between about 2,000 and 50,000 SNPs, wherein the DNA profile is from a person of interest; and calculating a degree of relatedness between the DNA profile and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included in a reference set of DNA profiles that includes one or more reference DNA profiles from relatives of the person of interest.
[0022] In some of such embodiments, the relatedness is calculated using a kinship model. In some of such embodiments, the relatedness is calculated using a kinship model trained using a PCA method. In some embodiments, the PCA method for training the kinship model is or includes PCA. In some of such embodiments, the PCA method is PC-AiR. In some embodiments, PC-AiR includes the steps of: (1) estimating relatedness coefficients between all pairs of samples in a training database, and optionally, in a training DNA profile, where pairings with a relatedness coefficient >0.025 are confirmed as closely related and pairings with a relatedness coefficient <-0.025 are confirmed as ancestrally diverged; (2) initializing an unrelated sample set containing all samples; and (3) iteratively: (i) identifying a set in the unrelated sample set that has the most related samples in the unrelated sample set, thereby designating this set as X; (ii) identifying a set of samples in X that has the fewest ancestrally diverged pairings compared to the samples in the unrelated sample set, thereby designating this set as Y; and (iii) terminating the process if Y has 0 samples, or randomly selecting one sample from Y and removing it from U if Y has at least one sample, and repeating beginning with step (3)(i).
[0023] In some of such embodiments, the PCA method is a modified PC-AiR. In some of such embodiments, the modified PC-AiR method includes: (1) estimating relatedness coefficients between all pairs of samples, optionally training DNA profiles, in a training database, where pairings with a relatedness coefficient >0.01 are identified as closely related and pairings with a relatedness coefficient <-0.025 are identified as ancestrally diverged; (2) removing all DNA profiles with ≥5% missing data; and (3) ranking all DNA profiles by assigning each DNA profile a ranking value. In some embodiments, the ranking value is determined based on the number of related DNA profiles in the complete database ranked from smallest to largest, broken down by the number of ancestrally diverged DNA profiles in the complete database ranked from largest to smallest. In some embodiments, step (3) includes iterating through the ranked DNA profiles, and for each DNA profile, (i) if the DNA profile is not already in the relevant sample set, adding it to the unrelated sample set and adding all relevant DNA profiles to the relevant sample set, and (ii) if the DNA profile is already in the relevant sample set, skipping to the next DNA profile and repeating starting at step (3)(i).
[0024] In some of such embodiments, calculating the degree of relatedness includes calculating a coefficient of relatedness using PC-Relate. In some of such embodiments, the degree of relatedness is calculated by providing a DNA profile of the person of interest as an input to PC-Relate. In some of such embodiments, the degree of relatedness is calculated by providing a kinship model and the DNA profile of the person of interest as input to PC-Relate. In some of such embodiments, one or more reference DNA profiles are further provided as inputs to PC-Relate.
[0025] In some of any such embodiments, calculating the degree of relatedness includes calculating the coefficient of relatedness using a whole-genome kinship algorithm as follows:
number
number
number
number
number
number
number
[0026] In some of such embodiments, calculating the degree of association comprises calculating a likelihood ratio. In some embodiments, calculating the likelihood ratio comprises comparing a plurality of SNPs between the DNA profile and one or more reference DNA profiles. In some embodiments, calculating the likelihood ratio comprises comparing a set of SNPs including related SNPs from among the plurality of SNPs between the DNA profile and one or more reference DNA profiles.
[0027] In some of any of such embodiments, calculating the likelihood ratio includes dividing the probability that the DNA profile and a reference DNA profile from among the one or more reference DNA profiles are related by the probability that the DNA profile and the reference DNA profile are unrelated, based on the genotypes of the plurality of SNPs.
[0028] In some of such embodiments, the likelihood ratio (LR) is calculated as follows:
number
[0029] In some of such embodiments, LR is calculated as follows:
number
[0030] In some of any of such embodiments, the person of interest is biologically male, and the method further includes calculating a likelihood ratio of sharing a Y chromosome between the DNA profile and one or more reference DNA profiles. In some embodiments, calculating the likelihood ratio of sharing a Y chromosome includes comparing a set of SNPs, including one or more Y-SNPs, between the DNA profile and the one or more reference DNA profiles.
[0031] In some of such embodiments, the one or more Y-SNPs are included within the plurality of SNPs. In some embodiments, the one or more Y-SNPs include at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 81, 82, 83, 84, or 85 Y-SNPs. In some of such embodiments, the one or more Y-SNPs include 85 Y-SNPs.
[0032] In some of any of such embodiments, calculating the likelihood ratio of sharing a Y chromosome includes dividing the probability that the DNA profile and a reference DNA profile from among the one or more reference DNA profiles share a Y chromosome by the probability that the DNA profile and the reference DNA profile do not share a Y chromosome, based on the genotypes of the one or more Y-SNPs.
[0033] In some of such embodiments, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the DNA profiles in the reference set of DNA profiles are from relatives of missing persons or victims of disaster or conflict. In some of such embodiments, each of the DNA profiles in the reference set of DNA profiles is from a relative of a missing person or victim of disaster or conflict.
[0034] In some of any of these embodiments, the reference set of DNA profiles comprises up to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 reference DNA profiles. In some of any of these embodiments, the reference set of DNA profiles comprises up to 100 reference DNA profiles.
[0035] In some of any of such embodiments, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from blood relatives of the person of interest. In some of any of such embodiments, at least 50% of the reference DNA profiles in the reference set of DNA profiles are from blood relatives of the person of interest.
[0036] In some of any of such embodiments, each relative of the person of interest in the reference set of DNA profiles is a first-degree, second-degree, third-degree, fourth-degree, or fifth-degree relative of the person of interest, respectively.
[0037] In some of such embodiments, the identity of each relative of the person of interest in the reference set of DNA profiles is known. In some of such embodiments, the identity of each relative of the person of interest in the reference set of DNA profiles is known. In some of such embodiments, the identity of each of the one or more reference DNA profiles in the reference set of DNA profiles is known.
[0038] In some of such embodiments, the reference set of DNA profiles is in a database. In some embodiments, the database is not publicly accessible. In some of such embodiments, the database is not accessible by third-party genealogy services.
[0039] Also provided herein are nucleic acid libraries constructed using any of the methods described herein.
[0040] Also provided herein are a plurality of primers that specifically hybridize to a plurality of target sequences comprising at least or about 2,000-50,000 single nucleotide polymorphisms (SNPs) in a nucleic acid sample from a person of interest, wherein amplification of the nucleic acid sample using the plurality of primers in one or more multiplex PCR reactions results in amplification products.
[0041] Also provided herein are a plurality of primers that specifically hybridize to a plurality of target sequences comprising between at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs) in a nucleic acid sample from a person of interest and one or more reference samples, wherein the one or more reference samples comprise samples from blood relatives of the person of interest, and wherein amplifying the nucleic acid sample from the person of interest and the one or more reference samples using the plurality of primers in one or more multiplex PCR reactions results in amplification products.
[0042] In some of any such embodiments, the nucleic acid sample from the person of interest comprises genomic DNA.
[0043] In some of such embodiments, the nucleic acid sample from the person of interest comprises one or more enzyme inhibitors. In some of such embodiments, the one or more enzyme inhibitors comprise one or more inhibitors selected from the group consisting of hematin, heme, humic acid, indigo, tannic acid, collagen, calcium, and hydroxyapatite. In some of such embodiments, the nucleic acid sample from the person of interest comprises low-quality and / or low-abundance nucleic acid molecules. In some embodiments, the low-quality nucleic acid molecules are degraded and / or fragmented genomic DNA. In some of any of such embodiments, the low quality nucleic acid molecules are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200 or a Degradation Index (DI) of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200. In some of any of such embodiments, the low quality nucleic acid molecules have a DI of at least 1 and up to or below 158.3. In some of any of such embodiments, the nucleic acid sample from the person of interest and / or the nucleic acid sample from one or more reference samples comprises high quality nucleic acid molecules. In some of any such embodiments, high quality nucleic acid molecules have a DI of less than 1.
[0044] In some of such embodiments, the person of interest is a missing person. In some of such embodiments, the person of interest is a victim of a disaster or conflict.
[0045] In some of any of such embodiments, the nucleic acid sample from the person of interest is derived from a buccal swab, paper, fabric or other substrate or object impregnated with saliva, blood or other bodily fluid, or containing hair or skin cells.
[0046] In some of any such embodiments, the nucleic acid sample from the person of interest comprises at or about 3 pg to 100 ng of genomic DNA. In some of any such embodiments, the nucleic acid sample from the person of interest comprises at or about 100 pg to 5 ng of genomic DNA, at or about 100 pg to 5 ng of genomic DNA, at or about 50 pg to 5 ng of genomic DNA, or at or about 3 pg to 5 ng of genomic DNA. In some of any such embodiments, the nucleic acid sample from the person of interest comprises at or about 1 ng of genomic DNA.
[0047] In some of any such embodiments, the plurality of SNPs comprises kinship SNPs (kiSNPs). In some of any such embodiments, the plurality of SNPs comprises Y chromosome SNPs (Y-SNPs). In some of such embodiments, the plurality of SNPs comprises kiSNPs and Y-SNPs. In some of such embodiments, the plurality of SNPs comprises kiSNPs, biogeographic ancestry SNPs (aiSNPs), identity SNPs (iiSNPs), phenotypic SNPs (piSNPs), X chromosome SNPs (X-SNPs), and Y chromosome SNPs (Y-SNPs). In some of such embodiments, the plurality of SNPs comprises SNPs selected from one or more of the group consisting of kiSNPs, aiSNPs, iiSNPs, piSNPs, X-SNPs, and Y-SNPs.
[0048] In some of any of such embodiments, at least or at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the plurality of SNPs are related SNPs. In some of any of such embodiments, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the DNA profiles in the reference set of DNA profiles are from relatives of missing persons or victims of the disaster or conflict.
[0049] In some of any such embodiments, each of the one or more reference samples is from a relative of a missing person or a victim of a disaster or conflict. In some of any such embodiments, the one or more reference samples include up to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 reference samples. In some of any of such embodiments, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the one or more reference samples are from relatives of the person of interest. In some of any of such embodiments, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from blood relatives of the person of interest, and at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are each first-, second-, third-, fourth-, or fifth-degree relatives. In some of any of such embodiments, at least 50% of the one or more reference samples are from blood relatives of the person of interest. In some of any of such embodiments, each relative of the person of interest in the one or more reference samples is a first-, second-, third-, fourth-, or fifth-degree relative of the person of interest, respectively. In some of any of such embodiments, each relative of the person of interest in the one or more reference samples is a first-, second-, or third-degree relative of the person of interest, respectively. In some of any of such embodiments, the identity of each relative of the person of interest in the one or more reference samples is known.In some of any such embodiments, the identity of each of the one or more reference samples is known.
[0050] Also provided herein is a method for constructing a DNA profile, the method comprising: providing a nucleic acid sample from a person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of between at least or about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplified products, wherein the amplification is performed in one or more multiplex PCR reactions; sequencing the amplified products; and genotyping the plurality of SNPs, thereby generating the DNA profile.
[0051] Also provided herein is a method for constructing a DNA profile, the method comprising: providing a nucleic acid sample from a person of interest; providing a nucleic acid sample from a relative of the person of interest; amplifying the nucleic acid sample from the person of interest and the nucleic acid sample from the relative with a plurality of primers that specifically hybridize to a plurality of target sequences collectively comprising a plurality of at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplified products, wherein the amplification is carried out in one or more multiplex PCR reactions; sequencing the amplified products; and genotyping the plurality of SNPs, thereby generating a DNA profile for the person of interest and the relative of the person of interest.
[0052] In some of such embodiments, the sequencing does not include whole genome sequencing (WGS). In some of such embodiments, the nucleic acid sample comprises genomic DNA. In some of such embodiments, the nucleic acid sample of the person of interest and / or the nucleic acid sample of a relative of the person of interest comprises genomic DNA.
[0053] In some of any of such embodiments, the nucleic acid sample, the nucleic acid sample of the person of interest and / or the nucleic acid sample of the relative comprises one or more enzyme inhibitors, hi some embodiments, the one or more enzyme inhibitors comprise one or more inhibitors selected from the group consisting of hematin, heme, humic acid, indigo, tannic acid, collagen, calcium, and hydroxyapatite.
[0054] In some of any of such embodiments, the nucleic acid sample, the nucleic acid sample of the person of interest and / or the nucleic acid sample of the relative, comprises low quality and / or low abundance nucleic acid molecules. In some embodiments, the low quality nucleic acid molecules are degraded and / or fragmented genomic DNA. In some of such embodiments, the low quality nucleic acid molecules comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200 The nucleic acid sample may have a degradation index (DI) of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200. In some of such embodiments, the low-quality nucleic acid molecules have a DI of at least 1 and up to 158.3 or less. In some of such embodiments, the nucleic acid sample, the nucleic acid sample of the person of interest, and / or the nucleic acid sample of the relative comprises high-quality nucleic acid molecules. In some embodiments, the high-quality nucleic acid molecules have a DI of less than 1.
[0055] In some of such embodiments, the person of interest is a missing person. In some of such embodiments, the person of interest is a victim of a disaster or conflict.
[0056] In some of any such embodiments, the blood relatives of the person of interest are first-, second-, third-, fourth-, or fifth-degree relatives. In some of any such embodiments, the blood relatives of the person of interest are first-, second-, or third-degree relatives.
[0057] In some of any of such embodiments, the nucleic acid sample, the nucleic acid sample of the person of interest and / or the nucleic acid sample of the relative is derived from a buccal swab, paper, fabric or other substrate or object impregnated with saliva, blood or other bodily fluid, or containing hair or skin cells.
[0058] In some of any such embodiments, the nucleic acid sample, the nucleic acid sample of the person of interest, and / or the nucleic acid sample of the relative comprises at or about 3 pg to 100 ng of genomic DNA. In some of any such embodiments, the nucleic acid sample, the nucleic acid sample of the person of interest, and / or the nucleic acid sample of the relative comprises at or about 100 pg to 5 ng of genomic DNA, at or about 50 pg to 5 ng of genomic DNA, or at or about 3 pg to 5 ng of genomic DNA. In some of any such embodiments, the nucleic acid sample, the nucleic acid sample of the person of interest, and / or the nucleic acid sample of the relative comprises at or about 1 ng of genomic DNA.
[0059] In some of any such embodiments, the plurality of SNPs includes kinship SNPs. In some of any such embodiments, the plurality of SNPs includes Y chromosome SNPs (Y-SNPs). In some of such embodiments, the plurality of SNPs includes kiSNPs and Y-SNPs. In some of such embodiments, the plurality of SNPs includes kiSNPs, biogeographic ancestry SNPs (aiSNPs), identity SNPs (iiSNPs), phenotypic SNPs (piSNPs), X chromosome SNPs (X-SNPs), and Y chromosome SNPs (Y-SNPs). In some of such embodiments, the plurality of SNPs includes SNPs selected from one or more of the group consisting of kiSNPs, aiSNPs, iiSNPs, piSNPs, X-SNPs, and Y-SNPs. In some of any of such embodiments, at least or at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the plurality of SNPs are related SNPs.
[0060] In some of such embodiments, the sequencing comprises a sequencing plexity of up to 40-plex. In some of such embodiments, the sequencing comprises a sequencing plexity of up to 32-plex. In some of such embodiments, the sequencing comprises a sequencing plexity of 12-plex to 32-plex. In some of such embodiments, the sequencing comprises a sequencing plexity of 24-plex to 32-plex. In some of such embodiments, the sequencing comprises a sequencing plexity of 4-plex, 5-plex, 6-plex, 7-plex, 8-plex, 9-plex, 10-plex, 11-plex, 12-plex, 13-plex, 14-plex, 15-plex, 16-plex, 17-plex, 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex, 35-plex, 36-plex, 37-plex, 38-plex, 39-plex, 40-plex, 41-plex, 42-plex, 43-plex, 44-plex, or 45-plex, or about 4-plex, 5-plex, 6-plex, 7-plex, 8-plex, 9-plex, 10-plex, 11-plex, 12-plex, 13-plex, 14-plex, 15-plex, 16-plex, 17-plex including 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex, 35-plex, 36-plex, 37-plex, 38-plex, 39-plex, 40-plex, 41-plex, 42-plex, 43-plex, 44-plex or 45-plex sequencing complexities.In some of any of such embodiments, the sequencing is performed at or about 10-plex, 11-plex, 12-plex, 13-plex, 14-plex, 15-plex, 16-plex, 17-plex 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex, or 35-plex. In some such embodiments, the sequencing comprises a sequencing plex of 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex, or 35-plex. In some such embodiments, the sequencing comprises a sequencing plex of 8-16-plex or about 8-16-plex for post-mortem samples, and / or the sequencing comprises a sequencing plex of 24-40-plex or about 24-40-plex for ante-mortem samples. In some such embodiments, the sequencing comprises a sequencing plex of 12-plex or about 12-plex for post-mortem samples, and / or the sequencing comprises a sequencing plex of 32-plex or about 32-plex for ante-mortem samples. In some of any of such embodiments, the sequencing comprises a sequencing complexity of 30-plex, 31-plex, or 32-plex, or about 30-plex, 31-plex, or 32-plex.
[0061] Also provided herein is a method for identifying genetic relatives of a DNA profile, the method comprising: calculating a degree of relatedness between the DNA profile of any one of claims 127-161 and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included within a reference set of DNA profiles comprising one or more reference DNA profiles from relatives of the person of interest; and generating a family tree comprising the DNA profile in relation to the one or more reference DNA profiles.
[0062] In some embodiments, the one or more reference DNA profiles are part of a database.
[0063] In some of any of such embodiments, the reference set of DNA profiles includes up to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 reference DNA profiles. In some of any of such embodiments, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest. In some of any of such embodiments, at least 50% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest. In some of such embodiments, each relative of the person of interest is a first-, second-, third-, fourth-, or fifth-degree relative of the person of interest, respectively.
[0064] In some of such embodiments, the identity of each relative of the person of interest in the reference set of DNA profiles is known. In some of such embodiments, the identity of each of the one or more reference DNA profiles in the reference set of DNA profiles is known.
[0065] In some of such embodiments, the reference set of DNA profiles is in a database. In some embodiments, the database is not publicly accessible. In some of such embodiments, the database is not accessible by third-party genealogy services.
[0066] Also provided herein is a method for confirming the identity of a DNA profile, the method comprising: calculating a degree of relatedness between a DNA profile comprising genotypes for at least or between about 2,000 and 50,000 SNPs and one or more reference DNA profiles, wherein the DNA profile is from a person of interest, and the one or more reference DNA profiles are included within a reference set of DNA profiles that includes one or more reference DNA profiles from relatives of the person of interest; and generating a pedigree tree that includes the DNA profile in relation to the one or more reference DNA profiles.
[0067] In some embodiments, the DNA profile is generated by any of the methods for generating a DNA profile described herein.
[0068] In some of such embodiments, the relatedness is calculated using a kinship model. In some of such embodiments, the relatedness is calculated using a kinship model trained using a PCA method. In some embodiments, the PCA method for training the kinship model is or includes PCA. In some of such embodiments, the PCA method is PC-AiR. In some of any such embodiments, PC-AiR includes: (1) estimating relatedness coefficients between all pairs of samples in a training database, and optionally in a training DNA profile, where pairings with a relatedness coefficient >0.025 are confirmed as closely related and pairings with a relatedness coefficient <-0.025 are confirmed as ancestrally diverged; (2) initializing an unrelated sample set containing all samples; and (3) iteratively: (i) identifying a set in the unrelated sample set that has the most related samples in the unrelated sample set, thereby designating this set as X; (ii) identifying a set of samples in X that has the fewest ancestrally diverged pairings compared to the samples in the unrelated sample set, thereby designating this set as Y; and (iii) terminating the process if Y has 0 samples, or randomly selecting one sample from Y and removing it from U if Y has at least one sample, and repeating beginning with step (3)(i).
[0069] In some of such embodiments, the PCA method is a modified PC-AiR. In some embodiments, the modified PC-AiR method includes: (1) estimating relatedness coefficients between all pairs of samples, optionally training DNA profiles, in a training database, where pairings with a relatedness coefficient >0.01 are identified as closely related and pairings with a relatedness coefficient <-0.025 are identified as ancestrally diverged; (2) removing all DNA profiles with ≥5% missing data; and (3) ranking all DNA profiles by assigning each DNA profile a ranking value. In some embodiments, the ranking value is determined based on the number of related DNA profiles in the complete database ranked from smallest to largest, broken down by the number of ancestrally diverged DNA profiles in the complete database ranked from largest to smallest. In some embodiments, step (3) includes iterating through the ranked DNA profiles, and for each DNA profile, (i) if the DNA profile is not already in the relevant sample set, adding it to the unrelated sample set and adding all relevant DNA profiles to the relevant sample set, and (ii) if the DNA profile is already in the relevant sample set, skipping to the next DNA profile and repeating starting at step (3)(i).
[0070] In some of such embodiments, calculating the degree of relatedness includes calculating a coefficient of relatedness using PC-Relate. In some embodiments, the degree of relatedness is calculated by providing a DNA profile of the person of interest as input to PC-Relate. In some of such embodiments, the degree of relatedness is calculated by providing a kinship model and the DNA profile of the person of interest as input to PC-Relate.
[0071] In some of any of such embodiments, one or more reference DNA profiles are further provided as input to PC-Relate.
[0072] In some of any such embodiments, calculating the degree of relatedness includes calculating the coefficient of relatedness using a whole-genome kinship algorithm as follows:
number
number
number
number
number
number
number
[0073] In some of such embodiments, calculating the degree of association comprises calculating a likelihood ratio. In some embodiments, calculating the likelihood ratio comprises comparing a plurality of SNPs between the DNA profile and one or more reference DNA profiles. In some embodiments, calculating the likelihood ratio comprises comparing a set of SNPs including related SNPs from among the plurality of SNPs between the DNA profile and one or more reference DNA profiles.
[0074] In some of any of such embodiments, calculating the likelihood ratio includes dividing the probability that the DNA profile and a reference DNA profile from among the one or more reference DNA profiles are related by the probability that the DNA profile and the reference DNA profile are unrelated, based on the genotypes of the plurality of SNPs.
[0075] In some of such embodiments, the likelihood ratio (LR) is calculated as follows:
number
number
[0076] In some of such embodiments, the person of interest is biologically male, and the method further comprises calculating a likelihood ratio of sharing a Y chromosome between the DNA profile and one or more reference DNA profiles. In some embodiments, calculating the likelihood ratio of sharing a Y chromosome comprises comparing a set of SNPs comprising one or more Y-SNPs between the DNA profile and one or more reference DNA profiles. In some embodiments, the one or more Y-SNPs are included within the plurality of SNPs. In some of such embodiments, the one or more Y-SNPs comprise at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 81, 82, 83, 84, or 85 Y-SNPs. In some of such embodiments, the one or more Y-SNPs comprise 85 Y-SNPs.
[0077] In some of any of such embodiments, calculating the likelihood ratio of sharing a Y chromosome comprises dividing the probability that the DNA profile and a reference DNA profile from among the one or more reference DNA profiles share a Y chromosome by the probability that the DNA profile and the reference DNA profile do not share a Y chromosome, based on the genotypes of the one or more Y-SNPs.
[0078] In some of any of such embodiments, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the DNA profiles in the reference set of DNA profiles are from relatives of missing persons or victims of the disaster or conflict.
[0079] In some of any of such embodiments, each of the one or more reference samples is from a relative of a missing person or a victim of a disaster or conflict. In some of such embodiments, the one or more reference samples include up to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 reference DNA samples.
[0080] In some of any of such embodiments, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the one or more reference samples are from relatives of the person of interest. In some of any of such embodiments, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest, and at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles, respectively, are first-, second-, third-, fourth-, or fifth-degree relatives. In some of such embodiments, at least 50% of the reference DNA profiles in one or more reference samples are from relatives of the person of interest. In some of such embodiments, each relative of the person of interest in one or more reference samples is a first-, second-, third-, fourth-, or fifth-degree relative of the person of interest, respectively. In some of such embodiments, each relative of the person of interest in one or more reference samples is a first-, second-, or third-degree relative of the person of interest, respectively.
[0081] In some of such embodiments, the identity of each relative of the person of interest in one or more reference samples is known. In some of such embodiments, the identity of each of the one or more reference samples is known.
[0082] Also provided herein are kits comprising at least one container means, wherein the at least one container means comprises a plurality of any of the primers described herein. In some of any of such embodiments, the plurality of SNPs is between 2,000 and 11,000 SNPs, between 3,000 and 11,000 SNPs, between 4,000 and 11,000 SNPs, between 5,000 and 11,000 SNPs, between 5,500 and 11,000 SNPs, between 6,000 and 11,000 SNPs, between 7,000 and 15,000 SNPs, between 7,000 and 14,000 SNPs, between 7,000 and 13,000 SNPs, between 7,000 and 12,000 SNPs, or between 8,000 and 13,000 SNPs. SNPs between 7,000 and 11,000, SNPs between 8,000 and 15,000, SNPs between 8,000 and 14,000, SNPs between 8,000 and 13,000, SNPs between 8,000 and 12,000, SNPs between 8,000 and 11,000, SNPs between 9,000 and 15,000, SNPs between 9,000 and 14,000, SNPs between 9,000 and 13,000, SNPs between 9,000 and 12,000 or SNPs between 9,000 and 11, 000 SNPs, or approximately between 2,000 and 11,000 SNPs, between 3,000 and 11,000 SNPs, between 4,000 and 11,000 SNPs, between 5,000 and 11,000 SNPs, between 5,500 and 11,000 SNPs, between 6,000 and 11,000 SNPs, between 7,000 and 15,000 SNPs, between 7,000 and 14,000 SNPs, between 7,000 and 13,000 SNPs, between 7,000 and 12,000 SNPs, In some of such embodiments, the plurality of SNPs comprises between 10,230 SNPs.In some of any of such embodiments, the plurality of SNPs is between 2,000 and 11,000 SNPs, between 3,000 and 11,000 SNPs, between 4,000 and 11,000 SNPs, between 5,000 and 11,000 SNPs, between 5,500 and 11,000 SNPs, between 6,000 and 11,000 SNPs, between 7,000 and 15,000 SNPs, between 7,000 and 14,000 SNPs, between 7,000 and 13,000 SNPs, between 7,000 and 12,000 SNPs, SNPs between 7,000 and 11,000, SNPs between 8,000 and 15,000, SNPs between 8,000 and 14,000, SNPs between 8,000 and 13,000, SNPs between 8,000 and 12,000, SNPs between 8,000 and 11,000, SNPs between 9,000 and 15,000, SNPs between 9,000 and 14,000, SNPs between 9,000 and 13,000, SNPs between 9,000 and 12,000 or SNPs between 9,000 and 11, 000 SNPs, or approximately between 2,000 and 11,000 SNPs, between 3,000 and 11,000 SNPs, between 4,000 and 11,000 SNPs, between 5,000 and 11,000 SNPs, between 5,500 and 11,000 SNPs, between 6,000 and 11,000 SNPs, between 7,000 and 15,000 SNPs, between 7,000 and 14,000 SNPs, between 7,000 and 13,000 SNPs, between 7,000 and 12,000 SNPs, In some of such embodiments, the plurality of SNPs comprises between 10,230 SNPs.
[0083] In some of any of such embodiments, the method further includes generating a pedigree tree that includes the DNA profile in relation to one or more DNA profiles included in the reference set of DNA profiles, hi some embodiments, the pedigree tree includes the DNA profile in relation to one or more DNA profiles from relatives of the person of interest. [Brief explanation of the drawings]
[0084] [Figure 1] FIG. 1 shows an exemplary schematic of a method for generating a sequenceable library.
[0085] [Figure 2] Figure 2 shows the results of the number of loci identified using various input dosage adjustments of genomic DNA, including 5 ng, 2.5 ng, 1 ng, 500 pg, 250 pg, 100 pg, and 50 pg.
[0086] [Figure 3] FIG. 3 shows the percentage of loci (call rate) detected in degraded DNA using the assay described herein compared to microarray (GSA) call rate.
[0087] [Figure 4] FIG. 4 shows the number of loci detected in the presence of the inhibitors hematin, humic acid, indigo, and tannic acid compared to the reference control.
[0088] [Figure 5] FIG. 5 shows an exemplary pedigree generated by the methods described herein.
[0089] [Figure 6] FIG. 6 shows the expected and observed coefficients of kinship calculated using the algorithm described herein.
[0090] [Figure 7]FIG. 7 shows the results of the 1:many search algorithm in an exemplary case study.
[0091] [Figure 8] FIG. 8 shows an exemplary pedigree generated from the results of the 1:multiple search algorithm.
[0092] [Figure 9] Figure 9 is a table summarizing the number and types of loci detected using various input dosage adjustments of genomic DNA, including 5 ng, 2.5 ng, 1 ng, 500 pg, 250 pg, 100 pg, and 50 pg.
[0093] [Figure 10] Figure 10 is a table summarizing the number and types of loci detected using DNA in the presence of the inhibitors hematin, humic acid, tannic acid, and indigo compared to the positive amplification control in the absence of inhibitors.
[0094] [Figure 11] Figure 11 is a table summarizing the number and types of loci detected for two samples of DNA obtained 9 and 22 hours after simulated sexual assault. DNA was isolated from the sperm fraction by differential extraction, with 500 pg of DNA input.
[0095] [Figure 12] FIG. 12 shows the number of loci detected in saliva samples with increasing phenol (a known PCR amplification inhibitor) content from the phenol-chloroform-isoamyl alcohol (PCIA) extraction method.
[0096] [Figure 13] Figure 13 shows the number of loci detected in blood samples isolated from different substrates or methods commonly performed in forensic laboratories, including blood with rust, blood on denim, blood on a swab, and blood with varying levels of heme (a known PCR amplification inhibitor) carryover from Chelex™ extraction.
[0097] [Figure 14] FIG. 14 shows an exemplary schematic diagram of a method for assessing kinship for an individual of interest, e.g., a missing person or a victim of a conflict or disaster, comprising analyzing kinship using DNA profiles, e.g., SNP reports, uploaded to a local server containing at least one DNA profile from a kin.
[0098] [Figure 15] FIG. 15 shows the total number of SNPs detected for each individual sample in an exemplary set of true post-mortem samples.
[0099] [Figure 16] Figure 16 shows the total number of SNPs detected for individual samples within an exemplary related mock postmortem sample set, which includes samples artificially degraded by boiling or are low-input samples with a DI of 0 and less than 1 ng of input DNA.
[0100] [Figure 17] Figure 17 shows the total number of SNPs detected for individual samples in an exemplary related premortem private family sample set including samples from CEPH / Utah involving up to second-degree relationships validated with Coriell, as well as three unrelated samples.
[0101] [Figure 18-1] Figures 18A-18E show the resulting receiver operating characteristic (ROC) curves for 2,000 SNPs (Figure 18A), 4,000 SNPs (Figure 18B), 6,000 SNPs (Figure 18C), 8,000 SNPs (Figure 18D), and 10,000 SNPs (Figure 18E) of first-, second-, and third-degree relatives. [Figure 18-2]Figures 18A-18E show the resulting receiver operating characteristic (ROC) curves for 2,000 SNPs (Figure 18A), 4,000 SNPs (Figure 18B), 6,000 SNPs (Figure 18C), 8,000 SNPs (Figure 18D), and 10,000 SNPs (Figure 18E) of first-, second-, and third-degree relatives.
[0102] [Figure 19-1] Figures 19A-19E show the resulting ROC curves for 2,000 SNPs (Figure 19A), 4,000 SNPs (Figure 19B), 6,000 SNPs (Figure 19C), 8,000 SNPs (Figure 19D), and 10,000 SNPs (Figure 19E) of fourth- or fifth-degree relatives. [Figure 19-2] Figures 19A-19E show the resulting ROC curves for 2,000 SNPs (Figure 19A), 4,000 SNPs (Figure 19B), 6,000 SNPs (Figure 19C), 8,000 SNPs (Figure 19D), and 10,000 SNPs (Figure 19E) of fourth- or fifth-degree relatives.
[0103] [Figure 20-1] Figure 20A shows the distribution of the number of SNPs classified in downsampled datasets of libraries sequenced at 16-plex and 30-plex. [Figure 20-2] Figures 20B and 20C show the number of SNPs classified in sequenced library samples for simulated ante-mortem (AM) samples (Figure 20B), where libraries were generated using 1 ng of intact DNA, with 3 (3-plex), 12 (12-plex), 16 (16-plex), 24 (24-plex), and 32 (32-plex) sample runs, and for simulated post-mortem (PM) samples (Figure 20C), where libraries were generated from degraded DNA samples from cremated, buried, and incinerated bone, dental remains, whole blood, and low-input DNA samples of 50, 100, 250, and 500 pg input DNA, with 3 (3-plex) and 12 (12-plex) sample runs.
[0104] [Figure 21-1] Figures 21A and 21B show allele concordance and heterozygosity for simulated antemortem (mock AM) samples (Figure 21A) and simulated postmortem (mock PM) (Figure 21B) samples from modern teeth, blood, buried bone, modern bone, or low-input DNA. [Figure 21-2] Figures 21A and 21B show allele concordance and heterozygosity for simulated antemortem (mock AM) samples (Figure 21A) and simulated postmortem (mock PM) (Figure 21B) samples from modern teeth, blood, buried bone, modern bone, or low-input DNA.
[0105] [Figure 22-1] Figures 22A-22D show graphical representations of the sensitivity and specificity of the total relatedness coefficients for de-identified GEDMatch samples for first-, second-, third-, fourth-, and fifth-degree relationships where the number of classified SNPs is 2,000 (Figure 22A), 4,000 (Figure 22B), 6,000 (Figure 23C), or 8,000 (Figure 22D). [Figure 22-2] Figures 22A-22D show graphical representations of the sensitivity and specificity of the total relatedness coefficients for de-identified GEDMatch samples for first-, second-, third-, fourth-, and fifth-degree relationships where the number of classified SNPs is 2,000 (Figure 22A), 4,000 (Figure 22B), 6,000 (Figure 23C), or 8,000 (Figure 22D). [Figure 22-3] Figures 22A-22D show graphical representations of the sensitivity and specificity of the total relatedness coefficients for de-identified GEDMatch samples for first-, second-, third-, fourth-, and fifth-degree relationships where the number of classified SNPs is 2,000 (Figure 22A), 4,000 (Figure 22B), 6,000 (Figure 23C), or 8,000 (Figure 22D). [Figure 22-4]Figures 22A-22D show graphical representations of the sensitivity and specificity of the total relatedness coefficients for de-identified GEDMatch samples for first-, second-, third-, fourth-, and fifth-degree relationships where the number of classified SNPs is 2,000 (Figure 22A), 4,000 (Figure 22B), 6,000 (Figure 23C), or 8,000 (Figure 22D).
[0106] [Figure 23] Figure 23 shows the pedigree of the Utah / CEPH 1463 family, consisting of grandparents, parents, and siblings, thereby representing first- and second-degree relationships. Samples were sequenced in 12-, 16-, 24-, and 32-sample pool libraries.
[0107] [Figure 24] Figure 24 shows the distribution of the number of classified SNPs across samples within the Utah / CEPH 1463 family when sequenced with four different numbers of samples per run: 12 samples per run (12-plex), 16 samples per run (16-plex), 24 samples per run (24-plex), and 32 samples per run (32-plex).
[0108] [Figure 25] Figures 25A and 25B show the distribution of relatedness coefficients (Figure 25A) and log base 10 likelihood ratios (LogLR) (Figure 25B) for all pairwise combinations taken from the Utah / CEPH 1463 family and 100 randomly selected samples from the 1000 Genomes Project, which represented unrelated controls. Samples include samples from grandparents (G), parents (P), siblings (S), unrelated controls (U), unrelated grandparents (GU), unrelated parents (PU), and unrelated siblings (SU).
[0109] [Figure 26] Figure 26 depicts a private relative (RF) tree consisting of parents, aunts / aunts, cousins, cousins' children (1C1R), and second cousins.
[0110] [Figure 27] Figure 27A shows the distribution of the number of SNPs classified for private relatives (RF) carrying first-, second-, third-, fourth-, and fifth-degree relationships using 12-plex degraded / low-input DNA samples and 30-plex intact samples. Figure 27B shows the kinship coefficients for pairs of individuals from private relatives (RF), with the corresponding log-likelihood ratios shown above each bar. DETAILED DESCRIPTION OF THE INVENTION
[0111] Detailed Description The practice of the techniques described herein may employ, unless otherwise indicated, conventional techniques and descriptions of molecular biology, cell biology, biochemistry, and sequencing techniques, which are within the skill of those in the art. Specific illustrations of suitable techniques can be had by reference to the examples herein.
[0112] All publications, including patent documents, scientific articles, and databases, referred to in this application are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication was individually incorporated by reference. To the extent that a definition set forth herein contradicts or otherwise conflicts with a definition set forth in a patent, application, published application, or other publication incorporated herein by reference, the definition set forth herein takes precedence over the definition incorporated herein by reference.
[0113] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. overview
[0114] Samples from missing persons or victims of disasters or conflicts can be highly degraded and may not be suitable for whole-genome sequencing (WGS), microarray, or short tandem repeat (STR) analysis. Mitochondrial analysis, while sensitive and potentially suitable for certain situations, only considers maternal inheritance, thereby limiting its use in kinship analysis. Furthermore, current methods for generating DNA profiles for comparison in genetic databases, including genotyping using high-density SNP microarrays and WGS and subsequent linkage of evidentiary samples to distant relatives in the database, require large amounts of high-quality DNA samples and are not designed for use in family searches or for identifying missing persons or victims of disasters or conflicts. These samples may be of low volume and low quality, for example, containing degraded DNA, and data from current methods require extensive imputation to generate results that can be uploaded to search databases. Finally, relatives of missing persons or victims of disasters or conflicts often do not want their genetic data uploaded to public databases. The new and improved methods provided herein overcome these limitations by enabling the use of small amounts and low-quality, e.g., degraded, DNA for the generation of nucleic acid profiles, without the need to upload genetic data to publicly accessible genetic databases, for more efficient genetic analysis than alternative approaches such as WGS or SNP microarrays. Furthermore, the new and improved methods provided herein also include improved methods for performing kinship analysis, requiring fewer calculations to accurately calculate kinship.
[0115] Identification of victims of fatal accidents is necessary for both humanitarian and legal reasons: when the death is not accidental, as in civil and criminal cases, it provides families with a sense of resolution and legitimacy in their search for missing family members. Identification of victims of mass fatality incidents (MFIs) can be difficult given the large number of victims and the impact of the disaster on the victims' physical integrity. MFIs can be caused by accidental events such as disease / famine, earthquakes / tsunamis / hurricanes, plane / train / car crashes, or fires, or by human intent such as war, terrorist attacks, or human rights violations / genocide. Recent events such as the September 11, 2001, terrorist attacks on the World Trade Center in New York City or the December 26, 2004, Boxing Day tsunami caused by an earthquake off the west coast of northern Sumatra, caused unimaginable loss of life and demonstrated the need for efficient procedures and methods for recovering and cataloging the remains of victims, preserving information about the remains, and identifying them.
[0116] The most common methods for identification (called disaster victim identification, or DVI) are fingerprinting, dental comparison (dental examination or radiology), itemization, autopsy examination for evidence of surgical scars / procedures and tattoos, and DNA analysis. Traditional methods such as fingerprinting and dental comparison are often utilized as a first line of defense due to the low labor and cost required and speed of the procedure. However, these methods require antemortem records of fingerprints and dental images for identification. Items can be misleading because victims may have similar jewelry or other personal items. Comparison of surgical procedures and tattoos also requires antemortem medical records or documentation of tattoos or other alterations. Finally, these methods require the victim's remains to be relatively intact. Some MFIs can result in fragmentation and commingling of remains. In cases of fragmentation, DNA analysis of postmortem (PM) samples can not only aid in the identification of missing persons but also help assign multiple remains or body parts to specific individuals. DNA analysis requires antemortem (AM) samples, such as razors, shavers, toothbrushes, or hairbrushes, from the missing person for comparison. If AM samples from the missing person are not available, samples donated by close family members can assist in identification. DNA analysis takes longer than traditional methods and has specific requirements for laboratory cleanliness and chain of custody tracking for the samples analyzed, which can be difficult in field situations.
[0117] When traditional methods fail or the remains are not intact, the success of DNA analysis depends on sample collection, the timing of that collection, storage conditions, and the quantity of sample obtained (as is the case in cases of decomposition and decay). DNA identification relies on non-coding DNA markers used in forensic genomics, including short tandem repeats (STRs), such as the set of 20 autosomal core loci included in the Combined DNA Index System (CODIS), mitochondrial DNA for maternal lineages, or STRs on the Y chromosome (Y-STRs) for paternal lineages. Autosomal markers are preferred in large-scale disasters due to the complexity that multiple family members who share either the same maternal or paternal lineage may be missing.
[0118] STR analysis has been used successfully for many years for identification in criminal, missing person, and paternity cases. The success of this type of analysis is a result of the highly polymorphic nature of these markers and the number of markers that can be multiplexed together for a single analysis. These markers have also been successfully utilized for DVI. In MFI, software solutions are useful for supporting multiple pairwise comparisons of profiles from PM (victim) and AM (self or relative) samples and statistical calculations of the degree of relatedness. Several software packages are available for the analysis of autosomal and Y-STR and mitochondrial DNA data. Not all samples from MFI are suitable for STR analysis. Highly degraded DNA samples do not amplify larger STR markers in commonly used capillary electrophoresis-based (CE) kits, limiting the number of markers that can be typed to perform identification. The development of CE kits utilizing smaller amplicons has assisted CE analysis of degraded DNA (Butler paper). Next-generation sequencing (NGS or massively parallel sequencing) STR assays allow for smaller amplicon sizes and allow for the analysis of more markers within one assay, both of which improve the recovery of information from degraded DNA samples.
[0119] Utilizing STR data is particularly appropriate when AM DNA samples are available from the missing person or close family members, such as first-degree relatives (parents, children, or siblings). Due to the number of false-positive identifications that can occur in unrelated individuals, STR analysis is less successful when only more distant family members are available for comparison. This is especially true when the MFI occurred many years ago and close family members have died. When DNA from more distant relatives (second- and third-degree relatives, such as nieces, nephews, grandchildren, and great-grandchildren) is available for comparison, utilizing more markers, such as single nucleotide polymorphisms (SNPs), can aid in identification.
[0120] The method described herein was developed to interrogate 10,230 forensically relevant SNPs for purposes such as solving cold cases, including identifying missing persons. Previous approaches involved analyzing amplified DNA using NGS for up to three samples per sequencing run. Such approaches were initially designed to classify enough SNPs to detect relationships up to the fifth degree when searching DNA microarray databases such as GEDmatch PRO to assist law enforcement in solving cold cases. To facilitate DVI, which requires a cost-effective, high-throughput solution, a new and improved method was developed that can sequence amplified DNA at a higher complexity to classify enough SNPs with significant overlap to determine relationships to the third degree. As shown in Example 13, simulated PM samples sequenced at a multiplexing of 12 samples per run and simulated AM samples sequenced at a multiplexing of 32 samples per run generated enough classified locus data to confirm relationships up to the third degree without false positive identifications using the method and kinship algorithm described herein.
[0121] The goals of this kinship algorithm were threefold: to confirm relationships in the absence of genealogy or relationship information, which can be important when attempting to assign multiple remains to a single individual or when the victim's relatives are unavailable or unknown; to reduce the false positive rate when confirming relationships, which is often high when calculating likelihood ratios using STR (Alonso et al., Croat Med J 2005, 46, 540-548, the contents of which are incorporated herein by reference in their entirety); and to perform the algorithm in a programming language such as R, a local build of likelihood ratio software such as Familias (Kling et al., Forensic Sci Int Genet 2014, 13, 121-127, doi:10.1016 / j.fsigen.2014.07.004, the contents of which are incorporated herein by reference in their entirety), or Bonaparte's personal account (Slooten et al., Forensic Science International: Genetics Maintaining the privacy of MFI victims requires knowledge of the identity of the victim (only up to two degrees of kinship can be confirmed). Because MFI victims and their families often request privacy in the identification of remains, maintaining genotype information on a private server is necessary. Thus, in some embodiments, the kinship algorithm described herein is localized on a private server to confirm relationships between samples prepared as described herein. The local software did not upload results to a law enforcement database; rather, it maintained the results for review on the private server. Furthermore, in some embodiments, the kinship algorithm described herein can efficiently confirm relationships with full sensitivity and specificity, including up to three degrees of kinship, for degraded / low-input mock PM samples sequenced at 12 plexity and mock reference or AM samples sequenced at 32 plexity.
[0122] Thus, disclosed herein is a method of identifying a person of interest by performing a DNA-based kinship analysis using a DNA profile from the person of interest, e.g., a missing person or a victim of a disaster or conflict, to determine the degree of relatedness between the DNA profile and one or more reference DNA profiles comprising known relatives of the person of interest, thereby identifying the person of interest.
[0123] Provided herein is a method for performing DNA-based kinship analysis, the method comprising: providing a nucleic acid sample from a person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences collectively comprising a plurality of at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplified products, wherein the amplification is performed in one or more multiplex PCR reactions; generating a nucleic acid library from the amplified products; sequencing the nucleic acid library generated from the amplified products; analyzing the sequences of the amplified products; genotyping the plurality of SNPs, thereby generating a DNA profile; and calculating a degree of relatedness between the DNA profile and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included in a reference set of DNA profiles that includes one or more reference DNA profiles from relatives of the person of interest.
[0124] Also provided herein is a method for performing DNA-based kinship analysis, comprising: providing a nucleic acid sample from a person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences collectively comprising a plurality of at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplified products, wherein the amplification is performed in one or more multiplex PCR reactions; generating a nucleic acid library from the amplified products; sequencing the nucleic acid library generated from the amplified products; genotyping the plurality of SNPs, thereby generating a DNA profile; and calculating a degree of relatedness between the DNA profile and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included in a reference set of DNA profiles that includes one or more reference DNA profiles from relatives of the person of interest.
[0125] Also provided herein is a method of constructing a nucleic acid library of a person of interest, comprising: providing a nucleic acid sample from the person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of between at least or about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating a nucleic acid library comprising amplified products, wherein the amplification is performed in one or more multiplex PCR reactions.
[0126] Also provided herein are methods for constructing a nucleic acid library for a reference DNA sample, the method comprising: providing a nucleic acid sample from a relative of a person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences collectively comprising at least or about 2,000 to 50,000 single nucleotide polymorphisms (SNPs), thereby generating a nucleic acid library comprising the amplified products; wherein the amplification is performed in one or more multiplex PCR reactions. In some embodiments, the relative is a first-, second-, third-, fourth-, or fifth-degree relative of the person of interest. In some embodiments, the relative is a first-, second-, or third-degree relative of the person of interest.
[0127] Also provided herein is a method for calculating degree of relatedness, comprising obtaining a DNA profile comprising genotypes for at least or between about 2,000 and 50,000 SNPs, wherein the DNA profile is from a person of interest; and calculating a degree of relatedness between the DNA profile and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included in a reference set of DNA profiles that includes one or more reference DNA profiles from relatives of the person of interest.
[0128] Also provided herein is a method for calculating degree of relatedness, comprising generating a DNA profile comprising genotypes for at least or between about 2,000 and 50,000 SNPs, wherein the DNA profile is from a person of interest; and calculating a degree of relatedness between the DNA profile and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included in a reference set of DNA profiles that includes one or more reference DNA profiles from relatives of the person of interest.
[0129] Also provided herein are nucleic acid libraries constructed using any of the methods described herein, eg, any of the methods for constructing a nucleic acid library described herein.
[0130] Also provided herein are a plurality of primers that specifically hybridize to a plurality of target sequences comprising at least or about 2,000-50,000 single nucleotide polymorphisms (SNPs) in a nucleic acid sample from a person of interest, wherein amplification of the nucleic acid sample using the plurality of primers in one or more multiplex PCR reactions results in amplification products.
[0131] Also provided herein are a plurality of primers that specifically hybridize to a plurality of target sequences comprising between at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs) in a nucleic acid sample from a person of interest and one or more reference samples, wherein the one or more reference samples comprise samples from blood relatives of the person of interest, and wherein amplifying the nucleic acid sample from the person of interest and the one or more reference samples using the plurality of primers in one or more multiplex PCR reactions results in amplification products.
[0132] Also provided herein is a method for constructing a DNA profile, the method comprising: providing a nucleic acid sample from a person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of between at least or about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplified products, wherein the amplification is performed in one or more multiplex PCR reactions; sequencing the amplified products; and genotyping the plurality of SNPs, thereby generating the DNA profile.
[0133] Also provided herein is a method for constructing a DNA profile, the method comprising: providing a nucleic acid sample from a person of interest; providing a nucleic acid sample from a relative of the person of interest; amplifying the nucleic acid sample from the person of interest and the nucleic acid sample from the relative with a plurality of primers that specifically hybridize to a plurality of target sequences collectively comprising a plurality of at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplified products, wherein the amplification is carried out in one or more multiplex PCR reactions; sequencing the amplified products; and genotyping the plurality of SNPs, thereby generating a DNA profile for the person of interest and the relative of the person of interest.
[0134] Also provided herein are DNA profiles constructed using any of the methods described herein, eg, any of the methods for constructing a DNA profile described herein.
[0135] Also provided herein is a method for identifying genetic relatives of a DNA profile, the method comprising: calculating a degree of relatedness between any of the DNA profiles described herein and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included within a reference set of DNA profiles that includes one or more reference DNA profiles from relatives of the person of interest; and generating a family tree that includes the DNA profile in relation to the one or more reference DNA profiles.
[0136] Also provided herein is a method for confirming the identity of a DNA profile, the method comprising: calculating a degree of relatedness between a DNA profile comprising genotypes for at least or between about 2,000 and 50,000 SNPs and one or more reference DNA profiles, wherein the DNA profile is from a person of interest, and the one or more reference DNA profiles are included within a reference set of DNA profiles that includes one or more reference DNA profiles from relatives of the person of interest; and generating a pedigree tree that includes the DNA profile in relation to the one or more reference DNA profiles.
[0137] Also provided herein are kits comprising at least one container means, the at least one container means containing any of a plurality of primers described herein.
[0138] In some of any such embodiments, the relatedness is calculated using a kinship model. In some of such embodiments, the relatedness is calculated using a kinship model that is trained using a PCA method. In some of such embodiments, the PCA method for training the kinship model is or includes PCA.
[0139] In some of such embodiments, the PCA method is PC-AiR. In some embodiments, PC-AiR includes: (1) estimating relatedness coefficients between all pairs of samples in a training database, optionally in a training DNA profile, where pairings with a relatedness coefficient >0.025 are confirmed as closely related and pairings with a relatedness coefficient <-0.025 are confirmed as ancestrally diverged; (2) initializing an unrelated sample set containing all samples; and (3) iteratively: (i) identifying a set in the unrelated sample set that has the most related samples in the unrelated sample set, thereby designating this set as X; (ii) identifying a set of samples in X that has the fewest ancestrally diverged pairings compared to the samples in the unrelated sample set, thereby designating this set as Y; and (iii) terminating the process if Y has 0 samples, or randomly selecting one sample from Y and removing it from U if Y has at least one sample, and repeating starting with step (3)(i).
[0140] In some of such embodiments, the PCA method is a modified PC-AiR. In some embodiments, the modified PC-AiR method includes: (1) estimating relatedness coefficients between all pairs of samples, optionally training DNA profiles, in a training database, where pairings with a relatedness coefficient >0.01 are identified as closely related and pairings with a relatedness coefficient <-0.025 are identified as ancestrally diverged; (2) removing all DNA profiles with ≥5% missing data; and (3) ranking all DNA profiles by assigning each DNA profile a ranking value. In some embodiments, the ranking value is determined based on the number of related DNA profiles in the complete database ranked from smallest to largest, broken down by the number of ancestrally diverged DNA profiles in the complete database ranked from largest to smallest. In some embodiments, step (3) includes iterating through the ranked DNA profiles, and for each DNA profile, (i) if the DNA profile is not already in the relevant sample set, adding it to the unrelated sample set and adding all relevant DNA profiles to the relevant sample set, and (ii) if the DNA profile is already in the relevant sample set, skipping to the next DNA profile and repeating starting at step (3)(i).
[0141] In some of such embodiments, calculating the degree of relatedness includes calculating a coefficient of kinship using PC-Relate. In some embodiments, the degree of relatedness is calculated by providing a DNA profile of the person of interest as input to PC-Relate. In some of such embodiments, the degree of relatedness is calculated by providing a kinship model and the DNA profile of the person of interest as input to PC-Relate. In some of such embodiments, calculating the degree of relatedness includes calculating a likelihood ratio. In some embodiments, calculating the likelihood ratio includes comparing a plurality of SNPs between the DNA profile and one or more reference DNA profiles. In some embodiments, calculating the likelihood ratio includes comparing a set of SNPs comprising kinship SNPs from among the plurality of SNPs between the DNA profile and one or more reference DNA profiles. In some embodiments, the person of interest is biologically male, and the method further includes calculating a likelihood ratio of sharing a Y chromosome between the DNA profile and one or more reference DNA profiles. In some embodiments, calculating the likelihood ratio of sharing a Y chromosome includes comparing a set of SNPs comprising one or more Y-SNPs between the DNA profile and one or more reference DNA profiles. In some embodiments, the one or more Y-SNPs are included in the plurality of SNPs. In some embodiments, the one or more Y-SNPs include at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 81, 82, 83, 84, or 85 Y-SNPs. In some embodiments, the one or more Y-SNPs include 85 Y-SNPs. In some embodiments, calculating the likelihood ratio of sharing a Y chromosome comprises dividing the probability that the DNA profile and the reference DNA profile from among one or more reference DNA profiles share a Y chromosome by the probability that the DNA profile and the reference DNA profile do not share a Y chromosome, based on the genotype of one or more Y-SNPs.
[0142] In some embodiments, calculating the likelihood ratio comprises dividing the probability that the DNA profile and a reference DNA profile from among the one or more reference DNA profiles are related by the probability that the DNA profile and the reference DNA profile are unrelated, based on the genotypes of the plurality of SNPs.
[0143] In some of any such embodiments, calculating the likelihood of sharing chromosome Y includes calculating the log-likelihood by providing a DNA profile of the person of interest and a match profile as input to identify matching chromosomes. Samples and sample processing
[0144] In some embodiments, the samples disclosed herein can be or include any suitable biological sample, or a sample derived therefrom. In some embodiments, the samples described herein are processed and amplified using any known suitable method to complement the methods described herein. Exemplary samples, sample processing methods, and sample amplification methods are described below. A. Nucleic Acid Sample
[0145] The nucleic acid sample disclosed herein can be derived from any biological sample, for example, any biological sample from a person of interest. The biological sample can be derived from blood, buccal swabs, hair, teeth, bone, skin, tissue, and / or semen, or any other source for obtaining DNA from a person of interest. In some embodiments, the nucleic acid sample is derived from a biological sample that is or contains blood, hair, teeth, bone, semen, skin, or sperm. In some embodiments, the nucleic acid sample is derived from a tissue sample. In some embodiments, the biological sample is a DNA sample. In some embodiments, the nucleic acid sample comprises DNA. In some embodiments, the DNA is genomic DNA (gDNA). In some embodiments, the nucleic acid sample from the person of interest comprises genomic DNA, and / or a reference DNA sample, for example, a nucleic acid sample from a relative of the person of interest, comprises genomic DNA. In some embodiments, the nucleic acid sample from the person of interest and / or the nucleic acid sample of a relative of the person of interest comprises genomic DNA. The DNA from which the nucleic acid sample can be derived can be intact or partially degraded. The DNA from which the nucleic acid sample can be obtained may be damaged, degraded, or inhibited due to, but not limited to, degradation of raw materials, variations in extraction, storage procedures, or environmental exposure. In some embodiments, the DNA is damaged due to calcium inhibition, cremation, incineration, and embalming. In some embodiments, the methods described herein include providing a nucleic acid sample from a person of interest.
[0146] In some embodiments, the DNA from which the nucleic acid sample can be obtained is a low-quality and / or low-quality DNA sample. In some embodiments, the DNA from which the nucleic acid sample can be obtained is a low-quality and low-quality DNA sample. In some embodiments, the low-quality DNA sample comprises low-quality nucleic acid molecules. In some embodiments, the low-quality nucleic acid molecules are degraded DNA, such as genomic DNA, and / or fragmented DNA, such as genomic DNA.
[0147] The quality of a nucleic acid, e.g., DNA, sample can be determined by calculating the degradation index (DI). DI is calculated by dividing the concentration of small DNA targets by the concentration of large DNA targets (DI = concentration of small DNA targets / concentration of large DNA targets). In general, a DI value less than 1 typically indicates that the nucleic acid, e.g., DNA, is not degraded, is not a low-quality sample, and / or is a high-quality sample; a DI value between 1 and 10 typically indicates that the nucleic acid, e.g., DNA, has a low to moderate amount of degradation; and a DI value greater than 10 typically indicates that the nucleic acid, e.g., DNA, is highly degraded.
[0148] In some embodiments, low quality nucleic acid molecules have a D of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200 or more 1 or has a DI of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195 or 200 or more. In some embodiments, the low quality nucleic acid molecules are at least 2 and 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200 or or greater, or less than 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200 or more. In some embodiments, low quality nucleic acid molecules have a DI of 2 or more, or at least 2 or more. In some embodiments, low quality nucleic acid molecules have a DI of 5 or more, or at least 5 or more. In some embodiments, low quality nucleic acid molecules have a DI of 10 or more, or at least 10 or more. In some embodiments, low quality nucleic acid molecules have a DI of 20 or more, or at least 20 or more.In some embodiments, low quality nucleic acid molecules have a DI between 1 and 200. In some embodiments, low quality nucleic acid molecules have a DI between 1 and 175. In some embodiments, low quality nucleic acid molecules have a DI of at least 1 and 158.3 or less than 158.3. In some embodiments, low quality nucleic acid molecules have a DI between 2 and 200. In some embodiments, low quality nucleic acid molecules have a DI between 2 and 175. In some embodiments, low quality nucleic acid molecules have a DI of at least 2 and 158.3 or less than 158.3. In some embodiments, low quality nucleic acid molecules have a DI between 5 and 200. In some embodiments, low quality nucleic acid molecules have a DI between 5 and 175. In some embodiments, low quality nucleic acid molecules have a DI of at least 5 and 158.3 or less than 158.3. In some embodiments, low quality nucleic acid molecules have a DI between 10 and 200. In some embodiments, low quality nucleic acid molecules have a DI between 10 and 175. In some embodiments, low quality nucleic acid molecules have a DI of at least 10 and 158.3 or less than 158.3.
[0149] In some embodiments, low quality nucleic acid molecules have a DI of between or about 1-10, between or about 1-50, between or about 1-50, between or about 1-100, between or about 1-200, between or about 2-10, between or about 2-10, between or about 2-50, between or about 2-50, between or about 2-100, between or about 2-100, between or about 200, between or about 5-10, between or about 5-10, between or about 5-50, between or about 5-50, between or about 5-100, between or about 5-100, between or about 5-200.
[0150] In some embodiments, the DNA from which the nucleic acid sample is obtained is a high quality nucleic acid sample. In some embodiments, a high quality nucleic acid sample has a DI of less than 1.
[0151] In some embodiments, the nucleic acid sample comprises one or more enzyme inhibitors. In some embodiments, the one or more enzyme inhibitors comprise one or more inhibitors selected from the group consisting of hematin, humic acid (e.g., heme), humic acid, indigo, tannic acid, collagen, calcium, and hydroxyapatite. In some embodiments, the one or more enzyme inhibitors comprise heme.
[0152] In some embodiments, the nucleic acid sample is from a person of interest, such as a missing person or a victim of a disaster or conflict. In some embodiments, the person of interest is a missing person. A missing person may be missing for any reason and may be voluntarily or involuntarily missing. For example, in some embodiments, the missing person is involuntarily missing and has been kidnapped or abducted. In some embodiments, the missing person is voluntarily missing and has run away, evaded detection, or is otherwise in hiding.
[0153] In some embodiments, the nucleic acid sample is from a reference DNA sample. In some embodiments, the reference DNA sample is from a relative of the person of interest. Thus, in some embodiments, the nucleic acid sample is from a relative of the person of interest, such as a first-degree relative, a second-degree relative, a third-degree relative, a fourth-degree relative, or a fifth-degree relative of the person of interest. In some embodiments, one or more of the one or more reference DNA profiles are derived from a reference DNA sample, for example, a reference DNA sample from a relative of the person of interest.
[0154] In some embodiments, the person of interest is a victim of a disaster or conflict. The victim of a disaster or conflict can be a victim of any type of disaster or conflict. For example, in some embodiments, the victim of a disaster or conflict is a victim of a disaster such as a hurricane, tornado, storm, fire including wildfire / forest fire, tsunami, earthquake, flood, volcanic eruption, avalanche, etc. In some embodiments, the disaster is a natural disaster. As used herein, "natural disaster" refers to any disaster resulting from the Earth's natural processes, such as those associated with meteorological and / or geological events, e.g., hurricanes, floods, storms, tsunamis, earthquakes, volcanic eruptions, etc. In some embodiments, the disaster is a non-natural disaster. As used herein, "non-natural disaster" refers to any disaster other than a natural disaster, including those caused by human influence, including disasters involving motor vehicles, airplanes, ships, and trains, disasters involving the collapse of buildings, roads, mines, and bridges, and disasters involving the burning of buildings, among other disasters caused by human influence. In some embodiments, the victim of a disaster or conflict is a victim of a conflict, such as war or other conflict between groups of people. As used herein, "conflict" refers to any conflict between different countries or states, or between different groups within a country or state, for example, a military conflict, e.g., war, or terrorist attack, or any other conflict between groups that results in the death and / or injury of persons.
[0155] In some embodiments, the person of interest is biologically female. In some embodiments, the person of interest is biologically male.
[0156] In some embodiments, the nucleic acid sample is derived from a buccal swab, paper, fabric, e.g., denim, or other substrate or object impregnated with saliva, blood, sperm, or other bodily fluids, or containing hair or skin cells. In some embodiments, the object impregnated with saliva, blood, sperm, or other bodily fluids, or containing hair or skin cells, is a personal object, such as a toothbrush or hairbrush. In some embodiments, the nucleic acid sample is derived from an object containing hair or skin cells, e.g., a hairbrush or toothbrush. In some embodiments, the nucleic acid sample is derived from a personal object, e.g., a toothbrush or hairbrush. In some embodiments, the nucleic acid sample is derived from a toothbrush or hairbrush. In some embodiments, the personal object is an object used by and / or associated with the person from whom the nucleic acid sample is derived, such that the person's nucleic acids are present on or within the object.
[0157] In some embodiments, the nucleic acid sample is from a crime scene, such as a murder, an assault, such as a sexual assault, or a burglary, or any other crime where participant identification is required, hi some embodiments, the nucleic acid sample is from a sexual assault.
[0158] In some embodiments, the nucleic acid sample is obtained at or about 30 minutes, at or about 1 hour, or at or about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 hours or more after the sample containing the nucleic acid sample is deposited by its source, e.g., a human subject. In some embodiments, the nucleic acid sample is purified within about 3 hours, 9 hours, 12 hours, 15 hours, 18 hours, 21 hours, 22 hours, 24 hours, 36 hours, 48 hours, 3 days, 4 days, 5 days, 6 days, 7 days, 2 weeks, 3 weeks, 4 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 1 year after the sample containing the nucleic acid sample is deposited by its source, e.g., a human subject. , 2, 3, or 4 years or more, or less than about 3 hours, 9 hours, 12 hours, 15 hours, 18 hours, 21 hours, 22 hours, 24 hours, 36 hours, 48 hours, 3 days, 4 days, 5 days, 6 days, 7 days, 2 weeks, 3 weeks, 4 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 1 year, 2 years, 3 years, or 4 years or more. In some embodiments, the nucleic acid sample is obtained 24 hours or less, e.g., 22 hours or less, after the sample containing the nucleic acid sample is deposited by its source, e.g., a human subject.
[0159] In some embodiments, the nucleic acid sample comprises at or about 3 pg to 100 ng of DNA, e.g., genomic DNA, or at or about 50 pg to 100 ng of DNA, e.g., genomic DNA. In some embodiments, the nucleic acid sample comprises at or about 100 pg to 5 ng of DNA, e.g., genomic DNA. In some embodiments, the nucleic acid sample comprises at or about 1 ng of DNA, e.g., genomic DNA. In some embodiments, the nucleic acid sample comprises at or about 3 pg to 100 ng of DNA, e.g., genomic DNA.
[0160] In some embodiments, the nucleic acid sample contains 3 pg to 100 ng or about 3 pg to 100 ng of DNA, e.g., genomic DNA. In some embodiments, the nucleic acid sample contains 10 pg to 100 ng or about 10 pg to 100 ng of DNA, e.g., genomic DNA, or 10 pg to 5 ng or about 10 pg to 5 ng of DNA, e.g., genomic DNA. In some embodiments, the nucleic acid sample contains 10 pg to 10 ng or about 10 pg to 10 ng, 10 pg to 5 ng or about 10 pg to 5 ng, 25 pg to 10 ng or about 25 pg to 10 ng, 25 pg to 5 ng or about 25 pg to 5 ng, 50 pg to 10 ng or about 50 pg to 10 ng, or 50 pg to 5 ng or about 50 pg to 5 ng of DNA, e.g., genomic DNA.
[0161] In some embodiments, the nucleic acid sample contains 3 pg to 5 ng or about 3 pg to about 5 ng of DNA, e.g., genomic DNA. In some embodiments, the nucleic acid sample contains 50 pg to 5 ng or about 50 pg to about 5 ng of DNA, e.g., genomic DNA. In some embodiments, the nucleic acid sample comprises 2.5 pg, 3 pg, 4 pg, 5 pg, 6 pg, 7 pg, 8 pg, 9 pg, 10 pg, 15 pg, 20 pg, 25 pg, 30 pg, 35 pg, 40 pg, 45 pg, 50 pg, 55 pg, 60 pg, 70 pg, 75 pg, 80 pg, 85 pg, 90 pg, 95 pg, 100 pg, 125 pg, 150 pg, 175 pg, 200 pg, 225 pg, 250 pg, 275 pg, 300 pg, 325 pg, 350 pg, 375 pg, 400 pg, 450 ... pg, 420pg, 425pg, 450pg, 475pg, 500pg, 600pg, 700pg, 800pg, 900pg, 1ng, 1.1ng, 1.2ng, 1.3ng, 1.4ng, 1.5ng, 1.6ng, 1.7ng, 1.8n g, 1.9ng, 2ng, 2.1ng, 2.2ng, 2.3ng, 2.4ng, 2.5ng, 2.6ng, 2.7ng, 2.8ng, 2.9ng, 3ng, 3.25ng, 3.5ng, 3.75ng, 4ng, 4.25ng, 4.5ng, 4.75ng or 5ng, or approximately 2.5pg, 3pg, 4pg, 5pg, 6pg, 7pg, 8pg, 9pg, 10pg, 15pg, 20pg, 25pg, 30pg, 35pg, 40pg, 45pg, 50pg, 55pg, 60pg, 70pg, 75pg, 80pg, 85pg, 90pg, 95pg, 100pg, 125pg, 150pg, 175pg, 200pg, 225pg, 250pg, 275pg, 300pg, 325pg, 350pg, 375pg, 400pg , 420pg, 425pg, 450pg, 475pg, 500pg, 600pg, 700pg, 800pg, 900pg, 1ng, 1.1ng, 1.2ng, 1.3ng, 1.4ng, 1.5ng, 1.6ng, 1.7ng, 1.8ng, 1.9ng, 2ng, 2.1ng, 2.2ng, 2.3ng, 2.4ng, 2.5ng, 2.6ng, 2.7ng, 2.8ng, 2.9ng, 3ng, 3.25ng, 3.5ng, 3.75ng, 4ng, 4.25ng, 4.5ng, 4.In some embodiments, the nucleic acid sample contains 75 ng or 5 ng of DNA, e.g., genomic DNA, or any of the values between. In some embodiments, the nucleic acid sample contains 3 pg to 10 ng or about 3 pg to 10 ng, 3 pg to 5 ng or about 3 pg to 5 ng, 3 pg to 4 ng or about 3 pg to 4 ng, 3 pg to 3 ng or about 3 pg to 3 ng, 3 pg to 2 ng or about 3 pg to 2 ng, 10 pg to 10 ng or about 10 pg to 10 ng, 10 pg to 5 ng or about 10 pg to 5 ng, 10 pg to 4 ng or about 10 pg to 4 ng, 10 pg to 3 ng or about 10 pg to 3 ng, 10 pg to 2 ng or about 10 pg to 2 ng, 25 pg g~10ng or about 25pg~10ng, 25pg~5ng or about 25pg~5ng, 25pg~4ng or about 25pg~4ng, 25pg~3ng or about 25pg~3ng, 25pg~2ng or about 25pg~2ng, 40pg~10ng or about 40pg~10ng, 40pg~5ng or about 40pg~5ng, 40pg~4ng or about 40pg~4ng, 40pg~3ng or about 40pg~3ng, 40pg~2ng or about 40pg~2ng, 50pg~10ng or about 50 pg~10ng, 50pg~5ng or approximately 50pg~5ng, 50pg~4ng or approximately 50pg~4ng, 50pg~3ng or approximately 50pg~3ng, 50pg~2ng or approximately 50pg~2ng, 10pg~2ng or approximately 10pg~2ng, 10pg~1.5ng or approximately 10pg~1.5ng, 10pg~1ng or approximately 10pg~1ng, 20pg~2ng or approximately 20pg~2ng, 20pg~1.5ng or approximately 20pg~1.5ng, 20pg~1ng or approximately 20pg~1ng, 2 5pg~2ng or approximately 25pg~2ng, 25pg~1.5ng or approximately 25pg~1.5ng, 25pg~1ng or approximately 25pg~1ng, 30pg~2ng or approximately 30pg~2ng, 30pg~1.5ng or approximately 30pg~1.5ng, 30pg~1ng or approximately 30pg~1ng, 35pg~2ng or approximately 35pg~2ng, 35pg~1.5ng or approximately 35pg~1.5ng, 35pg~1ng or approximately 35pg~1ng, 40pg~2ng or approximately 40pg~2ng, 40pg~1.This includes 5ng or about 40pg to 1.5ng, 40pg to 1ng or about 40pg to 1ng, 45pg to 2ng or about 45pg to 2ng, 45pg to 1.5ng or about 45pg to 1.5ng, 45pg to 1ng or about 45pg to 1ng, 50pg to 2ng or about 50pg to 2ng, 50pg to 1.5ng or about 50pg to 1.5ng, and 50pg to 1ng or about 50pg to 1ng. B. Sample Processing and Amplification
[0162] In some embodiments, methods provided herein are methods for performing DNA-based kinship analysis, comprising providing a nucleic acid sample from a person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplification products, wherein the amplification is performed in one or more multiplex PCR reactions.
[0163] Various steps were performed to prepare or process nucleic acid samples for and / or during the assay. Unless otherwise indicated, the preparation or processing steps described below generally can be combined in any manner and in any order to appropriately prepare or process a particular sample for analysis and / or sequencing as disclosed herein.
[0164] In some embodiments, the amount of nucleic acid sample provided is 1 ng of genomic DNA, about 1 ng of genomic DNA, or less than 1 ng of genomic DNA. In some embodiments, the methods disclosed herein include amplifying the genomic DNA. In some embodiments, the amplification of the genomic DNA includes one or more multiplex polymerase chain reactions (PCRs) including a plurality of primers, thereby generating an amplification product. In some embodiments, the amplification of the genomic DNA includes a single multiplex PCR reaction. In some embodiments, the amplification of the genomic DNA includes two multiplex PCR reactions. In some embodiments, the amplification of the genomic DNA includes three multiplex PCR reactions. In some embodiments, the amplification of the genomic DNA includes four multiplex PCR reactions.
[0165] In some embodiments, one or more primers in the plurality of primers are designed according to the non-existent design strategy described in International Publication No. WO 2015 / 126766, which is incorporated herein by reference in its entirety. In some embodiments, one or more primers in the plurality of primers are at least 24 nucleotides in length, and / or have a melting temperature of less than 60°C, and / or are AT-rich with an AT content of at least 60%. In some embodiments, one or more primers in the plurality of primers comprise at least 24 nucleotides in length that hybridize to the target sequence, and / or have a melting temperature of 50°C to 60°C, and / or are AT-rich with an AT content of at least 60%. In some embodiments, one or more primers in the plurality of primers have a melting temperature of less than 58°C or less than 54°C.
[0166] In some embodiments, genomic DNA may be amplified for multiple cycles using multiple primers that hybridize to and / or tag multiple target sequences collectively comprising a plurality of at least or about 2,000-50,000 single nucleotide polymorphisms (SNPs), or at least or about 5,000-50,000 SNPs. In some embodiments, genomic DNA may be amplified for multiple cycles using multiple primers that hybridize to and / or tag multiple target sequences that collectively contain at least between 2,000 and 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, or 50,000 SNPs, or between about 2,000 and 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, or 50,000 SNPs. In some embodiments, genomic DNA may be amplified for multiple cycles using multiple primers that hybridize to and / or tag multiple target sequences that collectively contain at least, or between about, 5,000-15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, or 50,000 SNPs. In some embodiments, genomic DNA can be amplified for multiple cycles using multiple primers that hybridize to and / or tag multiple target sequences that collectively contain at least between 10,000-11,000 SNPs, or between about 10,000-11,000 SNPs.In some embodiments, the genomic DNA contains at least between 2,000 and 11,000 SNPs, between 3,000 and 11,000 SNPs, between 4,000 and 11,000 SNPs, between 5,000 and 11,000 SNPs, between 5,500 and 11,000 SNPs, between 6,000 and 11,000 SNPs, between 7,000 and 15,000 SNPs, between 7,000 and 14,000 SNPs, between 7,000 and 13,000 SNPs, between 7,000 and 12,000 SNPs, between 7,000 and 11,000 SNPs, 000 SNPs, 8,000 to 15,000 SNPs, 8,000 to 14,000 SNPs, 8,000 to 13,000 SNPs, 8,000 to 12,000 SNPs, 8,000 to 11,000 SNPs, 9,000 to 15,000 SNPs, 9,000 to 14,000 SNPs, 9,000 to 13,000 SNPs, 9,000 to 12,000 SNPs or 9,000 to 11,000 SNPs, or about 2,000 to 11,000 SNPs. NPs, SNPs between 3,000 and 11,000, SNPs between 4,000 and 11,000, SNPs between 5,000 and 11,000, SNPs between 5,500 and 11,000, SNPs between 6,000 and 11,000, SNPs between 7,000 and 15,000, SNPs between 7,000 and 14,000, SNPs between 7,000 and 13,000, SNPs between 7,000 and 12,000, SNPs between 7,000 and 11,000, SNPs between 8,000 and 15,000, SNPs between 8,000 and 14,000 The target sequence may be amplified for multiple cycles using multiple primers that hybridize to and / or tag multiple target sequences that collectively contain between 0 SNPs, between 8,000 and 13,000 SNPs, between 8,000 and 12,000 SNPs, between 8,000 and 11,000 SNPs, between 9,000 and 15,000 SNPs, between 9,000 and 14,000 SNPs, between 9,000 and 13,000 SNPs, between 9,000 and 12,000 SNPs, or between 9,000 and 11,000 SNPs.In some embodiments, genomic DNA can be amplified multiple times using a plurality of primers that hybridize to and / or tag a plurality of target sequences that collectively comprise at least 6,000-11,000 SNPs, or about 6,000-11,000 SNPs. In some embodiments, the plurality of SNPs comprises 2,639 or about 2,639 SNPs. In some embodiments, genomic DNA can be amplified multiple times using a plurality of primers that hybridize to and / or tag a plurality of target sequences that collectively comprise at least 10,230 or about 10,230 SNPs.
[0167] In some embodiments, the plurality of SNPs comprises at least between 2,000 and 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, or 50,000 SNPs, or between about 2,000 and 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, or 50,000 SNPs. In some embodiments, the plurality of SNPs comprises at least between 5,000 and 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, or 50,000 SNPs, or between about 5,000 and 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, or 50,000 SNPs. In some embodiments, the plurality of SNPs comprises at least between 6,000 and 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, or 50,000 SNPs, or between about 6,000 and 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, or 50,000 SNPs. In some embodiments, the plurality of SNPs is at least between 2,000 and 11,000 SNPs, between 3,000 and 11,000 SNPs, between 4,000 and 11,000 SNPs, between 5,000 and 11,000 SNPs, between 5,500 and 11,000 SNPs, between 6,000 and 11,000 SNPs, between 7,000 and 15,000 SNPs, between 7,000 and 14,000 SNPs, between 7,000 and 13,000 SNPs, between 7,000 and 12,000 SNPs, between 7,000 and 13,000 SNPs, between 7,000 and 14,000 SNPs, between 7,000 and 15,000 SNPs, between 7,000 and 16,000 SNPs, between 7,000 and 17,000 SNPs, between 7,000 and 18,000 SNPs, between 7,000 and 19 ... Between 11,000 SNPs, 8,000 and 15,000 SNPs, 8,000 and 14,000 SNPs, 8,000 and 13,000 SNPs, 8,000 and 12,000 SNPs, 8,000 and 11,000 SNPs, 9,000 and 15,000 SNPs, 9,000 and 14,000 SNPs, 9,000 and 13,000 SNPs, 9,000 and 12,000 SNPs, or 9,000 and 11,000 SNPs, or between about 2,000 and 11,000 SNPs.000 SNPs, 3,000-11,000 SNPs, 4,000-11,000 SNPs, 5,000-11,000 SNPs, 5,500-11,000 SNPs, 6,000-11,000 SNPs, 7,000-15,000 SNPs, 7,000-14,000 SNPs, 7,000-13,000 SNPs, 7,000-12,000 SNPs, 7,000-11,000 SNPs The plurality of SNPs may comprise 8,000-15,000 SNPs, 8,000-14,000 SNPs, 8,000-13,000 SNPs, 8,000-12,000 SNPs, 8,000-11,000 SNPs, 9,000-15,000 SNPs, 9,000-14,000 SNPs, 9,000-13,000 SNPs, 9,000-12,000 SNPs, or 9,000-11,000 SNPs. In some embodiments, the plurality of SNPs comprises at or about 2,639 SNPs. In some embodiments, the plurality of SNPs comprises at or about 10,230 SNPs. In some embodiments, the plurality of SNPs is at least between 2,000 and 50,000 SNPs, between 5,000 and 50,000 SNPs, between 5,000 and 45,000 SNPs, between 5,000 and 40,000 SNPs, between 5,000 and 35,000 SNPs, between 5,000 and 30,000 SNPs, between 5,000 and 25,000 SNPs, between 5,000 and 20,000 SNPs, between 6,000 and 50,000 SNPs, between 6,000 and 45,000 SNPs, between 6,000 and 40,000 SNPs, SNPs between 6,000 and 35,000, SNPs between 6,000 and 30,000, SNPs between 6,000 and 25,000, SNPs between 6,000 and 20,000, SNPs between 7,000 and 50,000, SNPs between 7,000 and 45,000, SNPs between 7,000 and 40,000, SNPs between 7,000 and 35,000, SNPs between 7,000 and 30,000, SNPs between 7,000 and 25,000, SNPs between 7,000 and 20,000, SNPs between 8,000 and 50,000, SNPs between 8,000 and 45,000 SNPs, 8,000-40,000 SNPs, 8,000-35,000 SNPs, 8,000-30,000 SNPs, 8,000-25,000 SNPs, 8,000-20,000 SNPs, 9,000-50,000 SNPs, 9,000-45,000 SNPs, 9,000-40,000 SNPs, 9,000-35,000 SNPs, 9,000-30,000 SNPs, 9,000-25,000 SNPs or 9,000-20,000 SNPs 0 SNPs, or approximately between 2,000 and 50,000 SNPs, between 5,000 and 50,000 SNPs, between 5,000 and 45,000 SNPs, between 5,000 and 40,000 SNPs, between 5,000 and 35,000 SNPs, between 5,000 and 30,000 SNPs, between 5,000 and 25,000 SNPs, between 5,000 and 20,000 SNPs, between 6,000 and 50,000 SNPs, between 6,000 and 45,000 SNPs, between 6,000 and 40,000 SNPs, between 6,000 and 35,000 SNPs SNPs between 6,000 and 30,000, SNPs between 6,000 and 25,000, SNPs between 6,000 and 20,000, SNPs between 7,000 and 50,000, SNPs between 7,000 and 45,000, SNPs between 7,000 and 40,000, SNPs between 7,000 and 35,000, SNPs between 7,000 and 30,000, SNPs between 7,000 and 25,000, SNPs between 7,000 and 20,000, SNPs between 8,000 and 50,000, SNPs between 8,000 and 45,000 , 8,000 to 40,000 SNPs, 8,000 to 35,000 SNPs, 8,000 to 30,000 SNPs, 8,000 to 25,000 SNPs, 8,000 to 20,000 SNPs, 9,000 to 50,000 SNPs, 9,000 to 45,000 SNPs, 9,000 to 40,000 SNPs, 9,000 to 35,000 SNPs, 9,000 to 30,000 SNPs, 9,000 to 25,000 SNPs, or 9,000 to 20,000 SNPs.
[0168] In some embodiments, the plurality of SNPs is at least between 2,000 and 11,000 SNPs, between 2,500 and 11,000 SNPs, between 3,000 and 11,000 SNPs, between 3,500 and 11,000 SNPs, between 4,000 and 11,000 SNPs, between 4,500 and 11,000 SNPs, between 5,000 and 11,000 SNPs, between 5,550 and 11,000 SNPs. SNPs, between 6,000 and 11,000 SNPs, between 6,500 and 11,000 SNPs, between 7,000 and 11,000 SNPs, between 7,500 and 11,000 SNPs, between 8,000 and 11,000 SNPs, between 8,500 and 11,000 SNPs, between 9,000 and 11,000 SNPs, between 9,500 and 11,000 SNPs or between 10,000 and 11,000 SNPs 0 SNPs, or approximately between 2,000 and 11,000 SNPs, between 2,500 and 11,000 SNPs, between 3,000 and 11,000 SNPs, between 3,500 and 11,000 SNPs, between 4,000 and 11,000 SNPs, between 4,500 and 11,000 SNPs, between 5,000 and 11,000 SNPs, between 5,550 and 11,000 SNPs, between 6,000 and 1 Includes between 1,000 SNPs, between 6,500 and 11,000 SNPs, between 7,000 and 11,000 SNPs, between 7,500 and 11,000 SNPs, between 8,000 and 11,000 SNPs, between 8,500 and 11,000 SNPs, between 9,000 and 11,000 SNPs, between 9,500 and 11,000 SNPs, or between 10,000 and 11,000 SNPs.
[0169] In some embodiments, the plurality of SNPs comprises SNPs selected from one or more of the group consisting of kinship SNPs, ancestry SNPs, identity SNPs, phenotypic SNPs, X-SNPs, and Y-SNPs. In some embodiments, the plurality of SNPs comprises kinship SNPs, ancestry SNPs, identity SNPs, phenotypic SNPs, X-SNPs, and Y-SNPs. In some embodiments, the plurality of SNPs comprises kinship SNPs. In some embodiments, the plurality of SNPs comprises Y-SNPs. In some embodiments, the plurality of SNPs comprises kinship SNPs and Y-SNPs.
[0170] In some of such embodiments, the plurality of SNPs comprises one or more microhaplotypes.Therefore, in some embodiments, microhaplotype is a type of SNP that is included in the plurality of SNPs.In some embodiments, each microhaplotype comprises one or more SNPs that are shared in a single amplicon or that are shared close to each other on the genome.Generally, microhaplotype is a biomarker that represents a combination of multiple alleles, for example, a plurality of SNP-based allele markers, and is typically less than 300 nucleotides in length.
[0171] In some embodiments, the SNPs do not include SNPs with known medical associations, such as SNPs associated with known medical conditions, or SNPs with low minor allele frequency. By excluding SNPs with known medical associations, such as SNPs associated with known medical conditions, or SNPs with low minor allele frequency, privacy concerns are limited and genetic health data is protected.
[0172] In some embodiments, the SNPs include SNPs filtered in multiple genotype samples. In some embodiments, the SNPs are selected from categories including ancestral SNPs, identity SNPs, kinship SNPs, phenotypic SNPs, X-SNPs, and Y-SNPs. In some embodiments, the ancestral SNPs include 10-100 or about 10-100 SNPs. In some embodiments, the identity SNPs include 10-200 or about 10-200 SNPs. In some embodiments, the kinship SNPs include 7,000-12,000 or about 7,000-12,000 SNPs. In some embodiments, the phenotypic SNPs include 1-50 or about 1-50 SNPs. In some embodiments, the X-SNPs include 10-200 or about 10-200 SNPs. In some embodiments, the Y-SNPs include 10-200 or about 10-200 SNPs. In some embodiments, ancestral SNPs comprise at or about 0-10% of the total number of SNPs. In some embodiments, identity SNPs comprise at or about 0-10% of the total number of SNPs. In some embodiments, related SNPs comprise at or about 80-100% of the total number of SNPs. In some embodiments, at least or at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the plurality of SNPs are related SNPs. In some embodiments, at least 85% or at least about 85% of the plurality of SNPs are related SNPs. In some embodiments, at least 90% or at least about 90% of the plurality of SNPs are related SNPs. In some embodiments, at least 95% or at least about 95% of the plurality of SNPs are related SNPs. In some embodiments, at least 99% or at least about 99% of the plurality of SNPs are related SNPs. In some embodiments, 100% of the plurality of SNPs are related SNPs.In some embodiments, phenotypic SNPs comprise 0-5% or about 0-5% of the total number of SNPs. In some embodiments, X-SNPs comprise 0-5% or about 0-5% of the total number of SNPs. In some embodiments, Y-SNPs comprise 0-5% or about 0-5% of the total number of SNPs. In some embodiments, the SNPs do not include medically beneficial SNPs or minor allele frequency SNPs. The tag region can be any sequence, such as a universal tag region, capture tag region, amplification tag region, sequencing tag region, or UMI tag region.
[0173] In some embodiments, target sequences are purified and enriched to generate a library of the original DNA sample, also referred to as a nucleic acid library. In some embodiments, purification involves combining purification beads with enzymes to purify the amplified target from other reaction components. In some embodiments, the purified target sequences are enriched by amplifying the DNA and adding UDI adapters and sequences required for cluster generation. The UDI adapters can tag the DNA with a unique combination of sequences that identify each sample for analysis.
[0174] In some embodiments, a nucleic acid library is generated from amplification products, including amplification products produced by any of the methods or embodiments described herein. Thus, in some embodiments, the nucleic acid library comprises amplification products generated by amplifying a nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of at least 5,000 to 50,000, or about 5,000 to 50,000, or at least 2,000 to 50,000, or about 2,000 to 50,000 SNPs.
[0175] In some embodiments, nucleic acid or DNA libraries are normalized for quantification and quality checks, and pooled by combining equal amounts of normalized libraries to create a pool of libraries that can be sequenced together on the same flow cell. In some embodiments, quantification involves the use of fluorescent quantification methods. In some embodiments, quantification involves quantitative PCR. After pooling the DNA libraries, they can be denatured and diluted using sodium hydroxide (NaOH)-based methods, and sequencing controls can be added.
[0176] In some embodiments, the nucleic acid library is quantified, normalized, denatured, and diluted according to the instructions set forth in the Forenseq Kintelligence Kit User Guide (Verogen PN:V16000120, the contents of which are incorporated herein by reference in their entirety).
[0177] In some embodiments, the nucleic acid library of the DNA library is prepared for sequencing using massively parallel sequencing using any known suitable method to complement the methods described herein.
[0178] Also provided herein, in some embodiments, is a nucleic acid library constructed using any of the methods described herein.
[0179] In some embodiments, the methods provided herein include generating a nucleic acid library from the amplification products. Sequencing and analysis
[0180] In some embodiments, the nucleic acid library or DNA library described herein can be sequenced using any known suitable method to complement the method described herein, and is not limited to any specific sequencing platform.In some embodiments, the sample disclosed herein can be analyzed using any known suitable method to complement the method described herein.The exemplary method of sequencing and method analysis is described below. A. Sequencing
[0181] In some embodiments, the methods provided herein include sequencing a nucleic acid library generated from the amplification products.
[0182] In some embodiments, techniques for sequencing nucleic acid or DNA libraries produced by practicing the methods described herein include the use of polymerase-based sequencing by synthesis, ligation-based sequencing, pyrosequencing, or polymerase-based sequencing methods.
[0183] In some embodiments, the nucleic acid library is sequenced according to the instructions in the MiSeq FGx Sequencing System Reference Guide (e.g., Document No. VD2018006, the entire contents of which are incorporated herein by reference). In some embodiments, the nucleic acid library to be sequenced according to the instructions in the MiSeq FGx Sequencing System Reference Guide (e.g., Document No. VD2018006) is denatured.
[0184] In some embodiments, the sequencing methods disclosed herein include the use of massively parallel sequencing (MPS). In some embodiments, the sequencing methods disclosed herein do not include the use of whole genome sequencing (WGS). In some embodiments, the sequencing methods disclosed herein do not include the use of microarrays.
[0185] In some embodiments, the sequencing methods disclosed herein detect 90% or about 90% of the SNP loci.
[0186] In some embodiments, the sequencing methods disclosed herein generate an output report that includes the results of sequencing an amplification product that includes multiple SNPs.
[0187] In some embodiments, the sequencing comprises a sequencing plexity of up to 40-plex. In some embodiments, the sequencing comprises a sequencing plexity of 2-plex to 40-plex. In some embodiments, the sequencing comprises a sequencing plexity of 12-plex to 40-plex. In some embodiments, the sequencing comprises a sequencing plexity of 12-plex to 32-plex. In some embodiments, the sequencing comprises a sequencing plexity of 12-plex to 30-plex. In some embodiments, the sequencing comprises a sequencing plexity of 24-plex to 40-plex. In some embodiments, the sequencing comprises a sequencing plexity of 24-plex to 32-plex. In some embodiments, the sequencing comprises a sequencing plexity of 28-plex to 32-plex. In some embodiments, the sequencing comprises a sequencing plexity of 2-plex, 3-plex, 4-plex, 5-plex, 6-plex, 7-plex, 8-plex, 9-plex, 10-plex, 11-plex, 12-plex, 13-plex, 14-plex, 15-plex, 16-plex, 17-plex, 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, or 32-plex. In some embodiments, the sequencing comprises a sequencing plexity of 30-plex or about 30-plex. In some embodiments, the sequencing comprises a sequencing complexity of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40-plex.In some embodiments, the sequencing comprises a sequencing plexity of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45-plex. Sequencing plexity refers to the number of individual samples that are sequenced together, for example, in a flow cell.
[0188] In some embodiments, sequencing comprises sequencing the postmortem sample at a sequencing plexity of between or between about 6-plex and 16-plex. In some embodiments, sequencing comprises sequencing the postmortem sample at a sequencing plexity of between or between about 8-plex and 14-plex. In some embodiments, sequencing comprises sequencing the postmortem sample at a sequencing plexity of between or between about 10-plex and 14-plex. In some embodiments, sequencing comprises sequencing the postmortem sample at a sequencing plexity of 10-plex, 11-plex, 12-plex, 13-plex, or 14-plex. In some embodiments, sequencing comprises sequencing the postmortem sample at a sequencing plexity of 10-plex, 11-plex, 12-plex, 13-plex, or 14-plex. In some embodiments, sequencing comprises sequencing the postmortem sample at a sequencing plexity of 12-plex or about 12-plex.
[0189] In some embodiments, sequencing comprises sequencing the ante-mortem sample at a sequencing plexity of between or between about 24-plex and 40-plex. In some embodiments, sequencing comprises sequencing the post-mortem sample at a sequencing plexity of between or between about 26-plex and 38-plex. In some embodiments, sequencing comprises sequencing the post-mortem sample at a sequencing plexity of between or between about 28-plex and 36-plex. In some embodiments, sequencing comprises sequencing the post-mortem sample at a sequencing plexity of at or between about 28-plex and 36-plex. In some embodiments, sequencing comprises sequencing the post-mortem sample at a sequencing plexity of 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, or 34-plex. In some embodiments, the sequencing comprises sequencing the post-mortem sample at a sequencing complexity of 32-plex or about 32-plex. B. Analysis
[0190] In some embodiments, the methods provided herein include analyzing the sequence of the amplification product.
[0191] In some aspects, the methods disclosed herein involve the use of an analysis module that automatically initiates analysis upon completion of sequencing of a sample (i.e., amplification product). In some embodiments, the analysis module comprises universal analysis software (UAS).
[0192] In some embodiments, the analytical methods disclosed herein generate an output report that includes the results of sequencing an amplification product that includes multiple SNPs.
[0193] In some embodiments, the sequencing results are analyzed using any suitable sequence analysis software available in the art.
[0194] In some embodiments, the sequencing results are analyzed using Forenseq Universal Analysis Software, such as version 2.2 or later (Verogen, San Diego, CA), according to the instructions outlined in the Forenseq Universal Analysis Software Reference Guide, such as version 2.1 or 2.2 or later, provided in, e.g., Document No. VD2019002, the entire contents of which are incorporated herein by reference. Genotype and DNA profile determination
[0195] In some embodiments, the methods provided herein comprise genotyping a plurality of SNPs, thereby generating a DNA profile.
[0196] In some embodiments, the DNA profile is generated by genotyping multiple SNPs.
[0197] In some embodiments, an output report containing the results of sequencing an amplification product containing multiple SNPs generated by any of the methods described herein can be used to genotype a sample using any known suitable method to complement the methods described herein. In some embodiments, an output report containing the results of sequencing an amplification product containing multiple SNPs generated by any of the methods described herein can be used to generate a DNA profile using any known suitable method to complement the methods described herein.
[0198] In some embodiments, the DNA profile comprises genotypes for each of a plurality of SNPs. In some embodiments, the DNA profile comprises the genotypes of at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the SNPs, or at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the SNPs. In some embodiments, the DNA profile comprises the genotypes of at least 85% or at least about 85% of the SNPs. In some embodiments, the DNA profile comprises the genotypes of at least 90% or at least about 90% of the SNPs. In some embodiments, the DNA profile comprises the genotypes of at least 95% or at least about 95% of the SNPs. In some embodiments, the DNA profile comprises the genotypes of at least 99%, or at least about 99%, or about 100% of the SNPs.
[0199] In some embodiments, the methods disclosed herein include determining hair color, eye color, and biogeographic ancestry. Determining relevance
[0200] In some embodiments, the methods provided herein include calculating the degree of relatedness of the DNA profile to one or more reference DNA profiles, wherein the one or more reference DNA profiles are included in a reference set of DNA profiles that includes one or more reference DNA profiles from relatives of the person of interest.
[0201] In some embodiments, the relevance of the DNA profiles described herein can be calculated with reference to one or more reference DNA profiles using any known suitable method to complement the methods described herein.
[0202] In some embodiments, the one or more reference DNA profiles are included within a reference set of DNA profiles that includes one or more reference DNA profiles from relatives of the person of interest.
[0203] In some embodiments, a reference set of DNA profiles includes up to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 reference DNA profiles. In some embodiments, a reference set of DNA profiles includes up to 1000 reference DNA profiles. In some embodiments, a reference set of DNA profiles includes up to 500 reference DNA profiles. In some embodiments, a reference set of DNA profiles includes up to 250 reference DNA profiles. In some embodiments, a reference set of DNA profiles includes up to 150 reference DNA profiles. In some embodiments, a reference set of DNA profiles includes up to 100 reference DNA profiles. In some embodiments, a reference set of DNA profiles includes up to 75 reference DNA profiles. In some embodiments, the reference set of DNA profiles includes up to 50 reference DNA profiles. In some embodiments, the reference set of DNA profiles includes up to 25 reference DNA profiles. In some embodiments, the reference set of DNA profiles includes up to 15 reference DNA profiles. In some embodiments, the reference set of DNA profiles includes between 1 and 1,000 reference DNA profiles, between 1 and 500 reference DNA profiles, between 1 and 400 reference DNA profiles, between 1 and 300 reference DNA profiles, between 1 and 250 reference DNA profiles, between 1 and 200 reference DNA profiles, between 1 and 150 reference DNA profiles, between 1 and 100 reference DNA profiles, between 1 and 75 reference DNA profiles, between 1 and 50 reference DNA profiles, between 1 and 25 reference DNA profiles, between 1 and 20 reference DNA profiles, between 1 and 15 reference DNA profiles, between 1 and 10 reference DNA profiles, or between 1 and 5 reference DNA profiles.
[0204] In some embodiments, the reference set of DNA profiles includes at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 25 reference DNA profiles, and up to 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 reference DNA profiles.
[0205] In some embodiments, the reference set of DNA profiles comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 12, 14, 15, 16, 17, 18, 19, or 20 reference DNA profiles.
[0206] In some embodiments, the reference set of DNA profiles includes DNA profiles from blood relatives of the person of interest. In some embodiments, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from blood relatives of the person of interest. In some embodiments, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 1109%, 1110%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 145%, 146%, 147%, 148%, 149%, 150%, 151% of the reference DNA profiles in the reference set of DNA profiles. or 100%, or about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% are from relatives of the person of interest. In some embodiments, 100% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest. In some embodiments, at least 50% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest.
[0207] In some embodiments, each of the one or more reference DNA profiles from the blood relatives of the person of interest is a pre-mortem sample.In some embodiments, one or more of the one or more reference DNA profiles from the blood relatives of the person of interest is a pre-mortem sample.In some embodiments, one or more of the one or more reference DNA profiles from the blood relatives of the person of interest is a post-mortem sample.In some embodiments, the one or more reference DNA profiles from the blood relatives of the person of interest include a post-mortem sample and a pre-mortem sample.
[0208] In some embodiments, each relative of the person of interest in the reference set of DNA profiles is a first-, second-, third-, fourth-, or fifth-degree relative of the person of interest. For example, in an embodiment including a reference set of DNA profiles comprising three reference DNA profiles from relatives of the person of interest, each of the three reference DNA profiles can independently be from a first-, second-, third-, fourth-, or fifth-degree relative, e.g., the first reference DNA profile can be from a first-degree relative, the second reference DNA profile can be from a third-degree relative, and the third reference DNA profile can be from a first-degree relative.
[0209] In some embodiments, at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the one or more reference DNA profiles in the reference set of DNA profiles are relatives, and each of the one or more reference DNA profiles is, independently, from a first-degree relative, a second-degree relative, a third-degree relative, a fourth-degree relative, or a fifth-degree relative with respect to each of the other one or more reference DNA profiles in the reference set of DNA profiles.
[0210] In some embodiments, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest, and at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles each are first-, second-, third-, fourth-, or fifth-degree relatives.
[0211] In some embodiments, the identity of each relative of the person of interest in the reference set of DNA profiles is known. In some embodiments, the identity of each of the one or more reference DNA profiles in the reference set of DNA profiles is known.
[0212] In some embodiments, the reference set of DNA profiles includes DNA profiles derived from a sample from a person of interest. For example, in some embodiments, the reference set of DNA profiles includes DNA profiles derived from a sample from a person of interest, for example, before the disappearance of the person of interest if the person of interest is missing, or before the person of interest becomes a victim of a disaster or conflict. In some embodiments, for example, if the person of interest is missing, a DNA profile derived from a sample from a person of interest, for example, before the disappearance of the person of interest if the person of interest is missing, or before the person of interest becomes a victim of a disaster or conflict, is used as a positive control for the person of interest. This is because the sample was obtained before the person of interest's death or disappearance or victimization, and is a sample that has been confirmed to be derived from the person of interest. Thus, in some embodiments, the reference set of DNA profiles includes DNA profiles derived from a sample from a person of interest, for example, before the disappearance of the person of interest if the person of interest is missing, or before the person of interest becomes a victim of a disaster or conflict, and is known to be derived from the person of interest before being amplified and / or sequenced.
[0213] In some embodiments, the reference set of DNA profiles is in a database, such as a genetic database. In some embodiments, the database is not publicly accessible, i.e., not accessible by the public. In some embodiments, the database is not a public database, such as a public database accessible by law enforcement agencies or third-party genealogy services. In some embodiments, the database is not publicly accessible through a subscription service. In some embodiments, the database is not accessible by third-party genealogy services.
[0214] In some embodiments, calculating the degree of association between a DNA profile and one or more reference DNA profiles does not involve accessing a publicly accessible database, for example, a publicly accessible genetic database.In some embodiments, calculating the degree of association between a DNA profile and one or more reference DNA profiles does not require Internet access to access a database that contains a reference set of DNA profiles.In some embodiments, calculating the degree of association between a DNA profile and one or more reference DNA profiles involves using a local database that contains a reference set of DNA profiles.As used herein, "local database" refers to a database that is only stored and accessible locally, and cannot be accessed by the public, for example, a third party, who wishes to query the database.
[0215] In some embodiments, the reference set of DNA profiles comprises DNA profiles from two or more unrelated families, for example, two or more unrelated families (i.e., families that are not related to each other), each of which contains one or more relatives of the missing persons and / or victims of disasters or conflicts.For example, if a disaster or conflict results in multiple casualties from multiple unrelated families, one or more family members from each family can contribute a reference DNA profile to the reference set of DNA profiles.This local reference set of DNA profiles can then be used locally to identify victims of disasters or conflicts from multiple unrelated families.
[0216] In some embodiments, the reference set or database of DNA profile, for example, genetic database or local database, comprises one or more DNA profiles from individuals of target ethnicity.In some embodiments, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90% or 95% of the reference DNA profiles in the reference set of DNA profile are from target ethnicity.In some embodiments, at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90% or 95% of the reference DNA profiles in the reference set of DNA profile are from target ethnicity. In some embodiments, at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90% or 95% of the reference DNA profiles in the reference set of DNA profiles are from the ethnicity of interest.In some embodiments, at least 75%, 80%, 85%, 90% or 95% of the reference DNA profiles in the reference set of DNA profiles are from the ethnicity of interest.In some embodiments, at least 95%, 95%, 97%, 98% or 99% of the reference DNA profiles in the reference set of DNA profiles are from the ethnicity of interest.In some embodiments, 100% of the reference DNA profiles in the reference set of DNA profiles are from the ethnicity of interest.
[0217] In some embodiments, the person of interest is of an ethnicity of interest.
[0218] In some embodiments, the ethnicity of interest can be any ethnicity, for example, any ethnicity from any location.
[0219] In some embodiments, the ethnicity of interest is a rare ethnicity. In some embodiments, a rare ethnicity is represented by less than 0.01%, 0.05%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, or 5% of the population of the target country or worldwide, or less than 0.01%, 0.05%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, or 5% of the population of the target country or worldwide. In some embodiments, the ethnicity of interest is any ethnicity in the target country. In some embodiments, the ethnicity of interest is a dominant ethnicity in the target country. In some embodiments, the ethnicity of interest is a minority ethnicity in the target country. In some embodiments, the person of interest is from the target country.
[0220] The target country can be any target country. In some embodiments, the target country is Afghanistan, Albania, Algeria, Andorra, Angola, Antigua and Barbuda, Argentina, Armenia, Australia, Austria, Azerbaijan, Bahamas, Bahrain, Bangladesh, Barbados, Belarus, Belgium, Belize, Benin, Bhutan, Bolivia, Bosnia and Herzegovina, Botswana, Brazil, Brunei, Bulgaria, Burkina Faso, Burundi, Côte d'Ivoire, Cape Verde, Cambodia, Cameroon, Canada, Central African Republic, Chad , Chile, China, Colombia, Comoros, Congo (Republic of the Congo), Costa Rica, Croatia, Cuba, Cyprus, Czech Republic (Czech Republic), Democratic Republic of the Congo, Denmark, Djibouti, Dominica, Dominican Republic, Ecuador, Egypt, El Salvador, Equatorial Guinea, Eritrea, Estonia, Eswatini (formerly known as Swaziland), Ethiopia, Fiji, Finland, France, Gabon, Gambia, Georgia, Germany, Ghana, Greece, Grenada, Guatemala, Guinea, Guinea-Bissau, Guyana, Haiti, Holy See, Honduras Jurassic Park, Hungary, Iceland, India, Indonesia, Iran, Iraq, Ireland, Israel, Italy, Jamaica, Japan, Jordan, Kazakhstan, Kenya, Kiribati, Kuwait, Kyrgyzstan, Laos, Latvia, Lebanon, Lesotho, Liberia, Libya, Liechtenstein, Lithuania, Luxembourg, Madagascar, Malawi, Malaysia, Maldives, Mali, Malta, Marshall Islands, Mauritania, Mauritius, Mexico, Micronesia, Moldova, Monaco, Mongolia, Montenegro, Morocco, Mozambique Russia, Russia, Russia (Russia), ...The target country is selected from the group consisting of Seychelles, Sierra Leone, Singapore, Slovakia, Slovenia, Solomon Islands, Somalia, South Africa, South Korea, South Sudan, Spain, Sri Lanka, Sudan, Suriname, Sweden, Switzerland, Syria, Tajikistan, Tanzania, Thailand, Timor-Leste, Togo, Tonga, Trinidad and Tobago, Tunisia, Turkey, Turkmenistan, Tuvalu, Uganda, Ukraine, United Arab Emirates, United Kingdom, United States, Uruguay, Uzbekistan, Vanuatu, Venezuela, Vietnam, Yemen, Zambia, and Zimbabwe. In some embodiments, the target country is the United States.
[0221] In some embodiments, the DNA-based kinship analysis described herein includes the use of a local database. In some embodiments, the DNA-based kinship analysis described herein allows for report generation with minimal user input. In some embodiments, the DNA-based kinship analysis described herein includes the use of an algorithm to calculate a coefficient of kinship. In some embodiments, the coefficient of kinship determines the relatedness status of a sample or DNA profile with a reference DNA profile on a database. For example, in some embodiments, the coefficient of kinship indicates whether each of one or more identified genetic relatives is likely to be a great-great-grandmother, great-great-grandfather, great-grandfather, great-grandmother, grandmother, grandfather, cousin, child of a cousin, or second cousin based on the relative value of the coefficient of kinship. In some embodiments, the reference DNA profile is part of a genealogy database.
[0222] In some embodiments, the DNA-based kinship analysis described herein includes identifying genetic relatives at or about the first, second, third, fourth, or fifth degree of kinship. In some embodiments, the DNA-based kinship analysis described herein includes identifying genetic relatives at or about the first, second, third, fourth, or fifth degree of kinship. In some embodiments, the DNA-based kinship analysis described herein includes identifying genetic relatives at or above the first, second, third, fourth, or fifth degree of kinship. In some embodiments, the DNA-based kinship analysis described herein includes identifying the degree of relatedness between a person of interest and one or more of one or more reference DNA profiles in a reference set of DNA profiles. For example, in some embodiments, the method includes independently identifying the person of interest as a first-, second-, third-, fourth-, or fifth-degree relative of one or more of the one or more reference DNA profiles. A person's first-degree relatives are that person's parents (e.g., father or mother), full siblings (e.g., sisters or brothers), or children (e.g., sons or daughters). A person's second-degree relatives are people who share approximately 25% of the person's genes, such as the person's grandparents, aunts / aunts, uncles / uncles, nieces, nephews, grandchildren, or half-siblings. A person's third-degree relatives are people who share approximately 12.5% of the person's genes, such as great-grandparents, cousins, and great-grandchildren. Fourth-degree relatives include, for example, a cousin's child, a half-great-uncles, a half-great-aunt, a half-niece / grandchild, a half-nephew / grandchild, and a half-cousin. A fifth-degree relative includes, for example, a second cousin, a half-cousin's child, and a cousin's grandchild.
[0223] In some embodiments, the DNA-based kinship analysis described herein comprises generating a family tree comprising DNA profiles related to one or more DNA profiles. The family tree can be generated using any available means or methodology.
[0224] In some embodiments, the DNA-based kinship analysis described herein involves verifying suspects through common ancestry.
[0225] In some embodiments, calculating the degree of relatedness comprises calculating the degree of relatedness between the DNA profile, i.e., the DNA profile from the person of interest, and one or more reference DNA profiles contained within a reference set of DNA profiles, for example a reference set of DNA profiles comprising one or more reference DNA profiles from blood relatives of the person of interest.
[0226] In some embodiments, calculating the degree of association comprises calculating the degree of association between the DNA profile, i.e., the DNA profile from the person of interest, and one or more reference DNA profiles contained in a reference set of DNA profiles, such as a reference set of DNA profiles comprising one or more reference DNA profiles from the person of interest's relatives, by comparing a set of SNPs that are one or more Y-SNPs or that include one or more Y-SNPs. In some embodiments, the one or more Y-SNPs comprise 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 81, 82, 83, 84, or 85 Y-SNPs, or at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 81, 82, 83, 84, or 85 Y-SNPs. In some embodiments, the one or more Y-SNPs are or include 85 Y-SNPs. By comparing the set of SNPs that include the Y-SNPs between the DNA profile from the biological male and one or more reference DNA profiles from the biological male, a likelihood ratio for the male lineage can be determined.
[0227] The likelihood ratio (LR) and relatedness value can be calculated using any approach or algorithm(s) known in the art. In some embodiments, the likelihood ratio is calculated using the algorithms pedprobr (Brustad et al., Int. J. Legal Med., 2021, 135:117-129, the contents of which are incorporated herein by reference in their entirety) and dvir (Vigeland et al., Scientific Reports, 2021, 11:13661, the contents of which are incorporated herein by reference in their entirety). In some embodiments, the average population frequency from the Genome Aggregation Database (gnoMAD) (Karczewski et al., Nature, 2020, 581:434-443, the contents of which are incorporated herein by reference in their entirety) v3.0 is used for LR calculation. In some embodiments, no mutation model is used, and theta is set to 0 when the SNPs selected for analysis have low linkage disequilibrium (Karczewski et al., supra, the entire contents of which are incorporated herein by reference). In some embodiments, the LR is calculated as follows:
number
number
[0228] In some embodiments, LR is calculated as described in Galvan-Femenia et al., Heredity, 2021, 126:537-547, the entire contents of which are incorporated herein by reference.
[0229] In some embodiments, calculating the degree of relatedness comprises calculating a likelihood ratio of sharing a Y chromosome between a DNA profile, i.e., a DNA profile from the person of interest, and one or more reference DNA profiles contained within a reference set of DNA profiles, e.g., a reference set of DNA profiles comprising one or more reference DNA profiles from relatives of the person of interest. In some embodiments, calculating the likelihood ratio of sharing a Y chromosome comprises comparing a set of SNPs that are or comprise one or more Y-SNPs. In some embodiments, the one or more Y-SNPs comprise 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 81, 82, 83, 84 or 85 Y-SNPs, or at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 81, 82, 83, 84 or 85 Y-SNPs. In some embodiments, the one or more Y-SNPs are or comprise 85 Y-SNPs.
[0230] In some embodiments, calculating the likelihood ratio of sharing a Y chromosome comprises calculating a coefficient of kinship based on one or more Y-SNPs from among the plurality of SNPs. In some embodiments, calculating the likelihood of sharing a Y chromosome comprises calculating a log-likelihood by providing a DNA profile and one or more reference DNA profiles included in a reference set of DNA profiles as inputs, for example, to PC-Relate. In some embodiments, calculating the likelihood of sharing a Y chromosome can identify matching Y chromosomes shared between the DNA profiles, for example, the DNA profile of the person of interest, and one or more of the one or more reference DNA profiles, which can then be used to determine the likelihood ratio for the male lineage of the person of interest.
[0231] In some embodiments, this involves using a coefficient of kinship to calculate a measure of relatedness between two samples, for example, between a DNA profile and one of one or more reference DNA profiles. In some embodiments, the methods provided herein include calculating the coefficient of kinship using a kinship model constructed from, for example, a public genealogy database, for example, using PC-AiR or modified PC-AiR methods, and determining kinship on a local set of target samples, for example, using PC-Relate, rather than on the public database, using an expanded set of publicly accessible samples. This also, in some embodiments, involves calculating a likelihood ratio (LR) for each comparison. The likelihood ratio (LR) is a standard measure of relatedness, for example, in the field of forensic science.
[0232] In some embodiments, PC-Relate (Conomos et al., American Journal of Human Genetics, 2016, 98:127-148, the contents of which are incorporated herein by reference in their entirety) and PC-AiR (Conomos et al., Genetic Epidemiology, 2015, 39:276-293, the contents of which are incorporated herein by reference in their entirety), such as those described in Snedecor et al., Forensic Sci. Int. Genet., 2022, 61:102769, the contents of which are incorporated herein by reference in their entirety, are used to calculate the genome-wide relatedness coefficient, shared cM, and longest segment cM. This method is useful when the relationship between two individuals is unknown, thereby eliminating the need for a genealogy, and is applicable to situations such as missing persons or victims of conflict.
[0233] In some embodiments, the PC-AiR method first takes a set of genotyped individuals and separates them into two non-overlapping subsets: one set contains unrelated individuals representing the ancestry of all individuals (unrelated subset), and the other set contains individuals who have at least one relative in the first subset (related subset).To construct the unrelated subset, the original PC-AiR method is modified to improve computational efficiency in building the model.The unrelated subset is populated with samples that have no relatives or the fewest relatives, while samples with more relatives are excluded from the unrelated subset. This is done by calculating a kinship value for each pair and classifying each individual as related or unrelated based on a strict threshold, as in Conomos et al., Genetic Epidemiology, 2015, 39:276-293 (the contents of which are incorporated herein by reference in their entirety): kinship values greater than 0.01 are considered related, and kinship values less than -0.025 are considered unrelated. Samples with less than 5% missing SNP data are excluded. Next, principal component analysis (PCA) is performed on the unrelated subset, and values along components of variation are then predicted for all individuals in the related subset based on their genetic similarity to individuals in the unrelated subset. The resulting components represent a model that can be used in place of static population frequencies to confirm fit in an unknown set of individuals.
[0234] In some embodiments, the PC-Relate method uses principal components from PC-AiR to separate genetic correlation into two components: one for sharing of identical alleles by descent from a recent common ancestor and one for allele sharing by a more distant common ancestor. The components from PC-AiR are used to estimate allele frequencies based on an individual's ancestral background using linear regression instead of static population frequencies, such as those from gnoMAD. Then, for two individuals I and j, the relatedness coefficient
number
number
[0235]
number
number
number
number
number
[0236] Thus, in some embodiments, calculating the degree of relatedness comprises calculating the coefficient of relatedness using a whole-genome kinship algorithm as follows:
number
number
number
number
number
number
number
[0237] In some embodiments, calculating relatedness involves calculating kinship coefficients using a "windowed kinship" approach. See Snedecor et al., Forensic Sci Int Genet 2022, 61, 102769, doi:10.1016 / j.fsigen.2022.102769, the contents of which are incorporated herein by reference in their entirety. Windowed kinship involves calculating a genome-wide window of kinship to find shared related segments. This is performed by enumerating all possible windows within each chromosome and calculating the kinship coefficients for all windows. These windows are then filtered by a minimum kinship threshold and included in the shared cMs calculation. The filtered segments are then iterated, and stretches of SNPs that share at least one allele and two alleles are separately classified. The total shared cMs are then calculated across all segments. The total shared cMs and the longest segment of cMs are used to confirm the relationship when referring to the windowed kinship algorithm. If the number of shared SNPs between two individuals is between 6,000 and 8,000, the shared cM value must exceed 180 and the longest cM segment must exceed 30 to be considered related. If the number of shared SNPs between two individuals is between 8,000 and 9,000, the shared cM value must exceed 150 and the longest cM segment must exceed 30 to be considered related. If the number of shared SNPs between two individuals is 9,000 or more, the shared cM value must exceed 140 and the longest cM segment must exceed 30 to be considered related. The whole-genome relatedness coefficient can be used to filter by any number of shared SNPs. However, Snedecor et al. (2014) observed higher specificity when filtering by shared cM and longest cM segment (e.g., using windowed relatedness) when the SNP overlap is greater than 6,000, especially for higher degrees of relatedness.
[0238] More simply, the number of SNPs classified between two individuals (SNP overlap) can be used to determine when to use a genome-wide relatedness algorithm (<6000 SNP overlap) and when to use a windowed relatedness algorithm (>6000 SNP overlap). Once an algorithm is selected based on the SNP overlap, a value or set of values is used to filter the data and confirm the relationship, depending on the selected algorithm. As demonstrated in Snedecor et al., above, cutoffs for both genome-wide relatedness and windowed relatedness were selected to ensure high sensitivity, but more importantly, high specificity. Lowering these thresholds may capture more relationships (i.e., increasing sensitivity), but is expected to result in more false positive hits, especially for more distant relationships (e.g., fourth- and fifth-degree relatives).
[0239] In some embodiments, calculating relatedness comprises calculating the kinship coefficient for a DNA profile, for example, a DNA profile from a person of interest, and one of one or more reference DNA profiles. In some embodiments, relatedness, for example, the kinship coefficient, is calculated for each of the DNA profile and one or more reference DNA profiles. In some embodiments, the likelihood ratio is calculated by dividing the probability that a query, for example, a DNA profile from a person of interest, and a target, for example, one of one or more reference DNA profiles, are related by the probability that the query and the target are unrelated based on the observed genotypes in the two samples. The results can then be filtered based on the kinship coefficient and LR to confirm the most likely relationship(s) and eliminate false matches among one or more reference DNA profiles contained in the reference set of DNA profiles.
[0240] In some embodiments, calculating the likelihood ratio comprises comparing a plurality of SNPs between the DNA profile and one or more reference DNA profiles. In some embodiments, calculating the degree of relatedness comprises calculating a coefficient of relatedness based on kinship SNPs from within the plurality of SNPs. In some embodiments, calculating the degree of relatedness comprises calculating a coefficient of relatedness based on kinship SNPs from within the plurality of SNPs and calculating a coefficient of relatedness based on Y-SNPs from within the plurality of SNPs. In some embodiments, calculating the degree of relatedness comprises calculating a coefficient of relatedness based on Y-SNPs from within the plurality of SNPs. In some embodiments, calculating the likelihood ratio comprises comparing a set of SNPs including kinship SNPs from among the plurality of SNPs between the DNA profile and one or more reference DNA profiles.
[0241] In some embodiments, the person of interest is biologically male, and the method further comprises calculating a likelihood ratio of sharing a Y chromosome between the DNA profile and one or more reference DNA profiles. In some embodiments, calculating the likelihood ratio of sharing a Y chromosome comprises comparing a set of SNPs comprising one or more Y-SNPs between the DNA profile and one or more reference DNA profiles. In some embodiments, the one or more Y-SNPs are included within the plurality of SNPs. In some embodiments, the one or more Y-SNPs comprise at least 25, 50, 75, or 100 Y-SNPs. In some embodiments, the one or more Y-SNPs comprise at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 81, 82, 83, 84, or 85 Y-SNPs. In some embodiments, the one or more Y-SNPs comprise 85 Y-SNPs. In some embodiments, calculating the likelihood ratio of sharing a Y chromosome comprises dividing the probability that the DNA profile and a reference DNA profile from among the one or more reference DNA profiles share a Y chromosome by the probability that the DNA profile and the reference DNA profile do not share a Y chromosome, based on the genotypes of the one or more Y-SNPs.
[0242] In some embodiments, calculating the relatedness, e.g., the coefficient of kinship, comprises using a principal component analysis (PCA) method. In some embodiments, the relatedness is calculated using a kinship model. In some embodiments, the relatedness is calculated using a kinship model trained using a PCA method. In some embodiments, the PCA method for training the kinship model is PCA. In some embodiments, the PCA method for training the kinship model comprises PCA. In some embodiments, the PCA method for training the kinship model is a method that can account for relatedness of the samples, e.g., known or cryptic relatedness that may arise from the family structure of the entire sample. In some embodiments, the PCA method is PC-AiR, which can enable ancestry determination in the presence of known or cryptic relatedness. See, for example, Conomos et al., Robust Inference of Population Structure for Ancestry Prediction and Correction of Stratification in the Presence of Relatedness, Genet Epidemiol., 2015, 39(4):276-293, the entire contents of which are incorporated herein by reference. In some embodiments, the PCA method is a modified PC-AiR method as described herein.
[0243] In some embodiments, the kinship model is constructed using a training database. In some embodiments, the training database is a genetic database. In some embodiments, the training database is a genealogy database. In some embodiments, the training database is a publicly accessible database. In some embodiments, the training database includes between 1 and 10 million or more training DNA profiles. In some embodiments, the training database is 1, 5, 25, 50, 75, 100, 500, 1,000, 1,500, 2,000, 3,000, 4,000, 5,000, 10,000, 20,000, 30,000, 40,000, 50,000, 75,000, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, 275,000, 300,000, 400, 000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,250,000, 1,500,000, 1,750,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000 or 10,000,000 or about 1, 5, 25, 50, 75, 100, 500, 1,000, 1,500, 2,000, 3,000, 4,000, 5,0 00, 10,000, 20,000, 30,000, 40,000, 50,000, 75,000, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, 275,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,250,000, 1,500,000, 1, 750,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000 or 10,000,000 or at least 1, 5, 25, 50, 75, 100, 500, 1,000, 1,500, 2,000, 3,000, 4,000, 5,000, 10,000, 20,000, 30,000, 40,000, 50,000, 75,000, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, 275,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,250,000, 1,500,000, 1,750,000, 2,000 0,000, 3,000,000, 4,000,000, 5,000,000 or 10,000,000, or at least about 1, 5, 25, 50, 75, 100, 500, 1,000, 1,500, 2,000, 3,000, 4,000, 5,000, 10,000, 20,000, 30,000 0, 40,000, 50,000, 75,000, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, 275,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 9 The invention may comprise 00,000, 1,000,000, 1,250,000, 1,500,000, 1,750,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000 or 10,000,000 training DNA profiles, or a range between any two of the foregoing values. In some embodiments, the training database may contain up to or up to about 100, 500, 1,000, 1,500, 2,000, 3,000, 4,000, 5,000, 10,000, 20,000, 30,000, 40,000, 50,000, 75,000, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, In some embodiments, the training database comprises 275,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,250,000, 1,500,000, 1,750,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000, or 10,000,000 training DNA profiles. In some embodiments, the training database comprises between 5,000 and 500,000, or between 10,000 and 500,000, or between 15,000 and 500,000, or between 20,000 and 500,000, or between 25,000 and 500,000.Contains training DNA profiles between 25,000 and 500,000, or between 25,000 and 400,000, or between 25,000 and 300,000, or between 25,000 and 250,000, or between 50,000 and 500,000, or between 50,000 and 400,000, or between 50,000 and 300,000, or between 50,000 and 250,000.
[0244] In some embodiments, the PCA method is PC-AiR and the training database comprises at least 1 and up to 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,100, 2,200, 2,300, 2,400, 2,500, 2,600, 2,700, 2,800, 2,900, 3,000, 3,500, 4,000, 4,500, or 5,000 training DNA profiles, or a range between any two of the foregoing values.
[0245] In some embodiments, the PCA method is a modified PC-Air method and the training database is 0, 1,750,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000 or 10,000,000, or about 3,000, 4,000, 5,000, 10,000, 20,000, 30,000, 40,000, 50,000, 7 5,000, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, 275,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000 , 1,000,000, 1,250,000, 1,500,000, 1,750,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000 or 10,000,000, or at least 3,000, 4,000, 5,000, 10,000, 20,000, 30,000, 40,000, 50,000, 75,000, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, 275,000, 300,000, 400,000, 500,000 0, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,250,000, 1,500,000, 1,750,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000 or 10,000,000, or at least about 3,000, 4,000, 5,000, 10,000, 20,000, 30,000, 40,000, 50,000, 75,000, 100,000, 125,000, 150,000, 175,000, 200,000, 225,000, 250,000, 275,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,250,000, 1,500,000, 1,750,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000 or 10,000,000 training DNA profiles, or a range between any two of the foregoing values.
[0246] In some embodiments, accessing the training database does not require internet access. In some embodiments, training the kinship model does not require internet access. In some embodiments, the training database is locally accessible.
[0247] In some embodiments, the kinship model is trained by applying a PCA method to a training database. In some embodiments, the training DNA profile includes genotypes for a plurality of SNPs. In some embodiments, the kinship model includes principal components (PCs) obtained for the training database using the PCA method.
[0248] In some aspects, given a training database of training DNA profiles, both PC-AiR and modified PC-AiR methods can identify a sufficiently acceptable, unrelated sample set of training DNA profiles from the training database that is as close to large as possible while still adequately sampling all ancestral backgrounds present in the training database. In some embodiments, both PC-AiR and modified PC-AiR methods can identify a set of unrelated samples, e.g., training DNA profiles, within the training database. In some embodiments, the set of unrelated samples samples all or nearly all ancestral backgrounds present in the training database.
[0249] In some embodiments, both the PC-AiR and modified PC-AiR methods include a first step of estimating the kinship between all pairs of samples in the training database. In some embodiments, the kinship coefficients are estimated. In some embodiments, the kinship coefficients are estimated using a simplified kinship estimation method called "KING-Robust."
[0250] In some embodiments, PC-AiR then proceeds to subsequent steps, including: (1) initializing a set "U" with all samples from the training database; (2) scanning the set and calculating, for each sample, how many samples in U it is related to (called "R") and how many samples in U it is "ancestrally branched" from (called "D"); (3) selecting the sample with the highest R, and if there are multiple samples with the highest R, selecting the sample with the highest R and lowest D; (4) removing the selected sample from U; and (5) repeating from step (2). For example, using PC-AiR, if there are 50,000 samples, the process may begin with 50,000 samples in the first iteration. 2 data points, 49,999 in the second iteration 2 data points and repeat until there are no relevant samples in the set, e.g., 20,000 2 or 10,000 2 data points. In some embodiments, this procedure continues until U contains only irrelevant samples.
[0251] In some embodiments, PC-AiR considers samples related based on estimated relatedness, ie, samples with an estimated relatedness coefficient of ≧0.025 are considered related.
[0252] In some embodiments, PC-AiR considers samples to be ancestrally branched based on estimated kinship, ie, samples with an estimated kinship coefficient <0.025 are considered to be ancestrally branched.
[0253] In some embodiments, PC-AiR includes the steps of: (1) estimating relatedness coefficients between all pairs of samples (e.g., training DNA profiles) in a training database, where pairings with a relatedness coefficient >0.025 are confirmed as closely related and pairings with a relatedness coefficient <-0.025 are confirmed as ancestrally diverged; (2) initializing an unrelated sample set containing all samples; and (3) iteratively: (i) identifying a set in the unrelated sample set that has the most related samples in the unrelated sample set, thereby designating this set as X; (ii) identifying a set of samples in X that has the fewest ancestrally diverged pairings compared to the samples in the unrelated sample set, thereby designating this set as Y; and (iii) terminating the process if Y has 0 samples, or randomly selecting one sample from Y and removing it from U if Y has at least one sample, and repeating starting with step (3)(i).
[0254] In some embodiments, the modified PC-AiR method includes one or more adjustments compared to PC-AiR. In some embodiments, whether a sample is related is more strictly defined in the modified PC-AiR method. In some embodiments, the modified PC-AiR method considers a sample to be related if the estimated coefficient of relatedness is ≧0.01.
[0255] In some embodiments, the modified PC-AiR method considers samples to be ancestrally branched based on an estimated relatedness coefficient, ie, samples with an estimated relatedness coefficient <0.025 are considered to be ancestrally branched.
[0256] In some embodiments, the modified PC-AiR method involves removing all samples with >5% missing genotypes (e.g., more than 5% of the SNPs in the DNA profile) to ensure that each sample is fully informative.
[0257] In some embodiments, the modified PC-AiR method includes the steps of: (1) calculating, for each sample, "R," the total number of related samples in the training database; "D," the number of ancestral branch samples in the database; and "S," the set of related samples; (2) ranking all samples by R (ascending) and D (descending); (3) iterating through the ranked list of samples, and (i) if the sample is not in the "related" set, adding it to the unrelated set and adding all samples from S (i.e., DNA profiles related to the sample) to the related set; or (ii) if the sample is in the "related" set, ignoring the sample and moving on to the next sample. In some aspects, this modified PC-AiR method allows for a process of near-linear complexity (i.e., execution time scales linearly with the number of samples), rather than exponential.
[0258] In some embodiments, the modified PC-AiR method includes: (1) estimating the coefficients of kinship between all pairs of samples (e.g., DNA profiles) in a training database, where pairings with a coefficient of kinship >0.01 are identified as closely related and pairings with a coefficient of kinship <-0.025 are identified as ancestrally diverged; (2) removing all DNA profiles with ≥5% missing data; and (3) ranking all DNA profiles by assigning each DNA profile a ranking value. In some embodiments, the ranking value is determined based on the number of related DNA profiles in the complete database ranked from smallest to largest, broken down by the number of ancestrally diverged DNA profiles in the complete database ranked from largest to smallest. In some embodiments, step (3) includes iterating through the ranked DNA profiles, and for each DNA profile, (i) if the DNA profile is not already in the relevant sample set, adding it to the unrelated sample set and adding all relevant DNA profiles to the relevant sample set, and (ii) if the DNA profile is already in the relevant sample set, skipping to the next DNA profile and repeating starting at step (3)(i).
[0259] In some embodiments, after determining the unrelated sample set using either the PC-AiR method or the modified PC-AiR method, PCA is applied to the unrelated sample set to train a kinship model. In some embodiments, the kinship model further includes PC values calculated for the related sample set. In some embodiments, the PC values of the related sample set are determined based on the PCs obtained for the unrelated sample set.
[0260] In some embodiments, PCA is applied to the entire training database to build the kinship model.
[0261] In some embodiments, a provided method includes training a kinship model.
[0262] In some embodiments, the provided method does not include training a kinship model, hi some embodiments, the kinship model is trained before calculating the relatedness, e.g., relatedness coefficient.
[0263] In some embodiments, accessing the kinship model does not require internet access, hi some embodiments, the kinship model is locally accessible.
[0264] In some embodiments, the degree of relatedness, e.g., the coefficient of relatedness, is calculated using a kinship model. In some embodiments, the degree of relatedness is calculated using the PCs of the kinship model. In some embodiments, calculating the degree of relatedness includes obtaining PC values of a DNA profile, e.g., a DNA profile of the person of interest. In some embodiments, calculating the degree of relatedness includes obtaining PC values of a reference DNA profile(s). In some embodiments, the degree of relatedness is calculated using the PC values of the DNA profile. In some embodiments, the degree of relatedness is calculated using the PC values of the DNA profile and the reference DNA profile(s).
[0265] In some embodiments, the relatedness, e.g., the coefficient of kinship, is calculated using PC-Relate. See, e.g., Conomos et al., Model-free Estimation of Recent Genetic Relatedness, Am. J. Hum. Genet., 98(1):127-148 (2016), the entire contents of which are incorporated herein by reference. In some embodiments, the relatedness is calculated by providing a DNA profile, e.g., a DNA profile of a person of interest, as input to PC-Relate. In some embodiments, the relatedness is calculated by providing a kinship model, e.g., a PC, and a DNA profile as input to PC-Relate. In some embodiments, a reference DNA profile(s) is / are further provided as input to PC-Relate.
[0266] In some embodiments, the relatedness, e.g., the coefficient of kinship, is calculated locally. In some embodiments, calculating the relatedness does not require internet access.
[0267] In some embodiments, the methods described herein further include identifying the person of interest. In some embodiments, identifying the person of interest includes identifying the person of interest by the person's legal name. In some embodiments, identifying the person of interest includes identifying the person of interest by the person's family relationship to one or more known persons in the reference set of DNA profiles. For example, in some embodiments, identifying the person of interest includes confirming that the person of interest is a son or daughter of a particular known person and / or a full sibling of a particular known person. kit
[0268] Provided herein are kits containing any of the primers, reagents, or compositions described herein, which may further include instructions on how to use the kit, such as the uses described herein. The kits described herein may also include other materials desirable from a commercial and user standpoint, including other buffers, diluents, filters, and package inserts containing instructions for performing the methods described herein.
[0269] In some embodiments, provided herein are kits comprising at least one container means containing any of a plurality of primers described herein. Illustrative Embodiments
[0270] Among the exemplary embodiments provided herein are the following: 1. A method for performing DNA-based kinship analysis, comprising: providing a nucleic acid sample from the person of interest; amplifying the nucleic acid sample using a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplification products, wherein the amplification is performed in one or more multiplex PCR reactions; generating a nucleic acid library from the amplification products; sequencing the nucleic acid library generated from the amplification products; analyzing the sequence of the amplification product; genotyping the plurality of SNPs, thereby generating a DNA profile; and calculating a degree of relatedness between said DNA profile and one or more reference DNA profiles, said one or more reference DNA profiles being included within a reference set of DNA profiles that includes one or more reference DNA profiles from blood relatives of said person of interest; A method comprising: 2. A method for performing DNA-based kinship analysis, comprising: providing a nucleic acid sample from the person of interest; amplifying the nucleic acid sample using a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplification products, wherein the amplification is performed in one or more multiplex PCR reactions; generating a nucleic acid library from the amplification products; sequencing the nucleic acid library generated from the amplification products; genotyping the plurality of SNPs, thereby generating a DNA profile; and calculating a degree of relatedness between said DNA profile and one or more reference DNA profiles, said one or more reference DNA profiles being included within a reference set of DNA profiles that includes one or more reference DNA profiles from blood relatives of said person of interest; A method comprising: 3. The method of embodiment 1 or embodiment 2, wherein the sequencing is performed using massively parallel sequencing (MPS). 4. The method of any one of embodiments 1 to 3, wherein the sequencing does not include whole genome sequencing (WGS). 5. The method of any one of embodiments 1 to 4, further comprising generating a family tree comprising the DNA profile associated with one or more DNA profiles. 6. A method for constructing a nucleic acid library for a person of interest, comprising: providing a nucleic acid sample from the person of interest; amplifying the nucleic acid sample using a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating a nucleic acid library comprising amplified products, wherein the amplification is performed in one or more multiplex PCR reactions; A method comprising: 7. The method of embodiment 6, further comprising sequencing the amplification products to produce a DNA profile of the person of interest. 8. A method for constructing a nucleic acid library for a reference DNA sample, comprising: providing nucleic acid samples from relatives of the person of interest; amplifying the nucleic acid sample using a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating a nucleic acid library comprising amplified products, wherein the amplification is performed in one or more multiplex PCR reactions; A method comprising: 9. The method of embodiment 8, wherein the blood relative is a first-, second-, third-, fourth-, or fifth-degree blood relative of the person of interest. 10. The method of embodiment 8 or embodiment 9, wherein the blood relative is a first-, second-, or third-degree blood relative of the person of interest. 11. The method of any one of embodiments 1 to 10, wherein the nucleic acid sample comprises genomic DNA. 12. The method of any one of embodiments 1 to 11, wherein the nucleic acid sample contains one or more enzyme inhibitors. 13. The method of embodiment 12, wherein the one or more enzyme inhibitors comprise one or more inhibitors selected from the group consisting of hematin, heme, humic acid, indigo, tannic acid, collagen, calcium, and hydroxyapatite. 14. The method of any one of embodiments 1 to 13, wherein the nucleic acid sample contains low-quality and / or low-abundance nucleic acid molecules. 15. The method of embodiment 14, wherein the low-quality nucleic acid molecules are degraded and / or fragmented genomic DNA. 16. The low-quality nucleic acid molecule has a degradation index (DI) of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200, or at least ...20, 25, 20, 25, 20, 25, 20, 25, , 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195 or 200. 17. The method of embodiment 14 or embodiment 15, wherein the low-quality nucleic acid molecules have a DI of at least 1 and at most 158.3 or less. 18. The method of any one of embodiments 1 to 13, wherein the nucleic acid sample contains high-quality nucleic acid molecules. 19. The method of embodiment 18, wherein the high-quality nucleic acid molecules have a DI of less than 1. 20. The method of any one of embodiments 1 to 19, wherein the person of interest is a missing person. 21. The method of any one of embodiments 1 to 19, wherein the person of interest is a victim of a disaster or conflict. 22. The method of any one of embodiments 1 to 21, wherein the nucleic acid sample is derived from saliva, blood, semen, hair, teeth, bone, or skin. 23. The method of embodiment 22, wherein the nucleic acid sample is derived from saliva, blood, or semen. 24. The method of embodiment 22, wherein the nucleic acid sample is derived from bone or hair. 25. The method of any one of embodiments 1 to 21, wherein the nucleic acid sample is derived from a buccal swab, paper, fabric, or other substrate or object impregnated with saliva, blood, semen, or other bodily fluid. 26. The method of any one of embodiments 1 to 25, wherein the nucleic acid sample comprises 3 pg to 100 ng or approximately 3 pg to 100 ng of genomic DNA. 27. The method of any one of embodiments 1 to 26, wherein the nucleic acid sample comprises between or about 100 pg and 5 ng of genomic DNA, between or about 50 pg and 5 ng of genomic DNA, or between or about 3 pg and 5 ng of genomic DNA. 28. The method of embodiment 26 or embodiment 27, wherein the nucleic acid sample comprises 1 ng or about 1 ng of genomic DNA. 29. The method of any one of embodiments 1 to 28, wherein the plurality of SNPs comprises kinship-related SNPs (kiSNPs). 30. The method of any one of embodiments 1 to 29, wherein the plurality of SNPs comprises Y chromosome SNPs (Y-SNPs). 31. The method of any one of embodiments 1 to 30, wherein the plurality of SNPs comprises kiSNPs and Y-SNPs. 32. The method of any one of embodiments 1 to 31, wherein the plurality of SNPs comprises kiSNPs, biogeographic ancestry SNPs (aiSNPs), identity SNPs (iiSNPs), phenotypic SNPs (piSNPs), X chromosome SNPs (X-SNPs), and Y chromosome SNPs (Y-SNPs). 33. The method of any one of embodiments 1 to 28, wherein the plurality of SNPs comprises SNPs selected from one or more of the group consisting of kiSNPs, aiSNPs, iiSNPs, piSNPs, X-SNPs, and Y-SNPs. 34. The method of any one of embodiments 1-33, wherein at least or at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the plurality of SNPs are related SNPs. 35. The method of any one of embodiments 1 to 34, wherein the reference set of DNA profiles comprises up to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 reference DNA profiles. 36. The method of any one of embodiments 1 to 35, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from blood relatives of the person of interest. 37. The method of any one of embodiments 1 to 36, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from blood relatives of the person of interest, and wherein each of the at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are first-, second-, third-, fourth-, or fifth-degree blood relatives. 38. A method according to any one of embodiments 1 to 37, wherein at least 50% of the reference DNA profiles in the reference set of DNA profiles are from blood relatives of the person of interest. 39. A method described in any one of embodiments 36 to 38, wherein each blood relative of the person of interest in the reference set of DNA profiles is a first-degree, second-degree, third-degree, fourth-degree, or fifth-degree blood relative of the person of interest, respectively. 40. The method described in embodiment 39, wherein each blood relative of the person of interest in the reference set of DNA profiles is a first-degree, second-degree, or third-degree blood relative of the person of interest, respectively. 41. A method according to any one of embodiments 1 to 40, wherein the identity of each relative of the person of interest in the reference set of DNA profiles is known. 42. A method according to any one of embodiments 1 to 41, wherein the identity of each of the one or more reference DNA profiles in the reference set of DNA profiles is known. 43. A method according to any one of embodiments 1 to 42, wherein the reference set of DNA profiles is in a database. 44. The method of embodiment 43, wherein the database is not publicly accessible. 45. A method according to any one of embodiments 1 to 44, wherein the sequencing comprises a sequencing complexity of up to 40-plex. 46. A method according to any one of embodiments 1 to 44, wherein the sequencing comprises a sequencing complexity of up to 32-plex. 47. A method according to any one of embodiments 1 to 44, wherein the sequencing comprises a sequencing complexity of 12-plex to 32-plex. 48. A method according to any one of embodiments 1 to 44, wherein the sequencing comprises a sequencing complexity of 24-plex to 32-plex. 49. The sequencing is performed at or about 10-plex, 11-plex, 12-plex, 13-plex, 14-plex, 15-plex, 16-plex, 17-plex, 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex, or 35-plex. 45. The method of any one of embodiments 1 to 44, comprising a sequencing complexity of 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex or 35-plex. 50. A method according to any one of embodiments 1 to 49, wherein the sequencing comprises a sequencing complexity of 8 to 16 plex or about 8 to 16 plex for postmortem samples, and / or the sequencing comprises a sequencing complexity of 24 to 40 plex or about 24 to 40 plex for antemortem samples. 51. A method according to any one of embodiments 1 to 50, wherein the sequencing comprises a sequencing complexity of 12-plex or about 12-plex for postmortem samples, and / or the sequencing comprises a sequencing complexity of 32-plex or about 32-plex for antemortem samples. 52. A method according to any one of embodiments 1 to 51, wherein the sequencing comprises a sequencing complexity of 30-plex, 31-plex or 32-plex, or about 30-plex, 31-plex or 32-plex. 53. The method of any one of embodiments 1 to 52, further comprising identifying the person of interest. 54. A method for calculating relatedness, comprising: obtaining a DNA profile comprising genotypes of at least or between about 2,000 and 50,000 SNPs, wherein the DNA profile is from a person of interest; and calculating a degree of relatedness between the DNA profile and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included in a reference set of DNA profiles that includes one or more reference DNA profiles from blood relatives of the person of interest. 55. A method for calculating relatedness, comprising: generating a DNA profile comprising genotypes of at least or between about 2,000 and 50,000 SNPs, wherein the DNA profile is from a person of interest; and calculating a degree of relatedness between the DNA profile and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included in a reference set of DNA profiles that includes one or more reference DNA profiles from blood relatives of the person of interest. 56. A method according to any one of embodiments 1 to 55, wherein the relatedness is calculated using a kinship model. 57. A method according to any one of embodiments 1 to 56, wherein the degree of relatedness is calculated using a kinship model trained using a PCA method. 58. The method described in embodiment 57, wherein the PCA method for training the kinship model is or includes PCA. 59. The method described in embodiment 57 or embodiment 58, wherein the PCA method is PC-AiR. 60. The method of embodiment 59, wherein the PC-AiR comprises: (1) estimating the relatedness coefficients between all pairs of samples in a training database, and optionally in a training DNA profile, wherein pairings with a relatedness coefficient >0.025 are confirmed as closely related and pairings with a relatedness coefficient <-0.025 are confirmed as ancestrally diverged; (2) initializing an unrelated sample set including all samples; and (3) iteratively: (i) identifying a set in the unrelated sample set that has the most related samples in the unrelated sample set, thereby designating it as X; (ii) identifying a set of samples in X that has the fewest ancestrally diverged pairings compared to samples in the unrelated sample set, thereby designating it as Y; and (iii) terminating the process if Y has 0 samples, or randomly selecting one sample from Y and removing it from U if Y has at least one sample, and repeating starting with step (3)(i). 61. The method described in embodiment 57 or embodiment 58, wherein the PCA method is modified PC-Air. 62. The method of embodiment 61, wherein the modified PC-AiR comprises: (1) estimating relatedness coefficients between all pairs of samples, optionally training DNA profiles, in a training database, where pairings with a relatedness coefficient >0.01 are identified as closely related and pairings with a relatedness coefficient <-0.025 are identified as ancestrally diverged; (2) removing all DNA profiles with ≥5% missing data; and (3) ranking all DNA profiles by assigning each DNA profile a ranking value. In some embodiments, the ranking value is determined based on the number of related DNA profiles in the complete database ranked from smallest to largest, broken down by the number of ancestrally diverged DNA profiles in the complete database ranked from largest to smallest. In some embodiments, step (3) includes iterating through the ranked DNA profiles, and for each DNA profile, (i) if the DNA profile is not already in the relevant sample set, adding it to the unrelated sample set and adding all relevant DNA profiles to the relevant sample set, and (ii) if the DNA profile is already in the relevant sample set, skipping to the next DNA profile and repeating starting with step (3)(i). 63. A method according to any one of embodiments 1 to 62, wherein calculating the degree of relatedness comprises calculating a coefficient of relatedness using PC-Relate. 64. The method of embodiment 63, wherein the degree of relevance is calculated by providing the DNA profile of the person of interest as input to PC-Relate. 65. A method according to any one of embodiments 57 to 64, wherein the degree of relatedness is calculated by providing the kinship model and the DNA profile of the person of interest as input to PC-Relate. 66. A method according to any one of embodiments 63 to 65, wherein the one or more reference DNA profiles are further provided as input to PC-Relate. 67. Calculating the degree of relatedness includes calculating the coefficient of relatedness using a whole-genome kinship algorithm as follows:
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0271] The following examples are included for illustrative purposes only and are not intended to limit the scope of the invention. Example 1: Generation of sequence libraries and determination of sensitivity
[0272] This example describes a method for determining the sensitivity of the multiplex polymerase chain reaction described herein to generate a sequenceable library. Figure 1 shows an exemplary schematic of the method for generating a sequenceable library described in this example. A. PCR amplification of genomic DNA targets
[0273] A multiplex polymerase chain reaction was performed to amplify 10,230 individual amplicons in a genomic DNA sample. Each primer pair was designed to selectively hybridize to and promote the amplification of a specific single nucleotide polymorphism (SNP) in the genomic DNA sample. Input genomic DNA was tested at 50 ng to 50 pg, more specifically, 5 ng, 2.5 ng, 1 ng, 500 pg, 250 pg, 100 pg, and 50 pg. Briefly, 18.5 ml of PCR master mix containing sufficient buffer, dNTPs, MgCl2, salts, and PCR additives, such as glycerol, was added to a single well of a 96-well PCR plate. 10,530 primer pairs, 2–4 units of DNA polymerase, such as Phusion hot start DNA polymerase (Thermo Fisher Scientific, catalog number F549L) or any other thermostable DNA polymerase, and 5 microliters of primer pool containing 50 ng to 50 pg of genomic DNA were also added.
[0274] The PCR plate was sealed and placed in a thermal cycler (Veriti 96-well thermal cycler, Thermo Fisher Scientific, 4413964) and run with the tempehrate profile described below to generate the amplicon library. 98°C for 3 minutes 18 cycles: 96°C for 45 seconds 80°C for 10 seconds 54°C for 4 minutes, applicable lamp mode 66°C for 90 seconds, applicable lamp mode 68°C for 10 minutes Keep at 4°C
[0275] After cycling, the amplicon library was kept at 2-8°C until proceeding to the purification steps outlined below. B. Purification of amplicons from input DNA and primers
[0276] Two rounds of purification using MagBind Total Pure NGS beads (Omega Biotek, M1378-02) binding, washing, and elution at 1.6x and 0.6x volume ratios were found to remove genomic DNA and unbound or excess primers. The amplification and purification steps outlined herein produce amplicons approximately 150-350 bp long. The purified amplicons are then used in a second round of PCR to add adapters for sequencing. C. Enrichment of Purified Amplicons to Generate a Sequencing-Ready Library
[0277] A second PCR amplification was performed in a 96-well PCR plate by combining 25 ml of purified amplicons from the above step with 5 ml of adapters provided in the Forenseq Kintelligence kit (Verogen PN:V16000120) and 20 ml of KPCR2 Master Mix provided in the Forenseq Kintelligence kit (Verogen PN:V16000120). The PCR plate was sealed and placed in a thermal cycler (Veriti 96-well thermal cycler, Thermo Fisher Scientific, 4413964) and run using the template profile described below to generate the amplicon library. 98°C for 30 seconds 15 cycles: 98°C for 20 seconds 66°C for 30 seconds 72°C for 30 seconds 72°C for 1 minute Keep at 4°C
[0278] The library was purified using MagBind Total Pure NGS beads (Omega Biotek, M1378-02), binding, washing, and elution at 1×. The purified library is quantified, normalized, denatured, and diluted according to the instructions provided in the Forenseq Kintelligence Kit User Guide (Verogen PN:V16000120, the entire contents of which are incorporated herein by reference).
[0279] The denatured libraries were sequenced according to the instructions in the MiSeq FGx Sequencing System Reference Guide (Document No. VD2018006, the entire contents of which are incorporated herein by reference). As shown in Figure 2, the number of loci detected was similar across the range of input genomic DNA dosage adjustments.
[0280] Results are analyzed using Forenseq Universal Analysis Software 2.1 (Verogen, San Diego, CA) according to the instructions outlined therein and provided in reference guide document number VD2019002, the entire contents of which are incorporated herein by reference. Example 2: Generation of sequence libraries using degraded DNA
[0281] This example describes the sequencing of DNA from small, highly degraded samples. Degraded DNA: A series of degraded blood DNA was obtained from Innogenomics (New Orleans, Louisiana). DNA samples were used to generate sequencing libraries as described in Example 1, except that primer pairs for 10,327 loci were used in this example. Figure 3 shows the percentage of loci detected in degraded DNA using the assay described herein compared to the microarray (GSA) call rate (call rate). The degradation index (DI) is shown on the x-axis, and the number of detected loci is shown on the y-axis. These results demonstrate that even with highly degraded DNA with a DI of 158.3, the assay detected 9,167 loci, sufficient for uploading to a genealogy database for relative search. Alternative technologies, such as microarrays, were unable to detect loci in samples with a high degradation index. Example 3: Evaluation of inhibitor activity on library preparations
[0282] This example describes the evaluation of the effects of PCR inhibitors on the preparation of the libraries disclosed herein. DNA samples from crime scenes often contain co-purifying impurities that inhibit PCR. PCR inhibition is the most common cause of PCR failure when adequate copies of DNA are present. Humic compounds are a group of substances produced during the decay process and have been implicated as contaminants of DNA in soil, natural waters, and recent sediments. Other common inhibitors include hematin (derived from blood), indigo (derived from blue jean), and tannic acid.
[0283] To assess the impact of inhibitors commonly found in forensic samples, library preparation was performed as described in Example 1, except that 200 uM hematin, 50 ng / uL humic acid, 133 uM indigo, and 16 uM tannic acid were spiked into the "amplify and tag target" step above, and primer pairs for 10,380 loci were used. The results are shown in Figure 4, with PCR reactions without any inhibitors labeled as controls. Example 4: Determining relevance
[0284] This example describes exemplary results from samples prepared as generally described in Example 1 above.
[0285] The Illumina Global Screening Array (GSA) 2.0 was run on 200 ng of each of 17 samples from Utah CEPH family 1463 DNA (Coriell Institute). SNP calls were uploaded to the GEDmatch database (Verogen). An exemplary pedigree is shown in Figure 5. One of the samples, NA12889 (paternal grandfather), was run through the library preparation protocol described in Example 1 and run on the ForenSeq UAS 2.1 module. The generated report was uploaded to a database and searched using the 1:multi tool for relationship searching. The relatedness coefficients from the algorithm in the database were compared to the expected relatedness coefficients. The expected and observed relatedness coefficients are shown in Figure 6. Example 5: Determining the Coefficient of Kinship in an Illustrative Case Study
[0286] This example describes the results of an exemplary case study using sample SNP profiles to determine relatedness coefficients. Ten established pedigrees with 12-28 family members in the GEDmatch database were used to test the ability of the 1:many search algorithm to detect potential relatives. The sample SNP profile from the assay disclosed herein was considered to be Mr. / Ms. X = POI (person of interest / unknown crime scene profile). Candidate hits, relatedness coefficients, and relative status are shown in Figure 7.
[0287] The results generated from the search algorithm were then used to generate a family tree for Mr. X, as shown in Figure 8. As shown in the family tree, Mr. X's third-degree relatives, a cousin (1C) and great-grandfather (G GF), were returned within the first 11 candidate hits. Mr. X's great-grandmother (GG GM), great-great-uncles (Uncle GG / Uncle GG), and cousin's child (1C1R), were fourth-degree relatives and were returned within the first 15 candidate hits. Mr. X's fifth-degree relative, a second cousin (2C), was the 12th hit. Example 6: Generation of sequence libraries and determination of sensitivity, including evaluation by locus type
[0288] This example involves a method for determining the sensitivity of the multiplex polymerase chain reaction described herein to generate a sequenceable library, including evaluation by locus type.
[0289] Sequence libraries (sequenced nucleic acid libraries), also called DNA profiles, were generated in the same manner as described in Example 1, except that the results were analyzed using Forenseq Universal Analysis Software version 2.2.
[0290] The results are shown in Figure 9, which is a table summarizing the number of loci detected (as the average of three replicates) based on the amount of input DNA (ng) for each of the different types of loci, e.g., Y-chromosome SNPs (Y-SNPs), X-chromosome SNPs (X-SNPs), phenotypic SNPs (piSNPs), kinship SNPs (kiSNPs), identity SNPs (iiSNPs), and biogeographic SNPs (aiSNPs), out of a total of 10,230 loci analyzed. Genomic DNA input dose adjustments tested included 5 ng, 2.5 ng, 1 ng, 0.5 ng (500 pg), 0.25 ng (250 pg), 0.10 ng (100 pg), and 0.05 ng (50 pg) of input genomic DNA. As shown in Figure 9, input DNA amounts ranging from 0.05 ng to 5 ng of total detected SNPs each detected at least 98.9% (10,117) of the loci, and input DNA amounts of 0.10 ng and greater detected at least 99.5% (10,179) of the loci.
[0291] The data demonstrate that over 10,000 loci can be detected with high efficiency and sensitivity using different types of SNPs and amounts of input DNA ranging from 0.05 ng (50 pg) to 5 ng. Example 7: Evaluation of inhibitor activity on sequence library preparations, including evaluation by locus type
[0292] This example describes the evaluation of the effects of specific inhibitors on the preparation of sequence libraries (sequenced nucleic acid libraries), also known as DNA profiles, as disclosed herein, including by the type of locus being detected and sequenced. DNA samples from crime scenes often contain co-purifying impurities that inhibit amplification. Common inhibitors include hematin, humic acid, and indigo.
[0293] To assess the impact of inhibitors commonly found in forensic samples, results were analyzed using Forenseq Universal Analysis Software version 2.2. Library preparation was performed as described in Example 1, and evaluation of the impact of specific inhibitors on amplification was performed as described in Example 3, except the inhibitors tested were as follows: 200 μM hematin, 100 μM hematin, 50 ng / μL humic acid, 25 ng / μL humic acid, 16 μM tannic acid, 8 μM tannic acid, 133 μM indigo, and 66.5 μM indigo, included in the amplification step as described in Example 1, and primer pairs for the 10230 locus were used. A positive control reaction without inhibitors was also performed. 1 ng of input DNA was used.
[0294] The results are shown in Figure 10, demonstrating that various SNPs, including kiSNPs, Y-SNPs, X-SNPs, piSNPs, iiSNPs, and aiSNPs, can be amplified and detected in combination with one another according to the methods described herein with high efficiency and detection rates, as demonstrated, for example, by all or nearly all of each type of SNP being detected even in the presence of inhibitors. For example, the number of kiSNPs, Y-SNPs, X-SNPs, piSNPs, iiSNPs, and aiSNPs detected is similar to the number detected in the positive control lacking inhibitors (Figure 10). This data demonstrates that the presence of common inhibitors in a sample does not adversely affect the ability to amplify more than 10,000 SNPs in a PCR reaction using the methods described herein. Example 8: Evaluation of sequencing library preparation using DNA samples obtained after simulated sexual assault
[0295] This example describes the generation of a sequence library (sequenced nucleic acid library), also referred to as a DNA profile, using DNA from a simulated sexual assault sample to determine whether a sequence library, e.g., a sequenced nucleic acid library, can be successfully generated using small amounts of input DNA, e.g., 500 pg, with the recommended amount being less than 1 ng.
[0296] Simulated sexual assault DNA was obtained from samples collected 9 and 22 hours after the simulated sexual assault. DNA was isolated from the sperm fraction using differential extraction, and sperm fractions from both time points were collected and stored for analysis. The amount of DNA from the sperm fraction that was available as input for the assay (to generate the sequencing library) was only 500 pg, half the recommended amount of 1 ng.
[0297] The DNA samples were used to generate sequence libraries (sequenced nucleic acid libraries) as described in Example 1, except that the results were analyzed using Forenseq Universal Analysis Software version 2.2. The percentage of loci detected in the assay (call rate) and the number of SNPs of each type present are shown in Figure 11. The results demonstrate that the majority of SNPs were detected with as little as 500 pg of input DNA, with 99.99% of all SNPs (10,229 of 10,230 SNPs) detected at 9 hours and 99.93% of all SNPs (10,223 of 10,230 SNPs) detected at 22 hours. Specifically, all aiSNPs, iiSNPs, piSNPs, X-SNPs, and Y-SNPs were detected at both 9 and 22 hours after the simulated sexual assault. Only 1 kiSNP out of 9,867 was detected at 9 hours, and only 7 kiSNPs out of 9,867 were detected at 22 hours. The number of loci detected is sufficient to allow uploading to a genealogy database for relative searching.
[0298] This data demonstrates that the methods described herein can be used to detect over 10,000 SNPs, including various kiSNPs, Y-SNPs, X-SNPs, piSNPs, iiSNPs, and aiSNPs, and that sequence libraries can be generated using as little as 500 pg of DNA 9 and 22 hours after a simulated sexual assault, with over 99.9% of all SNPs detected. Thus, the methods described herein are suitable for use in generating sequence libraries containing less than the recommended amount of DNA, e.g., 500 pg, following a criminal incident, including sexual assault, or from victims of disasters or conflicts, or from samples left behind by missing persons. Example 9: Evaluation of PCIA carryover during generation of sequence libraries from saliva samples
[0299] This example describes sequencing a nucleic acid library (e.g., to generate a DNA profile) from DNA derived from a saliva sample extracted using organic extraction with the phenol-chloroform-isoamyl alcohol (PCIA) extraction method.
[0300] Saliva DNA was obtained from saliva samples in which the extraction reagent PCIA (e.g., no PCIA, light PCIA, medium PCIA, and heavy PCIA) was intentionally left over in the extracted DNA as carryover, simulating an incomplete extraction. PCIA, which contains phenol, is a known inhibitor of PCR amplification.
[0301] DNA samples with no PCIA, light PCIA, moderate PCIA, and heavy PCIA were used to generate sequence libraries (sequenced nucleic acid libraries) as described in Example 1, except that the results were analyzed using Forenseq Universal Analysis Software version 2.2. The total number of SNPs detected for each sample was determined and is shown in Figure 12. The results indicate that PCIA carryover, even at high levels with heavy PCIA carryover, does not affect the assay's ability to detect SNPs, as over 10,170 SNPs were detected in each sample. Example 10: Generation of sequence libraries from blood samples using various substrates and evaluation of the effect of heme
[0302] This example describes the sequencing of nucleic acid libraries (e.g., to generate DNA profiles) of DNA derived from blood samples deposited on different substrates typically found at crime scenes, including rust and denim, as well as blood samples on swabs where only 420 pg of DNA was available, and blood samples extracted using CheleX™ where increasing levels of heme were carried with the DNA. Heme is a known inhibitor of PCR amplification. Denim contains indigo dye, a known inhibitor of PCR amplification.
[0303] Each of the DNA samples was used to generate sequence libraries (sequenced nucleic acid libraries) as described in Example 1, except that the results were analyzed using Forenseq Universal Analysis Software version 2.2, including the blood and rust-containing sample, two blood samples on denim, a 420 pg blood sample on a swab, and blood samples with low or moderate heme carryover or no heme as controls, as well as a positive control blood sample. The total number of SNPs detected for each sample and reference control was determined and is shown in Figure 13. The results show that even with blood samples deposited on different substrates, detection of more than 10,114 SNPs out of 10,230 total SNPs was still possible. While 9,563 SNPs were detected in a blood sample with only 420 pg, over 10,000 SNPs were detected in the sample with heme, and the number of SNPs detected was not affected by the amount of heme present in the sample. This demonstrates that DNA extracted from blood samples deposited on a variety of substrates commonly found at crime scenes can be used according to the methods provided herein to detect over 10,000 SNPs for forensic applications. Example 11: Kinship analysis using related samples, related antemortem samples, unrelated postmortem samples, and related mock postmortem samples
[0304] This example describes performing a kinship analysis as described herein to ascertain up to three degrees of kinship in four different sample sets, including undegraded samples, highly degraded samples, and low-input samples. Specifically, the goal of this example was to determine up to three degrees of kinship from degraded samples sequenced at high plexity, where potential matches existed in a local private database (rather than a publicly accessible database), while still allowing for enough SNPs to accurately predict such kinship. The methodology involved is outlined in Figure 14 and includes: (a) curating a set of forensically relevant SNP targets and selecting >10,000 SNPs, such as 10,230 SNPs; (b) preparing sequencing libraries from postmortem and antemortem DNA samples by tagging and copying targets, enriching targets, purifying targets, and normalizing target abundance; (c) performing next-generation sequencing at higher plexity, for example, 12-plex or higher; (d) generating SNP reports (also called DNA profiles); (e) uploading the SNP reports to a local server; (f) performing pairwise comparisons; and (g) calculating relatedness coefficients and likelihood ratios and filtering the most likely family relationships. In some embodiments, the curation in step (a) is performed in a previous workflow, and the same selected SNP targets, for example, a specific set of 10,230 SNP targets, are utilized in this workflow.
[0305] A set of 10,230 SNP targets was selected for detection in each of the four sets of samples. The four different sample sets were sequenced to generate the sequence library described in Example 1. These four different sample sets included: (1) A set of related antemortem samples from CEPH / Utah, including up to two-degree relatives validated in Coriell (referred to herein as "related antemortem CEPH / Utah samples"); (2) A set of related antemortem samples from private family members, including up to five-degree relatives (referred to herein as "related antemortem private family samples"); (3) A set of unrelated postmortem samples, including bone (cremated, embalmed, incinerated, and buried), dental remains / teeth, and deteriorated blood at various Deterioration Index (DI) levels (referred to herein as "true postmortem samples"); and (4) A set of related mock postmortem samples, including the same samples from set (2), but either (a) artificially deteriorated by boiling the DNA for 24 hours (DI range of 2.1 to 20) or (b) sequenced with low DNA input (50 pg) (referred to herein as "related mock postmortem samples").
[0306] The true postmortem samples were run in 12-plex using a MiSeq FGx sequencing system, and the results are shown in Figure 15, which shows the number of SNPs detected for each individual sample within this set of true postmortem samples. This includes true postmortem samples labeled as "teeth," "degraded blood," "buried bone," "low-input samples," "other degraded samples," and "other true postmortem samples." As shown in Figure 15, for the true postmortem samples, three of the four buried bone samples had the fewest number of detected (or called) SNPs, ranging from 248 to 1,319 SNPs. The degraded blood samples with the highest DI (DI 158 and DI 56) had the next fewest number of detected SNPs, with 4,603 and 5,069 SNPs, respectively. The remaining true postmortem samples all had a minimum of 6,737 SNPs detected and a maximum of 9,903 SNPs detected (Figure 15). The "total pass" count reflects the total number of SNPs detected for each sample out of the full set of 10,230 SNPs, while the "count pass" count reflects the total number of SNPs detected among the subset of 2,639 SNPs that are consistently called across samples. As shown in Figure 15, the number of SNPs detected among the subset of 2,639 SNPs ("count pass" SNPs) has less variation than the total number of SNPs called overall, so there is a core set of SNPs that are consistently called, i.e., detected, across samples.
[0307] Mock postmortem samples were also run in 12-plex using a MiSeq FGx sequencing system, and the results are shown in Figure 16. As shown in Figure 16, related mock postmortem samples that were artificially degraded by boiling, i.e., samples with a DI greater than 0, had a range of 1,470 to 8,999 SNPs detected, with an average of 6,462 SNPs detected for samples from related parents and daughters. Low-input DNA samples had a DI of 0 and an input of 0.05 ng of DNA (Figure 16).
[0308] Using the MiSeq FGx sequencing system, related antemortem CEPH / Utah samples were run at 12-plex, 16-plex, 24-plex, and 32-plex to determine the highest plexity that would yield a sufficiently high number of detected SNPs (i.e., SNP call rate) for kinship analysis. 24-plex sequencing runs resulted in the detection of 9,691 SNPs on average, ranging from 8,297 detected SNPs to a maximum of 9,982 detected SNPs, depending on the sample. 32-plex sequencing runs resulted in the detection of 9,048 SNPs on average, ranging from 6,894 detected SNPs to a maximum of 9,827 detected SNPs (data not shown). This demonstrated that 30-plex runs allow for a sufficiently high throughput of detected SNPs without significantly compromising the number of detected SNPs and the reliability of kinship analysis.
[0309] A related pre-mortem private family sample was sequenced along with three unrelated samples using a 30-plex sequencing run using a MiSeq FGx Sequencing System. The results are shown in Figure 17. As shown in Figure 17, with the exception of one replicate of the "self" sample (labeled "rep1*"), over 7,000 SNPs were detected in each sample, which could be attributed to library preparation errors for that particular sample. Samples in which over 7,000 SNPs were detected included the "self" sample and several blood relatives of the "self" individual, including samples from a cousin's child, daughter, sister, nephew, cousin, and husband.
[0310] We then performed kinship analysis using the DNA profiles generated after the sequencing run. We compared related pre-mortem private family samples (sequenced at 30-plex) with mock postmortem samples (sequenced at 12-plex) derived from the same original related sample, but where the related mock postmortem samples were artificially degraded or used at low input. Using a minimum relatedness coefficient value of 0.031, all expected relationships up to the third degree of kinship (e.g., cousins) were matched and no false matches were obtained (e.g., no false positives), thereby achieving 100% specificity and 100% sensitivity (data not shown). Some, but not all, expected fourth-degree matches (e.g., cousins' children) were obtained (data not shown). These data confirm that this method can accurately confirm relationships up to the third degree of kinship while ruling out all unrelated relationships, even with highly degraded, low-input samples such as those that may be available in missing persons and disaster / conflict victim situations.
[0311] Next, to assess the most appropriate relatedness threshold, related mock postmortem samples (sequenced at 12-plex) were compared to the GEDMatch database, assuming that related individuals represented by the related mock postmortem samples had no close relatives in the GEDMatch database. It was determined that a relatedness threshold of 0.062 was required to achieve 100% specificity (i.e., no false positives). Applying this relatedness threshold (0.062) to related mock postmortem samples (sequenced at 12-plex) compared to related premortem private family samples (sequenced at 30-plex) revealed reduced sensitivity by excluding some known third-degree relationships. Example 12: Assessing the minimum number of SNPs to accurately determine relatedness
[0312] To identify the minimum number of detected SNPs (i.e., SNP call rate) to accurately determine kinship, we compared known relationships in the GEDMatch database simulating the range of detected SNPs (i.e., called SNPs). This involved testing call rates of 2,000 SNPs, 4,000 SNPs, 6,000 SNPs, 8,000 SNPs, and 10,000 SNPs for sensitivity and specificity in confirming first-, second-, and third-degree kinship. The results are shown in Figures 18A-E, which show receiver operating characteristic (ROC) curves for the results for 2,000 SNPs (Figure 18A), 4,000 SNPs (Figure 18B), 6,000 SNPs (Figure 18C), 8,000 SNPs (Figure 18D), and 10,000 SNPs (Figure 18E). The lowest SNP call rate (n=2,000) resulted in reduced specificity in first-, second-, and third-degree relationships (Figure 18A), suggesting that the SNP call rate of 2,000 detected SNPs represents an absolute floor in SNP call rate for accurately identifying relationships.
[0313] The ability to sequence undegraded samples in a 3-plex fashion opens the possibility of confirming fourth- and fifth-degree kinship relationships. A similar analysis was performed using the GEDMatch database, but to confirm fourth- and fifth-degree kinship relationships. This involved testing the call rates of 2,000 SNPs, 4,000 SNPs, 6,000 SNPs, 8,000 SNPs, and 10,000 SNPs for sensitivity and specificity in confirming fourth- and fifth-degree kinship relationships. The results are shown in Figure 19A-E, which shows the ROC curves for results of 2,000 SNPs (Figure 19A), 4,000 SNPs (Figure 19B), 6,000 SNPs (Figure 19C), 8,000 SNPs (Figure 19D), and 10,000 SNPs (Figure 19E) for blood relatives. As shown in Figures 19A-E, a higher minimum number of called SNPs (approximately 6,000) is required to accurately confirm true fourth- and fifth-degree relationships. Example 13: High-plexity SNP sequencing for kinship analysis using related, simulated antemortem, and simulated postmortem samples
[0314] This example describes performing kinship analysis as described herein to confirm relatedness in different sample sets, including undegraded samples, highly degraded samples, and low-input postmortem (PM) and antemortem (AM) samples. Specifically, the goal of this example was to determine familial relationships up to the third degree from degraded samples sequenced at high plexity, with potential matches present in a local private database (rather than a publicly accessible database), while still allowing for enough SNPs to accurately predict such familial relationships. The methodology involved is outlined in Figure 14 and includes the following steps: (a) curating a set of forensically relevant SNP targets and selecting >10,000 SNPs, such as 10,230 SNPs; (b) preparing sequencing libraries from postmortem and antemortem DNA samples by tagging and copying targets, enriching targets, purifying targets, and normalizing target abundance; (c) performing next-generation sequencing at higher plexity, such as 12-plex or higher; (d) generating SNP reports (also called DNA profiles); (e) uploading the SNP reports to a local server; (f) performing pairwise comparisons; and (g) calculating relatedness coefficients and likelihood ratios and filtering the most likely family relationships. In some embodiments, the curation in step (a) is performed in a previous workflow, and the same selected SNP targets, for example, a specific set of 10,230 SNP targets, are utilized in this workflow. In some embodiments, a windowed relatedness algorithm is used. A. Methodology
[0315] A set of 10,230 SNP targets was selected for detection in each set of samples, including mock antemortem and mock postmortem samples.
[0316] Libraries were generated as described in Example 1 in these studies with 1 ng of NA24385 DNA as a positive control, and results were analyzed automatically in ForenSeq Universal Software version 2.3 (UAS) using expected quality control metrics.
[0317] Commercially available intact DNA used as simulated antemortem (AM) samples was purchased for these studies and included 81 DNA samples from the 1000 Genomes Project, CEPH collection, or Personal Genome Project DNA samples from the Coriell Institute for Medical Research (Camden, NJ, USA), and four DNA samples extracted from whole blood from Innogenomics Inc. (New Orleans, LA, USA). Low-input DNA samples included sample NA24385 at 0.05 ng, 0.1 ng, 0.25 ng, and 0.5 ng.
[0318] The simulated postmortem (PM) sample DNA extracts consisted of five modern tooth (CT) samples, designated CT1, CT2, CT3, CT4, and CT5, seven modern bone samples, and one DNA extract from an ancient bone of Eastern European origin. DNA from the seven modern bone (CB) samples was extracted using either the PrepFiler™ Forensic DNA Extraction Kit (Thermo Fisher, Waltham, MA, USA) for samples CB1, CB3, CB4, CB6, and CB7, or a decalcification protocol for bone samples CB2 and CB5. The degradation index and DNA concentration of the CB bone DNA samples were determined using the Quantifiler™ Trio DNA Quantification Kit (Thermo Fisher, Waltham, MA, USA). The DIs of the CB samples were 13.6, 4.3, 5.6, 1.1, 1.8, 2.5, and 6.5 for CB1, CB2, CB3, CB4, CB5, CB6, and CB7, respectively.
[0319] Additionally, intentionally degraded DNA was used as either mock AM or mock PM samples. Two series of degraded DNA were purchased from Innogenous Inc. (New Orleans, LA, USA): DNA was extracted from whole blood using an organic method from two different male donors and sheared by sonication at 50°C for times ranging from 0 to 16 hours (samples 1231 and 3551). The quantity and degradation state of human DNA was also determined. The 1231 samples had DIs of 26.3, 33.6, 48.6, 160.3, and 459.8, respectively, for the 1231 samples designated 7, 8, 10, 11, and 12. The 3551 samples had DIs of 56 and 158.3, respectively, for the 3551 samples designated 56 and 158.
[0320] Buccal samples were collected from volunteers from families with known pedigrees (RF004, RF016-021), referred to herein as relatives (RF), with the pedigree shown in Figure 26. DNA was extracted from buccal swabs and purified. Two of the DNA samples from relatives (RF004 and RF016) were artificially degraded using high-temperature treatment as follows: five replicates of purified buccal DNA from each individual were subjected to 21 cycles of heating and cooling at 98°C for 1 hour, followed by 4°C for 10 minutes. DNA-grade water was then added to the dried DNA to bring the DNA into solution. The degradation index and DNA concentration were determined for all relative DNA samples. The degradation index varied across replicates, with values of 1, 2.1, 2.6, 5.1, and 20 for sample RF004 and 1, 1.5, 2.0, 2.2, and 2.9 for sample RF016.
[0321] DNA sequencing libraries were prepared using the ForenSeq Kintelligence Kit (Verogen, San Diego, CA, USA) according to the manufacturer's instructions, and libraries were quantified using the QuantiFluor ONE dsDNA system (Promega, Madison, WI, USA). When sequencing libraries at higher plexities, unique dual-index adapters (UDIs) were utilized. Prior to library preparation, intact DNA samples were quantified for input into the library preparation. Mock PM DNA samples were quantified using qPCR. Unless otherwise noted, the DNA was diluted to 40 pg / μL, and 1 ng of total DNA was added to the library preparation reaction. The positive control DNA NA24385 was serially diluted to 20, 10, 4, and 2 pg / μL for total DNA inputs of 500, 250, 100, and 50 pg to mimic low-input PM samples. The artificially degraded samples purchased had sufficient DNA concentration to allow the addition of 1 ng of DNA to the library preparation reaction. Not all degraded kin samples had sufficient DNA concentrations to input 1 ng of DNA into the library preparation reaction. For sample RF004, degraded replicates with DIs of 2.1, 20, 5.1, and 2.6 were added at 600 pg, 600 pg, 700 pg, and 250 pg, respectively, to the library preparation reaction. For sample RF016, degraded replicates with DIs of 2.0, 2.2, and 2.9 had sufficient DNA concentrations to input 1 ng into the library preparation reaction. To mimic low-input samples, 50 pg of RF004 with DIs of 1 and 50 and 250 pg of RF016 with DIs of 1 and 1.5 were added to the library preparation reaction, respectively. Based on mtDNA quantification of approximately 1,400 mtDNA copies / µL, the DNA concentration of the ancient bone was estimated to be 390 pg. Each set of library preps included one positive amplification control of 1 ng of NA24385 DNA and one negative template control (NTC).
[0322] After target amplification and purification, libraries were normalized to 0.75 ng / μL. If the library yield was less than 0.75 ng / μL, the libraries were pooled undiluted. Mock AM libraries generated from commercially obtained intact DNA showed library yields >0.75 ng / μL, with the exception of one at 0.67 ng / μL. Several libraries generated from mock PM, low DNA input, and commercially degraded DNA samples also had yields >0.75 ng / μL. Libraries were pooled at various plexities by pipetting 8 μL of each normalized or intact library into a 1.7 ml microcentrifuge tube. Libraries generated from mock AM samples were pooled at sample plexities of 3, 12, 16, 24, 30, or 32 total libraries for denaturation and sequencing. Libraries generated from mock PM DNA samples were pooled at sample plexities of 3 or 12 total libraries for denaturation and sequencing. The pooled libraries were denatured with freshly diluted NaOH (HP3) by incubation at room temperature for 5 minutes and then diluted with HT1. A human sequencing control (HSC) (a library consisting of 33 STRs serving as a positive sequencing control) was similarly denatured and diluted with HT1. The denatured library pool, combined with the denatured HSC, was then sequenced on a MiSeq FGx instrument using the MiSeq FGx Reagent Kit according to the manufacturer's recommendations. Where possible, sequencing runs were generated with ForenSeq Universal Analysis Software v2.2. Sequencing utilized 151 cycles of paired-end reads for all libraries. The sequencing run included two 8-cycle indexing reads required to demultiplex the libraries using the indexes present in the UDI adapters.
[0323] To generate sequencing results of mock AM libraries with very high reads per sample for simulation studies, 30 libraries from mock AM samples were sequenced on a NextSeq 500 instrument using the NextSeq 500 / 550 High Output Kit v2.5 (300 Cycles) kit (Illumina, San Diego, CA, USA) according to the manufacturer's recommendations (see NextSeq and Denaturation / Pooling Guide).
[0324] The sequencing data were then analyzed using secondary and tertiary data analysis as follows. Metrics are set using UAS for the quality of a MiSeq FGx™ run. These metrics include cluster density, cluster pass filter, phasing, prephasing, and Q-score threshold. Cluster density is the number of clusters per square millimeter of the run (K), and the metric should be between 400 and 1650 K / mm for optimal sequencing results. 2 The filter metric measures base call quality via the percentage of clusters that pass the Illumina chastity filter (ref), and the metric was set to ≥ 80%. Failing this metric impacts the number of usable reads, but not the quality of those passing reads. The phasing metric represents the percentage of DNA strands within a cluster that are behind the current cycle in the read, with a value of ≤ 0.25% passing. Alternatively, prephasing represents molecules within a cluster that are performed prior to the current cycle in the read, with a value of ≤ 0.15% passing. If phasing or prephasing is out of specification, a higher percentage of sequencing errors may be present. It is important to determine whether the HSC passes the metric before using data from a run.
[0325] All library samples prepared using the six UDI adapters provided with the ForenSeq Kintelligence Kit and sequenced on a MiSeq FGx were analyzed for allele and genotype calling using ForenSeq Universal Analysis Software (UAS) v2.3 (Verogen, San Diego, CA) as previously described (Jager et al., Forensic Sci. Int. Genet., 2017, 28:52-70, the contents of which are incorporated herein by reference in their entirety). For library samples prepared with additional UDI adapters for higher complexity, sequencing runs were analyzed on a separate server using the same bioinformatics pipeline utilized in UAS, but via a command-line tool. This pipeline (run within UAS or via a command-line tool) has the same basic algorithm for SNP genotype calling present in UAS v1.3 used for ForenSeq DNA Signature analysis (Jager ref) (Verogen, San Diego, CA, USA). First, samples were demultiplexed based on the supplied index sequences found on the UDI adapter by demultiplexing the binary base call (BCL) file and generating a FASTQ file. Reads 1 and 2 were aligned to the primer sequences using the Smith-Waterman-Gotoh algorithm (Gotoh, O., J. Mol. Biol., 1982, 162:705-708). Reads aligned to a particular primer pair were assigned to the locus corresponding to that pair. The alignments were then written in BAM format. At each SNP position, matches to the reference base call and matches to the alternate base call were counted and filtered with a minimum base quality of 30. The number of reads was then summed for each type of call (reference or alternate) at each locus, requiring a minimum coverage of more than 10 reads.SNP genotypes were then determined by filtering with an analysis threshold (AT) and interpretation threshold (IT) both set at 3%. The AT and IT thresholds were determined by multiplying the sum of the read counts at that locus by 3%. When low coverage occurred, a minimum of 650 reads was used to calculate the threshold. The resulting AT and IT values were then compared to the total read counts for the reference and alternative alleles at each locus. If the call passed both the AT and IT thresholds, the genotype was determined for each locus. The genotyping results were then written in a variant call format (VCF).
[0326] Several studies performed GEDMatch data simulations for whole-genome kinship algorithm testing. For these studies, GEDMatch database profiles were downloaded and analyzed as described in Snedecor et al., Forensic Sci. Int. Genet., 2022, 61, 102769, doi:10.1016 / j.fsigen.2022.102769 (the contents of which are incorporated herein by reference in their entirety). A set of 1,000 de-identified samples, referred to as "query samples," was randomly selected from GEDMatch. These samples were then queried for relatives in the GEDMatch database, and any hits, referred to as "target samples," were selected based on the shared centimorgan (cM) values calculated by the GEDMatch one-to-many tool. This search yielded 2,954 target samples along with matched query samples. These results included query-target pairs with 0 shared cM, representing truly unrelated pairs. Therefore, the number of target samples is greater than the number of query samples, as it includes both related and unrelated target samples. Degrees of relatedness were determined by comparing the resulting shared cM values generated from the one-to-many tool with the expected range of shared cM per degree of kinship provided by DNA Painter, accessible at https: / / dnapainter.com / tools / sharedcmv4. Loci classified for samples included in the query and target sets were first filtered for the 10,230 SNPs in the panel and then randomly filtered to 80%, 60%, 40%, and 20% call rates, resulting in 8,000, 6,000, 4,000, and 2,000 loci called for each query-target pair, respectively. Genome-wide relatedness coefficients were calculated for each query-target sample pair for each level of reduced locus call rate using the relatedness algorithm. Pairs with a genome-wide relatedness coefficient greater than 0.031 were considered related. Pairs with a genome-wide relatedness coefficient less than or equal to 0.031 were considered unrelated.Sensitivity and specificity were calculated by comparing these results with the one-to-many tool query results. In other words, the one-to-many query results were considered the true set, and the results generated by the kinship algorithm were considered the test set.
[0327] The likelihood ratio (LR) and relatedness value were then calculated as follows: LR was calculated using the algorithms pedprobr (Brustad et al., Int. J. Legal Med., 2021, 135:117-129, the contents of which are incorporated herein by reference in their entirety) and dvir (Vigeland et al., Scientific Reports, 2021, 11:13661, the contents of which are incorporated herein by reference in their entirety). The population frequency average from the Genome Aggregation Database (gnoMAD) (Karczewski et al., Nature, 2020, 581:434-443, the contents of which are incorporated herein by reference in their entirety) v3.0 was used for LR calculation. No mutation model was used, and theta was set to 0 because the SNPs selected for analysis have low linkage disequilibrium (Karczewski et al., supra). LR was calculated as follows:
number
number
[0328] PC-Relate (Conomos et al., American Journal of Human Genetics, 2016, 98:127-148, the contents of which are incorporated herein by reference in their entirety) and PC-AiR (Conomos et al., Genetic Epidemiology, 2015, 39:276-293, the contents of which are incorporated herein by reference in their entirety) were used to calculate whole-genome relatedness coefficients, shared cM, and longest segment cM, as previously described in Snedecor et al., Forensic Sci. Int. Genet., 2022, 61:102769, the contents of which are incorporated herein by reference in their entirety. This method is useful when the relationship between two individuals is unknown, thereby eliminating the need for a pedigree.
[0329] The PC-AiR method first takes a set of genotyped individuals and separates them into two non-overlapping subsets: one set contains unrelated individuals representing the ancestry of all individuals (unrelated subset), and the other set contains individuals who have at least one relative in the first subset (related subset). To construct the unrelated subset, modifications were made to the original PC-AiR method to improve computational efficiency in building the model. Samples with no relatives or the fewest relatives are added to the unrelated subset, while samples with more relatives are excluded from the unrelated subset. This was done by calculating a kinship value for each pair and classifying each individual as related or unrelated based on a strict threshold, as described by Conomos et al., Genetic Epidemiology, 2015, 39:276-293 (the contents of which are incorporated herein by reference in their entirety): kinship values greater than 0.01 were considered related, and kinship values less than -0.025 were considered unrelated. Samples with less than 5% missing SNP data were excluded. Next, principal component analysis (PCA) was performed on the unrelated subset, and values along components of variation were then predicted for all individuals in the related subset based on their genetic similarity to individuals in the unrelated subset. The resulting components represented a model that could be used in place of static population frequencies to confirm fit in an unknown set of individuals. The model used in this study was built using the GEDMatch database.
[0330] PC-Relate uses the principal components from PC-AiR to separate genetic correlations into two components: one for sharing of identical alleles by descent from a recent common ancestor, and one for allele sharing by a more distant common ancestor. The components from PC-AiR were used to estimate allele frequencies based on an individual's ancestral background using linear regression instead of static population frequencies such as those from gnoMAD. Then, for two individuals i and j, the relatedness coefficient is
number
number
number
number
number
number
number
[0331] Snedecor et al. (2013) introduced an additional step called "windowed kinship" to more accurately confirm distant relationships. Windowed kinship involves calculating a genome-wide window of kinship to find shared related segments. This is performed by enumerating all possible windows within each chromosome and calculating the kinship coefficients for all windows. These windows are then filtered by a minimum kinship coefficient threshold and included in the shared cMs calculation. The filtered segments are then iterated, and stretches of SNPs that share at least one allele and two alleles are separately classified. The total shared cMs are then calculated across all segments. The total shared cMs and the longest cM segment are used to confirm relationships when referring to the windowed kinship algorithm. If the number of shared SNPs between two individuals is between 6,000 and 8,000, the shared cM value must exceed 180, and the longest cM segment must exceed 30, to be considered related. If the number of shared SNPs between two individuals is between 8,000 and 9,000, the shared cM value must exceed 150 and the longest cM segment must exceed 30 to be considered related. If the number of shared SNPs between two individuals is 9,000 or more, the shared cM value must exceed 140 and the longest cM segment must exceed 30 to be considered related. The whole-genome relatedness coefficient can be used to filter for any number of shared SNPs. However, Snedecor et al. (2014) observed higher specificity, especially for higher degrees of relatedness, when filtering by shared cM and longest cM segment (e.g., using windowed relatedness) when the SNP overlap was greater than 6,000.
[0332] More simply, the number of SNPs classified between two individuals (SNP overlap) can be used to determine when to use a genome-wide relatedness algorithm (<6000 SNP overlap) and when to use a windowed relatedness algorithm (>6000 SNP overlap). Once an algorithm is selected based on the SNP overlap, a value or set of values is used to filter the data and confirm the relationship, depending on the selected algorithm. As demonstrated in Snedecor et al., above, cutoffs for both genome-wide relatedness and windowed relatedness were selected to ensure high sensitivity, but more importantly, high specificity. Lowering these thresholds may capture more relationships (i.e., increasing sensitivity), but is expected to result in more false positive hits, especially for more distant relationships (e.g., fourth- and fifth-degree relatives). Using the above method, studies were conducted to sequence libraries at higher plexity, thereby reducing total sample reads, in order to generate stable but smaller sorted SNP sets for a cost-effective method for kinship analysis. B. Results
[0333] To demonstrate the feasibility of high plexity using sequenced libraries, we simulated high plexity (>3 samples in a sequencing run) with a set of 30 libraries generated from high-quality DNA samples from among the mock AM samples. The 30 mock AM samples were prepared and sequenced together as described above to generate sample data with a large number of reads. The BCL files were demultiplexed using a custom, local build of the ForenSeq UAS secondary analysis pipeline, which differed only in the generation of FASTQ files due to differences in the raw data generated by NextSeq. To determine the number of reads to simulate 16 sequenced libraries, we divided 25,000,000 (the total estimated number of reads) by 16 (the number of samples) and multiplied by the average percent aligned (96%), resulting in a total of 1,500,000 reads per sample at this plexity. The same calculations were then performed to simulate 30 sequenced libraries, but instead of dividing by 16, we divided by 30 to obtain 800,000 reads. Simulating a sequencing run with fewer reads is called "downsampling." Bioinformatics tools such as seqtk (https: / / github.com / lh3 / seqtk) can be used to simulate runs with fewer reads or downsample sequencing runs. Using Seqtk, we downsampled each sample to 1.5 million reads to simulate a plexity of 16 samples / run or to 800,000 reads to simulate a plexity of 30 samples / run. seqtk randomly selects reads from the FASTQ file and outputs a new FASTQ with the desired number of reads. The subsequent downsampled FASTQs were processed through a locally built ForenSeq UAS pipeline, which analyzed the FASTQs as described above.
[0334] The total reads generated per sample ranged from 8,086,090 to 32,707,490, with an average of 23,186,251 reads. To simulate high plexity, we randomly selected reads from the FASTQ file for each sample until the desired number of reads was met, a process known as downsampling. Downsampling is defined as the random selection of reads from a FASTQ file until the desired number of reads is achieved. The randomly selected reads were then exported to a new FASTQ file. The resulting FASTQ file was analyzed with the bioinformatics algorithm described above. Sequencing plexities of 16 and 30 were simulated by downsampling the data to 1.5M and 800,000 reads for each sample, respectively. Reducing the number of reads per sample resulted in an expected decrease in the number of classified SNPs.
[0335] To determine the SNP call rate across samples, the number of classified SNPs was determined for each sample at each simulated sequencing complexity (16 and 30). The 16-sample run (16-plex) data was generated by downsampling the raw reads in the FASTQ file to 1.5 million reads per sample, and the 30-sample run (30-plex) data was generated by downsampling the raw reads to 800,000 reads per sample. Genotyping was then performed, and the number of classified SNPs per sample was summed. For the 16-sample run, the minimum was 7781, the first quartile was 8472, the median was 8630, the third quartile was 8708, and the maximum was 8848. For the 30-plex run, the minimum was 6375, the first quartile was 7179, the median was 7299, the third quartile was 7382, and the maximum was 7516. The distribution of sorted SNPs for two simulated plexities of 30 libraries is shown in Figure 20A. The mean recovery for 16 plexities was 8586 SNPs (ranging from 7781 to 8848 SNPs), with a median of 8630 SNPs, the first quartile of 8472 SNPs, and the third quartile of 8708 SNPs. The mean recovery for 30 plexities was 7234 SNPs (ranging from 6375 to 7516 SNPs), with a median of 7299 SNPs, the first quartile of 7179 SNPs, and the third quartile of 7382 SNPs (Figure 20A).
[0336] Determining relatedness using the above-described kinship algorithm and likelihood ratio relies significantly on the number of classified SNPs of each sample in a one-to-one comparison, also called SNP overlap.
[0337] The greater the number of SNPs classified in both samples, the more certain the kinship and likelihood ratio values. Therefore, we calculated the average SNP overlap between samples in both simulated 16- and 30-plexity sequencing runs to determine whether sufficient common classified SNP loci were shared between samples to confirm true kinship and likelihood ratio values. The average common classified SNP loci across all sample combinations in both simulated 16- and 30-plexity sequencing runs was 6,998 loci, with a minimum overlap of 6,058 and a maximum overlap of 7,322. The simulation demonstrates that sequencing libraries to obtain fewer total reads for each sample allows for classification of a smaller, more stable set of SNPs for each sample with sufficient overlap for kinship determination. The number of common classified SNPs was less than 8,000, which may not be sufficient to confirm higher-order relationships (e.g., fourth- and fifth-degree kinship), but may be sufficient for confirming relationships up to the third degree.
[0338] It was expected that sequencing libraries at increased plexity would result in a lower number of classified SNPs, potentially leading to increased allele dropout and subsequent decreased heterozygosity, especially for postmortem (PM) samples with degraded and / or low DNA content. To assess the level of locus and allele dropout and heterozygosity at higher plexities, libraries generated from mock antemortem (AM) and PM samples, as described above, were sequenced at the recommended plexity of 3 samples per sequencing run, followed by sequencing at four higher plexities of 12, 16, 24, and 32 samples per sequencing run. The resulting number of classified SNPs is shown in Figure 20B. As shown in Figure 20B, the minimum, 1st quartile, median, 3rd quartile, and maximum were 9853, 9976, 10009, 10059, and 10135 for 3-plex; 9332, 9394, 9419, 9520, and 9945 for 12-plex; 8881, 9091, 9303, 9419, and 9901 for 16-plex; and 7653, 8348, 8515, 8706, and 9753 for 32-plex. As shown in Figure 20C, the minimum, first quartile, median, third quartile, and maximum were 7215, 8677, 9724, 9923, and 9991 for 3-plex; and 4603, 8261, 9360, 9664, and 9903 for 12-plex. As shown in Figure 20B, the number of loci with reads below the accepted threshold (AT) increased as the plexity for the reference sample increased. Furthermore, the distribution of loci below the AT broadened as the sequencing plexity increased, with an average of 10,111 (minimum 9,853, maximum 10,135) classified SNPs for samples sequenced with three samples per run and an average of 8,528 (minimum 7,653, maximum 9,753) classified SNPs for samples sequenced with 32 samples per run.The minimum number of typed SNPs remained above 7,000 at the highest plexity of 32 samples per sequencing run tested, indicating that sequencing these libraries at high plexity resulted in more typed SNP loci compared to the simulated results discussed above and shown in Figure 20A.
[0339] Next, the impact of higher-plexity sequencing on sister allele loss for heterozygous loci was assessed by comparing SNP genotypes determined for sample libraries sequenced at the recommended plexity of three samples per run with SNP genotypes determined for the same sample libraries sequenced at 32 samples per run. For each sample and each locus, genotypes were considered concordant if both runs classified the same allele; otherwise, they were discordant. The average overlap between 3-plex and 32-plex sequencing was 8,610 SNPs, with a minimum of 7,808 SNPs and a maximum of 9,667 SNPs. For autosomal loci, both alleles must match to be considered concordant. Allelic concordance for each sample was calculated by dividing the number of matching alleles by the total number of alleles at the locus classified in both sequencing runs. For mock antemortem samples, allelic discordance (alleles dropping below AT) increased by an average of 1.9% between libraries sequenced at a plexity of 3 compared to 32, with a minimum of 0.50% and a maximum of 2.8% (Figure 21A, left y-axis). Heterozygosity was determined by summing the number of heterozygous loci per sample and dividing that value by the total number of loci called. Samples sequenced at 32-plex (32-plex) showed the greatest difference in heterozygosity compared to the standard plexity of 3 samples per run (3-plex), with a mean difference of 6.8%, a minimum difference of 2.0%, and a maximum difference of 10.2%. 3-plex sequencing exhibits greater heterozygosity by these values (Figure 21A, right y-axis).
[0340] A characteristic of samples from MFI victims is that they are often degraded and may contain low levels of genomic DNA. However, while increasing the number of samples per sequencing run is also advantageous in a cost-effective and overall context, this ma...
Claims
1. 1. A method for performing DNA-based kinship analysis, comprising: providing a nucleic acid sample from the person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of between at least or about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplification products, wherein the amplification is performed in one or more multiplex PCR reactions; generating a nucleic acid library from the amplification products; sequencing the nucleic acid library generated from the amplification products; analyzing the sequence of the amplification product; genotyping the plurality of SNPs, thereby generating a DNA profile; and calculating a degree of relatedness between said DNA profile and one or more reference DNA profiles, said one or more reference DNA profiles being included in a reference set of DNA profiles comprising one or more reference DNA profiles from blood relatives of said person of interest; A method comprising:
2. 1. A method for performing DNA-based kinship analysis, comprising: providing a nucleic acid sample from the person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of between at least or about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplification products, wherein the amplification is performed in one or more multiplex PCR reactions; generating a nucleic acid library from the amplification products; sequencing the nucleic acid library generated from the amplification products; genotyping the plurality of SNPs, thereby generating a DNA profile; and calculating a degree of relatedness between said DNA profile and one or more reference DNA profiles, said one or more reference DNA profiles being included in a reference set of DNA profiles comprising one or more reference DNA profiles from blood relatives of said person of interest; A method comprising:
3. 3. The method of claim 1 or claim 2, wherein the sequencing is performed using massively parallel sequencing (MPS).
4. The method of any one of claims 1 to 3, wherein said sequencing does not include whole genome sequencing (WGS).
5. The method of any one of claims 1 to 4, further comprising generating a family tree comprising said DNA profile associated with one or more DNA profiles.
6. 1. A method for constructing a nucleic acid library for a person of interest, comprising: providing a nucleic acid sample from the person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of between at least or about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating a nucleic acid library comprising amplified products, wherein the amplification is performed in one or more multiplex PCR reactions; A method comprising:
7. 7. The method of claim 6, further comprising sequencing the amplification products to produce a DNA profile of the person of interest.
8. 1. A method for constructing a nucleic acid library for a reference DNA sample, comprising: providing nucleic acid samples from relatives of the person of interest; amplifying the nucleic acid sample with a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of between at least or about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating a nucleic acid library comprising amplified products, wherein the amplification is performed in one or more multiplex PCR reactions; A method comprising:
9. 9. The method of claim 8, wherein the relative is a first-, second-, third-, fourth-, or fifth-degree relative of the person of interest.
10. 10. The method of claim 8 or claim 9, wherein the relative is a first, second or third degree relative of the person of interest.
11. The method of any one of claims 1 to 10, wherein the nucleic acid sample comprises genomic DNA.
12. The method of any one of claims 1 to 11, wherein the nucleic acid sample contains one or more enzyme inhibitors.
13. 13. The method of claim 12, wherein the one or more enzyme inhibitors comprise one or more inhibitors selected from the group consisting of hematin, heme, humic acid, indigo, tannic acid, collagen, calcium, and hydroxyapatite.
14. The method of any one of claims 1 to 13, wherein the nucleic acid sample comprises low quality and / or low abundance nucleic acid molecules.
15. 15. The method of claim 14, wherein the low-quality nucleic acid molecules are degraded and / or fragmented genomic DNA.
16. the low quality nucleic acid molecules have a Degradation Index (DI) of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195 or 200, or a Degradation Index (DI) of at least 1, 16. The method of claim 14 or claim 15, wherein the coating has a Deterioration Index (DI) of 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195 or 200.
17. 16. The method of claim 14 or claim 15, wherein the low quality nucleic acid molecules have a DI of at least 1 and at most 158.3 or less.
18. The method of any one of claims 1 to 13, wherein the nucleic acid sample comprises high quality nucleic acid molecules.
19. 20. The method of claim 18, wherein the high quality nucleic acid molecules have a DI of less than 1.
20. The method of any preceding claim, wherein the person of interest is a missing person.
21. The method of any one of claims 1 to 19, wherein the person of interest is a victim of a disaster or conflict.
22. The method of any one of claims 1 to 21, wherein the nucleic acid sample is derived from saliva, blood, semen, hair, teeth, bone or skin.
23. 23. The method of claim 22, wherein the nucleic acid sample is derived from saliva, blood, or semen.
24. 23. The method of claim 22, wherein the nucleic acid sample is derived from bone or hair.
25. 22. The method of any one of claims 1 to 21, wherein the nucleic acid sample is derived from a buccal swab, paper, fabric or other substrate or object impregnated with saliva, blood, semen or other bodily fluid, or containing hair or skin cells.
26. 26. The method of any one of claims 1 to 25, wherein the nucleic acid sample comprises between 3 pg and 100 ng of genomic DNA or between about 3 pg and 100 ng.
27. 27. The method of any one of claims 1 to 26, wherein the nucleic acid sample comprises between or about 100 pg and 5 ng of genomic DNA, between or about 50 pg and 5 ng of genomic DNA, or between or about 3 pg and 5 ng of genomic DNA.
28. 28. The method of claim 26 or claim 27, wherein the nucleic acid sample comprises 1 ng or about 1 ng of genomic DNA.
29. The method of any one of claims 1 to 28, wherein the plurality of SNPs comprises kinship SNPs (kiSNPs).
30. 30. The method of any one of claims 1 to 29, wherein the plurality of SNPs comprises Y chromosome SNPs (Y-SNPs).
31. The method of any one of claims 1 to 30, wherein the plurality of SNPs comprises a kiSNP and a Y-SNP.
32. 32. The method of any one of claims 1 to 31, wherein the plurality of SNPs comprises kiSNPs, biogeographic ancestry SNPs (aiSNPs), identity SNPs (iiSNPs), phenotypic SNPs (piSNPs), X chromosome SNPs (X-SNPs), and Y chromosome SNPs (Y-SNPs).
33. 29. The method of any one of claims 1 to 28, wherein the plurality of SNPs comprises SNPs selected from one or more of the group consisting of kiSNPs, aiSNPs, iiSNPs, piSNPs, X-SNPs and Y-SNPs.
34. 34. The method of any one of claims 1 to 33, wherein at least or at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% of the plurality of SNPs are related SNPs.
35. 35. The method of any one of claims 1 to 34, wherein the reference set of DNA profiles comprises up to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 reference DNA profiles.
36. 36. The method of any one of claims 1 to 35, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest.
37. 37. The method of any one of claims 1-36, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest, and wherein each of said at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are first-, second-, third-, fourth-, or fifth-degree relatives.
38. 38. The method of any one of claims 1 to 37, wherein at least 50% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest.
39. 39. The method of any one of claims 36 to 38, wherein each relative of the person of interest in the reference set of DNA profiles is a first-, second-, third-, fourth-, or fifth-degree relative of the person of interest, respectively.
40. 40. The method of claim 39, wherein each relative of the person of interest in the reference set of DNA profiles is a first-degree, second-degree, or third-degree relative of the person of interest, respectively.
41. 41. The method of any one of claims 1 to 40, wherein the identity of each relative of the person of interest in the reference set of DNA profiles is known.
42. 42. The method of any one of claims 1 to 41, wherein the identity of each of the one or more reference DNA profiles in the reference set of DNA profiles is known.
43. The method of any one of claims 1 to 42, wherein the reference set of DNA profiles is in a database.
44. 44. The method of claim 43, wherein the database is not publicly accessible.
45. 45. The method of any one of claims 1 to 44, wherein said sequencing comprises a sequencing complexity of up to 40-plex.
46. 45. The method of any one of claims 1 to 44, wherein said sequencing comprises a sequencing complexity of up to 32-plex.
47. 45. The method of any one of claims 1 to 44, wherein said sequencing comprises a sequencing complexity of 12-plex to 32-plex.
48. 45. The method of any one of claims 1 to 44, wherein said sequencing comprises a sequencing complexity of 24-plex to 32-plex.
49. The sequencing may be at or about 10-plex, 11-plex, 12-plex, 13-plex, 14-plex, 15-plex, 16-plex, 17-plex 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex, or 35-plex.
45. The method of any one of claims 1 to 44, comprising a sequencing complexity of 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex or 35-plex.
50. 50. The method of any one of claims 1 to 49, wherein the sequencing comprises a sequencing complexity of at or about 8 to 16 plex for post-mortem samples, and / or wherein the sequencing comprises a sequencing complexity of at or about 24 to 40 plex for ante-mortem samples.
51. 51. The method of any one of claims 1 to 50, wherein the sequencing comprises a sequencing complexity of at or about 12-plex for post-mortem samples, and / or wherein the sequencing comprises a sequencing complexity of at or about 32-plex for ante-mortem samples.
52. 52. The method of any one of claims 1 to 51, wherein the sequencing comprises a sequencing complexity of, or about, 30-plex, 31-plex, or 32-plex.
53. The method of any one of claims 1 to 52, further comprising identifying the person of interest.
54. 1. A method for calculating relatedness, comprising: obtaining a DNA profile comprising genotypes of at least between or about 2,000 and 50,000 SNPs, wherein the DNA profile is from a person of interest; and calculating a degree of relatedness between the DNA profile and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included in a reference set of DNA profiles that includes one or more reference DNA profiles from blood relatives of the person of interest.
55. 1. A method for calculating relatedness, comprising: generating a DNA profile comprising genotypes of at least or between about 2,000 and 50,000 SNPs, wherein the DNA profile is from a person of interest; and calculating a degree of relatedness between the DNA profile and one or more reference DNA profiles, wherein the one or more reference DNA profiles are included in a reference set of DNA profiles that includes one or more reference DNA profiles from blood relatives of the person of interest.
56. 56. The method of any one of claims 1 to 55, wherein the relatedness is calculated using a kinship model.
57. A method according to any preceding claim, wherein the degree of relatedness is calculated using a kinship model trained using a PCA method.
58. 58. The method of claim 57, wherein the PCA method for training the kinship model is or comprises PCA.
59. 59. The method of claim 57 or claim 58, wherein the PCA method is PC-AiR.
60. 60. The method of claim 59, wherein the PC-AiR comprises: (1) estimating relatedness coefficients between all pairs of samples in a training database, and optionally in a training DNA profile, wherein pairings with a relatedness coefficient >0.025 are confirmed as closely related and pairings with a relatedness coefficient <-0.025 are confirmed as ancestrally diverged; (2) initializing an unrelated sample set containing all samples; and (3) iteratively: (i) identifying a set in the unrelated sample set that has the most related samples in the unrelated sample set, thereby designating it as X; (ii) identifying a set of samples in X that has the fewest ancestrally diverged pairings compared to samples in the unrelated sample set, thereby designating it as Y; and (iii) terminating the process if Y has 0 samples, or randomly selecting one sample from Y and removing it from U if Y has at least one sample, and repeating beginning with step (3)(i).
61. 59. The method of claim 57 or claim 58, wherein the PCA method is modified PC-Air.
62. 62. The method of claim 61, wherein the modified PC-AiR comprises: (1) estimating relatedness coefficients between all pairs of samples of training DNA profiles, as appropriate, in a training database, wherein pairings with a relatedness coefficient >0.01 are identified as closely related and pairings with a relatedness coefficient <-0.025 are identified as ancestrally diverged; (2) removing all DNA profiles with ≥ 5% missing data; and (3) ranking all DNA profiles by assigning each DNA profile a ranking value. In some embodiments, the ranking value is determined based on the number of related DNA profiles in the complete database ranked from smallest to largest, broken down by the number of ancestrally branched DNA profiles in the complete database ranked from largest to smallest. In some embodiments, step (3) includes iterating through the ranked DNA profiles, and for each DNA profile, (i) if the DNA profile is not already in the related sample set, adding it to the unrelated sample set and adding all related DNA profiles to the related sample set, and (ii) if the DNA profile is already in the related sample set, skipping to the next DNA profile and repeating starting with step (3)(i).
63. 63. The method of any one of claims 1 to 62, wherein said calculating said degree of relatedness comprises calculating a coefficient of relatedness using PC-Relate.
64. 64. The method of claim 63, wherein the degree of relatedness is calculated by providing the DNA profile of the person of interest as input to PC-Relate.
65. 65. The method of any one of claims 57 to 64, wherein the degree of relatedness is calculated by providing the kinship model and the DNA profile of the person of interest as input to PC-Relate.
66. 66. The method of any one of claims 63 to 65, wherein the one or more reference DNA profiles are further provided as input to PC-Relate.
67. Calculating the degree of relatedness includes calculating a coefficient of relatedness using a whole-genome kinship algorithm as follows: [Number 62] the reference DNA profiles of the person of interest and the one or more reference DNA profiles are i and j; [Number 63] is the coefficient of kinship, [Number 64] is the estimated allele frequency, s is the SNP in S SNPs classified in both individuals, [Number 65] and [Number 66] are the numbers of reference alleles at i and j at SNP s, respectively; [Number 67] and [Number 68] 67. The method of any one of claims 1 to 66, wherein i and j are the expected allele frequencies calculated by PC-AiR for i and j at SNP s, respectively.
68. The method of any one of claims 1 to 67, wherein said calculating said relevance comprises calculating a likelihood ratio.
69. 69. The method of Claim 68, wherein said calculating said likelihood ratio comprises comparing said plurality of SNPs between said DNA profile and said one or more reference DNA profiles.
70. 69. The method of Claim 68, wherein said calculating said likelihood ratio comprises comparing a set of SNPs comprising kinship SNPs from among said plurality of SNPs between said DNA profile and said one or more reference DNA profiles.
71. 71. The method of any one of claims 68-70, wherein calculating the likelihood ratio comprises dividing the probability that the DNA profile and a reference DNA profile from among the one or more reference DNA profiles are related by the probability that the DNA profile and the reference DNA profile are unrelated, based on the genotypes of the plurality of SNPs.
72. The likelihood ratio (LR) is calculated as follows: [Number 69] In the formula, D represents the genotype, and H r represents the hypothesis that the individuals are related, and H u 72. The method of any one of claims 68 to 71, wherein represents the hypothesis that the individuals are unrelated.
73. The LR is calculated as follows: [Number 70] 72. The method of any one of claims 68 to 71, wherein 0.001 represents the genotyping error rate, p is the allele frequency of allele 1, and q is the allele frequency of allele 2.
74. 74. The method of any one of claims 1-73, wherein the person of interest is biologically male, and the method further comprises calculating a likelihood ratio of sharing a Y chromosome between the DNA profile and the one or more reference DNA profiles.
75. 75. The method of claim 74, wherein said calculating a likelihood ratio of sharing a Y chromosome comprises comparing a set of SNPs comprising one or more Y-SNPs between said DNA profile and said one or more reference DNA profiles.
76. 76. The method of claim 75, wherein said one or more Y-SNPs are comprised within said plurality of SNPs.
77. 76. The method of claim 75, wherein the one or more Y-SNPs comprise at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 81, 82, 83, 84, or 85 Y-SNPs.
78. 78. The method of any one of claims 75 to 77, wherein the one or more Y-SNPs comprise 85 Y-SNPs.
79. 79. The method of any one of claims 75 to 78, wherein calculating the likelihood ratio of sharing a Y chromosome comprises dividing the probability that the DNA profile and a reference DNA profile from among the one or more reference DNA profiles share a Y chromosome by the probability that the DNA profile and the reference DNA profile do not share a Y chromosome, based on the genotypes of the one or more Y-SNPs.
80. 80. The method of any one of claims 1 to 79, wherein at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the DNA profiles in the reference set of DNA profiles are from relatives of missing persons or victims of disaster or conflict.
81. 81. The method of any one of claims 1 to 80, wherein each of the DNA profiles in the reference set of DNA profiles is from a relative of a missing person or a victim of a disaster or conflict.
82. 82. The method of any one of claims 1-81, wherein the reference set of DNA profiles comprises up to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 reference DNA profiles.
83. 83. The method of any one of claims 1 to 82, wherein the reference set of DNA profiles comprises up to 100 reference DNA profiles.
84. 84. The method of any one of claims 1-83, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest.
85. 85. The method of any one of claims 1 to 84, wherein at least 50% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest.
86. 86. The method of claim 84 or claim 85, wherein each relative of the person of interest in the reference set of DNA profiles is a first-, second-, third-, fourth-, or fifth-degree relative of the person of interest, respectively.
87. 87. The method of any one of claims 1 to 86, wherein the identity of each relative of the person of interest in the reference set of DNA profiles is known.
88. 88. The method of any one of claims 1 to 87, wherein the identity of each relative of the person of interest in the reference set of DNA profiles is known.
89. 89. The method of any one of claims 1 to 88, wherein the identity of each of the one or more reference DNA profiles in the reference set of DNA profiles is known.
90. 90. The method of any one of claims 1 to 89, wherein the reference set of DNA profiles is in a database.
91. 91. The method of claim 90, wherein the database is not publicly accessible.
92. 92. The method of claim 90 or claim 91, wherein the database is not accessible by a third party genealogy service.
93. A nucleic acid library constructed using the method of any one of claims 6 to 92.
94. A plurality of primers that specifically hybridize to a plurality of target sequences comprising at least or about 2,000 to 50,000 single nucleotide polymorphisms (SNPs) in a nucleic acid sample from a person of interest, wherein amplification of the nucleic acid sample using the plurality of primers in one or more multiplex PCR reactions results in amplification products.
95. a plurality of primers that specifically hybridize to a plurality of target sequences comprising between at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs) in a nucleic acid sample from the person of interest and one or more reference samples; the one or more reference samples comprise samples from relatives of the person of interest; A plurality of primers, wherein amplifying said nucleic acid sample from said person of interest and said nucleic acid sample from one or more reference samples using said plurality of primers in one or more multiplex PCR reactions results in amplification products.
96. 96. The plurality of primers of claim 94 or claim 95, wherein the nucleic acid sample from the person of interest comprises genomic DNA.
97. 97. The plurality of primers of any one of claims 94 to 96, wherein the nucleic acid sample from the person of interest comprises one or more enzyme inhibitors.
98. 98. The plurality of primers of claim 97, wherein the one or more enzyme inhibitors comprise one or more inhibitors selected from the group consisting of hematin, heme, humic acid, indigo, tannic acid, collagen, calcium, and hydroxyapatite.
99. 99. The plurality of primers of any one of claims 94 to 98, wherein the nucleic acid sample from the person of interest comprises low quality and / or low abundance nucleic acid molecules.
100. 100. The plurality of primers of claim 99, wherein the low quality nucleic acid molecules are degraded and / or fragmented genomic DNA.
101. the low quality nucleic acid molecules have a Degradation Index (DI) of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, or 200, or at least 1, 2, 3, 101. The plurality of primers of claim 99 or claim 100, having a Degradation Index (DI) of 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195 or 200.
102. 102. The plurality of primers of any one of claims 99 to 101, wherein the low quality nucleic acid molecules have a DI of at least 1 and at most 158.3 or less.
103. 103. The plurality of primers of any one of claims 94 to 102, wherein the nucleic acid sample from the person of interest and / or the nucleic acid sample from one or more reference samples comprises high quality nucleic acid molecules.
104. 104. The plurality of primers of Claim 103, wherein said high quality nucleic acid molecules have a DI of less than 1.
105. 105. The plurality of primers of any one of claims 94 to 104, wherein the person of interest is a missing person.
106. 105. The plurality of primers of any one of claims 94 to 104, wherein the person of interest is a victim of a disaster or conflict.
107. 107. The plurality of primers of any one of claims 94-106, wherein the nucleic acid sample from the person of interest is derived from a buccal swab, paper, fabric or other substrate or object impregnated with saliva, blood or other bodily fluid, or containing hair or skin cells.
108. 108. The plurality of primers of any one of claims 94-107, wherein the nucleic acid sample from the person of interest comprises at or about 3 pg to 100 ng of genomic DNA.
109. 109. The plurality of primers of any one of claims 94-108, wherein the nucleic acid sample from the person of interest comprises between or about 100 pg and 5 ng of genomic DNA, between or about 50 pg and 5 ng of genomic DNA, or between or about 3 pg and 5 ng of genomic DNA.
110. 110. The plurality of primers of claim 108 or claim 109, wherein the nucleic acid sample from the person of interest comprises 1 ng or about 1 ng of genomic DNA.
111. 111. The plurality of primers of any one of claims 94 to 110, wherein the plurality of SNPs comprises kinship SNPs (kiSNPs).
112. 112. The plurality of primers of any one of claims 94 to 111, wherein the plurality of SNPs comprises Y chromosome SNPs (Y-SNPs).
113. 113. The plurality of primers of any one of claims 94 to 112, wherein the plurality of SNPs comprises a kiSNP and a Y-SNP.
114. 114. The plurality of primers of any one of claims 94 to 113, wherein the plurality of SNPs comprises kiSNPs, biogeographic ancestry SNPs (aiSNPs), identity SNPs (iiSNPs), phenotypic SNPs (piSNPs), X chromosome SNPs (X-SNPs), and Y chromosome SNPs (Y-SNPs).
115. 112. The plurality of primers of any one of claims 94 to 111, wherein the plurality of SNPs comprises SNPs selected from one or more of the group consisting of kiSNPs, aiSNPs, iiSNPs, piSNPs, X-SNPs and Y-SNPs.
116. 116. The plurality of primers of any one of claims 94-115, wherein at least or at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% of the plurality of SNPs are related SNPs.
117. 117. The plurality of primers of any one of claims 94 to 116, wherein at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the DNA profiles in the reference set of DNA profiles are from relatives of missing persons or victims of disaster or conflict.
118. 117. The plurality of primers of any one of claims 94 to 116, wherein each of the one or more reference samples is from a relative of a missing person or a victim of a disaster or conflict.
119. 119. The plurality of primers of any one of claims 94-118, wherein the one or more reference samples comprise up to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 reference samples.
120. 120. The plurality of primers of any one of claims 94-119, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the one or more reference samples are from relatives of the person of interest.
121. 121. The plurality of primers of any one of claims 94-120, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest, and wherein each of said at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are first-, second-, third-, fourth-, or fifth-degree relatives.
122. 122. The plurality of primers of any one of claims 94 to 121, wherein at least 50% of said one or more reference samples are from relatives of said person of interest.
123. 123. The plurality of primers of any one of claims 120-122, wherein each relative of the person of interest in the one or more reference samples is a first-, second-, third-, fourth-, or fifth-degree relative of the person of interest, respectively.
124. 124. The plurality of primers of Claim 123, wherein each relative of the person of interest in the one or more reference samples is a first-degree, second-degree, or third-degree relative of the person of interest, respectively.
125. 125. The plurality of primers of any one of claims 94 to 124, wherein the identity of each relative of the person of interest in the one or more reference samples is known.
126. 126. The plurality of primers of any one of claims 94 to 125, wherein the identity of each of the one or more reference samples is known.
127. 1. A method for constructing a DNA profile, comprising: providing a nucleic acid sample from the person of interest; amplifying the nucleic acid sample using a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of between at least or about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplification products, wherein the amplification is performed in one or more multiplex PCR reactions; sequencing the amplification products; genotyping the plurality of SNPs, thereby generating a DNA profile; and A method comprising:
128. 1. A method for constructing a DNA profile, comprising: providing a nucleic acid sample from the person of interest; providing a nucleic acid sample from a relative of the person of interest; amplifying the nucleic acid sample from the person of interest and the nucleic acid sample from the relatives using a plurality of primers that specifically hybridize to a plurality of target sequences that collectively comprise a plurality of at least or between about 2,000 and 50,000 single nucleotide polymorphisms (SNPs), thereby generating amplification products, wherein the amplification is performed in one or more multiplex PCR reactions; sequencing the amplification products; genotyping the plurality of SNPs, thereby generating a DNA profile for the person of interest and the relatives of the person of interest; A method comprising:
129. 129. The method of claim 127 or claim 128, wherein said sequencing does not include whole genome sequencing (WGS).
130. 130. The method of claim 127 or claim 129, wherein the nucleic acid sample comprises genomic DNA.
131. 130. The method of claim 128 or claim 129, wherein the nucleic acid sample of the person of interest and / or the nucleic acid sample of the relative of the person of interest comprises genomic DNA.
132. 132. The method of any one of claims 127 to 131, wherein the nucleic acid sample, the nucleic acid sample of the person of interest and / or the nucleic acid sample of the relative comprises one or more enzyme inhibitors.
133. 133. The method of claim 132, wherein the one or more enzyme inhibitors comprise one or more inhibitors selected from the group consisting of hematin, heme, humic acid, indigo, tannic acid, collagen, calcium, and hydroxyapatite.
134. 134. The method of any one of claims 127 to 133, wherein the nucleic acid sample, the nucleic acid sample of the person of interest and / or the nucleic acid sample of the relative comprises low quality and / or low abundance nucleic acid molecules.
135. 135. The method of claim 134, wherein the low quality nucleic acid molecules are degraded and / or fragmented genomic DNA.
136. The low quality nucleic acid molecules have a Degradation Index (DI) of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195 or 200, or at least 1, 2 , 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195 or 200.
137. 137. The method of any one of claims 134 to 136, wherein said low quality nucleic acid molecules have a DI of at least 1 and at most 158.3 or less.
138. 134. The method of any one of claims 127 to 133, wherein the nucleic acid sample, the nucleic acid sample of the person of interest and / or the nucleic acid sample of the relative comprises high quality nucleic acid molecules.
139. 139. The method of Claim 138, wherein said high quality nucleic acid molecules have a DI of less than 1.
140. 140. The method of any one of claims 127 to 139, wherein the person of interest is a missing person.
141. 140. A method according to any one of claims 127 to 139, wherein the person of interest is a victim of a disaster or conflict.
142. 142. The method of any one of claims 128, 129 and 131-141, wherein the relative of the person of interest is a first, second, third, fourth or fifth degree relative.
143. 142. The method of any one of claims 128, 129 and 131-141, wherein the relative of the person of interest is a first, second or third degree relative.
144. 144. The method of any one of claims 127 to 143, wherein the nucleic acid sample, the nucleic acid sample of the person of interest and / or the nucleic acid sample of the relative is derived from a buccal swab, paper, fabric or other substrate or object impregnated with saliva, blood or other bodily fluid, or containing hair or skin cells.
145. 145. The method of any one of claims 127 to 144, wherein the nucleic acid sample, the nucleic acid sample of the person of interest and / or the nucleic acid sample of the relative comprises at or about 3 pg to 100 ng of genomic DNA.
146. 146. The method of any one of claims 127 to 145, wherein the nucleic acid sample, the nucleic acid sample of the person of interest and / or the nucleic acid sample of the relative comprises between or about 100 pg and 5 ng of genomic DNA, between or about 50 pg and 5 ng of genomic DNA, or between or about 3 pg and 5 ng of genomic DNA.
147. 147. The method of claim 145 or claim 146, wherein the nucleic acid sample, the nucleic acid sample of the person of interest and / or the nucleic acid sample of the relative comprises 1 ng or about 1 ng of genomic DNA.
148. The method of any one of claims 127 to 147, wherein the plurality of SNPs comprises related SNPs.
149. 149. The method of any one of claims 127 to 148, wherein the plurality of SNPs comprises Y chromosome SNPs (Y-SNPs).
150. The method of any one of claims 127 to 149, wherein the plurality of SNPs comprises kiSNPs and Y-SNPs.
151. 151. The method of any one of claims 127 to 150, wherein the plurality of SNPs comprises kiSNPs, biogeographic ancestry SNPs (aiSNPs), identity SNPs (iiSNPs), phenotypic SNPs (piSNPs), X chromosome SNPs (X-SNPs), and Y chromosome SNPs (Y-SNPs).
152. 152. The method of any one of claims 127-151, wherein the plurality of SNPs comprises SNPs selected from one or more of the group consisting of kiSNPs, aiSNPs, iiSNPs, piSNPs, X-SNPs and Y-SNPs.
153. 153. The method of any one of claims 127-152, wherein at least or at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% of the plurality of SNPs are related SNPs.
154. 154. The method of any one of claims 1-92 and 127-153, wherein said sequencing comprises a sequencing complexity of up to 40-plex.
155. 154. The method of any one of claims 1-92 and 127-153, wherein said sequencing comprises a sequencing complexity of up to 32-plex.
156. 154. The method of any one of claims 1-92 and 127-153, wherein said sequencing comprises a sequencing complexity of 12-plex to 32-plex.
157. 154. The method of any one of claims 1-92 and 127-153, wherein said sequencing comprises a sequencing complexity of 24-plex to 32-plex.
158. The sequencing may be performed in a 4-plex, 5-plex, 6-plex, 7-plex, 8-plex, 9-plex, 10-plex, 11-plex, 12-plex, 13-plex, 14-plex, 15-plex, 16-plex, 17-plex, or 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex, 35-plex, 36-plex, 37-plex, 38-plex, 39-plex, 40-plex, 41-plex, 42-plex, 43-plex, 44-plex or 45-plex, or about 4-plex, 5-plex, 6-plex, 7-plex, 8-plex, 9-plex, 10-plex, 11-plex, 12-plex, 13-plex, 14-plex, 15-plex, 16-plex, 17-plex comprising a sequencing complexity of 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex, 35-plex, 36-plex, 37-plex, 38-plex, 39-plex, 40-plex, 41-plex, 42-plex, 43-plex, 44-plex or 45-plex; or The sequencing may be at or about 10-plex, 11-plex, 12-plex, 13-plex, 14-plex, 15-plex, 16-plex, 17-plex 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex, or 35-plex. including a sequencing complexity of 18-plex, 19-plex, 20-plex, 21-plex, 22-plex, 23-plex, 24-plex, 25-plex, 26-plex, 27-plex, 28-plex, 29-plex, 30-plex, 31-plex, 32-plex, 33-plex, 34-plex or 35-plex; 154. The method of any one of claims 1 to 92 and 127 to 153.
159. 159. The method of any one of claims 1-92 and 127-158, wherein the sequencing comprises a sequencing complexity of at or about 8-16 plex for post-mortem samples, and / or wherein the sequencing comprises a sequencing complexity of at or about 24-40 plex for ante-mortem samples.
160. 159. The method of any one of claims 1-92 and 127-158, wherein the sequencing comprises a sequencing complexity of at or about 12-plex for post-mortem samples, and / or the sequencing comprises a sequencing complexity of at or about 32-plex for ante-mortem samples.
161. 154. The method of any one of claims 1-92 and 127-153, wherein the sequencing comprises a sequencing complexity of, or about, 30-plex, 31-plex, or 32-plex.
162. 1. A method for identifying genetic relatives of a DNA profile, comprising: Calculating a degree of relatedness between the DNA profile of any one of claims 127 to 161 and one or more reference DNA profiles, wherein the one or more reference DNA profiles are comprised within a reference set of DNA profiles comprising one or more reference DNA profiles from blood relatives of the person of interest; generating a pedigree comprising said DNA profile in relation to said one or more reference DNA profiles.
163. 163. The method of claim 162, wherein the one or more reference DNA profiles are part of a database.
164. 164. The method of claim 162 or claim 163, wherein the reference set of DNA profiles comprises up to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 reference DNA profiles.
165. 165. The method of any one of claims 162-164, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest.
166. 166. The method of any one of claims 162 to 165, wherein at least 50% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest.
167. 167. The method of claim 165 or claim 166, wherein each relative of the person of interest is a first-degree, second-degree, third-degree, fourth-degree, or fifth-degree relative of the person of interest, respectively.
168. 168. The method of any one of claims 162 to 167, wherein the identity of each relative of the person of interest in the reference set of DNA profiles is known.
169. 169. The method of any one of claims 162 to 168, wherein the identity of each of said one or more reference DNA profiles in said reference set of DNA profiles is known.
170. 170. The method of any one of claims 162 to 169, wherein the reference set of DNA profiles is in a database.
171. 171. The method of claim 170, wherein the database is not publicly accessible.
172. 172. The method of claim 170 or claim 171, wherein the database is not accessible by a third party genealogy service.
173. 1. A method for confirming the identity of a DNA profile, comprising: calculating a degree of relatedness between a DNA profile comprising genotypes of at least or between about 2,000 and 50,000 SNPs and one or more reference DNA profiles, wherein the DNA profile is from a person of interest and the one or more reference DNA profiles are included within a reference set of DNA profiles comprising one or more reference DNA profiles from relatives of the person of interest; generating a pedigree comprising said DNA profile in relation to said one or more reference DNA profiles.
174. The method of claim 173, wherein the DNA profile is generated by the method of any one of claims 127 to 161.
175. 175. A method according to claim 173 or claim 174, wherein the degree of relatedness is calculated using a kinship model.
176. 176. A method according to any one of claims 173 to 175, wherein the degree of relatedness is calculated using a kinship model trained using a PCA method.
177. 177. The method of claim 176, wherein the PCA method for training the kinship model is or comprises PCA.
178. The method of claim 176 or claim 177, wherein the PCA method is PC-AiR.
179. 179. The method of claim 178, wherein the PC-AiR comprises: (1) estimating relatedness coefficients between all pairs of samples in a training database, and optionally in a training DNA profile, wherein pairings with a relatedness coefficient >0.025 are identified as closely related and pairings with a relatedness coefficient <-0.025 are identified as ancestrally diverged; (2) initializing an unrelated sample set containing all samples; and (3) iteratively: (i) identifying a set in the unrelated sample set that has the most related samples in the unrelated sample set, thereby designating this set as X; (ii) identifying a set of samples in X that has the fewest ancestrally diverged pairings compared to samples in the unrelated sample set, thereby designating this set as Y; and (iii) terminating the process if Y has 0 samples, or randomly selecting one sample from Y and removing it from U if Y has at least one sample, and repeating beginning with step (3)(i).
180. The method of claim 176 or claim 177, wherein the PCA method is modified PC-Air.
181. 181. The method of claim 180, wherein the modified PC-AiR comprises: (1) estimating relatedness coefficients between all pairs of samples of training DNA profiles, as appropriate, in a training database, wherein pairings with relatedness coefficients >0.01 are identified as closely related and pairings with relatedness coefficients <-0.025 are identified as ancestrally diverged; (2) removing all DNA profiles with ≧5% missing data; and (3) ranking all DNA profiles by assigning each DNA profile a ranking value. In some embodiments, the ranking value is determined based on the number of related DNA profiles in the complete database ranked from smallest to largest, broken down by the number of ancestrally branched DNA profiles in the complete database ranked from largest to smallest. In some embodiments, step (3) includes iterating through the ranked DNA profiles, and for each DNA profile, (i) if the DNA profile is not already in the related sample set, adding it to the unrelated sample set and adding all related DNA profiles to the related sample set, and (ii) if the DNA profile is already in the related sample set, skipping to the next DNA profile and repeating starting with step (3)(i).
182. 182. The method of any one of claims 173 to 181, wherein said calculating said degree of relatedness comprises calculating a coefficient of relatedness using PC-Relate.
183. 183. The method of claim 182, wherein the degree of relatedness is calculated by providing the DNA profile of the person of interest as input to PC-Relate.
184. 184. A method according to claim 182 or claim 183, wherein the degree of relatedness is calculated by providing the kinship model and the DNA profile of the person of interest as input to PC-Relate.
185. 185. The method of any one of claims 182 to 184, wherein the one or more reference DNA profiles are further provided as input to PC-Relate.
186. Calculating the degree of relatedness includes calculating a coefficient of relatedness using a whole-genome kinship algorithm as follows: [Number 71] the reference DNA profiles of the person of interest and the one or more reference DNA profiles are i and j; [Number 72] is the coefficient of kinship, [Number 73] is the estimated allele frequency, s is the SNP in S SNPs classified in both individuals, [Number 74] and [Number 75] are the numbers of reference alleles at i and j at SNP s, respectively; [Number 76] and [Number 77] are the expected allele frequencies calculated by PC-AiR for i and j at SNP s, respectively.
187. 187. The method of any one of claims 173 to 186, wherein said calculating said degree of association comprises calculating a likelihood ratio.
188. 188. The method of claim 187, wherein said calculating said likelihood ratio comprises comparing said plurality of SNPs between said DNA profile and said one or more reference DNA profiles.
189. 188. The method of claim 187, wherein said calculating said likelihood ratio comprises comparing a set of SNPs comprising kinship SNPs from among said plurality of SNPs between said DNA profile and said one or more reference DNA profiles.
190. 190. The method of any one of claims 187-189, wherein calculating the likelihood ratio comprises dividing the probability that the DNA profile and a reference DNA profile from among the one or more reference DNA profiles are related by the probability that the DNA profile and the reference DNA profile are unrelated, based on the genotypes of the plurality of SNPs.
191. The likelihood ratio (LR) is calculated as follows: [Number 78] In the formula, D represents the genotype, and H r represents the hypothesis that the individuals are related, and H u 191. The method of any one of claims 187 to 190, wherein represents the hypothesis that the individuals are unrelated.
192. The LR is calculated as follows: [Number 79] 191. The method of any one of claims 187 to 190, wherein 0.001 represents the genotyping error rate, p is the allele frequency of allele 1, and q is the allele frequency of allele 2.
193. 193. The method of any one of claims 173-192, wherein the person of interest is biologically male, and wherein the method further comprises calculating a likelihood ratio of sharing a Y chromosome between the DNA profile and the one or more reference DNA profiles.
194. 194. The method of claim 193, wherein said calculating a likelihood ratio of sharing a Y chromosome comprises comparing a set of SNPs comprising one or more Y-SNPs between said DNA profile and said one or more reference DNA profiles.
195. 195. The method of claim 194, wherein said one or more Y-SNPs are comprised within said plurality of SNPs.
196. 196. The method of claim 194 or claim 195, wherein the one or more Y-SNPs comprise at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 81, 82, 83, 84, or 85 Y-SNPs.
197. 197. The method of any one of claims 194 to 196, wherein the one or more Y-SNPs comprise 85 Y-SNPs.
198. 200. The method of any one of claims 193 to 197, wherein calculating the likelihood ratio of sharing a Y chromosome comprises dividing the probability that the DNA profile and a reference DNA profile from among the one or more reference DNA profiles share a Y chromosome by the probability that the DNA profile and the reference DNA profile do not share a Y chromosome, based on the genotypes of the one or more Y-SNPs.
199. 200. The method of any one of claims 173 to 198, wherein at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the DNA profiles in the reference set of DNA profiles are from relatives of missing persons or victims of disaster or conflict.
200. 200. The method of any one of claims 173 to 199, wherein each of the one or more reference samples is from a relative of a missing person or a victim of a disaster or conflict.
201. 201. The method of any one of claims 173-200, wherein the one or more reference samples comprise up to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 reference samples.
202. 202. The method of any one of claims 173 to 201, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the one or more reference samples are from relatives of the person of interest.
203. 203. The method of any one of claims 173-202, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are from relatives of the person of interest, and wherein each of said at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the reference DNA profiles in the reference set of DNA profiles are first-, second-, third-, fourth-, or fifth-degree relatives.
204. 204. The method of any one of claims 173 to 203, wherein at least 50% of said one or more reference samples are from relatives of said person of interest.
205. 205. The method of any one of claims 199 to 204, wherein each relative of the person of interest in the one or more reference samples is a first-, second-, third-, fourth-, or fifth-degree relative of the person of interest, respectively.
206. 206. The method of claim 205, wherein each relative of the person of interest in the one or more reference samples is a first-degree, second-degree, or third-degree relative of the person of interest, respectively.
207. 207. The method of any one of claims 173 to 206, wherein the identity of each relative of the person of interest in the one or more reference samples is known.
208. 208. The method of any one of claims 173 to 207, wherein the identity of each of the one or more reference samples is known.
209. A kit comprising at least one container means, said at least one A kit wherein the container means comprises a plurality of the primers of any one of claims 94 to 126.
210. The plurality of SNPs may be selected from the group consisting of: between 2,000 and 11,000 SNPs; between 3,000 and 11,000 SNPs; between 4,000 and 11,000 SNPs; between 5,000 and 11,000 SNPs; between 5,500 and 11,000 SNPs; between 6,000 and 11,000 SNPs; between 7,000 and 15,000 SNPs; between 7,000 and 14,000 SNPs; between 7,000 and 13,000 SNPs; between 7,000 and 12,000 SNPs; between 7,000 and 11,000 SNPs; SNPs between 8,000 and 15,000 SNPs, between 8,000 and 14,000 SNPs, between 8,000 and 13,000 SNPs, between 8,000 and 12,000 SNPs, between 8,000 and 11,000 SNPs, between 9,000 and 15,000 SNPs, between 9,000 and 14,000 SNPs, between 9,000 and 13,000 SNPs, between 9,000 and 12,000 SNPs or between 9,000 and 11,000 SNPs, or about 2,000 to 11,000 SNPs. 00 SNPs, 3,000-11,000 SNPs, 4,000-11,000 SNPs, 5,000-11,000 SNPs, 5,500-11,000 SNPs, 6,000-11,000 SNPs, 7,000-15,000 SNPs, 7,000-14,000 SNPs, 7,000-13,000 SNPs, 7,000-12,000 SNPs, 7,000-11,000 SNPs, 8,000-15,000 SNPs 209. The method of any one of claims 1 to 92 and 127 to 208, comprising between 8,000 and 14,000 SNPs, between 8,000 and 13,000 SNPs, between 8,000 and 12,000 SNPs, between 8,000 and 11,000 SNPs, between 9,000 and 15,000 SNPs, between 9,000 and 14,000 SNPs, between 9,000 and 13,000 SNPs, between 9,000 and 12,000 SNPs or between 9,000 and 11,000 SNPs.
211. The method of any one of claims 1 to 92 and 127 to 208, wherein the plurality of SNPs comprises 10,230 SNPs.
212. The plurality of SNPs may be selected from the group consisting of: between 2,000 and 11,000 SNPs; between 3,000 and 11,000 SNPs; between 4,000 and 11,000 SNPs; between 5,000 and 11,000 SNPs; between 5,500 and 11,000 SNPs; between 6,000 and 11,000 SNPs; between 7,000 and 15,000 SNPs; between 7,000 and 14,000 SNPs; between 7,000 and 13,000 SNPs; between 7,000 and 12,000 SNPs; between 7,000 and 11,000 SNPs; SNPs between 8,000 and 15,000 SNPs, between 8,000 and 14,000 SNPs, between 8,000 and 13,000 SNPs, between 8,000 and 12,000 SNPs, between 8,000 and 11,000 SNPs, between 9,000 and 15,000 SNPs, between 9,000 and 14,000 SNPs, between 9,000 and 13,000 SNPs, between 9,000 and 12,000 SNPs or between 9,000 and 11,000 SNPs, or about 2,000 to 11,000 SNPs. 000 SNPs, 3,000-11,000 SNPs, 4,000-11,000 SNPs, 5,000-11,000 SNPs, 5,500-11,000 SNPs, 6,000-11,000 SNPs, 7,000-15,000 SNPs, 7,000-14,000 SNPs, 7,000-13,000 SNPs, 7,000-12,000 SNPs, 7,000-11,000 SNPs, 8,000-15,000 SNPs 127. The plurality of primers of any one of claims 94 to 126, comprising between 8,000 and 14,000 SNPs, between 8,000 and 13,000 SNPs, between 8,000 and 12,000 SNPs, between 8,000 and 11,000 SNPs, between 9,000 and 15,000 SNPs, between 9,000 and 14,000 SNPs, between 9,000 and 13,000 SNPs, between 9,000 and 12,000 SNPs or between 9,000 and 11,000 SNPs.
213. 127. The plurality of primers of any one of claims 94 to 126, wherein the plurality of SNPs comprises 10,230 SNPs.
214. 214. The method of any one of claims 1-5, 11-93, 127-161, and 210-213, further comprising generating a pedigree comprising said DNA profile in relation to one or more DNA profiles contained within said reference set of DNA profiles.
215. 215. The method of claim 214, wherein said pedigree tree includes said DNA profile in association with one or more DNA profiles from blood relatives of said person of interest.