Methods and compositions to compare nucleic acid samples

The method of obtaining sequence information at multiple sites using specific primers and DNA polymerase generates a sample identity barcode to confirm nucleic acid sample identity, addressing swapping issues and ensuring accurate sample matching.

WO2025184044A1PCT designated stage Publication Date: 2025-09-04ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/017067
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-26
Filing Date
2025-02-24
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Inaccurate nucleic acid sample identification due to potential swapping during collection, data entry, or sequencing processes, leading to incorrect sequencing results.

Method used

A method involving obtaining sequence information at multiple sites using specific primers and DNA polymerase, generating a sample identity barcode (SIB) for nucleic acid samples, and comparing SIBs to confirm sample identity.

Benefits of technology

Enables rapid and cost-effective confirmation of nucleic acid sample identity, reducing the need for additional sequencing and identifying sample swaps, with applications in confirming individual sample matches and detecting contaminants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025017067_04092025_PF_FP_ABST
    Figure US2025017067_04092025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are methods and systems for comparing nucleic acids samples. Some embodiments include confirming a nucleic acid sample identity. Some embodiments include obtaining a nucleic acid sample from a human biological sample, determining sequence information for the nucleic acid sample at a plurality of sites, and based on the sequence information at each of the plurality of sites, confirming the nucleic acid sample identity. Further disclosed herein are kits for confirming a nucleic acid sample identity, and systems for confirming a nucleic acid sample identity. Further disclosed herein are methods of amplifying a human nucleic acid sample.
Need to check novelty before this filing date? Find Prior Art

Description

ILLINC.830WO PATENT METHODS AND COMPOSITIONS TO COMPARE NUCLEIC ACID SAMPLES CROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Prov. App. No. 63 / 570001 filed March 26, 2024, and to U.S. Prov. App. No.63 / 557781 filed February 26, 2024 which are each incorporated by reference in its entirety. REFERENCE TO SEQUENCE LISTING

[0002] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as a file entitled ILLINC.830WO.XML, created February 18, 2025, which is approximately 1,240,751 bytes in size. The information in the electronic format of the Sequence Listing is incorporated herein by reference in its entirety. FIELD

[0003] The present disclosure relates to the field of nucleic acid sample identification. More particularly, the present disclosure relates to the field of comparing nucleic acid samples. Some embodiments include confirming a nucleic acid sample identity by determining sequence information for the nucleic acid sample at a plurality of sites. BACKGROUND

[0004] There are many stages where a biological sample which includes nucleic acids, or the associated sequencing information, may be inadvertently swapped between individuals, leading to inaccurate results. For example, a sample may be swapped during sample collection where a physician takes samples from multiple members of the same family in one consultation, raising the possibility that swabs from each member did not go in the correctly labelled tube. Sequencing information may be swapped during data entry, as many locations rely on cutting and pasting data in a data table. Samples may also be inadvertently swapped when the samples are packaged and sent to a third party for sequencing. Furthermore, a sequencing pipeline may be a multistep process where there are opportunities for incorrect plating. Because of this potential for inaccurate results sequencing lab and health careprofessionals may need to check that the next-generation sequencing (NGS) test result they are about to return to an individual really belongs to that particular individual. SUMMARY

[0005] Some embodiments of the methods and compositions provided herein include a method for confirming a nucleic acid sample identity, comprising: obtaining a nucleic acid sample from a human biological sample; determining sequence information for the nucleic acid sample at a plurality of sites, wherein the plurality of sites comprises 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or each of the sites corresponding to positions: chr1:6693097, chr1:41204569, chr1:59147926, chr1:114515717, chr1:150808889, chr1:212911836, chr1:227069737, chr1:227071525, chr1:227935444, chr1:236716959, chr1:236719193, chr2:3392295, chr2:101624471, chr2:101638888, chr2:128939817, chr3:4403767, chr3:15737689, chr3:33434831, chr3:155481609, chr3:186509517, chr3:193374964, chr4:6600012, chr4:6606864, chr4:25419283, chr5:53815495, chr5:64881936, chr5:82834630, chr5:96503523, chr6:26598188, chr6:49403282, chr6:49425521, chr6:101166095, chr6:146755140, chr6:158517308, chr7:2577781, chr7:2578237, chr7:43916727, chr7:43917013, chr7:48004962, chr8:30973957, chr8:33356074, chr8:33369944, chr8:90995019, chr8:94935937, chr8:104427359, chr8:124448736, chr8:132982824, chr8:135612745, chr8:146067054, chr9:132400480, chr10:1046712, chr10:31138817, chr10:99219885, chr10:100219314, chr10:104572963, chr11:16133413, chr11:125763746, chr12:993930, chr12:62926398, chr13:28143229, chr13:39433606, chr14:20920250, chr14:31381351, chr15:34528948, chr15:75131959, chr16:624114, chr16:2028402, chr16:19680546, chr16:70303580, chr17:10614442, chr17:17168164, chr17:40714804, chr17:40716520, chr17:57290383, chr17:71196809, chr17:79514129, chr17:80008392, chr18:21413869, chr19:885818, chr19:1010691, chr19:36727365, chr19:41939297, chr19:50983930, chr19:58929052, chr20:6100088, chr20:61834695, chr21:47961711, chr22:32795641, chr22:32875190, chr22:50962208, chrY:2961349, chrY:7063980, chrY:2866789, chrY:2980414, chrY:7091532, chrY:13705715, chrY:14842155, or chrY:20592667 of hg 19 - GRCh37, and based on the sequence information at each of the plurality of sites, confirming the nucleic acid sample identity.

[0006] In some embodiments, determining sequence information comprises: contacting DNA or cDNA from the nucleic acid sample with a plurality of pairs of primers, a DNA polymerase buffer, and a DNA polymerase, amplifying the DNA or cDNA at the plurality of sites to generate amplified nucleic acid fragments, and sequencing the amplified nucleic acid fragments.

[0007] In some embodiments, the plurality of pairs of primers comprises at least 20 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392. In some embodiments, the plurality of pairs of primers comprises at least 50 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392. In some embodiments, the plurality of pairs of primers comprises at least 70 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392. In some embodiments, the plurality of pairs of primers comprises at least 90 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197- 392. In some embodiments, the plurality of pairs of primers comprises pairs of primers comprising a sequence as set forth in each of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.

[0008] In some embodiments, the DNA polymerase buffer comprises Illumina TTM buffer.

[0009] In some embodiments, the DNA polymerase comprises Thermo Phusion Hot Start II DNA Polymerase. In some embodiments, the nucleic acid sample comprises DNA. In some embodiments, the nucleic acid sample comprises RNA. Some embodiments also include obtaining cDNA based on the RNA from the nucleic acid sample.

[0010] In some embodiments, determining sequence information comprises determining whether the nucleic acid sample is homozygous for a reference allele, homozygous for an alternative allele, or heterozygous at each site of the plurality of sites.

[0011] In some embodiments, the human biological sample comprises a blood sample, a tumor sample, a saliva sample, a hair sample, a semen sample, a skin sample, a muscle sample, or an organ sample.

[0012] Some embodiments also include confirming a nucleic acid sample sequence identity for a second nucleic acid sample from a second human biological sample.

[0013] In some embodiments, the human biological sample and the second human biological sample are from an individual.

[0014] In some embodiments, the two human biological samples comprise a tumor sample and a normal sample from the individual.

[0015] Some embodiments also include excluding sequence information related to a site of the plurality of sites when the sequence information at the site indicates heterozygosity.

[0016] In some embodiments, amplifying DNA or cDNA from each of the one or more nucleic acid samples comprises multiplex PCR.

[0017] Some embodiments also include adding an adapter sequence to amplified nucleic acid fragments. In some embodiments, the adapter sequence is added by amplification.

[0018] Some embodiments include contacting nucleic acid fragments with a plurality of primers comprising an adapter sequence.

[0019] In some embodiments, determining sequence information comprises hybridizing the amplified nucleic acid fragments to a nucleic acid adapter comprising a sequence complementary to the adapter sequence.

[0020] In some embodiments, sequencing is performed in multiplex format.

[0021] Some embodiments of the methods and compositions provided herein include a computer-implemented method of comparing nucleic acid samples, comprising: (a) receiving sequence information for a first nucleic acid sample at a plurality of sites; (b) generating from the sequence information a first sample identity barcode (SIB) for the first nucleic acid sample, wherein the first SIB comprises an ordered list of identifiers for the plurality of sites; (c) generating a ratio of identity by comparing the first SIB and a second SIB, wherein the second SIB is generated from sequence information for a second nucleic acid sample; and (d) providing an indication of a match between the first and second nucleic acid samples based on the comparison.

[0022] In some embodiments, the plurality of sites comprises: (i) a plurality of single nucleotide polymorphisms (SNPs); (ii) locations distributed throughout a genome; (iii) locations in exons; (iv) biallelic sites; (v) locations within ubiquitously expressed genes; (vi)lack allele specific expression; and / or (vii) locations comprising a minor allele frequency in African, South and East Asian and European populations in a range from 0.3 to 0.7.

[0023] In some embodiments, the sequence information is obtained by any one of the foregoing methods.

[0024] In some embodiments, step (b) comprises obtaining a genotyped list for the plurality of sites; optionally, wherein a file comprises the genotyped list, wherein the file comprises a format selected from a variable call file (VCF), binary alignment map (BAM), and a compressed reference-oriented alignment map (CRAM).

[0025] In some embodiments, obtaining the genotyped list comprises: (i) aligning the sequence information with a reference; and (ii) genotyping each site of the plurality of sites to generate the genotyped list.

[0026] Some embodiments also include removing multi-nucleotide variant SNPs (MNVS) from the genotyped list.

[0027] Some embodiments also include encoding the genotyped list to obtain the first SIB, wherein each identifier is indicative that a site of the plurality of sites is: (i) homozygous to the reference, (ii) heterozygous to the reference, (iii) homozygous alternative to the reference, (iv) unexpected genotype, or (v) not determined; optionally, wherein the indicative identifier is 0, 1, 2, 3 or 4, respectively.

[0028] Some embodiments also include receiving the sequence information for the second nucleic acid sample at the plurality of sites, and generating the second SIB.

[0029] In some embodiments, step (c) comprises: (i) aligning the first SIB with the second SIB; (ii) comparing the same positions in each SIB with one another, excluding the same positions with an identifier in either SIB for unexpected genotype or for not determined; and (iii) determining a ratio for a number of matching identifiers at the same positions in each SIB to a number of the same positions compared, thereby generating the ratio of identity.

[0030] In some embodiments, the compared positions consist of positions for autosomal sites of the plurality of sites.

[0031] In some embodiments, the compared positions consist of positions with an identifier for homozygous to the reference, or for homozygous alternative to the reference.

[0032] Some embodiments also include detecting a kinship relationship between the nucleic acid samples, wherein each compared position has an identifier selected from: (i) aposition in the first SIB has an identifier for homozygous to the reference, and the same position in the second SIB has an identifier for homozygous to the reference or for heterozygous to the reference; (ii) a position in the first SIB has an identifier for heterozygous to the reference, and the same position in the second SIB has an identifier for homozygous to the reference, for heterozygous to the reference or for homozygous alternative to the reference; and (iii) a position in the first SIB has an identifier for homozygous alternative to the reference, and the same position in the second SIB has an identifier for heterozygous to the reference or for homozygous alternative to the reference.

[0033] Some embodiments also include detecting a contaminant in a nucleic acid sample, wherein a SIB of a nucleic acid sample comprising additional nucleic acids comprises an increased number of positions with an identifier for heterozygous positions compared to a SIB of the nucleic acid sample without the additional nucleic acids.

[0034] In some embodiments, the detecting comprises: (i) generating a variant allele frequency (VAF) at each position of the first SIB and the second SIB; (ii) binning each VAF in a bin of width 0.01 in a plurality of bins having a range from 0 to 1; (iii) classifying a VAF with a value between 0 and 0.1 or between 0.9 to 1 as homozygous, and a VAF with a value between 0.1 and 0.9 as heterozygous; (iv) generating a contamination ratio by dividing the number of VAF in the homozygous class by the number of VAF in the heterozygous class; wherein a contamination ratio greater than a threshold indicates a contaminated nucleic acid sample; optionally, wherein the threshold is 1.0.

[0035] In some embodiments, step (d) comprises providing an indication that the first and second nucleic acid samples are a match.

[0036] In some embodiments, at least 45%, 50%, 55%, or 60% of the positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another; optionally, wherein at least 45% of the positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another.

[0037] In some embodiments, at least 40, 45, 50, 55, 60, 65, 70, or 75 positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another; optionally, wherein at least 50 positions in each of the firstSIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another.

[0038] In some embodiments, the same positions of the first and second SIB are homozygous to a reference, heterozygous to the reference, or homozygous alternative to the reference.

[0039] In some embodiments, the ratio of identity is at least 0.90, 0.95, or 0.99; optionally, wherein the ratio of identity is at least 0.99.

[0040] Some embodiments also include providing an indication that the first or second nucleic acid sample comprises a Y chromosome, wherein at least 50% of sites of the plurality of sites located on the Y chromosome are identified.

[0041] Some embodiments of the methods and compositions provided herein include a system for comparing nucleic acid samples, comprising a processor configured with instructions for any one of the foregoing computer-implemented methods.

[0042] Some embodiments of the methods and compositions provided herein include a system for comparing nucleic acid samples, comprising: (a) a sequence identity barcode (SIB) generator module running on a processor and adapted to: (i) receive sequence information for a first nucleic acid sample at a plurality of sites, (ii) generate from the sequence information a first SIB for the first nucleic acid sample, wherein the first SIB comprises an ordered list of identifiers for the plurality of sites; and (b) a SIB comparator module adapted to: (i) generate a ratio of identity by comparing the first SIB and a second SIB, wherein the second SIB is generated from sequence information for a second nucleic acid sample, and (ii) provide an indication of a match between the first and second nucleic acid samples based on the comparison.

[0043] In some embodiments, the plurality of sites comprises: (i) a plurality of single nucleotide polymorphisms (SNPs); (ii) locations distributed throughout a genome; (iii) locations in exons; (iv) biallelic sites; (v) locations within ubiquitously expressed genes; (vi) lack allele specific expression; and / or (vii) locations comprising a minor allele frequency in African, South and East Asian and European populations in a range from 0.3 to 0.7.

[0044] In some embodiments, the sequence information is obtained by any one of the foregoing methods for confirming a nucleic acid sample identity.

[0045] In some embodiments, the generate from the sequence information a first SIB comprises receiving a genotyped list for the plurality of sites; optionally, wherein a file comprises the genotyped list, wherein the file comprises a format selected from a variable call file (VCF), binary alignment map (BAM), and a compressed reference-oriented alignment map (CRAM).

[0046] In some embodiments, receiving the genotyped list comprises: (i) aligning the sequence information with a reference; and (ii) genotyping each site of the plurality of sites to generate the genotyped list.

[0047] Some embodiments also include removing multi-nucleotide variant SNPs (MNVS) from the genotyped list.

[0048] Some embodiments also include encoding the genotyped list to obtain the first SIB, wherein each identifier is indicative that a site of the plurality of sites is: (i) homozygous to the reference, (ii) heterozygous to the reference, (iii) homozygous alternative to the reference, (iv) unexpected genotype, or (v) not determined; optionally, wherein the indicative identifier is 0, 1, 2, 3 or 4, respectively.

[0049] Some embodiments also include receiving the sequence information for the second nucleic acid sample at the plurality of sites, and generating the second SIB.

[0050] In some embodiments, the generate a ratio of identity comprises: (i) aligning the first SIB with the second SIB; (ii) comparing the same positions in each SIB with one another, excluding the same positions with an identifier in either SIB for unexpected genotype or for not determined; and (iii) determining a ratio for a number of matching identifiers at the same positions in each SIB to a number of the same positions compared, thereby generating the ratio of identity.

[0051] In some embodiments, the compared positions consist of positions for autosomal sites of the plurality of sites.

[0052] In some embodiments, the compared positions consist of positions with an identifier for homozygous to the reference, or for homozygous alternative to the reference.

[0053] Some embodiments also include detecting a kinship relationship between the nucleic acid samples, wherein each compared position has an identifier selected from: (i) a position in the first SIB has an identifier for homozygous to the reference, and the same position in the second SIB has an identifier for homozygous to the reference or forheterozygous to the reference; (ii) a position in the first SIB has an identifier for heterozygous to the reference, and the same position in the second SIB has an identifier for homozygous to the reference, for heterozygous to the reference or for homozygous alternative to the reference; and (iii) a position in the first SIB has an identifier for homozygous alternative to the reference, and the same position in the second SIB has an identifier for heterozygous to the reference or for homozygous alternative to the reference.

[0054] Some embodiments also include detecting a contaminant in a nucleic acid sample, wherein a SIB of a nucleic acid sample comprising additional nucleic acids comprises an increased number of positions with an identifier for heterozygous positions compared to a SIB of the nucleic acid sample without the additional nucleic acids.

[0055] In some embodiments, the detecting comprises: (i) generating a variant allele frequency (VAF) at each position of the first SIB and the second SIB; (ii) binning each VAF in a bin of width 0.01 in a plurality of bins having a range from 0 to 1; (iii) classifying a VAF with a value between 0 and 0.1 or between 0.9 to 1 as homozygous, and a VAF with a value between 0.1 and 0.9 as heterozygous; (iv) generating a contamination ratio by dividing the number of VAF in the homozygous class by the number of VAF in the heterozygous class; wherein a contamination ratio greater than a threshold indicates a contaminated nucleic acid sample; optionally, wherein the threshold is 1.0.

[0056] Some embodiments also include a module for providing an indication that the first and second nucleic acid samples are a match.

[0057] In some embodiments, at least 45%, 50%, 55%, or 60% of the positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another; optionally, wherein at least 45% of the positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another.

[0058] In some embodiments, at least 40, 45, 50, 55, 60, 65, 70, or 75 positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another; optionally, wherein at least 50 positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another.

[0059] In some embodiments, the same positions of the first and second SIB are homozygous to a reference, heterozygous to the reference, or homozygous alternative to the reference.

[0060] In some embodiments, the ratio of identity is at least 0.90, 0.95, or 0.99; optionally, wherein the ratio of identity is at least 0.99.

[0061] Some embodiments also include providing an indication that the first or second nucleic acid sample comprises a Y chromosome, wherein at least 50% of sites of the plurality of sites located on the Y chromosome are identified.

[0062] Some embodiments of the methods and compositions provided herein include a kit for confirming a nucleic acid sample identity, comprising: a plurality of pairs of primers, wherein the plurality of pairs of primers comprises at least 50 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 197-392. a DNA polymerase buffer; and a DNA polymerase. In some embodiments, the plurality of pairs of primers comprises at least 70 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392. In some embodiments, the plurality of pairs of primers comprises at least 90 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392. In some embodiments, the plurality of primers comprises pairs of primers comprising a sequence as set forth in each of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392. In some embodiments, the DNA polymerase buffer comprises Illumina TTM buffer. In some embodiments, the DNA polymerase comprises Thermo Phusion Hot Start II DNA Polymerase.

[0063] Some embodiments of the methods and compositions provided herein include a system for confirming a nucleic acid sample identity, comprising a processor configured to perform a method comprising: receiving sequence information for a nucleic acid sample at a plurality of sites, wherein the sites comprise 20 or more, 30 or more, 40 or more 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or each of the sites corresponding to positions: chr1:6693097, chr1:41204569, chr1:59147926, chr1:114515717, chr1:150808889, chr1:212911836, chr1:227069737, chr1:227071525, chr1:227935444, chr1:236716959, chr1:236719193, chr2:3392295, chr2:101624471, chr2:101638888, chr2:128939817, chr3:4403767, chr3:15737689, chr3:33434831, chr3:155481609,chr3:186509517, chr3:193374964, chr4:6600012, chr4:6606864, chr4:25419283, chr5:53815495, chr5:64881936, chr5:82834630, chr5:96503523, chr6:26598188, chr6:49403282, chr6:49425521, chr6:101166095, chr6:146755140, chr6:158517308, chr7:2577781, chr7:2578237, chr7:43916727, chr7:43917013, chr7:48004962, chr8:30973957, chr8:33356074, chr8:33369944, chr8:90995019, chr8:94935937, chr8:104427359, chr8:124448736, chr8:132982824, chr8:135612745, chr8:146067054, chr9:132400480, chr10:1046712, chr10:31138817, chr10:99219885, chr10:100219314, chr10:104572963, chr11:16133413, chr11:125763746, chr12:993930, chr12:62926398, chr13:28143229, chr13:39433606, chr14:20920250, chr14:31381351, chr15:34528948, chr15:75131959, chr16:624114, chr16:2028402, chr16:19680546, chr16:70303580, chr17:10614442, chr17:17168164, chr17:40714804, chr17:40716520, chr17:57290383, chr17:71196809, chr17:79514129, chr17:80008392, chr18:21413869, chr19:885818, chr19:1010691, chr19:36727365, chr19:41939297, chr19:50983930, chr19:58929052, chr20:6100088, chr20:61834695, chr21:47961711, chr22:32795641, chr22:32875190, chr22:50962208, chrY:2961349, chrY:7063980, chrY:2866789, chrY:2980414, chrY:7091532, chrY:13705715, chrY:14842155, or chrY:20592667 of hg 19 - GRCh37, and generating a sample identity barcode (SIB) comprising sequence information at each of the plurality of sites. Some embodiments also include comparing the SIBs of at least two nucleic acid samples of the one or more nucleic acid samples. In some embodiments, the sequence information comprises whether the nucleic acid sample is homozygous for a reference allele, homozygous for an alternative allele, or heterozygous at each site of the plurality of sites.

[0064] Some embodiments of the methods and compositions provided herein include a method of amplifying a human nucleic acid sample, comprising: contacting DNA or cDNA from a nucleic acid sample with a plurality of pairs of primers, a DNA polymerase buffer, and a DNA polymerase, wherein the DNA polymerase buffer comprises Illumina TTM buffer, and wherein the DNA polymerase comprises Thermo Phusion Hot Start II DNA Polymerase, and amplifying the DNA or cDNA at a plurality of sites.

[0065] In some embodiments, the plurality of sites comprises 20 or more, 30 or more, 40 or more 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or each of the sites corresponding to positions: chr1:6693097, chr1:41204569, chr1:59147926, chr1:114515717, chr1:150808889, chr1:212911836, chr1:227069737, chr1:227071525,chr1:227935444, chr1:236716959, chr1:236719193, chr2:3392295, chr2:101624471, chr2:101638888, chr2:128939817, chr3:4403767, chr3:15737689, chr3:33434831, chr3:155481609, chr3:186509517, chr3:193374964, chr4:6600012, chr4:6606864, chr4:25419283, chr5:53815495, chr5:64881936, chr5:82834630, chr5:96503523, chr6:26598188, chr6:49403282, chr6:49425521, chr6:101166095, chr6:146755140, chr6:158517308, chr7:2577781, chr7:2578237, chr7:43916727, chr7:43917013, chr7:48004962, chr8:30973957, chr8:33356074, chr8:33369944, chr8:90995019, chr8:94935937, chr8:104427359, chr8:124448736, chr8:132982824, chr8:135612745, chr8:146067054, chr9:132400480, chr10:1046712, chr10:31138817, chr10:99219885, chr10:100219314, chr10:104572963, chr11:16133413, chr11:125763746, chr12:993930, chr12:62926398, chr13:28143229, chr13:39433606, chr14:20920250, chr14:31381351, chr15:34528948, chr15:75131959, chr16:624114, chr16:2028402, chr16:19680546, chr16:70303580, chr17:10614442, chr17:17168164, chr17:40714804, chr17:40716520, chr17:57290383, chr17:71196809, chr17:79514129, chr17:80008392, chr18:21413869, chr19:885818, chr19:1010691, chr19:36727365, chr19:41939297, chr19:50983930, chr19:58929052, chr20:6100088, chr20:61834695, chr21:47961711, chr22:32795641, chr22:32875190, chr22:50962208, chrY:2961349, chrY:7063980, chrY:2866789, chrY:2980414, chrY:7091532, chrY:13705715, chrY:14842155, or chrY:20592667 of hg 19 - GRCh37.

[0066] In some embodiments, the plurality of pairs of primers comprises at least 10, at least 20, at least 30, at least 40, or at least 50 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392. In some embodiments, the plurality of pairs of primers comprises at least 70 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392. In some embodiments, the plurality of pairs of primers comprises at least 90 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197- 392. In some embodiments, the plurality of pairs of primers comprises pairs of primers comprising a sequence as set forth in each of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numerals correspond to similar, though perhaps not identical, components. For the sake of brevity, reference numerals or features having a previously described function may or may not be described in connection with other drawings in which they appear. In addition to the features described herein, additional features and variations will be readily apparent from the following descriptions of the drawings and exemplary embodiments. It is to be understood that these drawings depict typical embodiments, and are not intended to be limiting in scope.

[0068] FIG.1 schematically illustrates an embodiment of a workflow for methods of confirming a nucleic acid sample identity.

[0069] FIG.2 schematically illustrates an embodiment of amplification methods.

[0070] FIG. 3 schematically illustrates an embodiment of primer addition during amplification.

[0071] FIG.4A and FIG. 4B each depict a graph showing example TapeStation™ HSD1000 traces from Example 1 for high quality DNA and RNA control samples.

[0072] FIG.5 includes graphs illustrating control amplification yields for different enzyme and buffer combinations.

[0073] FIG. 6 includes graphs illustrating amplification yields with different enzymes and buffers and with control and with test oligos (IDPv1; primer oligos with sequences as set forth in SEQ ID NOs: 673-952).

[0074] FIG. 7 is a graph illustrating a comparison between IDPv1 amplification yields using PCR Amplification Enzyme (PAE) / PCR Amplification Buffer (PAB) versus Thermo PhusionTMHot Start II DNA Polymerase and buffer.

[0075] FIG. 8 includes graphs illustrating a comparison of amplification yields from different buffers.

[0076] FIG. 9 is a graph illustrating an overlay of results from PAB, TruSight Tumor Targeting Mix (TTM), and PAB*.

[0077] FIG. 10 includes graphs illustrating a comparison of amplification yields from different DNA polymerase enzymes.

[0078] FIG. 11 includes graphs which illustrate a comparison of amplification yields from different combinations of DNA polymerase enzyme and buffer.

[0079] FIG. 12 is a graph illustrating a comparison of yields from Thermo PhusionTMHot Start II DNA Polymerase and TTM buffer, compared to Thermo PhusionTMHot Start II DNA Polymerase and Thermo Buffer.

[0080] FIG. 13A and FIG. 13B schematically illustrate a target selection process (FIG.13A) and a multiplex PCR assay (FIG. 13B).

[0081] FIG. 14A illustrates an example pedigree. FIG. 14B depicts a Hankocompare output. FIG. 14C depicts a histogram of Hanko similarity ratios related to validation of Sample ID panel in the platinum genomes family.

[0082] FIG.15A and FIG.15B each depict a heat map for a Hankocompare output comparing tumor and normal samples related to the application of a sample ID panel to confirm tumor / normal pairs.

[0083] FIG. 16A and FIG. 16B each depict a Hankocompare output graph related to the validation of a sample ID panel for use in RNA. FIG. 16C and FIG. 16D each depict graphs for the frequency of successful hanko calls across different sample types.

[0084] FIG. 17A depicts a table related to a contamination detection feature. FIG. 17B, FIG. 17C, FIG. 17D each depict a graph related to a contamination detection feature.

[0085] FIG. 18 depicts an embodiment of a sequence identity barcode (SIB), such as a Hankoprint.

[0086] FIG. 19A depicts an alignment between two SIBs in which positions encoded as ‘4’ and ‘3’ are removed from the comparison. FIG. 19B depicts a comparison of SIBs between positions for autosomal sites and between positions identified as homozygous for autosomal sites, and determination of a ratio of identity.

[0087] FIG. 20 is a block diagram that schematically illustrates methods of comparing nucleic acid samples by generating and comparing SIBs of the nucleic acid samples.

[0088] FIG. 21A is a block diagram of an exemplary SIB generator / comparison system that may be used to perform the disclosed methods.

[0089] FIG. 21B is a block diagram of an exemplary computing device that may be used in connection with the exemplary SIB generator / comparison system of FIG.21A.DETAILED DESCRIPTION

[0090] The foregoing and other aspects of the present disclosure will now be described in more detail with respect to the description and methodologies provided herein. This description is not intended to be a detailed catalogue of all the ways in which the embodiments of the present disclosure may be implemented, or of all the features that may be added to the present disclosure. For example, features illustrated with respect to one embodiment may be incorporated into other embodiments, and features illustrated with respect to a particular embodiment may be deleted from that embodiment. In addition, numerous variations and additions to the various embodiments suggested herein, which do not depart from the instant disclosure, will be apparent to those skilled in the art in light of the instant detailed description, figures and claims. Hence, the following specification is intended to illustrate some particular embodiments, and not to exhaustively specify all permutations, combinations and variations thereof.

[0091] All patents, patent applications, and other publications, including all sequences disclosed within these references, referred to herein are expressly incorporated herein by reference, to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. All documents cited are, in relevant part, incorporated herein by reference in their entireties for the purposes indicated by the context of their citation herein. However, the citation of any document is not to be construed as an admission that it is prior art with respect to the present disclosure.

[0092] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.Overview

[0093] The present disclosure relates to methods, kits, and systems for comparing nucleic acids samples. Some embodiments include methods and systems for confirming the identity of a nucleic acid sample, and methods and kits for amplifying a nucleic acid.

[0094] Disclosed herein are methods for confirming a nucleic acid sample identity. In some embodiments, the methods comprise obtaining a nucleic acid sample from a human biological sample, determining nucleic acid sequence information for the nucleic acid sample at a plurality of sites, and based on the sequence information at each of the plurality of sites, confirming the nucleic acid sample identity. The plurality of sites may include at least 20 sites disclosed in TABLE 1. Each site in the plurality of sites may be a single nucleotide polymorphism (SNP) or any other one or more nucleotides at a predetermined site within the nucleic acids taken from the sample which are useful for identifying the identity of the sample. Other embodiments may include determining sequence information for the nucleic acid sample at a plurality of sites, wherein the plurality of sites comprises 30, 40, 50, 60, 70, 80, or 90 or more sites.

[0095] Further disclosed herein are kits for confirming a nucleic acid sample identity. The kits may include primers, DNA polymerase, and DNA polymerase buffer. The DNA polymerase buffer may be Illumina® TTM buffer, and the DNA polymerase may be Thermo PhusionTMHot Start II DNA Polymerase (Thermo Fisher Scientific, Carlsbad CA).

[0096] Further disclosed herein are systems for confirming a nucleic acid sample identity. The systems comprise a processor which is configured to receive sequence information for a nucleic acid sample at a plurality of sites, and generate a sample identity barcode (SIB) comprising sequence information at each of the plurality of sites.

[0097] The disclosed methods, kits, and systems may be useful in confirming the identity of a nucleic acid sample when it is possible that one or more samples have been inadvertently swapped during collection, testing, or data input processes. While sample swaps can sometimes be identified where multiple samples are processed from the same individual, the methods, kits, and systems of the present disclosure generate one or more data points for comparison when the sample enters the lab, creating the ability to identify when two samples have inadvertently been swapped without resort to an additional round of whole genome sequencing (WGS) or whole exome sequencing (WES), thereby saving time and money. Forexample, an aliquot of a sample from an individual may be taken at a first facility and a set of primer pairs may be used to amplify specific regions of the nucleic acids in the sample. The specific regions may be ones which are known to be highly polymorphic and thus have an increased likelihood of being different between various individuals. The nucleotide bases at each of the regions may be stored to a computer storage. The sample may then be shipped to a second facility for testing. At the second facility, another aliquot of the sample may be taken and the same set of primer pairs may be used to amplify the specific regions. If the sample at the first facility is the same as the sample at the second facility, then the specific regions should have the same nucleotide sequences. The second facility can access the stored data from the first facility and determine if the samples match or not.

[0098] The methods, kits, and systems of the present disclosure may have utility in other contexts as well. For example, for samples related to somatic cancer, a lab could confirm that tumor and normal samples are derived from the same individual, as the same individual will have the same sequences in the highly polymorphic regions. As another example, in multiomic studies the lab could confirm that the DNA and RNA samples are from the same individual. As a further example, sequencing labs could confirm that no processing issues had occurred from the plate that entered the lab. As a further example, relatedness, such as a match, of two samples could be confirmed. In some embodiments, the methods, kits and systems of the present disclosure enable rapid confirmation of nucleic acid sample identity, such as within 24 hours. In some embodiments, the methods, kits and systems of the present disclosure enable confirmation of nucleic acid sample identity that is more cost-effective than whole genome sequencing.

[0099] Further disclosed herein are methods and kits for amplifying a human nucleic acid sample. The method may include contacting DNA or cDNA from a nucleic acid sample with a plurality of pairs of primers, a DNA polymerase buffer, and a DNA polymerase, and amplifying the DNA or cDNA at a plurality of sites. The plurality of pairs of primers may include at least 20 primers comprising a sequence as set forth in any of SEQ ID NOs:01-196. In some embodiments, the plurality of pairs of primers may include at least 20 primers comprising a sequence as set forth in any of SEQ ID NOs: 197-392.

[0100] The methods, kits, and systems of the present disclosure may also more efficiently amplify a human nucleic acid sample. For example, the combination of Illumina®TTM buffer (Illumina Inc, San Diego CA) and Thermo PhusionTMHot Start II DNA Polymerase (Thermo Fisher Scientific, Carlsbad CA) may be unexpectedly effective in amplifying a nucleic acid sample compared to other combinations of DNA polymerase and buffer.

[0101] Some embodiments include methods and systems for comparing nucleic acid samples. In some such embodiments, a sample identity barcode (SIB) is generated from sequence information for a plurality of sites for a nucleic acid sample. The SIB of a nucleic acid sample is compared to the SIB of an additional nucleic acid sample to provide an indication that the samples are from the same individual. Advantageously, embodiments can readily identify a match between a normal sample and a tumor sample from an individual. In some such embodiments, the tumor sample may include multiple mutations including substitutions, deletions and / or insertions. Definitions

[0102] Although the following terms are believed to be well understood by one of skill in the art, the following definitions are set forth to facilitate understanding of the presently disclosed subject matter.

[0103] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques or substitutions of equivalent techniques that would be apparent to one of skill in the art.

[0104] As used herein, the terms “a” or “an” or “the” may refer to one or more than one. For example, “a” marker can mean one marker or a plurality of markers.

[0105] As used herein, the term “about,” when used in reference to a measurable value such as an amount of mass, dose, time, temperature, and the like, is meant to encompass variations of 20%, 10%, 5%, 1%, 0.5%, or even 0.1% of the specified amount.

[0106] As used herein, the term “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”).

[0107] Throughout this specification, unless the context requires otherwise, the words “comprise,” “comprises,” and “comprising” will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements.

[0108] As used herein, the term “consists essentially of” (and grammatical variants thereof), as applied to the compositions and methods of the present disclosure, means that the compositions / methods may contain additional components so long as the additional components do not materially alter the composition / method.

[0109] As used herein, the term “nucleic acid” refers to a polynucleotide sequence, or fragment thereof. A nucleic acid can comprise nucleotides. A nucleic acid can be exogenous or endogenous to a cell. A nucleic acid can exist in a cell-free environment. A nucleic acid can be a gene or fragment thereof. A nucleic acid can be DNA. A nucleic acid can be RNA. A nucleic acid can comprise one or more analogs (e.g., altered backbone, sugar, or nucleobase). Some non-limiting examples of analogs include: 5-bromouracil, peptide nucleic acid, xeno nucleic acid, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to the sugar), thiol containing nucleotides, biotin linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queuosine, and wyosine. "Nucleic acid", "polynucleotide, "target polynucleotide", and "target nucleic acid" can be used interchangeably.

[0110] As used herein the term “chromosome” refers to the heredity-bearing gene carrier of a living cell, which is derived from chromatin strands comprising DNA and protein components (especially histones). The conventional internationally recognized individual human genome chromosome numbering system is employed herein.

[0111] A “genome” refers to the complete genetic information of an organism or virus, expressed in nucleic acid sequences.

[0112] As used herein, the term “reference genome” or “reference sequence” refers to any particular known genome sequence, whether partial or complete, of any organism or virus which may be used to reference identified sequences from a subject. For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. In various embodiments,the reference sequence is significantly larger than the reads that are aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 105 times larger, or at least about 106 times larger, or at least about 107 times larger. In one example, the reference sequence is that of a full-length genome. Such sequences may be referred to as genomic reference sequences. For example, the reference sequence can be a reference human genome sequence, such as hg19 (for example, available at GenBank assembly accession GCA_000001405.1) or hg38 (for example, available at GenBank assembly accession GCA_000001405.15). In another example, the reference sequence is limited to a specific human chromosome such as chromosome 13. In some embodiments, a reference Y chromosome is the Y chromosome sequence from human genome version hg19. Such sequences may be referred to as chromosome reference sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes, sub-chromosomal regions (such as strands), etc., of any species. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a particular individual.

[0113] The term “nucleic acid sample” herein may refer to a sample, typically derived from one or more biological fluids, cells, tissues, organs, or organisms, comprising a nucleic acid or a mixture of nucleic acids comprising at least one nucleic acid sequence that is to be screened for copy number variation. In certain embodiments the nucleic acid sample comprises at least one nucleic acid sequence whose copy number is suspected of having undergone variation. Such samples may include, but are not limited to sputum / oral fluid, amniotic fluid, blood, a blood fraction, or fine needle biopsy samples (such as surgical biopsy, fine needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, and the like. Although the sample is often taken from a human subject (such as a patient), the sample may be from any mammal, including, but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc. The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivationof interfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employed with respect to the sample, such pretreatment methods are typically such that the nucleic acid(s) of interest remain in the test sample, sometimes at a concentration proportional to that in an untreated test sample (such as namely, a sample that is not subjected to any such pretreatment method(s)). Such “treated” or “processed” samples are still considered to be biological “test” samples with respect to the methods described herein. A “nucleic acid sample” may also include nucleic acid sequence information stored in a memory, and which was originally obtained from a source such as one or more biological fluids, cells, tissues, organs, or organisms.

[0114] As used herein, “high fidelity DNA polymerase” refers to a DNA polymerase with an low error rate, such as an error rate of about 1 in 1,000,000 nucleotides, or less. The error rate for a high fidelity DNA polymerase typically ranges from 10-6to 10-8errors per base pair replicated. High fidelity DNA polymerases may be designed to have low error rates, making them suitable for applications requiring accurate replication of DNA sequences, such as PCR (Polymerase Chain Reaction), DNA sequencing, and cloning. High fidelity DNA polymerases may achieve this low error rate through various mechanisms, including proofreading activity and selection for correct nucleotide incorporation. Examples of high fidelity DNA polymerases include Pfu (Pyrococcus furiosus) polymerase and Phusion™ DNA polymerase.

[0115] As used herein, “sample identity barcode” (SIB) refers to a composite of sequence information for a nucleic acid sample related to a plurality of sites, such as polymorphic sites. For example, a SIB may include genetic information about whether the sample is homozygous for a reference allele, homozygous for an alternative allele, or heterozygous, at each site of the plurality of sites. An embodiment of a SIB includes a Hankoprint.

[0116] As used herein, “adapter” or “adapter sequence” refers to a nucleic acid sequence which may facilitate sequencing, primer binding, and / or indexing. For example, an adapter sequence may be added to at least one end of a nucleic acid fragment.

[0117] As used herein, “splint,” “splint sequence” or “splint nucleotide sequence” refers to a sequence of nucleotides on an end of a nucleic acid molecule that are complementary to a second nucleic acid molecule. For example, SEQ ID NOs: 197-392 and 673-950 comprisesplints that are complementary to a portion of standard Illumina® adapters (with index sequences). Methods for Confirming a Nucleic Acid Sample Identity

[0118] In one aspect, disclosed herein are methods for confirming a nucleic acid sample identity. Nucleic acid sample

[0119] In some embodiments, the methods include obtaining a nucleic acid sample from a human biological sample. In some embodiments, the human biological sample comprises a blood sample, a tumor sample, a saliva sample, a hair sample, a semen sample, a skin sample, a muscle sample, or an organ sample. In some embodiments, the human biological sample comprises Formalin-Fixed Paraffin-Embedded (FFPE) tissue.

[0120] A nucleic acid sample may be obtained from a human biological sample using methods known to those of skill in the art. In some embodiments, the methods include steps of preparing a nucleic acid sample from a human biological sample.

[0121] In some embodiments of the method, a nucleic acid sample is obtained which was previously prepared from a human biological sample.

[0122] In some embodiments, the nucleic acid sample comprises DNA. In some embodiments, the nucleic acid sample comprises RNA. In some embodiments, an amount of nucleic acid molecules in the nucleic acid sample is less than about 100 ng, 50 ng, 30 ng, 20 ng, 10 ng, 5 ng, or 1 ng.

[0123] In some embodiments, the nucleic acid sample comprises RNA. In some embodiments, the method further comprises obtaining cDNA based on the RNA from the nucleic acid sample. cDNA may be obtained from RNA using methods known to those of skill in the art. For example, RNA may be reverse transcribed into cDNA using a reverse transcriptase. Plurality of Sites

[0124] In some embodiments, the methods include determining sequence information for the nucleic acid sample at a plurality of sites. In some embodiments, the plurality of sites comprises a plurality of polymorphic sites. In some embodiments, the plurality of sites comprises at least 10, at least 20, at least 30, at least 40, at least 50, at least60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, about 240 sites, or a range constructed from any of the aforementioned values.

[0125] In some embodiments, the plurality of sites comprises about 90 autosomal sites and about 8 Y chromosomal sites. In some embodiments, the plurality of sites comprises sites falling within exons present in both DNA and RNA. In some embodiments, the plurality of sites comprises sites which are expressed in multiple sample types, for example, present in genes expressed for blood and different tissues. For example, in some embodiments, each of the plurality of sites are located in regions which are expressed in one or more of blood, breast, or lung tissues. In some embodiments, the plurality of sites (or at least 90% of the plurality of sites) are constitutively expressed. In some embodiments, the plurality of sites is scattered throughout the human genome. For example, in some embodiments, the plurality of sites includes sites on each of the autosomal chromosomes, on the Y chromosomes, and / or different arms of several chromosomes.

[0126] In some embodiments, the plurality of sites comprises 10 or more, 20 or more, 30 or more, 40 or more 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or each of the sites described in TABLE 1. In TABLE 1, Hg 19 - GRCh37 refers to Genome Reference Consortium Human Build 37 (GRCh37), available at ncbi.nlm.nih.gov / datasets / genome / GCF_000001405.13 / . Grch38 refers to Genome Reference Consortium Human Build 38 (GRCh38), available at ncbi.nlm.nih.gov / datasets / genome / GCF_000001405.26 / . The sites of TABLE 1 are found in exomes, such that RNA or DNA sample may be used. TABLE 1 Name Hg 19 - chrGRCh37GRCh RefAlt Chr. (rsID)38 LabelcoordinateAlleleAllele Arm rs3174820 chr1 6693097 6633037 SNP A G p rs1057925 chr1 41204569 40738897 SNP C T p rs12139511 chr1 59147926 58682254 SNP T C p rs2358996 chr1 114515717 113973095 SNP G A pName Hg 19 - chrGRCh37GRC RefAlt Chr.(rsID)h38 LabelcoordinateAlleleAllele Arm rs2228099 chr1 150808889 150836413 SNP C G q rs15702 chr1 212911836 212738494 SNP T C q rs6759 chr1 227069737 226882036 SNP C T q rs1046240 chr1 227071525 226883824 SNP C T q rs2236359 chr1 227935444 227747743 SNP A G q rs2275685 chr1 236716959 236553659 SNP C T q rs1885533 chr1 236719193 236555893 SNP A G q rs11686212 chr2 3392295 3388524 SNP A G p rs746924 chr2 101624471 101008009 SNP T C q rs3739014 chr2 101638888 101022426 SNP A G q rs1699 chr2 128939817 128182243 SNP G A q rs2819561 chr3 4403767 4362083 SNP A G p rs2470548 chr3 15737689 15696182 SNP G A p rs2293250 chr3 33434831 33393339 SNP G A p rs7653384 chr3 155481609 155763820 SNP G A q rs187868 chr3 186509517 186791728 SNP G A q rs9851685 chr3 193374964 193657175 SNP T C q rs2301790 chr4 6600012 6598285 SNP A G p rs2301788 chr4 6606864 6605137 SNP A G p rs9174 chr4 25419283 25417661 SNP T C p rs2548612 chr5 53815495 54519665 SNP A C q rs27141 chr5 64881936 65586109 SNP A G q rs309557 chr5 82834630 83538811 SNP A G q rs160632 chr5 96503523 97167819 SNP C T q rs3800303 chr6 26598188 26597960 SNP A G p rs8589 chr6 49403282 49435569 SNP T C p rs2229384 chr6 49425521 49457808 SNP C T pName Hg 19 - chrGRCh37GRCh RefAlt Chr.(rsID)38 LabelcoordinateAlleleAllele Arm rs41288423 chr6 101166095 100718219 SNP G A q rs2942 chr6 146755140 146434004 SNP C T q rs2502601 chr6 158517308 158096276 SNP A G q rs1043291 chr7 2577781 2538147 SNP T C p rs4719552 chr7 2578237 2538603 SNP T C p rs2232108 chr7 43916727 43877128 SNP T G p rs2232105 chr7 43917013 43877414 SNP G A p rs1056663 chr7 48004962 47965365 SNP C T p rs1800392 chr8 30973957 31116441 SNP G T p rs6468171 chr8 33356074 33498556 SNP A G p rs2304748 chr8 33369944 33512426 SNP T C p rs1063045 chr8 90995019 89982791 SNP C T q rs4735258 chr8 94935937 93923709 SNP C T q rs3134295 chr8 104427359 103415131 SNP A C q rs7014678 chr8 124448736 123436496 SNP A G q rs1051221 chr8 132982824 131970577 SNP A G q rs894344 chr8 135612745 134600502 SNP A G q rs1735169 chr8 146067054 144841669 SNP G A q rs3739851 chr9 132400480 129638201 SNP G A q rs2306409 chr10 1046712 1000772 SNP G A p rs10160116 chr10 31138817 30849888 SNP G A p rs2152092 chr10 99219885 97460128 SNP G A q rs10883099 chr10 100219314 98459557 SNP A G q rs284860 chr10 104572963 102813206 SNP T C q rs4617548 chr11 16133413 16111867 SNP C T p rs3088241 chr11 125763746 125893851 SNP C G q rs7300444 chr12 993930 884764 SNP A G pName Hg 19 - chrGRCh37GRC RefAlt Chr.(rsID)h38 LabelcoordinateAlleleAllele Arm rs7957417 chr12 62926398 62532618 SNP G A q rs8002697 chr13 28143229 27569092 SNP A G q rs9532292 chr13 39433606 38859469 SNP A G q rs2275007 chr14 20920250 20452091 SNP T C q rs2273171 chr14 31381351 30912145 SNP T C q rs4577050 chr15 34528948 34236747 SNP C T q rs936227 chr15 75131959 74839618 SNP A G q rs2071979 chr16 624114 574114 SNP A G p rs8460 chr16 2028402 1978401 SNP A C p rs957676 chr16 19680546 19669224 SNP T C p rs2070203 chr16 70303580 70269677 SNP G A q rs406446 chr17 10614442 10711125 SNP A G p rs3182911 chr17 17168164 17264850 SNP G A p rs615942 chr17 40714804 42562786 SNP C A q rs598126 chr17 40716520 42564502 SNP A G q rs3744383 chr17 57290383 59213022 SNP G A q rs1026128 chr17 71196809 73200670 SNP A G q rs11552304 chr17 79514129 81547103 SNP G A q rs9916764 chr17 80008392 82050516 SNP G T q rs9962023 chr18 21413869 23833905 SNP A G q rs1060442 chr19 885818 885818 SNP A G p rs2240159 chr19 1010691 1010692 SNP A G p rs2070132 chr19 36727365 36236463 SNP G A q rs1043413 chr19 41939297 41433392 SNP C G q rs10409679 chr19 50983930 50480673 SNP C T q rs3764535 chr19 58929052 58417685 SNP G A q rs10373 chr20 6100088 6119441 SNP G T pName Hg 19 - chrGRCh37GR RefAlt Chr. (rsID)Ch38 LabelcoordinateAlleleAllele Arm rs6122103 chr20 61834695 63203343 SNP G A q rs2070435 chr21 47961711 46541798 SNP G A q rs5749426 chr22 32795641 32399654 SNP C T q rs11107 chr22 32875190 32479203 SNP G A q rs12148 chr22 50962208 50523779 SNP T G q ZFY chrY 2961349 2961448 region * TBL1Y chrY 7063980 7064064 region *rs35284970 chrY 2866789 2866876 region *rs770402645 chrY 2980414 2980509 region *TBL1Y-2 chrY 7091532 7091591 region *TMSB4Y chrY 13705715 13705809 region *rs3863116 chrY 14842155 14842253 region *EIF1AY chrY 20592667 20592766 region * *Presence or absence of a product indicates maleness

[0127] In some embodiments, the plurality of sites comprises 10 or more, 20 or more, 30 or more, 40 or more 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 110 or more, 120 or more, 130 or more, or each of the sites described in TABLE 2. The sites in TABLE 2 are found in intronic regions. TABLE 2 Name Chromosome GRCh38 coordinate rs10399787 chr1 6972569 rs1763926 chr1 69199017 rs3907134 chr1 105919040 rs4144397 chr1 115363836 rs891700 chr1 239718626 rs3014561 chr1 241414220 rs824448 chr2 19490683Name Chromosome GRCh38 coordinate rs6705138 chr2 25364982 rs848507 chr2 36440530 rs10196335 chr2 144571821 rs2389557 chr2 169354867 rs7591784 chr2 207637006 rs369751 chr3 32469805 rs34094301 chr3 59485848 rs4677108 chr3 72183280 rs4857284 chr3 94272660 rs1449445 chr3 146065025 rs2048961 chr4 18859767 rs35036278 chr4 30939693 rs279844 chr4 46327638 rs17025362 chr4 133490770 rs11733742 chr4 152754504 rs272735 chr5 33401364 rs55847630 chr5 33626006 rs2194236 chr5 54315149 rs164561 chr5 69044501 rs154214 chr5 107747099 rs7268 chr5 140332965 rs919320 chr5 165396474 rs6939421 chr6 6411712 rs9380610 chr6 36838109 rs1126476 chr6 39080715 rs2121881 chr6 127588832 rs7752169 chr6 135906420 rs6557440 chr6 155605318Name Chromosome GRCh38 coordinate rs12699957 chr7 18026199 rs4722897 chr7 29292617 rs1404719 chr7 46775939 rs2470961 chr7 104794766 rs7792485 chr7 144746078 rs11771560 chr7 156813341 rs4732838 chr8 28184670 rs921802 chr8 38576419 rs750961 chr8 58688336 rs6472862 chr8 74547684 rs4735258 chr8 93923709rs1494338chr9 14730126rs655258chr9 23542709rs2383812chr9 29417326rs1381532 chr9 97428498rs10114192chr9 108859784rs1999263chr9 112447393rs7903683 chr10 10162817 rs723211 chr10 12720304 rs3780962 chr10 17151347 rs10786349 chr10 97456309 rs10883099 chr10 98459557rs3123221chr10 131380359rs1498553 chr11 5687798rs7926887chr11 21545633rs1892953chr11 76561498rs7125361chr11 115209322Name Chromosome GRCh38 coordinaters687928chr11 122125553rs555172chr11 128696263rs11614639 chr12 1774523 rs4766232 chr12 4320852 rs2269355 chr12 6836750 rs2111980 chr12 105934476 rs36158849 chr12 122088381 rs10773760 chr12 130277151 rs1359215 chr13 37303798 rs9532292 chr13 38859469 rs9570447 chr13 61077561 rs7339162 chr13 65495593 rs4465496 chr13 72484553 rs354439 chr13 106286062 rs4982648 chr14 22602464 rs11156787 chr14 33372238 rs398745 chr14 36066975 rs7160372 chr14 76350033 rs17717270 chr14 89803535 rs1257263 chr14 99407920 rs8026015 chr15 36533049 rs4924176 chr15 37596648 rs3098168 chr15 50458853 rs28408562 chr15 60624880 rs8033162 chr15 92351052 rs2342747 chr16 5818699 rs893196 chr16 58997812Name Chromosome GRCh38 coordinate rs2070203 chr16 70269677 rs430046 chr16 77983154 rs9909684 chr17 13596937 rs8064562 chr17 14109705 rs9893096 chr17 34236812 rs4353533 chr17 62821857 rs8078417 chr17 82504059 rs11081037 chr18 3353014 rs2124299 chr18 9661018 rs9951171 chr18 9749882 rs1378329 chr18 25676256 rs17064378 chr18 58119127 rs2547071 chr19 8965061 rs9304650 chr19 29094082 rs10412617 chr19 31311644 rs3817 chr19 43586043 rs10373_1 chr20 6119441 rs6110405 chr20 14713943 rs1535199 chr20 21837670 rs2072973 chr20 45615203 rs2296241 chr20 54169680 rs2224176 chr20 55706196 rs1663501 chr21 26846988 rs914165 chr21 41044003 rs4148973 chr21 42903480 rs9637290 chr21 42981014 rs4675 chr22 20787012Name Chromosome GRCh38 coordinate rs987640 chr22 33163522 rs5768657 chr22 48478107 rs9616586 chr22 49370695 rs2404797 chrX 8827337 rs7888207 chrX 11898336 rs766117 chrX 43956960 rs4826609 chrX 54739480 rs5941047 chrX 92176386 rs916208 chrX 128807273 rs5975695 chrX 136186310 rs614511 chrX 150369567 rs11575897_M575_O2b chrY 2787139 rs35284970 chrY 2866813 rs16981311_M445_E chrY 7097174 rs2032680_M181_E1b1a6 chrY 12904625 rs2032658_M207_R chrY 13470103 rs17316592_M471_O3a chrY 16448125 rs17315772_M398_IJ chrY 16926422 rs9786197 chrY 18956002 rs17250163_M400_IJ chrY 19063884 rs16980406_M589_E1b1a8a1 chrY 19484172 rs3900_M9_K-R chrY 19568371 rs2032652_M89_F-R chrY 19755427

[0128] In some embodiments, the plurality of sites comprises 20 or more, 30 or more, 40 or more 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or each of the sites corresponding to positions: chr1:6693097, chr1:41204569, chr1:59147926, chr1:114515717, chr1:150808889, chr1:212911836, chr1:227069737, chr1:227071525,chr1:227935444, chr1:236716959, chr1:236719193, chr2:3392295, chr2:101624471, chr2:101638888, chr2:128939817, chr3:4403767, chr3:15737689, chr3:33434831, chr3:155481609, chr3:186509517, chr3:193374964, chr4:6600012, chr4:6606864, chr4:25419283, chr5:53815495, chr5:64881936, chr5:82834630, chr5:96503523, chr6:26598188, chr6:49403282, chr6:49425521, chr6:101166095, chr6:146755140, chr6:158517308, chr7:2577781, chr7:2578237, chr7:43916727, chr7:43917013, chr7:48004962, chr8:30973957, chr8:33356074, chr8:33369944, chr8:90995019, chr8:94935937, chr8:104427359, chr8:124448736, chr8:132982824, chr8:135612745, chr8:146067054, chr9:132400480, chr10:1046712, chr10:31138817, chr10:99219885, chr10:100219314, chr10:104572963, chr11:16133413, chr11:125763746, chr12:993930, chr12:62926398, chr13:28143229, chr13:39433606, chr14:20920250, chr14:31381351, chr15:34528948, chr15:75131959, chr16:624114, chr16:2028402, chr16:19680546, chr16:70303580, chr17:10614442, chr17:17168164, chr17:40714804, chr17:40716520, chr17:57290383, chr17:71196809, chr17:79514129, chr17:80008392, chr18:21413869, chr19:885818, chr19:1010691, chr19:36727365, chr19:41939297, chr19:50983930, chr19:58929052, chr20:6100088, chr20:61834695, chr21:47961711, chr22:32795641, chr22:32875190, chr22:50962208, chrY:2961349, chrY:7063980, chrY:2866789, chrY:2980414, chrY:7091532, chrY:13705715, chrY:14842155, or chrY:20592667 of hg 19 - GRCh37, or a corresponding position in GRCh38 as described in TABLE 1.

[0129] In some embodiments, the plurality of sites comprises 1 or more sites, 10 or more sites, 20 or more sites, 30 or more sites, 40 or more sites, 50 or more sites, 60 or more sites, 70 or more sites, 80 or more sites, 90 or more sites, or each of the sites listed in TABLE 1. In some embodiments, the plurality of sites comprises 1 or more sites, 10 or more sites, 20 or more sites, 30 or more sites, 40 or more sites, 50 or more sites, 60 or more sites, 70 or more sites, 80 or more sites, 90 or more sites, 100 or more sites, 110 or more sites, 120 or more sites, 130 or more sites, or each of the sites listed in TABLE 2. In some embodiments, the plurality of sites comprises 1 or more sites, 10 or more sites, 20 or more sites, 30 or more sites, 40 or more sites, 50 or more sites, 60 or more sites, 70 or more sites, 80 or more sites, 90 or more sites, 100 or more sites, 110 or more sites, 120 or more sites, 130 or more sites, 140 or more sites, 150 or more sites, 160 or more sites, 170 or more sites, 180 or more sites, 190 or moresites, 200 or more sites, 210 or more sites, 220 or more sites, 230 or more sites, or each of the sites listed in TABLE 1 and TABLE 2.

[0130] In some embodiments, the method includes determining whether the nucleic acid sample is homozygous for a reference allele, homozygous for an alternative allele, or heterozygous at each site of the plurality of sites. Plurality of Primers

[0131] In some embodiments, determining sequence information comprises contacting DNA or cDNA from the nucleic acid sample with a plurality of pairs of primers designed to bind adjacent the sites to be amplified. For example, a primer pair may bind adjacent to one or more of the sites listed in TABLE 1 or TABLE 2. TABLE 3 lists certain sequences that may be included in certain embodiments provided herein. TABLE 3 SEQ ID NOs Characteristic 01-196 Primer pairs for sites listed in TABLE 1 197-392 Sequences of SEQ ID NOs 01-196 with added splint sequences 393-670 Primer pairs for sites listed in TABLE 2 671-672 Bacterial contamination 673-950 Sequences of SEQ ID NOs 393-670 with added splint sequences 951-952 Bacterial contamination 953-954 Splint sequences Odd numbered SEQ ID NOs include a “forward primer (5 to 3 )”;Even numbered SEQ ID NOs include a “reverse primer (5 to 3 )”.

[0132] In some embodiments, the plurality of pairs of primers comprises at least 1, at least 10 at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, or at least 190 primers which comprise a sequence as set forth in any of SEQ ID NOs: 1-196. In some embodiments, the plurality of pairs of primers comprises at least 1, at least 10 at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, or at least 190 primers which comprise a sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least95%, at least 95%, at least 97%, at least 98%, or at least 99% sequence similarity to any of the sequences as set forth in any of SEQ ID NOs: 01-196. In some embodiments, the plurality of pairs of primers comprises 196 primers (98 pairs) as set forth in any of SEQ ID NOs: 01-196.

[0133] In some embodiments, the plurality of pairs of primers comprises at least 1, at least 10 at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, or at least 90 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196. In some embodiments, the plurality of pairs of primers comprises primers comprising a sequence as set forth in each of SEQ ID NOs: 01-196. SEQ ID NOs: 01-196 recite sequences for forward and reverse primers corresponding to the plurality of sites disclosed in TABLE 1.

[0134] In some embodiments, the plurality of pairs of primers comprises at least 1, at least 10 at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90 at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, or at least 190 primers which comprise a sequence as set forth in any of SEQ ID NOs: 197-392. In some embodiments, the plurality of pairs of primers comprises at least 1, at least 10 at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90 at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, or at least 190 primers which comprise a sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 95%, at least 97%, at least 98%, or at least 99% sequence similarity to any of the sequences as set forth in any of SEQ ID NOs: 197-392. In some embodiments, the plurality of pairs of primers comprises at least 1, at least 10 at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, or at least 90 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 197-392. In some embodiments, the plurality of pairs of primers comprises primers comprising a sequence as set forth in each of SEQ ID NOs: 197-392. SEQ ID NOs: 197-392 recite sequences for forward and reverse primers corresponding to the plurality of sites disclosed in TABLE 1.

[0135] In some embodiments, the plurality of primers comprises splint nucleotide sequences. For example, SEQ ID NOs: 197-392 are forward and reverse primers including the same nucleotide sequence as SEQ ID NOs: 01-196, respectively, however, SEQ ID NOs: 197- 392 additionally include splint nucleotide sequences.

[0136] In some embodiments, the plurality of pairs of primers comprises at least 1, at least 10 at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90 at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, or at least 240 primers which comprise a sequence as set forth in any of SEQ ID NOs: 393- 672. In some embodiments, the plurality of pairs of primers comprises at least 1, at least 10 at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90 at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, or at least 240 primers which comprise a sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 95%, at least 97%, at least 98%, or at least 99% sequence similarity to any of the sequences as set forth in any of SEQ ID NOs: 393-672. In some embodiments, the plurality of pairs of primers comprises at least 1, at least 10 at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, or at least 130 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 393-672. In some embodiments, the plurality of pairs of primers comprises primers comprising a sequence as set forth in each of SEQ ID NOs: 393-672. SEQ ID NOs: 393-670 recite sequences for forward and reverse primers to the plurality of sites disclosed in TABLE 2. SEQ ID NOs: 671-672 recite primers for detection of bacterial contamination in a nucleic acid sample.

[0137] In some embodiments, the plurality of pairs of primers comprises at least 1, at least 10 at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90 at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, or at least 240 primers which comprise a sequence as set forth in any of SEQ ID NOs: 673- 952. In some embodiments, the plurality of pairs of primers comprises at least 1, at least 10 at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90 at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, or at least 240 primers which comprise a sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 95%, at least 97%, at least 98%, or at least 99%sequence similarity to any of the sequences as set forth in any of SEQ ID NOs: 673-952. In some embodiments, the plurality of pairs of primers comprises at least 1, at least 10 at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, or at least 130 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 673-952. In some embodiments, the plurality of pairs of primers comprises primers comprising a sequence as set forth in each of SEQ ID NOs: 673-952. SEQ ID NOs: 673-952 are forward and reverse primers including the same nucleotide sequence as SEQ ID NOs: 393-672, respectively, however, SEQ ID NOs: 673-952 additionally include splint nucleotide sequences. SEQ ID NOs: 673-950 recite sequences for forward and reverse primers to the plurality of sites disclosed in TABLE 2. SEQ ID NOs: 951-952 recite primers for detection of bacterial contamination in a nucleic acid sample. DNA Polymerase and Buffer

[0138] In some embodiments, determining sequence information comprises contacting DNA or cDNA from the nucleic acid sample with a DNA polymerase. DNA polymerases for use in amplification reactions are known to those of skill in the art. In some embodiments, the DNA polymerase comprises a high-fidelity DNA polymerase. In some embodiments, the DNA polymerase comprises a Thermo Phusion™ Hot Start II DNA Polymerase.

[0139] In some embodiments, determining sequence information comprises contacting DNA or cDNA from the nucleic acid sample with a DNA polymerase buffer. DNA polymerase buffers for use in amplification reactions are known to those of skill in the art. In some embodiments, the DNA polymerase buffer comprises Illumina®TTM buffer. Amplification and Library Preparation

[0140] In some embodiments, determining sequence information comprises amplifying the DNA or cDNA at the plurality of sites to generate amplified nucleic acid fragments. Amplification may be accomplished by any method known to those of skill in the art, for example, polymerase chain reaction (PCR). In some embodiments, amplifying DNA or cDNA from each of the one or more nucleic acid samples comprises multiplex PCR. For example, the plurality of pairs of primers may hybridize to the DNA or cDNA of the nucleic acid sample, and DNA polymerase may be used to extend along the primers, creating amplified PCR product (i.e., amplified nucleic acid fragments, amplicons).

[0141] The plurality of primers may be as described above. For example, a plurality of pairs of primers (for example, comprising sequences as set forth in SEQ ID NOs: 1-196, 197-392, 393-672, or 673-952) may hybridize to the DNA or cDNA from the nucleic acid sample, and DNA polymerase may extend along the primers in PCR to create amplified nucleic acid sequences.

[0142] In some embodiments, the method includes generating amplified nucleic acid fragments comprising a splint sequence. In some embodiments, the splint sequence enables use with oligonucleotides comprising an index sequence, for example, index primers. For example, in some embodiments, the splint sequences include a sequence which allows hybridization to index primers. In some embodiments, the index primers comprise Illumina® UD index adapters.

[0143] In an alternative embodiment, an index sequence in included in the sequences of the plurality of primers. Alternatively, addition of an index sequence may be accomplished using methods known to those of skill in the art, including additional amplification reactions, ligation, tagmentation, etc.

[0144] In some embodiments, a splint sequence is added via amplification. For example, a plurality of pairs of primers with sequences as set forth in SEQ ID NOs: 197-392, which comprise splint sequences, may hybridize to the PCR product and be extended, to generate amplified nucleic acid fragments comprising a splint sequence. The amplification step may be accomplished using amplification methods known to those of skill in the art, such as with the DNA polymerases and buffers described above. For example, SEQ ID NOs: 197-392 and 673-952 include ‘splints’ on them that the standard Illumina® adapters (with index sequences) can bind to and amplify to generate a product that may be sequenced. In some embodiments, the plurality of pairs of primers comprising a splint sequence (for example, SEQ ID NOs: 197-392) and adapter primers (such as Illumina® UD index adapters) may be included in one container (such as a tube or well) during amplification reactions.

[0145] In some embodiments, the amplification and library preparation steps may advantageously be performed in multiplex format and / or within a single tube.

[0146] Before sequencing, the amplified nucleic acids may be prepared (such as size selection and / or clean up) using methods known to those of skill in the art. In some embodiments, clean-up is performed using SPRI beads.Sequencing

[0147] In some embodiments, determining sequence information comprises sequencing the amplified nucleic acid fragments. In some embodiments, determining sequence information is performed in multiplex format. In some embodiments, sequencing comprises high-throughput sequencing, such as by methods known to those of skill in the art.

[0148] In some embodiments, determining sequence information (i.e., sequencing the amplified nucleic acid fragments) comprises hybridizing the amplified nucleic acid fragments to a nucleic acid adapter (such as on a flow cell) comprising a sequence complementary to an adapter sequence on amplified nucleic acid fragments. Confirming Sample Identity Based on Sequence Information

[0149] In some embodiments, the methods include confirming the nucleic acid sample identity based on the sequence information at each of the plurality of sites. For example, the sequence information at each of the plurality of sites for the sample may be compared to independently-obtained sequence information at each of the plurality of sites, such as via a second sequencing process, or from a second sample.

[0150] For example, sample identity may be confirmed by comparing sequence information at each of the plurality of sites which was determined via a panel comprising primers targeted to those sites, with sequence information at each of the plurality of sites obtained via whole genome and transcriptome sequencing (WGTS).

[0151] As a further example, sample identity may be confirmed by comparing sequence information at each of the plurality of sites which was determined via a panel comprising primers targeted to those sites, with sequence information from a second sample which is presumed to have come from the same individual, whether from a similar panel or from WGTS.

[0152] In some embodiments, the method further includes confirming a nucleic acid sample sequence identity for a second nucleic acid sample from a second human biological sample. In some embodiments, the human biological sample and the second human biological sample are from an individual (i.e., the same individual). In some embodiments, the second human biological sample is from a second, different individual.

[0153] For example, the two human biological samples may comprise a tumor sample and a normal sample from an individual. In some embodiments, where a sampleidentity of a tumor sample is determined as compared to sample identity of a normal sample, sites of the plurality of sites which are heterozygous in the normal sample are excluded from the comparison. This may be advantageous where tumor cells may include copy number variation which may skew a sequencing result at a heterozygous site. Kits for Confirming a Nucleic Acid Sample Identity

[0154] In another aspect, disclosed herein are kits. In some embodiments, the kits comprise kits for confirming a nucleic acid sample identity. In some embodiments, the kit comprises a plurality of pairs of primers, a DNA polymerase buffer, and a DNA polymerase. The plurality of pairs of primers, DNA polymerase buffer, and DNA polymerase may be as described above with respect to the methods of confirming sample identity. In some embodiments, the plurality of pairs of primers includes at least 20 primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196. In some embodiments, the plurality of pairs of primers includes at least 20 primers comprising a sequence as set forth in any of SEQ ID NOs: 197-392. In some embodiments, the DNA polymerase buffer comprises an Illumina® TTM buffer. In some embodiments, the DNA polymerase comprises a Thermo Phusion™ Hot Start II DNA Polymerase. Systems and Methods

[0155] In another aspect, disclosed herein are systems for confirming a nucleic acid sample identity. In some embodiments, the system comprises a processor configured to perform a method.

[0156] In some embodiments, the method includes receiving sequence information for a nucleic acid sample at a plurality of sites. In some embodiments, the plurality of sites may be as described above with respect to the methods of confirming sample identity.

[0157] In some embodiments, the received sequence information includes whether the nucleic acid sample is homozygous for a reference allele, homozygous for an alternative allele, or heterozygous at each site of the plurality of sites.

[0158] In some embodiments, the method comprises, for each of the one or more nucleic acid samples, generating a sample identity barcode (SIB) comprising sequence information at each of the plurality of sites. In some embodiments, the SIB includesinformation related to whether the nucleic acid sample is homozygous for a reference allele, homozygous for an alternative allele, or heterozygous at each site of the plurality of sites.

[0159] The SIB may be displayed in a user-friendly graphical user interface to facilitate sample identity confirmation.

[0160] In some embodiments, the method comprises comparing the SIBs of at least two nucleic acid samples of the one or more nucleic acid samples. In some embodiments, the system is configured to indicate whether the SIBs match, based on the comparison. Systems and Methods For Comparing Nucleic acid Samples

[0161] Some embodiments provided herein relate to systems and methods for the generation and comparison of sample identity barcodes (SIB) for nucleic acid samples to detect a relatedness, such as a match, between the nucleic acid samples. Some embodiments can be used to confirm identity of nucleic acid samples, for example, that nucleic acids are from the same individual, to detect nucleic acid sample swaps, to detect mislabeled metadata, and to detect related nucleic acid samples. Some embodiments may be integrated into nucleic acid sequencing and sequence analysis workflows. For example, to check inputs before sequencing, to check outputs before delivery, to match DNA-RNA sample pairs, to match tumor-normal sample pairs.

[0162] Some embodiments include the generation of a SIB. In some embodiments, a SIB can include a Hankoprint. Example input nucleic acid samples include DNA, RNA, nucleic acid samples from normal tissue, nucleic acid samples from tumor. Example input sequence information to generate a SIB, such as a Hankoprint, can be in a file format such as a FASTQ format, a compressed reference-oriented alignment map (CRAM) format , a binary alignment map (BAM) format, variant call format (VCF), and a genomic VCF (gVCF).

[0163] In some embodiments, generation of a SIB, such as a Hankoprint, can include (i) receiving an input file, such as a FASTQ file; (ii) aligning the sequence data with a reference, such as reference data, such as a reference genome; (iii) genotyping the sequence information, such as calling / identifying a single nucleotide polymorphisms (SNP) variant in the sequence information as the same as the reference, a known variant of the reference, an unknown variant of the reference, or not determinable to generate a genotyped list; (iv)removing multi-nucleotide variant SNPs (mnvs) from the genotyped list; and (v) encoding the genotyped list to generate the SIB, such as a Hankoprint.

[0164] In some embodiments, the encoding can include scoring each item of the genotyped list such that (i) homozygous to the reference = ‘0’, (ii) heterozygous to the reference = ‘1’, (iii) homozygous alternative to the reference = ‘2’, (iv) unexpected genotype = ‘3’, and (v) not determined = ‘4’. TABLE 4 lists an embodiment for the encoding in which an allele of a SNP is either “1”, “0”, or “.”. For example, a genotype for a homozygous SNP would be “0” and “0”, or “1” and “1”. TABLE 4 Genotype Encoding Genotype Encoding0 0 . / . 40 / 0 0 . / 1 40|0 0 1|. 40 / 1 1 .|1 41 / 0 1 1 / . 40|1 1 0 / . 41|0 1 0|. 41 1 . / 0 41 / 1 2 .|0 41|1 2

[0165] An embodiment of a format for a Hankoprint can include a uniform resource name (URN); a custom namespace, such as ‘hanko’; a version identifier; and an ordered list of the encoded genotyped sequence information, such as SNPs, including encoded positions for sequence information from autosomal sites and allosomal sites. An embodiment of a Hankoprint is depicted in FIG.18.

[0166] Some embodiments include comparing two SIBs with one another. In some such embodiments, the SIBs can be aligned with one another. Positions in either SIB whichhave been encoded as ‘not determined’ or ‘unexpected variant’ can be removed from the comparison. A ratio of the number of matching positions to the total number of comparable positions can be determined to obtain a ratio of identity. The comparison can be made for positions in the SIBs for autosomal sites. In some embodiments, a threshold to determine a matching pair of SIBs can include a ratio of at least 0.99. In some embodiments, a threshold to determine a matching pair of SIBs can include a comparison between at least 50 positions.

[0167] In some embodiments, a comparison is determined between positions that have been identified as homozygous at autosomal sites. Such a comparison is particularly useful for SIBs generated from tumor / normal nucleic acid samples. Without being bound to any one theory, tumor samples are less likely to have mutations at homozygous sites than heterozygous sites. Thus, a comparison between positions in SIBs that have been identified as homozygous at autosomal sites can more readily detect a match between a tumor / normal nucleic acid samples from an individual. An example embodiment of an alignment between two SIBs is depicted in FIG. 19A. In FIG. 19A, positions encoded as ‘4’ and ‘3’ are removed from the comparison. FIG. 19B depicts a comparison between positions for autosomal sites, and between positions identified as homozygous for autosomal sites.

[0168] Some embodiments include a computer-implemented method of comparing nucleic acid samples, comprising: (a) receiving sequence information for a first nucleic acid sample at a plurality of sites; (b) generating from the sequence information a first SIB for the first nucleic acid sample, wherein the first SIB comprises an ordered list of identifiers for the plurality of sites; (c) generating a ratio of identity by comparing the first SIB and a second SIB, wherein the second SIB is generated from sequence information for a second nucleic acid sample; and (d) providing an indication of a match between the first and second nucleic acid samples based on the comparison.

[0169] In some embodiments, the plurality of sites comprises: (i) a plurality of SNPs; (ii) locations distributed throughout a genome; (iii) locations in exons; (iv) biallelic sites; (v) locations within ubiquitously expressed genes; (vi) lack allele specific expression; and / or (vii) locations comprising a minor allele frequency in African, South and East Asian and European populations in a range from 0.3 to 0.7. In some embodiments, the sequence information is obtained by any one of the methods provided herein for confirming a nucleic acid sample identity. For example, in some embodiments, the plurality of sites comprises 10or more, 20 or more, 30 or more, 40 or more 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or each of the sites listed in TABLE 1. In some embodiments the sequence information can be obtained by a primer or primer pair comprising a nucleotide sequence of any one of SEQ ID NOs: 01-196.

[0170] In some embodiments, step (b) comprises obtaining a genotyped list for the plurality of sites. In some embodiments, a file comprises the genotyped list, wherein the file comprises a format selected from a variable call file (VCF), binary alignment map (BAM), and a compressed reference-oriented alignment map (CRAM). In some embodiments, obtaining the genotyped list comprises: (i) aligning the sequence information with a reference, such as a reference genome; and (ii) genotyping each site of the plurality of sites to generate the genotyped list. Some embodiments also include removing multi-nucleotide variant SNPs (MNVS) from the genotyped list.

[0171] Some embodiments also include encoding the genotyped list to obtain the first SIB, wherein each identifier is indicative that a site of the plurality of sites is: (i) homozygous to the reference, (ii) heterozygous to the reference, (iii) homozygous alternative to the reference, (iv) unexpected genotype, or (v) not determined. For example, the indicative identifier can be 0, 1, 2, 3 or 4, respectively.

[0172] Some embodiments also include receiving the sequence information for the second nucleic acid sample at the plurality of sites, and generating the second SIB.

[0173] In some embodiments, step (c) comprises: (i) aligning the first SIB with the second SIB; (ii) comparing the same positions in each SIB with one another, excluding the same positions with an identifier in either SIB for unexpected genotype or for not determined; and (iii) determining a ratio for a number of matching identifiers at the same positions in each SIB to a number of the same positions compared, thereby generating the ratio of identity.

[0174] In some embodiments, the compared positions consist of positions for autosomal sites of the plurality of sites.

[0175] In some embodiments, the compared positions consist of positions with an identifier for homozygous to the reference, or for homozygous alternative to the reference.

[0176] Some embodiments also include detecting a kinship, such as a parent / child, relationship between the nucleic acid samples, wherein each compared position has an identifier selected from: (i) a position in the first SIB has an identifier for homozygous to thereference, and the same position in the second SIB has an identifier for homozygous to the reference or for heterozygous to the reference; (ii) a position in the first SIB has an identifier for heterozygous to the reference, and the same position in the second SIB has an identifier for homozygous to the reference, for heterozygous to the reference or for homozygous alternative to the reference; and (iii) a position in the first SIB has an identifier for homozygous alternative to the reference, and the same position in the second SIB has an identifier for heterozygous to the reference or for homozygous alternative to the reference.

[0177] Some embodiments also include detecting a contaminant in a nucleic acid sample, wherein a SIB of a nucleic acid sample comprising additional nucleic acids comprises an increased number of positions with an identifier for heterozygous positions compared to a SIB of the nucleic acid sample without the additional nucleic acids. In some embodiments, the detecting comprises: (i) generating a variant allele frequency (VAF) at each position of the first SIB and the second SIB; (ii) binning each VAF in a bin of width 0.01 in a plurality of bins having a range from 0 to 1; (iii) classifying a VAF with a value between 0 and 0.1 or between 0.9 to 1 as homozygous, and a VAF with a value between 0.1 and 0.9 as heterozygous; and (iv) generating a contamination ratio by dividing the number of VAF in the homozygous class by the number of VAF in the heterozygous class; wherein a contamination ratio greater than a threshold indicates a contaminated nucleic acid sample. In some embodiments, a contamination ratio greater than a threshold indicates a contaminated nucleic acid sample. In some embodiments, a contamination ratio greater than a threshold indicates a contaminated nucleic acid sample. In some embodiments, a contamination ratio greater than 1.0 indicates a contaminated nucleic acid sample.

[0178] In some embodiments, step (d) comprises providing an indication that the first and second nucleic acid samples are a match. In some embodiments, at least 45%, 50%, 55%, or 60% of the positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another. In some embodiments, at least 45% of the positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another. In some embodiments, the positions that are comparable to one another are the same positions of the first and second SIB are homozygous to a reference, heterozygous to the reference, or homozygous alternative to the reference.

[0179] In some embodiments, at least 40, 45, 50, 55, 60, 65, 70, or 75 positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another. In some embodiments, at least 50 positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another. In some embodiments, the positions that are comparable to one another are the same positions of the first and second SIB are homozygous to a reference, heterozygous to the reference, or homozygous alternative to the reference.

[0180] In some embodiments, the ratio of identity is at least 0.90, 0.95, or 0.99. In some embodiments, the ratio of identity is at least 0.99.

[0181] Some embodiments also include providing an indication that the first or second nucleic acid sample comprises a Y chromosome, wherein at least 50% of sites of the plurality of sites located on the Y chromosome are identified.

[0182] Some embodiments provided herein include a system for comparing nucleic acid samples, comprising a processor configured with instructions for any one of the computer- implemented methods provided herein.

[0183] Some embodiments provided herein include a system for comparing nucleic acid samples, comprising: (a) a sequence identity barcode (SIB) generator module running on a processor and adapted to: (i) receive sequence information for a first nucleic acid sample at a plurality of sites, (ii) generate from the sequence information a first SIB for the first nucleic acid sample, wherein the first SIB comprises an ordered list of identifiers for the plurality of sites; and (b) a SIB comparator module adapted to: (i) generate a ratio of identity by comparing the first SIB and a second SIB, wherein the second SIB is generated from sequence information for a second nucleic acid sample, and (ii) provide an indication of a match between the first and second nucleic acid samples based on the comparison.

[0184] In some embodiments, the plurality of sites comprises: (i) a plurality of single nucleotide polymorphisms (SNPs); (ii) locations distributed throughout a genome; (iii) locations in exons; (iv) biallelic sites; (v) locations within ubiquitously expressed genes; (vi) lack allele specific expression; and / or (vii) locations comprising a minor allele frequency in African, South and East Asian and European populations in a range from 0.3 to 0.7. In some embodiments, the sequence information is obtained by any one of the methods provided herein for confirming a nucleic acid sample identity. For example, in some embodiments, the pluralityof sites comprises 10 or more, 20 or more, 30 or more, 40 or more 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or each of the sites listed in TABLE 1. In some embodiments the sequence information can be obtained by a primer or primer pair comprising a nucleotide sequence of any one of SEQ ID NOs: 01-196.

[0185] In some embodiments, the generate from the sequence information a first SIB comprises receiving a genotyped list for the plurality of sites. In some embodiments, a file comprises the genotyped list, wherein the file comprises a format selected from a variable call file (VCF), binary alignment map (BAM), and a compressed reference-oriented alignment map (CRAM). In some embodiments, receiving the genotyped list comprises: (i) aligning the sequence information with a reference; and (ii) genotyping each site of the plurality of sites to generate the genotyped list. Some embodiments also include removing multi-nucleotide variant SNPs (MNVS) from the genotyped list.

[0186] Some embodiments also include encoding the genotyped list to obtain the first SIB, wherein each identifier is indicative that a site of the plurality of sites is: (i) homozygous to the reference, (ii) heterozygous to the reference, (iii) homozygous alternative to the reference, (iv) unexpected genotype, or (v) not determined. For example, the indicative identifier can be 0, 1, 2, 3 or 4, respectively.

[0187] Some embodiments also include receiving the sequence information for the second nucleic acid sample at the plurality of sites, and generating the second SIB.

[0188] In some embodiments, the generate a ratio of identity comprises: (i) aligning the first SIB with the second SIB; (ii) comparing the same positions in each SIB with one another, excluding the same positions with an identifier in either SIB for unexpected genotype or for not determined; and (iii) determining a ratio for a number of matching identifiers at the same positions in each SIB to a number of the same positions compared, thereby generating the ratio of identity.

[0189] In some embodiments, the compared positions consist of positions for autosomal sites of the plurality of sites.

[0190] In some embodiments, the compared positions consist of positions with an identifier for homozygous to the reference, or for homozygous alternative to the reference.

[0191] Some embodiments also include detecting a kinship, such as parent / child, relationship between the nucleic acid samples, wherein each compared position has anidentifier selected from: (i) a position in the first SIB has an identifier for homozygous to the reference, and the same position in the second SIB has an identifier for homozygous to the reference or for heterozygous to the reference; (ii) a position in the first SIB has an identifier for heterozygous to the reference, and the same position in the second SIB has an identifier for homozygous to the reference, for heterozygous to the reference or for homozygous alternative to the reference; and (iii) a position in the first SIB has an identifier for homozygous alternative to the reference, and the same position in the second SIB has an identifier for heterozygous to the reference or for homozygous alternative to the reference.

[0192] Some embodiments also include detecting a contaminant in a nucleic acid sample, wherein a SIB of a nucleic acid sample comprising additional nucleic acids comprises an increased number of positions with an identifier for heterozygous positions compared to a SIB of the nucleic acid sample without the additional nucleic acids. In some embodiments, the detecting comprises: (i) generating a variant allele frequency (VAF) at each position of the first SIB and the second SIB; (ii) binning each VAF in a bin of width 0.01 in a plurality of bins having a range from 0 to 1; (iii) classifying a VAF with a value between 0 and 0.1 or between 0.9 to 1 as homozygous, and a VAF with a value between 0.1 and 0.9 as heterozygous; (iv) generating a contamination ratio by dividing the number of VAF in the homozygous class by the number of VAF in the heterozygous class; wherein a contamination ratio greater than a threshold indicates a contaminated nucleic acid sample. In some embodiments, a contamination ratio greater than 1.0 indicates a contaminated nucleic acid sample.

[0193] Some embodiments also include a module for providing an indication that the first and second nucleic acid samples are a match. In some embodiments, at least 45%, 50%, 55%, or 60% of the positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another. In some embodiments, at least 45% of the positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another. In some embodiments, the positions that are comparable to one another are the same positions of the first and second SIB are homozygous to a reference, heterozygous to the reference, or homozygous alternative to the reference.

[0194] In some embodiments, at least 40, 45, 50, 55, 60, 65, 70, or 75 positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are availablefor comparison with one another. In some embodiments, at least 50 positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another. In some embodiments, the positions that are comparable to one another are the same positions of the first and second SIB are homozygous to a reference, heterozygous to the reference, or homozygous alternative to the reference.

[0195] In some embodiments, the ratio of identity is at least 0.90, 0.95, or 0.99. In some embodiments, the ratio of identity is at least 0.99.

[0196] Some embodiments also include providing an indication that the first or second nucleic acid sample comprises a Y chromosome, wherein at least 50% of sites of the plurality of sites located on the Y chromosome are identified.

[0197] FIG. 20 is a block diagram that schematically illustrates methods of comparing nucleic acid samples by generating and comparing SIBs of the nucleic acid samples. In some embodiments, the method 200 is implemented on a computer. The method 200 may be embodied in a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives, of a computing system. For example, the server device 3102 shown in FIGS.21A and 21B and described in greater detail below can execute a set of executable program instructions to implement the method 200. When the method 200 is initiated, the executable program instructions can be loaded into a memory, such as RAM, and executed by one or more processors of a server device 3102. Although the method 200 is described with respect to the server device 3102 shown in FIG. 21B, the description is illustrative only and is not intended to be limiting. In some embodiments, the method 200 or portions thereof may be performed serially or in parallel by multiple computing systems.

[0198] As shown in FIG. 20, the method 200 for comparing nucleic acid samples may start from block 201, wherein sequence reads for a plurality of sites are received. Next, the method 200 may proceed to block 202, wherein sequence information is aligned and compared to a reference, a genotyped list is obtained, and the genotype list is encoded to obtain the first SIB. The method 200 may then proceed to block 203, wherein the first SIB is compared to a second SIB obtained from a second nucleic acid sample. Next, the method 200 may proceed to block 204, wherein an indication of a match between the first and second nucleic acids is provided. For example, an indication that the first and second nucleic acids are from the same individual.

[0199] FIG. 21A illustrates a diagram of an environment in which a system for comparing nucleic acid samples, such as a SIB generator / comparison system, can operate in accordance with one or more implementations. The following paragraphs describe the SIB generator / comparison system with respect to illustrative figures that portray example implementations and embodiments. For example, FIG.21A illustrates a schematic diagram of a computing system 3000 in which SIB generator / comparison system 3106 operates in accordance with one or more implementations. As illustrated, the computing system 3000 includes one or more server device(s) 3102 connected to a user client device 3108, a local device 3118, and a sequencing device 3114 via a network 3112. The network 3112 can comprise any suitable network over which computing devices can communicate.

[0200] As shown in FIG. 21A, the computing system 3000 includes the server device(s) 3102. In various implementations, the server device(s) 3102 may generate, receive, analyze, store, and transmit digital data, such as data for nucleobase calls, sequenced nucleic- acid polymers, SIBs, and / or ratios of identity. In some implementations, the server device(s) 3102 receive various data from the sequencing device 3114, such as data from a sample genome and / or sequence reads. The server device(s) 3102 may also communicate with the user client device 3108. In particular, the server device(s) 3102 can send data for sequence reads, direct nucleobase calls, nucleobase calls, sequencing metrics, SIBs, and / or ratios of identity to the user client device 3108.

[0201] As shown, the server device(s) 3102 includes a sequencing application 3110. In general, the sequencing application 3110 analyzes the data (such as call data) received from the sequencing device 3114 or elsewhere to determine nucleobase sequences for nucleic- acid polymers. For example, the sequencing application 3110 can receive raw data from the sequencing device 3114 and determine a nucleobase sequence for a sample genome or a nucleic-acid segment. In some implementations, the sequencing application 3110 determines the sequences of nucleobases in DNA and / or RNA segments or oligonucleotides.

[0202] As also shown, the sequencing application 3110 includes the SIB generator / comparison system 3106. As described below, the SIB generator / comparison system 3106 can generate a SIB from sequence information for a nucleic acid sample, compare the SIB with an additional SIB generated from an additional nucleic acid sample, and generate aratio of identity for the comparison between SIBs to determine a match between the two nucleic acid samples.

[0203] Moreover, while the SIB generator / comparison system 3106 is described being implemented on the server device(s) 3102, as part of the sequencing application 3110, in some implementations, the SIB generator / comparison system 3106 is implemented by (such as located entirely or in part) on the user client device 3108, the sequencing device 3114, and / or the local device 3118. As mentioned, in some implementations, the SIB generator / comparison system 3106 is implemented by one or more other components of the computing system 3000, such as the sequencing device 3114. In particular, the SIB generator / comparison system 3106 can be implemented in a variety of different ways across the server device(s) 3102, the network 3112, the user client device 3108, the local device 3118, and the sequencing device 3114.

[0204] As further shown in FIG.21A, the computing system 3000 includes the user client device 3108. In various implementations, the user client device 3108 can generate, store, receive, and send digital data. In particular, the user client device 3108 can receive the data from the sequencing device 3114. As further illustrated, the user client device 3108 includes a sequencing application 3110. The sequencing application 3110 may be a web application or a native application stored and executed on the user client device 3108 (e.g., a mobile application, desktop application, or web application). The sequencing application 3110 can receive data from the sequencing application 3110 and / or SIB generator / comparison system 3106. For example, the user client device 3108 can receive variant call files, alignment files, SIB and / or ratio of identity data from the sequencing application 3110.

[0205] The sequencing application 3110 can also include instructions that (when executed) cause the user client device 3108 to receive data from SIB generator / comparison system 3106 and present data from the sequencing device 3114 and / or the server device(s) 3102. Furthermore, the sequencing application 3110 can instruct the user client device 3108 to display an indication of a match between nucleic acid samples.

[0206] As further shown in FIG. 21A, the computing system 3000 includes the sequencing device 3114. In various implementations, the sequencing device 3114 can sequence a genomic sample or other nucleic-acid polymer. For example, the sequencing device 3114 analyzes nucleic-acid segments or oligonucleotides extracted from nucleic acid samples to generate data either directly or indirectly on the sequencing device 3114. More particularly,the sequencing device 3114 receives and analyzes, within nucleotide-sample slides (such as flow cells), nucleic-acid sequences extracted from nucleic acid samples. In one or more implementations, the sequencing device 3114 utilizes SBS to sequence a genomic sample or other nucleic-acid polymers. In addition to, or in the alternative to communicating across the network 3112, in some implementations, the sequencing device 3114 bypasses the network 3112 and communicates directly with the user client device 3108.

[0207] As further depicted in FIG. 21A, in some implementations, the server device(s) 3102 includes a distributed collection of servers, where the server device(s) 3102 include several server devices distributed across the network 3112 and located in the same or different physical locations. For instance, the server device(s) 3102 can be implemented, in whole or in part, on the local device 3118. To illustrate, the local device 3118 may implement the sequencing application 3110 and / or the SIB generator / comparison system 3106. Further, the server device(s) 3102 and / or the local device 3118 can include a content server, an application server, a communication server, a web-hosting server, or another type of server.

[0208] The user client device 3108 illustrated in FIG. 21A can include various types of client devices. For example, in some implementations, the user client device 3108 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In various implementations, the user client device 3108 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones.

[0209] Though FIG.21A illustrates the components of the computing system 3000 communicating via the network 3112, in certain implementations, the components of computing system 3000 can also communicate directly with each other, bypassing the network 3112. For instance, in some implementations, the user client device 3108 communicates directly with the sequencing device 3114. Additionally, in some implementations, the user client device 3108 communicates directly with the SIB generator / comparison system 3106 and / or the server device(s) 3102. In some implementations, the user client device 3108 communicates directly with the local device 3118. Moreover, the SIB generator / comparison system 3106 can access one or more databases housed on or accessed by the server device(s) 3102 or elsewhere in the computing system 3000.

[0210] FIG. 21B is a block diagram of an exemplary server device 3102 that may be used in connection with the illustrative sequencing system 3000 of FIG. 21A. The serverdevice 3102 may be configured to generate a SIB from sequencing date for a plurality of sites for a nucleic acid sample, and to compare SIBs with one another. The general architecture of the server device 3102 depicted in FIG. 21B includes an arrangement of computer hardware and software components. The server device 3102 may include many more (or fewer) elements than those shown in FIG. 21B. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the server device 3102 includes a processing unit 310, a network interface 320, a computer readable medium drive 330, an input / output device interface 340, a display 350, and an input device 360, all of which may communicate with one another by way of a communication bus. The network interface 320 may provide connectivity to one or more networks or computing systems. The processing unit 310 may thus receive information and instructions from other computing systems or services via a network. The processing unit 310 may also communicate to and from memory 370 and further provide output information for an optional display 350 via the input / output device interface 340. The input / output device interface 340 may also accept input from the optional input device 360, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.

[0211] The memory 370 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 310 executes in order to implement one or more embodiments. The memory 370 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer-readable media. The memory 370 may store an operating system 372 that provides computer program instructions for use by the processing unit 310 in the general administration and operation of the server device 3102. The memory 370 may store a reference genome 373, such as for use by the sequencing application 3110. The memory 370 may further include computer program instructions and other information for implementing aspects of the present disclosure.

[0212] For example, in one embodiment, the memory 370 includes a sequencing application 3110, which may include a SIB generator / comparison system 3106. The SIB generator / comparison system 3106 can perform the methods disclosed herein. In addition, memory 370 may include or communicate with the data store 390 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results(including intermediate results) of detecting a relatedness, such as a match between nucleic acids samples.

[0213] In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis features and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genome data, or other types of biological data may be mediated via a central hub that stores and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data as well as distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates modification or annotation of sequence data by users. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand or on-line.

[0214] In some embodiments, software written to perform the methods as described herein is stored in some form of computer readable medium, such as memory, CD- ROM, DVD-ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system and the like.

[0215] In some embodiments, the methods may be written in any of various suitable programming languages, for example compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages could be script languages, such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java or Python. In some embodiments, the method may be an independent application with data input and data display modules. Alternatively, the method may be a computer software product and may include classes wherein distributed objects comprise applications including computational methods as described herein.

[0216] In some embodiments, the methods may be incorporated into pre-existing data analysis software, such as that found on sequencing instruments. Software comprising computer implemented methods as described herein are installed either onto a computer system directly, or are indirectly held on a computer readable medium and loaded as needed onto a computer system. Further, the methods may be located on computers that are remote to where the data is being produced, such as software found on servers and the like that are maintainedin another location relative to where the data is being produced, such as that provided by a third party service provider.

[0217] An assay instrument, desktop computer, laptop computer, or server which may contain a processor in operational communication with accessible memory comprising instructions for implementation of systems and methods. In some embodiments, a desktop computer or a laptop computer is in operational communication with one or more computer readable storage media or devices and / or outputting devices. An assay instrument, desktop computer and a laptop computer may operate under a number of different computer based operational languages, such as those utilized by Apple based computer systems or PC based computer systems. An assay instrument, desktop and / or laptop computers and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results and monitoring experimental progress. In some embodiments, an outputting device may be a graphic user interface such as a computer monitor or a computer screen, a printer, a hand-held device such as a personal digital assistant (i.e., PDA, Blackberry, iPhone), a tablet computer (such as iPAD), a hard drive, a server, a memory stick, a flash drive and the like.

[0218] A computer readable storage device or medium may be any device such as a server, a mainframe, a supercomputer, a magnetic tape system and the like. In some embodiments, a storage device may be located onsite in a location proximate to the assay instrument, for example adjacent to or in close proximity to, an assay instrument. For example, a storage device may be located in the same room, in the same building, in an adjacent building, on the same floor in a building, on different floors in a building, etc. in relation to the assay instrument. In some embodiments, a storage device may be located off-site, or distal, to the assay instrument. For example, a storage device may be located in a different part of a city, in a different city, in a different state, in a different country, etc. relative to the assay instrument. In embodiments where a storage device is located distal to the assay instrument, communication between the assay instrument and one or more of a desktop, laptop, or server is typically via Internet connection, either wireless or by a network cable through an access point. In some embodiments, a storage device may be maintained and managed by the individual or entity directly associated with an assay instrument, whereas in other embodiments a storage device may be maintained and managed by a third party, typically at a distal locationto the individual or entity associated with an assay instrument. In embodiments as described herein, an outputting device may be any device for visualizing data.

[0219] An assay instrument, desktop, laptop and / or server system may be used itself to store and / or retrieve computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. One or more of an assay instrument, desktop, laptop and / or server may comprise one or more computer readable storage media for storing and / or retrieving software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. Computer readable storage media may include, but is not limited to, one or more of a hard drive, a SSD hard drive, a CD-ROM drive, a DVD-ROM drive, a floppy disk, a tape, a flash memory stick or card, and the like. Further, a network including the Internet may be the computer readable storage media. In some embodiments, computer readable storage media refers to computational resource storage accessible by a computer network via the Internet or a company network offered by a service provider rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.

[0220] In some embodiments, computer readable storage media for storing and / or retrieving computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like, is operated and maintained by a service provider in operational communication with an assay instrument, desktop, laptop and / or server system via an Internet connection or network connection.

[0221] In some embodiments, a hardware platform for providing a computational environment comprises a processor (i.e., CPU) wherein processor time and memory layout such as random access memory (i.e., RAM) are systems considerations. For example, smaller computer systems offer inexpensive, fast processors and large memory and storage capabilities. In some embodiments, graphics processing units (GPUs) can be used. In some embodiments, hardware platforms for performing computational methods as described herein comprise one or more computer systems with one or more processors. In some embodiments, smaller computer are clustered together to yield a supercomputer network.

[0222] In some embodiments, computational methods as described herein are carried out on a collection of inter- or intra-connected computer systems (i.e., grid technology) which may run a variety of operating systems in a coordinated manner. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available through United Devices are exemplary of the coordination of multiple stand-alone computer systems for the purpose dealing with large amounts of data. These systems may offer Perl interfaces to submit, monitor and manage large sequence analysis jobs on a cluster in serial or parallel configurations. Methods of Amplifying a Human Nucleic Acid Sample

[0223] In another aspect, disclosed herein are methods of amplifying a human nucleic acid sample.

[0224] In some embodiments, the method includes contacting DNA or cDNA from a nucleic acid sample with a plurality of pairs of primers, a DNA polymerase buffer, and a DNA polymerase. The plurality of pairs of primers, DNA polymerase buffer, and DNA polymerase may be as described above with respect to the methods of confirming sample identity. In some embodiments, the plurality of pairs of primers includes at least 20 primers comprising a sequence as set forth in any of SEQ ID Nos: 1-196. In some embodiments, the plurality of pairs of primers includes at least 20 primers comprising a sequence as set forth in any of SEQ ID Nos: 197-392. In some embodiments, the DNA polymerase buffer comprises Illumina® TTM buffer. In some embodiments, the DNA polymerase comprises Thermo Phusion™ Hot Start II DNA Polymerase.

[0225] In some embodiments, the method includes amplifying the DNA or cDNA at a plurality of sites. The plurality of sites may be as described above with respect to the methods of confirming sample identity. Amplification may proceed as described above with respect to the methods of confirming sample identity. Kits for Amplifying a Nucleic Acid sample

[0226] In another aspect, disclosed herein are kits. In some embodiments, the kits comprise kits for amplifying a nucleic acid sample. In some embodiments, the nucleic acid sample is a human nucleic acid sample. In some embodiments, the kit comprises a plurality ofpairs of primers, a DNA polymerase buffer, and a DNA polymerase. The plurality of pairs of primers, DNA polymerase buffer, and DNA polymerase may be as described above with respect to the methods of confirming sample identity. In some embodiments, the plurality of pairs of primers includes at least 20 primers comprising a sequence as set forth in any of SEQ ID Nos: 1-196. In some embodiments, the plurality of pairs of primers includes at least 20 primers comprising a sequence as set forth in any of SEQ ID Nos: 197-392. In some embodiments, the DNA polymerase buffer comprises Illumina® TTM buffer. In some embodiments, the DNA polymerase comprises Thermo Phusion™ Hot Start II DNA Polymerase. EXAMPLES

[0227] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not in any way intended to limit the scope of the present disclosure. Those in the art will appreciate that many other embodiments also fall within the scope of the disclosure, as it is described herein above and in the claims. Example 1:

[0228] In the following example, sample identity was confirmed for samples including RNA and samples including DNA. FIG.1 illustrates the workflow, where the overall workflow is as follows. The process starts with a sample and reverse transcriptase being incubated in a 5 μl reaction for 13 minutes. The process then includes a step of adding in Polymerase Chain Reaction (PCR) reagents to initiate a PCR process for approximately 2 hours. The process then moves to a solid-phase reversible immobilization (SPRI) process which includes pooling with SPRI beads two times. The process then moves to a sequencing the sequences isolated from the SPRI step. Additional details are provided below.

[0229] RNA samples: RNA reverse Transcription Reaction.

[0230] 1.7 μl of total RNA was added to a new 96 well PCR plate. (Example RNA @29.4 ng / μl, 1.7 μl volume makes an input of 50 ng. Assay will work down to 15 ng input high quality RNA). Reverse Transcription MasterMix “RT-MM” was made as listed in TABLE 5A and placed on ice.TABLE 5A RT-MM μl per 6 sample FSM (First Strand Master Mix, Illumina) 9 RVT (Reverse Transcriptase, Illumina) 1 Total Volume 10

[0231] RNA MasterMix “RMM” was made up as listed in TABLE 5B. TABLE 5B RMM μl per sample EPH3 (Elute, Prime, Fragment High Concentration Mix, Illumina) 1.7 “RT-MM” 1.6 Total 3.3

[0232] 3.3 μl of “RMM” per RNA sample well was pipetted making a total volume of 5 μl per well. Plates were sealed with Microseal ‘B’. The plates were shaken at 2100 rpm for 2 minutes, centrifuged briefly, and placed on thermocycler for the following “RT_FAST” protocol in TABLE 6. TABLE 6 RT_FAST 25°C 2 minutes 42°C 2 minutes 70°C 2 minutes 10°C

[0233] Once “RT_FAST” was completed, plates were centrifuged at 300 xg for 1 minute.

[0234] DNA Samples (added to same plate as RNA after reverse transcription):

[0235] 2 μl of gDNA samples or DNA controls were added to the plate in new wells along with 3 μl DEPC(Diethylpyrocarbonate)-treated and sterile filtered water.

[0236] Prepare DNA and RNA-derived cDNA for Multiplex PCR:

[0237] The following "mPCR MasterMix" of TABLE 7 was made per well of sample, and kept on ice. TABLE 7 mPCR MasterMix μl per sample TTM (Illumina) 16.3 Thermo Phusion™ Hot Start II DNA Polymerase (Thermo Fisher) 1.0 Primers (SEQ ID NOs: 197-392) @ 50nM 2.8 DEPC water 15.0

[0238] The MasterMix was pulse vortexed and briefly centrifuged.35 μl of master mix was added to each sample well, using a trough and multichannel pipettes. 10 μl of IDT- ILMN DNA / RNA UD Index Set A (Illumina) index primers were added to the appropriate wells. The plates were sealed with Microseal 'B' and shaken at 2100 rpm for 2 minutes. At this stage, each sample well contained the reagents listed in TABLE 8. TABLE 8 μl volume DNA / RNA Sample 5 mPCR Master Mix 35 IDT-ILMN DNA / RNA UD Index Set A (20026121) 10 Total volume per sample well 50

[0239] FIG.2 illustrates amplification with gene-specific (GS) forward and reverse (GS-F and GS-R) primers, such as SEQ ID NOs: 197-392, followed by amplification with Illumina® standard adapters which include P5 and P7 adapters configured to bind with matching sequences on a flowcell. These steps are further illustrated in FIG.3. Plates were centrifuged to 300 g for 1 minute and placed on post-PCR thermocycler running the PCR program listed in TABLE 9A with preheat lid option to 100°C. TABLE 9A mPCR^FAST (2hr10min) Ramp Rate Cycling 98°C 3 min 96°C 15 sec 5mPCR^FAST (2hr10min) Ramp Rate Cycling 70°C 5 sec 53°C 60 sec 0.1°C / s** 72°C 30 sec 0.1°C / s** 96°C 15 sec 22 58°C 60 sec 72°C 60 sec 72°C 2 min 10°C HOLD

[0240] Indexing PCR Clean-up

[0241] SPRI beads were removed from 2-8°C storage and allowed to come to room temperature (at least 30 minutes). Resuspension buffer (RSB) was removed from 2-8°C storage. PCR plate was removed from thermal cycler and allowed to come to room temperature, and centrifuged at 1000 g for 30 seconds. An 80% ethanol solution was prepared.

[0242] Pool PCR Libraries: 45 μl of each sample was pooled into a 1.5 ml Eppendorf tube, and according to TABLE 9B. TABLE 9B Total number of Volume of PCR amplified Total volume of sample pool for samples sample to pool (μl) SPRI cleanup (μl) <=4 45 45-180 >4 <80Calculation: =160 / total numberof samples 160>=80 2 160

[0243] SPRI 1: The SPRI conducted is a X1 SPRI. Conduct a X1 according to the total volume of the sample pool according to TABLE 9B. Mix SPRI and pooled PCR Product: Pooled PCR libraries were taken from the "Pre-SPRI Pool" in the Eppendorf tube and placed into a labelled MIDI plate. SPRI beads were added. The plate was shaken on the plate shaker for 2 minutes @2100 rpm. The plate was incubated at room temperature (RT) for 5 minutes. The plate was centrifuged for 30 seconds @300 rpm. The MIDI Plate containing the bead / sample mixture was placed on Magnet. The beads were allowed to pellet 5 minutes oruntil liquid is clear. With the plate still on the magnet, the supernatant was removed and discarded from the tube.

[0244] EtOH Washes: 200 μl of the freshly prepared RT 80% EtOH was added slowly to each well containing pooled library of the 96-well MIDI plate. After 30 seconds, supernatant was removed and discarded without disturbing the beads. These steps were repeated to complete a second 80% EtOH wash. After a second wash, a 20 l pipette was used to remove residual EtOH. The plate was air dried for 5 minutes.

[0245] Resuspend & Elute: The plate containing the dried beads was removed from the magnetic stand and 27 μl of RSB was added. The plate was shaken on the plate shaker for 2 minutes @2100 rpm. The plate was incubated at room temperature (RT) for 5 minutes and centrifuged for 30 seconds @300rpm. The MIDI Plate containing the bead / sample mixture was placed on Magnet and the beads were allowed to pellet 5 minutes or until liquid is clear.

[0246] SPRI 2: 25 L was taken from the pool and was placed into a new clearly marked MIDI plate well. A second SPRI (SPRI 2) was performed similar to SPRI 1. The SPRI conducted is a X1 SPRI. After SPRI 2, 8 μl was taken from the pool and placed into a new 96 well plate or final library tube.

[0247] TapeStation™ D1000: The library pool was assessed with a TapeStation D1000 (Agilent, Santa Clara, CA) using standard protocols. FIG.4A and FIG.4B show typical TapeStation™ HSD1000 traces for a high quality DNA control sample and for a high quality RNA control sample.

[0248] Sequencing: The library pool was diluted to 2 nM in RSB and mixed with NaOH and HT1. The library pool was brought to 14 pM concentration and mixed with PhiX. The diluted, denatured library pool was loaded into a MiSeq® Cartridge. The library pool was sequenced using a MiSeq® system (Illumina). Example 2:

[0249] A human biological sample from a patient arrives in a lab. A nucleic acid sample is obtained from the human biological sample. If it is previously known that the sample is derived from the same individual as another previously- or simultaneously-received sample, the sample(s) are annotated appropriately.

[0250] A Sample ID panel is run on the sample by contacting DNA from the sample with the sets of primers comprising sequences as set forth in SEQ ID NOs: 197-392, Thermo Phusion™ Hot Start II DNA Polymerase, Illumina® UD index adapters and Illumina® TTM buffer; amplifying; cleaning up with amplified product with two sequential SPRI steps; and Illumina® sequencing. Amplification can proceed with primers (SEQ ID NOs: 197-392).

[0251] The sample is submitted to whole genome & transcriptome sequencing (WGTS). The Sample ID check is run to compare the WGTS result to the original Sample ID panel result before returning the data to the patient. A sample identity barcode (SIB) from the Sample ID panel and from the WGTS result is generated and is compared to confirm sample identity. Example 3:

[0252] In the following example, different combinations of DNA polymerases and buffers were tested for ability to amplify oligonucleotides.

[0253] Enzymes: PCR Amplification Enzyme (PAE), also known as FEM, also known as Phusion Hot Start II DNA polymerase without Triton X-200 (PN#15054836 Illumina); TruSight Tumor Targeting Enzyme (TTE) (PN# 15069896, Illumina); Phusion™ Hot Start II DNA Polymerase (“Thermo Enzyme”) (PN#F549L, Thermo).

[0254] Buffers: PCR Amplification Buffer (PAB) (PN#20019096, Illumina); TruSight Tumor Targeting Mix (TTM) (PN# 15069879, Illumina); 5x Phusion™ HF Buffer (“Thermo Buffer”) (contains 7.5 mM MgCl2) for the Hot Start II Polymerase - within the enzyme kit (PN#F549L, Thermo).

[0255] Conditions: Standard multiplex PCR settings. 16ng total DNA input. TCO (PN#20018993 Illumina) used as a control.

[0256] Protocol: A MasterMix was created with 13 μl of PCR buffer, 0.8 μl PCR Enzyme, and 2.2 μl PCR oligonucleotides. This was combined with target DNA (16 ng total in 16 μl) and i5 and i7 indexes (4 μl each, 8 μl total). A 2-step PCR reaction and SPRI cleanup were performed.

[0257] Results: As seen in FIG.5, yield for TCO amplified with PAE and PAB was much greater than yield with test oligos (IDPv1; primer oligos with sequences as set forth in SEQ ID NOs: 673-952) amplified with Thermo Enzyme (Phusion™ Hot Start II DNAPolymerase) and Thermo Buffer (5x Phusion™ HF Buffer (contains 7.5 mM MgCl2). Similarly, in FIG.6, library products yield increases using PAE and PAB in both TCO and test oligos. This is further illustrated in FIG.7, where it is shown that yield for test oligos was much greater in amplification reactions using PAE and PAB than in reactions using Thermo Enzyme and Thermo Buffer.

[0258] It was then examined whether it was the enzyme or buffer that increased the yield, or whether it was a combination of both the enzyme and the buffer.

[0259] FIG.8 illustrates a comparison of amplification yields from different buffers and the resulting yield of test oligos, where it is seen that PAB, TTM, and PAB* (made in house) had similar yields, while Thermo Buffer had a much lower yield. FIG.9 illustrates an overlay of results from PAB, TTM, and PAB*. Based on these results, it was observed that changing the buffer from Thermo Kit buffers to either PAB, TTM or in-house PAB* increases the library yield significantly, and that all three of the latter buffers obtain a similar yield.

[0260] FIG.10 illustrates a comparison of different DNA polymerase enzymes. Based on these results, it was observed that changing the enzyme does not significantly increase the library yield, or appear to have any positive significant changes. Thermo Enzyme was selected as having the best yield.

[0261] FIG.11 illustrates a comparison of different combinations of DNA polymerase enzyme and buffer. Based on these results, it was observed that maximum yields are obtained by using Thermo PhusionTMHot Start II DNA Polymerase in combination with TTM buffer or PAB* buffer made in house. FIG.12 illustrates a comparison of yields from Thermo PhusionTMHot Start II DNA Polymerase and TTM buffer compared to Thermo PhusionTMHot Start II DNA Polymerase and Thermo Buffer. Example 4: Generation and comparison of Hankoprints

[0262] The following example relates to rapid positive confirmation of sample identity in a WGTS workflow.

[0263] Next Generation Sequencing (NGS) can rapidly identify genetic variation in human samples and is a vital tool in both research and diagnostic laboratories. The data generated is important to the subject and can be significant, it is therefore useful to ensure that the correct data is returned in timely fashion. Described herein are two strategies employed torapidly confirm sample identity in a Whole Genome and Transcriptome Sequencing (WGTS) pipeline. The first is a multiplex PCR assay targeting 90 Single Nucleotide Polymorphisms (SNPs) that can be run on DNA or RNA to generate a sequencing library from the sample on the day of sample receipt. The second is software (called Hanko) that generates a ‘barcode’ from each sample based on these SNPs. Barcodes, or Hankoprints, can be generated from either multiplex PCR assay or WGTS data. The software compares pairs of Hankoprints to positively confirm sample identity. The combined system allows confirmation of sample identity pre- WGTS, to confirm information in an accompanying sample manifest, and post-WGTS, to confirm that the same Hankoprint is obtained from the data ultimately reported by the WGTS pipeline. Implementation of this workflow identifies any sample(s) sent in error earlier in the process than a standard post-WGTS check, preventing these from being sequenced.

[0264] Next Generation Sequencing (NGS) has become the gold standard for the identification of genetic variation in a variety of settings. The reduction in sequencing cost has allowed researchers to explore variation in a greater number of individuals, across larger regions of the genome, with studies sometimes involving multiple samples from the same individual to yield temporal or multi-omic information. A number of important applications involve comparison of two or more data sets derived from the same individual. Some of the most prominent examples come from the field of cancer genomics, where the sequencing of a tumor and matched normal sample identify somatic variants [1], and the addition of an RNA sample enables Whole Genome and Transcriptome Sequencing (WGTS) [2]. To generate meaningful data, it is essential that the samples processed are derived from the same individual. There are a number of places in the multi-step sample to answer workflow where samples may become mislabeled or swapped, particularly where samples are shipped between locations for processing. Indeed, a number of studies have sought to identify and estimate how frequently such events occur [3-6]. Therefore, a sample identity check would be a valuable quality control step in such workflows.

[0265] While data analyses post sequencing can resolve most issues, positive confirmation cannot happen until the end of the entire process [7-9]. The ability to identify discrepancies before samples enter a WGTS pipeline would be highly desirable, avoiding time and money being spent on sequencing incorrect samples. Moreover, the request for areplacement sample could be made many days sooner, minimizing the delay in reporting of the correct result.

[0266] Genetic data is by its nature highly identifiable, and the human genome contains many highly polymorphic sites between individuals. Perhaps the simplest of these is the Single Nucleotide Polymorphism (SNP) where a single base pair at a particular position in the genome differs in >1% of individuals. Large scale population sequencing efforts have catalogued many of these over recent years providing a rich source of well characterized genetic markers [10-12]. The size of these changes makes them highly amenable to detection by a variety of assays and means that they can be reliably identified even in highly fragmented or degraded DNA samples

[0013] . Thus, many groups have used panels of SNPs to help with the identification of individuals, although the specific SNPs included often differ based on the application [7, 9, 14-16].

[0267] To enable identification of samples sent in error at the start of the WGTS pipeline, a sample identity check was implemented. The check had two components, a multiplex PCR reaction that is run on samples when they first arrived in the laboratory, and an analysis software package, ‘Hanko’, that could generate a sample specific barcode or ‘Hankoprint’ from either multiplex PCR or WGTS data. Two or more Hankoprints were compared to determine whether a sample originates from the same individual, whether individuals could be a parent-child match, or even whether a sample contains genetic information from more than one individual. Methods Target selection

[0268] SNPs were selected from the genome aggregation database (gnomAD; v2.1

[0012] ). All SNPs with minor allele frequency (MAF) between 0.3 and 0.7 in African, South and East Asian and European populations were searched, then the search was restricted to variants contained in the 35,000 exons of 3,804 ubiquitously expressed genes

[0017] resulting in 275 SNPs. In addition, twenty-four SNPs were considered which were previously used in a comparison of exome products [7] and eight regions on chromosome Y.

[0269] Target SNPs were manually reviewed to confirm expression in a range of tissues using the GTEx portal and to identify the location of the SNP in the exon relative to intron / exon boundaries. Targets located toward the center of an exon were preferentiallyselected, as these would amplify DNA and RNA, allowing the same panel to be run on both input types. SNPs were confirmed to be biallelic and to show no evidence of allele-specific expression. Finally, targets were selected with the aim of achieving even target distribution across all autosome chromosome arms of the human genome.

[0270] Targets were converted to primer pairs with the help of Primer-BLAST

[0018] . Tails were added to primers to act as annealing sites for standard Illumina adapters. Forward primers had a ‘GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG’ (SEQ ID NO:953) splint while reverse primers had‘TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG’ (SEQ ID NO:954) added 5 to thetarget specific primer pairs. Multiplex PCR Sample ID panel assay For RNA samples:

[0271] A rapid, small volume reverse transcription step was performed to convert RNA to cDNA to use as template in the multiplex PCR. For each sample, 1.7 μl of RNA (minimum input 9 ng / μl) was transferred to a 96 well PCR plate (BioRad; HSP-9601). A FSM+RVT master mix (per 6 samples) was created by combining 9 μl FSM (Illumina PN#15026783) and 1 μl RVT (Illumina PN#20002175), this was briefly vortexed, centrifuged and placed on ice. A second master mix (RT-MM) was prepared by combining 1.7 μl of EPH3 (Illumina PN#20002174) and 1.6 μl of FSM+RVT per sample, this was briefly vortexed and spun down to mix before being placed on ice. RT-MM (3.3 μl) was added to each RNA sample containing well of the 96 well PCR plate to make a total volume of 5 μl per well. The plate was sealed (Microseal “B”; BioRad PN#MSB-1001), placed on a shaker at 2100 rpm for 2 min and centrifuged for 1 min at 300 g. The plate was then incubated for 2 minutes each at 25°C for random hexamer annealing, 42°C for reverse transcription and 70°C for RT enzyme inactivation, before being held at 10°C. The reaction plate was briefly centrifuged before addition of the PCR master mix in the Multiplex PCR assay section below. For DNA samples:

[0272] A 2 μl aliquot (minimum input 8 ng / μl; 25 ng / μl for FFPE) of each sample was added to a 96 well PCR plate and the volume made up to a total of 5 μl with nuclease free water.Multiplex PCR assay

[0273] A multiplex PCR mastermix (mPCR-MM) was made at room temperature consisting of 16.3 μl TTM (Illumina PN#15069879), 1.0 μl Thermo Phusion Hot Start II DNA Polymerase (Fisher Scientific PN# 10628439), 2.8 μl 50 nM Custom Oligo Pool (IDPv2), and 15 μl nuclease free water. The master mix was added to each sample well (35 μl), along with 10 μl of the appropriate unique index (Illumina IDT-ILMN DNA / RNA UD Index Set A PN#20026121, IDT-ILMN DNA / RNA UD Index Set B PN# 20026930, IDT-ILMN DNA / RNA UD Index Set C PN# 20026934, IDT-ILMN DNA / RNA UD Index Set D PN# 20026933) for a total PCR volume of 50 μl. The plate was transferred to a PCR machine and incubated for 3 minutes at 98°C, then 5 cycles of 96°C for 15 seconds, 70°C for 5 seconds, ramping down to 53°C for 60 seconds (ramp rate -0.1°C per second), 72°C for 30 seconds (ramp rate +0.1°C per second), followed by, 22 cycles of: 96°C for 15 seconds, 58°C for 60 seconds, and 72°C for 60 seconds. A final extension was the performed of 72°C for 2 min and the reaction held at 10°C.

[0274] Following PCR, all samples were uniquely indexed, and therefore can be combined for purification. For large numbers of samples a combined sample volume of 160 μl for purification was recommended. Two 1x Sample Purification Bead (SPB; Illumina PN# 15037172) purifications were performed sequentially. Samples were incubated for 5 minutes with an equal volume (1x) of SPB, placed on a magnetic stand for 5 minutes at room temperature, before being washed twice with 80% ethanol for 30 seconds. Pooled samples were resuspended by addition of 27 μl RSB (Illumina) for the first purification, and 10 μl RSB (Illumina) for the second.

[0275] Amplicon libraries were quantified using the Agilent D1000 ScreenTape on the Agilent 4200 TapeStation System (Agilent #5067-5582; #5067-5583). Up to 384 samples were pooled for sequencing on the MiSeq (Illumina). Libraries were denatured and loaded at 14 pM with a 1.5% PhiX control library using the standard Illumina MiSeq v3 protocol (2x 75 bp reads). WGTS

[0276] Whole Genome and Transcriptome sequencing was performed. Reads were aligned to Human Reference genome version 38 (GRCh38) and variants called using Illumina DRAGEN pipelines.Generation and comparison of unique sample barcodes Generation of Hankoprints

[0277] Hankoprints were generated using the Hankogenerate software hosted on the Illumina BaseSpace Sequencing Hub cloud platform. For Multiplex PCR data, reads were aligned to the Human Reference genome version 38 (GRCh38) with the Isaac aligner (Illumina), then variants and genotypes called at the positions of the known SNPs using the Pisces variant caller (Illumina). Genotypes were encoded as Hankoprints depending on whether the position was called as reference (0), heterozygous (1) or homozygous alternative (2), with specific encoding for unexpected genotypes (3) or where a call could not be made at this position (4). In the case of WGTS, where alignment and variant calling have already been performed by a DRAGEN pipeline, the bam, cram or variant call file (VCF) can be used as input.

[0278] For sex karyotype detection, if more than half of the positions on chromosome Y had reads aligned to them the sample was scored as XY, otherwise the sample was assumed to be XX. Comparison of Hankoprints

[0279] Hankoprints were compared using the Hankocompare software hosted on the Illumina BaseSpace Sequencing Hub cloud platform. The comparison generated a ratio of identity between any two Hankoprints. The output ratios returned were derived from the comparison of: i) all autosomal positions, ii) autosomal positions called as homozygous only (to remove uncertainty where one of the samples is a tumor sample with multiple regions of chromosomal gain or loss); and iii) a ratio based on the hypothesis that the samples being compared are a parent and child. Ratios were plotted in heatmaps and data used to generate these was supplied in csv format. Supplementary methods Kinship comparison methodology (parent and child checker)

[0280] All autosomal positions are compared according to the following rules: - If parent is 0 then child must be 0 or 1 - If parent is 1 then child may be 0, 1 or 2 - If parent is 2 then child must be 1 or 2 - Positions encoded as a 3 or 4 are ignored

[0281] The number of positions that meet the above criteria were divided by the total number of positions compared in the sample (i.e. those not a 3 or 4). The default threshold for samples to be called as a match is >0.99 and at least 45% of the panel positions must be available to make the comparison (40 positions in total). Contamination detection methodology

[0282] The variant allele frequency (VAF) observed for each position in the sample ID panel was binned in bins of width 0.01 over the range 0 to 1. Any positions with values between 0 and 0.1 and 0.9 and 1 were classified as homozygous, while values between 0.1 and 0.9 and classified as heterozygous. The contamination ratio was generated by dividing the number of positions in the homozygous group by the number in the heterozygous. Any sample with a ratio of >1.0 is flagged as potentially contaminated. Results Identification of candidate SNPs for a Sample ID panel

[0283] The process of identifying potential SNPs for the sample ID panel is summarized in FIG.13, panel (a). Candidate SNPs were searched in the genome aggregation database (gnomAD; v2.1

[0012] ), focusing the search on SNPs seen at significant frequency in African, South and East Asian and European populations. To ensure that the panel was RNA compatible, and would work in RNA samples derived from a variety of tissues, candidates were restricted to those contained in the exons of ubiquitously expressed genes

[0017] . In addition, 24 SNPs previously used in a comparison of exome products were considered [7]. The resultant 299 potential targets were manually reviewed to check that the position of the SNP in the exon would allow primer annealing on either side in both DNA and RNA. It was also confirmed that positions of interest were biallelic, showed no evidence of allele-specific expression, and were evenly distributed across all autosome chromosome arms of the human genome. Primer design was successful for 94 autosomal positions plus 8 target regions on chrY (FIG.13A, FIG.13B).

[0284] Primer performance was evaluated in multiplex PCR assays using DNA as input. Despite multiple rounds of optimization, four poorly performing (low yield / coverage) primer pairs were dropped from the multiplex pool resulting in a final assay targeting 90 autosomal positions and 8 positions on chrY (98-plex multiplex PCR).

[0285] FIG.13A and FIG.13B schematically illustrates the sample ID panel target selection process (FIG. 13A) and multiplex PCR assay (FIG. 13B). FIG. 13A illustrates the number of SNPs or target regions considered at each stage of the selection process as detailed in methods. MAF=Minor Allele Frequency. FIG. 13B illustrates a cartoon representation of the 4-primer style multiplex PCR assay. For each sample, 98 GS primer pairs and one Illumina adapter pair (for indexing) are added to the multiplex PCR assay. Target specific GS-F & GS- R primers with splint sequences amplify the target regions creating a product that standard Illumina adapter primers can anneal to and amplify, generating a full-length sequencing library. After PCR all samples are unique due to the adapter index sequence and so can be pooled for purification, quantification and sequencing. Validation of Sample ID panel and analysis software using DNA from a large family pedigree

[0286] A useful feature of a system for confirming sample identity is the ability to distinguish closely related individuals. To test this, the ID panel workflow on DNA samples derived from a number of individuals in CEPH pedigree 1463

[0019] (FIG. 14A). Data was analyzed using Hanko, a software tool that encodes the genotype of the sample at the positions of the Sample ID panel as a ‘Hankoprint’. Hankoprints can be thought of as barcodes and may be generated from either sample ID panel data or WGTS workflows. Two Hankoprints can be compared to determine whether samples are derived from the same individual. The results of multiple two-way comparisons are plotted as a heatmap displaying the similarity ratio between any two Hankoprints (FIG. 14B). No sample in CEPH pedigree 1463 matched another across more than 66 of the 94 autosomal positions (similarity ratio = 0.7), demonstrating the ability of these SNPs, and this assay, to uniquely identify closely related individuals. Identical results were observed with 30 ng and 60 ng input of DNA.

[0287] FIGs. 14A-14C illustrate the validation of the Sample ID panel in the platinum genomes family. FIG.14A depicts a CEPH pedigree 1463 family members. FIG.14B depicts Hankocompare output from sample ID panel. Each member of the platinum genomes family was run through the sample ID panel with 30 ng and 60 ng input for comparison. Parents, children, then grandparents are plotted in numerical order with the 30 ng input plotted first then the 60 ng input immediately after for each sample. Darker squares represent comparisons with a higher proportion of identical genotypes. Dark squares in the diagonal represent samples being compared to themselves and having 100% identity. FIG. 14Cillustrates a histogram of Hanko similarity ratios derived from all autosomal positions in 3,202 individuals in the 1000 genomes dataset. Validation of Sample ID panel positions in a 1000 genomes dataset

[0288] To further demonstrate the discriminative power of this panel of SNPs, Hankoprints were generated for 3,202 samples from the 1000 Genomes Dataset that had previously been reanalyzed using Illumina Dynamic Read Analysis for GENomics (DRAGEN) v3.5.7b and made available on AWS (aws.amazon.com / blogs / industries / dragen- reanalysis-of-the-1000-genomes-dataset-now-available-on-the-registry-of-open-data / ); accessed 12th February 2024, originally described

[0020] ). Hankoprints were generated and compared across all samples in this dataset. No sample matched another, the mean similarity ratio across all samples was 0.37, with no comparison yielding a higher similarity ratio than 0.72 across the 90 autosomal positions (FIG.14C). Application of Sample ID panel on FFPE breast tumor and matched blood normal

[0289] The identification of somatic variants by sequencing a tumor and matched normal sample is a powerful example of where the implementation of a sample identity check could bring significant benefit. Samples must be derived from the same individual to be informative, and significant costs result in sequencing the wrong sample. The sample ID panel was applied to a batch of samples received for sequencing in the WGTS pipeline. The accompanying sample manifest for these samples detailed paired samples from 195 individuals, with 6 individuals providing multiple tumor samples: a total of 201 FFPE tumor samples and 195 matched normal samples. Processing these samples through the workflow confirmed receipt of tumor normal pairs for the majority of individuals (194 / 195), however one normal sample was identified as having no matching FFPE tumor sample associated with it (FIG. 15A).

[0290] While the majority of comparisons of expected sample pairs gave a similarity ratio of 1 (all positions matching; darker / black boxes), some pairs yielded a lower similarity score (greater than or equal to 0.7; three left arrows above left heatmap, FIG. 15B). Lower similarity ratios would normally indicate a sample mismatch, however closer inspection revealed that discrepant positions were restricted to those determined to be heterozygous in the matched normal. This is significant because copy number gains and losses are frequently observed in cancer genomes, with the gain or loss of associated alleles. Thus, heterozygouspositions could be lost through loss of heterozygosity where a copy number change occurs. To provide confident matching in somatic tumor samples, Hankocompare performs an additional comparison restricted to homozygous positions in the matched normal (FIG.15B). While this reduces the number of positions in the comparison, focusing on these SNPs allows the correct resolution of tumor normal pairs in tumor samples where heterozygosity has been lost.

[0291] FIGs 15A-15B illustrate the application of sample ID panel to confirm tumor normal pairs. FIG. 15A illustrates Hankocompare plots from a subset of 31 of the 195 paired samples tested. Hankoprints derived from the normal samples (x axis) are compared to those derived from the FFPE tumors (y axis). Expected sample pairs (from the accompanying sample manifest) are plotted in the same order for ease of interpretation. The dark blue / black diagonal quickly highlights matching samples. Black arrow denotes an individual with two FFPE samples, red arrow indicates an instance where the tumor and normal samples received do not match (each other or any other in the 31 plotted). FIG. 15B illustrates a comparison of similarity plots for All Autosomal positions and Homozygous autosomal positions only. Zoomed in segment of the plot in FIG. 15A. The three arrows (FIG. 15B in left panel, at left) indicate sample pairs with lower similarity ratios due to the failure to match heterozygous positions. Plotting only homozygous positions reveals true matches (FIG. 15B in right panel, three arrows at left), while samples that do not match continue to have low similarity ratios (FIG.15B in right panel, arrow at right). Validation of Sample ID for use on RNA samples

[0292] To assess the performance of the sample ID panel on RNA samples, the assay was run on 13 unrelated RNA samples derived from breast tissue (FIG. 16A). All samples were identified as unique and showed very little similarity to one another. The assessment of panel performance was extended to sixteen unrelated RNA samples derived from blood, this time comparing the Hankoprints generated from the sample ID panel data to those generated from a set of thirty-nine samples for which Whole Transcriptome Sequencing (WTS) had been performed. This group included the sixteen samples run with the sample ID panel. The panel was able to uniquely identify the 16 samples run by both methods, confirming that the method can be used for RNA samples (FIG.16B).

[0293] Performing a sample identity check in RNA presents various challenges over DNA. While exons thought to be constitutively expressed were prioritized in targetselection, there may be some tissues or cells where this these are in fact absent, and the level of transcript expression can vary significantly from gene to gene and tissue to tissue. To evaluate how well the SNPs selected performed in real WTS data, performance was examined in RNAs derived from cohorts of 108 breast cancer samples, 216 glioma samples and 5,724 blood samples (FIG. 16C, FIG. 16D). The majority of samples had successful calls made at more than 50 positions (FIG. 16C). Plotting the frequency of Hankoprint values at each position in the panel reveals tissue specific differences for each position. As expected, some positions show poorer performance in some tissues but not all (highlighted by an excess of ‘4s’ at these positions signifying not enough data to make a genotype call), presumably related to the level of expression of that target in those tissues. Interestingly no autosomal target failed in all samples across all tissues. Taken together, WTS data confirms expression of the majority of this set of SNPs in samples derived from brain, breast and blood, suggesting robust panel performance in these tissues.

[0294] FIG. 16A illustrates Hankocompare output from 13 RNA breast cancer samples processed with the sample ID panel. FIG. 16B illustrates Hankocompare output comparing 16 sample ID panel results to 39 WTS results from samples derived from blood. FIG.16C illustrates a summary of Hankoprints generated from cohorts of WTS data in 3 tissue types. Called positions were successfully genotyped and are encoded as 0, 1 or 2 in the Hankoprint. FIG. 16D illustrates the frequency of Hankoprint encoding at each of the 98 sample ID positions in the WTS data presented in FIG. 16C. Flagging potential sample contamination

[0295] Potential further applications for the sample identity checking pipeline were explored, considering information that would be useful to know upfront of a costly WGTS experiment. One such application is sample contamination checking, where nucleic acid from one sample mistakenly gets combined with another resulting in a sample well containing genetic information from multiple individuals. When this occurs at a significant level, it can be challenging to identify or assign causative mutations in samples, prompting a request for a new sample. The pipeline should be able to identify such an occurrence, as the presence of additional alleles would result in an increase in the number of apparently heterozygous positions in such samples. DNA samples were artificially mixed from five well characterized family groups from the 1000 genomes project; four of which are included in from the NationalInstitute of Standard and Technology (NIST) genome in a bottle project

[0021] . In the majority of cases a female child was spiked into the mother’s DNA as this would be the most challenging contamination to detect. Hanko was adapted to consider the variant allele frequency (VAF) observed at each of the autosomal positions, generating a new ratio derived from dividing the number of variants of VAF less than 0.1 and greater than 0.9 by those seen with a VAF between 0.1 and 0.9. By applying a threshold of 1.0 on this ratio to the spike in samples the method correctly flagged contamination in all experiments where the simulated contamination level was >5% (FIGs. 17A-17D). Even lower levels of contamination were detected in some of the samples. Hankogenerate was configured to plot an allele frequency histogram of the variant allele frequencies called at ID panel positions for all samples where the zygosity ratio falls above the threshold, allowing the user to see population of homozygous reference and alternative alleles moving toward the 0.5 depending on the level of the contamination (FIGs. 17B -17D).

[0296] FIGs. 17A-17D depicts a table and graphs related to the contamination detection feature. Panel (a) illustrates a table showing DNA from child samples (daughter: NA12879, NA19240, HG00733, HG00514; son: HG002) were artificially spiked into DNA from mother (NA12878, NA19238, HG00732, HG00513 & HG004) samples in 5 well characterized reference sample families. Green box: Potential contamination flagged by Hankogenerate, Red box: no contamination detected. Panel (b) illustrates allele frequency histograms from a selection of the experiments summarized in (a). Variant allele Frequencies (VAFs) of variants called at sample ID panel positions are plotted. Samples with contamination see increasing numbers of variants deviating from the expected VAFs of 0 and 1.0 for homozygous reference and alternative positions. A dispersion of the ~0.5 VAF cluster can also be seen in the 20% spike in. Discussion

[0297] NGS has become routine in both scientific research and the clinic and the data associated with the samples under analysis, be they participants or patients, will only increase as new applications such as multiomic approaches become standard [22-26] (www.england.nhs.uk / publication / accelerating-genomic-medicine-in-the-nhs / ; accessed 6th March 2024). Whether testing is performed locally or samples are shipped to a central sequencing location, a sample, and an electronic or paper description of the sample, must passthrough a series of complex steps, including data entry, shipment, extraction and library preparation before sequencing data is returned. It is essential that quality control mechanisms are in place to avoid or flag any issues, ideally before the information becomes associated with other sample metadata.

[0298] The combined solution for confirming sample identity would make a valuable addition to any sequencing workflow or pipeline. The ability of a laboratory that returns NGS data to a customer to have documented evidence that the data matches the samples received allows rules out their laboratory as the source of any unexpected results. While labs that solely work in a 96 well format can introduce positive or negative controls at specific wells on the plate to confirm that no plate rotation has taken place, this may still leave some degree of uncertainty around individual wells on a plate if plating onto the 96 well format has also been performed by the laboratory. The ultimate check would be application of the sample ID panel to a completely separate sample, taken at a completely different point in time. As successful matching in this scenario would only occur if no errors had occurred at any of the steps in the process, including sample collection and extraction.

[0299] The past few years has seen a number of groups propose various sets of SNPs for sample identity checking. While the majority of these have slightly different use cases, all utilize the SNP as the genetic marker of choice. The identification of a suitable set SNPs has been made possible by the various endeavors to catalogue human variation [10, 11], highlighting the importance of these and future projects to better characterize genetic data in populations currently underrepresented the datasets. As has been highlighted by others, the SNP is in many ways the ideal unit of measurement. At a single base pair in length it is amenable to a wide variety of assays that can rapidly return a result, even in highly degraded material.

[0300] The SNPs selected as the basis of the sample identity panel differ from those selected by others. The system described herein is advantageously compatible with RNA, a requirement that rules out the vast majority of previously described sets of SNPs utilized in other projects [9, 14, 15]. Inclusion of RNA presented two significant concerns; firstly, whether an appropriate number of informative SNPs could be identified in exonic sequences to allow us to discriminate between individuals in a population and in family groupings, and secondly, would those exonic regions lie in genes that were expressed in the majority of tissues of interestto enable the detection of signal in RNA samples. While the vast majority of SNPs are not located in gene exons, and many genes are expressed only in a restricted subset of tissues, a significant number of SNPs were identified that met both criteria through their presence in constitutively expressed genes. This was important to ensure the RNA panel would work in as many tissue types as possible. While the SNPs identified herein work well in the tissues examined thus far, there are likely to be some RNA samples that are more challenging than others. One limitation discovered here was the failure to detect the presence of chrY in WTS of RNA samples derived from blood (contrasting the high level of success in glioma samples). Further testing may demonstrate whether this finding would be replicated in sample ID panel data, where the targeted sequencing data may more sensitively detect low level expression from chrY regions in these samples.

[0301] Perhaps the biggest advantage of the presently-disclosed panel over others is the fact that it has been converted into a multiplex PCR assay. This provides a valuable tool to the sequencing laboratory as samples can be looked at in advance of processing in the laboratory. This not only creates a timestamp of what the sample is when it arrives but allows some errors to be identified in advance of larger sequencing experiments. Identifying these errors early in a sequencing workflow has several advantages. References 1. Mandelker, D. and O. Ceyhan-Birsoy, Evolving Significance of Tumor-Normal Sequencing in Cancer Care. Trends Cancer, 2020.6(1): p.31-39. 2. Cuppen, E., et al., Implementation of Whole-Genome and Transcriptome Sequencing Into Clinical Cancer Care. JCO Precis Oncol, 2022.6: p. e2200245. 3. Lippi, G., et al., Managing the patient identification crisis in healthcare and laboratory medicine. Clin Biochem, 2017. 50(10-11): p.562-567. 4. Westra, H.-J., et al., MixupMapper: correcting sample mix-ups in genome-wide datasets increases power to detect small genetic effects. Bioinformatics, 2011.27(15): p.2104-2111. 5. Warmerdam, R., et al., Idéfix: identifying accidental sample mix-ups in biobanks using polygenic scores. Bioinformatics, 2022. 38(4): p.1059-1066. 6. Lynch, A.G., et al., Calling sample mix-ups in cancer population studies. PLoS One, 2012. 7(8): p. e41815. 7. Pengelly, R.J., et al., A SNP profiling panel for sample tracking in whole-exome sequencing studies. Genome Med, 2013.5(9): p.89. 8. Westphal, M., et al., SMaSH: Sample matching using SNPs in humans. BMC Genomics, 2019.20(Suppl 12): p. 1001. 9. Pakstis, A.J., et al., SNPs for a universal individual identification panel. Hum Genet, 2010. 127(3): p. 315-24.10. Auton, A., et al., A global reference for human genetic variation. Nature, 2015.526(7571): p.68-74. 11. Karczewski, K.J., et al., The ExAC browser: displaying reference data information from over 60000 exomes. Nucleic Acids Res, 2017.45(D1): p. D840-d845. 12. Karczewski, K.J., et al., The mutational constraint spectrum quantified from variation in 141,456 humans. Nature, 2020. 581(7809): p. 434-443. 13. Kayser, M. and P. de Knijff, Improving human forensics through advances in genetics, genomics and molecular biology. Nat Rev Genet, 2011. 12(3): p.179-92. 14. Børsting, C., et al., Performance of the SNPforID 52 SNP-plex assay in paternity testing. Forensic Science International: Genetics, 2008. 2(4): p.292-300. 15. Krjutškov, K., et al., Evaluation of the 124-plex SNP typing microarray for forensic testing. Forensic Science International: Genetics, 2009. 4(1): p.43-48. 16. Yousefi, S., et al., A SNP panel for identification of DNA and RNA specimens. BMC Genomics, 2018. 19(1): p.90. 17. Eisenberg, E. and E.Y. Levanon, Human housekeeping genes, revisited. Trends Genet, 2013.29(10): p.569-74. 18. Ye, J., et al., Primer-BLAST: a tool to design target-specific primers for polymerase chain reaction. BMC Bioinformatics, 2012.13: p.134. 19. Eberle, M.A., et al., A reference data set of 5.4 million phased human variants validated by genetic inheritance from sequencing a three-generation 17-member pedigree. Genome Res, 2017.27(1): p. 157-164. 20. Byrska-Bishop, M., et al., High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell, 2022.185(18): p.3426-3440.e19. 21. Zook, J.M., et al., Extensive sequencing of seven human genomes to characterize benchmark reference materials. Sci Data, 2016.3: p. 160025. 22. Smedley, D., et al., 100,000 Genomes Pilot on Rare-Disease Diagnosis in Health Care - Preliminary Report. N Engl J Med, 2021.385(20): p. 1868-1880. 23. Degasperi, A., et al., Substitution mutational signatures in whole-genome-sequenced cancers in the UK population. Science, 2022.376(6591). 24. Dimmock, D., et al., Project Baby Bear: Rapid precision care incorporating rWGS in 5 California children's hospitals demonstrates improved clinical outcomes and reduced costs of care. Am J Hum Genet, 2021.108(7): p.1231-1238. 25. Kingsmore, S.F. and T.B. Consortium, Dispatches from Biotech beginning BeginNGS: Rapid newborn genome sequencing to end the diagnostic and therapeutic odyssey. American Journal of Medical Genetics Part C: Seminars in Medical Genetics, 2022.190(2): p.243-256. 26. Capper, D., et al., DNA methylation-based classification of central nervous system tumours. Nature, 2018. 555(7697): p.469-474. Other Considerations

[0302] Conditional language used herein, such as, among others, “can,” “might,” “may,” “e.g.,” and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or states. Thus, suchconditional language is not generally intended to imply that features, elements and / or states are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements and / or states are included or are to be performed in any particular embodiment. The terms “comprising,” “including,” “having,” “involving,” and the like are synonymous and are used inclusively, in an open-ended fashion, and do not exclude additional elements, features, acts, operations, and so forth. Also, the term “or” is used in its inclusive sense (and not in its exclusive sense) so that when used, for example, to connect a list of elements, the term “or” means one, some, or all of the elements in the list.

[0303] Disjunctive language such as the phrase “at least one of X, Y or Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to present that an item, term, etc., may be either X, Y or Z, or any combination thereof (such as X, Y and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y or at least one of Z to each be present.

[0304] The terms “about” or “approximate” and the like are synonymous and are used to indicate that the value modified by the term has an understood range associated with it, where the range can be ±20%, ±15%, ±10%, ±5%, or ±1%. The term “substantially” is used to indicate that a result (such as a measurement value) is close to a targeted value, where close can mean, for example, the result is within 80% of the value, within 90% of the value, within 95% of the value, or within 99% of the value.

[0305] Unless otherwise explicitly stated, articles such as “a” or “an” should generally be interpreted to include one or more described items.

[0306] While the above detailed description has shown, described, and pointed out novel features as applied to illustrative embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms illustrated can be made without departing from the spirit of the disclosure. As will be recognized, certain embodiments described herein can be embodied within a form that does not provide all of the features and benefits set forth herein, as some features can be used or practiced separately from others. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

[0307] It should be appreciated that all combinations of the foregoing concepts (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein.

[0308] The scope of the present disclosure is not intended to be limited by the specific disclosures of examples in this section or elsewhere in this specification, and may be defined by claims as presented in this section or elsewhere in this specification or as presented in the future. The language of the claims is to be interpreted broadly based on the language employed in the claims and not limited to the examples described in the present specification or during the prosecution of the application, which examples are to be construed as non- exclusive.

Claims

WHAT IS CLAIMED IS:

1. A method for confirming a nucleic acid sample identity, comprising: obtaining a nucleic acid sample from a human biological sample; determining sequence information for the nucleic acid sample at a plurality of sites, wherein the plurality of sites comprises 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or each of the sites corresponding to positions: chr1:6693097, chr1:41204569, chr1:59147926, chr1:114515717, chr1:150808889, chr1:212911836, chr1:227069737, chr1:227071525, chr1:227935444, chr1:236716959, chr1:236719193, chr2:3392295, chr2:101624471, chr2:101638888, chr2:128939817, chr3:4403767, chr3:15737689, chr3:33434831, chr3:155481609, chr3:186509517, chr3:193374964, chr4:6600012, chr4:6606864, chr4:25419283, chr5:53815495, chr5:64881936, chr5:82834630, chr5:96503523, chr6:26598188, chr6:49403282, chr6:49425521, chr6:101166095, chr6:146755140, chr6:158517308, chr7:2577781, chr7:2578237, chr7:43916727, chr7:43917013, chr7:48004962, chr8:30973957, chr8:33356074, chr8:33369944, chr8:90995019, chr8:94935937, chr8:104427359, chr8:124448736, chr8:132982824, chr8:135612745, chr8:146067054, chr9:132400480, chr10:1046712, chr10:31138817, chr10:99219885, chr10:100219314, chr10:104572963, chr11:16133413, chr11:125763746, chr12:993930, chr12:62926398, chr13:28143229, chr13:39433606, chr14:20920250, chr14:31381351, chr15:34528948, chr15:75131959, chr16:624114, chr16:2028402, chr16:19680546, chr16:70303580, chr17:10614442, chr17:17168164, chr17:40714804, chr17:40716520, chr17:57290383, chr17:71196809, chr17:79514129, chr17:80008392, chr18:21413869, chr19:885818, chr19:1010691, chr19:36727365, chr19:41939297, chr19:50983930, chr19:58929052, chr20:6100088, chr20:61834695, chr21:47961711, chr22:32795641, chr22:32875190, chr22:50962208, chrY:2961349, chrY:7063980, chrY:2866789, chrY:2980414, chrY:7091532, chrY:13705715, chrY:14842155, or chrY:20592667 of hg 19 - GRCh37, and based on the sequence information at each of the plurality of sites, confirming the nucleic acid sample identity.

2. The method of claim 1, wherein determining sequence information comprises:contacting DNA or cDNA from the nucleic acid sample with a plurality of pairs of primers, a DNA polymerase buffer, and a DNA polymerase, amplifying the DNA or cDNA at the plurality of sites to generate amplified nucleic acid fragments, and sequencing the amplified nucleic acid fragments.

3. The method of claim 2, wherein the plurality of pairs of primers comprises at least 20 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.

4. The method of claim 2 or claim 3, wherein the plurality of pairs of primers comprises at least 50 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.

5. The method of any of claims 2-4, wherein the plurality of pairs of primers comprises at least 70 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.

6. The method of any of claims 2-5, wherein the plurality of pairs of primers comprises at least 90 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.

7. The method of any of claims 2-6, wherein the plurality of pairs of primers comprises pairs of primers comprising a sequence as set forth in each of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.

8. The method of any of claims 2-7, wherein the DNA polymerase buffer comprises Illumina TTM buffer.

9. The method of any of claims 2-8, wherein the DNA polymerase comprises Thermo Phusion Hot Start II DNA Polymerase.

10. The method of any of claims 1-9, wherein the nucleic acid sample comprises DNA.

11. The method of any of claims 1-10, wherein the nucleic acid sample comprises RNA.

12. The method of claim 11, wherein the method further comprises obtaining cDNA based on the RNA from the nucleic acid sample.

13. The method of any of claims 1-12, wherein determining sequence information comprises determining whether the nucleic acid sample is homozygous for a reference allele, homozygous for an alternative allele, or heterozygous at each site of the plurality of sites.

14. The method of any of claims 1-13, wherein the human biological sample comprises a blood sample, a tumor sample, a saliva sample, a hair sample, a semen sample, a skin sample, a muscle sample, or an organ sample.

15. The method of any of claims 1-14, further comprising confirming a nucleic acid sample sequence identity for a second nucleic acid sample from a second human biological sample.

16. The method of claim 15, wherein the human biological sample and the second human biological sample are from an individual.

17. The method of claim 16, wherein the two human biological samples comprise a tumor sample and a normal sample from the individual.

18. The method of claim 17, wherein the method further comprises excluding sequence information related to a site of the plurality of sites when the sequence information at the site indicates heterozygosity.

19. The method of any of claims 1-18, wherein amplifying DNA or cDNA from each of the one or more nucleic acid samples comprises multiplex PCR.

20. The method of any of claims 1-19, further comprising adding an adapter sequence to amplified nucleic acid fragments.

21. The method of claim 20, wherein the adapter sequence is added by amplification.

22. The method of claim 20 or claim 21, wherein the method comprises contacting nucleic acid fragments with a plurality of primers comprising an adapter sequence.

23. The method of any of claims 20-22, wherein determining sequence information comprises hybridizing the amplified nucleic acid fragments to a nucleic acid adapter comprising a sequence complementary to the adapter sequence.

24. The method of any of claims 1-23, wherein sequencing is performed in multiplex format.

25. A computer-implemented method of comparing nucleic acid samples, comprising:(a) receiving sequence information for a first nucleic acid sample at a plurality of sites; (b) generating from the sequence information a first sample identity barcode (SIB) for the first nucleic acid sample, wherein the first SIB comprises an ordered list of identifiers for the plurality of sites; (c) generating a ratio of identity by comparing the first SIB and a second SIB, wherein the second SIB is generated from sequence information for a second nucleic acid sample; and (d) providing an indication of a match between the first and second nucleic acid samples based on the comparison.

26. The method of claim 25, wherein the plurality of sites comprises: (i) a plurality of single nucleotide polymorphisms (SNPs); (ii) locations distributed throughout a genome; (iii) locations in exons; (iv) biallelic sites; (v) locations within ubiquitously expressed genes; (vi) lack allele specific expression; and / or (vii) locations comprising a minor allele frequency in African, South and East Asian and European populations in a range from 0.3 to 0.

7.

27. The method of claim 25 or 26, wherein the sequence information is obtained by the method of any one of claims 1-24.

28. The method of any one of claims 25-27, wherein step (b) comprises obtaining a genotyped list for the plurality of sites; optionally, wherein a file comprises the genotyped list, wherein the file comprises a format selected from a variable call file (VCF), binary alignment map (BAM), and a compressed reference-oriented alignment map (CRAM).

29. The method of claim 28, wherein obtaining the genotyped list comprises: (i) aligning the sequence information with a reference; and (ii) genotyping each site of the plurality of sites to generate the genotyped list.

30. The method of claim 29, further comprising removing multi-nucleotide variant SNPs (MNVS) from the genotyped list.

31. The method of any one of claims 28-30, further comprising encoding the genotyped list to obtain the first SIB, wherein each identifier is indicative that a site of the plurality of sites is: (i) homozygous to the reference, (ii) heterozygous to the reference, (iii) homozygous alternative to the reference, (iv) unexpected genotype, or (v) not determined; optionally, wherein the indicative identifier is 0, 1, 2, 3 or 4, respectively.

32. The method of any one of claims 25-31, further comprising receiving the sequence information for the second nucleic acid sample at the plurality of sites, and generating the second SIB.

33. The method of any one of claims 25-32, wherein step (c) comprises: (i) aligning the first SIB with the second SIB; (ii) comparing the same positions in each SIB with one another, excluding the same positions with an identifier in either SIB for unexpected genotype or for not determined; and (iii) determining a ratio for a number of matching identifiers at the same positions in each SIB to a number of the same positions compared, thereby generating the ratio of identity.

34. The method of claim 33, wherein the compared positions consist of positions for autosomal sites of the plurality of sites.

35. The method of claim 34, wherein the compared positions consist of positions with an identifier for homozygous to the reference, or for homozygous alternative to the reference.

36. The method of claim 34 or 35, further comprising detecting a kinship relationship between the nucleic acid samples, wherein each compared position has an identifier selected from: (i) a position in the first SIB has an identifier for homozygous to the reference, and the same position in the second SIB has an identifier for homozygous to the reference or for heterozygous to the reference; (ii) a position in the first SIB has an identifier for heterozygous to the reference, and the same position in the second SIB has an identifier for homozygous to the reference, for heterozygous to the reference or for homozygous alternative to the reference; and (iii) a position in the first SIB has an identifier for homozygous alternative to the reference, and the same position in the second SIB has an identifier for heterozygous to the reference or for homozygous alternative to the reference.

37. The method any one of claims 33-36, further comprising detecting a contaminant in a nucleic acid sample, wherein a SIB of a nucleic acid sample comprisingadditional nucleic acids comprises an increased number of positions with an identifier for heterozygous positions compared to a SIB of the nucleic acid sample without the additional nucleic acids.

38. The method of claim 37, wherein the detecting comprises: (i) generating a variant allele frequency (VAF) at each position of the first SIB and the second SIB; (ii) binning each VAF in a bin of width 0.01 in a plurality of bins having a range from 0 to 1; (iii) classifying a VAF with a value between 0 and 0.1 or between 0.9 to 1 as homozygous, and a VAF with a value between 0.1 and 0.9 as heterozygous; (iv) generating a contamination ratio by dividing the number of VAF in the homozygous class by the number of VAF in the heterozygous class; wherein a contamination ratio greater than a threshold indicates a contaminated nucleic acid sample; optionally, wherein the threshold is 1.

0.

39. The method of any one of claims 25-38, wherein step (d) comprises providing an indication that the first and second nucleic acid samples are a match.

40. The method of claim 39, wherein at least 45%, 50%, 55%, or 60% of the positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another; optionally, wherein at least 45% of the positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another.

41. The method of claim 39 or 40, wherein at least 40, 45, 50, 55, 60, 65, 70, or 75 positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another; optionally, wherein at least 50 positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another.

42. The method of claim 40 or 41, wherein the same positions of the first and second SIB are homozygous to a reference, heterozygous to the reference, or homozygous alternative to the reference.

43. The method of any one of claims 39-42, wherein the ratio of identity is at least 0.90, 0.95, or 0.99; optionally, wherein the ratio of identity is at least 0.99.

44. The method of any one of claims 25-43, further comprising providing an indication that the first or second nucleic acid sample comprises a Y chromosome, wherein at least 50% of sites of the plurality of sites located on the Y chromosome are identified.

45. A system for comparing nucleic acid samples, comprising a processor configured with instructions for the computer-implemented method of any one of claims 25- 44.

46. A system for comparing nucleic acid samples, comprising (a) a sequence identity barcode (SIB) generator module running on a processor and adapted to: (i) receive sequence information for a first nucleic acid sample at a plurality of sites, (ii) generate from the sequence information a first SIB for the first nucleic acid sample, wherein the first SIB comprises an ordered list of identifiers for the plurality of sites; and (b) a SIB comparator module adapted to: (i) generate a ratio of identity by comparing the first SIB and a second SIB, wherein the second SIB is generated from sequence information for a second nucleic acid sample, and (ii) provide an indication of a match between the first and second nucleic acid samples based on the comparison.

47. The system of claim 46, wherein the plurality of sites comprises: (i) a plurality of single nucleotide polymorphisms (SNPs); (ii) locations distributed throughout a genome; (iii) locations in exons; (iv) biallelic sites; (v) locations within ubiquitously expressed genes; (vi) lack allele specific expression; and / or (vii) locations comprising a minor allele frequency in African, South and East Asian and European populations in a range from 0.3 to 0.

7.

48. The system of claim 46 or 47, wherein the sequence information is obtained by the method of any one of claims 1-24.

49. The system of any one of claims 46-48, wherein the generate from the sequence information a first SIB comprises receiving a genotyped list for the plurality of sites; optionally, wherein a file comprises the genotyped list, wherein the file comprises a formatselected from a variable call file (VCF), binary alignment map (BAM), and a compressed reference-oriented alignment map (CRAM).

50. The system of claim 49, wherein receiving the genotyped list comprises: (i) aligning the sequence information with a reference; and (ii) genotyping each site of the plurality of sites to generate the genotyped list.

51. The system of claim 50, further comprising removing multi-nucleotide variant SNPs (MNVS) from the genotyped list.

52. The system of any one of claims 49-51, further comprising encoding the genotyped list to obtain the first SIB, wherein each identifier is indicative that a site of the plurality of sites is: (i) homozygous to the reference, (ii) heterozygous to the reference, (iii) homozygous alternative to the reference, (iv) unexpected genotype, or (v) not determined; optionally, wherein the indicative identifier is 0, 1, 2, 3 or 4, respectively.

53. The system of any one of claims 46-52, further comprising receiving the sequence information for the second nucleic acid sample at the plurality of sites, and generating the second SIB.

54. The system of any one of claims 46-53, wherein the generate a ratio of identity comprises: (i) aligning the first SIB with the second SIB; (ii) comparing the same positions in each SIB with one another, excluding the same positions with an identifier in either SIB for unexpected genotype or for not determined; and (iii) determining a ratio for a number of matching identifiers at the same positions in each SIB to a number of the same positions compared, thereby generating the ratio of identity.

55. The system of claim 54, wherein the compared positions consist of positions for autosomal sites of the plurality of sites.

56. The system of claim 55, wherein the compared positions consist of positions with an identifier for homozygous to the reference, or for homozygous alternative to the reference.

57. The system of claim 55 or 56, further comprising detecting a kinship relationship between the nucleic acid samples, wherein each compared position has an identifier selected from: (i) a position in the first SIB has an identifier for homozygous to the reference, and the same position in the second SIB has an identifier for homozygous to the reference or for heterozygous to the reference; (ii) a position in the first SIB has an identifier for heterozygous to the reference, and the same position in the second SIB has an identifier for homozygous to the reference, for heterozygous to the reference or for homozygous alternative to the reference; and (iii) a position in the first SIB has an identifier for homozygous alternative to the reference, and the same position in the second SIB has an identifier for heterozygous to the reference or for homozygous alternative to the reference.

58. The system any one of claims 54-57, further comprising detecting a contaminant in a nucleic acid sample, wherein a SIB of a nucleic acid sample comprising additional nucleic acids comprises an increased number of positions with an identifier for heterozygous positions compared to a SIB of the nucleic acid sample without the additional nucleic acids.

59. The system of claim 58, wherein the detecting comprises: (i) generating a variant allele frequency (VAF) at each position of the first SIB and the second SIB; (ii) binning each VAF in a bin of width 0.01 in a plurality of bins having a range from 0 to 1; (iii) classifying a VAF with a value between 0 and 0.1 or between 0.9 to 1 as homozygous, and a VAF with a value between 0.1 and 0.9 as heterozygous; (iv) generating a contamination ratio by dividing the number of VAF in the homozygous class by the number of VAF in the heterozygous class; wherein a contamination ratio greater than a threshold indicates a contaminated nucleic acid sample; optionally, wherein the threshold is 1.

0.

60. The system of any one of claims 46-59, further comprising a module for providing an indication that the first and second nucleic acid samples are a match.

61. The system of claim 60, wherein at least 45%, 50%, 55%, or 60% of the positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another; optionally, wherein at least 45% of the positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another.

62. The system of claim 60 or 61, wherein at least 40, 45, 50, 55, 60, 65, 70, or 75 positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another; optionally, wherein at least 50 positions in each of the first SIB and the second SIB for the same sites of the plurality of sites are available for comparison with one another.

63. The system of claim 61 or 62, wherein the same positions of the first and second SIB are homozygous to a reference, heterozygous to the reference, or homozygous alternative to the reference.

64. The system of any one of claims 60-63, wherein the ratio of identity is at least 0.90, 0.95, or 0.99; optionally, wherein the ratio of identity is at least 0.

99.

65. The system of any one of claims 46-64, further comprising providing an indication that the first or second nucleic acid sample comprises a Y chromosome, wherein at least 50% of sites of the plurality of sites located on the Y chromosome are identified.

66. A kit for confirming a nucleic acid sample identity, comprising: a plurality of pairs of primers, wherein the plurality of pairs of primers comprises at least 50 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 197-392. a DNA polymerase buffer; and a DNA polymerase.

67. The kit of claim 66, wherein the plurality of pairs of primers comprises at least 70 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.

68. The kit of claim 66 or claim 67, wherein the plurality of pairs of primers comprises at least 90 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.

69. The kit of any of claims 66-68, wherein the plurality of primers comprises pairs of primers comprising a sequence as set forth in each of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.

70. The kit of any of claims 66-69, wherein the DNA polymerase buffer comprises Illumina TTM buffer.

71. The kit of any of claims 66-70, wherein the DNA polymerase comprises Thermo Phusion Hot Start II DNA Polymerase.

72. A system for confirming a nucleic acid sample identity, comprising a processor configured to perform a method comprising: receiving sequence information for a nucleic acid sample at a plurality of sites, wherein the sites comprise 20 or more, 30 or more, 40 or more 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or each of the sites corresponding to positions: chr1:6693097, chr1:41204569, chr1:59147926, chr1:114515717, chr1:150808889, chr1:212911836, chr1:227069737, chr1:227071525, chr1:227935444, chr1:236716959, chr1:236719193, chr2:3392295, chr2:101624471, chr2:101638888, chr2:128939817, chr3:4403767, chr3:15737689, chr3:33434831, chr3:155481609, chr3:186509517, chr3:193374964, chr4:6600012, chr4:6606864, chr4:25419283, chr5:53815495, chr5:64881936, chr5:82834630, chr5:96503523, chr6:26598188, chr6:49403282, chr6:49425521, chr6:101166095, chr6:146755140, chr6:158517308, chr7:2577781, chr7:2578237, chr7:43916727, chr7:43917013, chr7:48004962, chr8:30973957, chr8:33356074, chr8:33369944, chr8:90995019, chr8:94935937, chr8:104427359, chr8:124448736, chr8:132982824, chr8:135612745, chr8:146067054, chr9:132400480, chr10:1046712, chr10:31138817, chr10:99219885, chr10:100219314, chr10:104572963, chr11:16133413, chr11:125763746, chr12:993930, chr12:62926398, chr13:28143229, chr13:39433606, chr14:20920250, chr14:31381351, chr15:34528948, chr15:75131959, chr16:624114, chr16:2028402, chr16:19680546, chr16:70303580, chr17:10614442, chr17:17168164, chr17:40714804, chr17:40716520, chr17:57290383, chr17:71196809, chr17:79514129, chr17:80008392, chr18:21413869, chr19:885818, chr19:1010691, chr19:36727365, chr19:41939297, chr19:50983930, chr19:58929052, chr20:6100088, chr20:61834695, chr21:47961711, chr22:32795641, chr22:32875190,chr22:50962208, chrY:2961349, chrY:7063980, chrY:2866789, chrY:2980414, chrY:7091532, chrY:13705715, chrY:14842155, or chrY:20592667 of hg 19 - GRCh37, and generating a sample identity barcode (SIB) comprising sequence information at each of the plurality of sites.

73. The system of claim 72, wherein the method further comprises comparing the SIBs of at least two nucleic acid samples of the one or more nucleic acid samples.

74. The system of claim 72 or claim 73, wherein the sequence information comprises whether the nucleic acid sample is homozygous for a reference allele, homozygous for an alternative allele, or heterozygous at each site of the plurality of sites.

75. A method of amplifying a human nucleic acid sample, comprising: contacting DNA or cDNA from a nucleic acid sample with a plurality of pairs of primers, a DNA polymerase buffer, and a DNA polymerase, wherein the DNA polymerase buffer comprises Illumina TTM buffer, and wherein the DNA polymerase comprises Thermo Phusion Hot Start II DNA Polymerase, and amplifying the DNA or cDNA at a plurality of sites.

76. The method of claim 75, wherein the plurality of sites comprises 20 or more, 30 or more, 40 or more 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or each of the sites corresponding to positions: chr1:6693097, chr1:41204569, chr1:59147926, chr1:114515717, chr1:150808889, chr1:212911836, chr1:227069737, chr1:227071525, chr1:227935444, chr1:236716959, chr1:236719193, chr2:3392295, chr2:101624471, chr2:101638888, chr2:128939817, chr3:4403767, chr3:15737689, chr3:33434831, chr3:155481609, chr3:186509517, chr3:193374964, chr4:6600012, chr4:6606864, chr4:25419283, chr5:53815495, chr5:64881936, chr5:82834630, chr5:96503523, chr6:26598188, chr6:49403282, chr6:49425521, chr6:101166095, chr6:146755140, chr6:158517308, chr7:2577781, chr7:2578237, chr7:43916727, chr7:43917013, chr7:48004962, chr8:30973957, chr8:33356074, chr8:33369944, chr8:90995019, chr8:94935937, chr8:104427359, chr8:124448736, chr8:132982824, chr8:135612745, chr8:146067054, chr9:132400480, chr10:1046712, chr10:31138817, chr10:99219885, chr10:100219314, chr10:104572963, chr11:16133413, chr11:125763746, chr12:993930, chr12:62926398, chr13:28143229, chr13:39433606, chr14:20920250, chr14:31381351,chr15:34528948, chr15:75131959, chr16:624114, chr16:2028402, chr16:19680546, chr16:70303580, chr17:10614442, chr17:17168164, chr17:40714804, chr17:40716520, chr17:57290383, chr17:71196809, chr17:79514129, chr17:80008392, chr18:21413869, chr19:885818, chr19:1010691, chr19:36727365, chr19:41939297, chr19:50983930, chr19:58929052, chr20:6100088, chr20:61834695, chr21:47961711, chr22:32795641, chr22:32875190, chr22:50962208, chrY:2961349, chrY:7063980, chrY:2866789, chrY:2980414, chrY:7091532, chrY:13705715, chrY:14842155, or chrY:20592667 of hg 19 - GRCh37.

77. The method of claim 75 or claim 76, wherein the plurality of pairs of primers comprises at least 10, at least 20, at least 30, at least 40, or at least 50 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.

78. The method of any of claims 75-77, wherein the plurality of pairs of primers comprises at least 70 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.

79. The method of any of claims 75-78, wherein the plurality of pairs of primers comprises at least 90 pairs of primers comprising a sequence as set forth in any of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.

80. The method of any of claims 75-79, wherein the plurality of pairs of primers comprises pairs of primers comprising a sequence as set forth in each of SEQ ID NOs: 01-196; optionally, a sequence as set forth in any of SEQ ID NOs: 197-392.