Methods, compositions and systems for identifying variant target molecules of interest

A method using recognition elements and concatemeric amplification with soft decision decoding addresses the inefficiencies of existing genetic variant detection, enabling cost-effective identification of pharmacogenomic targets for personalized medicine.

WO2026005800A1PCT designated stage Publication Date: 2026-01-02PLENO INC

Patent Information

Application Number
PCT/US2024/038517
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-27
Filing Date
2024-07-18
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing genetic variant detection technologies, such as next-generation sequencing and microarrays, are time and cost prohibitive, hindering the widespread adoption of pharmacogenomics and other fields that require comprehensive genetic variation analysis.

Method used

A method involving hybridization of biological samples to recognition elements with complementary sequences, followed by ligation, circularization, concatemeric amplification, and soft decision decoding to identify variant targets, enabling a streamlined and cost-effective workflow.

Benefits of technology

This approach allows for efficient and cost-effective identification of genetic variants, including pharmacogenomic targets, overcoming the limitations of existing technologies and facilitating personalized medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024038517_02012026_PF_FP_ABST
    Figure US2024038517_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides methods, compositions and systems for identifying variant targets of interest, for example, in clinically actionable genes. Disclosed herein are representative assays for identifying gene variants, for example, that are implicated in one or more drug metabolism pathways. The methods, composition and systems disclosed herein enable a highly streamlined and cost-effective workflow for moving forward the emerging field of personalized medicine and related fields of study such as pharmacogenomics.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] WSGR Docket No. 64100-744.601 METHODS, COMPOSITIONS AND SYSTEMS FOR IDENTIFYING VARIANT TARGET MOLECULES OF INTEREST CROSS-REFERENCE

[0001] This application claims the benefit of United States Provisional Patent Application Serial No.63 / 665,164, filed June 27, 2024, which is incorporated herein by reference in its entirety. INCORPORATION BY REFERENCE OF SEQUENCE LISTING

[0002] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing, which is submitted in XML format, is provided as a file entitled 64100_744_601_SL.xml, created on July 9, 2024, which is 3,148 bytes in size. The information in the electronic format of the Sequence Listing is incorporated by reference in its entirety. SUMMARY

[0003] The human genome is highly variable, and in some estimates, approximately four million variants can be found in the human genome, providing a genome of vast diversity. The frequency of genetic variants can depend on a multitude of factors, such as a subject’s ethnicity or a disease state, such that the genome of one subject may differ from another subject due to genetic variations. The variability in a subject’s genome can, in part, determine a subject’s response and / or susceptibility to diseases, responses to disease related therapies, and thus, clinical outcomes on a personal level.

[0004] As such, discerning a subject’s genetic makeup, such as identifying genetic variants that may or may not affect a subject’s ability to respond to disease, is an important field of study. Pharmacogenomics, or PGx, is the study of how a subject’s genetics can affect a subject’s response to drugs, for example how a subject’s chemistry and physiology metabolizes a given drug. To determine how a subject might respond to a drug, the genetic makeup of that subject can be analyzed to predict drug efficacy and any risk of adverse reactions a subject might have to a drug. One goal of PGx is to develop personalized medical initiatives and rational ways to optimize drug therapies by selecting medications and doses tailored to a subject’s genetic makeup. As drugs are metabolized by enzymes encoded by genes, a logical place to start identifying PGx targets that might affect a subject’s response to drugs is by studying genes that code for such drug metabolizing proteins.

[0005] Polymorphisms, or variations, in drug metabolizing enzymes can affect the efficacy and toxicity of a drug. For example, variants in the genes CYP2D6, CYP2C19 and CYP2C9 can WSGR Docket No. 64100-744.601 alter metabolism of many commonly prescribed drugs. The clinical implementation of PGx has the potential to improve patient care by maximizing therapeutic benefits while minimizing risks and making precision medicine a reality by enabling physicians to prescribe drugs focused on a patient’s specific genetic profile.

[0006] As such, studying variability of a genome, on a personal health level, is of high importance. Assays for doing so may require the use of multiple testing methodologies to provide a comprehensive view of a subject’s genetic variation. In addition, existing detection technologies such as next generation sequencing (NGS) or microarrays can be time and / or cost prohibitive.

[0007] As studying drug metabolism and genetic variants associated with drug metabolism has the potential to revolutionize the practice of medicine and clinical drug development through genomic based personalized medicine, what is needed are tools to support this highly impactful emerging field of study, as well as other fields of study that identify, on a personal level, a subject’s genetic variations, that provide alternatives to existing technologies.

[0008] The present disclosure provides methods, compositions and systems for identifying targets of interest associated with genetic variants in a genome, as an example pharmacogenomic relevant genes. Disclosed herein are assays for identifying gene variants of genes. Some of the genetic variants may be implicated in affecting, either positively or negatively, one or more drug metabolism pathways. The methods, composition and systems disclosed herein enable a highly streamlined, cost-effective workflow that overcomes many of the issues that have been holding back the widespread adoption of genetic variant detection in emerging fields such as pharmacogenomics.

[0009] Aspects disclosed herein provide methods for identifying one or more variant targets of interest from a biological sample, comprising: a) providing a biological sample suspected of having the one or more target variants of interest, b) hybridizing the biological sample to a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements comprises a 5’ and 3’ end, wherein the 3’ end comprises a complementary sequence to a variant target of interest of the one or more variant targets of interest, and wherein each recognition element of the plurality of recognition elements further comprises a code that is a proxy for the variant target of interest, c) ligating a recognition element that is hybridized to the biological sample to generate a circularized and ligated recognition element, d) generating a concatemeric amplification product from the circularized and ligated recognition element, and e) decoding the code from the concatemeric amplification product using soft decision decoding, thereby identifying the presence or absence of the one or more variant targets of interest from the biological sample. In some embodiments, the method further comprising hybridizing the WSGR Docket No. 64100-744.601 biological sample to a plurality of second recognition elements, wherein each recognition element of the plurality of second recognition elements comprises a 3’ end that is complementary to a wildtype nucleic acid of the one or more variant targets of interest, and wherein the recognition element that comprises a 3’ end that is complementary to the wildtype nucleic acid of the one or more variant targets of interest comprises the same code as the recognition element that comprises the 3’ end that comprises the complementary sequence to the one or more variant targets of interest. In some embodiments, the biological sample is from a blood sample, a buccal swab sample, a saliva sample, or a tissue sample. In further embodiments, the biological sample is an extracted and purified nucleic acid sample from the blood sample, the buccal swab, the saliva sample or the tissue sample. In some embodiments, the one or more variants of interest include a single nucleotide polymorphism, an insertion, a deletion and a copy number variant. In some embodiments, the one or more target variants of interest comprise from 10-1000 target variants of interest. In some embodiments, the one or more target variants of interest comprises at least 10 variants of interest, at least 100 variants of interest, or at least 1000 variants of interest. In some embodiments, the one or more target variants of interest comprise one or more pharmacogenomics target variants of interest, including, but not limited to, one or more star alleles. In some embodiments, the one or more target variants of interest comprise one or more target variants of interest from a CYP2D6 gene in the biological sample. In some embodiments, the one or more target variants of interest comprise one or more target variants of interest from one or more HLA genes in the biological sample. In some embodiments, the target variants of interest is from one or more genes selected from the group of ABCB1, ABCG2, ADRA2A, ALDH2, ANK3, ANKK1, APOE, BDNF, C11orf65, CACNA1C, CACNA1S, CEP72, CFTR, COMT, CYP1A2, CYP2C, CYP2C19, CYP2C9, CYP2D6, DBH, DPYD, DRD2, F2, F5, G6PD, GLP1R, GRIK1, GRIK4, GRIN2B, HTR2A, HTR2C, IFNL3, IFNL4, MC4R, MT-RNR1, MTHFR, OPRM1, PNPLA5, RYR1, SCL47A2, SLC6A4, SLCO1B1, SULT4A1, UGT1A1, UGT2B15, VKORC, YEATS4, CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, CYP4F2, DPYD, NUDT15, SLCO1B1, TPMT, and UGT1A1. In other embodiments, the one or more target variants of interest is from one or more genes selected from the group of ABCB1, ABCG2, ADRA2A, ALDH2, ANK3, ANKK1, APOE, BDNF, C11orf65, CACNA1C, CACNA1S, CEP72, CFTR, COMT, CYP1A2, CYP2C, CYP2C19, CYP2C9, CYP2D6, DBH, DPYD, DRD2, F2, F5, G6PD, GLP1R, GRIK1, GRIK4, GRIN2B, HTR2A, HTR2C, IFNL3, IFNL4, MC4R, MT-RNR1, MTHFR, OPRM1, PNPLA5, RYR1, SCL47A2, SLC6A4, SLCO1B1, SULT4A1, UGT1A1, UGT2B15, VKORC, and YEATS4. In some embodiments, the one or more target variants of interest comprise star allele variants from a gene of interest, WSGR Docket No. 64100-744.601 for example star alleles associated with CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, CYP4F2, DPYD, NUDT15, SLCO1B1, TPMT, UGT1A1. In some embodiments, a cytochrome P450 gene includes one or more of CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, and CYP4F2. In some embodiments, the one or more target variants of interest are selected from the one or more variants listed in Table 1 and Table 2. In some embodiments, the method comprises ligating the recognition elements comprising use of a thermostable ligase. In some embodiments, the method comprises generating the concatemeric amplification products comprises rolling circle amplification or multiple strand displacement. In some embodiments, the method comprises treating the circularized and ligated recognition element with an exonuclease prior to generating the concatemeric amplification product. In some embodiments, the circularized and ligated recognition elements are placed on a substrate prior to generating the concatemeric amplification products. In some embodiments, the substrate is one or more wells or one or more tubes that have been pre-treated with a nucleic acid immobilization composition, for example one or more wells of a 24, 48, 96 or 384 plate and each of the one or more wells is pre-treated with a polymer. In some embodiments, the polymer is selected from polyacrylamide, branched PEI, linear PEI, poly(β-aminoester) and poly(amidoamine), PEG, a gel, poly-L-lysine, silane, agarose, and muscle mimetic catecholamine polymer. In some embodiments, the method comprises decoding the code from the concatemeric amplification product comprises next generation sequencing or detection by hybridization. In some embodiments, the code comprises at least two nucleic acid segments, wherein the at least two nucleic acid segments are detected by hybridizing one or more detection polynucleotides to that at least two nucleic acid segments and imaging the hybridizing. In some embodiments, each of the one or more detection polynucleotides are labeled with a different fluorescent moiety.

[0010] Aspects disclosed herein provide methods for identifying pharmacogenomic targets of interest from a biological sample, comprising a) providing a biological sample suspected of having one or more pharmacogenomic target variants of interest, b) hybridizing the biological sample to a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements comprises a 5’ and 3’ end, wherein the 3’ end comprises a complementary sequence to a variant targets of interest of the one or more pharmacogenomic variant targets of interest, and wherein each recognition element of the plurality of recognition elements further comprises a code that is a proxy for the one or more pharmacogenomic variant targets of interest, c) ligating a recognition element that is hybridized to the biological sample to generate a circularized and ligated recognition element, d) generating a concatemeric amplification product from the circularized and ligated recognition element, e) decoding the WSGR Docket No. 64100-744.601 code from the concatemeric amplification product using soft decision decoding, thereby identifying the presence or absence of the one or more pharmacogenomic variant targets of interest from the biological sample. In some embodiments, the method further comprises hybridizing the biological sample to a plurality of second recognition elements, wherein each recognition element of the plurality of second recognition elements comprises a 3’ end that is complementary to a wildtype nucleic acid of the one or more pharmacogenomic variant targets of interest, and wherein the recognition element that comprises a 3’ end that is complementary to the wildtype nucleic acid of the one or more pharmacogenomic variant targets of interest comprises the same code as the recognition element that comprises the 3’ end that comprises the complementary sequence to the one or more pharmacogenomic variant targets of interest. In some embodiments, the biological sample is from a blood sample, a buccal swab sample, a saliva sample or a tissue sample, wherein the biological sample is an extracted and purified nucleic acid sample from the blood sample, the buccal swab sample, the saliva sample or the tissue sample. In some embodiments, one or more pharmacogenomic variants of interest include a single nucleotide polymorphism, an insertion, a deletion and a copy number variant. In some embodiments, the one or more pharmacogenomic target variants of interest comprise from 10- 1000 pharmacogenomic target variants of interest, at least 10 pharmacogenomic variants of interest, at least 100 pharmacogenomic variants of interest, or at least 1000 pharmacogenomic variants of interest. In some embodiments, the one or more pharmacogenomics target variants of interest comprise one or more star alleles. In some embodiments, the one or more pharmacogenomic target variants of interest comprise one or more pharmacogenomic target variants of interest from a CYP2D6 gene. In some embodiments, the one or more pharmacogenomic target variants of interest comprise one or more target variants of interest from one or more HLA genes. In some embodiments, the one or more pharmacogenomic target variants of interest are from one or more genes selected from the group comprising ABCB1, ABCG2, ADRA2A, ALDH2, ANK3, ANKK1, APOE, BDNF, C11orf65, CACNA1C, CACNA1S, CEP72, CFTR, COMT, CYP1A2, CYP2C, CYP2C19, CYP2C9, CYP2D6, DBH, DPYD, DRD2, F2, F5, G6PD, GLP1R, GRIK1, GRIK4, GRIN2B, HTR2A, HTR2C, IFNL3, IFNL4, MC4R, MT-RNR1, MTHFR, OPRM1, PNPLA5, RYR1, SCL47A2, SLC6A4, SLCO1B1, SULT4A1, UGT1A1, UGT2B15, VKORC, YEATS4, CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, CYP4F2, DPYD, NUDT15, SLCO1B1, TPMT, and UGT1A1. In some embodiments, the one or more pharmacogenomic target variants of interest is from one or more genes selected from the group comprising ABCB1, ABCG2, ADRA2A, ALDH2, ANK3, ANKK1, APOE, BDNF, C11orf65, CACNA1C, CACNA1S, CEP72, CFTR, COMT, CYP1A2, CYP2C, CYP2C19, CYP2C9, CYP2D6, DBH, WSGR Docket No. 64100-744.601 DPYD, DRD2, F2, F5, G6PD, GLP1R, GRIK1, GRIK4, GRIN2B, HTR2A, HTR2C, IFNL3, IFNL4, MC4R, MT-RNR1, MTHFR, OPRM1, PNPLA5, RYR1, SCL47A2, SLC6A4, SLCO1B1, SULT4A1, UGT1A1, UGT2B15, VKORC, and YEATS4. In some embodiments, the one or more pharmacogenomic target variants of interest comprise star allele variants from a cytochrome P450 gene such as CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, and CYP4F2. In some embodiments, the one or more pharmacogenomic target variants of interest comprise star allele variants from one or more genes selected from the group comprising CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, CYP4F2, DPYD, NUDT15, SLCO1B1, TPMT, UGT1A1. In some embodiments, the one or more pharmacogenomic target variants of interest are selected from the one or more variants listed in Table 1 and Table 2. In some embodiments, the method comprises ligating the recognition element using a thermostable ligase. In some embodiments, the method comprises generating the concatemeric amplification product using rolling circle amplification or multiple strand displacement. In some embodiments, the method comprises treating the circularized and ligated recognition element with an exonuclease prior to generating the concatemeric amplification product. In some embodiments, the circularized and ligated recognition elements are placed on a substrate prior to generating the concatemeric amplification products, wherein the substrate is one or more wells or one or more tubes that have been pre-treated with a nucleic acid immobilization composition such as a polymer selected from polyacrylamide, branched PEI, linear PEI, poly(β-aminoester) and poly(amidoamine), PEG, a gel, poly-L-lysine, silane, agarose, and muscle mimetic catecholamine polymer. In some embodiments, the method comprises decoding the code from the concatemeric amplification product comprises next generation sequencing or detection by hybridization. In some embodiments, the code comprises at least two nucleic acid segments. In some embodiments, the at least two nucleic acid segments are detected by hybridizing one or more detection polynucleotides to that at least two nucleic acid segments and imaging the hybridizing, wherein each of the one or more detection polynucleotides are labeled with a different fluorescent moiety.

[0011] Aspects disclosed herein provide methods for identifying the presence of one or more star alleles in a genome, comprising a) providing a biological sample suspected of having one or more star alleles, b) hybridizing the biological sample to a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements comprises a 5’ and 3’ end, wherein the 3’ end comprises a complementary sequence to a star allele of the one or more star alleles, and wherein each recognition element of the plurality of recognition elements further comprises a code that is a proxy for the one or more star alleles, c) ligating a recognition element that is hybridized to the biological sample to generate a circularized and ligated WSGR Docket No. 64100-744.601 recognition element, d) generating a concatemeric amplification product from the circularized and ligated recognition element, e) decoding the code from the concatemeric amplification product using soft decision decoding, thereby identifying the presence or absence of the one or more star alleles from the biological sample. In some embodiments, the method further comprising hybridizing the biological sample to a plurality of second recognition elements, wherein each recognition element of the plurality of second recognition elements comprises a 3’ end that is complementary to a wildtype sequence of the one or more star alleles, and wherein the recognition element that comprises a 3’ end that is complementary to the wildtype nucleic acid of the one or more star alleles comprises a different code as the recognition element that comprises the 3’ end that comprises the complementary sequence to the one or more star alleles. In some embodiments, one or more star alleles comprise a single nucleotide polymorphism, an insertion, a deletion and a copy number variant. In some embodiments, the one or more star alleles comprise from 10-1000 star alleles, or at least 10 star alleles, at least 100 star alleles, at least 1000 star alleles. In some embodiments, the one or more star alleles are present in a CYP2D6 gene from the biological sample. In some embodiments, the one or more star alleles are present in one or more HLA genes from the biological sample. In some embodiments, the one or more star alleles are from a cytochrome P450 gene from the biological sample such as CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, and CYP4F2. In some embodiments, the one or more star alleles are from one or more genes selected from the group comprising CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, CYP4F2, DPYD, NUDT15, SLCO1B1, TPMT, UGT1A1. In some embodiments, the one or more star alleles are selected from star alleles listed in Table 2.

[0012] Aspects disclosed herein provide methods for determining the presence of a genetic variant that alters drug metabolism, comprising a) providing a biological sample suspected of having one or more genetic variants that alter drug metabolism, b) hybridizing the biological sample to a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements comprises a 5’ and 3’ end, wherein the 3’ end comprises a complementary sequence to one of the one or more genetic variants, and wherein each recognition element of the plurality of recognition elements further comprises a code that is a proxy for the one or more genetic variants, c) ligating a recognition element that is hybridized to the biological sample to generate a circularized and ligated recognition element, d) generating a concatemeric amplification product from the circularized and ligated recognition element, e) decoding the code from the concatemeric amplification product using soft decision decoding, thereby identifying the presence or absence of the one or more genetic variants that alter drug metabolism from the biological sample. In some embodiments, the method further comprises WSGR Docket No. 64100-744.601 hybridizing the biological sample to a plurality of second recognition elements, wherein each recognition element of the plurality of second recognition elements comprises a 3’ end that is complementary to a wildtype sequence of the one or more genetic variants that alter drug metabolism, and wherein the recognition element that comprises a 3’ end that is complementary to the wildtype sequence of the one or more genetic variants comprises a different code as the recognition element that comprises the 3’ end that comprises the complementary sequence to the one or more genetic variants that alter drug metabolism. In some embodiments, the one or more genetic variants can alter the drug metabolism of one or more of Metoprolol, Propranolol, Timolol, Encainide, Flecainide, Perhexilene, Propafenone, Sparteine, Amitriptyline, Clomipramine, Desipramine, Fluoxetine, Fluvoxamine, Imipramine, Mianserin, Nortriptyline, Paroxetine, Venlafexine, Haloperidol, Perphenazine, Risperidone, Thioridazine, Zuclopenthixol, Codeine, Debrisoquine, Dextromethorphan, Phenoformin, Tolterodine and Tramadol. In some embodiments, the method may be used for determining the predisposition of altered drug metabolism in an individual.

[0013] Aspects disclosed herein provide methods for determining the presence of genetic variants in HLA genes, comprising a) providing a biological sample suspected of having one or more genetic variants in one or more HLA genes, b) hybridizing the biological sample to a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements comprises a 5’ and 3’ end, wherein the 3’ end comprises a complementary sequence to one of the one or more genetic variants, and wherein each recognition element of the plurality of recognition elements further comprises a code that is a proxy for the one or more genetic variants, c) ligating the recognition element that is hybridized to the biological sample to generate a circularized and ligated recognition element, d) generating a concatemeric amplification product from the circularized and ligated recognition element, e) decoding the code from the concatemeric amplification product using soft decision decoding, thereby identifying the presence or absence of the one or more genetic variants in the one or more HLA genes from the biological sample. In some embodiments, the method further comprises hybridizing the biological sample to a plurality of second recognition elements, wherein each recognition element of the plurality of second recognition elements comprises a 3’ end that is complementary to a wildtype sequence of the one or more genetic variants of the one or more HLA genes, and wherein the recognition element that comprises a 3’ end that is complementary to the wildtype sequence of the one or more genetic variants comprises a different code as the recognition element that comprises the 3’ end that comprises the complementary sequence to the one or more genetic variants of the one or more HLA genes. In some embodiments, the biological sample is from a blood sample, a buccal swab sample, a saliva sample or a tissue WSGR Docket No. 64100-744.601 sample that is an extracted and purified nucleic acid sample from the blood sample, the buccal swab sample, the saliva sample or the tissue sample. In some embodiments, the one or more genetic variants include a single nucleotide polymorphism, an insertion, a deletion and a copy number variant. In some embodiments, the one or more genetic variants comprise from 10-1000 genetic variants. In some embodiments, the method comprises ligating the recognition element comprises use of a thermostable ligase. In some embodiments, the method comprises generating the concatemeric amplification product using rolling circle amplification or multiple strand displacement. In some embodiments, the method further comprises treating the circularized and ligated recognition element with an exonuclease prior to generating the concatemeric amplification product. In some embodiments, the circularized and ligated recognition elements are placed on a substrate prior to generating the concatemeric amplification products, wherein the substrate is one or more wells or one or more tubes that have been pre-treated with a nucleic acid immobilization composition. In some embodiments, the substrate is one or more wells of a plate and each of the one or more wells is pre-treated with a polymer such as polyacrylamide, branched PEI, linear PEI, poly(β-aminoester) and poly(amidoamine), PEG, a gel, poly-L-lysine, silane, agarose, and muscle mimetic catecholamine polymer. In some embodiments, the decoding of the code from the concatemeric amplification product comprises next generation sequencing or detection by hybridization. In some embodiments, the code comprises at least two nucleic acid segments and the at least two nucleic acid segments are detected by hybridizing one or more detection polynucleotides to that at least two nucleic acid segments and imaging the hybridizing. In some embodiments, each of the one or more detection polynucleotides are labeled with a different fluorescent moiety.

[0014] Aspects disclosed herein provide methods for identifying a variant in a gene of interest in the presence of a high homology homolog, comprising: a) providing a biological sample suspected of having the variant in the gene of interest and the high homology homolog, b) hybridizing the biological sample to; i) a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements comprises a 5’ and 3’ end, wherein the 5’ end of a recognition element comprises a distinguishing nucleotide which is present in the gene of interest but not present in the high homology homolog and the 3’ end comprises a variant sequence to the variant in the gene of interest, and wherein each recognition element of the plurality of recognition elements further comprises a code that is a proxy for the variant of interest in the gene of interest, and ii) a third oligonucleotide, c) ligating the recognition elements that are hybridized to the biological sample to generate circularized and ligated recognition elements, wherein the recognition elements that are hybridized to the gene of interest are capable of being ligated compared to the high homology homolog that is associated WSGR Docket No. 64100-744.601 with the recognition elements which are not preferentially ligated, d) generating concatemeric amplification products from the circularized and ligated recognition elements, and e) decoding the code from the concatemeric amplification products using soft decision decoding, thereby identifying the presence or absence of the variant in the gene of interest from the biological sample in the presence of the high homology homolog. In some embodiments, the biological sample is from a blood sample, a buccal swab sample, a saliva sample or a tissue sample, and wherein the biological sample is an extracted and purified nucleic acid sample from the blood sample, the buccal swab sample, the saliva sample or the tissue sample. In some embodiments, the variant is a single nucleotide polymorphism, an insertion, a deletion and a copy number variant. In some embodiments, the ligating of the recognition elements comprises use of a thermostable ligase. In some embodiments, the generating of the concatemeric amplification products comprises rolling circle amplification or multiple strand displacement. In some embodiments, the method further comprises treating the circularized and ligated recognition elements with an exonuclease prior to generating the concatemeric amplification products. In some embodiments, the circularized and ligated recognition elements are placed on a substrate prior to generating the concatemeric amplification products. In some embodiments, the substrate is one or more wells or one or more tubes that have been pre-treated with a nucleic acid immobilization composition. In some embodiments, the substrate is one or more wells of a plate and each of the one or more wells is pre-treated with a polymer, selected from polyacrylamide, branched PEI, linear PEI, poly(β-aminoester) and poly(amidoamine), PEG, a gel, poly-L-lysine, silane, agarose, and muscle mimetic catecholamine polymer. In some embodiments, the decoding of the code from the concatemeric amplification products comprises next generation sequencing or detection by hybridization. In some embodiments, the code comprises at least two nucleic acid segments which are detected by hybridizing one or more detection polynucleotides to that at least two nucleic acid segments and imaging the hybridizing. In some embodiments, the one or more detection polynucleotides are labeled with a different fluorescent moiety.

[0015] Aspects disclosed herein provide compositions comprising a nucleic acid sample, wherein the nucleic acid sample comprises sequences to a gene of interest and a homolog to the gene of interest, wherein the gene of interest is hybridized to a recognition element comprising a differentiating base on the 5’ end and a variant nucleotide, or complement thereof, on the 3’ end, and a third oligonucleotide and the homolog is not hybridized to the differentiating base. In some embodiments, the present disclosure comprises a kit comprising a) a plurality of linear recognition elements, wherein each recognition element in the plurality comprises a 3’ end complementary to a genetic variant or the wildtype sequence of the genetic variant of a pharmacogenomics related gene target, b) one or more of a ligase, a DNA polymerase and an WSGR Docket No. 64100-744.601 exonuclease, c) a plurality of detection polynucleotides, d) optionally a third oligonucleotide, and d) instructions for practicing any of the methods disclosed herein. In some embodiments, the present disclosure provides a system for practicing any of the methods disclosed herein.

[0016] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative examples of the present disclosure are shown and disclosed. As will be realized, the present disclosure is capable of other and different examples, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. INCORPORATION BY REFERENCE

[0017] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The novel features of the inventive concepts are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present inventive concepts will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the inventive concepts are utilized, and the accompanying drawings of which:

[0019] Fig.1 is an example of a recognition element used in methods of the present disclosure.

[0020] Fig.2 is an example of a workflow for generating concatemeric amplification products used in methods of the present disclosure.

[0021] Fig.3A is an example of a detection polynucleotide.

[0022] Fig.3B is an example of a detection polynucleotide hybridized to an amplification product.

[0023] Fig.4 is a schematic diagram or an example of a soft decision decoding workflow for decoding a code of a recognition element.

[0024] Fig.5 is an example instrument that comprises a computer system for use with the methods described herein. WSGR Docket No. 64100-744.601

[0025] Fig.6 is an example of an application provision system for use with the methods described herein.

[0026] Fig.7 is an example of an application provision system for use with the methods described herein.

[0027] Fig.8 shows examples of recognition element designs demonstrating 5’ and 3’ end configurations for identifying target SNPs, deletions and insertions.

[0028] Fig.9 shows an example of a graph demonstrating high calls rates for plurality of SNVs and indels across a diversity of samples.

[0029] Fig.10 shows an example of a graph demonstrating the high degree of accuracy for the calls of Fig.9 which were highly reproducible when run independently across four different assay plates and two instruments.

[0030] Fig.11 shows an example of a graph demonstrating high call rates for a plurality of single nucleotide variants (SNVs) and indels for blood and saliva from one individual, gDNA served as the controls.

[0031] Fig.12 shows an example of a design for targeting recognition elements to seven different regions in the CYP2D6 gene for identifying copy number variations using recognition element constructs shown in Fig.8.

[0032] Fig.13A shows an example of a plot for CYP2D6 in different samples for normal complement of whole gene copy number. The dotted line represents WGS truth.

[0033] Fig.13B shows an example of a plot for CYP2D6 in different samples for whole gene deletion. The dotted line represents WGS truth.

[0034] Fig.13C shows an example of a plot for CYP2D6 in different samples for whole gene amplification following the design in Figs.14A-C. The dotted line represents WGS truth.

[0035] Fig.14A shows an example of a plot for CYP2D6 hybrid alleles in different samples for CYP2D6*10+*36 exon 9 deletion. The dotted line represents truth.

[0036] Fig.14B shows an example of a plot for CYP2D6 hybrid alleles in different samples for CYP2D6*10+*36 / *36 / *36 exon 9 deletion. The dotted line represents truth.

[0037] Fig.14C shows an example of a plot for CYP2D6 hybrid alleles in different samples for CYP2D6*68 intron 1 insertion. The dotted line represents truth.

[0038] Fig.14D shows an example of a plot for CYP2D6 hybrid alleles in different samples for CYP2D6*68 intron 1 insertion and CYP2D6*13 exon 9 deletion. The dotted line represents truth.

[0039] Fig.15A shows an example of a plot for hybrid allele CYP2D6*68 detection using the method of the present disclosure. The dotted line is WGS truth. WSGR Docket No. 64100-744.601

[0040] Fig.15B shows an example of a plot for hybrid allele CYP2D6*68 detection using a microarray assay. The dotted line is WGS truth.

[0041] Fig.16A shows an example of a recognition element design demonstrating 5’ and 3’ end configurations that include a gap oligonucleotide for a dual ligation scenario used in identifying a target single nucleotide variant in a region that shares homology with one or more homologs.

[0042] Fig.16B shows a graph showing the efficacy of the design in differentiating between the target variant of interest among one or more homologs.

[0043] Fig.17A shows an example of a plot demonstrating the efficacy of the single ligation recognition element shown in Fig.8.

[0044] Fig.17B shows an example of a plot demonstrating the efficacy dual ligation recognition element scenario from Fig.16A.

[0045] Fig.18A shows an example of a plot using methods disclosed herein for identifying MTHFR chr1:1794419_T>G.

[0046] Fig.18B shows and example of a plot using methods disclosed herein for identifying DPYD chr1:97883329_A>G.

[0047] Fig.18C shows an example of a plot using methods disclosed herein for identifying ABCG2 chr4:88131171_G>T.

[0048] Fig.18D shows an example of a plot using methods disclosed herein for identifying TPMT chr6:31386952_T>C.

[0049] Fig.18E shows an example of a plot using methods disclosed herein for identifying DBH chr9:133635393_T>C.

[0050] Fig.18F shows an example of a plot using methods disclosed herein for identifying GRIN2Bchr12:13800184_T>C.

[0051] Fig.18G shows an example of a plot using methods disclosed herein for identifying ABCB1 3435C>T.

[0052] Fig.18H shows an example of a plot using methods disclosed herein for identifying ADRA2A 1252G>C.

[0053] Fig.18I shows an example of a plot using methods disclosed herein for identifying VKORC1 -1639G>A.

[0054] Fig.18J shows an example of a plot using methods disclosed herein for identifying CYP3A5 981A>G.

[0055] Fig.18K shows an example of a plot using methods disclosed herein for identifying CYP1A2 -3860G>A.

[0056] Fig.18L shows an example of a plot using methods disclosed herein for identifying CYP2D6 1662G>C. WSGR Docket No. 64100-744.601

[0057] Fig.18M shows an example of a plot using methods disclosed herein for identifying CYP2D6 4402C>T.

[0058] Fig.18N shows an example of a plot using methods disclosed herein for identifying CYP2D6 4181G>C.

[0059] Fig.19A shows a design scenario for use of a third oligonucleotide in identifying more than one variant target in a sample where the variant nucleotide is located at the 5’ and 3’ ends of the recognition element.

[0060] Fig.19B shows a design scenario for use of a third oligonucleotide in identifying more than one variant target in a sample where one variant nucleotide is located at one arm of the recognition element and the second variant is located at the third oligonucleotide.

[0061] Fig.20 shows a design scenario for use of a third oligonucleotide in identifying more than one variant target in a sample, where the third oligonucleotide when hybridized to the target nucleic acid causes the target nucleic acid sequences between the two targets nucleotides of interest to loop out, thereby bringing two distant target nucleotide sequences in proximity.

[0062] Fig.21 shows examples of workflows for HLA imputation and HLA direct genotyping, which are combined in calling HLA variants.

[0063] Fig.22A show examples of graphs demonstrating HLA star allele calling and HLA imputation-based calling.

[0064] Fig.22B show examples of graphs demonstrating HLA star allele calling and HLA direct genotyping calling.

[0065] Fig.23 shows examples of graphs after correcting recognition elements for population variability at a variant of interest in the TNF gene.

[0066] Fig.24 shows scenarios for designing recognition elements for identifying phased variants in combination with a target of interest.

[0067] Fig.25 shows a recognition element design including a bridge element for identifying phase variants in CYP2D6 for genotype 130_132. SEQ ID NO: 1 in Fig.25 is the nucleic acid sequence for the antisense strand of CYP2D6 Exon 3. SEQ ID NO: 2 in Fig.25 is the nucleic acid sequence for the sense strand of CYP2D6 Exon 3.

[0068] Fig.26A shows an example of data when practicing the design scenario shown in Fig. 25.

[0069] Fig.26B shows an example of data when practicing the design scenario shown in Fig. 25.

[0070] Fig.26C shows an example of data when practicing the design scenario shown in Fig. 25. WSGR Docket No. 64100-744.601

[0071] Fig.26D shows an example of data when practicing the design scenario shown in Fig. 25.

[0072] Fig.26E shows an example of data when practicing the design scenario shown in Fig. 25.

[0073] Fig.26F shows an example of data when practicing the design scenario shown in Fig. 25. DETAILED DESCRIPTION

[0074] The present disclosure provides methods, compositions and systems for identifying targets of interest associated with pharmacogenomic related genes. Disclosed herein are assays for identifying gene variants of genes that are implicated in one or more drug metabolism pathways. The methods, composition and systems disclosed herein may enable a highly streamlined, cost-effective workflow that overcomes many of the issues that have been holding back the widespread adoption of pharmacogenomics.

[0075] The study of pharmacogenomics, or PGx, combines the study of drugs and the study of genes to develop effective, safe medications that can be provided to a patient based on the patient’s genetic makeup. Many drugs that are currently available are based on one size fits all. As every subject may respond to drugs differently, it may be difficult to predict which subjects will benefit from a drug and which subjects may not respond at all to the drug. Understanding how a subject will negatively react to a drug is important, wherein adverse drug reactions are a significant cause of hospitalizations and deaths. Pharmacogenomics is a growing field of study, and it is hoped that this field can be used to develop tailored drugs to treat a wide range of health problems, including, but not limited to, heart disease, neurological diseases, cancers, and respiratory diseases.

[0076] The present disclosure provides assays, methods, composition, systems and kits related to pharmacogenomics. I. METHODS Encoded assays

[0077] The methods and compositions as disclosed herein for determining a nucleic acid sequence may comprise use of an assay. In some embodiments, the assay is a solution-based assay. In some embodiments, the assay is a surface-bound assay. In some embodiments, the assay is a hybrid assay that includes a surface-bound component and a solution-based component. In some embodiments, the assay is performed in tubes or in a plate-based format, such as a multi-well plate, for example, a 96 well plate. In some embodiments, a multi-well plate WSGR Docket No. 64100-744.601 may include, for example, an array of nanowells. In some embodiments, the assay may be performed on a microfluidics device. In some embodiments, the assay may be performed partially in tubes and partially in a multi-well plate.

[0078] In some embodiments, a recognition element may be used in the assay. A recognition element used in the assay may include sequences that are complementary to a target sequence of interest, wherein the complementary sequence can hybridize a target sequence of interest, a code sequence that can be used to identify the target sequence of interest that has hybridized to its complement on the recognition element, and one or more functional sequences such as sequencing primer binding sites, one or more amplification primer binding sites, unique molecular identifier sequences (UMIs), sample indexes, or combinations thereof. In some embodiments, an amplification primer binding site may be adjacent to the code in a recognition element. The amplification primer binding site(s) may, in some cases, be universal primer sequence(s) that are common to all recognition elements in a set of recognition elements. Amplification primer binding site sequences may also be a code sequence or a portion thereof. A code sequence can be a combination of a number of subsequences, called nucleic acid segments, wherein their combination can identify a target sequence of interest that has hybridized to a recognition element. In some embodiments, an amplification primer can be a nucleic acid segment or a portion thereof. Unique identifier sequences (UMIs) and sample indexes, which oftentimes find utility in next generation sequencing reactions for counting, error correction and sample identification purposes, may also be part of code such as one or more nucleic acid segments.

[0079] In some embodiments, once a recognition element has recognized and hybridized to its target of interest, the recognition element may be circularized and ligated to generate a circular, ligated recognition element. The circular and ligated recognition element can then be amplified in anticipation of a decoding event to identify the code associated with the original target of interest that hybridized to the recognition element. Amplification may be by any method of amplification, including for example, nucleic acid extension, polymerase chain reaction (PCR), isothermal amplification, rolling circle amplification (RCA), and / or ultrarapid amplification. Surface based amplification may be performed using PCR with surface-anchored primers (e.g., Illumina® bridge amplification technology), or recombinase polymerase amplification (RPA) (e.g., ExAmp technology).

[0080] In one embodiment, the amplification operation comprises a rolling circle amplification (RCA) reaction to generate a concatemeric amplification products.

[0081] In one embodiment, a recognition element may include a sequence which may prevent RCA of the recognition element while allowing for linear double-stranded PCR products. The WSGR Docket No. 64100-744.601 non-extendable sequence may, for example, be located between a pair of amplification primer binding site sequences present on the recognition element. In one embodiment, a recognition element may include a restriction enzyme site that may be cleaved to yield a linear DNA molecule.

[0082] In some embodiments, a concatemeric amplification product may be sequenced to determine the nucleotide sequence of the code associated with the target molecule of interest. Any sequencing technology may be used to sequence the product. Non-limiting examples of sequencing technologies that may be used include sequencing by synthesis, avidity sequencing, sequencing by hybridization, sequencing by ligation, and nanopore sequencing.

[0083] In some embodiments, a sequencing library may be generated from a set of recognition elements or complements or amplicons thereof. The library may be sequenced to determine the code of the recognition element associated with a target molecule of interest. The code sequence may then be used as a digital count of the target molecule specific decoding event. In one embodiment, a sequencing library may be generated from a circularized recognition element. In another embodiment, a sequencing library may be generated from a concatemeric amplification product of a recognition element. In one embodiment, a concatemeric amplification product or a portion thereof that includes at least the code may be directly sequenced to determine the code associated with the target molecule of interest.

[0084] Fig.2 is a non-limiting example of an encoded assay for use with the detection polynucleotides disclosed herein. A linear recognition element 210 may comprise a 5' end 220a which is complementary to a portion of a target nucleic acid interest 222, a 3' end 220b which is complementary to another portion of a target nucleic acid of interest 222 from a sample, a code 216, and additional functional sequences 212, 214, 218 if desired such as amplification primer binding sites, capture sequencing, cleavage sites, sequencing primer binding sites, UMIs, and the like. A target nucleic acid of interest 222 which is complementary to the 5' 220a and 3' 220b ends of the linear recognition element 210 may hybridize to the linear recognition element, thereby bringing the ends in proximity for ligating to generate a circular and ligated recognition element 225. The circular and ligated recognition element 225 may be subjected to extension amplification using one of the functional sequences 212, 214, 218, or even the code 216 or a portion thereof, as a primer binding site. The result is a concatemeric amplification product 230 which can be decoded using the detection polynucleotides disclosed herein for identifying and determining the presence of the target nucleic acid of interest from a sample.

[0085] Additional examples of encoded assays can be found in WO2022 / 109496A2, which is incorporated herein by reference in its entirety. WSGR Docket No. 64100-744.601 Recognition elements

[0086] The methods and compositions described herein may include providing recognition elements to an encoded assay for identifying the presence of a target molecule of interest from a sample. In some embodiments, a plurality of recognition elements is provided. In some embodiments, each recognition element in the plurality of recognition elements comprises one or more target recognition regions. The target recognition regions of the recognition elements may comprise one or more nucleic acid sequence(s) configured to hybridize to a target nucleic acid molecule. In some embodiments, the one or more nucleic acid sequences hybridize to one or more target nucleic acid sequences of the target nucleic acid molecule. In some embodiments, the target recognition region is configured to hybridize to one or more regions of the target nucleic acid molecule flanking a target of interest (e.g., SNP, indel, and so on). In some embodiments, the target recognition region is configured to hybridize to a variant nucleic acid of interest (e.g., the target recognition region base pairs with the SNP when the target of interest is a SNP).

[0087] As shown in Fig.1, a recognition element may be used in encoded assays described herein and may comprise two target recognition regions, one at the 5’ prime end and another at the 3’ end. A recognition element further comprises a code. In Fig.1, the example code may be made up of four nucleic acid segments. However, the number of nucleic acid segments, be it one or more than one, is not limiting and the number of segments depends on the complexity of the assay (e.g., how many targets of interest are to be identified from a sample). In some embodiments, a target nucleic acid molecule of interest itself comprises the code, or an additional code such as a barcode that identifies another target molecule of interest such as a protein. In either embodiment, the code may be detected as a proxy for the target molecule, be it a nucleic acid or a protein, or both.

[0088] In some embodiments, the structure of the recognition element may vary. In some embodiments, the structure of the recognition element may configure into a specific structure when hybridized to a target nucleic acid. Non-limiting examples of a recognition element configuration may include a padlock probe, a molecular inversion probe, a hairpin oligonucleotide, a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a combination thereof. In some embodiments, the recognition element is linear prior to hybridization to its complementary target nucleic acid of interest. In some embodiments, the linear recognition element is circularized once hybridized to the respective target nucleic acid of interest and ligated thereafter. In some embodiments, the recognition element is circular prior to hybridization to the target nucleic acid of interest. In one embodiment, the target nucleic acid of WSGR Docket No. 64100-744.601 interest may serve as a primer for an extension reaction, for example to initiate rolling circular amplification of the recognition element.

[0089] In some embodiments, the recognition element may be configured to be a padlock probe once the recognition element is hybridized to the target nucleic acid of interest. Padlock probes may be referred to as linear oligonucleotides whose ends are complementary to adjacent target sequences, or to non-adjacent target sequences thereby leaving a gap between the ends of the hybridized recognition element. Upon hybridization to a target nucleic acid, the two ends (e.g., 5’ end and 3’ end) of the recognition element may be adjacently located, generating a padlock probe configuration for subsequent ligation. Alternatively, the two ends of the recognition element may be brought in proximity to, but not directly adjacent to, each other upon hybridization to a target nucleic acid. In this instance, a gap is left between the 5' and 3' hybridized ends of the recognition element which can be filled in several ways, for example by extension of the 5' end until it is adjacent to the 3' end, or by hybridizing a third oligonucleotide that fills the gap. In any scenario, the recognition element ends may be ligated together if hybridization, or hybridization and gap fill, occurs thereby generating circular and ligated recognition elements that are indicative of the hybridization event.

[0090] In some embodiments, a recognition element further comprises one or more functional sequences. Functional sequences include, but are not limited to, primer binding sites, cleavage sites, unique molecular identifiers, capture sequences, or combinations thereof. In some embodiments, a primer binding site and / or a cleavage site are universal in nature, such that a plurality of recognition elements shares the same sequence(s). A unique molecular identifier may be included in a recognition element to identify a source of material, for error correction, as known in the art. Codes

[0091] The methods described herein may relate to the use of a code for identifying a target nucleic acid of interest from a sample that hybridized to a recognition element to initiate a ligation event. A code in a recognition element may be used to associate the recognition element 5' and 3' end regions with a target nucleic acid of interest, thereby determining the presence of a target nucleic acid of interest from a sample without having to directly assay the target molecule itself. As such, a code in a recognition element may uniquely identify the presence of a target molecule from a sample. Using codes, any number of recognition elements can be multiplexed in one encoded assay as each code is unique and correlates to the presence of one target molecule. In some embodiments, the code may be selected from a set of codes wherein the set of codes make up a “code space”. In some embodiments, the code may comprise a plurality of nucleic acid segments, where each nucleic acid segment corresponds to one or more WSGR Docket No. 64100-744.601 computational symbols, or colors, that are used in a decoding process. Fig.1 shows an example where four nucleic acid segments make up the code of the recognition element, wherein each of the nucleic acid segments can be decoded using detection polynucleotides as disclosed herein and the combination of the decoded nucleic acid segments thereby builds the full code that is unique to the target nucleic acid of interest from a sample. The codes may be detected as proxies, thereby serving as an indirect analysis of the presence of a target molecule from a sample as the code correlates with the presence of the target molecule that hybridized to the recognition element allowing ligation, amplification and decoding. In some embodiments, if there is no hybridization of a target of interest to its complementary sequences of a recognition element, there is expected to be no ligation (e.g., as the 5’ and the 3’ ends of the recognition are not expected to be adjacent), no amplification and subsequently nothing to decode. As such, if there is no amplification product to decode, that may be an indication that the target molecule of interest was potentially absent from the sample, or at such a low incidence that hybridization resulted in too few amplification products to cross the threshold for detection by decoding.

[0092] In some embodiments, each code from the set of codes is from a predetermined set of codes. In some embodiments, each code from the set of codes may be selected to ensure that the selected code differs from other codes in the set of codes. As such, in some embodiments, several selection criteria may be implemented to generate a set of codes, wherein each code of a set of codes comprises from one to more than one nucleic acid segment. In some embodiments, selection of the codes, or nucleic acid segments that make up a code, may incorporate a Hamming distance criterion.

[0093] In some embodiments, to generate a code selected from a set of codes for use in a recognition element, a Hamming distance (HD) selection criterion may be implemented between any two codes of the set of codes, and also between any two nucleic acid segments that may be used in a code. A Hamming distance between two codes in a set of codes may refer to the number of symbols, or nucleotides, that differ between the two codes in the set of codes. In essence, the Hamming distance measures the number of changes that may need to be made to a first code sequence to change the string of symbols, in this case nucleotides, to the second code. As such, a Hamming distance criterion used to select a code may not be greater than the length of the code. For example, if the length of a code is measured by the number of cycles or flows of decoding runs or queries and that number being eight cycles, and if each cycle corresponds to one symbol or color, therefore eight symbols or colors, then the maximum Hamming distance is eight. In some embodiments, the Hamming distance may be a minimum Hamming distance. In some embodiments, the Hamming distance may be a maximum Hamming distance. In some embodiments, a minimum Hamming distance may be from about 2-10. In some embodiments, WSGR Docket No. 64100-744.601 the Hamming distance is between 2-7. In some embodiments, the Hamming distance is between 3-5. The Hamming distance may increase as the number of codes that can be used decreases, as one of the purposes of the code is to impart a way to uniquely identify one target molecule from another target molecule.

[0094] The code may have a certain length in nucleotides. In some embodiments, the code has a length of greater than or equal to about three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 contiguous nucleotides. In some embodiments, the code has a length of fewer than or equal to about 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 contiguous nucleotides. In some embodiments, the length of the code is about 5 to 200, 10 to 150, 15 to 100, 20 to 90, or 30 to 80 contiguous nucleotides. The length of the code may be further divided into a number of discrete nucleic acid segments, for example, four segments, as shown in Fig.1.

[0095] In some embodiments, each code from a set of codes is generated using a 4-ary nucleotide alphabet of A, C, G, and T. In some embodiments, each code of a set of codes is generated using a 3-ary nucleotide alphabet of a set of three of A, C, G, and T. In some embodiments, the codes can be generated from arbitrary symbols, or colors, 1 to 4, corresponding to the fluorophores that are associated with a unique string of nucleotides. The numbers then may be used in the abstract but may serve as a means to numerate colors or mixes of colors that are utilized to query the codes for decoding. Nucleic acid segments

[0096] A code of a recognition element can comprise one or more nucleic acid segments. For example, as shown in Fig.1, the example of the code comprises four nucleic acid segments.

[0097] In some embodiments, the recognition elements provided herein may comprise a code comprising one or more nucleic acid segments. The one or more nucleic acid segments, or the complements thereof, within the code may be used as a proxy for detection of the target molecule recognized by the recognition element.

[0098] The number of nucleic acid segments present in a code of a recognition element may be considered in the design of the recognition element. The number of segments in a code may help to determine the nucleotide length of the recognition element. For example, a recognition element that includes a code comprising five segments may comprise a greater nucleotide length than a recognition element that includes a code of two segments. A recognition element with a larger nucleotide length may run up against synthesis limits and may be at a greater risk of synthesis errors. Alternatively, a recognition element with a smaller nucleotide length may avoid synthesis limits and risks in synthesis errors. A recognition element with a larger nucleotide WSGR Docket No. 64100-744.601 length may include less space for other portions of the recognition element, such as the target recognition regions, functional sequences, universal sequences, etc.

[0099] In some embodiments, the code comprises about 2 to 10 nucleic acid segments. In some embodiments, the code comprises about 2 to 8 nucleic acid segments. In some embodiments, the code comprises about 3 to 5 nucleic acid segments. In some embodiments, the code comprises at least 4 nucleic acid segments, at least 5 nucleic acid segments, at least 6 nucleic acid segments, at least 7 nucleic acid segments, at least 8 nucleic acid segments, at least 9 nucleic acid segments, or at least 10 nucleic acid segments.

[0100] In some embodiments, each nucleic acid segment may comprise a length in nucleotides. In some embodiments, each nucleic acid segment may comprise a length of about 10 to 30 nucleotides. In some embodiments, each nucleic acid segment may comprise a length of about 10 to 25 nucleotides. In some embodiments, each nucleic acid segment may comprise a length of about 15 to 20 nucleotides. In some embodiments, each nucleic acid segment may comprise a length of 2 or more nucleotides, 4 or more nucleotides, 6 or more nucleotides, 8 or more nucleotides, 10 or more nucleotides, 12 or more nucleotides, 14 or more nucleotides, 16 or more nucleotides, 18 or more nucleotides, 20 or more nucleotides, or 22 or more nucleotides. In some embodiments, the nucleic acid segments in a code may be of the same length. In some embodiments, the nucleic acid segments in a code may not be the same length.

[0101] As with a code in a recognition element, a Hamming distance selection criterion may be implemented between any two nucleic acid segments of a code. A Hamming distance between two nucleic acid segments in a code may be referred to as the number of symbols that differ between the segments. In essence, the Hamming distance measures the number of changes that may need to be made to a first nucleic acid segment to change the string of symbols, or nucleotides, to a second nucleic acid segment. In some embodiments, the Hamming distance may be a minimum Hamming distance. In some embodiments, the Hamming distance may be a maximum Hamming distance. In some embodiments, a minimum Hamming distance may be from about 2-20, about 3-19, about 4-18, about 5-17, about 6-16, about 7-15, about 8-14, about 9-13, or about 10-12. In some embodiments, a minimum Hamming distance may be greater than or equal to about 2, greater than or equal to about 3, greater than or equal to about 4, greater than or equal to about 5, greater than or equal to about 6, greater than or equal to about 7, greater than or equal to about 8, greater than or equal to about 9, greater than or equal to about 10, greater than or equal to about 11, greater than or equal to about 12, greater than or equal to about 13, greater than or equal to about 14, greater than or equal to about 15, greater than or equal to about 16, greater than or equal to about 17, greater than or equal to about 18, greater than or equal to about 19, or greater than or equal to about 20. WSGR Docket No. 64100-744.601

[0102] In some embodiments, a nucleic acid segment may comprise a universal primer binding site for amplification. For example, a segment may comprise an amplification primer binding site for performing rolling circle amplification (RCA) for generating a plurality of concatemeric amplification products.

[0103] In some embodiments, a nucleotide or nucleic acid sequence of each segment may correspond to one or more computational symbols, such as a detection color, for performing a decoding process. For example, one or more nucleic acid segments of a code may be detected with a first pool of detection polynucleotide complexes to produce one or more detectable binding complexes, for example by using a fluorescent label. In some embodiments, the one or more detectable binding complexes, once imaged, may produce one or more optical signals such as fluorescence in a particular wavelength. When all or substantially all segments of the code are detected by iteratively applying additional pools of detection polynucleotide complexes to the amplification products, a series of optical signals may be observed and collated.

[0104] The application of detection polynucleotides to amplification products for decoding can be called a “flow” or "cycle" or “query”, wherein a flow, cycle or query is the number of times a particular segment of an amplification product is queried, or the number of times a detection polynucleotide is flowed over an amplification product in order to detect a nucleic acid segment sequence. If a nucleic acid segment is present, a detection polynucleotide that comprises a sequence complementary to that nucleic acid segment may hybridize to its complementary nucleic acid segment and the attached detectable label is detected, for example by imaging. In some embodiments, one or more optical signals observed from querying an amplification product with detection polynucleotide complexes may translate to one or more computational symbols such that each optical signal can be imaged and decoded. In some embodiments, a plurality of nucleic acid segments on a recognition element may correspond to at least three computational symbols. In some embodiments, the optical signal may be a color or a non-color. In some embodiments, the optical signal may be a combination of colors (e.g., when the detection polynucleotide complex comprises a plurality of detectable labels). In some embodiments, the computational symbols or colors can be referred to as numbers (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.). In some embodiments, each detection polynucleotide complex comprises a detectable label, as such for example when four different fluorescent moieties are used as detectable labels there may be four symbols, 1 to 4. However, the number of symbols can be larger depending on the combination of detectable labels with each unique detection polynucleotide complex. For example, for a set of 16 unique detection polynucleotide complexes wherein each detection polynucleotide has one of four fluorescent moieties, there may be 16 computational symbols used for decoding if all 16 unique detection polynucleotide WSGR Docket No. 64100-744.601 complexes are used to decode an amplification product. However, additional ways to increase the number of computational symbols for decoding include, but are not limited to, adding levels of identifiability associated with a particular detectable signal such as whether a detectable signal is brighter or dimmer compared to a normal level of signal, whether there is a combination of detectable colors that is used to identify a particular nucleotide. As such, the number of computational symbols that may be used may be limited by practicality for any given assay.

[0105] In some embodiments, the methods described herein may use a number of computational symbols. The number of computational symbols used in the methods and systems described herein may be considered in the design of the recognition elements. For example, in some embodiments, a detection scheme using a larger number of computational symbols may lead to a larger code space and a greater number of codes that may be generated, which may allow for a greater amount of information that may be detected thereby allowing for a higher degree of assay target molecule multiplexing. In some embodiments, a detection scheme using a smaller number of computational symbols may be limited in the amount of information that can be detected. In some embodiments, using a larger number of computational symbols may result in a faster detection process (e.g., less time to determine a target molecule compared to using a smaller number of computational symbols). In some embodiments, a detection scheme using a larger number of computational symbols may include greater instrument complexity, which may lead to potential drawbacks such as color crosstalk, wherein the computational symbols used in the detection scheme may become difficult to distinguish from other computational symbols. In some embodiments, a greater number of computational symbols may include that a more complex detection tool be used.

[0106] In some embodiments, each nucleic acid segment may correspond to a combination of computational symbols. In some embodiments, each nucleic acid segment may correspond to one or more computational symbols, two or more computational symbols, three or more computational symbols, four or more computational symbols, five or more computational symbols, six or more computational symbols, seven or more computational symbols, eight or more computational symbols, nine or more computational symbols, or 10 or more computational symbols. In some embodiments, each nucleic acid segment may correspond to 10 or less computational symbols, nine or less computational symbols, eight or less computational symbols, seven or less computational symbols, six or less computational symbols, five or less computational symbols, four or less computational symbols, three or less computational symbols, or two or less computational symbols. WSGR Docket No. 64100-744.601 Amplification

[0107] The methods described herein may include amplification of a circularized and ligated recognition element. In some embodiments, a target nucleic acid molecule is amplified. In some embodiments, the target nucleic acid molecule is a combination of a recognition element and a target nucleic acid molecule. In some embodiments, the amplification is selective amplification. For example, in some embodiments, amplification may occur if a target recognition region of a recognition element recognizes and binds to a complementary target nucleic acid of interest. In some embodiments, amplification may occur if a primer is used that is complementary to one or more of a portion of a recognition element, a portion of a nucleic acid segment of a code, or another sequence in the recognition element that is complementary to a primer used for amplification. In some embodiments, the amplification is non-selective. For example, in some embodiments, randomers can be used to prime amplification from a recognition element.

[0108] In some embodiments, the methods described herein may include selectively amplifying a subset of nucleic acids. For example, in some embodiments, a subset of a plurality of recognition elements hybridized to a plurality of target nucleic acid molecules may be amplified. The subset may comprise a percentage of the total amount of recognition elements hybridized to target nucleic acid molecules as described herein. In some embodiments, the subset may include 5% or more, 10% or more, 15% or more, 20% or more, 25% or more, 30% or more, 35% or more, 40% or more, 45% or more, 50% or more, 55% or more, 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, or 95% or more of the total amount of recognition elements hybridized to target nucleic acid molecules. In some embodiments, the subset may include 95% or less, 90% or less, 85% or less, 80% or less, 75% or less, 70% or less, 65% or less, 60% or less, 55% or less, 50% or less, 45% or less, 40% or less, 35% or less, 30% or less, 25% or less, 20% or less, 15% or less, 10% or less, or 5% or less of the total amount of recognition elements hybridized to target nucleic acid molecules.

[0109] In some embodiments, the amplification may include rolling circle amplification (RCA). In some embodiments, the RCA may generate a concatemer as an amplification product, wherein the concatemer comprises multiple copies of a circularized ligated recognition element, including associated codes, target recognition regions, and any other functional sequences that are included in the circularized and ligated recognition element. In some embodiments, RCA may be performed while the circularized and ligated recognition element is in solution. In some embodiments, RCA may be performed on a circularized recognition element while the circularized recognition element is immobilized, either reversibly or non-reversibly, on a substrate or surface. In some embodiments, the substrate or surface is a solid support and includes, but is not limited to, a bead, a flow cell, a microwell, a nanowell, a well, a slide. In WSGR Docket No. 64100-744.601 some embodiments, the substrate is glass such as optical glass of imaging quality. In some embodiments, the substrate is plastic, polycarbonate, etc. In some embodiments, the substrate is positively charged or negatively charged. In some embodiments, the substrate is an anionic substrate. In some embodiments, the substrate is a cationic substrate. In some embodiments, the substrate comprises an immobilization composition, such as polyacrylamide, branched PEI, linear PEI, poly(β-aminoester) and poly(amidoamine), PEG, a gel, poly-L-lysine, silane, agarose, muscle mimetic catecholamine polymer, and the like. In some embodiments, the substrate has no charge. In some embodiments, the substrate has no immobilization composition. In some embodiments, a substrate comprises a cationic polymer coated surface. An RCA reaction may be performed in the presence of a cationic polymer coated surface, resulting in simultaneous immobilization and amplification of a ligated recognition element. RCA primers may be supplied in solution or bound to the cationic polymer-coated surface prior to, or concurrent with, performing the RCA reaction.

[0110] In some embodiments, the amplification may include on-surface polymerase chain reaction (PCR), isothermal amplification, RCA, ultrarapid amplification, or a combination thereof. In some embodiments, amplification may include polymerase chain reaction (PCR). In some embodiments, PCR is multiplexed PCR. In some embodiments, PCR is ultrafast multiplexed PCR. The amplification methods disclosed herein may include isothermal amplification. Non-limiting examples of isothermal amplification include Nicking endonuclease amplification reaction (NEAR), Transcription mediated amplification (TMA), Loop-mediated isothermal amplification (LAMP), Helicase-dependent amplification (HDA), Nucleic Acid Sequence Based Amplification (NASBA), Strand displacement amplification (SDA), Multiple Displacement Amplification (MDA), Rolling Circle Amplification (RCA), bridge amplification, or Ramification (RAM) amplification method. In some embodiments, the amplification method is provided in Fakruddin M, Mannan KS, Chowdhury A, Mazumdar RM, Hossain MN, Islam S, Chowdhury MA. Nucleic acid amplification: Alternative methods of polymerase chain reaction. J Pharm Bioallied Sci.2013 Oct;5(4):245-52, which is hereby incorporated by reference in its entirety. Detection polynucleotides

[0111] The methods described herein may include introducing detection polynucleotides to amplified recognition elements. A detection probe may be a single stranded oligonucleotide or may be partially single stranded and partially double stranded as a detection polynucleotide. A detection oligonucleotide or a detection polynucleotide may comprise a detectable label.

[0112] In some embodiments, the methods described herein may include introducing one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight WSGR Docket No. 64100-744.601 or more, nine or more, 10 or more, 15 or more, 20 or more, 25 or more, 50 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1,000 or more detection polynucleotides (or single stranded oligonucleotides) to an amplification product. In some embodiments, the methods described herein may include introducing 1,000 or less, 900 or less, 800 or less, 700 or less, 600 or less, 500 or less, 400 or less, 300 or less, 200 or less, 100 or less, 50 or less, 25 or less, 20 or less, 15 or less, 10 or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less detection polynucleotides to an amplification product.

[0113] In some embodiments, a detection polynucleotide may comprise a detectable label (e.g., fluorescent molecule). In some embodiments, a detection polynucleotide may be a single stranded oligonucleotide with a portion that is complementary to a code, or a portion of a code, and a detectable label. In some embodiments, a detection polynucleotide may comprise two oligonucleotides. Fig.3A shows a non-limiting example of a structure of a detection polynucleotide 300 comprising two oligonucleotides. A first oligonucleotide 330 comprises a detectable label 340. A second oligonucleotide 350 comprises a portion that is complementary to the first oligonucleotide 310 and a second portion 320 that is complementary to a code or a portion of a code 360 (e.g., a nucleic acid segment of a code). The first oligonucleotide 330 hybridizes to the second oligonucleotide 350, thereby generating a detection polynucleotide. For decoding, a portion of the second oligonucleotide 320 hybridizes to its code complement 360 as seen in Fig.3B, and a signal is detected from the detectable label, thereby identifying the code which in turn is correlated back to the presence of a target nucleic acid of interest from a sample.

[0114] The detection polynucleotide may comprise various nucleotide lengths. In some embodiments, the detection polynucleotide may comprise a length of 5 to 25 nucleotides. In some embodiments, the detection polynucleotide may comprise a length of 5 to 20 nucleotides. In some embodiments, the detection polynucleotide may comprise a length of 5 to 15 nucleotides. In some embodiments, the detection polynucleotide may comprise a length of 5 to 10 nucleotides. In some embodiments, the detection polynucleotide may comprise a length of 5 to 8 nucleotides.

[0115] In some embodiments, the detection polynucleotide may comprise a length of between about 5-100 nucleotides, between about 10-80 nucleotides, between about 20-60 nucleotides, between about 30-50 nucleotides, or between about 15-30 nucleotides. In some embodiments, the detection polynucleotide may comprise one or more detectable labels. In some embodiments, the one or more detectable labels may comprise a fluorescent moiety. The fluorescent moiety may emit in the red, far-red, near-red, yellow, green, or blue wavelengths. In some embodiments, the fluorescent moiety comprises one or more of 6-FAM (6-carboxyfluorescein), WSGR Docket No. 64100-744.601 JOE (6-carboxy-4',5'-dichloro-2',7'-dimethoxyfluorescein), TAMRA (6- carboxytetramethylrhodamine), 5-Cy5 (5-carboxyrhodamine), 5-Cy5.5 (5-carboxylic acid succinimidyl ester), 5-Cy7 (5-carboxyrhodamine), (hexachlorofluorescein), Alexa Fluor 488 (AF488), Alexa Fluor 514 (AF514), Texas Red, Cyanine 3, Cyanine 5, Pacific Blue, Tetramethyl rhodamine, Oxazole Yellow, Atto647N, and Rhodamine 6G (R6G). In some embodiments, the detection polynucleotide may comprise two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or 10 or more fluorescent moieties. In some embodiments, the detection polynucleotide may comprise 10 or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less fluorescent moieties.

[0116] In some embodiments, the fluorescent moiety may comprise an organic dye, a biological fluorophore, a quantum dot, or a combination thereof. In some embodiments, the organic dye may comprise an organic molecule. In some embodiments, the organic dye may comprise a coumarin, a cyanine, a benzofuran, a quinoline, a quinazolinone, an indole, a benzazole, a borapolyazaindacene, a xanthene, or a combination thereof. The organic dye may correspond to a color. For example, the organic dye may correspond to a green color, a yellow color, a blue color, an indigo color, a red color, an orange color, a purple color, a pink color, a violet color, or a combination thereof. In some embodiments, the organic dye may correspond to no color. In some embodiments, the organic dye may correspond to a black color. In some embodiments, the organic dye may correspond to a white color.

[0117] In some embodiments, the detectable moiety can be identified by imaging. When the detectable label is a fluorophore, the fluorophore may emit a color in the visible light spectrum which can be captured by fluorescent imaging and associated filters. In some embodiments, the fluorophore may emit in a wavelength in the range between about 400 nm and 900 nm. In some embodiments, the fluorophore may emit in a wavelength between about 400 nm and 475 nm, about 475 nm and 490 nm, about 490 nm and 530 nm, about 530 nm and 575 nm, about 575 nm and 600 nm, about 600 nm and 700 nm, or about 700 nm and 800 nm. In some embodiments, the fluorophore may emit a wavelength of 400 nm or more, 425 nm or more, 450 nm or more, 475 nm or more, 500 nm or more, 525 nm or more, 550 nm or more, 575 nm or more, 600 nm or more, 625 nm or more, 650 nm or more, 675 nm or more, 700 nm or more, 725 nm or more, 750 nm or more, 775 nm or more, 800 nm or more, 825 nm or more, 850 nm or more, 875 nm or more, or 900 nm or more. In some embodiments, the fluorophore may emit a wavelength of 900 nm or less, 875 nm or less, 850 nm or less, 825 nm or less, 800 nm or less, 775 nm or less, 750 nm or less, 725 nm or less, 700 nm or less, 675 nm or less, 650 nm or less, 625 nm or less, 600 WSGR Docket No. 64100-744.601 nm or less, 575 nm or less, 550 nm or less, 525 nm or less, 500 nm or less, 475 nm or less, 450 nm or less, 425 nm or less, or 400 nm or less.

[0118] In some embodiments, the detectable labels (e.g., fluorescent moieties) may be optically distinct. The number of optically distinct detectable labels used in the methods described herein can impact the amount of information that is detected. For example, a detection scheme using a larger number of optically distinct detectable labels may allow for a higher amount of multiplexing of codes, which may in turn allow for a greater amount of target molecule related information to be detected and captured. A detection scheme using a smaller number of optically distinct detectable labels may allow for a lesser amount of target molecule related information to be detected and captured. In some embodiments, using a larger number of optically distinct detectable labels may lead to a detection process that identifies a target molecule in less time as compared to using a fewer number of optically distinct detectable labels when querying a concatemeric amplification product. In some embodiments, a detection scheme using a larger number of optically distinct detectable labels may lead to greater instrument complexity, which may lead to fluorescence detection crosstalk, whereby the fluorescence emission spectra of the optically distinct fluorescent moieties may not yield distinct fluorescence signals. In some embodiments, a detection scheme using a greater number of optically distinct fluorescent moieties may include use of a more complex detection tool.

[0119] In some embodiments, the detection polynucleotides may be provided in one or more detection pools. In some embodiments, the methods herein may use one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 45 or more, or 50 or more detection pools. In some embodiments, the methods herein may use 50 or less, 45 or less, 40 or less, 35 or less, 30 or less, 25 or less, 20 or less, 15 or less, ten or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less detection pools.

[0120] In some embodiments, each detection pool provided may comprise a number of detection polynucleotides. In some embodiments, each detection pool may comprise two or more, three or more, four or more, five or more, ten or more, 15 or more, 25 or more, 50 or more, 100 or more, 150 or more, 250 or more, 500 or more, 1,000 or more, 1,500 or more, 2,500 or more, or 5,000 or more detection polynucleotides. In some embodiments, each detection pool may comprise 5,000 or less, 2,500 or less, 1,500 or less, 1,000 or less, 500 or less, 250 or less, 150 or less, 100 or less, 50 or less, 25 or less, 15 or less, ten or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less detection polynucleotides. WSGR Docket No. 64100-744.601

[0121] The number of detection pools and the number of detection polynucleotides in each detection pool may be considered in the design of the recognition elements. For example, one advantage to using a smaller number of detection pools and detection polynucleotides in the methods described herein may be to lower design costs. Conversely, one advantage to using a larger number of detection pools and detection polynucleotides in the methods described herein may be the need for a higher degree of multiplexing for target molecule detection and larger amounts of information that may be detected. Imaging

[0122] The methods described herein may include imaging. The methods may include imaging a plurality of detection polynucleotides that have hybridized to their complementary code or a portion of a code in order to obtain identifiable signals which can be correlated back to the presence of a target molecule of interest. In some embodiments, the signals are associated with one or more segments of a code for each concatemeric amplification product. In some embodiments, the imaging is performed by an imaging system comprising a fluorescence detection system.

[0123] In some embodiments, the imaging may be conducted using an imaging system. The imaging system may comprise at the minimum a camera, a detector, an illuminator, a condenser, or a combination thereof. In some embodiments, the imaging may include images of fluorescence emission, luminescence, or a combination thereof. In some embodiments, the imaging systems comprise components or sub-systems of a larger system that may also include optics modules including when needed fluorescence filters, fluidics modules, temperature control modules, translation stages, robotic fluid dispensing and / or microplate handling, processors or computers, instrument control software, data analysis and display software, etc. In some embodiments, the imaging system is a fluorescence imaging system. In some embodiments, the imaging may include fluorescent images from the fluorescent moieties present on the labeled probes.

[0124] In some embodiments, the image may comprise fluorescence information from one or more wavelengths. In some embodiments, the fluorescence information may comprise emission data from a wavelength from about 220-830 nanometers (nm), about 230-820 nm, about 240- 810 nm, about 250-800 nm, about 260-790 nm, about 270-780 nm, about 280-770 nm, about 290-760 nm, about 300-750 nm, about 310-740 nm, about 320-730 nm, about 330-720 nm, about 340-710 nm, about 350-700 nm, about 360-690 nm, about 370-680 nm, about 380-670 nm, about 390-660 nm, about 400-650 nm, about 410-640 nm, about 420-630 nm, about 430- 620 nm, about 440-610 nm, about 450-600 nm, about 460-590 nm, about 470-580 nm, about WSGR Docket No. 64100-744.601 480-570 nm, about 490-560 nm, about 500-550 nm, about 510-540 nm, about 520-530 nm, or a combination thereof.

[0125] Image detection and capture may relate to iteratively repeating the operations of: (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof. In some embodiments, the iterative repetition of the operations is performed for each nucleic acid segment of a code one or more times.

[0126] In some embodiments, the iteratively repeating the operations of: (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof may comprise two or more iterative repetitions. For example, the methods described herein may comprise about 2-50 iterative repetitions, about 2-10 iterative repetitions, or about 2-8 iterative repetitions. In some embodiments, the method described herein may comprise one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, or 50 or more iterative repetitions of: (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof.

[0127] In some embodiments, the number of iterative repetitions of the operations of: (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof may correspond to the number of nucleic acid segments present in a code of the recognition element. In some embodiments, each nucleic acid segment of the code of the recognition element may undergo a number of iterative repetitions of the operations of: (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof. For example, the methods described herein may comprise iteratively repeating the operations two times per segment, three times per segment, or four times per segment. In some embodiments, the methods described herein may comprise iteratively repeating the operations two or more times per segment, three or more times per segment, four or more times per segment, five or more times per segment, six or more times WSGR Docket No. 64100-744.601 per segment, seven or more times per segment, eight or more times per segment, nine or more times per segment, 10 or more times per segment, 11 or more times per segment, 12 or more times per segment, 13 or more times per segment, 14 or more times per segment, 15 or more times per segment, 16 or more times per segment, 17 or more times per segment, 18 or more times per segment, 19 or more times per segment, or 20 or more times per segment.

[0128] Additional methods for imaging a detection polynucleotide can be found in WO2023 / 158993A2, which is incorporated herein by reference in its entirety. Decoding

[0129] Several models may be used to identify a code that is associated with a target molecule based on the fluorescence signals generated and images captured from detection polynucleotide hybridization to code sequences. In some embodiments, decoding may make use of a hard decision decoding model. In another embodiment, decoding may make use of a soft decision decoding model.

[0130] For soft decision decoding, it is not necessary to identify each nucleotide specifically. For example, signals generated during each detection event may be detected and recorded to produce a data set that may be used as input into a model to calculate a probability that a specific code is present without requiring that each nucleotide of a code be determined. Although it may not be necessary in a soft decision decoding model to make a hard decision about the identity of each nucleotide, a model may nevertheless include assigning a probability or identity to each nucleotide in the sequence of a code, wherein each nucleotide in the sequence of a code may be sequenced. Data gathered includes intensity readings for signals produced by the hybridized detection polynucleotide fluorescent moiety in various spectral bands. A set of intensity readings are detected by imaging, stored and used as input into a soft decision decoding model for determining a probability that a particular code is present, and hence a target nucleic acid is present in the sample.

[0131] A model may be developed or trained using data from known codes, such as signal intensity data across a predetermined spectrum. The model may be used to calculate a set of probabilities across a set of one or more codes, indicating, for example, for each code, a probability that it is present in a concatemeric amplification product.

[0132] The probability that a particular code is present may be indicative of the probability that a particular target molecule associated with the code is present in the sample of interest. Data indicating the probability that a particular target is present is, for example, to calculate probabilities relevant to diagnosis or screening of various medical conditions, or selection of drugs for treatment of various medical conditions. WSGR Docket No. 64100-744.601

[0133] A soft decoding decision model may include using an algorithm to predict the presence of target molecules from a sample. In some embodiments, the algorithm is a soft-decision decoding algorithm. In some embodiments, the algorithm is applied to the codes of the concatemeric amplification products for predicting the presence of a target molecule from a sample.

[0134] The methods disclosed herein may comprise soft decision decoding to predict the presence of the code in a recognition element or concatemeric amplification product thereof, wherein the presence of the code correlates and serves as a proxy for the presence of a target nucleic acid in a sample. In some embodiments, the methods described herein may use soft decision decoding. In some embodiments, the methods described herein may use hard decision decoding. For hard decision decoding, signals from queried concatemers may be extracted from images. This may be the same for soft decision decoding, in that signals that are generated and imaged are extracted from the images. For hard decision decoding, hard symbol calls are generated from the intensities of the signals, whereas with soft decision decoding no hard symbol calls are necessary as all of the signal range is retained. The code assignment for hard decision decoding is determined by matching symbol-to-symbol readouts of codewords to codes, whereas with soft decision decoding, the signals may be cross-correlated against the expected signals and the most likely code is assigned using a probabilistic methodology. When using soft decision decoding, it may not be necessary for the model to identify each symbol specifically. For example, signals (e.g., fluorescent signals) generated during each cycle of a detection process may be detected and recorded to produce a data set that may be used as input into a model to calculate a probability that a specific code is present.

[0135] The permutation space on a recognition element may be referred to herein as the totality of factors that determines the number of unique nucleotide possibilities at each nucleic acid segment. Factors may comprise the number of segments present on a recognition element, the number of incubation periods or times a segment is queried with a detection pool detection polynucleotides, or the number of computational symbols or colors.

[0136] Fig.4 details an example of a soft decision decoding workflow for determining the presence of a target molecule from a sample based on detection and decoding of a code associated with the target molecule that originally hybridized to a recognition element. Images of the sample may be acquired, aligned, and processed to extract the intensity of the features, or signals of interest across the imaged field of view in multiple spectral channels. The corrected intensities of said features may be fed through a series of algorithms that make up the soft decoder. At first, the intensity profiles of the codes may be learned based on features of high confidence or high intensity. This trained model may provide a template for each code from WSGR Docket No. 64100-744.601 which the rest of the features of interest are compared to in the second operation. Third, a confidence score may be computed from the difference between the intensity profile of each feature and the trained profiles. Several filters may be then applied to remove outliers, duplicates, and low confidence decoded concatemers. The output may be a table of decoded concatemers with associated filter status, confidence score, and most likely assignment to one of the codes of one or more concatemeric amplification products.

[0137] In some embodiments, a recognition element comprises a larger code, for example a code with four segments instead of two or three. In some embodiments, a recognition element comprising a larger code may result in a detection scheme with better error correction. Additionally, in some embodiments, a larger code may result in a lower signal-to-noise ratio.

[0138] In some embodiments, a recognition element comprises a smaller code, for example a code with two segments, or one segment. In some embodiments, a recognition element comprising a small code may result in a detection scheme with lower error correction abilities. Further, in some embodiments, a small code may result in a higher signal-to-noise ratio. II. SYSTEMS

[0139] Described herein are systems related to the methods, compositions, and kits described herein. In some embodiments, the systems comprise a solid substrate configured to immobilize one or more of a circularized and ligated recognition element, a concatemeric amplification product, a detection polynucleotide, and a hybridized complex of a concatemeric amplification product and a detection polynucleotide complex. In some embodiments, the systems comprise a welled plate or a flowcell. In some embodiments, the systems comprise a fluid flow controller, a temperature controller, an imaging system, a computer system, or any combination thereof.

[0140] In some embodiments, the systems disclosed herein may include a solid substrate or a solid surface. The solid substrates and surfaces disclosed herein may be referred to as a substrate, a support, a solid support, or a surface. The substrate may be modified for immobilizing circularized and ligated recognition elements or concatemeric amplification products, or both. Example solid substrates include, but are not limited to, glass, modified or functionalized glass, plastics, polysaccharides, nylon, nitrocellulose, ceramics, resins, silica, silica-based materials, carbon, metals, inorganic glasses, plastics, optical fiber bundles, optically clear glass, and other polymers. In some embodiments, the plastic solid substrates may include acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, or polyurethanes. In some embodiments, the silica-based solid substrates may include silicon or modified silicon.

[0141] In some embodiments, the substrate may be a welled plate. In some embodiments, the substrate may be a 96-well plate. In some embodiments, the substrate may be a 4-well plate, a 6- WSGR Docket No. 64100-744.601 well plate, an 8-well plate, a 12-well plate, a 24-well plate, a 48-well plate, a 384-well plate, an 864-well plate, or a 1,536-well plate. In some embodiments, the substrate may have greater than or equal to 96 wells. In some embodiments, the substrate may have less than or equal to 96 wells.

[0142] In some embodiments, the substrate may be a flowcell. In some embodiments, the flowcell may have two or more lanes. In some embodiments, the flowcell may have two or less lanes.

[0143] In some embodiments, the substrate may be a microarray, a slide, a chip, a microwell, a tube, a column, a particle, a bead, or a paramagnetic bead.

[0144] In some embodiments, the substrate may comprise a coating. In some embodiments, the coating may comprise a layer that may be charged. In some embodiments, the coating layer may be positively charged. In some embodiments, the coating layer may be negatively charged. In some embodiments, the coating may be non-charged. In some embodiments, the substrate may comprise a surface comprising a cation-coating layer. In some embodiments, the substrate may comprise a surface comprising an anion-coating layer. In some embodiments, the substrate may comprise a surface comprising a neutral-charged layer. In some embodiments, the substrate may be coated with streptavidin. In some embodiments, the substrate may be coated with avidin. In some embodiments, the substrate may be coated with one or more antibodies.

[0145] The systems disclosed herein may comprise a fluidics system. The fluidics system may comprise a fluid flow controller. In some embodiments, the fluid flow controller may comprise one or more pumps, valves, mixing manifolds, reagent reservoirs, waste reservoirs, or any combination thereof. In some embodiments, the fluidic system and subcomponents of the fluidics system are fluidically connected to the reaction vessel of the present disclosure. In some embodiments, the fluidic system and subcomponents of the fluidics system iteratively flow in reagents (e.g., buffers, detector polynucleotides, anchor polynucleotides, detection oligonucleotide complexes, etc.) to the reaction vessel. In some embodiments, the reaction vessel comprises a solid substrate configured to immobilize the circularized and ligated recognition elements or concatemeric amplification products thereof.

[0146] The systems disclosed herein may comprise a temperature system. The temperature system may comprise a temperature controller. The temperature controller may be incorporated into the systems described herein to facilitate accuracy of the methods and systems described herein. In some embodiments, the temperature controller may comprise temperature control components. Non-limiting examples of temperature control components include resistive heating elements, infrared light sources, heating or cooling devices, heat sinks, thermocouples, thermistors, or a combination thereof. In some embodiments, the temperature controller may WSGR Docket No. 64100-744.601 provide changes in temperature over specified time intervals. In some embodiments, the temperature controller may provide an increase in temperature. In some embodiments, the temperature controller may provide a decrease in temperature. In some embodiments, the temperature controller may provide for cycling of temperatures between two or more set temperatures so that thermocycling or amplification may be performed. In some embodiments, the temperature controller may provide a constant temperature.

[0147] The systems disclosed herein may comprise an imaging system. In some embodiments, signals produced by the labeled probes disclosed herein may be imaged by the imaging systems disclosed herein. The imaging system may comprise one or more light sources, one or more optical components, one or more filters, one or one or more imaging sensors for imaging and detection, or a combination thereof. In some embodiments, the one or more light sources may comprise light from a bulb. In some embodiments, the one or more optical components may comprise lenses, mirrors, digital mirror devices, prisms, optical filters, colored glass filters, narrowband interference filters, broadband interference filters, dichroic reflectors, diffraction gratings, apertures, optical fibers, optical waveguides, or a combination thereof. In some embodiments, the one or more imaging sensors may comprise a charge-coupled device (CCD) sensor or camera, a complementary metal-oxide-semiconductor (CMOS) imaging sensor or camera, a negative-channel metal-oxide semiconductor (NMOS) imaging sensor or camera, or a combination thereof.

[0148] Various operations of the methods and systems disclosed herein may be performed by a computer system of the present disclosure. Referring to Fig.5, an example of a block diagram is shown depicting an example of a machine that includes a computer system 500 (e.g., a processing or computing system) within which a set of instructions can execute for causing a device to perform or execute any one or more of the aspects and / or methodologies for static code scheduling of the present disclosure. The components in Fig.5 are examples and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or a combination of two or more such components implementing particular embodiments.

[0149] Computer system 500 may include one or more processors 501, a memory 503, and a storage 508 that communicate with each other, and with other components, via a bus (not shown). The bus may also link a display 532, one or more input devices 533 (which may, for example, include a keypad, a keyboard, a mouse, a stylus, etc.), one or more output devices 534, one or more storage devices 535, and various tangible storage media 536. All of these elements may interface directly or via one or more interfaces or adaptors to the bus. For instance, the various tangible storage media 536 can interface with the bus via storage medium interface 526. Computer system 500 may have any suitable physical form, including but not limited to one or WSGR Docket No. 64100-744.601 more integrated circuits (ICs), printed circuit boards (PCBs), mobile handheld devices (such as mobile telephones or PDAs), laptop or notebook computers, distributed computer systems, computing grids, or servers.

[0150] Computer system 500 may include one or more processor(s) 501 (e.g., central processing units (CPUs), general purpose graphics processing units (GPGPUs), or quantum processing units (QPUs)) that carry out functions. Processor(s) 501 optionally comprises a cache memory unit 502 for temporary local storage of instructions, data, or computer addresses. Processor(s) 501 may be configured to assist in execution of computer readable instructions. Computer system 500 may provide functionality for the components depicted in Fig.5 as a result of the processor(s) 501 executing non-transitory, processor-executable instructions embodied in one or more tangible computer-readable storage media, such as memory 503, storage 508, storage devices 535, and / or storage medium 536. The computer-readable media may store software that implements particular embodiments, and processor(s) 501 may execute the software. Memory 503 may read the software from one or more other computer-readable media (such as mass storage device(s) 535, 536) or from one or more other sources through a suitable interface, such as network interface 520. The software may cause processor(s) 501 to carry out one or more processes or one or more operations of one or more processes described or illustrated herein. Carrying out such processes or operations may include defining data structures stored in memory 503 and modifying the data structures as directed by the software.

[0151] The memory 503 may include various components (e.g., machine readable media) including, but not limited to, a random-access memory component (e.g., RAM 504) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM), phase- change random access memory (PRAM), etc.), a read-only memory component (e.g., ROM 505), and any combinations thereof. ROM 505 may act to communicate data and instructions unidirectionally to processor(s) 501, and RAM 504 may act to communicate data and instructions bidirectionally with processor(s) 501. ROM 505 and RAM 504 may include any suitable tangible computer-readable media described below. In one example, a basic input / output system 506 (BIOS), including basic routines that help to transfer information between elements within computer system 500, such as during start-up, may be stored in the memory 503.

[0152] Fixed storage 508 may be connected bidirectionally to processor(s) 501, optionally through storage control unit 507. Fixed storage 508 may provide additional data storage capacity and may also include any suitable tangible computer-readable media described herein. Storage 508 may be used to store operating system 509, executable(s) 510, data 511, applications 512 (application programs), and the like. Storage 508 can also include an optical disk drive, a solid- WSGR Docket No. 64100-744.601 state memory device (e.g., flash-based systems), or a combination of any of the above. Information in storage 508 may, in appropriate cases, be incorporated as virtual memory in memory 503.

[0153] In one example, storage device(s) 535 may be removably interfaced with computer system 500 (e.g., via an external port connector (not shown)) via a storage device interface 525. Particularly, storage device(s) 535 and an associated machine-readable medium may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 500. In one example, software may reside, completely or partially, within a machine-readable medium on storage device(s) 535. In another example, software may reside, completely or partially, within processor(s) 501.

[0154] Bus may connect a wide variety of subsystems. Herein, reference to a bus may encompass one or more digital signal lines serving a common function, where appropriate. Bus may be any of several types of bus structures including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combinations thereof, using any of a variety of bus architectures. As an example, and not by way of limitation, such architectures include an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a Micro Channel Architecture (MCA) bus, a Video Electronics Standards Association local bus (VLB), a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, an Accelerated Graphics Port (AGP) bus, HyperTransport (HTX) bus, serial advanced technology attachment (SATA) bus, and any combinations thereof.

[0155] Computer system 500 may also include an input device 533. In one example, a user of computer system 500 may enter commands and / or other information into computer system 500 via input device(s) 533. Examples of an input device(s) 533 include, but are not limited to, an alpha-numeric input device (e.g., a keyboard), a pointing device (e.g., a mouse or touchpad), a touchpad, a touch screen, a multi-touch screen, a joystick, a stylus, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), an optical scanner, a video or still image capture device (e.g., a camera), and any combinations thereof. In some embodiments, the input device is a Kinect, Leap Motion, or the like. Input device(s) 533 may be interfaced to bus via any of a variety of input interfaces 523 (e.g., input interface 523) including, but not limited to, serial, parallel, game port, USB, FIREWIRE, THUNDERBOLT, or any combination of the above.

[0156] In particular embodiments, when computer system 500 is connected to network 530, computer system 500 may communicate with other devices, specifically mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, and the like, connected to network 530. Communications to and from computer system WSGR Docket No. 64100-744.601 500 may be sent through network interface 520. For example, network interface 520 may receive incoming communications (such as requests or responses from other devices) in the form of one or more packets (such as Internet Protocol (IP) packets) from network 530, and computer system 500 may store the incoming communications in memory 503 for processing. Computer system 500 may similarly store outgoing communications (such as requests or responses to other devices) in the form of one or more packets in memory 503 and communicated to network 530 from network interface 520. Processor(s) 501 may access these communication packets stored in memory 503 for processing.

[0157] Examples of the network interface 520 may include, but are not limited to, a network interface card, a modem, and any combination thereof. Examples of a network 530 or network segment 530 include, but are not limited to, a distributed computing system, a cloud computing system, a wide area network (WAN) (e.g., the Internet, an enterprise network), a local area network (LAN) (e.g., a network associated with an office, a building, a campus or other relatively small geographic space), a telephone network, a direct connection between two computing devices, a peer-to-peer network, and any combinations thereof. A network, such as network 530, may employ a wired and / or a wireless mode of communication. In general, any network topology may be used.

[0158] Information and data can be displayed through a display 532. Examples of a display 532 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED) such as a passive-matrix OLED (PMOLED) or active-matrix OLED (AMOLED) display, a plasma display, and any combinations thereof. The display 532 can interface to the processor(s) 501, memory 503, and fixed storage 508, as well as other devices, such as input device(s) 533, via the bus. The display 532 is linked to the bus via a video interface 522, and transport of data between the display 532 and the bus can be controlled via the graphics control 521. In some embodiments, the display is a video projector. In some embodiments, the display is a head- mounted display (HMD) such as a VR headset. In further embodiments, suitable VR headsets include, by way of non-limiting examples, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOVE VR, Zeiss VR One, Avegant Glyph, Freefly VR headset, and the like. In still further embodiments, the display is a combination of devices such as those disclosed herein.

[0159] In addition to a display 532, computer system 500 may include one or more other peripheral output devices 534 including, but not limited to, an audio speaker, a printer, a storage device, and any combinations thereof. Such peripheral output devices may be connected to the bus via an output interface 524. Examples of an output interface 524 include, but are not limited WSGR Docket No. 64100-744.601 to, a serial port, a parallel connection, a USB port, a FIREWIRE port, a THUNDERBOLT port, and any combinations thereof.

[0160] In addition, or as an alternative, computer system 500 may provide functionality as a result of logic hardwired or otherwise embodied in a circuit, which may operate in place of or together with software to execute one or more processes or one or more operations of one or more processes described or illustrated herein. Reference to software in this disclosure may encompass logic, and reference to logic may encompass software. Moreover, reference to a computer-readable medium may encompass a circuit (such as an IC) storing software for execution, a circuit embodying logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware, software, or both.

[0161] Those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm operations described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and operations have been described above generally in terms of their functionality.

[0162] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general- purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0163] The operations of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by one or more processor(s), or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An example storage medium is coupled to the processor such the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The WSGR Docket No. 64100-744.601 ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.

[0164] In accordance with the description herein, suitable computing devices include, by way of non-limiting examples, server computers, desktop computers, laptop computers, notebook computers, sub-notebook computers, netbook computers, notepad computers, set-top computers, media streaming devices, handheld computers, Internet appliances, mobile smartphones, tablet computers, personal digital assistants, video game consoles, and vehicles. Those of skill in the art will also recognize that select televisions, video players, and digital music players with optional computer network connectivity are suitable for use in the system described herein. Suitable tablet computers, in various embodiments, include those with booklet, slate, and convertible configurations, known to those of skill in the art.

[0165] In some embodiments, the computing device includes an operating system configured to perform executable instructions. The operating system is, for example, software, including programs and data, which manages the device’s hardware and provides services for execution of applications. Those of skill in the art will recognize that suitable server operating systems include, by way of non-limiting examples, FreeBSD, OpenBSD, NetBSD®, Linux, Apple®Mac OS X Server®, Oracle®Solaris®, Windows Server®, and Novell®NetWare®. Those of skill in the art will recognize that suitable personal computer operating systems include, by way of non- limiting examples, Microsoft®Windows®, Apple®Mac OS X®, UNIX®, and UNIX-like operating systems such as GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Those of skill in the art will also recognize that suitable mobile smartphone operating systems include, by way of non-limiting examples, Nokia®Symbian®OS, Apple®iOS®, Research In Motion®BlackBerry OS®, Google®Android®, Microsoft®Windows Phone®OS, Microsoft®Windows Mobile®OS, Linux®, and Palm®WebOS®. Those of skill in the art will also recognize that suitable media streaming device operating systems include, by way of non-limiting examples, Apple TV®, Roku®, Boxee®, Google TV®, Google Chromecast®, Amazon Fire®, and Samsung®HomeSync®. Those of skill in the art will also recognize that suitable video game console operating systems include, by way of non-limiting examples, Sony®PS3®, Sony®PS4®, Microsoft®Xbox 360®, Microsoft Xbox One, Nintendo®Wii®, Nintendo®Wii U®, and Ouya®.

[0166] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more non-transitory computer readable storage media encoded with a program including instructions executable by the operating system of an optionally networked computing device. In further embodiments, a computer readable storage medium is a tangible component of a computing device. In further embodiments, a computer readable storage medium is optionally WSGR Docket No. 64100-744.601 removable from a computing device. In some embodiments, a computer readable storage medium includes, by way of non-limiting examples, CD-ROMs, DVDs, flash memory devices, solid state memory, magnetic disk drives, magnetic tape drives, optical disk drives, distributed computing systems including cloud computing systems and services, and the like. In some cases, the program and instructions are permanently, substantially permanently, semi-permanently, or non-transitorily encoded on the media.

[0167] In some embodiments, the platforms, systems, media, and methods disclosed herein include at least one computer program, or use of the same. A computer program includes a sequence of instructions, executable by one or more processor(s) of the computing device’s CPU, written to perform a specified task. Computer readable instructions may be implemented as program modules, such as functions, objects, Application Programming Interfaces (APIs), computing data structures, and the like, which perform particular tasks or implement particular abstract data types. In light of the disclosure provided herein, those of skill in the art will recognize that a computer program may be written in various versions of various languages.

[0168] The functionality of the computer readable instructions may be combined or distributed as desired in various environments. In some embodiments, a computer program comprises one sequence of instructions. In some embodiments, a computer program comprises a plurality of sequences of instructions. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from a plurality of locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or combinations thereof.

[0169] In some embodiments, the computer programs described herein may be used to perform at least one function. The computer programs described herein may perform functions related to storing data, receiving data, analyzing data, exporting data, or a combination thereof. In some embodiments, the computer programs described herein may perform functions related to applying selection criteria, including in silico selection criteria, functional selection criteria, or a combination thereof. In some embodiments, the computer programs may receive sequence information, including sequence information for nucleic acid segments. The sequence information may be configured as an array, a table, a list, or combination thereof. The sequence information may be formatted in a variety of ways, including, but not limited to a .txt file, a FASTA file, an .xls file, or a combination thereof. The computer programs described herein may apply selection criterion or selection criteria to a set of nucleic acid segments. The computer programs may sort the nucleic acid segments, determine or compute characteristics of the WSGR Docket No. 64100-744.601 nucleic acid segments, perform calculations, reorder the nucleic acid segments, or a combination thereof. In some embodiments, the computer programs described herein may store information related to the nucleic acid segments. In some embodiments, the computer program may use information stored related to the nucleic acid segments to apply selection criteria to the nucleic acid segments. In certain embodiments, the computer program may receive information and / or data related to nucleic acid segments, selection criteria, or a combination thereof. In some embodiments, the computer programs may perform functions related to analyzing data from functional assays, including, but not limited to functional assays described herein. In some embodiments, analyzing data from functional assays may comprise image analysis, image quantification, intensity quantification, feature identification, or a combination thereof. The computer programs described herein may also export information. In some embodiments, the exported information may comprise images, files, data tables, documents, folders, or a combination thereof.

[0170] In some embodiments, a computer program includes a web application. In light of the disclosure provided herein, those of skill in the art will recognize that a web application, in various embodiments, utilizes one or more software frameworks and one or more database systems. In some embodiments, a web application is created upon a software framework such as Microsoft®.NET or Ruby on Rails (RoR). In some embodiments, a web application utilizes one or more database systems including, by way of non-limiting examples, relational, non-relational, object oriented, associative, XML, and document oriented database systems. In further embodiments, suitable relational database systems include, by way of non-limiting examples, Microsoft®SQL Server, mySQL™, and Oracle®. Those of skill in the art will also recognize that a web application, in various embodiments, is written in one or more versions of one or more languages. A web application may be written in one or more markup languages, presentation definition languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In some embodiments, a web application is written to some extent in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or eXtensible Markup Language (XML). In some embodiments, a web application is written to some extent in a presentation definition language such as Cascading Style Sheets (CSS). In some embodiments, a web application is written to some extent in a client-side scripting language such as Asynchronous JavaScript and XML (AJAX), Flash®ActionScript, JavaScript, or Silverlight®. In some embodiments, a web application is written to some extent in a server-side coding language such as Active Server Pages (ASP), ColdFusion®, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tcl, Smalltalk, WebDNA®, or Groovy. In some embodiments, a web WSGR Docket No. 64100-744.601 application is written to some extent in a database query language such as Structured Query Language (SQL). In some embodiments, a web application integrates enterprise server products such as IBM®Lotus Domino®. In some embodiments, a web application includes a media player element. In various further embodiments, a media player element utilizes one or more of many suitable multimedia technologies including, by way of non-limiting examples, Adobe®Flash®, HTML 5, Apple®QuickTime®, Microsoft®Silverlight®, Java™, and Unity®.

[0171] Referring to Fig.6, in a particular embodiment, an application provision system may comprise one or more databases 600 accessed by a relational database management system (RDBMS) 610. Suitable RDBMSs include Firebird, MySQL, PostgreSQL, SQLite, Oracle Database, Microsoft SQL Server, IBM DB2, IBM Informix, SAP Sybase, Teradata, and the like. In this embodiment, the application provision system may further comprise one or more application severs 620 (such as Java servers, .NET servers, PHP servers, and the like) and one or more web servers 630 (such as Apache, IIS, GWS and the like). The web server(s) optionally expose one or more web services via app application programming interfaces (APIs) 640. Via a network, such as the Internet, the system provides browser-based and / or mobile native user interfaces.

[0172] Referring to Fig.7, in a particular embodiment, an application provision system may alternatively have a distributed, cloud-based architecture 700 and may comprise elastically load balanced, auto-scaling web server resources 710 and application server resources 720 as well as synchronously replicated databases 730.

[0173] In some embodiments, a computer program includes a mobile application provided to a mobile computing device. In some embodiments, the mobile application is provided to a mobile computing device at the time it is manufactured. In other embodiments, the mobile application is provided to a mobile computing device via the computer network described herein.

[0174] In view of the disclosure provided herein, a mobile application may be created by techniques known to those of skill in the art using hardware, languages, and development environments known to the art. Those of skill in the art will recognize that mobile applications are written in several languages. Suitable programming languages include, by way of non- limiting examples, C, C++, C#, Objective-C, Java™, JavaScript, Pascal, Object Pascal, Python™, Ruby, VB.NET, WML, and XHTML / HTML with or without CSS, or combinations thereof.

[0175] Suitable mobile application development environments are available from several sources. Commercially available development environments include, by way of non-limiting examples, AirplaySDK, alcheMo, Appcelerator®, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments are WSGR Docket No. 64100-744.601 available without cost including, by way of non-limiting examples, Lazarus, MobiFlex, MoSync, and Phonegap. Also, mobile device manufacturers distribute software developer kits including, by way of non-limiting examples, iPhone and iPad (iOS) SDK, Android™ SDK, BlackBerry®SDK, BREW SDK, Palm®OS SDK, Symbian SDK, webOS SDK, and Windows®Mobile SDK.

[0176] Those of skill in the art will recognize that several commercial forums are available for distribution of mobile applications including, by way of non-limiting examples, Apple®App Store, Google®Play, Chrome WebStore, BlackBerry®App World, App Store for Palm devices, App Catalog for webOS, Windows®Marketplace for Mobile, Ovi Store for Nokia®devices, Samsung®Apps, and Nintendo®DSi Shop.

[0177] In some embodiments, a computer program includes a standalone application, which is a program that is run as an independent computer process, not an add-on to an existing process, e.g., not a plug-in. Those of skill in the art will recognize that standalone applications are often compiled. A compiler is a computer program(s) that transforms source code written in a programming language into binary object code such as assembly language or machine code. Suitable compiled programming languages include, by way of non-limiting examples, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB .NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In some embodiments, a computer program includes one or more executable complied applications.

[0178] In some embodiments, the computer program may include a web browser plug-in (e.g., extension, etc.). In computing, a plug-in is one or more software components that add specific functionality to a larger software application. Makers of software applications support plug-ins to enable third-party developers to create abilities which extend an application, to support easily adding new features, and to reduce the size of an application. When supported, plug-ins enable customizing the functionality of a software application. For example, plug-ins are commonly used in web browsers to play video, generate interactivity, scan for viruses, and display particular file types. Those of skill in the art will be familiar with several web browser plug-ins including, Adobe®Flash®Player, Microsoft®Silverlight®, and Apple®QuickTime®. In some embodiments, the toolbar comprises one or more web browser extensions, add-ins, or add-ons. In some embodiments, the toolbar comprises one or more explorer bars, tool bands, or desk bands.

[0179] In view of the disclosure provided herein, those of skill in the art will recognize that several plug-in frameworks are available that enable development of plug-ins in various WSGR Docket No. 64100-744.601 programming languages, including, by way of non-limiting examples, C++, Delphi, Java™, PHP, Python™, and VB .NET, or combinations thereof.

[0180] Web browsers (also called Internet browsers) are software applications, designed for use with network-connected computing devices, for retrieving, presenting, and traversing information resources on the World Wide Web. Suitable web browsers include, by way of non- limiting examples, Microsoft®Internet Explorer®, Mozilla®Firefox®, Google®Chrome, Apple®Safari®, Opera Software®Opera®, and KDE Konqueror. In some embodiments, the web browser is a mobile web browser. Mobile web browsers (also called microbrowsers, mini-browsers, and wireless browsers) are designed for use on mobile computing devices including, by way of non- limiting examples, handheld computers, tablet computers, netbook computers, subnotebook computers, smartphones, music players, personal digital assistants (PDAs), and handheld video game systems. Suitable mobile web browsers include, by way of non-limiting examples, Google®Android®browser, RIM BlackBerry®Browser, Apple®Safari®, Palm®Blazer, Palm®WebOS®Browser, Mozilla®Firefox®for mobile, Microsoft®Internet Explorer®Mobile, Amazon®Kindle®Basic Web, Nokia®Browser, Opera Software®Opera®Mobile, and Sony®PSP™ browser.

[0181] In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or use of the same. In view of the disclosure provided herein, software modules are created by techniques known to those of skill in the art using machines, software, and languages known to the art. The software modules disclosed herein are implemented in a multitude of ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, a distributed computing resource, a cloud computing resource, or combinations thereof. In further various embodiments, a software module comprises a plurality of files, a plurality of sections of code, a plurality of programming objects, a plurality of programming structures, a plurality of distributed computing resources, a plurality of cloud computing resources, or combinations thereof. In various embodiments, the one or more software modules comprise, by way of non- limiting examples, a web application, a mobile application, a standalone application, and a distributed or cloud computing application. In some embodiments, software modules are in one computer program or application. In other embodiments, software modules are in more than one computer program or application. In some embodiments, software modules are hosted on one machine. In other embodiments, software modules are hosted on more than one machine. In further embodiments, software modules are hosted on a distributed computing platform such as a cloud computing platform. In some embodiments, software modules are hosted on one or more WSGR Docket No. 64100-744.601 machines in one location. In other embodiments, software modules are hosted on one or more machines in more than one location.

[0182] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases, or use of the same. In view of the disclosure provided herein, those of skill in the art will recognize that many databases are suitable for storage and retrieval of nucleic acid segment sequences or analysis thereof information. In various embodiments, suitable databases include, by way of non-limiting examples, relational databases, non-relational databases, object oriented databases, object databases, entity-relationship model databases, associative databases, XML databases, document oriented databases, and graph databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, Sybase, and MongoDB. In some embodiments, a database is Internet-based. In further embodiments, a database is web-based. In still further embodiments, a database is cloud computing-based. In a particular embodiment, a database is a distributed database. In other embodiments, a database is based on one or more local computer storage devices. III. KITS

[0183] Provided herein are kits related to the methods, compositions and systems described herein. In some embodiments, the kits may comprise a plurality of recognition elements, a plurality of detection polynucleotides, one or more buffers, one or more reagents, instructions for use, a manual, a protocol, or a combination thereof.

[0184] Provided herein are kits related to the methods, compositions, and systems described herein. The kits may comprise a plurality of recognition elements. Each recognition element in the plurality of recognition elements may comprises a 3’ end. The 3’ end may be complementary to a genetic variant or the wildtype sequence of the genetic variant of a pharmacogenomics related gene target. The kits may comprise a ligase. The kits may comprise a DNA polymerase. The kits may comprise an exonuclease. The kits may comprise a third oligonucleotide. The kits may comprise a plurality of detection polynucleotides. The kits may comprise instructions for practicing any one of the methods disclosed herein.

[0185] In some embodiments, a kit may comprise one or more buffers. In some embodiments, a kit may comprise two or more buffers. In some embodiments, a first buffer of a kit may be configured to promote hybridization. In some embodiments, a second buffer of a kit may be configured to promote de-hybridization, ligation, nucleic acid digestion, storage of a purified molecule, or combinations thereof. In some embodiments, a kit may comprise one or more reagents. In some embodiments, a kit comprises one or more enzymes. In some embodiments, a kit may comprise one or more of a ligase, a DNA polymerase, an exonuclease, or combinations thereof. In some embodiments, a kit may comprise instructions for use, a manual, a protocol, or WSGR Docket No. 64100-744.601 a combination thereof. In some embodiments, a kit may comprise one or more well plates, for example, one or more 96 well plates. In some embodiments, one of the 96 well plates of a kit may be configured to be assayed by an optical imaging device described herein. Detection of variant targets of interest

[0186] Genetic mutations, including single nucleotide polymorphisms (SNPs), insertions, deletions, copy number variations (CNVs), and the like, can affect how gene products are expressed. Some genetic mutations may be silent, such that a nucleotide change in a sequence may not affect the coding of the subsequent protein. However, many genetic mutations may affect how a gene protein is built and expressed, to the detriment of the subject.

[0187] Many drugs, for example lipophilic psychotropic drugs, may need to be metabolized before they are excreted from the body. Variants in the gene(s) that code for an enzyme that metabolizes drugs, such as psychotropic drugs, may be important targets in the field of pharmacogenomics. Pharmacogenomics has the potential to greatly improve therapeutic management regimens by increasing the efficacy of drugs and reducing their toxicity by understanding the underlying mechanisms and variables that affect drug metabolism.

[0188] As an example, the gene CYP2D6 produces an enzyme that metabolizes at least 100 drugs, and mutations in the CYP2D6 gene affects that metabolism. Drugs that are known to be metabolized by CYP2D6 include, but are not limited to, Metoprolol, Propranolol, Timolol, Encainide, Flecainide, Perhexilene, Propafenone, Sparteine, Amitriptyline, Clomipramine, Desipramine, Fluoxetine, Fluvoxamine, Imipramine, Mianserin, Nortriptyline, Paroxetine, Venlafexine, Haloperidol, Perphenazine, Risperidone, Thioridazine, Zuclopenthixol, Codeine, Debrisoquine, Dextromethorphan, Phenoformin, Tolterodine and Tramadol. As such, CYP2D6 may be an important gene to target to determine the existence of a variant in a sample from a patient when a drug treatment regimen, for example based on one or more of the aforementioned drugs, is being considered. Other cytochrome P450 genes may also be important pharmacogenomic targets for drug metabolism and may therefore be equally important targets for studying variants that may impact drug metabolism. Table 1 below lists a collection of gene targets that may be targeted in methods of the present disclosure, where variants of the genes may be important to identify in support of personalized health initiatives. WSGR Docket No. 64100-744.601 Table 1-Examples of gene targets and variant information WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601

[0189] Some genetic variants that may be relevant to the field of pharmacogenomics can be described by utilizing a special nomenclature, wherein a mutation is referred to as a “star allele”. In this nomenclature system, alleles are not identified by their cDNA or genomic position, but through the means of numbers and letters, separated from the gene name by a star*. Information on the star allele nomenclature system can be found at the Clinical Pharmacogenetics Implementation Consortium (https: / / cpicpgx.org / ).

[0190] Star alleles may be referred to as haplotype patterns at the gene level and may be associated with protein activity levels. A haplotype can have genetic variants including SNPs, indels, or copy number variants (CNVs). Knowing the combination of variants within a given haplotype, together with the diploid content in an individual, may provide important information for studying drug metabolism, drug response and adverse drug reactions.

[0191] For example, CYP3A5*3 identifies the haplotype of CYP3A5 gene comprising variant rs776746 at genomic position g.99672916T>C, which leads to a splicing defect, impacting protein function. The star allele nomenclature was initially used to identify alleles with the cytochrome P450 or CYP gene family and afterwards spread to other genes studied in pharmacogenomics.

[0192] A wildtype allele in a gene, also called an “extensive metabolizer” corresponds to *1. The numbers *2, *3, *4 etc. represent alleles with altered functionality which may lead to profiles of increased or reduced drug metabolism. More information on star alleles can be found in the PGx annotation database from 1000 Genomes Project at www.coriell.org / starallele / search. Examples of star alleles are found in Table 2 below. Table 2-Example of star alleles and their variants WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601

[0193] In some embodiments, a workflow of the present disclosure for identifying one or more variant target molecules of interest and / or variants therefrom, from a sample may include the use of one or more instruments. For example, an automated liquid handling system can be designed to perform a series of operations in a method for a relatively hands off workflow. By way of example, the company Hamilton provides a wide range of laboratory automated liquid handling systems which may be useful in performing operations in the disclosed methods. In some embodiments, an automated liquid handling system may perform any operation from nucleic acid extraction through concatemeric amplification of the presently disclosed methods. An automated liquid handling system may extract DNA from a sample, or a plurality of samples, hybridize the target molecules to recognition elements, ligate the recognition elements, remove linear DNA from the circularized recognition elements, amplify the circularized recognition elements, and hybridize the detection polynucleotides. In some embodiments, a fluorescent microscopy plate reader can be utilized to image the fluorescent concatemeric amplification products, decode the images to determine the presence of a target of interest, and output results. Such a workflow, which is largely automated with little hands-on time from a technician, can provide a seamless integration between the different operations of the workflow and imaging of all samples simultaneously thereby providing an accelerated path to actionable results.

[0194] In some embodiments, instruments may not be needed to perform the workflow. For example, a sample such as a buccal swab, a saliva sample, a blood sample, a tissue sample, etc. can be obtained and the nucleic acids may be extracted and optionally purified to generate a nucleic acid sample in solution. The extracted nucleic acids may be aliquoted to separate wells or tubes and recognition elements that include complementary 5’ and 3’ ends to the target variant in a sample can be added to the separate well or tube. Additionally, recognition elements that include 5’ and 3’ end sequences that are complementary to the wildtype sequence of the WSGR Docket No. 64100-744.601 variant of interest can also be added to the well or tube. The recognition elements that target both the variant and the wildtype comprise a different code, such that the code regardless of the presence of the variant or wildtype sequence of the target can be aligned with the target sequence. Hybridization between the target nucleic acid may result in a circularized recognition element, for example a padlock probe configuration. The circularized recognition elements can be ligated to yield a circular, ligated recognition element comprising the complements of the target nucleic acid sequence, a code unique to the target sequence, and other sequence that facilitate amplification. In some embodiments, amplification of the circular, ligated recognition element can be performed to yield a concatemeric amplification product.

[0195] In some embodiments, the amplification can be performed in solution. In some embodiments, the amplification can be performed when circular and ligated recognition elements are immobilized in one or more wells of a 12, 24, 48, 96, 384, etc. plate. In some embodiments, wells of a plate can be precoated with a compound. The compound may be configured to immobilize nucleic acids on the wells of the plate. The circular and ligated recognition elements can be amplified to generate concatemeric amplification products and the code of the concatemeric amplification products can be queried with labeled detection polynucleotides, and the detectable label can be detected for example via fluorescence imaging. Images can be subsequently decoded to identify the target variant nucleic acid that was present in the original sample based on the associated code. When a target nucleic acid is not present, there may be no circularization and ligation and therefore, there may be nothing to amplify. The code from the recognition element may serve as a proxy of the presence of the target molecule. Decoding can be performed using soft decision decoding as previously described.

[0196] Fig.8 demonstrates examples of recognition elements and their 5’ and 3’ ends directed to identifying different variant targets such as single nucleotide polymorphisms (SNPs), insertions and deletions or their wildtype sequences. For SNP detection, the 3’ end can include the specific variant complement of the target SNP. For known insertions and deletions, the recognition elements may include one or both 5’ and 3’ ends specific for the nucleotides that are the inserted sequences, or a subset thereof, or sequences that may remain in the target after sequences are deleted.

[0197] In some embodiments, a target of interest comprises a single nucleotide polymorphism (SNP). In some embodiments, a recognition element that recognizes a single nucleotide polymorphism in a target of interest comprises a 3’ end region wherein the 3’ end nucleotide is complementary to a SNP of interest in the target. In some embodiments, a recognition element may recognize a wildtype nucleotide in lieu of a single nucleotide polymorphism in a target of interest, wherein the 3’ end sequence of the 3’ end region comprises a nucleotide that is WSGR Docket No. 64100-744.601 complementary to the wildtype sequence of interest in the target of interest at a particular SNP site. In some embodiments, when the wildtype nucleotide sequence is present in the nucleic acid of interest the recognition element comprising the complementary wildtype 3’ end sequence may hybridize to the target sequence of interest such that ligation of the 5’ and 3’ ends of the recognition element is possible. In other embodiments, when a SNP is present in the target of interest a recognition element comprising a 3’ end sequence that is complementary to the SNP in the target sequence may hybridize to the SNP of the target sequence of interest such that ligation of the 5’ and 3’ ends of the recognition element is possible. In preferred embodiments, recognition elements that comprise the complementary wildtype nucleotide at the 3’ end at a known SNP location and recognition elements that comprise the complementary variant nucleotide at the 3’ end at the known SNP location may both be present in the assay, such that both wildtype and variant targets of interest can be detected. In some embodiments, for those recognition elements that are ligated, amplification can occur wherein multiple copies of the circularized and ligated recognition elements are generated to produce concatemeric amplification products which comprise copies of the variant or wildtype target sequence of interest, or complements thereof, and copies, or complements thereof, of the recognition element code which is aligned with the variant target sequence of interest or the wildtype sequence of interest, thereby correlating the presence (or absence) of the variant and / or the wildtype target sequence of interest from the sample.

[0198] In some embodiments, a target of interest comprises an insertion of one or more nucleotides. The one or more nucleotides may be at or in proximity to the target sequence of interest. In some embodiments, a recognition element that recognizes an insertion of one or more nucleotides at or in proximity to a target of interest comprises a 3’ end region wherein the 3’ end comprises one or more complementary sequences to the one or more inserted nucleotides of interest in the target of interest. In some embodiments, a recognition element may recognize a wildtype nucleotide in lieu of the one or more inserted sequences in a target of interest, wherein the 3’ end sequence of the 3’ end region comprises the wildtype sequence, or a complement thereof, in the target of interest at or in proximity to the variant insertion site at the target of interest. In some embodiments, when the wildtype nucleotide sequence is present in the target nucleic acid of interest the recognition element comprising the complementary wildtype 3’ end sequence may hybridize to the wildtype target sequence of interest such that ligation of the 5’ and 3’ ends of the recognition element is possible. In other embodiments, when one or more nucleotides of an inserted variant sequence is present in the target of interest a recognition element comprising a 3’ end sequence that is complementary to the one or more nucleotides of the inserted variant sequence, or a subset thereof, in the target sequence hybridizes to the WSGR Docket No. 64100-744.601 inserted variant sequence of the target sequence of interest such that ligation of the 5’ and 3’ ends of the recognition element is possible. In some embodiments, for those recognition elements that are ligated, amplification can occur wherein multiple copies of the circularized and ligated recognition elements are generated to produce concatemeric amplification products which comprise copies of the variant or wildtype target sequence of interest, or complements thereof, and copies, or complements thereof, of the recognition element code which is aligned with the variant or wildtype target sequence of interest, thereby correlating the presence (or absence) of the target variant or wildtype sequence of interest from the sample. In preferred embodiments, recognition elements that comprise the complementary wildtype sequence (e.g., if no insertion has occurred) at the 3’ end at a target insertion location and recognition elements that comprise the complementary insertion nucleotides at the 3’ end at the target insertion location are both present in the assay, such that both wildtype and variant targets of interest can be detected.

[0199] In some embodiments, a target of interest comprises a deletion of one or more nucleotides. The deletion of one or more nucleotides may be at or in proximity to the target sequence of interest. In some embodiments, a recognition element that recognizes a deletion of one or more nucleotides at or in proximity to a target of interest may comprise a 3’ end region wherein the 3’ end comprises one or more complementary sequences to one or more nucleotides that remain in the variant target of interest after a deletion event has occurred. In some embodiments, a recognition element may recognize the one or more deleted sequences from a target of interest, in effect the wildtype sequence with no deletion, wherein the 3’ end sequence of the 3’ end region comprises the wildtype sequence, or a complement thereof, in the target of interest. In some embodiments, when the wildtype nucleotide sequence is present in the target nucleic acid of interest, for example when there has been no deletion event, the recognition element comprising the complementary wildtype 3’ end sequence hybridizes to the wildtype target sequence of interest such that ligation of the 5’ and 3’ ends of the recognition element is possible. In other embodiments, when one or more nucleotides are deleted in the target of interest a recognition element comprising a 3’ end sequence that is complementary to one or more nucleotides adjacent to the deleted sequence, or a subset thereof, in the target sequence may hybridizes to the adjacent sequences post deletion event of the target sequence of interest such that ligation of the 5’ and 3’ ends of the recognition element is possible. In some embodiments, for those recognition elements that are ligated, amplification can occur wherein multiple copies of the circularized and ligated recognition elements are generated to produce concatemeric amplification products which comprise copies of the variant or wildtype target sequence of interest, or complements thereof, and copies, or complements thereof, of the WSGR Docket No. 64100-744.601 recognition element code which is aligned with the target variant or wildtype sequence of interest, thereby correlating the presence (or absence) of the target variant or wildtype sequence of interest from the sample. In preferred embodiments, recognition elements that comprise the complementary wildtype sequence (e.g., if no deletion has occurred) at the 3’ end at a target deletion location and recognition elements that comprise the sequence if the deletion is present at the 3’ end at the target deletion location are both present in the assay, such that both wildtype and variant targets of interest can be detected.

[0200] In some instances, there may be two or more variants in a target of interest that are in close proximity to each other. For example, locations in a sequence where there are two or more variants that are within a few nucleotides of each other. In some embodiments, there are two variants that are at least one nucleotide, at least two nucleotides, a least three nucleotides, at least four nucleotides, or at least five nucleotides apart. In some embodiments, there are three variants that are in proximity to each other at a location on a sequence of interest. Multiple variants in proximity to each other on a sequence of interest may be referred to as “hot spots” of variants. In some embodiments, the two or more variants are known, while in other embodiments one or more of the variants are not known. Indeed, in some embodiments when there are two variants in close proximity at a sequence of interest, one may be the variant of interest while the other may not be known. As such, a recognition element that recognizes the variant of interest, but is not designed to recognize the second variant may not hybridize to the target variant due to the mismatched nucleotide that is in close proximity to the end of the recognition element. A mismatch in the recognition element sequence that targets the variant of interest may affect hybridization and / or ligation of the target of interest to the recognition element complementary sequence. In some embodiments, the unknown variant can be an insertion, a deletion or a single nucleotide polymorphism (SNP), or simply a different base that was miscalled in a sequence database that may affect the hybridization and / or ligation of the recognition element thereby yielding a false negative for the presence of the sequence variant of interest. For example, a recognition element may be designed such that the complementary sequence of a target of interest, for example a single nucleotide polymorphism (SNP), may be created at the 3’ end of the recognition element.

[0201] However, in some embodiments, for example, three nucleotides away from the target SNP may be an unknown SNP that was not designed into the 3’ end of the recognition element. In this example scenario, the hybridization of the target sequence of interest to the recognition element may be weak or non-existent due to the unknown mismatch, and the 5’ and 3’ ends may not be adequately brought into proximity for subsequent ligation yielding little to no ligation product and no detection of the target SNP of interest. In order to correct for this event, in some WSGR Docket No. 64100-744.601 embodiments, recognition elements can be designed to include a degenerate base at one or more locations within the 3’ end or the 5’ ends of the recognition element, such that all potential mismatches may be corrected and hybridization and ligation can occur. The mismatch correction can be designed into the recognition elements that specifically target the SNP of interest, and / or the recognition elements that target the wildtype of the target SNP of interest. In some embodiments, the knowledge of the presence of these potential mismatches can be enhanced by researching the variants in the target sequence of interest and determining the possibility of the presence of high or low frequency variants in the sequences of interest, and designing the recognition elements with that knowledge in mind, such that multiple recognition elements can be designed to account for the multiple variants that may be present within the region of the variant of interest. Designs may include one or more of the incorporation of degenerate nucleotides at know locations, incorporation of pseudo-degenerate (e.g., using a subset of the canonical nucleotides, for example, A and G, or C and T, etc.) nucleotides at known locations, incorporation of degenerate or pseudo-degenerate nucleotides at random locations in the 5’ or 3’ regions of the recognition elements, in combination with the wildtype of variant targeting sequences that may be designed into the recognition elements.

[0202] Within the scenario of designing recognition elements to recognize additional, either known or unknown mismatches, it is assumed that all the permutations of the recognition elements that are designed to identify the variant sequence of interest may have a unique code, such that the variant sequence of interest and the mismatches may all be collated during data analysis for identifying the variant of interest in addition to the mismatches that are in proximity to the variant of interest. In some embodiments, one variant of interest is targeted, however in other embodiments an underlying variant is also of interest, such that the variant of interest in combination with an underlying variant represents a known star allele.

[0203] One example of genes with high variability and hence the possibility of inherent variability across a population are the Human Leukocyte Antigen (HLA) genes. Disclosed herein are additional methods for identifying HLA variants of interest, however the scenario of designing multiple recognition elements as described may be used alone or in combination with the additional HLA methods as described herein. Another example is located at chr 6:29974074, which is in the tumor necrosis factor (TNF) gene, which demonstrates the application of designing probes to mismatches to correct miscalling of variants (Fig.24).

[0204] In some embodiments, nucleic acid sequences that share high homology with a target gene of interest, for example homologs or pseudogenes, that are present in a sample with a target sequence of interest in a gene of interest may be problematic for identifying the target sequence of interest from the gene of interest in a sample. It can be difficult to identify a target of interest, WSGR Docket No. 64100-744.601 such as a variant or wildtype sequence of interest, when a gene shares high homology to other sequences. CYP2D6 is an example of a polymorphic gene that has high homology to the homolog CYP2D7 gene and the homolog CYP2D8 gene and surrounding sequences.

[0205] Cytochrome P4502D6, or CYP2D6, is involved in the metabolism of many drugs including, but not limited to psychotropic medications as previously described. As such, it is important to be able to identify variants of this gene for pharmacogenomics. Identification can be hampered by the high sequence similarity CYP2D6 shares with the aforementioned homologs, in part because of recombination events that may generate duplications, multiplications, deletions and gene conversions in this gene family. The highly polymorphic nature of CYP2D6 has resulted in over 100 star alleles for this gene alone, further underscoring the importance of this gene and the desire to identify its gene variants with accuracy and precision, in support of personalized health with regards to drug metabolism.

[0206] In some embodiments, the identification of variants of CYP2D6 in a background of potential homolog noise can be accomplished by practicing a targeted hybridization scenario as demonstrated in Fig.16A. In some embodiments, a dual ligation scenario comprising a recognition element may be utilized in identifying a target of interest from a homolog. In some embodiments, a dual ligation scenario can further comprise utilizing a gap fill method in combination with a dual ligation scenario comprising a recognition element in identifying a target of interest from a homolog. In some embodiments, a target of interest comprises a SNP of interest. In some embodiments, a target of interest is important in the field of pharmacogenomics. In some embodiments, the 3’ end of the 3’ region of a recognition element comprises the complementary sequence of the SNP of the target of interest. In further embodiments, in combination with the 3’ end comprising a complementary sequence of a SNP of interest the recognition element further comprises a differentiating base on the 5’ end of the 5’ region of recognition element, wherein the differentiating base hybridizes to a target sequence of interest and not a homolog of the target of interest. As such, a recognition element that can be used to differentiate a highly polymorphic target of interest from close homologs comprises two sequences specific to the target of interest, a complementary variant sequence and a complementary differentiating base sequence, wherein the two sequences hybridize the target of interest over that of a close homolog.

[0207] In some embodiments, the two specific 3’ and 5’ sequences of the recognition element may hybridize the target of interest while leaving a gap between the 5’ end and the 3’ end of the recognition element as it is hybridized to the target of interest, in that the 5’ end and the 3’ end are not adjacently located post hybridization to the target of interest. In some embodiments, a gap fill can be performed such that a third oligonucleotide complementary to target sequence of WSGR Docket No. 64100-744.601 interest sequence that comprises the gap can be added to the reaction well, tube, etc. such that the third oligonucleotide hybridizes to the target sequence that comprises the gap, thereby filling the gap between the 5’ and 3’ ends of the hybridized recognition element. When the gap is filled, the 5’ end and the 3’ end of the third oligonucleotide can be ligated and the 5’ end of the third oligonucleotide and the 3’ end of the recognition element can be ligated, for a double ligation event, thereby adding a level of specificity in circularizing and ligating the recognition element for downstream applications. As seen in Fig.16A, the differentiating base is specific to the target gene of interest and not the homolog, even if the SNP is shared by the target gene of interest and the homolog. This specificity allows for hybridization of the target gene of interest to the recognition element, whereas the homolog is not expected to hybridize as such no ligation and circularization of the recognition element if the homolog hybridizes to the variant sequence (Fig.16B). This is the case with or without gap fill, as the 5’ end sequence, in the absence of hybridization, is not conducive to ligation.

[0208] In some embodiments, the targets of interest can be extracted from a sample of interest, for example a buccal swab, a tissue sample such as a biopsy sample, a saliva sample, a blood sample, etc. In some embodiments, nucleic acids such as DNA and / or RNA are extracted and optionally purified from a sample. In some embodiments, an extracted and or optionally purified nucleic acid sample is combined with targeted recognition elements as described herein. In some embodiments, one or more target sequences of interest are amplified from an extracted and optionally purified nucleic acid sample prior to combining with targeted recognition elements. In some embodiments, target sequences of interest are amplified from a sample of interest wherein the amplified products comprise the targets of interest in addition to any homologs that share homology with the target of interest. In some embodiments, target sequences of interest are amplified from a sample of interest such that any homolog amplified products are significantly reduced or eliminated. If a target of interest is amplified in conjunction with homologs, a double ligation scenario as previously discussed and as shown in Fig.16A can be employed to specifically detect a target of interest in the presence of potential high homology homolog amplification.

[0209] In some embodiments, a third oligonucleotide, or a bridge oligonucleotide, can be between about 10 and 300 nucleotides long. In some embodiments, a third oligonucleotide can be between about 20-200 nucleotides, between about 40-100 nucleotides, or between about 50- 75 nucleotides long. In some embodiments, a third oligonucleotide can be between about 10-50 nucleotides long, between about 20-40 nucleotides long, or between about 20-35 nucleotides long. In some embodiments, a third oligonucleotide is included in the reaction vessel concurrent with a target of interest and one or more recognition elements, such that hybridization of the WSGR Docket No. 64100-744.601 target of interest to its complementary sequence(s) in the one or more recognition elements and hybridization of the third oligonucleotide and the target sequences occurs for the most part simultaneously contrary to a step wise hybridization scenario. In some embodiments, a third oligonucleotide is added to the reaction vessel after a target of interest and one or more recognition elements are added to the reaction vessel, such that hybridization of the target of interest and the one or more recognition elements occurs in a first operation, followed by addition of a third, or bridge, oligonucleotide and hybridization of the third oligonucleotide to its target sequence in a second operation.

[0210] The present disclosure provides several scenarios for identifying at least two variants using a recognition element which comprises two target sequences for variant detection, for example as found in Fig.19A or whenever a target of interest comprises at least two variants in need of identification in the target of interest. In some embodiments, identification of two variants in a target of interest additionally comprises hybridization of the targeted variants with the recognition element, wherein the hybridization leaves a gap between the 5’ and 3’ ends of the recognition element. In this scenario, a third oligonucleotide, a bridge oligonucleotide, is utilized to fill the space between the hybridized 5’ and 3’ ends of the recognition element to the targeted variants. In some embodiments, hybridization of two variant sequences to a recognition element causes a loop out of the nucleic acids that are between the two targeted variants, for example when the two target variants are more than 10, more than 20, more than 50, more than 100, more than 200 nucleotides apart on a target molecule. In some embodiments, a third oligonucleotide or bridge oligonucleotide can be at least 10 nucleotides long, at least 20 nucleotides long, at least 30 nucleotides long, at least 40 nucleotides long, at least at least 50 nucleotides long, at least 60 nucleotides long, at least 70 nucleotides long, at least 80 nucleotides long, at least 90 nucleotides long, or at least 100 nucleotides long. In some embodiments, a third oligonucleotide or bridge oligonucleotide can be between about 30-100 nucleotides in length, between about 40-90 nucleotides in length, between about 50-80 nucleotides in length, or between about 60-70 nucleotides in length. Fig.19A, Fig.19B and Fig.20 are non-limiting examples of different gap filling scenarios comprising the identification of two variants in a target of interest from a sample of interest.

[0211] In one embodiment, a target sequence of interest comprises two variants wherein the identification of the presence of each of the two variants may be desired. In some embodiments, both variants in the target sequence of interest may be relevant to drug metabolism. In other embodiments, the target sequence of interest includes a variant which is not relevant to drug metabolism. In some embodiments, the two variants are not adjacently located on the target of interest. In some embodiments, a target sequence of interest comprises more than two variants of WSGR Docket No. 64100-744.601 interest, for example it may be desired to identify the presence of more than or equal to two, three, four, five, six, seven, eight, nine or ten variant sequences of interest in a target of interest. In some embodiments, when a target of interest comprises two variant sequences, such as two single nucleotide polymorphisms (SNPs) of interest, a recognition element comprises a 5’ end of a 5’ region that is complementary to a first variant sequence and a 3’ end of a 3’ region that is complementary to a second variant sequence of the target of interest. In some embodiments, the first and second complementary sequences of interest in the recognition element hybridize to a first and second variant sequences in the target of interest, wherein the hybridization leaves a gap between the 5’ end and the 3’ end of the hybridized recognition element. A third oligonucleotide, complementary to the target sequence of interest that is located between the first and second variant sequences in the target of interest, can be hybridized to the target sequence of interest between the first and second variant sequences thereby filling the gap between the 5’ end and the 3’ end of the recognition element. In some embodiments, a double ligation event occurs such that the 5’ end of the recognition element and the 3’ end of the third oligonucleotide, and the 3’ end of the recognition element and the 5’ end of the third oligonucleotide are ligated together, thereby circularizing the recognition element for downstream applications.

[0212] In some embodiments, a target sequence of interest may comprise two variants wherein the identification of the presence or absence of each of the two variants may be desired. In some embodiments, the two variants are not adjacently located on the target sequence of interest. In some embodiments, a series of nucleotides on the target sequence of interest exist between the first and second variant. In some embodiments, a recognition element for use in identifying two variants of interest in a target of interest comprises a 5’ end of a 5’ region of a recognition element that is complementary to a nucleotide that is adjacently upstream of a first variant sequence of interest and a 3’ end of a 3’ region of a recognition element that is complementary to a second variant sequence of interest. In some embodiments, the target sequence of interest hybridizes to the 5’ end sequence and the 3’ end sequence, wherein the hybridization leaves a gap between the hybridized 5’ and 3’ end sequences of the recognition element due to the series of nucleotides on the target sequence that exist between the first and second variants. In some embodiments, a third oligonucleotide is hybridized to the sequence of the target of interest that is between the first and second variant sequences, thereby filling the gap between the 5’ and 3’ end sequences of the recognition element.

[0213] In some embodiments, a complementary sequence to a first variant sequence may be located on the 3’ end of the third oligonucleotide that is adjacent to the 5’ end of the recognition element, such that the first variant sequence of the target of interest can hybridize to the WSGR Docket No. 64100-744.601 complementary 3’ end of the third oligonucleotide. Concurrently, the 5’ end sequence of the third oligonucleotide that is adjacent to the 3’ end of the recognition element comprises a sequence that is adjacently located to the second variant sequence of interest, such that the 5’ end of the third oligonucleotide hybridizes to a sequence adjacently located to the second variant sequence in the target of interest which is hybridized or can hybridize to the 3’ end of the 3’ end region of the recognition element comprising a complement of the second variant sequence. The hybridization of the third oligonucleotide fills the gap between the 5’ and 3’ ends of the recognition element such that the 5’ end of the recognition element and the 3’ end of the third oligonucleotide and the 5’ end of the third oligonucleotide and the 3’ end of the recognition element can be ligated in a double ligation event generating a circularized and ligated recognition element available for downstream applications.

[0214] In some embodiments, a target sequence of interest may comprise two variants wherein the identification of the presence or absence of each of the two variants is desired. In some embodiments, the two variants are not adjacently located on the target sequence of interest. In some embodiments, a series of nucleotides on the target sequence of interest exist between the first and second variant. In some embodiments, a recognition element for use in identifying two variants of interest in a target of interest comprises a 5’ end of a 5’ region of a recognition element that is complementary to a first variant sequence of interest and a 3’ end of a 3’ region of a recognition element that is complementary to a nucleotide adjacently located to a second variant sequence of interest. In some embodiments, the target sequence of interest hybridizes to the 5’ end sequence and the 3’ end sequence, wherein the hybridization leaves a gap between the hybridized 5’ and 3’ end sequences of the recognition element due to the series of nucleotides on the target sequence that exist between the first and second variants. In some embodiments, a third oligonucleotide is hybridized to the sequence of the target of interest that is between the first and second variant sequences, thereby filling the gap between the 5’ and 3’ end sequences of the recognition element.

[0215] In some embodiments, a complementary sequence to a sequence directly adjacent to the first variant of interest may be located on the 3’ end of the third oligonucleotide that is adjacent to the 5’ end of the recognition element, such that the sequence adjacently located to the first variant of the target of interest can hybridize to the complementary 3’ end of the third oligonucleotide. Concurrently, the 5’ end sequence of the third oligonucleotide that is adjacent to the 3’ end of the recognition element comprises a second variant sequence, such that the 5’ end of the third oligonucleotide hybridizes to the second variant sequence in the target of interest which is hybridized or can hybridize to the 3’ end of the 3’ end region of the recognition element comprising a sequence adjacently located to the second variant of interest. The WSGR Docket No. 64100-744.601 hybridization of the third oligonucleotide fills the gap between the 5’ and 3’ ends of the recognition element such that the 5’ end of the recognition element and the 3’ end of the third oligonucleotide and the 5’ end of the third oligonucleotide and the 3’ end of the recognition element can be ligated in a double ligation event generating a circularized and ligated recognition element available for downstream applications.

[0216] In some embodiments, a target sequence of interest may comprise two variants wherein the identification of the presence or absence of each of the two variants may be desired. In some embodiments, the two variants are not adjacently located on the target of interest. In some embodiments, a series of nucleotides on the target of interest exist between the first and second variant. In some embodiments, the two variants are located at a distance on the target of interest such that a third oligonucleotide hybridizes to a portion of the sequence on the target of interest that exists between the first variant and the second variant, thereby generating a loop structure in the target of interest that is not hybridized to a third oligonucleotide for gap filling as previously described. In some embodiments, a first variant is located at the 5’ end of the 5’ region on a recognition element and a second variant is located at the 3’ end of the 3’ region of a recognition element. In some embodiments, the first variant in the target of interest hybridizes to the 5’ end of the recognition element and the second variant in the target of interest hybridizes to the 3’ end of the recognition element. In some embodiments, a third oligonucleotide comprising sequences complementary to adjacent sequences to the first and second variant is the target of interest is hybridized to the target of interest, such that said hybridization results in a series of nucleotides that is not hybridized to either end of the recognition element and is not hybridized to the third oligonucleotide. In some embodiments, after hybridization the 3’ end of the third oligonucleotide is adjacent to the 5’ end of the recognition element and the 5’ end of the third oligonucleotide is adjacent to the 3’ end of the recognition element such that two ligation events are able to occur, resulting in circularization and ligation of the recognition element and third oligonucleotide which can be used for downstream applications. In the absence of either the first variant or the second variant, or both, in the target of interest it is contemplated that there may be no hybridization at the site where the variant(s) is absent, as such no circularization and no ligation at the one or more junctions.

[0217] In some embodiments, a first complementary variant sequence of a target of interest may not be located at the 5’ end of the 5’ region of a recognition element but may instead be located at the 3’ end of a third oligonucleotide. In some embodiments, a second complementary variant sequence of a target of interest is not located at the 3’ end of the 3’ region of a recognition element but is instead located at the 5’ end of a third oligonucleotide. In some embodiments, when a first complementary variant sequence is located at the 3’ end of a third oligonucleotide, WSGR Docket No. 64100-744.601 the second complementary sequence is located at the 3’ end of the 3’ region of a recognition element. In some embodiments, when a second complementary variant sequence is located at the 5’ end of a third oligonucleotide the first complementary variant sequence is located at the 5’ end of the 5’ region of a recognition element. In some embodiments, if there are two variant sequences of interest in a target of interest one complementary variant sequence can be located at either both ends of a recognition element, or one complementary variant sequence is located at one of the ends of the recognition element while the second complementary variant sequence is located at the distal end of the third oligonucleotide, relative to the location of the first complementary variant sequence in the recognition element.

[0218] In some embodiments, a first and a second variant sequence are located on a target of interest, wherein the variant sequences are located a number of nucleotides apart. In some embodiments, a first and a second variant sequence are located on a target of interest, wherein the variant sequences are located at least 5 nucleotides apart, at least 10 nucleotides apart, at least 15 nucleotides apart, at least 20 nucleotides apart, at least 30 nucleotides apart, at least 40 nucleotides apart, at least 50 nucleotides apart, at least 60 nucleotides apart, at least 70 nucleotides apart, at least 80 nucleotides apart, at least 100 nucleotides apart, at least 200 nucleotides apart, at least 300 nucleotides apart, or at least 500 nucleotides apart. In some embodiments, a first and a second variant sequence are located on a target of interest, wherein the variant sequences are located between about 5-500 nucleotides apart, between about 10-400 nucleotides apart, between about 20-300 nucleotides apart, between about 30-200 nucleotides apart, between about 40-100 nucleotides apart, or between about 20-50 nucleotides apart.

[0219] Human Leukocyte Antigens (HLAs) are highly polymorphic, with different HLAs having different alleles. HLA testing and typing has been used for determining histocompatibility for organ transplantation to mitigate the risk of organ rejection. However, some HLA alleles are also known as star alleles and have been associated with drug metabolism and adverse drug reactions with certain medications, as such PGx testing of HLA alleles is also disclosed herein.

[0220] Some adverse reactions with HLA mutations include severe cutaneous adverse reactions (SCARs) and drug induced liver injury (DILI). For example, HLA-B*57:01 mutant is associated with Abacavir cutaneous hypersensitivity and Flucloxacillin induced liver injury. HLA-B*58:01 is associated with Allopurinol induced SCARs and Nevirapine induced liver injury along with HLA-DRB*01. HLA-A*31:01 is associated with Carbamazepine induced SCARs, HLA-A*13- 01 is associated with Dapsone hypersensitivity in leprosy patients and SCARS in non-leprosy patients. WSGR Docket No. 64100-744.601

[0221] Additionally, variants within the HLA genes may play a significant role in determining susceptibility to various diseases such as autoimmune diseases and conditions. The intricate linkage disequilibrium (LD) structures and long-range haplotypes of regional variants, along with the challenges posed by sequencing, may have made fine-mapping of causal variants in this region difficult, especially for large cohorts.

[0222] To address this challenge, HLA imputation provides a method to predict HLA types from regional single SNPs. Utilizing a reference panel, such as GRCh38 or other version, with both HLA and SNP genotypes, HLA imputation leverages LD structures and SNP information to infer HLA genotypes accurately.

[0223] In this disclosure, specific HLA star alleles are identified rather than typing all possible HLA alleles as performed when do doing HLA imputation. This disclosure provides methods for direct genotyping of HLA SNPs and using the information from direct genotyping to identify HLA star alleles. In some embodiments, both HLA imputation and direct genotyping of HLA SNPs are combined in identifying HLA star alleles.

[0224] HLA imputation may begin with identifying intergenic SNPs that exhibit strong LD with the alleles of interest. Combinations of SNPs may be prioritized based on LD with the target HLA allele. The performance of selected SNPs is validated using imputation algorithms through cross-validations. Once sets of SNPs are chosen for the desired HLA alleles, the alleles may be identified and reported using imputation algorithms alongside a reference panel, yielding imputation probabilities indicating the presence or absence of the desired HLA star alleles.

[0225] Conversely, direct genotyping for identifying specific sets of HLA star alleles (e.g., HLA-A*31:01, HLA-B*15:02, HLA-B*57:01 and HLA-B*58:01) may include querying HLA loci unique to each allele directly to distinguish them from other HLA star alleles and from the rest of the genome. These loci, whether single or in combination, may serve as targets for designing probes specific to the desired alleles. Genotyping results can then be converted into probability scores.

[0226] In preferred embodiments, the HLA results may be obtained by integrating both imputation and direct genotyping approaches. The combination rule for each star allele may depend on the accuracy of imputation and direct genotyping for that particular allele.

[0227] Samples for use in the present disclosure can derive from a variety of sources. For example, an HLA target molecule of interest can be genomic DNA, which may be derived from a buccal swab, a saliva, a blood, a tissue, a biopsy, cells, or the like. A sample may also be from a biobank, for example a tissue sample from which nucleic acids can be extracted or are already extracted and which their variants are known. Samples can be obtained from a known source, wherein the sample comprises nucleic acids in solution where the variant makeup is known. WSGR Docket No. 64100-744.601 IV. DEFINITIONS

[0228] Unless defined otherwise, all terms of art, notations and other technical and scientific terms or terminology used herein are intended to have the same meaning as is commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over what is generally understood in the art.

[0229] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.

[0230] As used in the specification and claims, the singular forms “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a sample” includes a plurality of samples, including mixtures thereof.

[0231] The terms “determining”, “measuring”, “evaluating”, “assessing”, “assaying”, and “analyzing” are often used interchangeably herein to refer to forms of measurement. The terms include determining if an element is present or not (for example, detection). These terms can include quantitative, qualitative or quantitative and qualitative determinations. Assessing can be relative or absolute. “Detecting the presence of” can include determining the amount of something present in addition to determining whether it is present or absent depending on the context.

[0232] The terms “subject”, “individual”, or “patient” are often used interchangeably herein. A “subject” can be a biological entity containing expressed genetic materials. The biological entity can be a plant, animal, or microorganism, including, for example, bacteria, viruses, fungi, and protozoa. The subject can be tissues, cells and their progeny of a biological entity obtained in vivo or cultured in vitro. The subject can be a mammal. The mammal can be a human. The subject may be diagnosed or suspected of being at high risk for a disease. In some cases, the subject is not necessarily diagnosed or suspected of being at high risk for the disease. WSGR Docket No. 64100-744.601

[0233] As used herein, the term “about” a number refers to that number plus or minus 10% of that number. The term “about” a range refers to that range minus 10% of its lowest value and plus 10% of its greatest value.

[0234] As used herein, the terms “treatment” or “treating” are used in reference to a pharmaceutical or other intervention regimen for obtaining beneficial or desired results in the recipient. Beneficial or desired results include but are not limited to a therapeutic benefit and / or a prophylactic benefit. A therapeutic benefit may refer to eradication or amelioration of symptoms or of an underlying disorder being treated. Also, a therapeutic benefit can be achieved with the eradication or amelioration of one or more of the physiological symptoms associated with the underlying disorder such that an improvement is observed in the subject, notwithstanding that the subject may still be afflicted with the underlying disorder. A prophylactic effect includes delaying, preventing, or eliminating the appearance of a disease or condition, delaying or eliminating the onset of symptoms of a disease or condition, slowing, halting, or reversing the progression of a disease or condition, or any combination thereof. For prophylactic benefit, a subject at risk of developing a particular disease, or to a subject reporting one or more of the physiological symptoms of a disease may undergo treatment, even though a diagnosis of this disease may not have been made.

[0235] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. V. EXAMPLES Example 1-Assay general workflow for identifying variants in target nucleic acids

[0236] This protocol is the general protocol used for generating concatemeric amplification products from targets of interest and recognition element hybridization events for identifying the presence of a target variant of interest in a biological sample. Variants of interest in the following examples followed this protocol unless otherwise stated.

[0237] Nucleic acids were extracted and / or purified from a biological sample. The nucleic acid sample, for example 100 ng per reaction, was added to a well of a 96 well plate. Approximately 15 µL of a master mix was added to nucleic acid sample for hybridization and ligation that includes 0.5 nM of linear recognition elements, a thermostable ligase (e.g., AmpLigase, Tth ligase, etc.), a buffer, any adjuvants for enzymatic enhancement as desired, and water for a total reaction volume of approximately 20 µL. The linear recognition elements included those that are targeting variants and their wildtype counterparts. Each recognition element may have a unique code so that both variant and wildtype targets can be assayed, detected and decoded. The hybridization and ligation reaction was heated to 95oC for 5 min, followed by six cycles of 60oC for 20 min / 95oC for 2 min, to allow for hybridization of target nucleic acids to their WSGR Docket No. 64100-744.601 complementary sequences in the recognition elements and ligation and circularization of the hybridized recognition elements. The reactions were held at 4oC. If a pause in the protocol is desired, the reactions can be stored for 3 days at -20oC.

[0238] After the hybridization and ligation reaction was completed, the reactions were treated with one or more exonucleases to remove linear nucleic acids from the reaction. While the reactions were on ice at 4oC, 30 µL of an exonuclease master mix was added, the master mix included, for example, a cocktail of Exo I and Exo III, a buffer, and water. The reactions were heated to 37oC for 30 min, followed by enzyme inactivation at 95oC for 5 min. The reactions were placed back on ice and held at 4oC.

[0239] After exonuclease treatment, the circularized recognition elements were amplified by rolling circle amplification to generate concatemeric amplification products. The exonuclease treated samples were transferred to new wells of a new plate that was pre-treated for nucleic acid immobilization. After transfer, the plate was sealed and heated to 42oC for 1 hr to immobilize the circularized recognition elements to the bottom of the well. After incubation, liquid from the reactions was removed, to the fullest extent possible without touching the bottom of the well.50 uL of an amplification master mix including a DNA polymerase for rolling circle amplification (e.g., EquiPhi polymerase), a buffer, dNTPs, an amplification primer oligonucleotide, a reducing agent, and water was added to the immobilized circularized recognition elements. The plate was again sealed and incubated at 42oC for 2 hrs.

[0240] Following amplification, the plate wells were washed and can be stored at room temperature in the dark for up to 3 days or stored at 4oC for longer storage.

[0241] Control reactions can also be run concurrent with test nucleic acid samples. The same protocol was followed, however known nucleic acid samples with known targets and their associated recognition elements were utilized as controls for troubleshooting an assay may problems arise, for example for evaluating whether reagents are performing properly and for use as baseline data across all runs and plates. The controls, including at least one negative control where water was used in lieu of either a nucleic sample and / or recognition elements, can be run in wells on the same plate as the test nucleic acid samples so that every plate that includes test nucleic acid samples for analysis also included positive and negative controls. Additionally, controls can be included within a test sample well as a quality control for a specific assay operation. For example, linear recognition elements with orthogonal targeting 5’ and 3’ ends and known codes can be added at ligation and / or circular recognition elements with known codes added at the exonuclease or RCA operations to detect for ligation, digestion and / or amplification failures. WSGR Docket No. 64100-744.601

[0242] After generation of concatemeric amplification products, detection and decoding of the codes present in the concatemeric amplification products can be performed to identify the presence or absence of a target variant that is associated with the code of the targeted recognition element that is used as a proxy for the presence or absence of the target variant in the original nucleic acid sample from the biological sample.

[0243] The assay plate was placed in an instrument for cycling of detection polynucleotide addition, detection polynucleotide hybridization to its targeted code in a recognition element, detection of the hybridization via fluorescence imaging and decoding of the concatenated images to generate results, including the determination of the code and using that code to report the probability of the presence of the original target sequence of interest from the sample. Example 2- Detection of Single Nucleotide Polymorphisms (SNPs), insertions and deletions in genomic DNA

[0244] Assays were performed as described in Example 1 unless otherwise stated. Multiple target genomic DNA (gDNA) samples (100 ng, Coriell) were aliquoted into separate wells of a 96 well, wherein each well included a ligation master mix (AmpLigase 0.33U), 1x AmpLigase buffer, 0.5 nM of each target recognition element (following the design scenario of Fig.8, depending on whether the target variant was a SNP, an insertion or a deletion), 800 mM of a hybridization / ligation enhancer and water. Table 3 lists the SNPs and insertion-deletion (indels) that were targeted in each of the Coriell samples evaluated, and data is shown in Figs.9-10. Table 3-Targeted SNPs, insertions and deletions queried in gDNA WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601

[0245] The samples were incubated at 95oC for 5 min and held at 60oC for 20 min followed by 95oC for 2 min. The 60oC hold for 20 min and 95oC for 2 min was repeated for a total of six times. The samples were held on ice or at 4oC in preparation of exonuclease digestion of any remaining linear DNA in the test wells.

[0246] An exonuclease digestion mix was added to the sample wells, the mix included 0.53U of Exonuclease I, 2.67U of Exonuclease II in a reaction buffer with water. The samples were incubated at 37oC for 30 min. followed by enzyme inactivation at 95oC for 5 min., followed by sample storage at 4oC.

[0247] After exonuclease digestion, an aliquot of the sample wells was transferred to a polylysine coated, optically clear bottom 96 well plate. The plate with the samples was sealed and incubated at either room temperature or 42oC for 1 hr to facilitate immobilization of the ligated recognition elements to the surface of the plate wells. After incubation, the solution in the wells was removed. Amplification master mix was added to each sample well, the master mix included water, EquiPhi buffer, 1 mM each dNTPs, 1mM DTT, 10 nM forward amplification primer, and 0.15U EquiPhi DNA polymerase. The plate was sealed, and the amplification reaction was incubated for 2 hr at 42oC, washed with TE-EDTA buffer several times and left in TE-EDTA buffer for decoding.

[0248] Samples were washed with 100 µl of 0.1N NaOH and three times with TE-EDTA.50 µl of a hybridization buffer, blocker DNA (to block polylysine on the plate where a concatemeric amplification product was not bound), 0.5 mM each of a detection oligonucleotide and an anchor oligonucleotide and water was added to the washed sample wells. Hybridization of the anchor oligonucleotide to the detection oligonucleotide and concurrent hybridization of the WSGR Docket No. 64100-744.601 detection polynucleotide to its target code sequence was performed at room temperature for 15 min. The sample wells were washed several times and stored in TE.

[0249] Images of the hybridized fluorescently labeled detection polynucleotides to their code sequences were taken using an Agilent BioTek Lionheart fluorescent microscope. Several rounds of application of fluorescently labeled detection polynucleotides and imaging were performed and the imaging data was collated and applied to the soft decision decoding pipeline for code determinations and subsequent target sequence alignment.

[0250] Figs.9 and 10 show examples of data for call rates and accuracy of the calls, respectively. Following the design scenarios from Fig.8 for the targeting recognition elements, the data shows high call rates for SNPs, insertions and deletions across all the gDNA samples tested. Additionally, high accuracy of the calls and concordance across all replicate samples was seen, running all samples across on four different assays plate and using two imaging instruments. As such, by designing the recognition elements to target SNPs, insertions and deletions as shown in Fig.8, the assays were highly reproducible in both calls and accuracy in identifying the target variant of interest across a diversity of gDNA samples and a diversity of variants. The same results were seen when saliva samples were used as the biological sample. As such, the design strategies for identifying variants in the methods herein may not be sample type, or variant type, dependent. Example 3- Detection of SNPs, insertions, and deletions in blood and saliva DNA

[0251] The assay workflow was performed as in Example 1 unless otherwise noted. For this experiment, matched blood and saliva samples from 10 different subjects and control gDNA Coriell samples were assayed. Recognition elements were designed to each target variant in Table 4 as shown in Fig.8. Each sample was queried for all of the variants, and the data is shown in Fig.11. Table 4-SNPs, insertions and deletions queried in blood, saliva and control gDNA WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601 WSGR Docket No. 64100-744.601

[0252] Fig.11 shows call rates for the three sample types. As in Example 4, high call rates were seen for most samples. While the quality of the extracted DNA was different for the different sample types, the data showed that all samples assayed yielded similar results regardless of origin or quality. Concordance and accuracy were also high, regardless of origin or quality. As such, the recognition element design strategies as seen in Fig.8 can be used with multiple sample types to accurately identify the presence of multiple different variants in a sample type and are not sample type or variant dependent. Example 4- Copy number variations and hybrid allele detection

[0253] This assay workflow follows that of Example 1, unless otherwise stated. This experiment demonstrates a method for determining the copy number of a gene of interest using the WT construct designs as shown in Fig.8.

[0254] The gene CYP2D6 was used as the target gene of interest. Recognition elements were designed, following the strategies from Fig.8, targeting different exons and introns within the WSGR Docket No. 64100-744.601 CYP2D6 as shown in Fig.12. DNA from Coriell samples that included normal gene complements (NA17658, NA18544 and NA17679), whole gene deletion (NA12873, NA18508 and NA19035) and whole gene amplification (NA18909 and NA17454) were utilized as control samples to demonstrate that the recognition elements as designed may be used to identify whole gene shifts if present in a sample. Figs.13 A-C demonstrate the outcome of the control assays. Fig.13A shows that the normal chromosomal complement was detected and decoded; Fig.13B shows that whole gene deletions were detected and decoded; and Fig.13C shows that whole gene amplifications were detected and decoded, thereby demonstrating that the design of the recognition elements was successful in querying all of the locations across the gene (Fig.12).

[0255] The next set of experiments was undertaken to determine whether the same design scenario may also be applied to hybrid alleles of CYP2D6:2D7. The sample locations were queried as shown in Fig.12. DNA from Coriell samples were assayed for known hybrid alleles having exon 9 conversions *10 +*36 (NA18553, NA18624, NA18612, Fig.14A), exon 9 conversions *10+*36 / *36+*36 (NA18757, Fig.14B), an intron 1 conversion *68 (NA12878, Fig.14C), and an intron 1 conversion *68 and exon 9 conversion *13 (NA19982, Fig.14D). Figs.14A-D demonstrate that all the hybrid alleles were correctly detected, decoded and CNVs assigned. The hybrid CYPD26 *68 sample was utilized in a comparison between the method described herein and an Illumina® Infinium Microarray assay. As shown in Figs.15A-B, the microarray assay was not able to detect the hybrid intron 1 amplification that is present in CYP2D6 *68 (Fig.15B) whereas previously demonstrated the present method is able to identify the intron 1 amplification (Fig.15A). As such, the present methods are able to accurately detect and decode copy number variances for whole gene deletions and amplifications and hybrid star alleles. Example 5- Identifying variants in gene regions with close homologs

[0256] The workflow was followed as in Example 1, except where otherwise noted. Homologs and other high homology genes such as pseudogenes can be challenging when the target molecules of interest share close homology to other loci, for example the cytochrome P450 genes. Fig.16A shows a strategy for identifying a target of interest in regions with close homologs. In this strategy, the 5’ end of the recognition element comprises one or more nucleotides that serve to differentiate the gene variant of interest from the homolog. The 3’ end of the recognition elements comprises the variant nucleotide of interest, in this example a SNP that is present in both the gene of interest and the homolog. Additionally, a third oligonucleotide, or bridge oligonucleotide is included such that two ligation events are performed, based on two hybridization events specific to the gene of interest and not the homolog, thereby providing additional specificity for the gene of interest and not the homolog. WSGR Docket No. 64100-744.601 Fig.16B demonstrates the efficacy of the strategy of Fig.16A, using CYP2D6 as the gene of interest. When the recognition element does not include a differentiating base and a bridge oligonucleotide, the target is not discernable from the homolog (data not shown). However, the incorporation of a differentiating base and the bridge oligonucleotide greatly decreases the off target counts (from the homolog), thereby allowing for the variant of interest in the gene of interest to be detected.

[0257] Figs.17—B supports the strategy of Fig.16A. In Fig.17A, for example, a strategy using a recognition element comprising just the variant sequence at the 3’ end of the recognition element (Fig.8) demonstrates the limited ability to differentiate the CYP2D6 from a close homolog as the high alt mutant counts in the HomRef genotype are due to the close sequence with the homolog. However, Fig.17B shows, for example, that using a recognition element with a differentiating base, the variant sequence, and the bridge oligonucleotide of strategy Fig.16A is able to differentiate between the gene variant in the gene of interest over the homolog, thereby demonstrating the power of the recognition design strategy for differentiating a gene variant in a gene of interest from a close homolog. Example 6- Implementation of SNP detection strategy and learning profiles

[0258] The workflow was followed as in Example 1, except where otherwise noted. This experiment demonstrates the ability of the disclosed strategies for SNP detection to correctly identify the presence of a SNP in a target gene.

[0259] As more data is obtained through assaying multiple samples and multiple variants, learning is enabled to enhance the probabilities of specific error profiles of cluster positions. Error profiles are learned for each recognition element, each different variant, using hundreds of samples. The learned error profiles can then be applied to assay data to correct for errors and miscalls during decoding. Figs.18A-18N show examples of plots of multiple target queries using the SNP detection strategy of Fig.8. The targets shown can be queried using any number of different samples to generate a body of data used for machine learning the error profiles of a particular target of interest, which can then be applied during decoding for error correction. Example 7- HLA imputation compared to direct genotyping for HLA allele identification

[0260] Due to high polymorphism in HLA loci, traditional methods for direct probing of SNPs may be problematic. Imputation methods and disclosed direct detection methods were compared for HLA alleles HLA-A*31-01, HLA-B*15:02, HLA-B*57:01 and HLA-B*58:01.

[0261] Assays were performed for identifying HLA target variants as described in Example 1.

[0262] For HLA imputation, the genetic structure of the MHC of a sample is characterized by high levels of linkage disequilibrium compared to the rest of the genome and long-range haplotypes of regional variants and their corresponding HLA allelic types. Linkage WSGR Docket No. 64100-744.601 disequilibrium measures the degree to which alleles at two loci are associated. HLA imputation uses a reference panel, such as a genomic reference database like 1000 Genomes thereof and / or internally derived data from multiple assay results that includes both HLA and SNP genotypes to infer HLA genotypes from SNP data.

[0263] For HLA imputation, the focus is identifying a set of SNPs in HLA alleles that are in strong linkage disequilibrium with the HLA alleles of interest, as such SNPs that are not directly the target variants in the alleles of interest. Conversely, the method of direct genotyping may focus on finding SNPs directly in the alleles of interest that are unique to that allele. Direct genotyping searches for the minimum number of SNPs that differentiate the allele of interest from the other alleles. Table 3 lists the SNP targets that were assayed for four different HLA star alleles.

[0264] For imputation, 10 or 12 sites in the HLA genes were genotypes and used for analysis. Imputation uses locations away from the HLA variant sites for analysis, whereas for direct genotyping the variant site itself is directly assayed and measured. Nucleic acid samples were obtained from Coriell and included mixed ethnic populations for each HLA star allele from Africa, Americas, Europe, Eastern Asia and South Asia. Four HLA star alleles that correlate to the Coriell SNP samples, the number of loci for each star allele and the expected precision and accuracy of the assay is found in Table 5. Table 5-HLA star alleles targeted

[0265] Fig.21 shows an example of a chart that details the analysis workflow for identifying HLA SNPs using HLA imputation in combination with direct genotyping of HLA alleles. For HLA imputation, SNPs from the 1000 Genomes project and SNPs from 1000 HLA types were evaluated for linkage disequilibrium between the two sets of SNPs. Those SNPs demonstrating the highest degree of linkage disequilibrium comprise a subset of SNPs on which is run an WSGR Docket No. 64100-744.601 imputation algorithm. The evaluating, subset definition and imputation can be performed multiple times to increase the accuracy of calling HLA variant alleles. For HLA direct genotyping, the HLA allele SNPs themselves are evaluated and aligned for determining which subset of SNPs have the maximum difference from one target allele to another. The imputation is done, in this example, using Beagle, a third-party open source software (2012, Ayres et al., Sys Biol 61(1):170-173). The imputation reference panel is imported from a public genome database, such as 1000 Genomes. The HLA direct genotyping and HLA imputation analysis data are combined and analyzed, and the probability of the presence of an HLA variant target is determined and reported.

[0266] Figs.22A-B demonstrate that, independently, direct genotyping probability is comparable to imputation probability, demonstrating the disclosed methods can be used to directly differentiate highly polymorphic HLA alleles. Results correlate with HLA typing results found in the literature. Generally, HLA-B*58:01 was a lower performer perhaps due to low signal:noise ratio. It is contemplated that this may be due to the lower LD in African population and more similarity with other alleles for this particular HLA allele. In this example, 99% calling accuracy of the HLA targeted star allele variant was realized, thereby matching the expected accuracy. It is contemplated that performing both imputation and direct genotyping for calling HLA star alleles may provide an even greater degree of accuracy in variant calling. Example 8- Correcting for underlying variants

[0267] This assay workflow follows that of Example 1, unless otherwise stated. The experiment demonstrates a method for correcting miscalls for the presence of a target variant of interest from a sample as a result of an underlying variant in the TNF gene, which was uncorrected for initially.

[0268] A SNP in the tumor necrosis factor, or TNF gene, at chr:2-9974074 was targeted; the reference wildtype sequence being chr:2-9974074G_G with the variant sequence being chr:2- 9974074G_T. The DNA sample assayed that includes this variant is NA18952 (Coriell). Two recognition elements were designed, one that included a 3’ end nucleotide targeting the wildtype sequence and the second that included a 3’ end targeting the variant sequence. However, upon data analysis miscalling of the SNP of interest was seen, as such suggesting there might be an underlying SNP in close proximity to the SNP of interest at this location in the TNF sequence. Indeed, an insertion of a single nucleotide six nucleotides away from the SNP of interest was found. As such, two new recognition elements were designed to include the underlying variant. Table 6 shows portions of the original and redesigned 3’ end sequences of the recognition elements, bolded letters are the wildtype (Ref) and variant nucleotides. WSGR Docket No. 64100-744.601 Table 6 Original and redesigned 3’ end sequences for correcting for an underlying variant

[0269] The redesigned recognition elements, one with the 3’ complementary variant and the other for the wildtype sequence, further included a one nucleotide insertion (bolded “A”) at the - 6 position from the 3’ end. This additional nucleotide was not present in the originally designed recognition elements and thus contributed to the miscalling of the variant. This design strategy is shown in Fig.24 (right panel), with Loc 1 being the variant of interest and Loc 2 being the phased variant. The design strategy can also be applied where a phased variant is known to exist, but perhaps the alternative nucleotide of the phased variant is not well understood. In this scenario, four different recognition elements can be concurrently used to query the sample, where the location of the phased variant includes any one of the four canonical nucleotides (e.g., A, C, G, T) in addition to the complement to the nucleotide of interest. It is contemplated this miscalling was due to hybridization and / or ligation inefficiency, thereby skewing the results and detecting either variant or wildtype, as seen in the left side graph of Fig.23. The right-side graph reflects results using the redesigned recognition elements. When utilizing the original recognition elements, 210 concatemeric amplification products were identified as wildtype and 2164 concatemeric amplification products were identified as having the variant of interest in the initial assay. However, when adding the redesigned recognition elements along with the original recognition elements, the skewed results were corrected such that the wildtype sequence was identified in 1520 concatemeric amplification products and the variants was identified in 2435 concatemeric amplification products. As such, the redesigned recognition elements that corrected for the underlying variant corrected the miscalls of both the wildtype and the variant target of interest.

[0270] While Fig.24 shows a design strategy for phased variant detection without a bridge element, Fig.25 shows the design strategy including a bridge element. In Fig.25, the target of interest is close to a phased variant (130 variant of interest and 132 phased variant, respectively) found in CYP2D6 exon 3. The 3’ targeted arm of the recognition element is designed such that four recognition elements 3’ targeted arms (recognition element #1-#4), cover the different potential genotypes and the bridge element is complementary to the region of the gene between the 3’ and 5’ targeted arms of the recognition element, with the 3’ nucleotide of the bridge element including a differentiating nucleotide that targets the CYP2D6 sequence over a homolog. As such, all potential 130_132 genotypes of CYP2D6 that may be present in a sample WSGR Docket No. 64100-744.601 can be identified using all of the differently targeted recognition elements in one assay, including the ability to differentiate between the gene of interest and a homolog.

[0271] Data shown in Figs.26A-F further shows the design strategy for identifying a target of interest when phased variants are present. Data was generated for targets from CYP2D6, at position 130 (target of interest) and 132 (phased variant). When designing recognition elements with a nucleotide at the 3’ end for querying the target of interest and degeneracy at the phased variant nucleotide, the star alleles were identified in CYP2D6; A) genotype *1 / *6 both loci ref, B) genotype *1 / *21 and *1 / *46 (left and right graph, respectively) both samples with het at position 130 and ref at position 132, C) genotype *68+*4 / *5, *2 / *4, and *2 / *41 (left, middle and right, respectively) all samples with alt at position 130 and ref at position 132, D) genotype *29 / *5 both loci alt, E) genotype *1 / *29 both loci het, and F) genotype *17 / *29 (2 or 1 copies, left or right, respectively) position 130 alt and position 132 het. While the experiments included recognition elements that were redesigned to include one underlying variant, this strategy may also apply to recognition elements redesigned to target, for example, hot spots of variability within a sequence. For example, a target sequence that comprises more than two variant sequences that are in close proximity. These two or more variant sequences may be known or unknown, such that recognition elements may be designed with known nucleotide possibilities at different locations along the target sequence, or unknown such that different location may include degenerate sequences at known locations. Any combination in between is also contemplated, such as a recognition element 3’ end that includes two or more known variant and associated wildtype nucleotides at a given location in combination with locations that are degenerate in nature, as such more than one fixed variant sequence in combination with degenerate nucleotides at any given location.

[0272] While certain examples of methods and systems have been shown and disclosed herein, one of skill in the art will realize that these are provided by way of example only and not intended to be limiting within the specification. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the scope disclosed herein. Furthermore, it shall be understood that all aspects of the disclosed methods and systems are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables and the description is intended to include such alternatives, modifications, variations or equivalents.

Claims

WSGR Docket No. 64100-744.601 CLAIMS WHAT IS CLAIMED IS:

1. A method for identifying one or more variant targets of interest from a biological sample, comprising: a) providing a biological sample suspected of having the one or more variant targets of interest; b) hybridizing the biological sample to a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements comprises a 5’ end and a 3’ end, wherein the 3’ end comprises a sequence that is complementary to a variant target of interest of the one or more variant targets of interest, and wherein each recognition element of the plurality of recognition elements further comprises a code that is a proxy for the variant target of interest, c) ligating a recognition element that is hybridized to the biological sample to generate a circularized and ligated recognition element; d) generating a concatemeric amplification product from the circularized and ligated recognition element; and e) decoding the code from the concatemeric amplification product using soft decision decoding, thereby identifying the presence or absence of the one or more variant targets of interest from the biological sample.

2. The method of claim 1, further comprising hybridizing the biological sample to a plurality of second recognition elements, wherein each second recognition element of the plurality of second recognition elements comprises a 3’ end that is complementary to a wildtype nucleic acid of the one or more variant targets of interest, and wherein the second recognition element that comprises the 3’ end that is complementary to the wildtype nucleic acid of the one or more variant targets of interest comprises a different code as the recognition element that comprises the 3’ end that comprises the complementary sequence to the one or more variant targets of interest.

3. The method of claim 1, wherein the biological sample is from a blood sample, a buccal swab sample, a saliva sample, or a tissue sample.

4. The method of claim 3, wherein the biological sample is an extracted or purified nucleic acid sample from the blood sample, the buccal swab, the saliva sample, or the tissue sample.

5. The method of claim 1, wherein the one or more variant targets of interest comprises a single nucleotide polymorphism (SNP), an insertion, a deletion, or a copy number variant (CNV).

6. The method of claim 5, wherein the variant targets of interest comprises the SNP.WSGR Docket No. 64100-744.601 7. The method of claim 5, wherein the variant targets of interest comprises the insertion.

8. The method of claim 5, wherein the variant targets of interest comprises the deletion.

9. The method of claim 5, wherein the variant targets of interest comprises the CNV.

10. The method of claim 1, wherein the one or more variant targets of interest comprises from 10 to 1,000 variant targets of interest.

11. The method of claim 1, wherein the one or more variant targets of interest comprises at least 10 variant targets of interest.

12. The method of claim 1, wherein the one or more variant targets of interest comprises at least 100 variant targets of interest.

13. The method of claim 1, wherein the one or more variant targets of interest comprises at least 1,000 variant targets of interest.

14. The method of claim 1, wherein the one or more variant targets of interest comprises one or more pharmacogenomics variant targets of interest.

15. The method of claim 14, wherein the one or more pharmacogenomics variant targets of interest comprises one or more star alleles.

16. The method of claim 1, wherein the one or more variant targets of interest comprise one or more variant targets of interest from a CYP2D6 gene in the biological sample.

17. The method of claim 1, wherein the one or more variant targets of interest comprises one or more variant targets of interest from one or more HLA genes in the biological sample.

18. The method of claim 1, wherein the one or more variant targets of interest is from one or more genes selected from the group consisting of: ABCB1, ABCG2, ADRA2A, ALDH2, ANK3, ANKK1, APOE, BDNF, C11orf65, CACNA1C, CACNA1S, CEP72, CFTR, COMT, CYP1A2, CYP2C, CYP2C19, CYP2C9, CYP2D6, DBH, DPYD, DRD2, F2, F5, G6PD, GLP1R, GRIK1, GRIK4, GRIN2B, HTR2A, HTR2C, IFNL3, IFNL4, MC4R, MT- RNR1, MTHFR, OPRM1, PNPLA5, RYR1, SCL47A2, SLC6A4, SLCO1B1, SULT4A1, UGT1A1, UGT2B15, VKORC, YEATS4, CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, CYP4F2, DPYD, NUDT15, SLCO1B1, TPMT, and UGT1A1.

19. The method of claim 1, wherein the one or more variant targets of interest is from one or more genes selected from the group consisting of: ABCB1, ABCG2, ADRA2A, ALDH2, ANK3, ANKK1, APOE, BDNF, C11orf65, CACNA1C, CACNA1S, CEP72, CFTR, COMT, CYP1A2, CYP2C, CYP2C19, CYP2C9, CYP2D6, DBH, DPYD, DRD2, F2, F5, G6PD, GLP1R, GRIK1, GRIK4, GRIN2B, HTR2A, HTR2C, IFNL3, IFNL4, MC4R, MT- RNR1, MTHFR, OPRM1, PNPLA5, RYR1, SCL47A2, SLC6A4, SLCO1B1, SULT4A1, UGT1A1, UGT2B15, VKORC, and YEATS4.WSGR Docket No. 64100-744.601 20. The method of claim 1, wherein the one or more variant targets of interest comprises star allele variants from a cytochrome P450 gene.

21. The method of claim 20, wherein the cytochrome P450 gene comprises one or more of CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, or CYP4F2.

22. The method of claim 1, wherein the one or more variant targets of interest comprises star allele variants from one or more genes selected from the group consisting of CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, CYP4F2, DPYD, NUDT15, SLCO1B1, TPMT, and UGT1A1.

23. The method of claim 1, wherein the one or more variant targets of interest are selected from the one or more variants listed in Table 1 or Table 2.

24. The method of claim 1, wherein ligating the recognition element comprises use of a thermostable ligase.

25. The method of claim 1, wherein generating the concatemeric amplification product comprises rolling circle amplification (RCA) or multiple strand displacement.

26. The method of claim 1, further comprising treating the circularized and ligated recognition element with an exonuclease prior to generating the concatemeric amplification product.

27. The method of claim 1, wherein the circularized and ligated recognition element is placed on a substrate prior to generating the concatemeric amplification product.

28. The method of claim 27, wherein the substrate comprises one or more wells or one or more tubes that have been pre-treated with a nucleic acid immobilization composition.

29. The method of claim 27, wherein the substrate is one or more wells of a 24, 48, 96 or 384 well plate and each of the one or more wells is pre-treated with a polymer.

30. The method of claim 29, wherein the polymer is selected from the group consisting of: polyacrylamide, branched PEI, linear PEI, poly(β-aminoester) and poly(amidoamine), PEG, a gel, poly-L-lysine, silane, agarose, and muscle mimetic catecholamine polymer.

31. The method of claim 1, wherein decoding the code from the concatemeric amplification product comprises next generation sequencing or detection by hybridization.

32. The method of claim 1, wherein the code comprises at least two nucleic acid segments.

33. The method of claim 32, wherein the at least two nucleic acid segments are detected by hybridizing one or more detection polynucleotides to the at least two nucleic acid segments and imaging the hybridizing.

34. The method of claim 33, wherein each of the one or more detection polynucleotides are labeled with a different fluorescent moiety.WSGR Docket No. 64100-744.601 35. A method for identifying pharmacogenomic targets of interest from a biological sample, comprising: a) providing a biological sample suspected of having one or more pharmacogenomic variant targets of interest; b) hybridizing the biological sample to a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements comprises a 5’ end and a 3’ end, wherein the 3’ end comprises a sequence that is complementary to a variant target of interest of the one or more pharmacogenomic variant targets of interest, and wherein each recognition element of the plurality of recognition elements further comprises a code that is a proxy for the one or more pharmacogenomic variant targets of interest; c) ligating a recognition element that is hybridized to the biological sample to generate a circularized and ligated recognition element; d) generating a concatemeric amplification product from the circularized and ligated recognition element; and e) decoding the code from the concatemeric amplification product using soft decision decoding, thereby identifying the presence or absence of the one or more pharmacogenomic variant targets of interest from the biological sample.

36. The method of claim 35, further comprising hybridizing the biological sample to a plurality of second recognition elements, wherein each recognition element of the plurality of second recognition elements comprises a 3’ end that is complementary to a wildtype nucleic acid of the one or more pharmacogenomic variant targets of interest, and wherein the recognition element that comprises the 3’ end that is complementary to the wildtype nucleic acid of the one or more pharmacogenomic variant targets of interest comprises the same code as the recognition element that comprises the 3’ end that comprises the complementary sequence to the one or more pharmacogenomic variant targets of interest.

37. The method of claim 35, wherein the biological sample is from a blood sample, a buccal swab sample, a saliva sample, or a tissue sample.

38. The method of claim 37, wherein the biological sample is an extracted or purified nucleic acid sample from the blood sample, the buccal swab sample, the saliva sample, or the tissue sample.

39. The method of claim 35, wherein the one or more pharmacogenomic variant targets of interest comprises a single nucleotide polymorphism (SNP), an insertion, a deletion, or a copy number variant (CNV).

40. The method of claim 39, wherein the one or more pharmacogenomic variant targets of interest comprises the SNP.WSGR Docket No. 64100-744.601 41. The method of claim 39, wherein the one or more pharmacogenomic variant targets of interest comprises the insertion.

42. The method of claim 39, wherein the one or more pharmacogenomic variant targets of interest comprises the deletion.

43. The method of claim 39, wherein the one or more pharmacogenomic variant targets of interest comprises the CNV.

44. The method of claim 35, wherein the one or more pharmacogenomic variant targets of interest comprise from 10 to 1,000 pharmacogenomic variant targets of interest.

45. The method of claim 35, wherein the one or more pharmacogenomic variant targets of interest comprise at least 10 pharmacogenomic variant targets of interest.

46. The method of claim 35, wherein the one or more pharmacogenomic variant targets of interest comprises at least 100 pharmacogenomic variant targets of interest.

47. The method of claim 35, wherein the one or more pharmacogenomic variant targets of interest comprises at least 1,000 pharmacogenomic variant targets of interest.

48. The method of claim 35, wherein the one or more pharmacogenomics variant targets of interest comprises one or more star alleles.

49. The method of claim 35, wherein the one or more pharmacogenomic variant targets of interest comprises one or more pharmacogenomic variant targets of interest from a CYP2D6 gene.

50. The method of claim 35, wherein the one or more pharmacogenomic variant targets of interest comprises one or more variant targets of interest from one or more HLA genes.

51. The method of claim 35, wherein the one or more pharmacogenomic variant targets of interest are from one or more genes selected from the group consisting of: ABCB1, ABCG2, ADRA2A, ALDH2, ANK3, ANKK1, APOE, BDNF, C11orf65, CACNA1C, CACNA1S, CEP72, CFTR, COMT, CYP1A2, CYP2C, CYP2C19, CYP2C9, CYP2D6, DBH, DPYD, DRD2, F2, F5, G6PD, GLP1R, GRIK1, GRIK4, GRIN2B, HTR2A, HTR2C, IFNL3, IFNL4, MC4R, MT-RNR1, MTHFR, OPRM1, PNPLA5, RYR1, SCL47A2, SLC6A4, SLCO1B1, SULT4A1, UGT1A1, UGT2B15, VKORC, YEATS4, CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, CYP4F2, DPYD, NUDT15, SLCO1B1, TPMT, and UGT1A1.

52. The method of claim 35, wherein the one or more pharmacogenomic variant targets of interest are from one or more genes selected from the group consisting of: ABCB1, ABCG2, ADRA2A, ALDH2, ANK3, ANKK1, APOE, BDNF, C11orf65, CACNA1C, CACNA1S, CEP72, CFTR, COMT, CYP1A2, CYP2C, CYP2C19, CYP2C9, CYP2D6, DBH, DPYD, DRD2, F2, F5, G6PD, GLP1R, GRIK1, GRIK4, GRIN2B, HTR2A, HTR2C, IFNL3,WSGR Docket No. 64100-744.601 IFNL4, MC4R, MT-RNR1, MTHFR, OPRM1, PNPLA5, RYR1, SCL47A2, SLC6A4, SLCO1B1, SULT4A1, UGT1A1, UGT2B15, VKORC, and YEATS4.

53. The method of claim 35, wherein the one or more pharmacogenomic variant targets of interest comprise star allele variants from a cytochrome P450 gene.

54. The method of claim 53, wherein the cytochrome P450 gene comprises one or more of CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, or CYP4F2.

55. The method of claim 35, wherein the one or more pharmacogenomic variant targets of interest comprise star allele variants from one or more genes selected from the group consisting of: CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, CYP4F2, DPYD, NUDT15, SLCO1B1, TPMT, and UGT1A1.

56. The method of claim 35, wherein the one or more pharmacogenomic variant targets of interest are selected from the one or more variants listed in Table 1 or Table 2.

57. The method of claim 35, wherein ligating the recognition element comprises use of a thermostable ligase.

58. The method of claim 35, wherein generating the concatemeric amplification product comprises rolling circle amplification or multiple strand displacement.

59. The method of claim 35, further comprising treating the circularized and ligated recognition element with an exonuclease prior to generating the concatemeric amplification product.

60. The method of claim 35, wherein the circularized and ligated recognition element is placed on a substrate prior to generating the concatemeric amplification products.

61. The method of claim 60, wherein the substrate is one or more wells or one or more tubes that have been pre-treated with a nucleic acid immobilization composition.

62. The method of claim 60, wherein the substrate is one or more wells of a plate and each of the one or more wells is pre-treated with a polymer.

63. The method of claim 62, wherein the polymer is selected from the group consisting of: polyacrylamide, branched PEI, linear PEI, poly(β-aminoester) and poly(amidoamine), PEG, a gel, poly-L-lysine, silane, agarose, and muscle mimetic catecholamine polymer.

64. The method of claim 35, wherein decoding the code from the concatemeric amplification product comprises next generation sequencing or detection by hybridization.

65. The method of claim 35, wherein the code comprises at least two nucleic acid segments.

66. The method of claim 65, wherein the at least two nucleic acid segments are detected by hybridizing one or more detection polynucleotides to the at least two nucleic acid segments and imaging the hybridizing.WSGR Docket No. 64100-744.601 67. The method of claim 66, wherein each of the one or more detection polynucleotides are labeled with a different fluorescent moiety.

68. A method for identifying the presence or absence of one or more star alleles in a genome, comprising: a) providing a biological sample suspected of having the one or more star alleles; b) hybridizing the biological sample to a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements comprises a 5’ end and a 3’ end, wherein the 3’ end comprises a sequence that is complementary to a star allele of the one or more star alleles, and wherein each recognition element of the plurality of recognition elements further comprises a code that is a proxy for the one or more star alleles; c) ligating a recognition element that is hybridized to the biological sample to generate a circularized and ligated recognition element; d) generating a concatemeric amplification product from the circularized and ligated recognition element; and e) decoding the code from the concatemeric amplification product using soft decision decoding, thereby identifying the presence or absence of the one or more star alleles from the biological sample.

69. The method of claim 68, further comprising hybridizing the biological sample to a plurality of second recognition elements, wherein each recognition element of the plurality of second recognition elements comprises a 3’ end that is complementary to a wildtype nucleic acid of the one or more star alleles, and wherein the recognition element that comprises a 3’ end that is complementary to the wildtype nucleic acid of the one or more star alleles comprises the same code as the recognition element that comprises the 3’ end that comprises the complementary sequence to the one or more star alleles.

70. The method of claim 68, wherein the biological sample is from a blood sample, a buccal swab sample, a saliva sample, or a tissue sample.

71. The method of claim 70, wherein the biological sample is an extracted or purified nucleic acid sample from the blood sample, the buccal swab sample, the saliva sample, or the tissue sample.

72. The method of claim 68, wherein the one or more star alleles comprise a single nucleotide polymorphism (SNP), an insertion, a deletion, or a copy number variant (CNV).

73. The method of claim 72, wherein the one or more star alleles comprises the SNP.

74. The method of claim 72, wherein the one or more star alleles comprises the insertion.

75. The method of claim 72, wherein the one or more star alleles comprises the deletion.

76. The method of claim 72, wherein the one or more star alleles comprises the CNV.WSGR Docket No. 64100-744.601 77. The method of claim 68, wherein the one or more star alleles comprises from 10 to 1,000 star alleles.

78. The method of claim 68, wherein the one or star alleles comprises at least 10 star alleles.

79. The method of claim 68, wherein the one or more star alleles comprises at least 100 star alleles.

80. The method of claim 68, wherein the one or more star alleles comprises at least 1,000 star alleles.

81. The method of claim 68, wherein the one or more star alleles are present in a CYP2D6 gene from the biological sample.

82. The method of claim 68, wherein the one or more star alleles are present in one or more HLA genes from the biological sample.

83. The method of claim 68, wherein the one or more star alleles are from a cytochrome P450 gene from the biological sample.

84. The method of claim 83, wherein the cytochrome P450 gene comprises one or more of CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, or CYP4F2.

85. The method of claim 68, wherein the one or more star alleles are from one or more genes comprising CYP1A2, CYP2B6, CYP2C19, CYP2C8, CYP2C9, CYP2D6, CYP3A4, CYP3A5, CYP4F2, DPYD, NUDT15, SLCO1B1, TPMT, or UGT1A1.

86. The method of claim 68, wherein the one or more star alleles are selected from star alleles listed in Table 2.

87. The method of claim 68, wherein ligating the recognition element comprises use of a thermostable ligase.

88. The method of claim 68, wherein generating the concatemeric amplification product comprises rolling circle amplification or multiple strand displacement.

89. The method of claim 68, further comprising treating the circularized and ligated recognition element with an exonuclease prior to generating the concatemeric amplification product.

90. The method of claim 68, wherein the circularized and ligated recognition element is placed on a substrate prior to generating the concatemeric amplification product.

91. The method of claim 90, wherein the substrate is one or more wells or one or more tubes that have been pre-treated with a nucleic acid immobilization composition.

92. The method of claim 90, wherein the substrate is one or more wells of a plate and each of the one or more wells is pre-treated with a polymer.WSGR Docket No. 64100-744.601 93. The method of claim 92, wherein the polymer is selected from the group consisting of: polyacrylamide, branched PEI, linear PEI, poly(β-aminoester) and poly(amidoamine), PEG, a gel, poly-L-lysine, silane, agarose, and muscle mimetic catecholamine polymer.

94. The method of claim 68, wherein decoding the code from the concatemeric amplification product comprises next generation sequencing or detection by hybridization.

95. The method of claim 94, wherein the code comprises at least two nucleic acid segments.

96. The method of claim 95, wherein the at least two nucleic acid segments are detected by hybridizing one or more detection polynucleotides to the at least two nucleic acid segments and imaging the hybridizing.

97. The method of claim 96, wherein each of the one or more detection polynucleotides is labeled with a different fluorescent moiety.

98. A method for determining the presence or absence of one or more genetic variants that alter drug metabolism, comprising: a) providing a biological sample suspected of having the one or more genetic variants that alter drug metabolism; b) hybridizing the biological sample to a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements comprises a 5’ end and a 3’ end, wherein the 3’ end comprises a sequence that is complementary to one of the one or more genetic variants, and wherein each recognition element of the plurality of recognition elements further comprises a code that is a proxy for the one or more genetic variants; c) ligating a recognition element that is hybridized to the biological sample to generate a circularized and ligated recognition element; d) generating a concatemeric amplification product from the circularized and ligated recognition element; and e) decoding the code from the concatemeric amplification product using soft decision decoding, thereby identifying the presence or absence of the one or more genetic variants that alter drug metabolism from the biological sample.

99. The method of claim 98, further comprising hybridizing the biological sample to a plurality of second recognition elements, wherein each recognition element of the plurality of second recognition elements comprises a 3’ end that is complementary to a wildtype nucleic acid of the one or more genetic variants that alter drug metabolism, and wherein the recognition element that comprises a 3’ end that is complementary to the wildtype nucleic acid of the one or more genetic variants comprises the same code as the recognition element that comprises the 3’ end that comprises the complementary sequence to the one or more genetic variants that alter drug metabolism.WSGR Docket No. 64100-744.601 100. The method of claim 98, wherein the biological sample is from a blood sample, a buccal swab sample, a saliva sample, or a tissue sample.

101. The method of claim 100, wherein the biological sample is an extracted or purified nucleic acid sample from the blood sample, the buccal swab sample, the saliva sample, or the tissue sample.

102. The method of claim 98, wherein the one or more genetic variants comprises a single nucleotide polymorphism (SNP), an insertion, a deletion and a copy number variant (CNV).

103. The method of claim 98, wherein the one or more genetic variants comprise from 10 to 1,000 genetic variants.

104. The method of claim 98, wherein the one or more genetic variants are found in a cytochrome P450 gene from the biological sample.

105. The method of claim 104, wherein the cytochrome P450 gene is CYP2D6.

106. The method of claim 98, wherein the one or more genetic variants can alter the drug metabolism of one or more of: Metoprolol, Propranolol, Timolol, Encainide, Flecainide, Perhexilene, Propafenone, Sparteine, Amitriptyline, Clomipramine, Desipramine, Fluoxetine, Fluvoxamine, Imipramine, Mianserin, Nortriptyline, Paroxetine, Venlafexine, Haloperidol, Perphenazine, Risperidone, Thioridazine, Zuclopenthixol, Codeine, Debrisoquine, Dextromethorphan, Phenoformin, Tolterodine, and Tramadol.

107. The method of claim 98, wherein ligating the recognition element comprises use of a thermostable ligase.

108. The method of claim 98, wherein generating the concatemeric amplification product comprises rolling circle amplification or multiple strand displacement.

109. The method of claim 98, further comprising treating the circularized and ligated recognition element with an exonuclease prior to generating the concatemeric amplification product.

110. The method of claim 98, wherein the circularized and ligated recognition element is placed on a substrate prior to generating the concatemeric amplification product.

111. The method of claim 110, wherein the substrate is one or more wells or one or more tubes that have been pre-treated with a nucleic acid immobilization composition.

112. The method of claim 110, wherein the substrate is one or more wells of a plate and each of the one or more wells is pre-treated with a polymer.

113. The method of claim 112, wherein the polymer is selected from the group consisting of: polyacrylamide, branched PEI, linear PEI, poly(β-aminoester) and poly(amidoamine), PEG, a gel, poly-L-lysine, silane, agarose, and muscle mimetic catecholamine polymer.WSGR Docket No. 64100-744.601 114. The method of claim 98, wherein decoding the code from the concatemeric amplification product comprises next generation sequencing or detection by hybridization.

115. The method of claim 114, wherein the code comprises at least two nucleic acid segments.

116. The method of claim 115, wherein the at least two nucleic acid segments are detected by hybridizing one or more detection polynucleotides to the at least two nucleic acid segments and imaging the hybridizing.

117. The method of claim 116, wherein each of the one or more detection polynucleotides is labeled with a different fluorescent moiety.

118. The method of claim 98, wherein the method is used for determining the predisposition of altered drug metabolism in a subject.

119. A method for determining the presence or absence of one or more genetic variants in one or more HLA genes, comprising: a) providing a biological sample suspected of having the one or more genetic variants in the one or more HLA genes; b) hybridizing the biological sample to a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements comprises a 5’ end and a 3’ end, wherein the 3’ end comprises a sequence that is complementary to one of the one or more genetic variants, and wherein each recognition element of the plurality of recognition elements further comprises a code that is a proxy for the one or more genetic variants; c) ligating the recognition element that is hybridized to the biological sample to generate a circularized and ligated recognition element; d) generating a concatemeric amplification product from the circularized and ligated recognition element; and e) decoding the code from the concatemeric amplification product using soft decision decoding, thereby identifying the presence or absence of the one or more genetic variants in the one or more HLA genes from the biological sample.

120. The method of claim 119, further comprising hybridizing the biological sample to a plurality of second recognition elements, wherein each recognition element of the plurality of second recognition elements comprises a 3’ end that is complementary to a wildtype nucleic acid of the one or more genetic variants of the one or more HLA genes, and wherein the recognition element that comprises a 3’ end that is complementary to the wildtype nucleic acid of the one or more genetic variants comprises the same code as the recognition element that comprises the 3’ end that comprises the complementary sequence to the one or more genetic variants of the one or more HLA genes.WSGR Docket No. 64100-744.601 121. The method of claim 119, wherein the biological sample is from a blood sample, a buccal swab sample, a saliva sample, or a tissue sample.

122. The method of claim 121, wherein the biological sample is an extracted or purified nucleic acid sample from the blood sample, the buccal swab sample, the saliva sample, or the tissue sample.

123. The method of claim 119, wherein the one or more genetic variants comprises a single nucleotide polymorphism (SNP), an insertion, a deletion, or a copy number variant (CNV).

124. The method of claim 119, wherein the one or more genetic variants comprise from 10 to 1,000 genetic variants.

125. The method of claim 119, wherein ligating the recognition element comprises use of a thermostable ligase.

126. The method of claim 119, wherein generating the concatemeric amplification product comprises rolling circle amplification or multiple strand displacement.

127. The method of claim 119, further comprising treating the circularized and ligated recognition element with an exonuclease prior to generating the concatemeric amplification product.

128. The method of claim 119, wherein the circularized and ligated recognition element is placed on a substrate prior to generating the concatemeric amplification product.

129. The method of claim 128, wherein the substrate comprises one or more wells or one or more tubes that have been pre-treated with a nucleic acid immobilization composition.

130. The method of claim 128, wherein the substrate comprises one or more wells of a plate and each of the one or more wells is pre-treated with a polymer.

131. The method of claim 130, wherein the polymer is selected from the group consisting of: polyacrylamide, branched PEI, linear PEI, poly(β-aminoester) and poly(amidoamine), PEG, a gel, poly-L-lysine, silane, agarose, and muscle mimetic catecholamine polymer.

132. The method of claim 119, wherein decoding the code from the concatemeric amplification product comprises next generation sequencing or detection by hybridization.

133. The method of claim 132, wherein the code comprises at least two nucleic acid segments.

134. The method of claim 133, wherein the at least two nucleic acid segments are detected by hybridizing one or more detection polynucleotides to the at least two nucleic acid segments and imaging the hybridizing.

135. The method of claim 134, wherein each of the one or more detection polynucleotides is labeled with a different fluorescent moiety.WSGR Docket No. 64100-744.601 136. The method of claim 119, wherein the one or more genetic variants comprise one or more of HLA-A*31-01, HLA-B*15:02, HLA-B*57:01, or HLA-B*58:

01.

137. A method for identifying a variant in a gene of interest in the presence of a high homology region in a second gene, comprising: a) providing a biological sample suspected of having the variant in the gene of interest and the second gene; b) hybridizing the biological sample to: i) a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements comprises a 5’ end and a 3’ end, wherein the 5’ end of a recognition element comprises a distinguishing nucleotide which is present in the gene of interest but is not present in the second gene, and the 3’ end comprises a variant sequence to the variant in the gene of interest, and wherein each recognition element of the plurality of recognition elements further comprises a code that is a proxy for the variant of interest in the gene of interest; and ii) a third oligonucleotide; c) ligating the recognition elements that are hybridized to the biological sample to generate circularized and ligated recognition elements, wherein the recognition elements that are hybridized to the gene of interest are capable of being ligated compared to the second gene that is associated with the recognition elements which is not preferentially ligated; d) generating concatemeric amplification products from the circularized and ligated recognition elements; and e) decoding the code from the concatemeric amplification products using soft decision decoding, thereby identifying the presence or absence of the variant in the gene of interest from the biological sample in the presence of the high homology second gene.

138. The method of claim 137, wherein the 3’ end of the third oligonucleotide comprises the distinguishing nucleotide which is present in the gene of interest but not present in the second gene and the 5’ end of the recognition element comprises a wildtype nucleotide.

139. The method of claim 137, wherein the biological sample is from a blood sample, a buccal swab sample, a saliva sample, or a tissue sample, and wherein the biological sample is an extracted or purified nucleic acid sample from the blood sample, the buccal swab sample, the saliva sample, or the tissue sample.

140. The method of claim 137, wherein the variant comprises a single nucleotide polymorphism (SNP), an insertion, a deletion, or a copy number variant (CNV).

141. The method of claim 137, wherein ligating the recognition elements comprises use of a thermostable ligase.WSGR Docket No. 64100-744.601 142. The method of claim 137, wherein generating the concatemeric amplification product comprises rolling circle amplification or multiple strand displacement.

143. The method of claim 137, further comprising treating the circularized and ligated recognition elements with an exonuclease prior to generating the concatemeric amplification products.

144. The method of claim 137, wherein the circularized and ligated recognition elements are placed on a substrate prior to generating the concatemeric amplification products.

145. The method of claim 144, wherein the substrate comprises one or more wells or one or more tubes that have been pre-treated with a nucleic acid immobilization composition.

146. The method of claim 145, wherein the substrate comprises one or more wells of a plate and each of the one or more wells is pre-treated with a polymer, wherein the polymer is selected from the group consisting of: polyacrylamide, branched PEI, linear PEI, poly(β- aminoester) and poly(amidoamine), PEG, a gel, poly-L-lysine, silane, agarose, and muscle mimetic catecholamine polymer.

147. The method of claim 137, wherein decoding the code from the concatemeric amplification products comprises next generation sequencing or detection by hybridization.

148. The method of claim 147, wherein the code comprises at least two nucleic acid segments.

149. The method of claim 148, wherein the at least two nucleic acid segments are detected by hybridizing one or more detection polynucleotides to the at least two nucleic acid segments and imaging the hybridizing.

150. The method of claim 149, wherein each of the one or more detection polynucleotides is labeled with a different fluorescent moiety.

151. The method of any one of claims 137-150, wherein the second gene is a homology or a high homology pseudogene.

152. A composition comprising a nucleic acid sample, wherein the nucleic acid sample comprises sequences to a gene of interest and a high homology region in a second gene, wherein the gene of interest allows for a recognition element and a third oligonucleotide to be ligated and the high homology region in the second gene does not allow for the recognition element and the third oligonucleotide to be ligated.

153. A kit comprising: a) a plurality of recognition elements, wherein each recognition element in the plurality of recognition elements comprises a 3’ end that is complementary to a genetic variant or the wildtype sequence of the genetic variant of a pharmacogenomics related gene target; b) one or more of a ligase, a DNA polymerase and an exonuclease;WSGR Docket No. 64100-744.601 c) optionally a third oligonucleotide; d) a plurality of detection polynucleotides; and e) instructions for practicing the method of any one of claims 1-151.

154. A system for practicing the method of any one of claims 1-151.

Citation Information

Patent Citations

  • Multiscale lens systems and methods for imaging well plates and including event-based detection

    WO2023158993A2

  • Multi-high speed pharmacogenomic diagnostic kit for personalized pharmacotherapy and predicting drug side effects of multiple prescription drugs associated with cancer and chronic diseases

    WO2021029473A1

  • Encoded assays

    WO2022109496A2

  • Multiplexed detection of target biomolecules

    WO2023096672A1

  • Encoded assays

    WO2023096674A1

Cited By

  • Methods and compositions for increasing detection of structural variants

    WO2026161646A1