Identification of methylated nucleic acids

WO2026178412A2PCT designated stage Publication Date: 2026-08-27PLENO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/016115
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-02-20
Publication Date
2026-08-27

Smart Images

  • Figure US2026016115_27082026_PF_FP_ABST
    Figure US2026016115_27082026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure is directed to identifying the presence of methylated nucleic acid targets in a sample. Additionally, identifying the presence of other nucleic acid target sequences, such as variants or wild type in a nucleic acid sequence, can be combined with identifying the presence of methylation nucleic acid targets. Therefore, the present disclosure also provides an assay that can identify multiple different types of target sequences concurrently in one assay.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] IDENTIFICATION OF METHYLATED NUCLEIC ACIDS

[0002] CROSS-REFERENCE

[0003] The present application claims the benefit of United States Provisional Patent Application Serial No. 63 / 761,484, filed on February 21, 2025, and United States Provisional Patent Application Serial No. 63 / 778,269, filed on March 26, 2025, both of which are incorporated herein by reference in their entireties.

[0004] INTRODUCTION DNA methylation is an epigenetic mechanism that regulates gene expression and tissue differentiation by adding methyl groups to double stranded DNA with the help of DNA methyl transferases. Methylation plays a vital role in controlling organismal development, cell differentiation, gene regulation, genomic imprinting where genes can be expressed differently depending on their inheritance lineage, memory storage, biological clocks and aging, energy production in cells, and other biological processes.

[0005] However, abnormal methylation patterns can contribute to diseases like cancer, autoimmune disorders, cardiovascular diseases, and neurological conditions. DNA methylation basically acts like a molecular switch that can determine which genes are turned on and off in a cell depending on the environmental cues and developmental stages.

[0006] There is an abundance of evidence that is emerging indicating the importance of DNA methylation. As such, the study of DNA methylation and how it relates to human development and disease continues to be an important field of scientific study.

[0007] There are many challenges in studying epigenetic changes such as DNA methylation. For example, DNA methylation profiles can differ for each tissue as such epigenetic studies can be difficult depending on the tissue type. Each tissue type is usually made up of heterogeneous cell populations, leading to difficulties in studying epigenomic wide association studies. Additionally, DNA methylation can be influenced by genetic factors and environmental changes and therefore its role in disease etiology can be complicated to tease out.

[0008] SUMMARY

[0009] In some embodiments, the present disclosure provides a method for determining the methylation status of a sample, comprising cleaving a target nucleic acid with a restrictionendonuclease, wherein the restriction endonuclease cleaves one or more non-methylated nucleic acid sequences of interest of the target nucleic acid and does not cleave methylated nucleic acid sequences of the target nucleic acid, hybridizing one or more recognition elements to the target nucleic acid methylation sites of interest thereby generating one or more hybridized recognition elements, wherein the one or more recognition element comprises a 5’ end and a 3’ end, and wherein the 5’ end of the one or more recognition elements and the 3’ end of the one or more recognition elements are complementary to the target nucleic acid methylation sites of interest, ligating the 5’ end of the one or more hybridized recognition elements and the 3’ end of the one or more hybridized recognition element to produce one or more ligated recognition elements, and determining the methylation status of the sample based on the presence or absence of the one or more ligated recognition elements.

[0010] In some embodiments, the one or more recognition elements further comprise a hypercode. In some embodiments, the unique hypercode of a recognition elements can be used as a proxy to identify a unique target nucleic acid sequence.

[0011] In some embodiments, the restriction endonuclease for use in a method described herein is selected from the group consisting of Hhal, BstUI, Hpall. AccI, MspI, MspJI, Aatll, Acil, AcII, Afel, Agel, Asci, AsiSI, Aval, BceAI, BmgBI, BsaAI, BsaHI, BsiEI, BsiWI, BsmBI-v2, BspDI, BsrFI-v2, BssHII, BstBI, Clal, Eagl-HF, Esp3I, Paul, Fsel, FspI, Haell, Hgall, Hhal, HinPlI. HpyCH4IV, Hpy99I, KasI, Mlul, Nael, Narl, NgoNIV, Notl, Nt.BsmAI, Nt.CviPII. PaeR71, PluTi, Pmll, Pvul, SacII, Sall, Sfol, SgrAI, Smal, SnaBI, Srfl, TspMI, and Zral. In some embodiments, the preferred restriction endonuclease is Hhal.

[0012] In some embodiments, the restriction endonuclease for use in a method described herein is selected from the group consisting of Alwl, Bell, BclI-HF, DpnII, HphI, Mbol, Nt.AlwI, PspGI, and SexAI.

[0013] In some embodiments, the methylated nucleotides being detection comprise 5mC. 5hmC. or 6-mA.

[0014] In some embodiments, a method as described further comprises providing one or more additional recognition elements, wherein the one or more additional recognition elements hybridizes to one or more of a second type of target of interest in the target nucleic acid, wherein the one or more of the second type of target of interest comprises a nucleic acid variant. In some embodiments, the nucleic acid variant comprises a single nucleotide polymorphism, an insertion,a deletion, a copy number variant, or a splicing variant. Tn some embodiments, the nucleic acid variant is preferably a single nucleotide polymorphism.

[0015] In some embodiments, a method described herein further comprising amplifying the one or more ligated recognition elements, for example amplification by rolling circle amplification or multiple strand displacement amplification.

[0016] In some embodiments, the determining step of a method described herein comprises determining the methylation status of the sample by detecting the hypercode of the one or more ligated recognition elements, wherein detecting the hypercode of one or more ligated recognition elements comprises hybridizing two or more detection polynucleotide complexes to the hypercode, or a portion thereof, to generate a hypercode profile. In some embodiments, the hypercode profile is decoded using a soft decision decoding algorithm (pipeline).

[0017] In some embodiments, the method does not comprise chemically converting unmethylated cytosines prior to detecting the hypercode.

[0018] In some embodiments, the present disclosure comprises providing a composition comprising a nucleic acid sample and a first plurality of recognition elements hybridized to a plurality of methylated target nucleic acid sequences of interest in the nucleic acid sample and a second plurality of recognition elements hybridized to a plurality of variant nucleic acid target sequences of interest in the nucleic acid sample. In some embodiments, a composition can further comprise a ligase and / or one or more exonucleases. In some embodiments, one or more of the plurality of variant nucleic acid target sequences of interest of a composition described herein comprises one or more of a single nucleotide polymorphism, an insertion, a deletion, a copy number variant, or a splicing variant.

[0019] In some embodiments, the present disclosure provides a method for identifying the presence of a methylated nucleic acid in a sample, comprising providing an antibody-oligonucleotide conjugate to the sample, wherein the oligonucleotide of the antibody-oligonucleotide conjugate comprises a first sequence and a second sequence and wherein the antibody of the antibody-oligonucleotide conjugate binds a methylated nucleotide, binding the antibody of the antibody-oligonucleotide conjugate to the methylated nucleotide in the sample, providing two oligonucleotides to the sample, wherein the first oligonucleotide of the two oligonucleotides comprises a first portion that is complementary to a first sequence in the sample that is in proximity to the methylated nucleotide and a second portion that is complementary tothe first sequence of the oligonucleotide of the antibody-oligonucleotide conjugate and the second oligonucleotide of the two oligonucleotides comprises a first portion that is complementary to a second sequence in the sample that is in proximity to the methylation nucleotide and a second portion that is complementary to the second sequence of the oligonucleotide of the antibody-oligonucleotide conjugate, hybridizing the two oligonucleotides to their complementary sequences in the sample and hybridizing the oligonucleotide of the antibody-oligonucleotide conjugate to its complementary sequences on the two oligonucleotides, ligating the two ends of the oligonucleotides hybridized on the sample and the two ends of the oligonucleotides hybridized to the oligonucleotide of the antibody-oligonucleotide conjugate, thereby generating a circularized recognition element, amplifying the circularized recognition element to generate a concatemeric amplification product, and detecting the presence of the methylated nucleotide in the sample by the presence of the concatemeric amplification product. In some embodiments, the steps of the method as described can be performed substantially simultaneously or sequentially, or a combination of the two. In some embodiments, when the two oligonucleotides are ligated together a unique hypercode is generated. In some embodiments, the two oligonucleotides that hybridize to the oligonucleotide of the antibody-oligonucleotide conjugate generate an amplification primer binding site when ligated.

[0020] In some embodiments, the methylated nucleotides being detection comprise 5mC, 5hmC, or 6-mA. In some embodiments, the antibody of the antibody-oligonucleotide conjugate comprises Anti-5mC, Anti-5hmC, or Anti-6mA. In some embodiments, the antibody of the antibody-oligonucleotide conjugate is conjugated to the oligonucleotide of the antibody-oligonucleotide conjugate by a linker.

[0021] In some embodiments, when the two oligonucleotides are fully ligated they can be amplified, for example by rolling circle amplification thereby generating concatemeric amplification products. In some embodiments, detecting comprises detecting the presence of a hypercode in the concatemeric amplification product, wherein the hypercode is indicative of the presence of the methylated nucleotide in the sample.

[0022] In some embodiments, the present disclosure provides a composition comprising an antibody-oligonucleotide conjugate, wherein the antibody of the antibody-oligonucleotide conjugate is bound to a methylated nucleotide on a nucleic acid sample and a first portion of the oligonucleotide of the antibody-oligonucleotide conjugate is hybridized to a 5’ portion of a firstoligonucleotide and a second portion of the oligonucleotide of the antibody-oligonucleotide conjugate is hybridized adjacent to a 3’ portion of a second oligonucleotide.

[0023] 32. The composition of claim 31. further wherein a 3’ portion of the first oligonucleotide is hybridized to the nucleic acid sample in proximity to the methylated nucleotide and a 5’ portion of the second oligonucleotide is hybridized to the nucleic acid sample adjacent to the 3’ portion of the first oligonucleotide. In some embodiments, the adjacently hybridized portions of the first oligonucleotide and the second oligonucleotide are ligated to each other.

[0024] In some embodiments, the present disclosure provides a kit for practicing any of the methods disclosed herein, comprising a plurality of recognition elements that are capable of recognizing and hybridizing to a plurality of target nucleic acid methylation sites of interest, one or more enzymes, wherein the one or more enzymes comprise a methylation resistant restriction endonuclease, a ligase, a DNA polymerase, or an exonuclease, and instructions for performing any one of the methods disclosed herein. In some embodiments, a kit further comprises a plurality of recognition elements that are capable of recognizing and hybridizing to a plurality of variant nucleic acid target sites of interest.

[0025] In some embodiments, the present disclosure provides a system for determining methylation status in a sample, comprising an assay module comprising a hybridization reaction or a ligation reaction, or a combination thereof, wherein the hybridization reaction or ligation reaction, or the combination thereof, comprises one or more recognition elements that recognize and hybridize to one or more methylated nucleotides in the sample, and an analysis module configured to determine a presence of the one or more methylated nucleotides in the sample based on products generated from the assay module. In some embodiments, the assay module of a system further comprises one or more additional recognition elements that recognize and hybridize to one or more variant nucleotides in the sample and the analysis module determines the presence of the one or more variant nucleotides in the sample in addition to the one or more methylated nucleotides in the sample. In some embodiments, each of the one or more recognition elements of a system recognizes and hybridizes to the one or more methylated nucleotides or the one or more variant nucleotides comprises a unique hypercode. In some embodiments, the analysis module of a system comprises a plurality of detection polynucleotide complexes, where the determining the presence of the one or more methylated nucleotides or the one or more variant nucleotides comprises hybridizing the plurality of detection polynucleotide complexes tothe unique hypercode of the one or more recognition elements, or a portion thereof, imaging the hybridization to generate a hypercode profile for each of the one or more recognition elements, and determining the presence of the one or more methylated nucleotides or the one or more variant nucleotides by performing soft decision decoding on the hypercode profiles for each of the one or more recognition elements.

[0026] In some embodiments, the assay module of a system comprises a plurality of recognition elements that recognize and hybridize to a plurality of methylated nucleotides in the sample and the analysis module determines the presence of the plurality of methylated nucleotides in the sample. In some embodiments, the assay module of a system further comprises a plurality of recognition elements that recognize and hybridize to a plurality of variant nucleotides in the sample and the analysis module determines the presence of the plurality of variant nucleotides in the sample in addition to the plurality of methylated nucleotides in the sample. In some embodiments, each of the plurality of recognition elements of a system recognizes and hybridizes to the plurality of methylated nucleotides or the plurality of variant nucleotides comprises a unique hypercode. In some embodiments the analysis module of a system comprises a plurality of detection polynucleotide complexes, where the determining the presence of the plurality of methylated nucleotides or the plurality of variant nucleotides comprises hybridizing the plurality of detection polynucleotide complexes to the unique hypercode of each of the plurality of recognition elements, or a portion thereof, imaging the hybridization to generate a hypercode profile for each of the plurality of recognition elements, and determining the presence of the plurality of methylated nucleotides or the plurality of variant nucleotides by performing soft decision decoding on the hypercode profiles for each of the plurality of recognition elements.

[0027] In some embodiments, a system described herein further includes a step where a sample is pre-treated with a methylation resistant restriction endonuclease prior to the hybridization reaction. In some embodiments, the assay module of a system further comprises an amplification reaction, such as rolling circle amplification, for generating one or more of, or a plurality of, concatemeric amplification products.

[0028] In some aspects, the disclosure provides a method for determining methylation status of a target nucleic acid, the method comprising: obtaining a plurality of amplification products, each amplification product comprising a hypercode associated with a target nucleic acid methylationsite; iteratively hybridizing a plurality of detection polynucleotide complexes to the plurality of amplification products over a plurality of detection cycles, wherein each detection polynucleotide complex comprises a detection oligonucleotide with a detectable label and an anchor oligonucleotide complementary to a segment of the hypercode; imaging the detectable labels after each detection cycle to obtain signal intensities associated with segments of the hypercode; and applying a soft decision decoding algorithm to the signal intensities obtained across the plurality of detection cycles to determine a probability that the hypercode is present, wherein the probability determination is performed without assigning a nucleotide identity to individual positions within the hypercode; thereby determining methylation status of the target nucleic acid.

[0029] In some embodiments, the hypercode is selected from a codespace having a Hamming distance of 3. In some embodiments, the hypercode is selected from a codespace having a Hamming distance of 4. In some embodiments, the hypercode is selected from a codespace having a Hamming distance of 5.

[0030] In some embodiments, the signal intensities are obtained from amplification products immobilized on a substrate.

[0031] In some embodiments, applying the soft decision decoding algorithm comprises crosscorrelating the signal intensities against expected hypercode signal profiles and assigning a most likely hypercode based on a probabilistic metric.

[0032] In some aspects, the disclosure provides a method for determining methylation status of one or more target nucleic acid sites in a sample, the method comprising: providing a sample comprising a target nucleic acid containing one or more target methylation sites; hybridizing a plurality of recognition elements to the target nucleic acid, wherein each recognition element is configured to preferentially interact with a methylated or unmethylated form of at least one target methylation site, wherein each recognition element comprises a unique hypercode; generating assay data for the plurality of recognition elements, wherein the assay data comprises one or more of binding kinetics, amplification product counts, or signal intensities associated with the hypercodes; processing the assay data using a machine learning model comprising a database of methylation signatures derived from differential hybridization behavior of methylated and unmethylated target sites; and determining, based on the processing, a methylation status for at least one target methylation site.In some embodiments, the methylation status is determined based on a probability of methylated nucleotides.

[0033] In some embodiments, the methylation status is determined based on a percentage of methylated nucleotides.

[0034] In some embodiments, the methylation status is determined based on a number of methylated nucleotides.

[0035] In some embodiments, generating assay data comprises iteratively detecting binding kinetics associated with the hypercodes over a plurality of detection cycles.

[0036] In some embodiments, generating assay data comprises iteratively detecting amplification product counts associated with the hypercodes over a plurality of detection cycles.

[0037] In some embodiments, generating assay data comprises iteratively detecting signal intensities associated with the hypercodes over a plurality of detection cycles.

[0038] In some embodiments, processing the assay data comprises applying a soft decision decoding algorithm to binding kinetics obtained across the plurality of detection cycles to determine a probability that the hypercode is present, wherein the probability determination is performed without assigning a nucleotide identity to individual positions within the hypercode.

[0039] In some embodiments, processing the assay data comprises applying a soft decision decoding algorithm to amplification product counts obtained across the plurality of detection cycles to determine a probability that the hypercode is present, wherein the probability determination is performed without assigning a nucleotide identity to individual positions within the hypercode.

[0040] In some embodiments, processing the assay data comprises applying a soft decision decoding algorithm to signal intensities obtained across the plurality of detection cycles to determine a probability that the hypercode is present, wherein the probability determination is performed without assigning a nucleotide identity to individual positions within the hypercode. In some embodiments, the hypercode is selected from a codespace having a Hamming distance of 3. In some embodiments, the hypercode is selected from a codespace having a Hamming distance of 4. In some embodiments, the hypercode is selected from a codespace having a Hamming distance of 5.

[0041] In some embodiments, the binding kinetics are obtained from amplification products immobilized on a substrate.In some embodiments, the amplification product counts are obtained from amplification products immobilized on a substrate.

[0042] In some embodiments, the signal intensities are obtained from amplification products immobilized on a substrate.

[0043] In some embodiments, processing the assay data comprises cross-correlating the signal intensities against expected hypercode signal profiles and assigning a most likely hypercode based on a probabilistic metric.

[0044] In some embodiments, the plurality of recognition elements comprises at least two recognition elements that are complementary to different numbers of methylated nucleotides within a target region of the target nucleic acid, wherein the target region comprises a plurality of target methylation sites, and wherein determining the methylation status comprises determining a methylation signature for the target region based on assay data associated with different hypercodes of the plurality of recognition elements.

[0045] BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The features of the inventive concepts are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present inventive concepts will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the inventive concepts are utilized, and the accompanying drawings of which:

[0047] FIG. 1 shows an illustration of a recognition element used in methods of the present disclosure.

[0048] FIG. 2 shows an illustration of an assay workflow used in methods of the present disclosure.

[0049] FIG. 3 shows a detection polynucleotide used in methods of the present disclosure; A) an example of a detection polynucleotide, and B) an example of a detection polynucleotide hybridized to its complementary sequence in a recognition element.

[0050] FIG. 4 shows a schematic diagram of an example of a soft decision decoding algorithm (pipeline) for decoding a hypercode of a recognition element.

[0051] FIG. 5 shows an illustration of a methylation assay building on the assay workflow as shown in FIG.2.FIG. 6 shows an illustration of amplification of a recognition element.

[0052] FIG. 7 shows gene and chromosomal locations for methylation targets of interest and the recognition element 5’ and 3’ end sequences complementary to the target sequences of interest.

[0053] FIG. 8 shows gene and nucleotide reference (WT, REF) and alternate single nucleotide variant (SNV, ALT) locations for targets of interest and the recognition element 5’ and 3’ end sequences complementary to the target sequences of interest.

[0054] FIG. 9 shows an example of a graph for the three SNV recognition elements that were chosen for use as internal calibrators to normalize methylation data, as they targeted chrl2:65121253C_C.

[0055] FIG. 10 shows an example of a graph for the three SNV recognition elements that were chosen for use as internal calibrators to normalize methylation data, as they targeted chr7:24284322C-C.

[0056] FIG. 11 shows an example graph demonstrating that detection of SNV targets of interest were not inhibited by the presence of a methylation resistant restriction endonuclease.

[0057] FIG. 12 shows an example graph of multiomic detection of methylation status of targets of interest and SNVs of interest in a sample.

[0058] FIGs. 13A-D show example graphs of methylation detection; A) 1%, B) 2.5%, C) 5% and D) 10% methylation at a site in a hypermethylated NA12878 sample using a Student’s t-test.

[0059] FIG. 14 shows exemplary data for range of detection of a methylated target site in MLH1 using an All Pairs, Tukey Kramer test.

[0060] FIG. 15 shows exemplary data for range of detection of a methylated target site in TWIST1 using an All Pairs. Tukey Kramer test.

[0061] FIG. 16 shows exemplary data for range of detection of a methylated target site in APC using an All Pairs, Tukey Kramer test.

[0062] FIG. 17 shows an example of a diagram illustrating a recognition element complementary to a variant sequence.

[0063] FIG. 18 shows an example of a diagram illustrating a recognition element complementary to a variant sequence.

[0064] FIG. 19 shows a scenario illustrating methylation detection by oligonucleotide-antibody conjugate binding to methylated targets.FIG. 20 shows a scenario wherein differential hybridization kinetics of a recognition element to either a methylated or unmethylated nucleotide can be utilized for machine learning training of datasets for generating a methylation database for differentially methylated CpG sites for methylation calling in a sample.

[0065] FIG. 21 shows another scenario wherein differential hybridization kinetics of a recognition element to either a methylated or unmethylated nucleotide at a plurality of potential methylated sites can be utilized for machine learning training of datasets for generating a methylation database for differentially methylated regions for methylation calling in a sample.

[0066] DETAILED DESCRIPTION DNA methylation, in the mammalian genome, is an epigenetic process that primarily involves the transfer of a methyl group onto a cytosine at the C5 position, thereby forming 5-methylcytosine. The epigenetic process of DNA methylation is catalyzed by a family of enzymes, DNA methyltransferases, which transfer a methyl group from S-adenosylmethionine (SAM) to the fifth carbon of a cytosine, thereby generating 5-methylcytosine. The majority of DNA methylation occurs in mammals on a cytosine that precedes a guanine in a nucleic acid sequence, these sites are called CpG islands. Cytosine methylation is also present in plants and in prokaryotic species, albeit to a lesser extent in prokaryotes. Cytosine methylation is known to regulate gene expression by recruiting gene expression related proteins, or by inhibiting the binding of DNA to transcription factors. The pattern of DNA methylation changes due to dynamic processes in DNA methylation and DNA de-methylation, resulting in DNA methylation patterns that can regulate tissue specific gene transcription and other cellular processes that affect mammalian gene and protein metabolic processes.

[0067] Adenine methylation may also be present in mammals; however it is mostly prevalent in prokaryotic species. Adenine methylation is a process where a methyl group is added to the N6 position of an adenine base via a DNA adenine methyltransferase and transfer of a methyl group from a donor such as SAM to the adenine N6 location. Adenine methylation is considered an epigenetic marker in bacteria which is thought to regulate cellular functions such as DNA replication and repair and transcription. As such, the study of adenine methylation can also be important when studying pathogenic bacteria and how it might relate to human disease and pathology.The study of changes in gene activity caused by direct alterations of a DNA sequence, such as by epigenetic processes like methylation, is an emerging field of study and important to understand the human genome and how cellular processes can lead to human phenotypes and disease states, such as cancer.

[0068] One problem is it can be challenging to identify changes in methylation states in a sample with sensitivity and specificity. The present disclosure provides innovative methods and compositions for studying DNA methylation, thereby providing additional tools and solutions for use in deciphering the complex field of epigenetics.

[0069] The present disclosure provides methods and compositions to help researchers study DNA methylation and thereby gain deeper understanding of how epigenetics, in this case DNA methylation events, can effect changes in a genome which might be related to diseases in humans and other species.

[0070] The present disclosure provides methods, systems and compositions for a platform that enables high-plex detection and quantitation of targets of interest from a biological sample, in this disclosure the targets of interest comprise DNA methylation targets of interest which can optionally be multiplexed with detection of additional targets of interest such as variant targets of interest from the same sample. Detection of the targets of interest comprise use of a recognition element which recognizes and hybridizes to target sequences of interest and wherein each recognition element comprises a hypercode, also known as a code, which serves as a proxy for the presence of a target of interest from a biological sample for a fast, simple and highly-plexed assay. The methods and system presented herein provide for detecting and quantifying multiple biological samples and multiple targets of interest in those biological samples, in parallel and at high plexity.

[0071] The present disclosure provides methods, systems and compositions for multiplexed target molecule detection utilizing a detection and decoding approach. The target molecule may be a nucleic acid molecule from a sample (e.g., a biological sample) or a nucleic acid molecule serving as a surrogate of a target molecule that is other than a nucleic acid molecule (e.g., polypeptide, protein, sugar, metabolite, etc.). Methods, systems and compositions of the present disclosure provide encoded assays (or components of the encoded assays) comprising a recognition element that uniquely recognizes and hybridizes to a target molecule from a sample under conditions sufficient that the recognition element undergoes a molecular transformation inthe presence (but not in the absence) of the target molecule to produce a circularized and ligated recognition element.

[0072] Such modified recognition elements comprise a nucleic acid code or hypercode, which is associated or correlated with the 5’ and 3’ regions of a recognition element that is complementary to a target molecule. The circularized and ligated recognition element may be amplified to produce multiple copies of the recognition element including its hypercode. For example, where a ligation event adjoining a 3’ probe arm and a 5’ probe arm of the recognition element (in the presence of the target of interest), the amplification may be rolling circle amplification (RCA). The code or hypercode may be made up of one or more segments wherein each segment may hybridize to a detectable oligonucleotide or a detection polynucleotide complex to determine the presence of the one or more segments of the hypercode, which can be correlated to the presence of the target molecule. The detection polynucleotide complexes of the present disclosure may include a detection oligonucleotide having a detectable label and a nucleic acid sequence configured to hybridize to an anchor oligonucleotide, wherein the anchor oligonucleotide is configured to hybridize to a segment of the hypercode and the detection oligonucleotide, as shown for example in FIG. 3. Diverse pools of the detection polynucleotide complexes utilized over multiple rounds of detection, also called flows or cycles, (e.g., for each segment, or subsegment, of a hypercode) are able to deliver a plurality of signals (e.g., fluorescent signals such as fluorescent color signals) that may be combined to generate a pattern of detected signals, which can be subsequently decoded (e.g., via soft decision decoding) to determine the probability of the presence of the hypercode which serves as a proxy or surrogate for the presence of the target molecule in the sample.

[0073] Recognition Elements and Hypercode Design

[0074] As described herein, methods for identifying methylation presence and patterns in a target of interest, optionally in combination with SNP detection, comprise hypercoding of targets which can be enabled using a multi-functional nucleic acid molecule called a “recognition element”. A recognition element, as illustrated in FIG. 1, comprises a circularizable linear DNA molecule that comprises a code or hypercode unique to a particular target of interest and 5’ and 3’ ends that are complementary to a target sequence of interest, such that when hybridized to target sequences the recognition element can conform into a padlock probe configuration. In the presence of the complementary genomic sequence, methylated, wild type or variant, arecognition element can hybridize to its intended target sequences of interest, undergo a conformation change into a circular DNA molecule, and the two adjacent ends of the circularized recognition element can be ligated together to form a circularized and ligated recognition element. In the absence of target sequences of interest, there is expected to be no or minimal hybridization and therefore no circularized DNA molecule and no ligation of adjacent ends. After ligation, the circularized recognition element can be amplified to increase the number of hypercodes for downstream fluorescence detection.

[0075] The methods described herein relate to determining the presence of one or more methylated targets of interest, and optionally one or more SNV targets of interest, from a target of interest in a sample. In some embodiments, the presence of a target molecule is determined by introducing a plurality of recognition elements to the sample, wherein each recognition element comprises a target recognition region specific to a methylation target, a wild type target, or a SNV target in the sample under conditions sufficient to hybridize the recognition elements to the respective targets. In some embodiments, the target recognition regions are complementary to the respective target nucleic acid molecule. In some embodiments, the recognition element comprises a hypercode that is associated with target molecules by way of alignment with the 5’ and 3’ ends of the recognition element that may be detected and used as a surrogate, or proxy, for the presence of the methylated target molecule or the SNV target molecule. In some embodiments, the hypercode comprises a plurality of segments that each correspond to one or more signals (e.g., fluorescent colors, signals) that are used in a decoding process to determine the presence of the hypercode and therefore the presence of the methylation, wild type and SNV target nucleic acid molecules. In some embodiments, the decoding process comprises a soft decision decoding algorithm.

[0076] The recognition elements hybridized to the respective target molecules may be selectively amplified to produce a plurality of amplification products comprising amplified hypercodes. In some embodiments, the amplification products are immobilized on a substrate (e.g., welled plate, flow cell). In some embodiments, amplification to generate amplification products is performed in solution. In some embodiments, amplification to generate amplification products is performed on a substrate.

[0077] In some embodiments, a plurality of detection polynucleotide complexes are introduced to the plurality of amplification products. In some embodiments, a detection oligonucleotide andan anchor oligonucleotide are added to the amplification products, wherein the detection oligonucleotide and the anchor oligonucleotide assemble to form a detection polynucleotide complex substantially simultaneously to the anchor oligonucleotide hybridizing to a hypercode or a portion of a hypercode.

[0078] In some embodiments, the detection polynucleotide complex comprises a detection oligonucleotide and an anchor oligonucleotide. A portion of an anchor oligonucleotide may be complementary to at least a portion of a hypercode of the amplification product. Another portion of the anchor oligonucleotide may be complementary to at least a portion of the detection oligonucleotide. When the detection polynucleotide complex has an anchor oligonucleotide complementary to a portion of a hypercode of an amplification product, detectable binding complexes with the portion of the hypercode are formed. In some embodiments, the detectable binding complexes are imaged with an imaging system to obtain signals associated with the hypercode for each amplification product, thereby generating a hypercode profile which can be decoded. This process may be repeated for each segment, or a portion thereof, of each hypercode, thereby building a color profile for each hypercode. A decoding process (e.g., soft decision decoding) may be applied to the hypercode profile to determine or predict the probability of the presence of each hypercode, thereby identifying the presence of the target molecules in the sample.

[0079] In some embodiments, provided herein are methods comprising analyzing a plurality of methylated target nucleic acid molecules from a sample, providing a plurality of recognition elements, wherein each recognition element of the plurality comprises one or more target recognition regions complementary to a corresponding methylated target nucleic acid molecule of the plurality of methylated target nucleic acid molecules; and a hypercode from a set of hypercodes, wherein the hypercode is associated with one or more methylated target nucleic acid molecules of the plurality from the sample including the corresponding methylated target nucleic acid molecule.

[0080] In some embodiments, the methods comprise selectively amplifying a subset of the plurality of recognition elements to produce a plurality of amplification products, wherein each amplification product comprises a hypercode, introducing a plurality of detection polynucleotide complexes to the plurality of amplification products, wherein each detection polynucleotide complex of the plurality comprises a detection oligonucleotide and an anchor oligonucleotide,wherein a portion of each anchor oligonucleotide is complementary to at least a portion of the amplification product of the plurality of amplification products, and another portion of the anchor oligonucleotide is complementary to at least a portion of the detection oligonucleotide forming a plurality of detectable binding complexes. In some embodiments, each detectable binding complex comprises a detection polynucleotide complex bound to an amplification product of the plurality of amplification products, further including imaging the plurality of detectable binding complexes to obtain signals associated with the different segments of the plurality of segments of each hypercode of the plurality of hypercodes for each amplification product of the plurality of amplification products, iteratively repeating the operations of the introducing, forming, and imaging for each segment of each hypercode of the plurality of hypercodes and applying a soft decision decoding algorithm to the plurality of hypercodes to predict a presence of the one or more methylated target nucleic acid molecules from the sample. The same can be done concurrently for SNV target nucleic acid molecules from the sample and / or wild type target nucleic acid molecules from the sample.

[0081] In some embodiments, provided herein are methods comprising analyzing a plurality of methylated target nucleic acid molecules from a sample, providing a plurality of recognition elements, wherein each recognition element of the plurality comprises one or more target recognition regions complementary to a corresponding methylated target nucleic acid molecule of the plurality of methylated target nucleic acid molecules; and a hypercode from a set of hypercodes, wherein the hypercode is associated with one or more methylated target nucleic acid molecules of the plurality from the sample including the corresponding target nucleic acid molecule. In some embodiments, the hypercode comprises a plurality of segments that corresponds to at least two signals of a set of signals, selectively amplifying a subset of the plurality of recognition elements bound to the plurality of target nucleic acid molecules to produce a plurality of amplification products, wherein each amplification product comprises a unique hypercode. In some embodiments, the methods comprise introducing a plurality of detection oligonucleotides to the plurality of amplification products, wherein each detection oligonucleotide comprises a sequence complementary to a portion of a hypercode and further comprises a detectable moiety and when bound to a segment of the hypercode is called a detectable binding complex, wherein each detectable binding complex comprises a detection oligonucleotide bound to an amplification product of the plurality of amplification products,imaging the plurality of detectable binding complexes to obtain signals associated with the different segments of the plurality of segments, or portions thereof, of each hypercode of the plurality of hypercodes for each amplification product of the plurality of amplification products, iteratively repeating the operations of the introducing, forming, and imaging for each segment, or a portion thereof, of each hypercode of the plurality of hypercodes and applying a soft decision decoding algorithm to the plurality of hypercodes to predict a presence of the one or more methylated target nucleic acid molecules from the sample. The same can be done concurrently for detection of SNV target nucleic acid molecules, or wild type nucleic acid molecules, from the sample.

[0082] In some embodiments, the methods described herein may comprise iteratively repeating the operations of: (i) introducing detection polynucleotide complexes: (ii) forming detectable binding complexes; and (iii) imaging the detectable binding complexes to obtain signals, in order to generate a pattern of states indicative of each hypercode. In some embodiments, the method is performed for each segment, or a portion thereof, within a hypercode. In some embodiments, each operation comprises a cycle whereby a first pool of detection polynucleotide complexes is introduced to a plurality of modified recognition elements associated with respective target molecules, or amplification products thereof, under conditions sufficient to bind a detection polynucleotide complex to a modified recognition element or an amplification product thereof to facilitate detection of the segment of a hypercode. In some embodiments, the iteratively repeating operations (i) to (iii) comprises adding and cycling additional pools of detection polynucleotide complexes in a sequential manner until all or substantially all of the segments for each of the hypercodes are detected. In some embodiments, methods further comprise a wash step between operations (ii) and (iii) to remove one or more of unbound detection polynucleotide complexes, unbound detection oligonucleotides, or anchor oligonucleotides. In some embodiments, methods further comprise a dehybridization operation in between cycles to destabilize and remove a first pool of detection polynucleotide complexes prior to the addition of a second pool of detection polynucleotide complexes. The number of segments of a hypercode is dependent on the complexity of the assay, as such a hypercode can comprise two, three, four, five, six, seven, eight, nine, ten and more segments that are combined to provide one hypercode, with the detection cycles increasing depending on the number of segments that comprise a hypercode.FIG. 2 provides a non-limiting example of a method for detecting the presence of a target nucleic acid molecule comprising the assay proper followed by detection and decoding.

[0083] The assays disclosed herein are capable of multiplex target detection, in this disclosure target detection is methylation and SNVs. The readout of the assays can be concurrent for the target types, thereby enabling a platform for the analysis of different target molecules from a sample.

[0084] Encoding

[0085] An assay workflow, according to some embodiments herein, may comprise the following operations. A sample may be collected or provided. The sample may be whole blood, lymphatic fluid, serum, plasma, sweat, tear, saliva, sputum, cerebrospinal fluid, amniotic fluid, seminal fluid, vaginal excretion, serous fluid, synovial fluid, pericardial fluid, peritoneal fluid, pleural fluid, transudates, exudates, cystic fluid, bile, urine, gastric fluid, intestinal fluid, fecal samples, liquids containing single or multiple cells, liquids containing organelles, fluidized tissues, cell free DNA (cfDNA), circulating tumor cell DNA (ctcDNA), circulating tumor cells, fluidized organisms, liquids containing multi-celled organisms, tissues, cells, biopsy samples, biological swabs or biological washes. For example, nucleic acids from a sample can be extracted and optionally purified and said sample can be used in the methods described herein.

[0086] Target nucleic acid extraction, concentration, and / or purification processes may be performed. The target molecule may be DNA from any source as previously described. DNA in a plasma sample, for example cfDNA, may be extracted, purified, and concentrated for analysis. A proteinase K digestion step may be used to digest proteins present in the plasma sample or a tissue sample. In some cases, a heat denaturation step (e.g., 94-98°C for 20-30 seconds) may be used to denature double- stranded DNA into single- stranded DNA. A bead-based extraction and concentration protocol may be used to capture single- stranded DNA from the plasma sample or any other sample. In some embodiments, the bead-based extraction protocol uses magnetically responsive nucleic acid capture beads. The bead-bound DNA may be released from the capture beads using an elution buffer (or other elution means suitable to the capture bead used) to produce a processed DNA sample for analysis. In some embodiments, target molecules can be extracted from cells or tissues, or any other sample of interest, using known protocols. A skilled artisan will understand the methods that can be used to extract and / or purify target molecules such as nucleic acids or proteins, from a sample.The DNA sample may be transferred into an analysis device, welled plate, tube or other container according to some embodiments herein. The container may comprise a reaction vessel. Non-limiting examples of reaction vessels include a plate, a well, a container, a tube, a flow cell, a microfluidic chip, and the like. The plate may be a welled plate, such as a 12-well plate, a 24-well plate, a 48-well plate, a 96-well plate, a 384 well plate, a 1536-well plate, and the like. The reaction vessel, or a reaction surface thereof, may be optically clear to enable optical target detection directly in the reaction vessel. The reaction vessel may comprise a glass surface. The reaction vessel may comprise a glass-bottomed, well plate. The reaction vessel may comprise a surface coating that promotes sequestration of nucleic acid amplification products and immobilization of the same. For example, the reaction vessel may comprise a cationic coating which serves to immobilize nucleic acids to the vessel.

[0087] A recognition event for each target molecule in a set of target molecules may be performed. FIG.2 provides an example of an assay workflow 200. A recognition element 201 (see also FIG. 1) comprises 5’ and 3’ ends that are complementary to target sequences of interest. A recognition element further comprises a code, also known as a hypercode, in this instance there are four segments to the hypercode, wherein the hypercode can be used as a proxy for indirect detection of the presence of a target of interest. Target nucleic acids of interest are incubated with the recognition element wherein they can hybridize, if present, to their complementary sequences in the recognition element 202. After hybridization (e.g., in the presence of the target sequence of interest), the ends of the recognition element which can be adjacently located can be ligated 203. If there is no target of interest, there is no hybridization and no ligation. The reactions can be treated with one or more exonucleases 204, thereby digesting any linear nucleic acids present in the reaction that did not participate in the hybridization and ligation events.

[0088] Still referring to FIG.2, after exonuclease digestion the circularized recognition elements 205 can be aliquoted into a welled plate 206 where they can be immobilized onto the surface of the welled plate. The addition of a polymerase (e.g., DNA polymerase) allows for amplification and concatenation of the circularized recognition elements. The concatenated amplified products 207 can be queried with fluorescent detection complexes 208 (see also FIG. 3) and imaged 209.

[0089] A hypercode profile can be generated upon multiple query events, also called cycles or flows.The resulting hypercode profile can be decoded to identify the hypercode which can be used as a proxy for the presence of the target of interest.

[0090] The target molecule can be uniquely recognized by and hybridized to a recognition element associated with a hypercode (and optionally other elements). In one example, the recognition event for the set of target molecules uses a plurality of hypercoded recognition elements. In another example, the recognition event for the set of target molecules uses a panel of molecular inversion probes. In another example, the recognition event for the set of target molecules uses a panel of padlock probes. The recognition event yields a set of hypercoded target molecules comprising the target molecule and the recognition element comprising the hypercode.

[0091] The recognition event may include sequence- specific hybridization between a 5’ end and a 3’ end of the recognition element to the target nucleic acid molecule under conditions sufficient to form a hybridization complex comprising the recognition element and the target nucleic acid molecule. In embodiments where the recognition element forms a padlock probe configuration, the 5’ end and the 3’ end comprises a target recognition element that hybridizes to two adjacent sequences in the target molecule. In another embodiment, the 5’ end and the 3’ end hybridize to the target nucleic acid molecule at 3’ and 5’ regions flanking the target region thereby leaving a gap between the 5’ end and the 3’ end of the recognition element. If a gap is present, an extension reaction can be performed to extend the 3’ end of the gap to reach the 5’ end, thereby once again putting the two ends in proximity of each other for ligation.

[0092] A ligation reaction can ligate together the end of the recognition element, comprising one or more enzymes under conditions sufficient to circularize the recognition element. The ligating enzyme may be a DNA ligase or catalytically active portion thereof. Non-limiting examples of ligases include any ligase which can ligate the 3’ hydroxyl group to a 5’ phosphate group of a DNA molecule, for example T4 DNA ligase, thermostable T4 DNA ligase. AmpLigase Thermostable DNA ligase, HiFi Taq DNA ligase, and the like. A gap-fill ligation reaction may occur where there is a gap between the 3’ end and the 5’ end of the recognition element following hybridization of the recognition element to the target nucleic acid molecule. In addition to the ligase, a polymerizing enzyme to synthesize DNA in the gap may be used. The polymerizing enzyme may be a DNA polymerase or a catalytically active portion thereof. The DNA polymerase may be a thermostable polymerase, a thermolabile polymerase, Bstpolymerase, a Bst-like polymerase, a Therminator X polymerase, a Bst3.0 polymerase, an ArcticZymes polymerase, or a Bsm DNA polymerase. The ligation event may produce a circularized recognition element (e.g„ a version of the recognition element that is ligated or gap-filled and ligated). In one example, a ligation or gap-fill ligation reaction generates a circular and ligated recognition element.

[0093] An exonuclease cleanup step may be used following ligation of the recognition element to digest any remaining single stranded nucleic acids, such as unhybridized recognition elements, amplification primers, and single-stranded target molecules. Non-limiting examples of exonucleases useful for digesting remaining single-stranded nucleic acids include Exonuclease I, Exonuclease III, Exonuclease VII, Msz Exonuclease I, T5 exonuclease, Exonuclease V, DNase I, or any combination thereof.

[0094] Decoding

[0095]

[0096] An amplification reaction on the circularized and ligated recognition elements may be performed. In one example, the amplification event may be a rolling circle amplification (RCA) reaction to generate a set of concatenated amplification products. The amplification reaction thereby yields a set of concatenated amplified recognition elements including their unique hypercodes (e.g., hypercodes present in the circularized recognition elements) that can be correlated to the target molecule. An amplification reaction could further be a multiple strand displacement reaction to generate a set of target molecule- specific amplification products.

[0097] A detection event followed by a decoding event for each amplified recognition element may be performed to identify the hypercode of an amplified recognition element. In one example, the hypercode may be detected by hybridization of one or more segments, or portions thereof, of the hypercode (and optionally other elements) to a detection polynucleotide complex or a detection oligonucleotide. The detection events detect the hypercode as a surrogate or proxy for identifying the presence of the target molecule in the sample. Decoding the detection event profile may in some cases make use of a soft decision decoding algorithm.

[0098] A bioinformatics analysis of the hypercode information (and optionally other elements) from the detection operation may be performed. The bioinformatic analysis may be performed by one or more computer systems as described herein.In some embodiments, the amplification reaction and the detection event may occur in a step wise manner, such that first the modified recognition elements are amplified followed by one or more washes, followed by the detection event and the decoding event. In some embodiments, presence of the hypercodes may be determined with a detection and decoding by hybridization process disclosed herein. For example, a plurality of detection polynucleotide complexes or detection oligonucleotides may be introduced to the amplified recognition elements iteratively for detection of each segment, or a portion thereof, within all or substantially all amplified hypercodes in the amplified recognition elements. FIG. 5 is a schematic diagram illustrating an example of a methylation detection workflow 500 for detecting a methylated target nucleic acid of interest. In this disclosure, there is no bisulfite treatment of a sample prior to assay, as the methods described herein are directed to identifying nascent methylation of a nucleic acid from a sample. As such, A DNA sample may include a methylated or unmethylated target. The extracted DNA sample 501 is subjected to a restriction endonuclease 502 that is resistant to DNA methylation, as such a methylated target site of interest will not be cleaved by the restriction endonuclease and a non-methylated target site of interest will be cleaved.

[0099] Methylation resistant restriction enzymes cleave double stranded DNA at unmethylated cytosine residues. If there is a methylated cytosine, for example CpG methylation in mammals, the methylation resistant restriction enzyme will not cleave that site thereby leaving the methylated site intact. Additional methylation sites such as methyladenine (6mA) can also be queried using methylation resistant restriction enzymes.

[0100] In methods described herein, the restriction endonuclease Hhal was utilized, however there are myriad other methylation resistant restriction endonucleases that could also be used, including but not limited to, BstUI, Hpall. AccI, MspI, MspJI, Aatll, Acil, AcII, Afel, Agel, Alwl, Asci, AsiSI, Aval, BceAI, BmgBI, BsaAI, BsaHI, BsiEI, BsiWI, BsmBI-v2, BspDI, BsrFI-v2, BssHII. BstBI. Clal. Eagl-HF, Esp3I, Faul, Fsel, FspI, Haell, Hgall, Hhal, HinPH, HpyCH4IV, Hpy99I, KasI, Mlul, Nael, Narl, NgoNIV, Notl, Nt.BsmAI, Nt.CviPII, PaeR71, PluTi. Pmll, Pvul, SacII, Sall, Sfol, SgrAI, Smal, SnaBI, Srfl, TspMI, Zral, Bell, DpnII, HphI, Mbol, Nt.AlwI, PspGI, and SexAI.

[0101] If desired, an optional amplification 503 of the methylated targets can be performed to increase the amount of methylated targets of interest, whereas the cleaved non-methylated targets are not capable of being amplified.Following the restriction digest 502 and optional amplification 503, a recognition element comprising a hypercode and targeted 5’ end and 3’ end sequences can be hybridized to the methylated target of interest 504, and the adjacent ends ligated together. When a target sequence has been cleaved 502, a recognition element can still hybridize 504, however there is no ligation event.

[0102] After ligation and amplification as previously described in FIG.2, detection decoding and analysis 506 can be performed to identify the presence of a methylated target of interest. FIG. 5 shows example bar graphs of what decode counts would be expected should a methylation target of interest be present. Additionally, if a multiomics assay is performed concurrently on the same sample, for example if SNVs are also queried with targeted recognition elements, a bar graph representing what decode counts might be expected is also represented. When assaying for SNV targets in combination with methylation targets, it should be noted that restriction digest is not expected to affect the detection of SNVs, unless the target SNV sequence is part of the restriction endonuclease cut site, which should be taken into account when determining which restriction endonuclease to use.

[0103] In one embodiment of the workflow 500, the recognition element may generate a padlock probe configuration as shown in 504. In another embodiment, a molecular inversion probe configuration that includes a 3 '-terminal single base gap at a target site of interest may be used. A gap-fill and ligation event using only a single added nucleotide (at a minimum) may be used to generate the circular ligated recognition element comprising the hypercode only when the nucleotide corresponding to the target site of interest is incorporated. This approach provides two forms of specificity to the assay: (i) the 3 '-terminus of the recognition element recognizes and binds the interrogated site; and (ii) a single base extension reaction that incorporates the nucleotide corresponding to the target site of interest occurs.

[0104] In some embodiments, a plurality of recognition elements is provided in assays disclosed herein. Each recognition element in the plurality of recognition elements comprises target recognition regions. The target recognition regions of the recognition elements comprises one or more nucleic acid sequence(s) complementary to a target molecule of interest. The complementary regions of a recognition element may hybridize to its complementary nucleic acid sequences of the target molecule. In some embodiments, the target recognition region is configured to hybridize to a methylated, non-cleaved target of interest. In some embodiments,the target recognition region is configured to bind to a variant sequence of interest. Tn some embodiments, the recognition element comprises a hypercode. The hypercode may be detected as a surrogate or proxy for the presence of the target molecule. A non-liming example of a recognition element comprising target recognition regions and a hypercode is depicted in FIG. 1.

[0105] In this non-limiting depiction, the recognition element comprises two target recognition regions (e.g., one at the 5’ end and another on the 3’ end of the recognition element) and a hypercode comprising four segments (e.g., as an example).

[0106] In some embodiments, the structure of the recognition elements may vary. In some embodiments, the structure of the recognition element may configure into a specific structure when hybridized to a target nucleic acid. Non-limiting examples of a recognition element configuration may include a padlock probe, a molecular inversion probe, a hairpin oligonucleotide, a single- stranded oligonucleotide, a double- stranded oligonucleotide, or a combination thereof. In some embodiments, the recognition element is linear. In some embodiments, the linear recognition element is circularized and ligated during the ligation event once hybridized to the respective target molecule. In some embodiments, the recognition element is circular prior to the ligation event. In one embodiment, the target molecule may serve as a primer binding site for an amplification reaction (e.g., rolling circle amplification, multiple strand displacement amplification, etc.). In another embodiment, a segment of the hypercode may serve as a primer binding site for an amplification reaction. In still another embodiment, a universal sequence that is included in the recognition element may serve as a primer binding site for amplification.

[0107] In some embodiments, the recognition element is configured to be a padlock probe once the recognition element is hybridized to the target molecule of interest sequences. Padlock probes may be referred to as linear oligonucleotides whose ends are complementary to adjacent target sequences. Upon hybridization to a target molecule, the two ends (e.g. 5’ end and 3’ end) of the recognition element may be brought together adjacently, generating a padlock probe configuration for subsequent ligation.

[0108] Each recognition element provided comprises target recognition regions. The target recognition regions of the recognition element are configured to hybridize to a complementary target molecule sequence. As disclosed herein, the target recognition regions of the recognition elements used in methods of the present application comprise regions complementary tomethylated targets of interest or variant targets of interest. For example, as depicted in FIG. 1, the recognition element may comprise two target recognition regions. One target recognition region may be present at the 5’ end of the recognition element, and another target recognition region may be present at the 3’ end of the recognition element. Alternatively, the target recognition region may be interposed between the 5’ end and the 3’ end of the recognition element.

[0109] The target recognition regions are selected based on thermodynamic predictions for hybridization to the intended target sequences of interest. As such, each target recognition region of the recognition element may comprise a plurality of nucleotides. In some embodiments, each target recognition region may comprise a length of 2 or more nucleotides, 3 or more nucleotides, 4 or more nucleotides. 5 or more nucleotides, 10 or more nucleotides, 15 or more nucleotides, 20 or more nucleotides, 25 or more nucleotides, 30 or more nucleotides, 35 or more nucleotides, 40 or more nucleotides, 45 or more nucleotides, or 50 or more nucleotides. In some embodiments, each target recognition region comprises a length of 50 or less nucleotides, 45 or less nucleotides, 40 or less nucleotides, 35 or less nucleotides, 30 or less nucleotides, 25 or less nucleotides, 20 or less nucleotides, 15 or less nucleotides, 5 or less nucleotides, 4 or less nucleotides, 3 or less nucleotides, or 2 or less nucleotides. FIG. 7 and FIG. 8 include examples of 5’ end and 3’ end sequences of recognition elements that are complementary to either methylation target sequences of interest (FIG. 7) or variant target sequences of interest (FIG.8).

[0110] After both the 5’ target sequence end and a 3’ target sequence end for a recognition element are determined a hypercode unique to those sequences, and hence the target of interest, is assigned to the recognition element.

[0111] The methods described herein may include amplification of a nucleic acid. In some embodiments, the nucleic acid is a recognition element. In some embodiments, the nucleic acid is a target nucleic acid molecule. In some embodiments, the nucleic acid is a combination of a recognition element and a target nucleic acid molecule or complements thereof. In some embodiments, the amplification is selective amplification. For example, in the assay workflow for identifying methylation targets as illustrated in FIG. 5, after digestion of the extracted DNA with a methylation resistant restriction endonuclease, there is an optional PCR step.

[0112] Amplification of the methylated target of interest can help boost the interaction between the methylated target of interest and its targeted recognition element by providing more targets forhybridization. This can be useful if the methylated target of interest is contemplated to not be highly represented in the sample, for example.

[0113] In some embodiments, amplification occurs if a primer is used that is complementary to one or more of a portion of a target recognition region, a portion of a segment of a hypercode, or another sequence in the recognition element that is complementary to a primer used for amplification. In some embodiments, the amplification is non-selective. For example, in some embodiments randomers can be used to prime amplification from one or more recognition elements. In another example, a universal primer that is universal for all recognition elements, or a subset of recognition elements, can be used to prime amplification from a plurality of recognition elements.

[0114] In some embodiments, the methods described herein may include selectively amplifying a subset of nucleic acids. For example, in some embodiments, a subset of a plurality of recognition elements that were hybridized to a plurality of target nucleic acid molecules and ligated may be amplified. The subset may comprise a percentage of the total amount of recognition elements. In some embodiments, the subset may include 5% or more, 10% or more, 15% or more, 20% or more, 25% or more, 30% or more, 35% or more, 40% or more, 45% or more, 50% or more, 55% or more, 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, or 95% or more of the total amount of recognition elements bound to target nucleic acid molecules. In some embodiments, the subset may include 95% or less, 90% or less, 85% or less, 80% or less, 75% or less, 70% or less, 65% or less, 60% or less, 55% or less, 50% or less, 45% or less, 40% or less, 35% or less, 30% or less, 25% or less, 20% or less, 15% or less, 10% or less, or 5% or less of the total amount of recognition elements bound to target nucleic acid molecules.

[0115] Rolling Circle Amplification (RCA)

[0116] In some embodiments, the amplification may include rolling circle amplification (RCA). In some embodiments, the amplification may include multiple strand displacement amplification. In some embodiments, RCA may generate a concatemer as an amplification product, wherein the concatemer comprises multiple copies of the circularized ligated recognition element, including associated hypercodes, target recognition regions, and any other sequences that are included in the circular ligated recognition element. In some embodiments, RCA may be performed while the circularized and ligated recognition element is in solution. In some embodiments, RCA maybe performed on a circularized and ligated recognition element while the circularized and ligated recognition element is immobilized, either reversibly or non-reversibly, on a solid substrate or surface. In some embodiments, RCA is performed on a circularized and ligated recognition element that is still hybridized to the target molecule. The terms “solid substrate” and “solid surface” may be referred to herein as a surface or substrate. In some embodiments, the substrate is a bead, a flow cell, a microwell, or a nanowell. In some embodiments, the substrate is coated with a composition that enhances target molecule immobilization. In some embodiments, the substrate is charged. In some embodiments, the substrate is positively charged or negatively charged. In some embodiments, the substrate is an anionic substrate. In some embodiments, the substrate is a cationic substrate. In some embodiments, the substrate comprises an immobilization composition, such as polyacrylamide, branched PEI, linear PEI, poly( -aminoester) and poly(amidoamine), PEG, a gel, poly-L-lysine, silane, agarose, muscle mimetic catecholamine polymer, and the like. In some embodiments, the substrate has no charge. FIG. 6 shows a schematic illustrating RCA amplification of a recognition element to yield a concatemeric amplification product. Rolling circle amplification using primer 616b that is complementary to recognition element sequence 616 and is hybridized to circular and ligated recognition element 625 and used to initiate the RCA reaction to generate an amplification product 630. Amplification product 630 is a polymeric concatemeric molecule that includes multiple repeated copies of circular and ligated recognition element 625, wherein each copy includes primer 616, hypercode 614, a functional sequence 612, target recognition regions, and a second functional sequence 618. In this example, the complement of modified recognition element 625 is indicated by the dashed line. In some embodiments, one or more functional sequences which may be included in a recognition element include, but are not limited to, a unique molecular identifier (UMI) sequence, a sequencing primer sequence, an index sequence, a restriction endonuclease sequence, a cleavage sequence, a unique molecular identifier, or combinations thereof. An RCA reaction may be performed in the presence of the cationic polymer coated surface, resulting in simultaneous immobilization and amplification of an amplification product. RCA primers may be supplied in solution or bound to the cationic polymer-coated surface prior to, or concurrent with, performing the RCA reaction.Alternative Amplification Methods

[0117] In some embodiments, amplification may include on-surface polymerase chain reaction (PCR), isothermal amplification, RCA, or a combination thereof. In some embodiments, amplification may include polymerase chain reaction (PCR). In some embodiments, PCR is multiplexed PCR. The amplification methods disclosed herein may include isothermal amplification. Non-limiting examples of isothermal amplification include Nicking endonuclease amplification reaction (NEAR), Transcription mediated amplification (TMA), Loop-mediated isothermal amplification (LAMP), Helicase-dependent amplification (HD A), Nucleic Acid Sequence Based Amplification (NASBA), Strand displacement amplification (SDA), Multiple Displacement Amplification (MDA), Rolling Circle Amplification (RCA), bridge amplification, or Ramification (RAM) amplification method. In some embodiments, the amplification method is provided in Lakruddin M, Mannan KS, Chowdhury A, Mazumdar RM, Hossain MN, Islam S, Chowdhury MA. Nucleic acid amplification: Alternative methods of polymerase chain reaction. J Pharm Bioallied Sci. 2013 Oct;5(4):245-52, which is hereby incorporated by reference in its entirety.

[0118] Codespaces and Hypercodes

[0119] The methods described herein relate to the use of a hypercode, also known as a code, as part of a recognition element. The terms “hypercode” and “code” are used interchangeably herein. In the presence of the complementary sequence, a recognition element hybridizes at the 5’ and 3’ ends to the complementary sequences in a target of interest. The incorporation of a hypercode or code into a recognition element provides a unique way to identify the originally hybridized target sequences of interest to the recognition element, thereby serving as a proxy for the presence of the target of interest in a sample. In some embodiments, the hypercode is selected from a set of hypercodes wherein the set of hypercodes make up a codespace. In some embodiments, the hypercode is associated with one or more target nucleic acid molecules via the complementary 5’ and 3’ ends of the recognition element. In some embodiments, the hypercode comprises a plurality of segments, where each segment corresponds to one or more signals that are used in a decoding process of the present disclosure. In some embodiments, the decoded hypercodes may be used as surrogates or proxies of target molecules, thereby serving as an indirect analysis of the presence of a target molecule from a sample as the hypercodes correlate with the presence of a target molecule that hybridized to a recognition element.In some embodiments, the hypercode is associated with one or more target nucleic acid molecules. In some embodiments, the hypercode is associated with two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or 10 or more target nucleic acid molecules. In some embodiments, the hypercode is associated with 10 or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less target nucleic acid molecules.

[0120] In some embodiments, the hypercode present on the recognition element is selected from a set of hypercodes which comprise a codespace. In some embodiments, each hypercode from the set of hypercodes may be from a predetermined set of hypercodes. In some embodiments, each hypercode from the set of hypercodes may be selected to ensure that the selected hypercode differs from other hypercodes in the set of hypercodes. As such, in some embodiments, several selection criteria may be implemented to generate a set of hypercodes. In some embodiments, selection of the hypercodes may comprise a Hamming distance.

[0121] In some embodiments, to generate a hypercode selected from a set of hypercodes for use in a recognition element, a Hamming distance (HD) selection criterion may be implemented between any two hypercodes of the set of hypercodes. A Hamming distance between two hypercodes in a set of hypercodes may refer to the number of states that differ between two hypercodes in the set of hypercodes. In essence, the Hamming distance measures the number of changes that would need to be made to a first hypercode to change the string of nucleotides to the second hypercode. As such, the hypercodes cannot have a Hamming distance greater than the length of the hypercode. For example, if the length of a hypercode being the number of cycles or flows of decoding runs being eight, and if each cycle or flow corresponds to one state, therefore eight states, then the maximum Hamming distance is eight. In some embodiments, the Hamming distance may be a minimum Hamming distance. In some embodiments, the Hamming distance may be a maximum Hamming distance. In some embodiments, a minimum Hamming distance may be from about 2 to about 10. In some embodiments, the Hamming distance is from about 2 to about 7. In some embodiments, the Hamming distance is from about 3 to about 5. In some embodiments, the Hamming distance is at least 3. In some embodiments, the Hamming distance is 3.

[0122] As the HD increases the number of hypercodes that can be used decreases, whereas as the number of cycles that a hypercode is exposed to increases the number of hypercodes that can beused increases. The hypercode may have a certain length in nucleotides. Tn some embodiments, the hypercode has a length of greater than or equal to about three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18. 19. 20. 21. 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 contiguous nucleotides. In some embodiments, the hypercode has a length of fewer than or equal to about 200, 190, 180, 170, 160, 150, 140, 130, 120. 110, 100, 90, 80, 70, 60, 50, 40, 30, 20. or 10 contiguous nucleotides. In some embodiments, the length is about 5 to 200, 10 to 150, 15 to 100, 20 to 90, or 30 to 80 contiguous nucleotides. The hypercode can be the combination of two or more nucleic acid segments, wherein their combination comprises a hypercode. The flexibility of using recognition elements with hypercodes allows for assay optimization on various levels. For example, the number of detection events can be reduced, thereby lowering costs and runtimes for low-plexity assays (e.g„ quantifying a small number of targets of interest over a maximal dynamic range), or increasing the number of detection events when higher plexity assays are needed. Further, detection can be performed with fewer fluorescent moieties and more queries to maximize signal-to-noise and minimize error rates, or conversely with more fluorescent moieties to further reduce data output time. The present disclosure demonstrates the scalability of the disclosed methods where over 10,000 hypercodes and measurements of absolute analyte concentration spanning up to at least a 10 fold dynamic range with sensitivities as low as 1 fM. Each code or hypercode used in a recognition element is generated based on a codespace design. A “codespace” is a collection of “codewords” that can be used to uniquely identify a recognition element in an assay and is based on colors and / or their combination. Each codeword comprises a string of states drawn from a collection of possible states, for example a four-color state codespace is encoded using a number of different fluorescent moieties, in this example four different fluorescent moieties, which emit at different wavelengths.

[0123] Given 5 states or colors (corresponding to fluorescence states or different emission spectra) and F readout flows or detection cycles (e.g., queries), there can be sFpossible codewords or different possible color state combinations. The present disclosure reports data corresponding to =4 and F=8, so 48potential codewords. However, the number of flows or detection events or cycles can be decreased (for faster data output) or increased (for higher target plexity and / or higher minimum Hamming distance between codewords). A strength of decoding by hybridization as described herein is that the number of states or colors 5 can be increased byusing more fluorescent moieties for detection events, combinations of fluorescent moieties, or fluorescence levels, for example where each amplification product is detected with only one of several colors that maximizes ease of decoding with high signal-to-noise ratio (SNR). As such, the number of distinctive hypercodes can scale with the number of cycles and resolvable optical signatures at each cycle thereby expanding the potential assay complexity.

[0124] The list of possible codewords can be filtered using heuristics derived from data output. For example, low-complexity codewords that include long runs of a single color symbol can be excluded, out of concern that such codewords may be more vulnerable to being misread through biochemistry or optical artifacts. For example, a single color symbol could be given the number 1, and the fluorescent moiety could be FITC, such that using the same color multiple times in a row for detection events would result in a symbol profile of 1111111. Such a long run of the same color detection could be misread by an instrument thereby misidentifying a target of interest.

[0125] Given the collection of codewords of different fluorescent profiles that are unique, a codespace can be generated, wherein the codespace comprises a collection of codewords with a minimum pairwise distance threshold. Selecting a maximum-size codespace (enabling the largest possible assay plexity) can be challenging since, for example, enumerating all codeword combinations can become computationally intractable for nontrivial cases. As such, one strategy is to heuristically generate multiple codespaces and select the candidate codespace with the largest number of codewords. One strategy in designing a codespace is to begin with an empty list and an available list of all valid codewords. One codeword can be selected at random from the list of available codewords and added to the empty list. Any candidates whose Hamming distance from the chosen codeword is smaller than a chosen cutoff can be removed from the list of available codewords. This selection method strategy can be repeated until no more codewords can be added to the previously empty list. As such, the codespace generation process can be repeated many times to generate many potential codeword candidates.

[0126] In an alternative strategy, codeword sets can be generated by deliberately choosing “snug codewords” which are as close as possible to the codewords already in the set, without violating the Hamming distance cutoff. The alternative strategy is similar to the random construction strategy; however the difference is in how the next available codeword is selected. For example, for each available codeword, the number of already- selected codewords at the minimum-allowable Hamming distance is tracked. The next codeword is chosen at random from the subset of available codewords with the largest number of Nearest Neighbors. A snug codeword based codespace design strategy can yield higher codespace size (e.g., 1,166 valid codewords for a four-color state eight-flow data output) compared to the random selection strategy (e.g., 966 valid codewords).

[0127] The implementation of the codespace generation strategies described herein evaluated 10,000 candidate four-flow / eight- symbol codespaces in 3.7 hours on a 3.6 GHz desktop PC, with the maximum codespace size obtained by the 4721stiteration after 1.75 hours. Runtime increases for still larger symbol or flow counts but can readily be accelerated by using additional threads. The candidate codewords of a codespace can be used and can serve as possible inclusions into a hypercode used in a recognition element.

[0128] Once the codespace is defined, sequences for each hypercode can be applied to the codewords in the codespace for generating unique hypercodes for each recognition element. Each hypercode comprises a unique nucleotide sequence, comprising a number of short distinct segments (e.g., nucleic acid segments), which are unique in their locational position in a recognition element and which correlate with the 5’ and 3’ ends of the recognition element, which are in turn specific for a target of interest in a biological sample. As such, a hypercode can be used as an indirect surrogate or proxy for the presence of a target of interest, or the absence thereof. Collectively, all of the possible hypercode sequences that could be incorporated into recognition elements is referred to as the hypercode space. The design of a highly sensitive and specific hypercode space comprises careful selection of nucleic acid segments with favorable biochemical properties. A nucleic acid segment, or portion thereof, of a hypercode hybridizes specifically to an anchor oligonucleotide that in turn hybridizes to a detection oligonucleotide (thereby generating a detection polynucleotide complex), while exhibiting low affinity for hybridizing to other anchor oligonucleotide sequences that are used to detect other nucleic acid segments of a hypercode.

[0129] Given p segment positions each of nucleotide length Lin a recognition element, there can be 4pLpossible hypercodes (e.g., if there are four positions in the hypercode to be filled). In some embodiments, there can be fewer or more positions to be filled in a hypercode, as desired for plexity. From this set of 4pLpossible hypercodes, a subset of sequences with favorable biochemical properties can be selected. As a first step, segments that are anticipated to bevulnerable to readout failure due to mis-hybridization are removed from the subset. For example, nucleic acid elements with repeated nucleotide sequences such as AAAAAAAA are excluded because failure to dehybridize a detection polynucleotide could cause some other sequence, such as AAAAACGT, to be easily misread as the original sequence. As a second step, to minimize off-target hybridization, the number of nucleic acid segments in a set can be further reduced by removing nucleic acid segments whose complements have a high predicted melting temperature (Tm) when hybridized to other nucleic acid segments used in the hypercode set.

[0130] After reduction for biochemical suitability, the largest possible set of nucleic acid segments S„ such that the minimum Levenshtein distance (e.g„ number of nucleotide positions that are different) between any pair of segments is larger than a determined cutoff value is chosen. From the resulting subset of nucleic acid segments, the set of hypercodes can be built, wherein one of N segments is chosen for each segment location in a hypercode. A nucleic acid segment for a given position is determined using the flows, or detection events, that hybridize to that position. For example, given F total number of flows and F / p flows per position, N=s‘'vl / ,>nucleic acid segments at each position can be determined. The number of potential combinations constructed in this way is c = Np, for example, with p = 4, and N= 44= 16, a set of 65,536 candidate hypercodes is possible, from which all the hypercodes or codes, whose corresponding codewords fall within a designed codespace as previously described, can be selected.

[0131] In some embodiments, the hypercodes comprise one or more segments. For example, as shown in FIG. 1, the hypercode comprises four segments.

[0132] In some embodiments, the recognition elements provided herein comprise a hypercode comprising one or more segments. The one or more segments, or the complements thereof, within the hypercode may be used as a proxy for detection of the target nucleic acid molecules recognized by the recognition element.

[0133] The number of segments present in a hypercode of a recognition element may be considered in the design of the recognition element. The number of segments in a hypercode of a recognition element helps to determine the nucleotide length of the recognition element. For example, a recognition element that includes a hypercode comprising five segments may comprise a greater nucleotide length than a recognition element that includes a hypercode of only two segments. A recognition element with a larger nucleotide length may run up against synthesis limits and can lead to a greater risk of synthesis errors. Alternatively, a recognitionelement with a smaller nucleotide length may avoid synthesis limits and risks in synthesis errors. A recognition element with a larger nucleotide length may include less space for other portions of the recognition element, such as the target recognition 5’ and / or 3’ regions. A recognition element with a smaller nucleotide length may include more space for other portions of the recognition element, such as the target recognition 5’ and / or 3’ regions or additional functional sequences.

[0134] In some embodiments, the hypercode comprises 2 to 10 segments. In some embodiments, the hypercode comprises 2 to 8 segments. In some embodiments, the hypercode comprises 3 to 5 segments. In some embodiments, the hypercode comprises at least 4 segments, at least 5 segments, at least 6 segments, at least 7 segments, at least 8 segments, at least 9 segments, at least 10 segments.

[0135] In some embodiments, each segment may comprise a length in nucleotides. In some embodiments, each segment may comprise a length of 10 to 30 nucleotides. In some embodiments, each segment may comprise a length of 10 to 25 nucleotides. In some embodiments, each segment may comprise a length of 15 to 20 nucleotides. In some embodiments, each segment may comprise a length of 2 or more nucleotides, 4 or more nucleotides, 6 or more nucleotides, 8 or more nucleotides, 10 or more nucleotides, 12 or more nucleotides, 14 or more nucleotides, 16 or more nucleotides, 18 or more nucleotides, 20 or more nucleotides, 22 or more nucleotides. In some embodiments, the segments in a hypercode are of the same length. In some embodiments, the segments in a hypercode are not the same length.

[0136] As with a hypercode in a recognition element being compared to another hypercode in another recognition element, a Hamming distance selection criterion may be implemented between any two segments of a hypercode. A Hamming distance between two segments in a hypercode refers to the number of nucleotides that differ between the segments. In essence, the Hamming distance measures the number of changes that would need to be made to a first segment to change the string of nucleotides to the second segment. In some embodiments, the Hamming distance may be a minimum Hamming distance. In some embodiments, the Hamming distance may be a maximum Hamming distance. In some embodiments, a minimum Hamming distance may be from about 2-20, about 3-19, about 4-18, about 5-17, about 6-16, about 7-15, about 8-14, about 9-13, or about 10-12. In some embodiments, a minimum Hamming distance may be greater than or equal to about 2, greater than or equal to about 3, greater than or equal toabout 4, greater than or equal to about 5, greater than or equal to about 6, greater than or equal to about 7, greater than or equal to about 8, greater than or equal to about 9, greater than or equal to about 10, greater than or equal to about 11, greater than or equal to about 12, greater than or equal to about 13, greater than or equal to about 14, greater than or equal to about 15, greater than or equal to about 16, greater than or equal to about 17, greater than or equal to about 18, greater than or equal to about 19. greater than or equal to about 20.

[0137] In some embodiments, a segment may further serve as a primer binding site for amplification. For example, a segment may comprise an amplification primer binding sequence for rolling circle amplification (RCA) for generating a plurality of amplification products. In some embodiments, a plurality of segments is present on the recognition elements provided herein. In some embodiments, the nucleotide or nucleic acid sequence of each segment corresponds to one or more signals for performing a decoding process of the present disclosure. For example, one or more segments of a hypercode may be detected with a first pool of detection polynucleotide complexes to produce one or more detectable binding complexes. In some embodiments, the one or more detectable binding complexes, once imaged, produce one or more optical signals. When all or substantially all segments of the hypercode are detected by iteratively applying additional pools of detectable polynucleotide complexes to the amplification products, a series of optical signals may be observed (e.g., a hypercode profile). The application of detection oligonucleotides to amplification products for detection is called “flow” or “cycle”, wherein “flow” or “cycle” is the number of times a particular segment of an amplification product is queried, or the number of times detection oligonucleotides or detection polynucleotide complexes are flowed or cycled over an amplification product in order to detect a segment sequence. In some embodiments, one or more optical signals observed from querying an amplification product with detection polynucleotide complexes signals may be used to decode a hypercode. In some embodiments, the optical signal may be a color or a non-color. In some embodiments, the optical signal may be a combination of colors (e.g., when the detection polynucleotide complex comprises a plurality of detectable labels).

[0138] In some embodiments, the signals can be depicted as numbers (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.). In some embodiments, each detection polynucleotide complex that comprises a detectable label comprises a signal, as such when four different fluorescent moieties are used as detectable labels there are four signals, 1 to 4, each corresponding to the emitted lightwavelength of the detectable label. However, the number of signals can be larger depending on the combination of detectable labels with each unique detection polynucleotide complex. For example, there can be 16 signals used for decoding if all 16 unique detection polynucleotide complexes are used to detect a corresponding amplification product. However, additional ways to increase the number of signals for decoding include, but are not limited to, adding levels of identifiability associated with a particular detectable signal such as whether a detectable signal is brighter or dimmer compared to a normal level of signal, whether there is a combination of detectable colors that is used to identify a particular nucleotide. As such, the number of signals that could be used is only limited by practicality for any given assay.

[0139] In some embodiments, the methods described herein may use a number of signals. The number of signals used in the methods and systems described herein may be considered in the design of the recognition elements. For example, in some embodiments, a detection scheme using a larger number of signals may lead to a larger codespace, which may allow for a greater amount of information that may be detected. In some embodiments, a detection scheme using a smaller number of signals may be limited in the amount of information that can be detected. In some embodiments, using a larger number of signals may result in a faster detection process (less time to determine a target molecule compared to using a smaller number of signals). In some embodiments, a detection scheme using a larger number of signals may require greater instrument complexity, which may lead to potential drawbacks such as color crosstalk, wherein the signals used in the detection scheme may become difficult to distinguish from other signals. In some embodiments, a greater number of signals may require that a more complex detection tool be used.

[0140] In some embodiments, three or more signals may be used in the methods described herein. In some embodiments, the methods described herein may use one or more signals, five or more signals. 10 or more signals. 15 or more signals. 20 or more signals, 25 or more signals, 30 or more signals, 35 or more signals, 40 or more signals, 45 or more signals, or 50 or more signals. In some embodiments, the methods described herein may use 50 or less signals. 45 or less signals, 40 or less signals, 35 or less signals, 30 or less signals, 25 or less signals, 20 or less signals, 15 or less signals, 10 or less signals, or five or less signals.

[0141] In some embodiments, each segment of a hypercode may correspond to a combination of signals. In some embodiments, each segment may correspond to one or more signals, two ormore signals, three or more signals, four or more signals, five or more signals, six or more signals, seven or more signals, eight or more signals, nine or more signals, or 10 or more signals. In some embodiments, each segment may correspond to 10 or less signals, nine or less signals, eight or less signals, seven or less signals, six or less signals, five or less signals, four or less signals, three or less signals, or two or less signals. It is the combinations of detected signals that are used to build a hypercode profile which can be decoded for identifying the presence of a target molecule.

[0142] Detection Polynucleotide Complexes

[0143] The methods described herein may include introducing detection polynucleotide complexes or a detection oligonucleotide to the circularized, ligated and amplified recognition elements. The methods described herein may include introducing a plurality of detection polynucleotide complexes to the plurality of concatemeric amplification products. Each detection polynucleotide complex may comprise a detection oligonucleotide and an anchor oligonucleotide.

[0144] In some embodiments, the methods described herein may include introducing one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, 10 or more, 15 or more, 20 or more, 25 or more, 50 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1,000 or more detection polynucleotide complexes to an amplification product. In some embodiments, the methods described herein may include introducing 1,000 or less. 900 or less, 800 or less, 700 or less, 600 or less, 500 or less, 400 or less, 300 or less, 200 or less, 100 or less, 50 or less, 25 or less, 20 or less, 15 or less, 10 or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less detection polynucleotide complexes to an amplification product.

[0145] Detection Oligonucleotide

[0146] In some embodiments, each detection polynucleotide complex may comprise a detection oligonucleotide. In some embodiments, the detection oligonucleotide may comprise a portion comprising a detectable label (e.g., fluorescent molecule) and another portion configured to bind to at least a portion of an anchor oligonucleotide. In some embodiments, the portion configured to bind to the anchor oligonucleotide comprises a nucleic acid sequence complementary to a portion of the nucleic acid sequence of the anchor oligonucleotide. FIG. 3A shows a nonlimiting example of a structure of a detection oligonucleotide comprising a fluorescent molecule 340 and a portion complementary to at least a portion of an anchor oligonucleotide 330.

[0147] The detection oligonucleotide may comprise various nucleotide lengths. In some embodiments, the detection oligonucleotide may comprise a length of 5 to 25 nucleotides. In some embodiments, the detection oligonucleotide may comprise a length of 5 to 20 nucleotides. In some embodiments, the detection oligonucleotide may comprise a length of 5 to 15 nucleotides. In some embodiments, the detection oligonucleotide may comprise a length of 5 to 10 nucleotides, hi some embodiments, the detection oligonucleotide may comprise a length of 5 to 8 nucleotides.

[0148] In some embodiments, the detection oligonucleotide may comprise a length of between about 5-100 nucleotides, between about 10-80 nucleotides, between about 20-60 nucleotides, between about 30-50 nucleotides, and between about 15-30 nucleotides. In some embodiments, the detection oligonucleotide may comprise one or more detectable labels. In some embodiments, one or more detectable labels comprise a fluorescent moiety. The fluorescent moiety may emit in the red, far-red, near-red, yellow, green, blue, or ultraviolet wavelengths. In some embodiments, the fluorescent moiety comprises one or more of 6-FAM (6-carboxyfluorescein), JOE (6-carboxy-4',5'-dichloro-2',7'-dimethoxyfluorescein), TAMRA (6-carboxytetramethylrhodamine), 5-Cy5 (5-carboxyrhodamine), 5-Cy5.5 (5-carboxylic acid succinimidyl ester), 5-Cy7 (5-carboxyrhodamine), (hexachlorofluorescein), Alexa Fluor 488 (AF488), Alexa Fluor 514 (AF514), Texas Red, Cyanine 3, Cyanine 5, Pacific Blue, Tetramethyl rhodamine, Oxazole Yellow, Atto647N, and Rhodamine 6G (R6G). In some embodiments, the detection oligonucleotide may comprise one or more fluorescent moieties. In some embodiments, the detection oligonucleotide may comprise two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or 10 or more fluorescent moieties. In some embodiments, the detection oligonucleotide may comprise 10 or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less fluorescent moieties.

[0149] In some embodiments, the fluorescent moiety may comprise an organic dye, a biological fluorophore, a quantum dot, or a combination thereof. In some embodiments, the organic dye may comprise an organic molecule. In some embodiments, the organic dye may comprise a coumarin, a cyanine, a benzofuran, a quinoline, a quinazolinone, an indole, a benzazole, aborapolyazaindacene, a xanthene, or a combination thereof. The organic dye may correspond to a color. For example, the organic dye may correspond to a green, a yellow, a blue, an indigo, a red, an orange a purple, a pink, a violet, or a combination thereof. In some embodiments, the organic dye may correspond to no color. In some embodiments, the organic dye may correspond to a black color. In some embodiments, the organic dye may correspond to a white color.

[0150] In some embodiments, during imaging, the fluorophore may emit a color in the visible light spectrum. In some embodiments, the fluorophore may emit in a wavelength in the range between 400 nm and 900 nm. In some embodiments, the fluorophore may emit in a wavelength between 400 nm and 475 nm, 475 nm and 490 nm, 490 nm and 530 nm, 530 nm and 575 nm, 575 nm and 600 nm, 600 nm and 700 nm, or 700 nm and 800 nm. In some embodiments, the fluorophore may emit a wavelength of 400 nm or more, 425 nm or more, 450 nm or more, 475 nm or more, 500 nm or more, 525 nm or more, 550 nm or more, 575 nm or more, 600 nm or more, 625 nm or more, 650 nm or more, 675 nm or more, 700 nm or more, 725 nm or more, 750 nm or more, 775 nm or more. 800 nm or more, 825 nm or more. 850 nm or more, 875 nm or more, or 900 nm or more. In some embodiments, the fluorophore may emit a wavelength of 900 nm or less, 875 nm or less, 850 nm or less, 825 nm or less, 800 nm or less, 775 nm or less, 750 nm or less, 725 nm or less, 700 nm or less, 675 nm or less, 650 nm or less, 625 nm or less, 600 nm or less, 575 nm or less, 550 nm or less, 525 nm or less, 500 nm or less, 475 nm or less, 450 nm or less, 425 nm or less, or 400 nm or less.

[0151] The wavelength of light that the fluorophore emits may correspond to a color on the visible spectrum. Examples of colors include, but are not limited to, green, blue, red, yellow, orange, pink, purple, or a combination thereof. For example, in some embodiments, a fluorophore emitting light in a wavelength between 400 nm and 475 nm may produce a purple color. In some embodiments, a fluorophore emitting in a wavelength between 420 nm and 530 nm may produce a blue color. In some embodiments, a fluorophore emitting in a wavelength between 490 nm and 575 nm may produce a green color. In some embodiments, a fluorophore emitting in a wavelength between 530 nm and 600 nm may produce a yellow color. In some embodiments, a fluorophore emitting in a wavelength between 575 nm and 750 nm may produce an orange color. In some embodiments, fluorophore emitting in a wavelength between 600 nm and 800 nm may emit a red color.The methods described herein may use a number of detectable labels (e.g., fluorescent moieties). The detectable labels (e.g., fluorescent moieties) may be optically distinct. The number of optically distinct detectable labels used in the methods described herein may impact the amount of information that may be detected. For example, a detection scheme using a larger number of optically distinct detectable labels may allow for multiplexing of hypercodes, which may allow for a greater amount of target molecule related information to be detected. In some embodiments, a detection scheme using a smaller number of optically distinct detectable labels may be limited in the amount of information that may be detected from an amplification product. In some embodiments, using a larger number of optically distinct detectable labels may lead to a detection process that identifies a target molecule in less time as compared to using a fewer number of optically distinct detectable labels when querying an amplification product. In some embodiments, a detection scheme using a larger number of optically distinct detectable labels may lead to greater instrument complexity, which may lead to fluorescence detection crosstalk, whereby the fluorescence emission spectra of the optically distinct fluorescent moieties may not yield distinct fluorescence signals. In some embodiments, a detection scheme using a greater number of optically distinct fluorescent moieties may require use of a more complex detection tool.

[0152] Anchor Oligonucleotide

[0153] In some embodiments, each detection polynucleotide complex may comprise an anchor oligonucleotide. In some embodiments, each anchor oligonucleotide comprises a portion that is complementary to at least a portion of a modified recognition element or an amplification product thereof, and another portion that is complementary to at least a portion of a detection oligonucleotide. FIG.3A shows a non-limiting example of the structure of an anchor oligonucleotide comprising a portion 310 that may be complementary to a detection oligonucleotide 330 and a portion 320 that may be complementary to an amplification product 360 (FIG. 3B). In some embodiments, the portion of the amplification product that the anchor oligonucleotide may be complementary to is a segment of a hypercode. For example, the anchor oligonucleotide may be complementary to a segment 360 of a hypercode on an amplified recognition element.

[0154] In some embodiments, the anchor oligonucleotide 350 may comprise various nucleotide lengths. In some embodiments, the anchor oligonucleotide may comprise a length of 20 to 100nucleotides. Tn some embodiments, the anchor oligonucleotide may comprise a length of 30 to 70 nucleotides. In some embodiments, the anchor oligonucleotide may comprise a length of 40 to 50 nucleotides. In some embodiments, the anchor oligonucleotide may comprise a length of 10-20 nucleotides.

[0155] In some embodiments, the anchor oligonucleotide may comprise a length of one or more nucleotides, two or more nucleotides, three or more nucleotides, four or more nucleotides, five or more nucleotides, six or more nucleotides, seven or more nucleotides, eight or more nucleotides, nine or more nucleotides, 10 or more nucleotides, 15 or more nucleotides, 20 or more nucleotides, 25 or more nucleotides, 30 or more nucleotides, 35 or more nucleotides, 40 or more nucleotides, 45 or more nucleotides, 50 or more nucleotides, 55 or more nucleotides, 60 or more nucleotides, 65 or more nucleotides, 70 or more nucleotides, 75 or more nucleotides, 80 or more nucleotides, 85 or more nucleotides, 90 or more nucleotides, 95 or more nucleotides, or 100 or more nucleotides. In some embodiments, the anchor oligonucleotide may comprise 100 or less nucleotides, 95 or less nucleotides. 90 or less nucleotides, 85 or less nucleotides. 80 or less nucleotides, 75 or less nucleotides, 70 or less nucleotides, 65 or less nucleotides, 60 or less nucleotides, 55 or less nucleotides, 50 or less nucleotides, 45 or less nucleotides, 40 or less nucleotides, 35 or less nucleotides, 30 or less nucleotides, 25 or less nucleotides, 20 or less nucleotides, 15 or less nucleotides, 10 or less nucleotides, nine or less nucleotides, eight or less nucleotides, seven or less nucleotides, six or less nucleotides, five or less nucleotides, four or less nucleotides, three or less nucleotides, or two or less nucleotides.

[0156] FIG. 3A shows an example of a detection polynucleotide complex 300 comprising a detection oligonucleotide 330 and an anchor oligonucleotide 350. Referring to FIG.3B, the detection oligonucleotide and the anchor oligonucleotide form a detection polynucleotide complex 300. The detection polynucleotide complex, when associated with the target molecule, forms a detectable binding complex that is suitable for detection using the imaging system of the present disclosure. A plurality of detectable binding complexes are formed when a pool of detection oligonucleotides and anchor oligonucleotides, or detection polynucleotide complexes, are introduced to a plurality of concatemeric amplification products. Upon formation of detectable binding complexes, a plurality of signals may be observed, wherein one or more signals correspond to a single detectable binding complex, thereby generating a hypercode signal profile for a hypercode that can be decoded.Multiplexing and Throughput

[0157] Detection Pools

[0158] In some embodiments, the methods herein relate to providing detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides to concatemeric amplification products. In some embodiments, the detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides, may be provided in one or more detection pools. In some embodiments, the methods herein may use one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 45 or more, or 50 or more detection pools. In some embodiments, the methods herein may use 50 or less, 45 or less, 40 or less, 35 or less. 30 or less. 25 or less. 20 or less, 15 or less, ten or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less detection pools.

[0159] In some embodiments, each detection pool provided may comprise a number of detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides. In some embodiments, each detection pool may comprise two or more, three or more, four or more, five or more, ten or more, 15 or more, 25 or more, 50 or more, 100 or more, 150 or more, 250 or more, 500 or more, 1,000 or more, 1,500 or more, 2,500 or more, 5,000 or more detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides. In some embodiments, each detection pool may comprise 5,000 or less, 2,500 or less, 1,500 or less, 1,000 or less, 500 or less, 250 or less, 150 or less, 100 or less. 50 or less, 25 or less, 15 or less, ten or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides.

[0160] The number of detection pools and the number of detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides, provided in each detection pool may be considered in the design of the recognition elements. In some embodiments, one advantage to using a smaller number of detection pools in the methods described herein may be to lower costs. In some embodiments, one advantage to using a larger number of detection pools in the methods described herein for detection may be a larger amount of information that may be detected. Forexample, a larger codespace may allow for the simultaneous detection of a larger number of hypercodes and therefore a larger number of target molecules which may be present in a sample.

[0161] To identify a hypercode from a concatemeric amplification product, a segment is queried, for example, twice and sequentially with a unique set of detection polynucleotide complexes, also known as a read out cycle. An array of fluorescent images is generated in each of four color channels, for example for 10 read out cycles (for example, see FIG.2 at 209). Subsequent image analysis and signal processing generates an intensity vector for each amplification product, which can then be matched to the nearest consensus hypercode profile that is associated with the target of interest.

[0162]

[0163] The methods described herein may include imaging a plurality of detectable binding complexes to obtain identifiable signals, such as fluorescent signals. In some embodiments, the signals are associated with the segments of a hypercode, or a portion thereof, for each amplification product. In some embodiments, the imaging is performed by an imaging system such as a fluorescent microscope or fluorescent plate or slide reader.

[0164] In some embodiments, the imaging may be conducted using an imaging system. The imaging system may comprise at the minimum a camera, a detector, an illuminator, a condenser, or a combination thereof. In some embodiments, the imaging may include images of fluorescence emission, luminescence, or a combination thereof. In some embodiments, the imaging systems comprise components or sub-systems of a larger system that may also include fluidics modules, temperature control modules, translation stages, robotic fluid dispensing and / or microplate handling, processors or computers, instrument control software, data analysis and display software, etc. In some embodiments, the imaging system is a fluorescence imaging system. In some embodiments, the imaging may include fluorescent images from the fluorescent moieties present on the detection polynucleotide on the detection oligonucleotide complex.

[0165] In some embodiments, the image may comprise fluorescence information from a variant of wavelengths. In some embodiments, the fluorescence information may comprise emission data from a wavelength from about 220-830nm, about 230-820nm, about 240-8 lOnm, about 250- 800nm, about 260-790nm, about 270-780nm, about 280-770nm, about 290-760nm, about 300- 750nm, about 310-740nm, about 320-730nm, about 330-720nm, about 340-710nm, about 350- 700nm, about 360-690nm, about 370-680nm, about 380-670nm, about 390-660nm, about 400-650nm, about 410-640nm, about 420-630nm, about 430-620nm, about 440-61 Onm, about 450-600nm, about 460-590nm, about 470-580nm, about 480-570nm, about 490-560nm, about 500-550nm. about 510-540nm, about 520-530nm, or a combination thereof. In some embodiments, the fluorescence data may comprise emission data from a wavelength of about 220nm, about 230nm, about 240nm, about 250nm, about 260nm, about 270nm, about 280nm, about 290nm, about 300nm, about 310nm, about 320nm, about 330nm, about 340nm, about 350nm, about 360nm, about 370nm, about 380nm, about 390nm, about 400nm, about 410nm, about 420nm, about 430nm, about 440nm, about 450nm, about 460nm, about 470nm, about 480nm, about 490nm, about 500nm, about 510nm, about 520nm, about 530nm, about 540nm, about 550nm, about 560nm, about 570nm, about 580nm, about 590nm, about 600nm, about 610nm, about 620nm, about 630nm, about 640nm, about 650nm, about 660nm, about 670nm, about 680nm, about 690nm, about 700nm, about 710nm, about 720nm, about 730nm, about 740nm, about 750nm, about 760nm, about 770nm, about 780nm, about 790 nm, about 800 nm, about 810 nm, about 820 nm, about 830 nm, or a combination thereof.

[0166] The imaging of fluorescently labeled detection polynucleotide complexes for detection of circularized and amplified recognition elements can be performed on a fluorescent instrument. The decoding of the images begins with collectively analyzing the images of all the detection events across all of the fluorescent channels for an amplified product. Subsequent analysis steps can involve intensity normalizations, feature identification, intensity extractions, and data conditioning operations. Minimizing optical, electrical, biochemical, thermal and motion- induced noise sources is preferential for optimizing decoding of the fluorescent detection profile of an amplified product. As such, the imaging pipeline can employ subpixel feature detection methods. Subsequently, subpixel registration techniques can be applied to pinpoint the same amplified product in all readout flows and channels. Optical and biochemical imperfections, such as non-linear distortion and non-uniform illumination, can be corrected for in-situ or via prior calibrated correction factors. The extraction of fluorescent intensities can be accomplished by convolving the feature image raw pixels with an appropriately tuned extraction kernel. These computationally intensive processing steps can be implemented using, for example, C++ and Halide software thereby enabling pseudo-real-time processing.

[0167] In some embodiments, following extraction the raw intensity values from the population of amplified recognition elements the data can undergo normalization. For example, mapping allintensities into a standard range to eliminate outliers, for example into the [0.5, 0.995]% range of all intensity values can be performed. Spatial variability in foreground and background intensities resulting from factors such as non-uniform illumination can be corrected for. For example, an image can be divided into a grid of sub-images, and high and low percentiles of intensities in each sub-region can be measured. Fluorescence intensities from the amplified products can be re-scaled such that, after correction, the background and foreground intensities of the various sub-regions are comparable. To minimize edge effects across sub-regions, bilinear interpolation can be applied to the correction factors. In some embodiments, the addition of color-balanced control recognition elements can help ensure that all fluorescent color channels observe some bright objects in each readout flow, regardless of the true application sample plexity and the sample assayed. In some embodiments, color crosstalk caused by the overlap of emission spectra of different fluorophores used in the assay can be corrected.

[0168] In some embodiments, analyzing crosstalk as a linear mixing operation and correcting for it by “unmixing” the observed intensities can be an effective approach for electrophoretic sequencing. The linear models can be measured for a particular combination of fluorescent dye moieties, excitation lasers, and filters used. Once the crosstalk between colors is quantified, it can be corrected by multiplying the received intensities by the inverse of the 4 x 4 matrix of color crosstalk coefficients. Finally, the intensity footprint of each amplified recognition element can be normalized to a unit norm.

[0169] Soft Decision

[0170]

[0171] The methods described herein may relate to using an algorithm to predict the presence of target nucleic acid molecules in a sample based on the hypercode profile data. In some embodiments, the algorithm is a soft decision decoding algorithm. In some embodiments, the algorithm is applied to the detectable signals of the hypercodes, the hypercode profile, for predicting the presence of a target nucleic acid in a sample.

[0172] The methods disclosed herein may comprise soft decision decoding to predict, or determine the probability of, the presence of the hypercode in a recognition element or amplification product thereof, wherein the presence of the hypercode correlates and serves as a proxy for the presence of a target nucleic acid in a sample. In some embodiments, the methods described herein may use soft decision decoding.In some embodiments, the methods described herein may use hard decision decoding. For hard decision decoding, signals from queried concatemers are extracted from images. This is the same for soft decision decoding, in that signals that are generated and imaged are extracted from the images. Conversely, for hard decision decoding, hard basecalls for each nucleotide of a hypercode are generated from the intensities of the signals, whereas with soft decision decoding no hard basecalls are necessary as all of the signal range is retained. The hypercode assignment for hard decision decoding is determined by matching the nucleotide reads to hypercodes, whereas with soft decision decoding, the signals are cross correlated against the expects signals and the most likely hypercode is assigned, as such soft decision decoding is a probabilistic methodology. When using soft decision decoding techniques, it is not necessary for the model to identify each base specifically. For example, signals (e.g., fluorescent signals) generated during each cycle of a detection process may be detected and recorded to produce a data set that may be used as input into a model to calculate a probability that a specific hypercode is present without requiring a nucleotide by nucleotide base call that results from a hard decision decoding model. Although it is not necessary to use a soft decision decoding model to make a hard decision about the identity of each nucleotide, a soft decision decoding model developed according to the methods of the disclosure may nevertheless include assigning a probability or identity to each nucleotide in the sequence of a hypercode.

[0173] FIG. 4 details a soft decision decoding algorithmic pipeline for determining the presence of a hypercode and hence the presence of a target nucleic acid from a sample. Images of the detection polynucleotide complex queried sample are acquired, aligned, and processed to extract the intensity of the features of interest across the imaged field of view in multiple spectral channels. The corrected intensities of said features are then fed through a series of algorithms that make up the soft decision decoder. At first, the intensity profiles of the hypercodes are learned based on features of high confidence or high intensity. This trained model provides a template for each hypercode from which the rest of the features of interest are compared to in the second step. A confidence score is computed from the difference between the intensity profile of each feature and the trained profiles. Several filters are applied to remove outliers, duplicates, and low confidence decoded concatemers. The final output is a table of decoded concatemers with associated filter status, confidence score, and most likely assignment to one of the hypercodes of the codeset used in the recognition elements of the assay.In some embodiments, the methods described herein determine methylation status without chemical conversion of cytosines, including without bisulfite treatment.

[0174] In some embodiments, a recognition element comprises a larger hypercode, for example a hypercode with four segments instead of two or three. In some embodiments, a recognition element comprising a larger hypercode may result in a detection scheme with improved error correction compared to a smaller hypercode. Additionally, in some embodiments, a larger hypercode may result in a lower signal-to-noise ratio.

[0175] In some embodiments, a recognition element comprises a smaller hypercode, for example a hypercode with two segments. In some embodiments, a recognition element comprising a smaller hypercode may result in a detection scheme with lower error correction abilities. Further, in some embodiments, a small hypercode may result in a higher signal-to-noise ratio.

[0176] In some embodiments, soft decision decoding comprises determining the presence of a hypercode based on probabilistic matching of a hypercode profile to expected signal profiles associated with a codespace, without requiring assignment of a definitive nucleotide identity to each position of the hypercode. In such embodiments, rather than performing nucleotide-by-nucleotide basecalling, the decoded signals obtained from iterative detection cycles are treated as a signal pattern that is compared to one or more expected hypercode profiles corresponding to recognition elements used in the assay.

[0177] In some embodiments, the hypercode profile for an amplification product comprises signal intensities obtained across a plurality of detection cycles and one or more detectable labels. The hypercode profile may be represented as a vector, matrix, or other multidimensional data structure that encodes the signal states observed for each segment, or portion thereof, of a hypercode across the detection cycles.

[0178] In some embodiments, soft decision decoding comprises comparing the hypercode profile of an amplification product against a library of expected hypercode profiles associated with the set of hypercodes used in the assay. The comparison may generate a confidence score or probability value for each candidate hypercode, reflecting the degree of similarity between the observed hypercode profile and the expected profile for that hypercode. The hypercode having the highest confidence score or probability may be assigned to the amplification product, thereby identifying the target molecule associated with the recognition element comprising that hypercode.In some embodiments, the confidence score or probability is derived using one or more probabilistic or statistical metrics, including but not limited to likelihood estimation, maximum likelihood estimation, Bayesian inference, distance-based scoring, correlation-based scoring, or combinations thereof. In some embodiments, the confidence score represents a relative rather than an absolute probability.

[0179] In some embodiments, soft decision decoding comprises cross-correlating the hypercode profile of an amplification product with expected hypercode signal profiles. Cross-correlation may be performed using normalized or unnormalized correlation metrics to assess similarity between observed and expected signal patterns across detection cycles. In some embodiments, the resulting correlation values are used directly as confidence scores. In other embodiments, the correlation values are further transformed, normalized, or combined with additional metrics to generate a confidence score or probability distribution across candidate hypercodes.

[0180] In some embodiments, the expected hypercode signal profiles are generated based on the designed detection scheme, including the number of detection cycles, the detection polynucleotide complexes used in each cycle, and the detectable labels associated therewith. In some embodiments, the expected hypercode profiles are learned or refined using amplification products, calibration samples, internal controls, or a combination thereof.

[0181] In some embodiments, the soft decision decoding disclosed herein provides advantages for hypercode identification in multiplexed assays. Because decoding is based on evaluation of a hypercode profile across multiple detection cycles rather than discrete identification at individual nucleotide positions, the decoding process is tolerant to signal variability arising from issues such as optical noise, photobleaching, or non-uniform hybridization.

[0182] In some embodiments, the soft decision decoding disclosed herein supports correct hypercode identification even when one or more detection cycles yield weak, ambiguous, or missing signals, by leveraging information accumulated across the remaining cycles. This property supports robust decoding in multiplex assays in which large numbers of amplification products are analyzed in parallel.

[0183] In some embodiments, the soft decision decoding disclosed herein supports flexibility in hypercode and codespace design, as hypercodes may be optimized for distinguishability at the signal pattern level across detection cycles rather than at individual nucleotide positions, thereby supporting increased error tolerance and scalable multiplexing.Iterative detection

[0184] The methods described herein may relate to iteratively repeating the operations of: (i) introducing detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides; (ii) forming detectable binding complexes; and (iii) imaging the detectable binding complexes. In some embodiments, the iterative repetition of the operations may be performed for each segment of a hypercode, or a portion thereof.

[0185] In some embodiments, the iteratively repeating the operations of: (i) introducing detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides; (ii) forming detectable binding complexes; and (iii) imaging the detectable binding complexes may comprise a number of iterative repetitions. For example, the methods described herein may comprise 2-50 iterative repetitions, 2-10 iterative repetitions, 2-8 iterative repetitions, or four iterative repetitions of the operations of: (i) introducing detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides; (ii) forming detectable binding complexes; and (iii) imaging the detectable binding complexes. In some embodiments, the method described herein may comprise one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, or 50 or more iterative repetitions of : (i) introducing detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides; (ii) forming detectable binding complexes; and (iii) imaging the detectable binding complexes may comprise a number of iterative repetitions. In some embodiments, the method described herein may comprise 50 or less, 45 or less, 40 or less, 35 or less, 30 or less, 25 or less. 20 or less. 15 or less, 10 or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less iterative repetitions of the operations of : (i) introducing detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides; (ii) forming detectable binding complexes; and (iii) imaging the detectable binding complexes may comprise a number of iterative repetitions.

[0186] In some embodiments, the number of iterative repetitions of the operations of: (i) introducing detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides; (ii) forming detectable binding complexes; and (iii) imaging the detectable binding complexes may comprise a number of iterative repetitions that may correspond to thenumber of segments present in a hypercode of the recognition element. As described above, the hypercode of the recognition element may comprise a number of segments.

[0187] In some embodiments, each segment of the hypercode of the recognition element may undergo a number of iterative repetitions of the operations of: (i) introducing detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides; (ii) forming detectable binding complexes: and (iii) imaging the detectable binding complexes may comprise a number of iterative repetitions. For example, the methods described herein may comprise iteratively repeating the operations of: (i) introducing detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides; (ii) forming detectable binding complexes; and (iii) imaging the detectable binding complexes two times per segment, three times per segment, or four times per segment. In some embodiments, the methods described herein may comprise iteratively repeating the operations of: (i) introducing detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides; (ii) forming detectable binding complexes: and (iii) imaging the detectable binding complexes two or more times per segment, three or more times per segment, four or more times per segment, five or more times per segment, six or more times per segment, seven or more times per segment, eight or more times per segment, nine or more times per segment, 10 or more times per segment, 11 or more times per segment, 12 or more times per segment, 13 or more times per segment, 14 or more times per segment, 15 or more times per segment, 16 or more times per segment, 17 or more times per segment, 18 or more times per segment, 19 or more times per segment, or 20 or more times per segment. In some embodiments, the methods described herein may comprise iteratively repeating the operations of: (i) introducing detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides; (ii) forming detectable binding complexes; and (iii) imaging the detectable binding complexes 20 or less times per segment, 19 or less times per segment. 18 or less times per segment, 17 or less times per segment, 16 or less times per segment, 15 or less times per segment, 14 or less times per segment, 13 or less times per segment, 12 or less times per segment, 11 or less times per segment, 10 or less times per segment, nine or less times per segment, eight or less times per segment, seven or less times per segment, six or less times per segment, five or less times per segment, four or less times per segment, three or less times per segment, or two or less times per segment.In some embodiments, a detection scheme comprising a smaller number of iterative repetitions and iterative repetitions per segment may require a greater number of detection pools and detection polynucleotide complexes, or detection oligonucleotides and anchor oligonucleotides per detection pool, which may result in greater costs and resources needed for the detection scheme. In some embodiments, a detection scheme comprising a greater number of iterative repetitions per segment may allow for a greater amount of information that may be detected and for a larger codespace. In some embodiments, a detection scheme comprising a greater number of iterative repetitions and iterative repetitions per segment may result is greater error correction mechanisms.

[0188]

[0189] The methods described herein relate to providing a sample. The sample may be a biological sample. The sample may comprise target nucleic acid molecules. The target nucleic acid molecules may be used to detect a nucleotide sequence. In some embodiments, the sample comprises nucleic acid methylation targets of interest. In some embodiments, the sample comprises nucleic acid variant targets of interest. In some embodiments, a sample for use in the methods disclosed herein comprise both methylation targets of interest and nucleic acid variants of interest.

[0190] In some embodiments, the sample may comprise a biological sample. In some embodiments, the sample may comprise whole blood, lymphatic fluid, serum, plasma, sweat, tears, saliva, sputum, cerebrospinal fluid, amniotic fluid, seminal fluid, vaginal excretion, serous fluid, synovial fluid, pericardial fluid, peritoneal fluid, pleural fluid, transudates, exudates, cystic fluid, bile, urine, gastric fluid, intestinal fluid, fecal samples, liquids containing single or multiple cells, liquids containing organelles, tissues, organisms, liquids containing multi-celled organisms, biological swabs, biological washes, or a combination thereof. In some embodiments, the sample may comprise whole blood. A whole blood may be veinous blood or capillary blood (e.g., obtained by a fingerstick).

[0191] In some embodiments, the sample may be from a subject. In some embodiments, the subject may be a mammal. In some embodiments, the subject may be a human. The subject may be male. The subject may be female. The subject may be an adult. The subject may be a child. The subject may be a vertebrate. In some embodiments, a sample is from a eukaryote or a prokaryote. In some embodiments, the sample is a viral sample, for example a DNA virus or anRNA virus. Tn some embodiments, the sample is a microorganism and could be one or more of a pathogenic microorganism or a bioterrorism related microorganism. In some embodiments, the sample may be from a plant, for example from a crop plant or agricultural species such as corn, soybeans, cotton, wheat, etc. In some embodiments, the sample may be from a non-crop plant species, for example a horticultural and / or landscape plant species. In some embodiments, the sample may be from a livestock animal including, but not limited to, a sheep, cow, chicken, duck, goose, llama, alpaca, emu, or pig.

[0192] In some embodiments, a sample may be from an animal that has been, or is thought to have been, infected with a virus or bacterium. In some embodiments, the methylation status of a sample may be a direct result of an animal being infected with a virus or a bacterium which, once infected, is able to alter the epigenetic profile of an animal. In some embodiments, a sample may be from a mammal, or animal, that has been infected with a parasite, wherein the parasite may alter the epigenetic profile of the animal. Some notable viral diseases that may affect livestock and may alter the animal’s epigenetic signature include, but are not limited to, bovine viral diarrhea, foot and mouth disease, African swine fever, bluetongue, infections bovine rhinotracheitis, avian influenza, lumpy skin disease, peste des petits ruminants, equine influenza, Newcastle disease, classical swine fever, infectious bursal disease, equine infectious anemia, and rabies. Some notable bacterial diseases that may alter an animal’s epigenetic signature include, but are not limited to, bovine respiratory disease, mastitis, Johne’s disease, leptospirosis, anthrax, brucellosis, salmonellosis, campylobacteriosis, fowl cholera, pullorum disease, tuberculosis, clostridial myositis, actinomycosis. Notable parasites, especially in bovine species, includes but is not limited to Haematobia irritans and Rhipicephalus microplus (2022, de Soutello, RVG et al., Sci Rep 12, 18135, incorporated herein by reference in its entirety). The ability to identify a change in an animal’s epigenetic signature due to a viral or bacterial infection can significantly impact animal health and productivity. As such, the methods disclosed herein find utility in disease detection that may impact crop and livestock, or general animal, health and rapid and accurate diagnosis of these diseases is essential for effective control and prevention, helping to reduce economic losses and improve animal welfare.

[0193] In some embodiments, epigenetic signatures may also be affected by bacterial infection and hence detected using the methods described herein. Examples of bacterial infections in humans that may be detected by their epigenetic profiles include, but are not limited to,Mycobacterium tuberculosis, and Helicobacter pylori. Examples of bacterial infections in mammals (including humans) that may be detected by their epigenetic profiles include, but are not limited to, Legionella pneumophila, Burkholderia thailandensis, Chlamydia trachomitis, Chlamydophila pneumoniae, Bacillus anthracis, Mycoplasma hyorhinis, Klebsiella pneumoniae, Salmonella enterica, Methanosarcina maze, Bacillus fragilis, and Streptococcus pneumonia.

[0194] The methods described herein relate to providing one or more samples. In some embodiments, the methods relate to providing one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, or 20 or more samples. In some embodiments, the methods relate to providing 25 or more, 50 or more, 75 or more, 100 or more, 125 or more, 150 or more. 175 or more, 200 or more, 225 or more, 250 or more, 275 or more, or 300 or more samples. In some embodiments, the methods relate to providing 20 or less, 19 or less, 18 or less, 17 or less, 16 or less, 15 or less, 14 or less, 13 or less, 12 or less, 11 or less, 10 or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less samples. In some embodiments, the methods comprise providing 300 or less, 275 or less, 250 or less, 225 or less, 200 or less, 175 or less, 150 or less, 125 or less, 100 or less, 75 or less, 50 or less, or 25 or less samples.

[0195] In some embodiments, the sample that is provided may comprise a plurality of target nucleic acid molecules. In some embodiments, the sample may comprise one or more target nucleic acid molecules, two or more target nucleic acid molecules, three or more target nucleic acid molecules, four or more target nucleic acid molecules, five or more target nucleic acid molecules, six or more target nucleic acid molecules, seven or more target nucleic acid molecules, eight or more target nucleic acid molecules, nine or more target nucleic acid molecules, 10 or more target nucleic acid molecules, 15 or more target nucleic acid molecules, 25 or more target nucleic acid molecules, 50 or more target nucleic acid molecules, 100 or more target nucleic acid molecules, 250 or more target nucleic acid molecules, 500 or more target nucleic acid molecules, 750 or more target nucleic acid molecules, or 1,000 or more target nucleic acid molecules. In some embodiments, the sample may comprise 1,000 or less target nucleic acid molecules, 750 or less target nucleic acid molecules, 500 or less target nucleic acid molecules, 250 or less target nucleic acid molecules, 100 or less target nucleic acid molecules, 50 or less target nucleic acid molecules, 25 or less target nucleic acid molecules, 15 or less targetnucleic acid molecules, 10 or less target nucleic acid molecules, nine or less target nucleic acid molecules, eight or less target nucleic acid molecules, seven or less target nucleic acid molecules, six or less target nucleic acid molecules, five or less target nucleic acid molecules, four or less target nucleic acid molecules, three or less target nucleic acid molecules, or two or less target nucleic acid molecules.

[0196] In some embodiments, the target nucleic acid molecules may include DNA, The target nucleic acid molecules may include RNA. The target nucleic acid molecules may include a combination of DNA and RNA. In some embodiments, the target nucleic acid molecules are fragments or components of DNA. In some embodiments, the target nucleic acid molecules are fragments or components of RNA. In some embodiments, the target nucleic acid molecules comprise complementary DNA or cDNA. In some embodiments, the target nucleic acid molecules comprise mRNA.

[0197] The target nucleic acid molecules may comprise DNA. The DNA may be genomic DNA. The DNA may include one or more single nucleotide variants (SNVs), single nucleotide polymorphisms (SNPs), insertions / deletions (indels), copy number variants (CNVs), methylated nucleotides, or any combination thereof. In some embodiments, the DNA may include cell-free DNA (cfDNA). The cfDNA may include maternal cfDNA, fetal cfDNA, or combinations thereof, or cfDNA from a tumor. In some embodiments, the DNA may include circulating tumor cell DNA or ctcDNA. In some embodiments, the DNA may include a synthetic DNA target, such as a product of a polymerase chain reaction (PCR). In some embodiments, the DNA may be transcribed from single- stranded RNA templates, such as complementary DNA (cDNA) from a first strand or second strand synthesis reaction. The DNA may be PCR-amplified or RT-PCR amplified DNA or complementary DNA (cDNA). In some embodiments, DNA is genomic DNA or fragmented genomic DNA. The target nucleic acid molecules may comprise methylated cytosines. The target nucleic acid molecules may comprise methylated adenines. The target nucleic acid molecules may comprise RNA. The RNA may include messenger RNA (mRNA). The mRNA may be a splice variant. In some embodiments, methylation state specific amplification primers can be utilized that extend a target sequence post-bisulfite conversion or post-TET-assisted pyridine borane conversion. In this scenario, methylation state specific primers could be specific for nucleotides that were either converted or not converted, depending on their methylation status and desired application, wherein the primer extension would occurfollowed by subsequent circularization and ligation of a recognition element as described herein if the desired methylation state was present.

[0198] In some embodiments, maintaining methylation status of a PCR amplification reaction is contemplated. For example, during canonical PCR amplification epigenetic markers are lost. However, utilizing DNA methyltransferases such as DNMT1, DNMT3a and DNMT3b could be utilized after each PCR amplification cycle to copy the methylation status of the template DNA to the new complementary strand.

[0199] In some embodiments, it may be desirable to convert non-methylated cytosines in a sample prior to assaying the sample for methylated nucleotides, for example to gain knowledge on which cytosines are not methylated relative to those that are, thereby generating an epigenetic signature of the sample which may be used to indicate a disease state for a subject, be it a human subject or any animal subject as described herein. In some embodiments, a sample can be treated with bisulfite, which converts non-methylated cytosines to uracils while not affecting methylated cytosines. As such, an assay can be performed following the methods described herein where uracils can be detected in combination with methylated CpG sites. For example, recognition elements can be designed comprising a guanidine at the 5’ end or 3’ end to detect the presence of a cytosine, while another recognition element can be designed with an adenine at the 5’ or 3’ end to detect the presence of a uracil (due to a non-methylated cytosine present in the sample), thereby determining the epigenetic signature or methylation level of the sample.

[0200] However, bisulfite conversion can be harsh and can oftentimes result in unwanted DNA fragmentation. Therefore, in some embodiments, TET-assisted pyridine borane conversion (2021, Siejka-Zielinska et al.. Sci. Adv. 7:eabh0534, incorporated herein by reference in its entirety) can be utilized. Conversely to bisulfite conversion of non-methylated cytosines to uracil, TET-assisted pyridine borane conversion converts methylated cytosines to uracil. For example, TET-assisted pyridine borane conversion oxidizes 5mC or 5hmC by mTetlCD enzyme to 5-carboxycytosine followed by reduction to dihydrouracil, which can be detected in a sample. A recognition element can be designed taking into account that cytosines that are not methylated are assumed to retain their native identity (i.e., they remain cytosines) and a guanidine at the 5’ end or 3’ end of the recognition element can detect the cytosine. Another recognition element can be designed with an adenine at the 5’ end or 3’ end of the recognition element which can detectthe uracil as a result of conversion of the methylated CpG motif. As such, the epigenetic signature or methylation level of the sample can be determined.

[0201] In some embodiments, multiple methylated cytosines may be present in close proximity, or adjacently located, known as a “hot spot” for methylation. In this case, the methylation status of neighboring methylated nucleotides may affect hybridization and subsequent ligation at the methylation site of interest. In some embodiments, a recognition element can be designed that differentiates between one or more methylation sites, which can be used to identify degrees of methylation. FIG.20 and FIG.21 provide scenarios for addressing the identification of one or more methylated sites and how recognition elements can be used to quantify percentage of methylation at a location in a sample. In some embodiments, recognition elements can comprise universal bases such as deoxyinosine which can hybridize with both cytosine and uracil, which may be used to detect neighboring methylation sites, be they adjacently located or in close proximity.

[0202] In some embodiments, the methylation detection workflow as described herein and exemplified in FIG. 5500 uses extracted DNA 501. DNA can be extracted and optionally purified from any eukaryotic or prokaryotic source. In some embodiments, the DNA can be from a mammal, for example a human, a primate, a dog, a cat, and the like. In some embodiments, the DNA can be from a bacteria, such as a gram positive bacteria, a gram negative bacteria, an encapsulated bacteria, a lab generated and possibly enhanced bacteria. A skilled artisan will understand that DNA can be from any source for practicing methods as described herein.

[0203] In some embodiments, the DNA, extracted and optionally purified, can be exposed to a restriction endonuclease 502, whereas the restriction endonuclease is methylation resistant and cleaves only unmethylated targets. The methylation site of interest may comprise either a methylcytosine or a methyladenine, or two or more target sites for a combination of both target sites in a DNA sample. The restriction endonuclease Hhal was utilized in experiments as found in the Examples, however there are other methylation resistant restriction endonucleases that could also be used, including but not limited to, BstUI, Hpall. AccI, MspI, MspJI, Aatll, Acil, AcII, Afel, Agel, Alwl, Asci, AsiSI, Aval, BceAI, BmgBI, BsaAI, BsaHI, BsiEI, BsiWI, BsmBI-v2, BspDI, BsrFI-v2, BssHII, BstBI, Clal, Eagl-HF, Esp3I, Faul, Fsel, FspI, Haell, Hgall, Hhal, HinPlI. HpyCH4IV. Hpy99I, KasI, Mlul, Nael, Narl, NgoNIV, Notl, Nt.BsmAI. Nt.CviPII.PaeR71, PluTi, Pmll, Pvul, SacII, Sall, Sfol, SgrAI, SmaT, SnaBT, Srfl, TspMI, Zral, Bell, DpnII, HphI, Mbol, Nt.AlwI, PspGI, and SexAI.

[0204] When digesting a DNA sample of interest with a methylation resistant restriction endonuclease, if the enzyme recognition site is methylated there is expected to be no cleavage. However, if the enzyme recognition site is not methylated, the restriction endonuclease is expected to cleave the DNA at that recognition site, thereby generating a break in the target DNA sample at that location. While the cleaved target DNA can hybridize to the complementary sequences of a recognition element, there is expected to be minimal to no ligation of the ends of the recognition element due to the cleaved site. As such, the methylated target of interest is enriched over that of the non-methylated target of interest.

[0205] In some embodiments, it may be advantageous to identify other targets of interest in a sample. For example, in addition to identifying methylated target sites of interest it may also be beneficial to identify target variant sequences that might also be present in a sample. Variants could include, but are not limited to, single nucleotide polymorphisms, insertions, deletions, copy number variances, splicing variants and the like. Other targets of interest could also include proteins or polypeptides of interest if the DNA sample is extracted but not purified away from the protein component. As such, the present disclosure provides assays for multiomics detection, in this case methylation in combination with variant detection and / or protein detection targets of interest in combination with methylation targets of interest.

[0206] In some embodiments, a method as disclosed herein could include identification of methylated targets of interest and variants of interest in a sample. In some embodiments, it may be advantageous to identify a variant of interest even if the target methylation site of interest is unmethylated. In some instances, the identification of a variant of interest when the sample is unmethylated and cleaved may not be an issue as the variant site of interest is not in proximity to the cleavage site of the unmethylated target site thereby leaving the variant target sequence untouched and able to hybridize to its target recognition element. However, in other instances a variant could be in such proximity to a methylation site of interest that cleavage by a restriction endonuclease can compromise the hybridization of the variant target recognition element to the variant sequence of interest. FIG. 17 illustrates such an event. In FIG. 17, the top illustration shows a recognition element 1700 that is complementary to the variant sequence of interest 1703 in a target sequence 1701. The recognition element can be complementary to either the sense orantisense strand of the target sequence 1701, the scenario is the same regardless. Tn this scenario, the variant sequence of interest 1703 is in close proximity to the methylation site of interest 1702, such that cleavage of the unmethylated target site of interest compromises one of complementary sequence ends of the recognition element and hybridization of the recognition element 1700 to the variant target sequence 1703 is unable to occur, as such no ligation of the ends of the recognition element 1700, no amplification and no detection of the variant.

[0207] The strategy as illustrated in FIG. 18 offers an alternative workflow such that detection of a variant that is in close proximity to a cleaved sequence is possible. The strategy illustrated in FIG. 18. A recognition element 1800 is complementary to a variant sequence 1803 in a target of interest 1801 that is in close proximity to a methylation site of interest 1802, such that cleavage of an unmethylated site of interest 1802 could compromise the hybridization of the recognition element 1800 to the variant sequence of interest 1803, such that no hybridization, ligation, amplification or detection of the variant would occur. However, addition of an oligonucleotide 1804 that is complementary to the sense or antisense strand of the methylation site of interest, and therefore the restriction endonuclease site of interest, could serve to complete the restriction endonuclease site which can be cleaved by the restriction endonuclease while leaving the opposite strand intact. The uncleaved, intact strand with the variant sequence of interest 1803 can hybridize to the variant recognition element 1800, which can be ligated, amplified and detected regardless of cleavage of an unmethylated target site. In this scenario, the variant can be detected even though the unmethylated target of interest is not detected.

[0208] In some embodiments, the methods disclose herein and as exemplified in FIG. 5 comprise an optional amplification 503 of the methylated target of interest, for example after restriction endonuclease digestion. An optional amplification step can be, for example, PCR where the methylated target of interest is exponentially amplified using a thermostable DNA polymerase thereby increasing the incidence of the target of interest in the assay over that of the unmethylated target site (e.g., which is not expected to be detected). It is contemplated that PCR would amplify the methylated, non-cleaved sequence and not amplify the cleaved, unmethylated target sequence, thereby generating an enriched pool of methylated targets for hybridizing to target recognition elements in an assay. By way of example, for low allelic frequency targets amplification could be used to increase the incidence of the target for hybridizing to itsrecognition element thereby increasing the sensitivity in detecting the low allelic frequency targets.

[0209] In some embodiments, the methods disclosed herein comprise hybridization of a targeted recognition element that is complementary to a methylated target site 504. It is contemplated that targeted recognition elements may also hybridize to some extent to a cleaved unmethylated target site, however it is not expected that the ends of the recognition element would be in such proximity that ligation would occur and therefore no, or minimal, amplification or detection would result.

[0210] In some embodiments, the methods disclosed herein comprise hybridizing a plurality of recognition elements to multiple targets of interest in a sample. A recognition element can include, but is not limited to, 5’ and 3’ ends that are complementary to adjacent sequences at the methylated target site of interest. In some embodiments, the 5’ and 3’ ends of the recognition element are not adjacently located, such that a gap exists between the two ends. A gap fill reaction can be performed, wherein the methylated site of interest is located within the gap and a gap fill reaction can be performed to extend the 5’ end of the recognition element to the 3’ end of the recognition element incorporating the methylation site of interest in the extension reaction. A ligation reaction can ligate the extended 5’ end to the 3’ end of the recognition element. Whether adjacent or non-adjacent followed by gap fill, the ends of the recognition element can be ligated to generate a circularized and ligated recognition elements comprising the sequence, or complement thereof, of the target of interest. Additionally, each recognition element comprises a code or hypercode that identifies the 5’ and 3’ ends, therefore also the target of interest, wherein the code or hypercode is used as an indirect detection of the presence of the target of interest in the sample. In some embodiments, the recognition element comprises one or more of an amplification primer binding site, a unique molecular identifier, a universal primer binding site, a cleavage site, and sequences that could facilitate using the recognition element in next generation sequencing methods (e.g., sequence by synthesis, nanopore sequencing, sequencing by hybridization, sequencing by ligation, etc.).

[0211] In some embodiments, the plurality of recognition elements target different types of targets of interest in the same sample in the same assay. For example, one assay can comprise a subset of recognition elements that target methylated target sites of interest in a sample and another subset of recognition elements that target variants of interest in the same sample, forexample a single nucleotide polymorphism, an insertion, a deletion, a copy number variant. As such, one sample may be interrogated for the presence of more than one type of target of interest.

[0212] In some embodiments, after ligation and circularization of the recognition elements in an assay, the reactions can be treated with one or more exonucleases. Treatment of the reactions with exonucleases digests left over linear single stranded recognition elements and target nucleic acids that remain after recognition element / target of interest hybridization thereby decreasing any potential background interference from the unreacted components. Exonucleases of use include, but a not limited to, exonuclease I, exonuclease II, exonuclease III, or any combination of exonucleases.

[0213] In some embodiments, after ligation and circularization of the recognition elements in an assay, the recognition elements can be further amplified thereby generating multiple copies of the recognition element and all of its parts. In some embodiments, the ligated and circularized recognition elements can be isothermally amplified by rolling circle amplification where multiple copies of the ligated recognition element are concatenated end to end. In some embodiments, the ligated and circularized recognition elements can be amplified by multiple strand displacement. In some embodiments, the ligated and circularized are amplified by other means as disclosed herein. In preferred embodiments, the ligated and circularized recognition elements are amplified by rolling circle amplification, generating concatemeric amplification products. In some embodiments, the amplification is performed in solution. In some embodiments, the amplification is performed on a substrate, such as in a well of a multi-well plate. In some embodiments, the substrate can be treated with an immobilization compound such that amplification can be performed on the ligated recognition elements that are immobilized on the substrate resulting in immobilization of the concatemeric amplification products on the substrate. In some embodiments, amplification is performed by a DNA polymerase, including but not limited to EquiPhi DNA polymerase, or Phi29 DNA polymerase.

[0214] In some embodiments, the concatemeric amplification products are queried for the presence of the hypercodes that are associated with the target of interest 506. Each recognition element has a unique hypercode that can serve as a proxy for the presence of the target of interest in a sample, regardless of how many targets of interest there are in one sample. For example, a recognition element that targets a methylation target site of interest can have one hypercode and a recognition element that targets a variant target site of interest in the same sample can have asecond, and different, hypercode. As such, both the methylated target site and the variant target site can be detected in one sample by way of the unique hypercode for each target sequence of interest. A concatemeric amplification product hypercode can be queried using a detection oligonucleotide or a detection polynucleotide complex as previously described. In some embodiments, as each hypercode comprises a plurality of nucleic acid segments that make up the hypercode as previously described, each hypercode and hence each concatemeric amplification product can be queried with a detection polynucleotide complex one or more times, two or more times, three or more times, four or more times, five or more times, six or more times seven or more times, eight or more times, nine or more times, ten or more times, eleven or more times, twelve or more times, thirteen or more times or more than thirteen times. For example, each nucleic acid segment can be queried at least once, at least twice, at least three times, at least four times. The number of times a hypercode, or a subset of a hypercode, is queried can depend on the number of nucleic acid segments that comprise a hypercode. The more complex the hypercode, the greater the number of detection events needed to generate a hypercode profile that will differentiate one hypercode from the next. As such, the multiple rounds, or cycles of detection of the hypercode by detection polynucleotide complexes reveals a hypercode profile that is unique to a recognition element.

[0215] In some embodiments, the hypercode profile that is identified for a recognition element can be decoded 506 such that the profile is aligned with the known hypercode profiles of each recognition element. A soft decoding pipeline can perform the decoding, resulting in a probability that the decoded hypercode profile is aligned with a target of interest. FIG. 5, step 506 demonstrates an exemplary decoding and analysis result output, where the decode counts are high for the methylated target site identification compared to an unmethylated target site. In addition, if a second type of target of interest is detected and decoded, for example if a variant such as a single nucleotide polymorphism is a target of interest, the exemplary decoding and analysis results can include whether the variant, or ALT, allele is present or whether the wild type or REF allele for both methylated and unmethylated samples.

[0216] In some embodiments, the present disclosure includes kits for practicing the methods described herein for detection of methylated target sites of interest, with or without detecting the presence of a second type of target of interest in the sample. In some embodiments, a kit may comprise nucleic acid extraction reagents for extracting and purifying DNA from otherenvironmental components, such as proteins, cellular debris such as that found in a cellular or tissue lysate. Nucleic acid extraction reagents could comprise beads that preferentially bind nucleic acids in a solution or suspension. In some embodiments, a kit may comprise one or more enzymes such as one or more ligases, one or more polymerases, one or more restriction endonucleases, one or more exonucleases and buffers and other reagents needed to practice enzymatic reactions such as ligations, amplifications, and digestions. In some embodiments, a kit may comprise one or more recognition elements designed to hybridize to a target of interest. In some embodiments, a kit comprises a first plurality of recognition elements, wherein each recognition element in the first plurality hybridizes to one methylation target site of interest. In some embodiments, a kit can comprise a second plurality of recognition elements, wherein each recognition element in the second plurality hybridizes to one variant target of interest. In some embodiments, a kit can comprise one or more substrates for practicing one or more steps in an assay. In some embodiments, a kit can comprise instructions for practicing a method as described herein.

[0217] In some embodiments, methods for identifying one or more methylated nucleotides comprises the use of antibodies that can bind methylated nucleotides, such as Anti-5mC antibodies for 5 -methylcytosine detection, Anti-6mA antibodies for 6-methyladenine detection, Anti-m6A antibodies for methylated RNA adenine detection, and Anti-5hmC antibodies for 5-hydroxymethylcytosine detection (see FIG. 19). Methylation specific antibodies can be found commercially, for example from abeam, ThermoFisher, ActiveMotif, Zymo Research, Diagenode and EpigenTek, to name a few. In some embodiments, an antibody that recognizes and binds to methylated nucleotides can be conjugated to an oligonucleotide, thereby generating an antibody-oligonucleotide complex. For example, in FIG. 19, an antibody-oligonucleotide conjugate 1900 that recognizes and binds to a CpG site can comprise Anti-5mC 1902 antibody that is linked via a linker to an oligonucleotide 1901. Additionally, an antibody-oligonucleotide conjugate 1910 that recognizes and binds to a hydroxy methylated cytosine can comprise Anti-5hmC 1912 antibody that is linked via a linker to an oligonucleotide 1911. By way of example, Anti-5mC 1902 can bind to a methylated cytosine if present on a DNA strand 1920 and Anti-5hmC 1912 can bind to a hydroxymethylated cytosine if present on a DNA strand 1920.

[0218] However, methods using an antibody-oligonucleotide conjugate that recognizes methylated cytosines does not provide context as to where on the DNA strand the methylatedcytosines are located. While information with regards to a general epigenetic signature is important, the ability to identify target locations of one or more methylated nucleotides is also very important. As such, when referring to FIG. 19, recognition elements 1903 and 1913 can provide locational context for methylated targets of interest, thereby providing a method for identifying methylated targets of interest. In some embodiments, a recognition element, 1903 and 1913. comprise two oligonucleotides, 1903a and 1903b or 1913a and 1913b. respectively, wherein once ligated the two oligonucleotides reconstitute one complete circularized recognition element, either 1903 or 1913, respectively, each comprising a unique hypercode and possibly additional functional sequences. As the methylation targeted antibody, 1902 or 1912, recognizes and binds their 5mC or 5hmC targets, respectively, a portion of each of the oligonucleotides recognizes and hybridizes to its complementary DNA sequence in proximity to the methylated nucleotide. For example, as antibody 1902 binds to the target methylated nucleotide on the DNA 1920, 1905a and 1905b hybridize to their target sequences on the DNA in proximity to the methylated nucleotide that is bound to the antibody. It is contemplated that the target nucleic acids that hybridize to 1905a and 1905b are in proximity to, but not necessarily immediately adjacent to, the methylated target nucleotide of interest which may need to be taken into account with the binding of the antibody 1902, as antibodies can be nanometers in size. The same scenario can hold true for the hybridization of 1915a and 1915b in proximity to the binding of antibody 1912 to its methylated target nucleotide of interest.

[0219] In some embodiments, upon hybridization of the oligonucleotides to their nucleic acid complementary sequences on the DNA, the two ends can be ligated together (indicated by a star). In preferred embodiments, the ends of the oligonucleotides that hybridize to target nucleic acid complementary sequences hybridize adjacently on the DNA, thereby allowing for ligation to occur. In other embodiments, the ends of the oligonucleotides that hybridize to target nucleic acid complementary sequences hybridize non- adjacently on the DNA, wherein a gap fill event would occur prior to ligation.

[0220] In some embodiments, substantially simultaneously with the hybridization and ligation of the oligonucleotides the target nucleic acid sequences on the DNA, hybridization can also occur between the oligonucleotide 1901 or 1911 linked to the antibody 1902 or 1912, and a second set of complementary sequences 1904a and 1904b or 1914a and 1914b, respectively, such that the ends of the hybridized oligonucleotides can be ligated together (indicated by a star). In someembodiments, the oligonucleotide linked to the antibody hybridizes to adjacent sequences of the second set of target nucleic acid sequences. In other embodiments, the ends of the oligonucleotides that hybridize to target nucleic acid complementary sequences hybridize non-adjacently, wherein a gap fill event would occur prior to ligation. Once hybridized, the ends 1904a and 1904b or 1914a and 1914b, respectively, can be ligated together in a second ligation event (indicated by a star). The dual ligations of the two sets of ends of recognition element 1903 and 1913 provide circularized recognition elements comprising a hypercode (not shown) that differentiates the two recognition elements. If one of the two hybridization and ligation events fails to provide a circularized recognition element, subsequent amplification is not expected to occur and no detection or decoding would follow. Such an event could occur if either the methylated nucleotide is not present and the antibody does not bind, therefore the oligonucleotide linked to the antibody does not hybridize to the ends of the recognition element and no ligation event occurs, or if the sequences in the DNA are not present and the ends of the recognition element do not hybridize and the ligation event does not occur.

[0221] In preferred embodiments, the oligonucleotides 1901 or 1911 comprise two different universal sequences specific to whether the target of interest is a 5mC or 5hmC antibody target. For example, the oligonucleotide 1901 can be universal to every oligonucleotide-antibody conjugate 1900, wherein the antibody binds to a target 5mC nucleotide. A universal sequence can be an amplification primer binding site that, once reconstituted by ligation, can serve as a complementary sequence for an amplification primer for amplifying the ligated and circularized recognition element 1903. The universality of an amplification primer binding site can allow for multiple recognition elements to be concurrently amplified in an assay, thereby providing a plurality of amplified products, such as concatenated amplification products, resulting from RCA.

[0222] The same scenario can occur with the recognition element 1913. For example, the oligonucleotide 1911 can be universal to every oligonucleotide-antibody conjugate 1910, wherein the antibody binds to a target 5hmC nucleotide. A universal sequence can be an amplification primer binding site that, once reconstituted by ligation, can serve as a complementary sequence for an amplification primer for amplifying the ligated and circularized recognition element 1913. The universality of an amplification primer binding site can allow for multiple recognition elements to be concurrently amplified in an assay, thereby providing aplurality of amplified products, such as concatenated amplification products, resulting from RCA.

[0223] In the scenario illustrated in FIG. 19, the two oligonucleotides 1901 and 1911 can represent two different universal sequences that are complementary to two different universal sequences on the oligonucleotides of the two oligonucleotides of each recognition element 1903 and 1913, thereby providing each methylated site, either 5mC or 5hmC, a different universal amplification primer binding site. However, in some embodiments the sequences of the oligonucleotides that are linked to different antibodies are the same and not different, such that the oligonucleotides recognize and hybridize to the same complementary amplification primer binding site. As such, only one amplification primer would be necessary for amplifying both ligated and circularized 1903 and 1913 recognition elements. As with all recognition elements described herein, unique hypercodes (not shown) are also included in the recognition elements 1903 and 1913, such that identification and correlation of the targeted methylated nucleotide is possible.

[0224] As described above, the scenario illustrated in FIG. 19 would also apply in all its embodiments if a methylated adenine nucleotide, 6-mA, or RNA methylated adenine, m6A, is the methylated target of interest. Additional modified nucleotides which may be of interest in detecting include, but are not limited to, RNA modifications like N6,2-O-dimethyladenosine (m6Am), 8-oxo-7,8-dihydroguanosine (8-oxoG), pseudouridine. 5-methylcytidine (m5C) and N4-acetylcytidine (ac4C).

[0225] In some embodiments, the nucleic acid strand 1920 comprises RNA rather than DNA, for example mRNA. In some embodiments, the nucleic acid strand 1920 comprises cDNA from transcribed mRNA, for example using RT-PCR methods to generate the cDNA. In some embodiments, the nucleic acid strand 1920 is cell-free DNA, for example harvested from plasma or serum and may comprise methylated targets that are associated with cancers or disease states.

[0226] In some embodiments, the 5’ and / or 3’ ends of a recognition element can include locked nucleic acids or other modified bases which may lend stability in the face of hybridization mismatches. In some embodiments, the 5’ and / or 3’ ends of a recognition element can comprise RNA sequences for hybridizing to transcribed methylated DNA sequences, if they be a target of interest.If gap filling is desired, such that the 5’ and 3’ ends of a recognition element do not adjacently hybridize, a DNA polymerase can be utilized to extend the 5’ end of the hybridized recognition element until it reaches the 3’ end of the recognition element. Alternatively, a bridge oligonucleotide could be used to fill the gap between the 5’ end and the 3’ end of a non-adjacently hybridized recognition element. In some instances, a gap may be present due to a homopolymer region that may be present in a sequence, wherein gap fill of the homopolymer region would also provide valuable information as to the potential number of repeats and / or the methylation status of the homopolymer region. In fact, in some embodiments, homopolymeric regions may be the target(s) site of interest, wherein the epigenetic signature of the homopolymeric region is the research focus.

[0227] As FIG. 19 finds utility in identifying differentially methylated CpG sites, or adenine methylation at a single site, FIG.20 and FIG. 21 find utility in identifying differentially methylated regions, or hot spots of methylation, or a plurality of methylated nucleotides in a given area or location.

[0228] Referring to FIG.20, a machine learning scenario is illustrated using methylated and unmethylated recognition element binding kinetics for developing training data sets and generating a database that can be used in identifying methylation at a target location in a sample. In FIG.20, two recognition elements are designed, one that recognizes a methylated nucleotide at a given location on a DNA strand, in this instance a 5mC (but could be 5hmC, 6mA, etc.)(top left) and a second that recognizes a non-methylated nucleotide at a given location on a DNA strand (bottom left). The 3’ end of the recognition elements can hybridize to either the methylated or unmethylated nucleotide, wherein the dynamics of the hybridization event can change depending on the methylation thereby providing differences in final concatemeric amplification product counts upon detection and decoding as previously described. A multitude of assays can be performed to generate one or more data sets for one or more target methylation sites. Through machine learning, the data sets can be used to populate a database which can be used to train a system to predict whether a target site of interest is methylated, and a call on methylation can be effected. As such, based on the recognition element count following detection of the concatemeric amplification products, percent methylation can be assigned at a differentially methylated CpG site (or other methylated nucleotide site) as learned from a trained database.Referring to FIG.21, a machine learning scenario is illustrated using methylated binding kinetics for developing training data sets and generating a database that can be used in identifying levels of methylation for differentially methylated regions, or hot spots, in a sample. In this scenario, three recognition elements each with a unique hypercode are illustrated wherein each of the recognition elements recognizes and hybridizes to a different epigenetic signature. For example, the top illustration recognition element 2100, which comprises hypercode 1, is illustrated as being complementarity to four target methylation sites (C); recognition element 2101, which comprises hypercode 2, is illustrated as being complementarity to three target methylation sites and recognition element 2102, which comprises hypercode 3, is illustrated as being complementarity to two target methylation sites. In this scenario, the methylated targets indicate three methylation hot spots and the question to be answered is how many of the nucleotides available for methylation are methylated. The recognition elements 2100, 2101 and 2102 comprise 5’ and 3’ arms that target the differentially target different numbers of target methylation sites, wherein the hypercode 1, 2 or 3 is unique to that recognition element and therefore the number of target methylation sites to be identified. The binding kinetics of the three different recognition elements can change depending on the number of methylation sites identified and hybridized to, thereby providing differences in final concatemeric amplification product counts upon detection and decoding. A multitude of assays can be performed to generate one or more data sets for the different numbers of nucleic acids that are methylated. Through machine learning, the data sets can be used to populate a database which can be used to train a system to predict the number of methylated sites present in a sequence, in this example four (2100), three (2101) or two (2102) methylated nucleotides in a DNA strand, wherein the data can be sorted by the unique hypercodes to provide a percent methylation by hypercode (right side graph) and a call on an epigenetic signature can be effected. As such, based on the recognition element count following detection of the concatemeric amplification products and alignment with the unique hypercodes, percent methylation can be assigned for differentially methylated regions, or hot spots, as assigned from a trained database.

[0229] In some embodiments, the present disclosure provides for a system, or a platform, that comprises an assay module for performing any one of the methods described herein in conjunction with an analysis module for determining whether the methylated nucleotide(s) of interest and / or the variant nucleotide(s) of interest are present in a sample of interest. An assaymodule of a system can be configured to perform one or more steps in an assay workflow, for example a hybridization step, a ligation step, one or more digestion steps, an amplification step, and a detection step as described herein. For example, an assay module could include one or more of 1) subjecting an extracted nucleic acid sample to a restriction digest step wherein methylated sequences of interest are not digested but non-methylated sequences are not digested, 2) optionally amplifying, for example by PCR, the restriction digest products such that only the non-digested methylated sequences are amplified, 3) hybridizing the amplification products to target recognition elements that comprise unique hypercodes specific to each target methylated sequence or each variant sequence, 4) ligating the hybridized recognition elements, 5) amplifying the ligated recognition elements to generate concatemeric amplification products, and 6) detecting the amplified recognition elements. As described herein, detection of the amplified recognition elements in an assay module can comprise hybridizing fluorescently labeled detection polynucleotide complexes to the hypercodes, or portions thereof, of the concatemeric amplification products, in multiple cycles, to generate a hypercode profile for each concatemeric amplification product.

[0230] A system would also include an analysis module, wherein the analysis module would use the hypercode profiles and, for example by soft decision decoding, determine whether the target methylated nucleotide and / or the target variant nucleotide was present in the sample. In some embodiments, the hypercode profile is decoded utilizing a soft decision decoding algorithm to determine the probability of the presence of the target, either a methylated target and / or a variant sequence target, present in an original nucleic acid sample.

[0231] In some embodiments, decoded hypercode assignments and associated confidence scores are used as inputs for downstream analysis, including machine-leaming-based methylation status determination.

[0232] Various modifications and variations of the disclosed methods, compositions and uses as disclosed herein will be apparent to the skilled person without departing from the scope and spirit of the inventive concepts. Although methods have been disclosed in connection with specific preferred aspects or embodiments, it should be understood that methods, compositions and systems as claimed should not be unduly limited to such specific aspects or embodiments. Although the foregoing subject matter has been described in some detail by way of illustration and example for purposes of clarity of understanding, it will be understood bythose skilled in the art that certain changes and modifications can be practiced within the scope of the appended claims.

[0233] EXAMPLES

[0234] The following examples are provided for illustrative purposes only and are not intended to limit the scope of the present disclosure.

[0235] Example 1 -Detection of both methylation status and variants in a single sample

[0236] This example presents a general workflow used in all the experiments where methylated targets were evaluated. A positive DNA control NA 12878 (Coriell Institute of Medical Research) and a negative water control were run in conjunction with the test experiments.

[0237] The general workflow is exemplified in FIG. 2 and includes 201 combining a sample comprising target nucleic acids, that has or has not been hypermethylated and cleaved, with one or more recognition elements that include complementary ends to one or more target nucleic acid sequences, wherein each recognition element further comprises a code that is unique to the recognition element ends, and thus can be used as an indirect method for determining the presence of the target nucleic acid; 202 hybridizing the target nucleic acids and the recognition elements; 203 ligating the ends of the recognition elements thereby generating circularized recognition elements; 204 exonuclease treating the ligation reaction to remove any linear oligonucleotides such as unhybridized target nucleic acids and unhybridized recognition elements; 207 amplifying the circularized recognition elements; 208 detecting the code of the amplified recognition elements; and 209 decoding the detection output to determine methylation status of the target nucleic acid.

[0238] The DNA sample NA 12878 was hypermethylated using CpG methyltransferase. Briefly, 2 ug of NA12878 was combined in a reaction with CpG methyltransferase, buffer and S-adenosylmethionine (SAM) and the methylation reaction was performed at 37°C for 60 min, 65°C for 20 min. and put on ice.

[0239] Following methylation, the hypermethylated DNA sample was digested with Hhal, a restriction endonuclease that is methylation sensitive, as such Hhal will not cleave a CpG methylated sequence but will cleave a nucleic acid sequence that is not methylated, thereby preserving intact a nucleic acid target sequence that has been methylated. Briefly, 100 ng of the test CpG methyltransferase treated sample was incubated with an Hhal restriction enzyme and an associated buffer for 37°C for 60 min., 65°C for 20 min. and put on ice. Sample cleavage wasverified by agarose gel electrophoresis and quantified by Qubit Fluorometric Quantification (ThermoFisher Scientific).

[0240] The test samples and controls (100 ng) were added to separate wells of an assay plate and were combined with a reaction mixture that include 0.5 nM of each recognition element, including recognition elements targeting methylation targets of interest and SNPs of interest, Ampligase and ligase buffer, a hybridization / ligation enhancer additive, and water to a total reaction volume of 20 pL. All samples and controls were run in triplicate in one plate, and two total plates were run. The methylation targets of interest (locations based on GRCh38), SNV targets of interest and their WT counterparts, and the complementary recognition element 5’ and 3’ end sequences that recognize and hybridize to the target sequences are listed in FIG. 7 for the methylation targets and FIG.8 for the variant targets.

[0241] The reactions were incubated for concurrent hybridization and ligation; 95°C / 2 min., 60°C / 15 min. followed by 95°C / 1 min. for nine cycles, followed by placing on ice.

[0242] An endonuclease reagent master mix was created, including Exonuclease I, Exonuclease III, an exonuclease buffer and water to bring the master mix to a total of 30 pL for each test well, and the master mix was aliquoted into each sample well. The exonuclease digestion reactions were incubated at 37°C for 30 min. followed by an enzyme inactivation at 95°C for 5 min. After the exonuclease digestion, the samples were placed on ice.

[0243] Following exonuclease digestion, the reactions were transferred to a new well plate which had been pre-treated for DNA immobilization and further included an optically clear bottom for subsequent detection reactions. After sample transfer, the plate was sealed and incubated at 42°C for 1 hr.

[0244] Following incubation, the circularized and ligated recognition elements were amplified by rolling circle amplification. An amplification master mix reagent was created, including 1 mM each dNTP, EquiPhi DNA polymerase and buffer, 10 nM of an amplification primer and water for a total well dispense volume of 50 pL. To each well was added the 50 uL amplification reagent, and the amplification reactions were incubated at 42°C for 2 hrs. The amplification reactions were washed several times with TE buffer, incubated with 0.1 N NaOH for 1 min., washed several times with TE buffer, and left in TE buffer for subsequent detection.

[0245] Detection reagents were generated that included a hybridization buffer, 0.5 pM of an anchor oligonucleotide, 0.5 pM of a fluorescently labeled detection oligonucleotide, a DNAblocker and water for a total well dispense volume of 50 pL. Detection reagent was added to each well, the reactions were incubated at room temperature for 15 min. and washed several times to remove any unhybridized anchor and detection oligonucleotides. The wells were imaged on a fluorescent imager (in these experiments a RAPTOR fluorescent imaging instrument was used for detection and decoding). Each well was queried multiple times, thereby generating a fluorescent profile indicative of the unique code for each of the amplified recognition elements in each of the wells. Four different fluorescently labeled detection oligonucleotides were used in conjunction with their complementary anchor oligonucleotides. A fluorescent moiety AZdye532 (ex / em 532 / 554 nM)), AZdye568 (ex / em 578 / 602 nM), AZdye647 (ex / em 648 / 671 nM) or AZdye680 (ex / em 678 / 701 nM) was attached to a detection oligonucleotide such that four colors were used for detection of code sequences, and the final fluorescent profile built for each code following multiple queries, which was then decoded, was built for each amplified recognition element.

[0246] After the amplified recognition element detection cycles were completed and a fluorescent profile generated, the fluorescent profiles were bioinformatically decoded and matched to the known profile for the hypercodes used in the recognition elements and the resulting data was used to determine whether the methylated target of interest was present, or not, and whether the SNP of interest was present, or not.

[0247] In some embodiments, internal calibrators are used for normalization. In some embodiments, internal calibrators comprise variant recognition elements, wild-type recognition elements, synthetic controls, or combinations thereof.

[0248] Methylation data was normalized using average signal from the internal calibrators made up of the SNP recognition elements. FIG. 9 and FIG. 10 show example graphs of improved response due to normalization by internal calibrators, based on the increased R2values for chrl6:72622501C_C and chr7:27165267C_C, respectively.

[0249] The effect of the presence of a restriction endonuclease on the ability of the assay to detect the target variants was evaluated. As demonstrated by an exemplary graph in FIG. 11, in the presence of a restriction endonuclease Hhal, the ability of the assay to detect the target variants was not compromised. In FIG. 11, a hypermethylated NA12878 sample that was restricted with Hhal prior to hybridization of the sample with the recognition elements is shown by the right hand column (red) for each of the target sets. A negative control, the left handcolumn (blue) for each target set was a sample of NA12878 that was not restricted with HhaT. For each target variant, two recognition elements were used, one with complementary sequences to the wild type target sequence and a second with complementary sequences to the variant target sequence. The expected variants, either WT or ALT, were identified in the assay regardless of the presence of a restriction endonuclease.

[0250] The results of an exemplary multiomics assay where both methylation status of target sites and variant detection was performed on the same sample is shown in FIG. 12. The left hand column (blue) for each methylation target represents the variant allele frequency of a NA 12878 sample that was not hypermethylated but was endonuclease digested, thereby serving as a control for the methylated, endonuclease digested sample of the right hand column (red). While there is some background frequency for the negative control samples, it was contemplated that while the non-hypermethylated sample was not treated with CpG methyltransferase, some of the target sites may have been nascently methylated. Regardless, it is apparent that the methods described herein can detect and differentiate methylated target sites of interest when they are methylated compared to the samples that were cleaved due to the lack of methylation at the target site. As well, the variant targeted sites that were targeted in both the hypermethylated NA12878 and the non-methylated NA12878 control were basically the same, as expected. While the nonmethylated and cleaved NA 12878 control may negatively affect the ligation of the recognition element targeting the methylated site, it is not expected to affect the hybridization of the recognition elements that were targeting variants in the same sample as those variant sites did not participate in the cleavage reaction, as such they were able to hybridize to their target recognition elements, ligation and amplification events could occur and detection was possible. As with the methylated targets of interest providing the expected results, so to the variant targets were successfully identified.

[0251] As such, the methods described herein were successful in combining and detecting two different types of targets of interest, methylation targets and variant targets, in a sample in one assay.

[0252] Example 2- Determination of dynamic range for sample input amounts

[0253] Experiments were performed to evaluate the dynamic range of detection for the methylated targets and variant targets. The assay workflow was performed as found in Example1. Methylation targets were assayed and dynamic range of detection were evaluated using two different analysis methods.

[0254] FIGs. 13A, B, C and D shows exemplary graphs supporting that the methods used in this disclosure can detect a methylated target as low as 1% in a background of 0%, depending on the recognition elements used in the assay. For example, FIG. 13A shows that when utilizing three different recognition elements to interrogate methylation at chr3:158105726C_C in NA12878 there was significant separation from no methylation at 1% using a Student’s t-test. FIG. 13B shows better separation of methylation target detection of 2.5% methylation at chr7:24284265C_C with 12 different recognition elements, FIG. 13C demonstrates still greater separation at 5% methylation with 14 different recognition elements and FIG. 13D demonstrates the greatest amount of separation between no methylation and 10% methylation at chr3:36992857C-C using 22 different recognition elements.

[0255] Additional examples of dynamic range are seen in FIG. 14, FIG. 15 and FIG. 16, when detecting a methylated target in MLH1 (chr3:36992857C_C), TWIST1 (chr7:19118295C_C). and APC (chr5: 112707442C_C), respectively, using the All Pairs, Tukey Kramer method for multiple comparisons. The bar graph to the right of each range of detection graph on the left shows the decode count for an unmethylated target site (0% methylated) compared to the decode count for the methylated target (100% methylated) at the point on the horizontal line in the range graph on the left. These data demonstrate that, while each different recognition element used to inteiTogate a methylated site can differ in its limit of detection, generally the limit of detection for the recognition elements used in assays was around 5% to 25%.

[0256] As such, the methods disclosed herein demonstrate the detection of methylated targets, even when the target in question was present at 1% depending on the recognition element used in the assay. Further, assays were able to detect methylated targets in a sample in addition to other target sequences, such as detecting variants in a sample in combination with detection of targeted methylation sites thereby enabling multiomics detection of multiple, different targets in one sample.

Claims

CLAIMS1. A method for determining the methylation status of a sample, comprising:a) cleaving a target nucleic acid with a restriction endonuclease, wherein the restriction endonuclease cleaves one or more non-methylated nucleic acid sequences of interest of the target nucleic acid and does not cleave methylated nucleic acid sequences of the target nucleic acid. b) hybridizing one or more recognition elements to the target nucleic acid methylation sites of interest thereby generating one or more hybridized recognition elements, wherein the one or more recognition element comprises a 5’ end and a 3’ end, and wherein the 5’ end of the one or more recognition elements and the 3’ end of the one or more recognition elements are complementary to the target nucleic acid methylation sites of interest,c) ligating the 5’ end of the one or more hybridized recognition elements and the 3’ end of the one or more hybridized recognition element to produce one or more ligated recognition elements, andd) determining the methylation status of the sample based on the presence or absence of the one or more ligated recognition elements.

2. The method of claim 1, wherein the one or more recognition elements further comprise a hypercode.

3. The method of claim 1 or claim 2, wherein each of the one or more recognition elements comprises a unique hypercode and targets a unique target nucleic acid sequence.

4. The method of any of claims 1-3, wherein the restriction endonuclease is selected from the group consisting of Hhal, BstUI, Hpall. AccI, MspI, MspJI, Aatll. Acil, AcII, Afel, Agel, Asci, AsiSI, Aval, BceAI, BmgBI, BsaAI, BsaHI, BsiEI, BsiWI, BsmBI-v2, BspDI, BsrFI-v2, BssHII, BstBI, Clal, Eagl-HF, Esp3I, Faul, Fsel, FspI, Haell, Hgall, Hhal, HinPlI, HpyCH4IV, Hpy99I, KasI, Mini. Nael, Narl, NgoNIV, Notl. Nt.BsmAI, Nt.CviPII, PaeR71, PluTi. Pmll, Pvul, SacII, Sall, Sfol, SgrAI, Smal, SnaBI, Srfl, TspMI, and Zral.

5. The method of claim 4, wherein the restriction endonuclease is Hhal.

6. The method of claim 1, wherein the restriction endonuclease is selected from the group consisting of Alwl, Bell, BclI-HF, DpnII, HphI, Mbol, Nt.AlwI, PspGI, and SexAI.

7. The method of any one of claims 1-6, further comprising providing one or more additional recognition elements, wherein the one or more additional recognition elements hybridizes to one or more of a second type of target of interest in the target nucleic acid.

8. The method of claim 7, wherein the one or more of the second type of target of interest comprises a nucleic acid variant.

9. The method of claim 8, wherein the nucleic acid variant comprises a single nucleotide polymorphism, an insertion, a deletion, a copy number variant, or a splicing variant.

10. The method of claim 9, wherein the nucleic acid variant comprises a single nucleotide polymorphism.

11. The method of any one of claims 1-10, further comprising amplifying the one or more ligated recognition elements.

12. The method of claim 11, wherein the amplifying comprises rolling circle amplification.

13. The method of claim 11, wherein the amplifying comprises multiple strand displacement amplification.

14. The method of any of claims 1-13, wherein determining the methylation status of the sample comprises detecting the hypercode of the one or more ligated recognition elements.

15. The method of claim 14, wherein detecting the hypercode of the one or more ligated recognition elements comprises hybridizing two or more detection polynucleotide complexes to the hypercode, or a portion thereof, to generate a hypercode profile.

16. The method of claim 15, wherein the hypercode profile is decoded using a soft decision decoding algorithm.

17. The method of any one of claims 1-16, wherein the method does not comprise chemically converting unmethylated cytosines prior to detecting the hypercode.

18. A composition comprising a nucleic acid sample and a first plurality of recognition elements hybridized to a plurality of methylated target nucleic acid sequences of interest in the nucleic acid sample and a second plurality of recognition elements hybridized to a plurality of variant nucleic acid target sequences of interest in the nucleic acid sample.

19. The composition of claim 18. further comprising a ligase.

20. The composition of claim 18 or claim 19, further comprising one or more exonucleases.

21. The composition of any one of claims 18-20, wherein one or more of the plurality of variant nucleic acid target sequences of interest comprises one or more of a single nucleotide polymorphism, an insertion, a deletion, a copy number variant, or a splicing variant.

22. A method for identifying the presence of a methylated nucleic acid in a sample, comprising:a) providing an antibody-oligonucleotide conjugate to the sample, wherein the oligonucleotide of the antibody-oligonucleotide conjugate comprises a first sequence and a second sequence and wherein the antibody of the antibody-oligonucleotide conjugate binds a methylated nucleotide;b) binding the antibody of the antibody-oligonucleotide conjugate to the methylated nucleotide in the sample;c) providing two oligonucleotides to the sample, wherein the first oligonucleotide of the two oligonucleotides comprises a first portion that is complementary to a first sequence in the sample that is in proximity to the methylated nucleotide and a second portion that is complementary to the first sequence of the oligonucleotide of the antibody-oligonucleotide conjugate and the second oligonucleotide of the two oligonucleotides comprises a first portion that is complementary to a second sequence in the sample that is in proximity to the methylation nucleotide and a second portion that is complementary to the second sequence of the oligonucleotide of the antibody-oligonucleotide conjugate;d) hybridizing the two oligonucleotides to their complementary sequences in the sample and hybridizing the oligonucleotide of the antibody-oligonucleotide conjugate to its complementary sequences on the two oligonucleotides;e) ligating the two ends of the oligonucleotides hybridized on the sample and the two ends of the oligonucleotides hybridized to the oligonucleotide of the antibody-oligonucleotide conjugate, thereby generating a circularized recognition element;f) amplifying the circularized recognition element to generate a concatemeric amplification product; andg) detecting the presence of the methylated nucleotide in the sample by the presence of the concatemeric amplification product.

23. The method of claim 22, wherein steps (a) through (d) occur substantially simultaneously.

24. The method of claim 22, wherein first step (a) and step (b) occur substantially simultaneously and second step (c) and step (d) occur substantially simultaneously.

25. The method of any one of claims 22-24, wherein the methylated nucleotide comprises 5mC, 5hmC, or 6-mA.

26. The method of any of claims 22-25, wherein the two oligonucleotides when ligated comprise a unique hypercode.

27. The method of any one of claims 22-26, wherein the two oligonucleotides that hybridize to the oligonucleotide of the antibody-oligonucleotide conjugate generate an amplification primer binding site when ligated.

28. The method of any one of claims 22-27, wherein the amplifying comprises rolling circle amplification.

29. The method of any one of claims 22-28. wherein the detecting comprises detecting the presence of a hypercode in the concatemeric amplification product, and wherein the hypercode is indicative of the presence of the methylated nucleotide in the sample.

30. The method of any one of claims 22-29, wherein the antibody of the antibody-oligonucleotide conjugate comprises Anti-5mC, Anti-5hmC, or Anti-6mA.

31. The method of any one of claims 22-30, wherein the antibody of the antibody-oligonucleotide conjugate is conjugated to the oligonucleotide of the antibody-oligonucleotide conjugate by a linker.

32. A composition comprising an antibody-oligonucleotide conjugate, wherein the antibody of the antibody-oligonucleotide conjugate is bound to a methylated nucleotide on a nucleic acid sample and a first portion of the oligonucleotide of the antibody-oligonucleotide conjugate is hybridized to a 5’ portion of a first oligonucleotide and a second portion of the oligonucleotide of the antibody-oligonucleotide conjugate is hybridized adjacent to a 3’ portion of a second oligonucleotide.

33. The composition of claim 32, further wherein a 3’ portion of the first oligonucleotide is hybridized to the nucleic acid sample in proximity to the methylated nucleotide and a 5’ portion of the second oligonucleotide is hybridized to the nucleic acid sample adjacent to the 3’ portion of the first oligonucleotide.

34. The composition of claim 33, wherein the adjacently hybridized portions of the first oligonucleotide and the second oligonucleotide are ligated to each other.

35. A kit, comprising:a) a plurality of recognition elements that are capable of recognizing and hybridizing to a plurality of target nucleic acid methylation sites of interest;b) one or more enzymes, wherein the one or more enzymes comprise a methylation resistant restriction endonuclease, a ligase, a DNA polymerase, or an exonuclease;c) instructions for performing any one of the methods of claims 1-16 or claims 22-31.

36. The kit of claim 35, further comprising a plurality of recognition elements that are capable of recognizing and hybridizing to a plurality of variant nucleic acid target sites of interest.

37. A system for determining methylation status in a sample, comprising:a) an assay module comprising a hybridization reaction or a ligation reaction, or a combination thereof, wherein the hybridization reaction or ligation reaction, or the combination thereof, comprises one or more recognition elements that recognize and hybridize to one or more methylated nucleotides in the sample, andb) an analysis module configured to determine a presence of the one or more methylated nucleotides in the sample based on products generated from the assay module.

38. The system of claim 37. wherein the assay module further comprises one or more additional recognition elements that recognize and hybridize to one or more variant nucleotides in the sample and the analysis module determines the presence of the one or more variant nucleotides in the sample in addition to the one or more methylated nucleotides in the sample.

39. The system of claim 37 or 38, wherein each of the one or more recognition elements that recognize and hybridize to the one or more methylated nucleotides or the one or more variant nucleotides comprises a unique hypercode.

40. The system of claim 39, wherein the analysis module comprises a plurality of detection polynucleotide complexes, where the determining the presence of the one or more methylated nucleotides or the one or more variant nucleotides comprises hybridizing the plurality of detection polynucleotide complexes to the unique hypercode of the one or more recognition elements, or a portion thereof, imaging the hybridization to generate a hypercode profile for each of the one or more recognition elements, and determining the presence of the one or more methylated nucleotides or the one or more variant nucleotides by performing soft decision decoding on the hypercode profiles for each of the one or more recognition elements.

41. The system of claim 37, wherein the assay module comprises a plurality of recognition elements that recognize and hybridize to a plurality of methylated nucleotides in the sample and the analysis module determines the presence of the plurality of methylated nucleotides in the sample.

42. The system of claim 41, wherein the assay module further comprises a plurality of recognition elements that recognize and hybridize to a plurality of variant nucleotides in thesample and the analysis module determines the presence of the plurality of variant nucleotides in the sample in addition to the plurality of methylated nucleotides in the sample.

43. The system of claim 42, wherein each of the plurality of recognition elements that recognizes and hybridizes to the plurality of methylated nucleotides or the plurality of variant nucleotides comprises a unique hypercode.

44. The system of claim 43. wherein the analysis module comprises a plurality of detection polynucleotide complexes, where the determining the presence of the plurality of methylated nucleotides or the plurality of variant nucleotides comprises hybridizing the plurality of detection polynucleotide complexes to the unique hypercode of each of the plurality of recognition elements, or a portion thereof, imaging the hybridization to generate a hypercode profile for each of the plurality of recognition elements, and determining the presence of the plurality of methylated nucleotides or the plurality of variant nucleotides by performing soft decision decoding on the hypercode profiles for each of the plurality of recognition elements.

45. The system of any one of claims 37-44, wherein the sample is pre-treated with a methylation resistant restriction endonuclease prior to the hybridization reaction.

46. The system of any one of claims 37-45, wherein the assay module further comprises an amplification reaction.

47. The system of claim 46, wherein the amplification reaction comprises rolling circle amplification.

48. A method for determining methylation status of a target nucleic acid, comprising:a) obtaining a plurality of amplification products, each amplification product comprising a hypercode associated with a target nucleic acid methylation site;b) iteratively hybridizing a plurality of detection polynucleotide complexes to the plurality of amplification products over a plurality of detection cycles, wherein each detection polynucleotide complex comprises a detection oligonucleotide with a detectable label and an anchor oligonucleotide complementary to a segment of the hypercode;c) imaging the detectable labels after each detection cycle to obtain signal intensities associated with segments of the hypercode; andd) applying a soft decision decoding algorithm to the signal intensities obtained across the plurality of detection cycles to determine a probability that the hypercode is present, whereinthe probability determination is performed without assigning a nucleotide identity to individual positions within the hypercode;thereby determining methylation status of the target nucleic acid.

49. The method of claim 48, wherein the hypercode is selected from a codespace having a Hamming distance of 3, 4, or 5.

50. The method of claim 48 or 49, wherein the signal intensities are obtained from amplification products immobilized on a substrate.

51. The method of any one of claims 48-50, wherein step (d) comprises cross correlating the signal intensities against expected hypercode signal profiles and assigning a most likely hypercode based on a probabilistic metric.

52. A method for determining methylation status of one or more target nucleic acid sites in a sample, comprising:a) providing a sample comprising a target nucleic acid containing one or more target methylation sites;b) hybridizing a plurality of recognition elements to the target nucleic acid, wherein each recognition element is configured to preferentially interact with a methylated or unmethylated form of at least one target methylation site, wherein each recognition element comprises a unique hypercode;c) generating assay data for the plurality of recognition elements, wherein the assay data comprises one or more of binding kinetics, amplification product counts, or signal intensities associated with the hypercodes;d) processing the assay data using a machine learning model comprising a database of methylation signatures derived from differential hybridization behavior of methylated and unmethylated target sites; ande) determining, based on the processing, a methylation status for at least one target methylation site.

53. The method of claim 52, wherein the methylation status is determined based on a probability, percentage, or number of methylated nucleotides.

54. The method of claim 52 or 53, wherein step (c) comprises iteratively detecting binding kinetics, amplification product counts, or signal intensities associated with the hypercodes over a plurality of detection cycles.

55. The method of claim 54, wherein step (d) comprises applying a soft decision decoding algorithm to the binding kinetics, amplification product counts, or signal intensities obtained across the plurality of detection cycles to determine a probability that the hypercode is present, wherein the probability determination is performed without assigning a nucleotide identity to individual positions within the hypercode.

56. The method of any one of claims 52-55. wherein the hypercode is selected from a codespace having a Hamming distance of 3, 4, or 5.

57. The method of any one of claims 52-56, wherein the binding kinetics, amplification product counts, or signal intensities are obtained from amplification products immobilized on a substrate.

58. The method of any one of claims 52-57, wherein step (d) comprises cross correlating the signal intensities against expected hypercode signal profiles and assigning a most likely hypercode based on a probabilistic metric.

59. The method of any one of claims 52-58. wherein the plurality of recognition elements comprises at least two recognition elements that are complementary to different numbers of methylated nucleotides within a target region of the target nucleic acid, wherein the target region comprises a plurality of target methylation sites, and wherein determining the methylation status comprises determining a methylation signature for the target region based on assay data associated with different hypercodes of the plurality of recognition elements.