RNA and DNA analysis using artificial surfaces
Patent Information
- Application Number
- JP2024531038
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-11
- Filing Date
- 2022-11-23
- Publication Date
- 2025-12-01
AI Technical Summary
Current methods for analyzing epitranscriptomic and epigenetic modifications in nucleic acids are not sensitive, cause degradation, and fail to identify modifications with single-base resolution or multiplex detection.
A method involving compositions with substrates, binding domains, and adapters that specifically bind to non-canonical features of DNA or RNA, using nucleic acid barcode sequences for identification and analysis, enabling high-parallelized, sensitive profiling of multiple modifications.
Enables accurate, fast-throughput analysis of DNA and RNA modifications with single-base resolution and multiplex detection, facilitating the discovery of biological control mechanisms and therapeutic paradigms.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Application No. 63 / 282,808, filed November 24, 2021, and U.S. Provisional Application No. 63 / 388,036, filed July 11, 2022, the disclosures of which are incorporated herein by reference in their entireties.
[0002] Technical Field The present invention relates generally to the identification and analysis of epitranscriptomic, epigenetic, and other modifications or non-canonical features to the structure of nucleic acids, including RNA and DNA.
[0003] Federal funding provisions This invention was made with United States Government support under Grant No. 1R43HG012170-01 awarded by the National Human Genome Research Institute. The United States Government has certain rights in the invention. [Background technology]
[0004] background Epigenetic changes, involving chemical alterations of nucleotides, are widespread and play a major role in biological processes such as gene expression, gene silencing, and response to DNA damage. Similarly, chemical modifications of RNA, known as epitranscriptomic (hereafter referred to as epitranscriptome) modifications, frequently occur in cells during or after transcription. RNA modifications play important roles in translation initiation, translation error rate, alternative splicing, RNA stability, folding and transport.
[0005] A wide variety of diseases, behaviors, and other health indicators are correlated with epigenetic changes in DNA, including almost all types of cancer, cognitive dysfunction, and respiratory, cardiovascular, reproductive, autoimmune, and neurobehavioral diseases. However, little is known about the distribution of epigenetic changes across the genome, particularly in relation to health and disease. Although some functions of epitranscriptomic modifications are known, much remains unknown due to the lack of analytical methods to localize and quantify these modifications across cellular RNA. Currently, very little is known about the correlated levels of epitranscriptomic RNA modifications and their changes within cells.
[0006] A combination of chemical derivatization, molecular recognition (usually using antibodies for both enrichment and detection), and reverse transcription sequencing have provided profiling methods for a limited number of DNA and RNA modifications. However, these methods are not sensitive, cause nucleic acid degradation / fragmentation, and often cannot identify the location of modifications with single-base resolution. Furthermore, these methods do not allow for simultaneous multiplex detection of multiple DNA or RNA modifications. Existing methods for sequencing common epitranscriptomic RNA modifications often give conflicting findings, both in the number of modifications detected (which differ by more than an order of magnitude) and in the location of the modifications.
[0007] Thus, there is a need in the art for improved compositions and methods for identifying, analyzing, quantifying and localizing DNA and RNA modifications. Such advances could pave the way for the discovery of important control mechanisms of biology in health and disease, and for the development of new therapeutic paradigms in medicine. Summary of the Invention
[0008] overview The present invention provides compositions and methods for identifying and analyzing epitranscriptomic, epigenetic and other chemical modifications to the structure of nucleic acids, including RNA and DNA. The present invention provides highly parallel, sensitive, highly accurate, high-throughput methods for simultaneously profiling a potentially unlimited number of DNA and / or RNA modifications.
[0009] The present invention provides a composition comprising: i) a substrate; ii) a binding domain that binds to the substrate via a first linker; and iii) an adaptor that binds to the substrate via a second linker, wherein the binding domain specifically binds to a non-canonical feature of DNA or RNA, and the adaptor comprises a nucleic acid barcode sequence unique to the non-canonical feature.
[0010] The present invention also provides a composition comprising: i) a substrate; ii) a second recognition element bound to the substrate; iii) an adaptor that binds to the second recognition element; and iv) a binding domain, wherein the binding domain is configured to specifically bind to a non-canonical feature of DNA or RNA, and the binding domain is immobilized by the second recognition element, wherein the adaptor comprises a nucleic acid barcode sequence unique to the non-canonical feature. In one aspect, the composition comprises a plurality of second recognition elements, wherein the plurality of second recognition elements comprises second recognition elements that are different from each other, the adaptor is bound to one of the plurality of second recognition elements, and the binding domain is bound to a different second recognition element. In one aspect, the composition comprises a plurality of second recognition elements, wherein the adaptor is bound to one of the plurality of second recognition elements, and the binding domain is bound to another instance of the same second recognition element.
[0011] Also provided herein is a method for determining whether a target molecule comprises: i) a substrate; ii) a second recognition element bound to the substrate; Also provided is a composition comprising iii) a binding domain that binds to a substrate via a linker; and iv) an adaptor that binds to the substrate via a second recognition element, where the binding domain is configured to specifically bind to a non-canonical feature of DNA or RNA, and the adaptor comprises a nucleic acid barcode sequence unique to the non-canonical feature.
[0012] The invention also provides a composition comprising: i) a substrate; ii) a binding domain that binds to the substrate via a first linker or a second recognition element; iii) a mosaic end (ME) adapter that binds to the substrate via a second linker or a second recognition element; and iv) a transposase, wherein the transposase is loaded onto an immobilized ME adapter, the binding domain specifically binds to a non-canonical feature of DNA or RNA, and at least one of the ME adapters comprises a nucleic acid barcode sequence unique to the non-canonical feature; or a composition comprising: i) a substrate; ii) a binding domain that binds to the substrate via a linker or a second recognition element; and iii) a transposase bound to the binding domain, wherein the transposase is loaded onto an ME adapter, the binding domain specifically binds to a non-canonical feature of DNA or RNA, and at least one of the ME adapters comprises a nucleic acid barcode sequence unique to the non-canonical feature.
[0013] The invention also provides a composition comprising: i) a substrate; ii) a plurality of second recognition elements bound to the substrate; iii) an adaptor bound to one of the plurality of second recognition elements; and iv) a binding domain bound to another one of the plurality of second recognition elements, where the binding domain specifically binds to a non-canonical feature of DNA or RNA, and the adaptor comprises a nucleic acid barcode sequence unique to the non-canonical feature specifically bound by the binding domain.
[0014] The invention also provides a complex comprising one or more of the compositions comprising a binding domain described herein and a target nucleic acid bound to the binding domain.
[0015] Also provided are methods for making the compositions and composites disclosed herein and depicted in the drawings.
[0016] The present invention also provides a method for analyzing a plurality of target nucleic acids, the method comprising: (i) contacting a solution containing a plurality of target nucleic acids with a composition described herein, wherein the target nucleic acid containing a non-canonical feature binds to the binding domain; (ii) performing any of the following steps: (a) transferring a nucleic acid barcode to the target nucleic acid containing a non-canonical feature to generate a barcoded target nucleic acid, or (b) generating a barcoded copy of the target nucleic acid containing a non-canonical feature; (iii) amplifying the barcoded target nucleic acid; and (iv) determining the sequence of the barcoded target nucleic acid, wherein steps (i) and (ii) are performed sequentially or simultaneously. In one aspect, an adaptor with a 3' degenerate base randomly primes the target nucleic acid. In one aspect, step (ii) further comprises introducing a modification-specific barcode, wherein the 3' end of the adaptor is extended by reverse transcriptase or DNA polymerase.
[0017] The present invention also provides a method for analyzing multiple target nucleic acids, the method comprising: (i) contacting a solution containing multiple target nucleic acids with a composition described herein, wherein the target nucleic acids containing non-canonical features bind to the binding domain; (ii) performing one of the following: (a) transferring a nucleic acid barcode to the target nucleic acid containing non-canonical features to generate a barcoded target nucleic acid, or (b) generating a barcoded copy of the target nucleic acid containing non-canonical features; (iii) amplifying the barcoded target nucleic acid; and (iv) sequencing the barcoded target nucleic acid, wherein steps (i) and (ii) are performed sequentially or simultaneously. In one aspect, an adapter having a 3' spacer sequence is site-specifically bound to a synthetic spacer sequence exhibited by the target nucleic acid. In one aspect, step (ii) further comprises introducing a modification-specific barcode, wherein one or both 3' ends are extended by reverse transcriptase or DNA polymerase.
[0018] The present invention also provides a method of analyzing a plurality of target nucleic acids, the method comprising: (i) contacting a solution containing the plurality of target nucleic acids with a composition described herein, wherein the target nucleic acids comprising the non-canonical feature bind to a binding domain; (ii) performing one of the following: (a) transferring a nucleic acid barcode to the target nucleic acid comprising the non-canonical feature to generate a barcoded target nucleic acid, or (b) generating a barcoded copy of the target nucleic acid comprising the non-canonical feature; (iii) amplifying the barcoded target nucleic acid; and (iv) sequencing the barcoded target nucleic acid, wherein steps (i) and (ii) are performed sequentially or simultaneously.
[0019] The present invention also provides a method of analyzing a plurality of target nucleic acids, the method comprising the steps of: (i) providing a plurality of target nucleic acids by reverse transcribing a target RNA molecule to form a DNA-RNA heteroduplex molecule or providing a target double-stranded DNA molecule; (ii) contacting a solution containing the plurality of target nucleic acids with a composition described herein, wherein the target nucleic acid comprising a non-canonical feature is bound to a binding domain; (iii) using a transposase to transfer two adapters, at least one of which comprises a nucleic acid barcode, to the double-stranded target nucleic acid comprising the non-canonical feature to generate a barcoded target nucleic acid; (iv) amplifying the barcoded target nucleic acid; and (v) sequencing the barcoded target nucleic acid, wherein steps (ii) and (iii) are performed simultaneously or sequentially.
[0020] The invention also provides a method of detecting a plurality of non-canonical features in a plurality of target nucleic acids, the method comprising: (i) contacting a solution comprising a plurality of target nucleic acids with a plurality of compositions described herein, wherein the number of compositions contacted in step (i) is equal to or greater than the number of non-canonical features, wherein the binding domains of the plurality of compositions each bind to a different non-canonical feature of DNA or RNA, or the binding domains of the plurality of compositions bind to the same non-canonical feature of DNA or RNA, and wherein the adapters of the plurality of compositions each comprise a nucleic acid barcode sequence unique to, or unique to, the non-canonical feature specifically bound by the binding domain of that composition; (ii) performing one of the following: (a) transferring the nucleic acid barcode sequence of each of the plurality of compositions to the plurality of target nucleic acids, or (b) generating barcoded copies of the plurality of target nucleic acids; (iii) amplifying the barcoded target nucleic acids; and (iv) sequencing the barcoded target nucleic acids. In one aspect, adapter introduction by a transposase is also included.
[0021] The present invention also provides a method for detecting a plurality of non-canonical features in a plurality of target nucleic acids, the method comprising the steps of: (i) providing a microarray, bead, and / or fluidic device comprising a plurality of compositions as described herein, wherein the number of compositions provided in step (i) is equal to or greater than the number of non-canonical features, and wherein the binding domains of the plurality of compositions each bind to a different non-canonical feature of DNA or RNA, or the binding domains of the plurality of compositions each bind to the same non-canonical feature of DNA or RNA, and the binding domains of the plurality of compositions each bind to the same non-canonical feature of DNA or RNA, and the binding domains of the plurality of compositions each bind to a different non-canonical feature of DNA or RNA. (ii) contacting a plurality of target nucleic acids with a plurality of compositions and performing one of the following: (a) transferring the nucleic acid barcode sequence of each of the plurality of compositions to the plurality of target nucleic acids, or (b) generating barcoded copies of the plurality of target nucleic acids; (iii) amplifying the barcoded target nucleic acids; and (iv) sequencing the barcoded target nucleic acids. In one aspect, adapter introduction by a transposase is also included.
[0022] These and other aspects of the present invention may become evident upon reference to the following detailed description, drawings, claims, embodiments, procedures, compounds, and / or compositions, as well as the associated background information and references. [Brief description of the drawings]
[0023] [Figure 1]1A-1H show various molecular structures for attaching DNA adapters and binding domains to a surface (e.g., a substrate). In an exemplary embodiment, the DNA adapter contains an RNA modification-specific barcode that is transferred to a target RNA for the purpose of identifying the modification. In FIG. 1A, both the DNA adapter and the binding domain are covalently attached to the surface by the same or orthogonal chemical reaction. A linker may be included to increase the flexibility and accessibility of the surface-attached molecule. In FIG. 1B, the adapter is attached to a second recognition element that binds to the RNA-specific binding domain, such as an RNA-specific primary antibody immobilized via protein G, A, or L, or a secondary antibody. FIG. 1C shows the use of a layer of biotinylated adapter molecules bound to a different second recognition element, such as streptavidin, for adapter immobilization. In this example, the binding domain is immobilized via a linker. Alternatively, both the binding domain and the adapter may be immobilized via streptavidin, or the adapter may be covalently attached to the substrate while the binding domain is bound to the second recognition element. FIG. 1D shows the use of two different second recognition elements to immobilize the binding domain and the adapter. For example, an antibody binding domain is immobilized to a substrate via protein G, and a biotinylated adapter is immobilized to a substrate via streptavidin. FIG. 1E shows two antibodies immobilized to a substrate via protein G. One antibody species is labeled with an adapter and does not bind to nucleic acids, while the other antibody species is specific for a non-canonical feature of the nucleic acid and is unlabeled. FIG. 1F shows the immobilization of the binding domain to a surface, where the binding domain is bound to a nucleic acid complementary to the capture sequence. When the capture sequence is immobilized on the substrate (e.g., via a linker), it hybridizes with the nucleic acid sequence bound to the binding domain, resulting in the immobilization of the binding domain on the substrate. In this example, the adapter contains a cleavage site for enzymatic binding of the target RNA or cDNA to the surface-bound adapter followed by release from the surface.Cleavage can occur within a uracil-modified adaptor using the USER enzyme, or within an 8-oxoguanine-modified adaptor using the FpG enzyme, or it can be part of a linker, such as a photocleavable PC or disulfide linker. Figure 1G is a substrate showing a mosaic end (ME) adaptor for transposition adjacent to a binding domain. Each Tn5 transposase dimer contains two adaptor molecules. DNA library preparation by tagmentation includes a Tn5 dimer containing an ME adaptor with forward and reverse primer sites, respectively. Figure 1H shows another method of linking a Tn5 molecule adjacent to an antibody binding domain. A dimer of Tn5-Protein A fusion protein is loaded onto the ME adaptor and binds to an antibody via affinity binding between Protein A and the Fc region of the antibody. [Diagram 2]Figures 2A-2G show various methods for attaching an adapter to an RNA molecule or its corresponding cDNA. Figure 2A shows ligation between the 3'OH of RNA (acceptor) and the 5'-phosphate of DNA or RNA (donor) catalyzed by T4 RNA ligase 1. Related exemplary formats include ligation of a preadenylated RNA or DNA donor to an RNA acceptor by T4 RNA ligase 2, and ligation of the 3'phosphate of RNA to the 5'OH of RNA by RtcB ligase. Two single-stranded DNA fragments can be ligated by CircLigase. Figure 2B shows ligation of a nick structure by T4 RNA ligase 2. Both the donor and acceptor can be RNA, and the donor can be DNA. The nick in the double-stranded DNA is sealed by T4 DNA ligase. Figure 2C shows splint extension using reverse transcriptase with the target RNA as a template. In this format, barcoded cDNA is generated. In FIG. 2D, the target RNA acts as a primer and a barcode is added. Extension by DNA polymerase requires ligation of a short spacer sequence (SP) with a known sequence. In one aspect, the invention also includes methods that include multiple successive barcode transfers, for example, where the barcode is directly attached to the target nucleic acid. An adapter with two spacer regions as shown in FIG. 2D is an example of an adapter suitable for such an iterative barcoding process. Reverse transcriptase can extend the adapter as shown in FIG. 2G, thereby synthesizing a cDNA copy of the RNA target. FIG. 2E shows barcoding by double-stranded ligation of blunt or sticky-ended DNA by T4 DNA ligase. FIG. 2F shows chemical ligation that occurs between two parts A and B. Part A is part of a short spacer that is ligated to the RNA target to make it suitable for chemical ligation. FIG. 2H is similar to FIG. 2G, but does not rely on ligation of a spacer sequence to the RNA target.The 3' end of the adapter displays degenerate bases to allow random priming of the target RNA, and the barcode is then transferred by single or bidirectional primer extension. [Diagram 3] Figure 3 shows a general overview of RNA profiling using artificial surfaces (e.g. beads). Multiple RNA strands are chemically fragmented. Modified RNA fragments (modifications shown as hexagons) are enriched on the surface via interactions with RNA modification-specific binding domains. Multiple beads can also be used, with each bead displaying the same binding domain and a copy of the barcode. Any number of different beads can be used in the reaction to interrogate the RNA modifications. The RNA modifications are converted into a DNA code by transferring the barcode to the target RNA. The cDNA library is sequenced to determine the modification state of each RNA fragment.
[0024] [Figure 4]Figures 4A-4C show several surface-based assay formats for simultaneously interrogating multiple non-canonical features (e.g., RNA modifications) on different strands in the same reaction. These formats aim to spatially separate different types of binding domains and associated barcodes and expose them to the same analyte to enable multiplexed analysis. Figure 4A shows different types of beads combined in a pool, where each type of bead captures and barcodes a specific RNA modification. The beads can be filtered through a frit column or retrieved by magnetization. Figure 4B shows the use of a DNA array for surface-mediated barcoding and capture of binding domains. Each spot of the array contains at least one unique barcoded adapter and captures only one type of binding domain by hybridization with the DNA tag displayed by the binding domain. In Figure 4C, monoclonal patches with co-immobilized barcodes and binding domains are incorporated into a microfluidic chip forming individual channels. Each channel contains an immobilized barcode and a binding domain for one DNA / RNA modification or non-canonical feature. The analyte is delivered by sample partitioning. [Diagram 5] Figure 5 shows the complete RNA modification profiling workflow utilizing 3'-immobilized adapters and barcoding by ligation. The workflow steps include modification-specific RNA capture, barcoding by single-strand ligation, first-strand cDNA synthesis, and second-strand synthesis by template switching. The DNA adapter contains a 3' amine for surface immobilization, a universal priming site, a unique molecular barcode, a modification-specific barcode, and a 5' phosphate. [Figure 6]Figure 6 shows a complete RNA modification profiling workflow utilizing 5' immobilized adapters and barcoding by primer extension. In this non-limiting example, a short spacer (SP) is ligated upstream. The spacer is complementary to the surface-bound adapter, and annealing of the RNA target to the surface-bound adapter forms a priming site for reverse transcriptase. To ensure RNA modification-specific pull-down of the target, the spacer interaction is weak and not stable alone in the absence of antibody binding. Simultaneous antibody and spacer binding is shown. The DNA adapter contains a 5' amine for surface immobilization, a universal priming site, a unique molecular barcode, a modification-specific barcode, and a 3' spacer. Extension of the surface-bound adapter by reverse transcriptase in the presence of template switching oligos generates barcoded first strand cDNA, a second sequencing adapter is introduced, and the cDNA is covalently attached to the surface. Amplification of the cDNA is performed in a separate reaction by bead-input PCR or in situ on the surface (shown in Figures 7 and 8, respectively). [Figure 7]Figure 7 shows a schematic of surface-based cDNA amplification that can be used to generate clusters of identical copies of a target nucleic acid on a substrate. Similar to solution PCR, the process uses temperature cycles to anneal, extend, and melt DNA strands, resulting in exponential amplification. Surface-based amplification generates monoclonal clusters of identical copies of the first cDNA strand. Each cluster is seeded by a binding domain that is bound to the substrate and recognizes a non-canonical feature. The surface density of the binding domains is sparse to avoid polymerization of neighboring clusters. The first cDNA strand is generated according to the workflow shown in Figure 6, using a surface displaying P5 and P7 primers. At low temperature, the cDNA strand is annealed to the complementary surface primer. When DNA polymerase extends the primer at intermediate temperature, a copy of the parent strand is generated. The resulting duplex is separated by heat and / or by the addition of a chaotrope, which serves as the starting point for the next cycle. There may be one or multiple clusters of identical copies. The methods of the invention involve in situ sequencing of clusters of identical copies of a target nucleic acid on a substrate.
[0025] [Figure 8] Figure 8 illustrates the process of generating monoclonal cDNA clusters suitable for sequencing by synthesis, with each cluster representing a modified RNA strand. The fragmented RNA is partitioned and seeded onto a flow cell based on the interaction of the RNA modification with a binding domain (see Figure 6). The flow cell is segmented, with each segment targeting a different modification. For example, to detect 10 modifications, the flow cell contains 10 modified regions with appropriate binding domains and adapter pairs. The surface density of antibodies is low to prevent contamination with neighboring sequences during amplification. RNA strands are captured based on their modifications, covalently attached to the surface, and barcoded by primer extension, followed by clonal amplification (see Figure 7). The clonally amplified barcoded cDNA is then linearized and directly sequenced using sequencing-by-synthesis (SBS) chemistry. [Figure 9] Figure 9 shows a method for rapid profiling of RNA modifications using Tn5 transposase for barcoding. RNA is reverse transcribed into DNA / RNA heteroduplexes. The heteroduplexes are immunoprecipitated onto a surface (e.g., beads) presenting an antibody and an adapter with mosaic ends (ME). Transposomes containing Tn5 transposase molecules bound to ME adapters are assembled, and Tn5 transposase inserts the barcoded adapters in a one-step cut-and-paste mechanism in the presence of Mg2+ ions. Gap-fill followed by PCR completes the library preparation workflow. [Figure 10] Figure 10 shows the process of position marking of multiple m6A modifications in the same RNA strand by base editing with ADAR enzymes. After position marking, individual RNA strands are barcoded by transposase as shown in Figure 9. NGS (next generation sequencing) reads derived from the same parent RNA molecule share the same barcode. "A>I" refers to the mutation of adenine to inosine catalyzed by ADAR enzymes. [Figure 11] Figure 11 shows the concept of the long read phase. Position marking and barcoding by the process shown in Figure 10 allows the reconstruction of long transcripts from short sequence reads. To uniquely barcode each short nucleic acid fragment derived from the same parent molecule, each bead displays multiple unique barcodes representing RNA modifications and individual beads. The surface of the beads is small and on average can capture only one full-sized parent molecule. The immobilized transposomes cleave the parent molecules into short fragments and insert specific barcodes on the beads. The short reads are aligned to the reference genome and joined at junctions that display the same barcode. “A>I” means adenine to inosine mutation. [Figure 12]Figures 12A-12D show the structure of various DNA adapters. Figure 12A shows adapters containing UFPs or URPs. Figure 12B shows adapters that can be used for library preparation by circularization. Figure 12C shows adapters that can be used for barcode introduction by ligation. Figure 12D shows adapters that can be used for single or multiple barcode transcription by primer extension. The spacer can be a specific base sequence or can be composed of random bases. As shown in the legend, "UFP" is an abbreviation for universal forward primer, "URP" is an abbreviation for universal reverse primer, "MBC" is an abbreviation for modified encoded barcode, "UMI" is an abbreviation for unique molecular identifier, and "CLS" is an abbreviation for cleavage site. "SP" is an abbreviation for spacer. [Figure 13] FIG. 13 shows an exemplary mosaic end adapter molecule (ME and ME'). The grey lines are portions of DNA and the sequences are the ME and ME' adapters. Each transposase loads two adapters (in this example, Tn5ME- / ME and Tn5ME-B / ME), which are ligated to both ends of the ds-DNA. The following sequences are shown: SEQ ID NO: 14 GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG 15 TCTACACATATTCTCTGTC 16 CTGTCTCTTATACACATCT 17 GACAGAGAATATGTGTAGACTGCGACGGCTGCT
[0026] [Figure 14] Figures 14A-14G show binding and dissociation curves of different antibodies measured by Biolayer Interferometry (BLI). The solid lines show the binding of the antibodies to a modified RNA oligo with a central base modification. The dotted lines show the binding of the antibodies to an unmodified RNA oligo of the same length. [Figure 15]Figure 15A shows the products generated by labeling a reporter antibody with different amounts of adapter oligo and separating them by denaturing gel electrophoresis. The reaction produces a distribution of labeled stoichiometries. The average labeling stoichiometry increases as the molar excess of oligo to antibody increases. Figure 15B shows that the barcoding rate increases as the number of reporter antibodies on the surface increases. Data are obtained by loading a mixture of RNA modification-specific and reporter antibodies onto the beads and then immunoprecipitating the modified RNA with a terminal dye label to initiate the barcoding reaction. Barcoding is quantified by denaturing gel electrophoresis of the eluted RNA and densitometry of the gel bands. Figure 15C shows a schematic of the configuration of "monoclonal" beads. Monoclonal beads display a single RNA modification-specific antibody and a single adapter sequence representing that antibody. To effectively barcode the immunoprecipitated RNA, the adapter must be present at a density that allows interaction of the RNA with the adapter molecule. [Figure 16] Figure 16 shows the results of analysis of fragmented RNA before and after spacer ligation using capillary electrophoresis. The fragment sizes are normally distributed around 104 and 109 nucleotides, respectively. [Figure 17]Figure 17A shows the molecular structure of the barcode by reverse transcription. Three types of beads are included in the 3-plex experiment. One bead type shows the m6A antibody and m6A-specific adapter (MBC3-Ab05(m6A)), a second bead type shows the inosine antibody and inosine-specific adapter (MBC4-Ab10(I)), and a third bead type shows the m5C antibody and m5C-specific adapter (MBC5-Ab16(m5C)). Spacer hybridization (SP-SP') between the target RNA and the adapter allows bidirectional extension by reverse transcriptase to copy the modified barcode (MBC) and produce cDNA. The inclusion of a template switching oligo (TSO) in the reverse transcription reaction results in the attachment of the second sequencing adapter. Figures 17B and 17C summarize the sequencing results obtained in a 3-plex experiment using modified RNA obtained by in vitro transcription (IVT) from four different genomes in the presence of the indicated modified nucleotides. The experiment summarized in Figure 17B used SuperScript IV reverse transcriptase, whereas Figure 17C used Maxima Minus reverse transcriptase. The normalized fraction of each MBC is plotted for each genome and the modifications are indicated. [Figure 18] Figures 18A-18G show the sequencing results of a singleplex experiment using a single bead type and a target pool containing modified IVT RNA from four different genomes. The purpose of the experiment was to compare the efficiency of barcoding with different antibodies. The MBC fraction associates the RNA modification with the correct genome. The antibodies are shown above the plot along with the modified target. [Figure 19] Figure 19A shows the nucleic acid architecture required for barcoding with DNA polymerase. The bead nomenclature is the same as in Figure 17A, but the 3' end of the adapter is blocked to prevent bottom strand extension (light grey dot). Figure 19B reports the associated sequence data. RNA modifications are represented in the majority of barcodes.
[0027] [Figure 20]FIG. 20A shows splint ligation as a barcoding method. A splint (black line) bridges the RNA target and the adapter. A ligase fills the gap and ligates the adapter to the RNA target. Two beads are shown, targeted to m6A and m5C. FIG. 20B-20C summarize the corresponding sequencing data showing the simultaneous detection of m6A and m5C. The portion of the splint hybridizing to the adapter is 7 nt in length, whereas the portion facing the RNA target is 6 nt (7-6 splint) or 3 nt (7-3 splint). The following sequences are shown: SEQ ID NO: 18 AAAGCTGCACTCA / 3SpC3 / 19 ATATAGGCACTCA / 3SpC3 20 AAAGCTGCAC / 3SpC3 / 21 ATATAGGCAC / 3SpC3 / [Figure 21] Figure 21A shows an alternative method of ligating a universal spacer to an RNA target. For barcoding by primer extension, the RNA is A-tailed (polyA tail (AAAAAAAAAAAAAAA (SEQ ID NO: 22)) and hybridized to an adapter sequence ending with the sequence NVTTTTTTTT. Reverse transcription and template switching are performed as above. Figure 21B shows the proof of concept of a singleplex data set. [Figure 22]Figure 22A shows a rapid method to profile RNA modifications using Tn5 transposase for barcoding. RNA is reverse transcribed into DNA / RNA heteroduplexes. A surface (e.g., beads) contains an antibody and ME adapters carried on it. The heteroduplexes are immunoprecipitated onto the surface presenting the antibody and adapters. After washing the beads, Tn5 transposase is loaded onto the ME adapters in the absence of Mg2+. Tagmentation buffer containing Mg2+ is then added to insert the adapters into the DNA-RNA duplexes securely captured on the beads. Gap filling followed by PCR completes the library preparation workflow. Figure 22B shows a coverage plot obtained in an experiment using m6A-specific beads and a target pool containing modified IVT RNA from four different genomes. The plot shows a significant enrichment of m6A-containing fragments, proving selective tagmentation of m6A-modified RNAs. [Diagram 23] Figure 23A shows the global barcode representation measured in MBC fractions for technical triplicates of barcoded IP RNA and non-enriched (input) samples. Figure 23B shows the location of called peaks within genes. Figure 23C shows the number of called peaks for each modification and each replicate sample in a Venn diagram. [Figure 24]Figure 24A shows a method for tagmentation of DNA / RNA heteroduplexes using immobilized conjugates containing antibodies and protein A-Tn5 fusion proteins. Protein G is bound to a surface (e.g., beads) to bind conjugates containing antibodies and protein A-Tn5 molecules. Each Tn5 dimer carries a pair of mosaic end (ME) adapters containing a barcode. To create a substrate for transfer, RNA is reverse transcribed into DNA / RNA heteroduplexes and immunoprecipitated onto beads. The beads are washed and Mg2+-containing tagmentation buffer is added to initiate the tagmentation reaction. The tagged DNA / RNA heteroduplexes are gap-filled and PCR amplified. Library preparation is then performed to complete the workflow. Figure 24B shows a comparison of read coverage plots for input (control) and immunoprecipitated samples from an experiment targeting m6A. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0028] Detailed Description The present invention provides compositions and methods for multiplex profiling of RNA and DNA modifications across transcriptomes and genomes. The methods combine molecular recognition of non-canonical features of target nucleic acids (e.g., base modifications, backbone modifications, lesions, and / or structural elements) with writing information from this recognition event into the adjacent genetic sequence of the target nucleic acid using barcodes. The resulting barcoded nucleic acids are then converted into a sequencing library and read, for example, by DNA / RNA sequencing or other methods. This process reveals the sequence of the barcode, which correlates with the non-canonical features of the target nucleic acid. Sequencing can also confirm the localization of the non-canonical features in the target nucleic acid. The rapid profiling methods described herein allow the nature and location of multiple or all DNA / RNA modifications to be identified in parallel. These methods can also determine the abundance and stoichiometry of DNA / RNA modifications.
[0029] In some embodiments, the methods described herein are used not only to identify modifications on a target nucleic acid, but also to localize modifications on a target nucleic acid with resolution as high as one base.
[0030] The present invention will be described in more detail below in an illustrative and non-limiting manner and with reference to the accompanying drawings. However, the present invention can be embodied in many different forms and should not be construed as being limited to the following embodiments. Rather, these embodiments are provided so that the present specification will be complete and will convey the scope of the present specification to those skilled in the art.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention belongs. The terms used in the detailed description of this specification are for the purpose of describing particular embodiments only and are not intended to be limiting.
[0032] All publications, patent applications, patents, GenBank / Uniprot or other accession numbers, and other references mentioned herein are incorporated by reference in their entirety for all purposes.
[0033] definition The following terminology is used in the specification and appended claims.
[0034] The singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0035] Additionally, as used herein, when referring to measurable values such as length of polynucleotide or polypeptide sequences, dosage, time, temperature, and the like, the term "about" is meant to encompass variations of ±20%, ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of the stated amount.
[0036] Also, as used herein, "and / or" refers to and includes any possible combination of one or more of the associated listed items, as well as, when interpreted in the alternative, the absence of a combination ("or").
[0037] Unless otherwise stated, it is specifically intended that the various features described herein can be used in any combination. Moreover, in some embodiments, any of the features or combinations of features described herein can be excluded or omitted. To further explain, for example, when the present specification indicates that a particular DNA base can be selected from A, T, G and / or C, the term also indicates that the base can be selected from any subset of these bases, such as A, T, G or C; A, T or C; T or G; C only, etc., and each such subcombination is described as if expressly described herein. Furthermore, the term also indicates that one or more of the specified bases can be excluded. For example, in some embodiments, the nucleic acid is described as not A, T or G; not A; not G or C, etc., and each such possible exclusion is described as if expressly described herein.
[0038] As used herein, the terms "reduce," "reduce," "decrease," and similar terms refer to a decrease of at least about 10%, about 15%, about 20%, about 25%, about 35%, about 50%, about 75%, about 80%, about 85%, about 90%, about 95%, about 97% or more.
[0039] As used herein, the terms "increase," "improvement," "enhancement," "enhancement," and similar terms refer to an increase of at least about 10%, about 15%, about 20%, about 25%, about 50%, about 75%, about 100%, about 150%, about 200%, about 300%, about 400%, about 500% or more.
[0040] The term "epigenetic change" is used herein to mean a phenotypic change in a living cell, organism, etc. that is not encoded in the primary sequence (i.e., A, T, C, and G) of the DNA of that cell or organism. Epigenetic changes can include, for example, chemical changes to nucleotides and / or histones (i.e., proteins involved in coiling and packaging of DNA in the nucleus). Exemplary DNA nucleotide modifications include the common epigenetic marker 5-methylcytidine (5mC) and its oxidation products 5-hydroxymethylcytidine (5hmC), 5-formylcytidine (5fC), 5-carboxymethylcytidine (5caC), etc. While 5mC is best known for its role in gene silencing, there is growing evidence suggesting metabolic functions for the oxidative intermediates 5hmC, 5fC, and 5caC in the demethylation pathway of 5mC. Further metabolically relevant DNA modifications include oxidation, alkylation, dimerization, crosslinking, and other chemically modified nucleotides associated with DNA damage. Although such DNA modifications are important for understanding toxicity, their distribution throughout the genome upon damage is poorly understood and they may play additional regulatory roles, for example by participating in G-quadruplex dynamics in promoters and other regions of the genome.
[0041] As used herein, the term "epitranscriptome changes" refers to chemical modifications of RNA that occur during or after transcription. Over 170 different RNA modifications are known, including chemical changes to nucleobases, ribose, and phosphodiester backbones. RNA modifications are found in all RNA types, including mRNA, tRNA, rRNA, lncRNA, and miRNA, and can alter cellular phenotypes by altering RNA structure and dynamics, or by altering the molecular recognition of RNA by other biomolecules, such as proteins. Natural chemical RNA modifications in the epitranscriptome control a wide range of functions in RNA metabolism, including RNA processing, splicing, polyadenylation, editing, structure, stability, localization, translation initiation, and gene expression. The epitranscriptome varies across cell types, metabolic states, and health states, and plays a critical (but poorly understood) role in cellular phenotypic and functional differentiation, helping to explain dramatic phenotypic differences between cells of the same organism that have identical primary genetic sequences. Epitranscriptome changes correlate with disease. For example, mRNA and ncRNA modifications are known to control spatiotemporal changes in gene expression during cancer stem cell differentiation, thereby playing an orchestrated role in disease progression. Furthermore, RNA modifications are strongly suspected to be an important mechanism by which RNA viruses (e.g., Coronaviridae and Flaviviridae) outwit their hosts and evade the innate immune system.
[0042] The term "genome" refers to all the DNA contained in a cell or cell population, or a selection of a particular type of DNA molecule (e.g., coding DNA, non-coding DNA, mitochondrial DNA, or chloroplast DNA). The term "transcriptome" refers to all the RNA molecules produced in a cell or cell population, or a selection of a particular type of RNA molecule contained in the complete transcriptome (e.g., mRNA vs. ncRNA, or a particular mRNA within an mRNA transcriptome). In some embodiments, the transcriptome includes multiple different types of RNA, such as coding RNA (i.e., RNA that is translated into protein, e.g., mRNA) and non-coding RNA. A non-limiting list of different types of RNA molecules found in the transcriptome include those that may contain modified nucleosides: 7SK RNA, signal recognition particle RNA, antisense RNA, CRISPR RNA, guide RNA, long non-coding RNA, microRNA, messenger RNA, piwi-interacting RNA, repeat-associated siRNA, retrotransposons, ribonuclease MRP, ribonuclease P, ribosomal RNA, small Cajal body-specific RNA, small interfering RNA, smY RNA, small nuclear RNA, and trans-acting siRNA.
[0043] As used herein, the term "non-canonical feature" of a nucleic acid refers to a feature of a nucleic acid that is distinct from its primary sequence. For example, a non-canonical feature can be a DNA or RNA base, or a chemical modification to the DNA or RNA backbone. In some embodiments, a non-canonical feature can be a structural sequence such as a hairpin or loop. Other exemplary non-canonical structures include, but are not limited to, Z-DNA structures, G-quadruplexes, i-motifs, bulges, abasic sites, triplexes, three-way bonds, cruciform structures, quadruple loops, ribose zippers, pseudoknots, and the like. Nucleic acids, including DNA and RNA, contain many non-canonical features. The frequency of these modifications varies widely depending on the RNA and type of feature, although clusters of modifications may occur. In some embodiments, a non-canonical feature can result from DNA and / or RNA damage. As used herein, the terms "non-canonical feature" and "modification" can be used interchangeably, as would be understood by one of skill in the art in the context.
[0044] As used herein, the term "target nucleic acid" refers to a nucleic acid that contains one or more non-canonical features. The nucleic acid binding molecules described herein can bind to a target nucleic acid when the binding domain of the molecule recognizes the non-canonical feature.
[0045] As used herein, the term "substrate" is used to mean any solid support. For example, a substrate can be a bead, a chip, a plate, a slide, a dish, a gel, a tube, a flow cell, a matrix, an array, a microfluidic device or component thereof, a well, a cartridge, or a three-dimensional matrix. As described herein, a binding domain described herein can be bound to one or more substrates, and a substrate can be bound to one or more binding domains. Additionally, an adaptor described herein can be bound to one or more substrates, and a substrate can be bound to one or more adaptors. Substrates can be formed of a variety of materials. In some embodiments, a substrate is a resin, a membrane, a fiber, a polymer. In some embodiments, a substrate comprises sepharose, agarose, cellulose, polystyrene, polymethacrylate, and / or polyacrylamide. In some embodiments, a substrate comprises a polymer, such as a synthetic polymer. A non-limiting list of synthetic polymers includes poly(ethylene) glycol, polyisocyanopeptide polymers, polylactic-co-glycolic acid, poly(ε-caprolactone) (PCL), polylactic acid, poly(3-hydroxybutyrate-co-3-hydroxyvalerate) (PHBV), chitosan and cellulose.
[0046] The term "barcode" as used herein refers to a synthetically produced nucleic acid. A unique barcode may be assigned to a particular nucleic acid modification to allow for specific identification of those modifications in the methods described herein. Thus, a barcode is "unique" to a non-canonical modification when it is specifically used to identify that modification in one or more of the methods described herein. Barcodes can be produced using methods known in the art, such as solid-phase oligonucleotide synthesis. In some embodiments, the barcode may be a DNA barcode (i.e., may include a DNA sequence). In some embodiments, the barcode may include a synthetic DNA structure, such as a peptide nucleic acid (PNA) or a locked nucleic acid (LNA). In some embodiments, the synthetic DNA structure may include one or more modified bases. In some embodiments, the synthetic DNA structure may include one or more modified bases. In some embodiments, the barcode may be an RNA barcode (i.e., may include an RNA sequence). Barcodes may be of any length, for example, ranging from about 4 to about 150 nucleotides in length. In some embodiments, the barcode is about 4 to about 20 nucleotides in length, e.g., about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length. Typically, the barcode may comprise a rationally designed sequence that is not in the genome of a known organism. However, in some embodiments, the barcode may comprise a known sequence. For example, the barcode sequence may comprise a signature associated with a pathogen or other biological material. In some embodiments, the barcode may comprise a sequence configured to facilitate a sequencing reaction. The terms "barcode" and "adapter" may be used interchangeably herein. As understood in the art, an adapter may be comprised of a barcode in some embodiments.In some embodiments, the adapter can include a barcode and one or more additional elements, as described below and shown in Figures 12A-12D.
[0047] The term "amplification" when used in reference to a nucleic acid means to generate copies of that nucleic acid. Nucleic acids can be amplified, for example, using polymerase chain reaction (PCR). Alternative methods of nucleic acid amplification include helicase-dependent amplification (HAD), recombinase polymerase amplification (RPA), loop-mediated isothermal amplification (LAMP), nucleic acid sequence-based amplification (NASBA), self-sustained sequence replication (3SR), and rolling circle amplification (RCA).
[0048] As used herein, the term "coupled" may be used to describe two or more components that are associated with one another, such as a first component that is bound to a second component, either covalently or non-covalently bonded or otherwise attached.
[0049] As used herein, the term "intracomplex adapter transfer" or "intracomplex barcode transfer" refers to the transfer of an adapter and / or barcode to a target nucleic acid (e.g., DNA or RNA) while the binding domain and adapter are bound. Thus, the term "complex" as used herein refers to the complex formed between a target nucleic acid, a binding domain, and its cognate adapter.
[0050] As used herein, the terms "crosstalk," "barcode crosstalk," and similar terms refer to off-target transfer of nucleic acid barcodes. For example, barcode crosstalk can occur when the barcode of an adapter is transferred to a nucleic acid that is not bound to the binding domain of a nucleic acid binding molecule.
[0051] The term "DNA address" refers to a DNA or RNA sequence and / or its complement that is used as a programmable binding element to promote a specific binding event. For example, a deaminase can bind to and direct a DNA or RNA sequence (i.e., a first DNA address) that binds to a target DNA or RNA sequence (e.g., a second DNA address) to which the deaminase can be directed.
[0052] “Nucleic acid damage”, such as “DNA damage” or “RNA damage”, is a chemical modification of nucleic acids that can occur as a result of endogenous processes and / or exogenous agents. For example, DNA damage can be caused by oxidative damage (e.g., 8-oxoguanine), reaction with electrophiles and alkylating agents such as those found in charred meat and cigarette smoke (benzo[a]pyrene adducts and alkylated nucleobases), UV damage (cyclobutane pyrimidine dimers and 6-4 pyrimidine-pyrimidine photoproducts), metal complex formation (mercury complexes and platinum bridges), etc. DNA damage caused by endogenous processes occurs frequently. DNA damage is usually repaired by various repair enzymes or bypassed by damage-bypassing polymerases during replication of the genetic code. Mutations that result in unnatural cell proliferation and growth drive cancer development. Although mutations can be easily detected by routine DNA sequencing, the damage itself cannot be detected by standard DNA sequencing workflows. Damage is not uniformly distributed throughout the genome, and the effect of repair is related to DNA locus and cellular state. Furthermore, the most common cancer chemotherapy drugs (e.g., cisplatin, gemcitabine) induce DNA damage, therefore mapping DNA damage throughout the human genome offers great potential for understanding the pathogenesis of aging and cancer, improving the efficacy and reducing the toxicity of cancer chemotherapy drugs.
[0053] Surface Structure and Composition Described herein are compositions comprising adaptors and binding domains for identifying non-canonical features on nucleic acids. The compositions described herein comprise distinct surface structures of the binding domains and adaptors spatially separated on a substrate.
[0054] In some embodiments, the binding domains described herein are bound to a substrate. In some embodiments, the binding domains are directly bound to a substrate. In some embodiments, the binding domains are bound to a linker, and the linker is bound to a substrate. In some embodiments, the binding domains are covalently bound to a substrate. In some embodiments, the binding domains are non-covalently bound to a substrate.
[0055] In some embodiments, the adapters described herein are attached to a substrate. In some embodiments, the adapters are attached directly to a substrate. In some embodiments, the adapters are attached to a linker, and the linker is attached to a substrate. In some embodiments, the adapters are covalently attached to a substrate. In some embodiments, the adapters are non-covalently attached to a substrate.
[0056] In some embodiments, the present invention provides a composition comprising a substrate, an adaptor, and a binding domain. In some embodiments, the composition comprises a substrate, a binding domain, and an adaptor, as shown in FIG. 1A. In some embodiments, the composition comprises a binding domain directly bound to the substrate, and an adaptor directly bound to the substrate. In some embodiments, the composition comprises a binding domain that binds to the substrate via a linker, and an adaptor directly bound to the substrate. In some embodiments, the composition comprises a binding domain that binds to the substrate via a linker, and an adaptor that binds to the substrate via a linker. In some embodiments, the composition comprises a binding domain that binds to the substrate via a first linker, and an adaptor that binds to the same substrate via a second linker.
[0057] In one embodiment, the composition comprises: i) substrate; ii) a binding domain that binds to the substrate via a first linker, and iii) an adaptor that binds to the substrate via a second linker; In one aspect, the binding domain specifically binds to a non-canonical feature of DNA or RNA, and the adapter comprises a nucleic acid barcode sequence unique to the non-canonical feature.
[0058] In some embodiments, the invention provides a composition comprising a second recognition element, a substrate, a binding domain, and an adaptor.
[0059] In certain aspects, the invention includes one or more methods of making the compositions and conjugates disclosed herein and shown in the figures. In one aspect, the method includes directly or indirectly binding one or more adaptors to a substrate, and directly or indirectly binding one or more binding domains to a substrate, where any indirect binding may be via a linker. See, e.g., FIG. 1A.
[0060] In one aspect, the method includes directly or indirectly attaching one or more second recognition elements to a substrate, and directly or indirectly attaching one or more binding domains to the one or more second recognition elements, and directly or indirectly attaching one or more adaptors to the one or more second recognition elements, where any indirect attachment may be via a linker. See, e.g., FIG. 1B.
[0061] In one aspect, the method of manufacture includes directly or indirectly binding one or more second recognition elements to a substrate, and directly or indirectly binding one or more binding domains to a substrate, and directly or indirectly binding one or more adaptors to one or more second recognition elements, or directly or indirectly binding one or more adaptors to a substrate, where any indirect binding may be via a linker. See, e.g., FIG. 1C. In one aspect, the method of manufacture includes directly or indirectly binding one or more second recognition elements to a substrate, and directly or indirectly binding one or more binding domains to a second recognition element, and directly or indirectly binding one or more adaptors to a substrate, where any indirect binding may be via a linker.
[0062] In one aspect, the method of production includes directly or indirectly attaching two or more second recognition elements to a substrate, and directly or indirectly attaching one or more binding domains to at least one of the second recognition elements, and directly or indirectly attaching one or more adapters to the one or more second recognition elements, where any indirect attachment may be via a linker. See, e.g., FIG. 1D.
[0063] In one aspect, the method of production includes directly or indirectly attaching one or more second recognition elements to a substrate, and directly or indirectly attaching two or more binding domains to the second recognition elements, and directly or indirectly attaching one or more adaptors to a portion of the binding domains, such that one binding domain species is adaptor-labeled and does not bind to nucleic acids, while one or more other binding domain species are specific for non-canonical features of nucleic acids and are not labeled, where any indirect attachment may be via a linker. See, e.g., FIG. 1E.
[0064] In one aspect, the method of production includes directly or indirectly attaching two or more different types of cleavable adaptors to a substrate, and directly or indirectly attaching one or more capture molecules to the substrate, and providing one or more binding domains attached to a nucleic acid complementary to a capture sequence of the capture molecule such that the nucleic acid complementary to the capture sequence of the capture molecule hybridizes with the capture molecule, where any indirect attachment may be via a linker. See, e.g., FIG. IF.
[0065] In one aspect, the method of manufacture includes forming a transposome comprising a transposase dimer carrying two mosaic end (ME)-containing adapter molecules, binding the transposome directly or indirectly to a substrate, binding one or more second recognition elements directly or indirectly to the substrate, and binding one or more binding domains directly or indirectly to the second recognition elements, where any indirect binding may be via a linker. See, e.g., FIG. 1G. In one aspect, the method of manufacture includes forming a transposome comprising a transposase dimer carrying two mosaic end (ME)-containing adapter molecules, binding the transposome directly or indirectly to a substrate, and binding one or more binding domains directly or indirectly to the substrate, where any indirect binding may be via a linker.
[0066] In one aspect, the method of production includes directly or indirectly binding a second recognition element to a substrate, fusing Tn5 to Protein A to form a Tn5-Protein A fusion protein, forming a dimer of the fusion proteins, loading the Tn5-Protein A fusion protein dimer with an ME adaptor, binding a binding domain to the second recognition element, and binding Protein A of the fusion protein to a binding domain (e.g., an Fc region of an antibody), where any indirect binding may be via a linker. See, e.g., FIG. 1H. In one aspect, the method of production includes directly or indirectly binding the second recognition element to a substrate, fusing Tn5 to Protein A to form a Tn5-Protein A fusion protein, forming a dimer of the fusion proteins, loading the Tn5-Protein A fusion protein dimer with an ME adaptor, directly or indirectly binding the binding domain to the substrate, and binding Protein A of the fusion protein to a binding domain (e.g., an Fc region of an antibody), where any indirect binding is via a linker.
[0067] In some embodiments, the composition comprises a second recognition element, a substrate, a binding domain, and an adaptor, the adaptor being bound to the second recognition element as shown in FIG. 1B. In some embodiments, the composition comprises a second recognition element, a substrate, a binding domain, and an adaptor, the adaptor being bound to the second recognition element as shown in FIG. 1C. In some embodiments, the composition comprises a second recognition element directly bound to the substrate (FIG. 1C). In some embodiments, the composition comprises a second recognition element indirectly bound to the substrate, for example, via a linker (FIG. 1B). In some embodiments, the composition comprises a second recognition element directly bound to the substrate, and an adaptor directly bound to the second recognition element. In some embodiments, the composition comprises a second recognition element bound to the substrate via a linker, and an adaptor directly bound to the substrate. In some embodiments, the composition comprises a second recognition element bound to the substrate via a first linker, and an adaptor bound to the substrate via a second linker.
[0068] In some embodiments, the composition comprises: i) substrate; ii) a second recognition element bound to the substrate; iii) an adaptor that binds to a second recognition element, and iv) Binding Domain In one aspect, the binding domain specifically binds to a non-canonical feature of DNA or RNA, the binding domain is immobilized by a second recognition element, and the adapter comprises a nucleic acid barcode sequence specific for the non-canonical feature specifically bound by the binding domain.
[0069] In some embodiments, the second recognition element can bind to a single binding domain. In some embodiments, the second recognition element can bind to multiple different types of binding domains. In some aspects, the second recognition element is streptavidin, avidin, neutravidin, or a similar molecule. In some aspects, the second recognition element is protein G, protein A, protein L, a variant thereof, or an antibody.
[0070] In some embodiments, the composition comprises: i) substrate; ii) a second recognition element bound to the substrate; iii) an adaptor that binds to a second recognition element, and iv) Binding Domain In one aspect, the binding domain is configured to specifically bind to a non-canonical feature of DNA or RNA, the binding domain is immobilized by a second recognition element, and the adapter comprises a nucleic acid barcode sequence unique to the non-canonical feature. In one aspect, the composition alternatively or additionally comprises an adapter attached to the substrate directly or via a linker.
[0071] In some embodiments, the composition comprises a plurality of second recognition elements, the plurality of second recognition elements comprising different second recognition elements from one another, wherein the adapter is attached to one of the plurality of second recognition elements and the binding domain is attached to a different second recognition element.
[0072] In some embodiments, the composition comprises a plurality of second recognition elements, where the adaptor is attached to one of the plurality of second recognition elements and the binding domain is attached to another instance of the same second recognition element.
[0073] In some embodiments, the composition comprises: i) substrate; ii) a second recognition element bound to the substrate; iii) a binding domain that binds to the substrate via a second recognition element; iv) an adaptor that binds to a substrate via a linker; In some embodiments, the binding domain is configured to specifically bind to a non-canonical feature of DNA or RNA, and the adapter comprises a nucleic acid barcode sequence unique to the non-canonical feature.
[0074] In some embodiments, the composition comprises: i) substrate; ii) a second recognition element bound to the substrate; iii) a binding domain that binds to a substrate via a linker; iv) an adaptor that binds to the substrate via a second recognition element; In some embodiments, the binding domain is configured to specifically bind to a non-canonical feature of DNA or RNA, and the adapter comprises a nucleic acid barcode sequence unique to the non-canonical feature.
[0075] In some embodiments, the composition comprises a substrate, a capture molecule, an adaptor, and a binding domain. In some embodiments, the composition comprises a substrate, a capture molecule, an adaptor, and a binding domain as shown in FIG. IF. In some embodiments, the binding domain is immobilized by the capture molecule.
[0076] In some embodiments, the composition comprises: i) substrate; ii) a capture molecule bound to a substrate; iii) an adaptor bound to a substrate; and iv) a binding domain immobilized on a substrate via a capture molecule Includes.
[0077] In some embodiments, the capture molecule is a capture molecule as shown in FIG. IF. In some embodiments, the capture molecule is directly attached to the substrate. In some embodiments, the capture molecule is attached to the substrate via a linker. In some embodiments, the capture molecule is an oligonucleotide, e.g., an oligonucleotide that can capture the binding domain by binding to a complementary oligonucleotide sequence attached thereto. In some embodiments, the capture molecule is polyethylene glycol, with a pendant click chemistry group such as DBCO, azide, alkyne, mTET or TCO.
[0078] In some embodiments, the capture molecule may effect capture of the binding domain by covalent or non-covalent mechanisms. For example, covalent capture can be achieved by using bioorthogonal chemistry (DBCO / azide, alkyne / azide, mTet / TCO, etc.). Non-covalent capture is achieved by protein-based capture molecules that target specific binding sites on the binding domain.
[0079] In some embodiments, the composition comprises: i) substrate; ii) a capture molecule bound to a substrate; iii) an adaptor bound to a substrate; and iv) Binding Domain wherein the binding domain specifically binds to a non-canonical feature of DNA or RNA, the binding domain is immobilized by the capture molecule, and the adapter comprises a nucleic acid barcode sequence unique to the non-canonical feature specifically bound by the binding domain.
[0080] In one aspect, the invention includes a composition comprising: i) substrate; ii) a binding domain that binds to the substrate via a first linker or a second recognition element; iii) a mosaic end (ME) adaptor that binds to the substrate via a second linker or second recognition element; and iv) transposase, wherein the transposase is loaded onto immobilized ME adapters, the binding domain specifically binds to a non-canonical feature of DNA or RNA, and at least one of the ME adapters comprises a nucleic acid barcode sequence unique to the non-canonical feature.
[0081] In one aspect, the invention includes a composition comprising: i) substrate; ii) a binding domain that binds to the substrate via a linker or a second recognition element; and iii) a transposase bound to a binding domain; wherein the transposase is loaded onto ME adapters, the binding domain specifically binds to a non-canonical feature of DNA or RNA, and at least one of the ME adapters comprises a nucleic acid barcode sequence unique to the non-canonical feature.
[0082] In some embodiments, the composition comprises a substrate, a binding domain attached to the substrate via a first linker or attached to a second recognition element directly or indirectly attached to the substrate, a mosaic end (ME) adapter attached to the substrate via a second linker, and a transposase, where the transposase is loaded onto the ME adapter, the binding domain specifically binds to a non-canonical feature of DNA or RNA, and the adapter comprises a nucleic acid barcode sequence specific to the non-canonical feature specifically bound by the binding domain. See, for example, FIG. 1G. In some embodiments, the transposase is a Tn5 transposase. FIG. 1H shows a dimer of Tn5-Protein A fusion protein loaded with an ME adapter and bound to an antibody via affinity binding between Protein A and the Fc region of the antibody. In some embodiments, the composition comprises a substrate, a plurality of second recognition elements bound to the substrate, an adaptor bound to one of the plurality of second recognition elements, and a binding domain bound to another of the plurality of second recognition elements, where the binding domain specifically binds to a non-canonical feature of DNA or RNA, and the adaptor comprises a nucleic acid barcode sequence unique to the non-canonical feature specifically bound by the binding domain. See, e.g., FIG. ID. In some aspects, the composition comprises beads as a substrate, e.g., as shown in FIG. 15C. According to any of the aspects of the compositions described herein, the composition may comprise 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, or 15x excess of adaptor relative to binding domain. Such ratios provide efficient barcoding yields while minimizing by-products.
[0083] Also provided herein is a composition comprising one or more binding domains of the present invention.In some embodiments, the composition comprises two or more different binding domains.For example, the composition may comprise a first binding domain that binds to a first non-canonical feature, and a second binding domain that binds to a second non-canonical feature. In some embodiments, the composition comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 60, 71, 80, 90, 100, 125, 150, 175, or 200 or more different types of binding domains.
[0084] Also provided herein are compositions comprising one or more binding domains and one or more adaptors, each adaptor comprising a nucleic acid barcode sequence specific to a non-canonical feature specifically bound by the respective binding domain. For example, in a composition comprising two binding domains and two adaptors, a first adaptor comprises a nucleic acid barcode sequence specific to a non-canonical feature specifically bound by the first binding domain, and a second adaptor comprises a nucleic acid barcode sequence specific to a non-canonical feature specifically bound by the second binding domain.
[0085] In some embodiments, the compositions of the invention comprise one or more substrates. In some embodiments, the compositions comprise two substrates. In some embodiments, the compositions comprise one, two, three, four, five or more substrates.
[0086] The compositions described herein may further comprise a base editing enzyme in some embodiments. In some embodiments, the base editing enzyme is an adenosine deaminase, cytosine deaminase, glycosylase, methylase, demethylase, or dioxygenase. In some embodiments, the base editing enzyme is an enzyme that removes bases, e.g., a glycosylase. The base editing enzyme can be bound to, for example, a binding domain. Having the base editing enzyme bound to the binding domain brings the enzyme into proximity with the target nucleic acid bound to the binding domain. The base editing enzyme can then edit the target nucleic acid. After the nucleic acid is amplified and sequenced, the location of the edited base can be determined and used to determine the location where the binding domain bound to the target nucleic acid (i.e., the location of the non-canonical feature on the target nucleic acid).
[0087] In some embodiments, the base editing enzyme is covalently linked to the binding domain. For example, the base editing enzyme can be fused to the binding domain (i.e., as a fusion protein). In certain embodiments, the base editing enzyme can be covalently linked to the binding domain via a linker fused to both the base editing enzyme and the binding domain. In some embodiments, the base editing enzyme is linked to the binding domain via a targeting moiety. The targeting moiety can be selected from, for example, a peptide tag, a protein tag, a secondary antibody, a nucleic acid sequence, or a bioorthogonal reactive group. In one exemplary embodiment, the base editing enzyme can be linked to a secondary antibody, where the secondary antibody recognizes the binding domain (e.g., the primary antibody). In some embodiments, the targeting moiety is protein A, protein L, or protein G. In some embodiments, the targeting moiety is a nucleic acid bound to the base editing enzyme, where the nucleic acid bound to the base editing enzyme is complementary to the nucleic acid bound to the binding domain.
[0088] In some embodiments, the compositions described herein include one or more carriers, excipients, buffers, etc. The compositions may have a pH of about 0.5, about 1.0, about 1.5, about 2.0, about 2.5, about 3.0, about 3.5, about 4.0, about 4.5, about 5.0, about 5.5, about 6.0, about 6.5, about 7.0, about 7.5, about 8.0, about 8.5, about 9.0, about 9.5, about 10.0, about 10.5, about 11.0, about 11.5, about 12.0, about 12.5, about 13.0, about 13.5, or about 14.0. In some embodiments, the compositions are pharmaceutical compositions.
[0089] adapter As used herein, the term "adapter" refers to any short nucleic acid sequence that can be attached to the end of a DNA or RNA molecule and confers some function. For example, in some embodiments, the adapter can facilitate sequencing and / or identification of the DNA or RNA molecule. In some embodiments, the adapter is a DNA, RNA, or a mixed sequence of DNA and RNA. In some examples, the nucleic acid adapter comprises one or more backbone modifications selected from a locked nucleic acid (LNA), peptide nucleic acid (PNA), glycol nucleic acid (GNA), phosphorothioate, 2'-fluororibose, 2'-methoxyribose phosphorodithioate, methylphosphonate, phosphoramidate, guanidinopropyl phosphoramidate, triazole, guanidinium, morpholino, threose nucleic acid (TNA), or hexitol nucleic acid (HNA).
[0090] In certain embodiments, the adapter comprises a 5' phosphate. In some embodiments, the adapter comprises a 3' phosphate. In some embodiments, the adapter comprises a 5' phosphate and a 3' phosphate. In some embodiments, the adapter is single stranded. In some embodiments, the adapter is double stranded. In some embodiments, the double stranded adapter may comprise a single stranded adapter hybridized to a complementary oligonucleotide.
[0091] In some embodiments, the adapter is cleavable. For example, the adapter may include one or more cleavage sites. The cleavage site may include, for example, one or more uracil bases, a sequence recognized by an enzyme (e.g., a restriction enzyme or other nuclease), or a synthetic chemical moiety. In some embodiments, the adapter is cleavable, as shown in FIG. IF. In some embodiments, the linker is cleaved by chemical or enzymatic cleavage, for example, using a disulfide, a cathepsin B cleavage site, photocleavage, etc. In some embodiments, the adapter is cleaved at a site within the adapter, for example, at a restriction site (requiring duplex formation), with a uracil / USER enzyme, with an 8-oxoG / FPG enzyme, or via a photocleavable phosphate backbone modification, etc.
[0092] In some embodiments, the adapter comprises a universal forward primer (UFP). In some embodiments, the adapter comprises a universal reverse primer (URP). In some embodiments, the adapter comprises a UFP and a URP. In some embodiments, the adapter comprises a UFP or a URP. The UFP and URP sequences are non-naturally occurring DNA sequences that can selectively amplify only those sequences that are introduced into the target nucleic acid (or its copy). During sequencing, the UFP and / or URP anneal to the DNA target and provide a start site for the extension of a new DNA molecule (i.e., its copy). A list of exemplary UFPs and URPs can be found at the following World Wide Web address: lslabs.com / resources / universal-primer-list. In some embodiments, the universal primer sequence used in the adapter (and transferred to the target nucleic acid) is compatible with established DNA sequencing platforms and can be used to introduce surface adapters such as Illumina P5 and P7 in downstream PCR reactions.
[0093] In some embodiments, the adapter may include a barcode, such as a modification coded barcode (MBC). The MBC is a short unique nucleic acid sequence. Each MBC is used in association with a particular epigenetic or epitranscriptomic modification to aid in its identification and / or analysis. For example, the MBC may be used in adapters bound to binding domains specific for particular non-canonical features. In some embodiments, the adapter may be comprised of a barcode. In some embodiments, the adapter may be comprised of an MBC.
[0094] In some embodiments, the adapter may include a unique molecular barcode (UMI). [UMIの長さ] For example, a 10-mer UMI consists of a short random sequence with 1,048,576 (4 10 ) unique molecules. UMIs are used for absolute quantification of sequencing reads to correct for PCR amplification bias and errors. For example, an RNA sample may contain 100 copies of transcript A and 100 copies of transcript B. After PCR amplification, 1M copies of transcript A and 2M copies of transcript B may be detected because transcript B amplifies more efficiently. When using UMIs for transcript A, 10,000 copies of the 100 UMI variant are detected, and 20,000 copies of the 100 UMI variant are detected for transcript B. Counting the number of UMI variants instead of counting the number of reads gives the absolute number of molecules.
[0095] In some embodiments, the adaptor comprises one or more unnatural nucleobases. In some embodiments, the one or more unnatural nucleobases are selected from the group consisting of G clamp (9-(2-aminoethoxy)-3H-benzo[b]pyrimido[4,5-e][1,4]oxazin-2(10H)-one), tC (3H-benzo[b]pyrimido[4,5-e][1,4]oxazin-2(10H)-one), tC O(3H-benzo[b]pyrimido[4,5-e][1,4]oxazin-2(10H)-one), inosine, Super T (5-hydroxybutynyl-2'-deoxyuridine), Super G (8-aza-7-deazaguanosine), uracil, or 8-oxo-G;
[0096] In some embodiments, the adapter comprises two or more random bases at its 3' end, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more, or 2-12, or 3-8, or 4-6 random bases. In some aspects, the invention includes a method of random priming RNA to introduce a barcode using such random bases. This method does not require ligation of a spacer sequence to the target nucleic acid prior to the barcoding step.
[0097] In one aspect, the adaptor comprises a 3' or 5' blocking group. In one aspect, the 3' or 5' blocking group is independently selected from a dideoxyribose, a phosphate, a reverse base, or a linker.
[0098] 12A-12D show exemplary nucleic acid adaptor structures, with legends providing descriptions of each element used therein. These adaptors are labeled Type A, Type B, Type C, and Type D for ease of reference.
[0099] The adapter shown in FIG. 12A (Type A) represents a minimal adapter that may contain either a UFP or URP sequence. Type A adapters do not contain sequences that can be used to identify or analyze non-canonical nucleic acid functions, and are instead used for library construction. In some embodiments, Type A adapters are attached to nucleic acid molecules that do not contain non-canonical features. In some embodiments, Type A adapters are attached to nucleic acid molecules that contain non-canonical features after introducing a barcoded adapter at the other end of the target nucleic acid. For example, Type A adapters can be used to cap and prepare nucleic acids for PCR amplification after one or more barcodes have been added.
[0100] Each of the adapters shown in Figures 12B-12D contains an MBC specific for one feature of non-canonical DNA / RNA (e.g., modified bases). As shown in Figure 12B, type B adapters can be used in library preparation workflows involving circularization of cDNA. They contain a cleavage site (CLS). Cleavage of type B adapters may be performed before PCR amplification. As shown in Figure 12C, type C adapters lack a CLS and contain only one universal primer region. Type C adapters can be used for barcode introduction, for example, by ligation reactions. They may be combined with second strand synthesis methods such as template switching oligonucleotides or other adapter ligation by Smart-Seq technology. Type D adapters are specifically designed for encoding by primer extension, as shown in Figure 12D. Type D adapters have one 3'-terminal spacer (SP) or two spacer regions (e.g., SP1, SP2) at both ends. The reaction is initiated by ligating a short spacer region (SP) to the 3' end of the target nucleic acid and attaching a type D adapter with a complementary spacer. The spacer may be universal across all nucleic acid binding molecules and cycles, may be specific to each type of nucleic acid binding molecule, or may be specific to each cycle of barcoding. In some embodiments, the adapter comprises one, two, three, or four spacers. In some embodiments, the adapter comprises one spacer. In some embodiments, the adapter comprises two spacers. In some embodiments, the spacer is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides long. In some embodiments, the spacer is 6 nucleotides long. In some embodiments, the spacer is 7 nucleotides long. In some embodiments, the spacer is 8 nucleotides long.Type D adapters can be used for one barcode transcription, for example by primer extension reaction, or multiple consecutive barcode transcriptions. Multiple cycles of barcoding can be used to interrogate only one or a subset of non-canonical features in each cycle. For example, the first encoding cycle can use a nucleic acid binding molecule specific for m5C. The second encoding cycle can use a nucleic acid binding molecule specific for m6A. The third encoding cycle can use a nucleic acid binding molecule specific for inosine, etc. In another embodiment, the first cycle can interrogate m5C and m6A, and the second cycle can interrogate inosine. In another embodiment, the first encoding cycle can try all non-canonical features, and the second encoding cycle can try all non-canonical features a second time.
[0101] In some embodiments, the adapter comprises a UFP, a URP, or a UFP and a URP. In some embodiments, the adapter comprises a UFP and / or a URP and further comprises an MBC. In some embodiments, the adapter comprises a UFP and / or a URP, an MBC, and a UMI. In some embodiments, the adapter comprises a UFP and / or a URP, an MBC, a UMI, and a CLS. In some embodiments, the adapter comprises a UFP and / or a URP, an MBC, a UMI, a CLS, and a SP. In some embodiments, the adapter comprises a UFP, a CLS, a URP, a UMI, and a MBC. In some embodiments, the adapter comprises a UFP, a UMI, and a MBC. In some embodiments, the adapter comprises a URP, a UMI, and a MBC. In some embodiments, the adapter comprises a first SP, an MBC, a UMI, and a second SP.
[0102] The adaptors described herein may, in some embodiments, include one or more linkers, such as linkers that aid in linking the binding domain to the adaptor. Linkers include polyethylene glycol, carbohydrates, peptides, DNA, or RNA. The length of the linker may vary. Longer linkers can be used when the non-canonical DNA or RNA features are located away from the 5' or 3' end of the nucleic acid sequence. Shorter linkers are used when the non-canonical DNA or RNA features are located relatively close to the 5' or 3' end of the nucleic acid sequence.
[0103] In some embodiments, the adapter, or a linker sequence contained therein, is cleavable. For example, the adapter may include one or more cleavage sites. The adapter may be chemically, photochemically, or enzymatically cleavable. The cleavage site may include, for example, one or several uracil bases, a sequence recognized by an enzyme (e.g., a restriction enzyme or other nuclease), or a synthetic chemical moiety, such as a disulfide, carbonate, hydrazone, cis-aconityl, or (β-glucuronide).
[0104] As described in more detail below, adapters can be fused to single- or double-stranded target nucleic acids (e.g., DNA or RNA) using a barcode transfer reaction.
[0105] In some embodiments, barcoding by primer extension includes adding a 3' poly-rA tail to the RNA target. The 3' poly-rA tail is added by polyadenylation using any known poly(A) polymerase (e.g., E. coli poly(A) polymerase). In some embodiments, the RNA target is incubated with poly(A) polymerase and a competing poly-dT oligonucleotide. Co-treatment of poly(A) polymerase with a competing poly-dT oligonucleotide controls the length of the added 3' poly-rA tail. Typically, polyadenylation results in an average 3' poly-rA tail length of about 150 bases. In some embodiments, the length of the 3' poly-rA tail is about 5 bases, about 10 bases, about 15 bases, about 20 bases, about 25 bases, about 30 bases, about 35 bases, about 40 bases, about 45 bases, about 50 bases, about 55 bases, or about 60 bases.
[0106] In some embodiments, primer extension comprises adding a 3' poly-U tail to the RNA target. The 3' poly-U tail is added by polyuridylation using a known poly(U) polymerase (e.g., fission yeast; Schizosaccharomyces pombe Cid1). In some embodiments, the length of the 3' poly-U tail is about 5 bases, about 10 bases, about 15 bases, about 20 bases, about 25 bases, about 30 bases, about 35 bases, about 40 bases, about 45 bases, about 50 bases, about 55 bases, or about 60 bases.
[0107] In some embodiments, the adapter comprises any one of SEQ ID NOs: 1-5 provided in Table 4. In some embodiments, the adapter comprises SEQ ID NO: 1. In some embodiments, the adapter comprises SEQ ID NO: 2. In some embodiments, the adapter comprises SEQ ID NO: 3. In some embodiments, the adapter comprises SEQ ID NO: 4. In some embodiments, the adapter comprises SEQ ID NO: 5. In some embodiments, the adapter comprises an adapter as set forth in Table 4, or a sequence having one, two, three, four or five amino acid substitutions thereto.
[0108] In some embodiments, the adapters described herein comprise a 5'-amine moiety (5AmMC6). In some embodiments, the adapters comprise a 3' amino moiety (3AmMO). In some embodiments, the adapters comprise an 18 atom hexaethylene glycol spacer (iSp18). In some embodiments, the adapters comprise a single uracil surrounded by filler AT repeats for release from the substrate surface by cleavage by the USER enzyme (NEB). In some embodiments, the adapters comprise an 8-base barcode.
[0109] In some embodiments, the adaptors described herein are functionalized to the substrate with TCO-PEG4-NHS ester. In some embodiments, the adaptors are immobilized to the substrate using Protein G, A or L.
[0110] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4]
[0111] Binding domain The term "binding domain" as used herein refers to a nucleic acid, polypeptide, etc., that binds to a non-canonical feature of a target nucleic acid, such as a modified nucleoside. The term "binding domain" may be used interchangeably herein with terms such as "binding agent," "recognition element," "antibody," etc., as would be understood by one of skill in the art from the context. In some embodiments, the binding domain binds to a non-canonical feature of a target nucleic acid. In some embodiments, the binding domain does not bind to any nucleic acid feature adjacent to the non-canonical feature. In some embodiments, the binding domain binds to both (i) the non-canonical feature of the target nucleic acid, and (ii) one or more nucleic acid features (e.g., nucleobase, sugar, phosphate, or combinations thereof) adjacent to the non-canonical feature. In some embodiments, the binding domain may bind to a conserved sequence motif. For example, m 6 A frequently occurs in the motif: GG(m 6 A) CT. Therefore, the binding domain is m 6 When binding to A, it may also bind to one or more nucleic acids adjacent to it (e.g., GG or CT). As another example, the binding domain can bind to all or part of the anticodon loop of a tRNA.
[0112] The nucleic acid binding molecules described herein comprise one or more binding domains, and the binding domains specifically bind to non-canonical features of DNA or RNA. The binding domains described herein are proteins, nucleic acids, or fragments or derivatives thereof that can recognize and bind to non-canonical features of target nucleic acids. For example, in some embodiments, the binding domain comprises an antibody, an aptamer, a reader protein, a writer protein, an eraser protein, an endonuclease V, an artificial polymer scaffold, an artificial protein scaffold, or a selective covalent capture reagent, or a fragment or derivative thereof. In some embodiments, the binding domain comprises a catalytically inactive mutant of a writer protein or an eraser protein. In some aspects, the reader protein is NUDT16, YTHDC1, YTHDC2, YTHDF1, YTHDF2, or a fragment or derivative thereof. In some aspects, the writer protein is a DNMT protein, a NAT10 protein, a METTL protein, a TRM protein, a BMT protein, a DUS protein, a PUS protein, an ADAR protein, or an NSUN protein, or a fragment or derivative thereof. In some aspects, the writer protein is DNTM1, DNTM3A / B, NAT10, METTL3, METTL8, METTL14, METTL16, TRM, BMT, DUS2, PUS, or NSUN2, or a fragment or derivative thereof. In some aspects, the eraser protein is FTO protein, ALKBH protein, TET protein, or a fragment or derivative thereof. In some aspects, the eraser protein is FTO, ALKBH3, or ALKBH5, or a fragment or derivative thereof. In some embodiments, the binding domain is an IgG antibody, an antigen binding fragment (Fab), a single chain variable fragment (scFv), or a heavy or light chain single domain (V H and V L In some embodiments, the binding domain comprises a heavy chain antibody (hcAb) or a V HIn some embodiments, the binding domain comprises an artificial protein scaffold such as an adectin, an affibody, an affilin, an anticalin, an atrimer, an avimer, a bicyclic peptide, a sentinin, a cis-knot, a darpin, a finomer, a Kunitz domain, an obody, or a pronectin.
[0113] IgG antibodies are the major immunoglobulin isotype. IgG contains two identical heavy chains and two identical light chains, which are covalently stabilized through disulfide bonds. IgG consists of heavy chains (V H ) and light chain (V L) and recognizes the antigen via its variable N-terminal domain and six complementarity determining regions (CDRs). Antibodies that bind to several modified DNA and RNA bases are commercially available. For example, several companies, including Active Motif and Sigma, sell antibodies specific for 5-methylcytidine (m5C), 5-hydroxymethylcytidine (hm5C), and N6-methyladenosine (m6A). Eurogentec (Belgium) sells a monoclonal antibody that binds to m5C. Monoclonal antibodies that bind to inosine can be purchased, for example, from Diagenode. Megabase Research Products (USA) sells a rabbit polyclonal serum that binds to m5C 6-methyladenosine and 7-methylguanosine. Abcam (USA) sells recombinant antibodies against the RNA modifications m6A, ac4C, m1A, m2,2G, m4C, m2A, m6,6A, and m8A. Antibodies that bind to modified bases can be developed according to methods known to those skilled in the art. In some embodiments, the antibody is a monoclonal antibody, a polyclonal antibody, or a functional fragment or variant thereof. As used herein, the term "antibody" refers to any specific binding substance having a binding domain with the required specificity. Thus, the term covers antibody fragments, derivatives, functional equivalents, and homologs of antibodies, including any polypeptide containing an immunoglobulin binding domain, whether natural or synthetic, monoclonal or polyclonal. Also included are chimeric molecules consisting of a fusion of an immunoglobulin binding domain, or equivalent, with another polypeptide.
[0114] In some embodiments, the binding domain may comprise a nanobody. Nanobodies are antibodies that contain a single variable domain (V) of a heavy chain antibody, such as those produced by camelids and some cartilaginous fish. H H) is included. HThe H domain contains three CDRs that are enlarged compared to those of IgG antibodies, providing an antigen interaction surface of similar size to IgG (i.e., about 800 Å2). Nanobodies bind antigens with similar affinity to IgG antibodies, but have several advantages over them: they are small (15 kDa), have fewer disulfide bonds and are therefore less susceptible to reducing environments, are more soluble, and lack post-translational glycosylation. Nanobodies can be produced in bacterial expression systems, allowing affinity and specificity maturation by phage and other display techniques. Other advantages include improved thermostability and solubility, and an easy approach to site-specific labeling. Due to their small size, nanobodies can form convex paratopes, making them suitable for binding to antigens that are difficult to access. Exemplary methods for producing nanobodies include immunizing the respective animal (e.g., camel) with the antigen of interest, further evolving existing naive libraries, or a combination thereof.
[0115] In some embodiments, the binding domain comprises a reader protein, a writer protein, or an eraser protein. A "reader protein" is a protein that selectively recognizes and binds to a specific chemical modification on DNA or RNA. A "writer protein" is a protein that adds a specific chemical modification to DNA or RNA. An "eraser protein" is an enzyme that removes a specific chemical modification from DNA or RNA. In some embodiments, the binding domain comprises a fragment or derivative of a reader protein, a writer protein, or an eraser protein. In some embodiments, the binding domain comprises an engineered form of a reader, writer, or eraser protein, such as a form engineered to retain nucleic acid binding but lack enzymatic activity. In some embodiments, the binding domain comprises a catalytically inactive mutant of a writer or eraser protein. Exemplary reader proteins, writer proteins, and eraser proteins that may be used in the binding domains described herein are shown in Tables 1 and 2. Other reader proteins, writer proteins, and eraser proteins are listed at the following World Wide Web address: rnawre.bio2db.com.
[0116] [Table 2]
[0117] [Table 3-1] [Table 3-2] [Table 3-3] Legend: W: writer, E: eraser, R: reader, TS: tumor suppressor, Onc: oncogene. RNA modifications: m1A: 1-methyladenosine, ms2i6A: 2-methylthio-N6-isopentenyladenosine, i6A: N6-isopentenyladenosine, m6A: N6-methyladenosine, m3C: 3-methylcytosine, m5C: 5-methylcytosine, ac4C: N4-acetylcytosine, m7Gpp(pN): 7-methylguanosine cap, m7G: 7-methylguanosine internal, m2,2G: N2,N2,-dimethylguanosine, m2G: N2-methylguanosine, Q: queuosine, yW et al: wibbutosine and derivatives, m5U: 5-methyluridine, ncm5U: 5-carbamoyl-methyluridine, mcm5U: 5-methoxycarbonyl-methyluridine, mcm5s2U: 5-methoxycarbonylmethyl-2-thiouridine, D. dihydrouridine, Ψ: pseudouridine, Nm: 2'-O-methyl nucleotide, m(pN): 5' phosphate monomethylation, A-to-I: adenosine deamination, C-to-U: cytosine deamination.RNA-modifying enzymes: ADAR1-3: adenosine deaminase RNA-specific 1-3, ALKBH1 / 3 / 5 / 8: AlkB homolog 1 / 3 / 5 / 8, APOBEC1 / 3G: apolipoprotein B mRNA editing catalytic subunit 1 / 3G, BCDIN3D: BCDIN3 domain-containing RNA methyltransferase, BUD23: rRNA methyltransferase and ribosome maturation factor, CDK5RAP1: CDK5 regulatory subunit-associated protein 1, CMTR1 / 2: cap methyltransferase 1 / 2, CTU1 / 2: cytoplasmic thiouridylase subunit 1 / 2, DKC1: dyskerin pseudouridine synthase 1, DNMT2: tRNA Aspartate methyltransferase 1, DUS2: dihydrouridine synthase 2, ELP3: elongator acetyltransferase complex subunit 3, FTO: FTO α-ketoglutarate-dependent dioxygenase, HENMT1: HEN methyltransferase 1, METTL1 / 2 / 3 / 6 / 8 / 14 / 16: methyltransferase-like-1 / 2 / 3 / 6 / 8 / 16, NAT10: N-acetyltransferase 10, NSUN1-5: NOP2 / Sun RNA methyltransferase 1-5, NUDT16: Nudix hydrolase 16, RNMT: RNA guanosine-7 methyltransferase, TGT: Queuine TRNA-ribosyltransferase catalytic subunit 1, TRIT1: tRNA isopentenyltransferase 1, TRMT1 / 2A / 2B1 / 5 / 6 / 10C / 11 / 61A / 61B / 112: tRNA methyltransferase subunits, TYW2: tRNA-YW synthesis protein 2 homolog.
[0118] In some embodiments, the binding domain comprises a reader protein. In some embodiments, the binding domain comprises a reader protein selected from NUDT16, YTHDC1, YTHDC2, YTHDF1 or YTHDF2. NUDT is a U8 snoRNA decapping enzyme (see, e.g., Uniprot Accession No. Q96DE0). YTHDC1 is a regulator of alternative splicing that specifically recognizes and binds to N6-methyladenosine (m6A)-containing RNA (see, e.g., Uniprot Accession No. Q96MU7). YTHDC2 is a 3'-5' RNA helicase (see, e.g., Uniprot Accession No. Q9H6S0). YTHDF1 specifically recognizes and binds to mRNAs containing N6-methyladenosine (m6A) and regulates their stability (see, e.g., Uniprot Accession No. Q9BYJ9). YTHDF2 specifically recognizes and binds to mRNAs containing N6-methyladenosine (m6A) and regulates their stability (see, e.g., Uniprot Accession No. Q9Y5A9). In some embodiments, the binding domain comprises a fragment or derivative of NUDT16, YTHDC1, YTHDC2, YTHDF1 or YTHDF2.
[0119] In some embodiments, the binding domain comprises a writer protein. In some embodiments, the binding domain comprises a writer protein selected from DNTM1, DNTM3 A / B, NAT10, METTL3, METTL8, METTL15, TRM, BMT, DUS2, PUS, and NSUN2. DNMT1 and DNTM3A / B are DNA (cytosine-5) methyltransferases. NAT10 is an RNA cytidine acetyltransferase (see, e.g., Uniprot Accession No. Q9H0A0). METTL3 is an N6-adenosine-methyltransferase catalytic subunit (see, e.g., Uniprot Accession No. Q86U44). NSUN2 is an RNA cytosine C(5)-methyltransferase (see, e.g., Uniprot Accession No. Q08J23). In some embodiments, the binding domain comprises a writer protein that is a fragment or derivative of NAT10, METTL3, or NSUN2. In certain aspects, the writer protein is a DNMT protein, a NAT10 protein, a METTL protein, a TRM protein, a BMT protein, a DUS protein, a PUS protein, an ADAR protein, or an NSUN protein, or a fragment or derivative thereof.
[0120] In some embodiments, the binding domain comprises an eraser protein. In some embodiments, the binding domain comprises an engineered eraser protein selected from FTO, ALKBH3, and ALKBH5. FTO is an α-ketoglutarate-dependent dioxygenase (see, for example, Uniprot Accession No. Q9C0B1). ALKBH3 is an α-ketoglutarate-dependent dioxygenase alkB homolog 3 (see, for example, Uniprot Accession No. Q96Q83). ALKBH5 is an RNA demethylase (see, for example, Uniprot Accession No. Q6P6C2). In some embodiments, the binding domain comprises a writer protein that is a fragment or derivative of FTO, ALKBH3, or ALKBH5.
[0121] The binding domain can be selected and / or engineered to bind to a non-canonical feature of DNA or RNA. For example, the non-canonical feature is a modified base, a modified backbone, or a structural element. In some embodiments, the binding domain can bind to more than one non-canonical feature.
[0122] In some embodiments, the binding domain binds to modified bases and / or nucleosides. In some embodiments, the binding domain contacts at least one, at least two, or at least three modified nucleosides. In some embodiments, the binding domain contacts at least one modified nucleoside. In some embodiments, the binding domain contacts at least one modified nucleoside and one or more adjacent nucleotides. Exemplary modified nucleosides that may occur in humans and other organisms are shown in Table 3A. Modified nucleosides known to be expressed in humans are shown in Table 3B. Other modified bases and nucleosides are listed at the world wide web address genesilico.pl / modomics / modifications.
[0123] [Table 4] *As will be appreciated by those of skill in the art, modified bases / nucleosides that are typically found in RNA may be present in DNA and modified bases / nucleosides that are typically found in DNA may be present in RNA.
[0124] [Table 5-1] [Table 5-2]
[0125] In some embodiments, the binding domain binds to one or more of the following modified nucleosides: 3-methylcytidine (mC), 5-methylcytidine (mC), N 4 -Acetylcytidine (ac4C), pseudouridine (Ψ), 1-methyladenosine (m1A), N 6 -Methyl adenosine (m6A), inosine (I), 7-methyl guanosine (m7G), 7-methyl guanosine (m7G)-cap, dihydrouridine (D), 3-methyl uridine (m3U), 5-methyl uridine (m5U), 1-methyl guanosine (m1G), N 2 -Methylguanosine (m2G), 5-methyldeoxycytidine (m5dC), N 4 -methyldeoxycytidine, 5-hydroxymethylcytidine (5-hmC), 5-hydroxymethyldeoxycytidine (5hmdC), 5-carboxydeoxycytidine (5cadC), 5-carboxycytidine (5caC), 5-formylcytidine (5fC), 5-formyldeoxycytidine (5fdC), 6-methyldeoxyadenosine, N 7 -methylguanosine (m7G), 2,7,2'-methylguanosine, ribose methylated (Nm), N2,N2-dimethylguanosine (m 2 2G), 5-carbamoylmethyl-2'-O-methyluridine (ncm5Um), 5-methoxycarbonylmethyluridine (ncm5mU), 5-methoxycarbonylmethyl-2-thiouridine (mcm5s2U), queuosine (Q), 2-thiouridine (s2U), 5-taurinomethyluridine (τm5U), 5-taurinomethyl-2-thiouridine (τm5s2U), N6-isopentenyladenosine (I6A), 2-methylthio-N6-threonylcarbamoyladenosine (ms2t6A).
[0126] In some embodiments, the non-canonical features are: 3-methylcytidine (mC), 5-methylcytidine (mC), N 4 -Acetylcytidine (ac4C), pseudouridine (Ψ), 1-methyladenosine (m1A), N 6-Methyl adenosine (m6A), inosine (I), 7-methyl guanosine (m7G), 7-methyl guanosine (m7G)-cap, dihydrouridine (D), 3-methyl uridine (m3U), 5-methyl uridine (m5U), 1-methyl guanosine (m1G), N 2 -Methylguanosine (m2G), 5-methyldeoxycytidine (m5dC), N 4 -methyldeoxycytidine, 5-hydroxymethylcytidine (5-hmC), 5-hydroxymethyldeoxycytidine (5hmdC), 5-carboxydeoxycytidine (5cadC), 5-carboxycytidine (5caC), 5-formylcytidine (5fC), 5-formyldeoxycytidine (5fdC), 6-methyldeoxyadenosine, N 7 -methylguanosine (m7G), 2,7,2'-methylguanosine, or ribose methylation (Nm).
[0127] In some embodiments, the binding domain binds to nucleic acid damage resulting from naturally occurring oxidative or UV-induced damage, or from bulky adduct formation or base alkylation by exogenous agents. In some embodiments, the nucleic acid damage is determined to be 8-oxoguanine (8-oxoG), one or more abasic sites, cis-platin crosslinks, benzo(a)pyrene diol epoxide (BPDE) adducts, cyclobutene pyrimidine dimers (CPDs), pyrimidine-pyrimidone (6-4) photoproducts (6-4PPs), 6-O-methylguanine (O-methylguanine), 1-methyl-1-methylguanine (1-methyl-1 ... (1- 6 In some embodiments, the non-canonical feature is nucleic acid damage resulting from naturally occurring oxidative or UV-induced damage, or from bulky adduct formation or base alkylation by exogenous agents. In some embodiments, the nucleic acid damage is 8-oxoguanine (8-oxoG), one or more abasic sites, cis-platin crosslinks, benzo(a)pyrene diol epoxide (BPDE) adducts, cyclobutene pyrimidine dimers (CPDs), pyrimidine-pyrimidone (6-4) photoproducts (6-4PPs), 6-O-methylguanine (O 6-MedG), or O6-(carboxymethyl)-2'-deoxyguanosine (O6-CMdG).
[0128] In some embodiments, the binding domain binds to a structural element. The structural element may be, for example, a hairpin or a loop. Other exemplary structural elements include, but are not limited to, Z-DNA structures, G-quadruplexes, I-motifs, bulges, abasic sites, triplexes, three-way bonds, cruciform structures, tetraloops, ribose zippers, pseudoknots, and the like. In some embodiments, a plurality of compositions are provided, each composition comprising a binding domain, and each binding domain binds to a different type of non-canonical feature. This allows for a multiplexing approach that can simultaneously detect multiple non-canonical features.
[0129] The binding domains described herein may specifically bind to RNA or specifically bind to DNA. In some embodiments, the binding domains bind to both RNA and DNA. In some embodiments, the binding domains specifically bind to double-stranded nucleic acids with one or more non-canonical features. In some embodiments, the binding domains specifically bind to single-stranded nucleic acids with one or more non-canonical features.
[0130] In some embodiments, the binding domain binds to a non-canonical feature of the target nucleic acid, such that the DNA adaptor is proximate to the 5' or 3' end of the target nucleic acid. For example, FIG. 3 shows a target nucleic acid bound to a binding domain, with the adaptor positioned proximate to the 3' end of the target nucleic acid. FIG. 5 shows a binding domain immobilized to a second recognition element, with the target nucleic acid bound to the binding domain, with the adaptor positioned proximate to the 3' end of the target nucleic acid. In some embodiments, the target nucleic acid is bound to the binding domain, with the adaptor positioned proximate to the 3' end of the target nucleic acid. In some embodiments, the target nucleic acid is bound to the binding domain, with the adaptor positioned proximate to the 5' end of the target nucleic acid.
[0131] Binding domains can be produced using standard molecular biology, protein engineering, and / or chemical techniques.
[0132] Adapters (e.g., adapters including linkers) can be attached to substrates using several different methods. In some embodiments, adapters can be covalently attached to second recognition elements or intermediate proteins by random tagging (see, e.g., FIG. 1B and FIG. 1E). For example, NHS-activated residues on the adapters can react with one or more amine groups of surface-exposed protein lysine residues of the second recognition element or intermediate protein. Similarly, maleimide-activated adapters can react with natural or artificial cysteines of the second recognition element or intermediate protein. As will be appreciated by those skilled in the art, the number of adapters linked to the second recognition element or intermediate protein can vary depending on the number of reactive lysine or cysteine residues, respectively, and the choice of reaction conditions. In some embodiments, adapters can be non-covalently attached to the second recognition element (see, e.g., FIG. 1C and FIG. 1D). For example, 5'-biotinylated adapters can be attached to substrate-anchored streptavidin, avidin, neutravidin, or variants thereof.
[0133] Site-selective coupling methods can also be used to couple adapters to second recognition elements (see, e.g., Figures 1B and 1E). Site-selective methods can also be used to couple Tn5 transposase to binding domains (see, e.g., Figure 1H) or nucleic acid editing enzymes to binding domains (see, e.g., Figure 10). Site-specific coupling avoids affecting the function of the binding domain, the second recognition element, and intermediate proteins, allowing for reproducible material production. Site-selective internal tagging of the second recognition element or intermediate protein can be achieved by genetically incorporating unnatural amino acids using cell lines engineered with aminoacyl-tRNA synthetase / tRNA pairs. The incorporated unnatural amino acids represent sites where bioorthogonal reactions can occur. Commonly used are amino acids that have sites capable of copper-catalyzed azide-alkyne cycloaddition (CuAAC), photoactivated 1,3-dipolar cycloaddition, strain-promoted azide-alkyne cycloaddition (SPAAC), or inverse electron-demand Diels-Alder cycloaddition (IEDDA). Exemplary and versatile methods for C- or N-terminal or internal tagging of binding domains, second recognition elements, or intermediate proteins include the use of protein or peptide tags. Protein tags, such as SNAP-tag, Halo-tag, Spy-tag, Snoop-tag, Isopeptag, Dog-tag, Sdy-tag, and Clip-tag, are small proteins or peptides that can be cloned into any protein gene to express binding domains, second recognition elements, or intermediate proteins as protein-tag fusion proteins. Such protein tags can autocatalyze covalent bond formation with specific peptides or substrates. For example, SpyCatcher is a 113 residue protein that recognizes the 13 residue peptide SpyTag, which can readily bind to any DNA sequence. In some embodiments, SpyCatcher comprises SEQ ID NO: 12. In some embodiments, SpyTag comprises SEQ ID NO: 10. Depending on the binding domain, the second recognition element, and the molecular weight of the intermediate protein, smaller peptide tags may be preferred.Peptide tags are usually 10-12 amino acids long and act in an enzyme-mediated coupling reaction. Examples of enzyme-mediated reactions for attaching a second recognition element or intermediate protein to an adaptor include, but are not limited to, (a) using biotin ligase to attach an AP-peptide-labeled binding domain and biotin-DNA (e.g., biotin linker), (b) using lipoic acid ligase to link an LAP-peptide-labeled second recognition element and lipoic acid-DNA (e.g., lipoic acid-linker), (c) using tubulin tyrosine ligase to link a Tub-tag-labeled second recognition element and tyrosine-modified DNA (e.g., tyrosine-modified linker), (d) using Sortase-A to react with LPxTG peptide and glycine-modified DNA (e.g., glycine-modified linker), etc. In some embodiments, a Tn5 transposase-protein A fusion protein can be generated and attached to the Fc region of an RNA modification-specific antibody (see, for example, FIG. 1H). In some embodiments, the ADAR enzyme, adenosine deaminase, can be genetically fused to protein L and attached to the Fc region of an RNA modification-specific antibody (see, for example, FIG. 10). In some embodiments, SpyTag can be engineered into a binding domain, and SpyCatcher can be engineered into a nucleic acid editing enzyme or Tn5 transposase. Mixing SpyTag-modified binding domains with SpyCatcher-modified nucleic acid editing enzymes can produce covalent conjugates that include the binding domain and the nucleic acid editing enzyme, which can be used to mark the location of non-canonical features. Mixing SpyTag-modified binding domains with Tn5 transposase produces covalent conjugates that include the binding domain and the Tn5 transposase. In addition, a group of metal ion recognition tags and small molecule binding motifs can also be used. Another method of peptide tagging is to redirect endogenous cellular machinery to introduce aldehydes into recombinant proteins. This method utilizes formylglycine synthase (FGE), which co-translationally converts cysteine to formylglycine (FGly) within a conserved 13-residue consensus sequence.The resulting aldehyde tag can be easily modified with reactive amines for attachment to DNA.
[0134] In some embodiments, the adaptor is attached to a second recognition element or intermediate protein via bioorthogonal chemistry. In some embodiments, the second recognition element or intermediate protein comprises a DNA oligonucleotide that facilitates the attachment of the barcode. DNA oligonucleotides are readily available commercially with amino, azide, biotin, and alkyne modifications. Alkyne and azide oligos can be coupled to unnatural amino acids in copper-catalyzed azide-alkyne cycloaddition reactions or strain-promoted azide-alkyne cycloaddition reactions. Amino oligonucleotides react with formylglycine and are introduced into the binding domain within a conserved sequence of 13 amino acids (aa) by formylglycine generating enzyme (FGE).
[0135] When the binding domain described herein binds to a target nucleic acid, a complex is formed. In some embodiments, the binding domain of the complex may be covalently bound to the target nucleic acid. For example, the binding domain may be chemically and / or photochemically bound to the target nucleic acid.
[0136] Second Recognition Element The second recognition element is an antibody, protein, or peptide that is used to anchor the binding domain described herein to the surface of a substrate. In some embodiments, the second recognition element described herein is bound to a linker, and the linker is bound to a substrate. In some embodiments, the second recognition element is bound to an antibody binding domain. In some embodiments, the second recognition element is protein G, protein L, protein A, protein AG, protein AL, protein LG, or an antibody. In some embodiments, the antibody is a species-specific antibody. In some embodiments, the species-specific antibody is selected from, but not limited to, mouse, rat, rabbit, human, or non-human primate.
[0137] In some embodiments, an adaptor is attached to the second recognition element. For example, in some embodiments, the second recognition element is an antibody and the adaptor is attached to the Fc region of the antibody. The adaptor may be attached to a lysine of the protein using an N-hydroxysuccinimidyl ester (NHS ester). The adaptor may be attached to a cysteine of the protein using a maleimide group or an iodoacetyl group. The adaptor may react with a carbohydrate group of an antibody or other glycoprotein. In some embodiments, one adaptor is attached to the second recognition element. In some embodiments, two adaptors are attached to the second recognition element. In some embodiments, multiple adaptors are attached to the second recognition element.
[0138] In some embodiments, the second recognition element is a protein. In some embodiments, the second recognition element is a peptide tag. Examples of peptide tags include, but are not limited to, Flag, Avi, HA, His, Myc and Strep-tag. In some embodiments, the second recognition element is a covalent peptide tag. Examples of peptide tags include, but are not limited to, Spy tag, Snoop tag or Dog tag. In some embodiments, the second recognition element is a protein tag. Examples of protein tags include, but are not limited to, MBD, CLIP and Halo.
[0139] In some embodiments, the second recognition element is an avidin protein, such as streptavidin, neutravidin, or related variants. For example, the substrate can be coated with streptavidin and co-functionalized with a biotin-labeled adaptor and biotinylated Protein G, where Protein G is further conjugated to an antibody binding domain.
[0140] Adapter / barcode transfer reaction The nucleic acid binding molecules described herein can be used to transfer an adapter, such as an adapter that includes a barcode, to a target nucleic acid. Thus, in some embodiments, the binding domains described herein can be used to transfer a barcode to a target nucleic acid. The barcode can be an MBC, i.e., a barcode specific to the non-canonical function that is specifically bound by the binding domain. The target nucleic acid to which the adapter has been introduced is referred to herein as a "labeled target nucleic acid," "labeled target," or similar term. The target nucleic acid to which the barcode has been transcribed is referred to herein as a "barcoded target nucleic acid," "barcoded target," or similar term. The reaction in which the adapter is transferred to the target nucleic acid is referred to herein as an "adapter transfer reaction." Similarly, the reaction in which the barcode is transferred to the target nucleic acid is referred to herein as a "barcode transfer reaction."
[0141] The purpose of adapter / barcode transfer is to covalently link the adapter / barcode to the target nucleic acid molecule or to a copy of the target nucleic acid molecule. For example, in some embodiments, the barcode is chemically or enzymatically ligated to the 5' or 3' end of the target nucleic acid. In some embodiments, barcoding is achieved by extending the 3' end of the nucleic acid by DNA polymerase, RNA polymerase or reverse transcriptase using the adapter as a template to introduce the barcode. In some embodiments, the 3' end of the target nucleic acid and the 3' end of the adapter are each hybridized and simultaneously extended by reverse transcriptase. In some embodiments, an adapter with a degenerate base at the 3' end randomly primes the DNA or RNA target and is extended by DNA polymerase or reverse transcriptase. The labeled / barcoded nucleic acid molecule can be sequenced in a downstream process in some embodiments. In some embodiments, a copy of the labeled target nucleic acid can be sequenced. Figures 2A-2G show examples of adapter / barcode transfer reactions.
[0142] The enzymes used for adapter transfer are different for DNA and RNA target nucleic acids and depend on the structure of the adapter. Introduction of the adapter / barcode to the target DNA can be achieved using one or more enzymes such as T4 DNA ligase, Circ ligase, Klenow fragment, Bst DNA polymerase or Bsu DNA polymerase. Transfer of the adapter / barcode to the target RNA can be achieved using, for example, T4 RNA ligase 1, T4 RNA ligase 2 or RtcB ligase. Reverse transcriptase can also be used to simultaneously copy the barcode and synthesize cDNA. This reaction can be achieved using M-MLV reverse transcriptase, AMV reverse transcriptase or a group II intron-encoded reverse transcriptase such as Induro (商標) It is catalyzed by reverse transcriptase (NEB). Some commercially available M-MLV variants, such as Superscript II RT (Thermo Fisher), Superscript IV RT (Thermo Fisher), and Maxima H Minus RT (Thermo Fisher), can catalyze the template switching reaction, which can be used to introduce a second adapter after barcode transcription (see, for example, Figures 5 and 6).
[0143] For example, Figure 5 shows ligation of a single-stranded DNA adapter (e.g., an adapter that includes or consists of a barcode) to a single-stranded target nucleic acid. In some embodiments where the target nucleic acid is RNA, the adapter includes a 5' phosphate and is catalyzed by T4 RNA ligase 1. Alternatively, the adapter can be 5'-preadenylated and transferred by T4 RNA ligase 2, eliminating the need for ATP and limiting the reaction to one turnover. Alternatively, an unphosphorylated adapter can be used and transferred to a 3'-phosphorylated RNA using RtcB ligase. In some embodiments where the target nucleic acid is DNA, the adapter / barcode can be transferred in a reaction catalyzed by Circ ligase.
[0144] Figure 6 shows reverse transcriptase-mediated barcoding. Ligation of a universal spacer sequence (SP) allows hybridization of the target RNA to the adapter, while the binding domain captures the RNA modification. Hybridization occurs in the configuration shown in Figure 2G. Reverse transcriptase extends the 3' end of the RNA, copying the modified barcode. At the same time, the enzyme extends the 3' end of the adapter, generating cDNA. Template switching oligos are included in the reaction to introduce a universal region of choice, e.g., an Illumina sequencing adapter.
[0145] Splint ligation can also be used to introduce adapters / barcodes into target nucleic acids. In splint ligation, a bridging DNA or RNA oligonucleotide is used to join two nucleic acids, which are then joined by one or more enzymes. For example, splint ligation of two RNAs (e.g., a target RNA and an adapter / barcode) can be performed using T4 ligase 1 and a bridging RNA oligonucleotide that is complementary to the RNA. For example, the splint nucleic acid construct shown in Figure 2B can be created using splint ligation. SplintR ligase is used to join the 3' end of the RNA to the 5'-pDNA, which is annealed to the DNA or RNA complement. When the target molecule is DNA, splint DNA ligation can be performed using enzymes such as T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, or E. coli DNA ligase.
[0146] Splint extension is another method that can be used to transfer adapters / barcodes to target nucleic acids. A "sprint" is a sequence that spans a ligation junction. When a primer is used, the primer generally does not span the ligation junction. The splint can display random bases or universal synthetic bases to facilitate binding to a target nucleic acid of unknown sequence. Figure 2C shows adapter transfer by splint extension, where a copy of the sequence of a target nucleic acid molecule is made using a double-stranded adapter with a 3'-base overhang. The 3'-base overhang can include random bases or synthetic universal bases that promiscuously base-pair. When the target nucleic acid molecule is RNA, the reaction is catalyzed by reverse transcriptases such as Avian Myeloblastosis Virus (AMV) reverse transcriptase and Moloney Murine Leukemia Virus (M-MuLV, MMLV). When the target molecule is DNA, the primer can be extended by any suitable DNA polymerase, whether or not it has 3' to 5' exonuclease activity.
[0147] In some embodiments, template extension can be used to transfer adapters / barcodes to target nucleic acids. FIG. 2G shows direct adapter transfer by primer extension, initiated by hybridization of the adapter to the target RNA. Reverse transcriptase can be used to extend the adapter and target nucleic acid, thereby introducing a barcode into the target RNA and its cDNA copy ("bidirectional extension"). The adapter can hybridize via a short spacer sequence (SP) that can be ligated upstream to the target nucleic acid (FIG. 2G) or can hybridize randomly via degenerate bases that are part of the adapter sequence (FIG. 2H). Blocking either one of the 3' ends controls whether primer extension is unidirectional or bidirectional. Unidirectional extension as shown in FIG. 2D can be performed as part of a multicycle encoding process using adapters with two spacer sequences or as a single cycle. In the case of DNA adapters / barcodes, extension of the target nucleic acid is catalyzed by a DNA polymerase, such as Klenow fragment, T7, T4, Bst or Bsu DNA polymerase. In some embodiments, the generated barcoded nucleic acids are capped with a universal primer for downstream amplification as a final step.
[0148] In addition, double-stranded ligation can also be used to transfer adapter / barcode to target nucleic acid. For example, FIG. 2E shows double-stranded ligation for adapter / barcode transfer. In some embodiments, the target nucleic acid molecule can be double-stranded DNA, or an RNA / DNA hybrid, and can have either blunt or sticky ends. Ligation of double-stranded DNA blunt and sticky ends is catalyzed by T4, T3, T7 or E. coli ligase.
[0149] In some embodiments, chemical ligation can be used to introduce the adapter / barcode into the target nucleic acid.
[0150] In some embodiments, target nucleic acids can be barcoded by enzymatic transposition using Tn5 transposase (Figures 9 and 13). Tn5 and transposase use a one-step cut-and-ligate mechanism to insert mosaic end (ME) adapters into double-stranded nucleic acid targets. Suitable targets are genomic DNA or DNA / RNA heteroduplexes. ME adapters can include a 19 bp ME sequence, MBC, UMI, and universal sequences such as UFP or URP. An example of a product of the transposition reaction is shown in Figure 13. Transposase forms homodimers, with each transposase monomer loading one ME adapter. A transposase dimer containing two ME adapters is required to cleave the target nucleic acid and barcode both ends released at the cleavage site.
[0151] Methods for facilitating barcode transfer to immobilized target nucleic acids by binding domains - Patent Application 20070123633 Transfer of the adapter / barcode can be facilitated by spatially separating the molecules involved in the reaction (e.g., binding domain, adapter, second recognition element, intermediate protein, etc.) Specifically, transfer can be facilitated by positioning a complex that includes a molecule (e.g., adapter and binding domain), a target nucleic acid, and / or a binding domain bound to a target nucleic acid such that the binding domain bound to the target nucleic acid is in close proximity to the adapter, allowing transfer of the adapter to the target nucleic acid.
[0152] In some embodiments, spatial arrangement can be achieved by surface immobilization. For example, the binding domains described herein can be immobilized by binding to a substrate (see Figures 1A-1H). In most substrate formats, there is only one type of binding domain. The format shown in Figure 1B, in which the adapter is bound to a second recognition element that is also bound to the binding domain, may further include at least two, at least three, at least four, at least five, or more types of binding domains, provided that the binding domains are configured at single molecular intervals. Each "type" of binding domain binds to a different non-canonical feature and / or includes a different barcode. In some embodiments, the binding domain is positioned on the substrate adjacent to the adapter, allowing for transfer of the adapter to a target nucleic acid bound to the binding domain. In some embodiments, the binding domain is positioned on the substrate adjacent to the adapter, allowing for transfer of a copy of the adapter sequence to the target nucleic acid. For barcoding by ligation, the adapter is transferred. However, for barcoding by primer extension, a copy of the adapter is transferred. For example, the distance between the binding domain and the adapter is 200 nm or less, 190 nm or less, 180 nm or less, 170 nm or less, 160 nm or less, 150 nm or less, 140 nm or less, 130 nm or less, 120 nm or less, 110 nm or less, 100 nm or less, 90 nm or less, 80 nm or less, 70 nm or less, 60 nm or less, 50 nm or less, 40 nm or less, 30 nm or less, 20 nm or less, 10 nm or less, or 5 nm or less.
[0153] Exemplary substrates to which the binding domain, adaptor, second recognition element, and / or intermediate protein may be attached include, for example, beads, chips, plates, slides, dishes, or three-dimensional matrices. In some embodiments, the substrate is a resin, a membrane, a fiber, or a polymer. In some embodiments, the substrate is a bead, such as beads comprising sepharose, agarose, cellulose, polystyrene, polymethacrylate, and / or polyacrylamide. In some embodiments, the substrate is a magnetic bead. In some embodiments, the support is a polymer, such as a synthetic polymer. A non-limiting list of synthetic polymers includes polystyrene, poly(ethylene)glycol, polyisocyanopeptide polymer, polylactic-co-glycolic acid, poly(ε-caprolactone) (PCL), polylactic acid, poly(3-hydroxybutyrate-co-3-hydroxyvalerate) (PHBV), chitosan, cellulose, and the like.
[0154] The molecules (e.g., binding domains, adaptors, second recognition elements, and / or intermediate proteins) can be directly bound to the substrate surface. For example, the molecules can be directly bound to the substrate by one or more covalent or non-covalent bonds. In embodiments where the substrate is a 3D matrix or other 3D structure, the nucleic acid binding molecules can be bound to multiple surfaces of the substrate.
[0155] In some embodiments, the nucleic acid binding molecule can be indirectly bound to the surface of the substrate. For example, the binding molecule can be indirectly bound to the substrate surface via a capture molecule, and the capture molecule is directly bound to the substrate. The capture molecule can be a nucleic acid, a protein, a sugar, a chemical linker, etc., that can bind or link both the substrate and the nucleic acid binding molecule and / or the target nucleic acid. In some embodiments, the capture molecule is bound to a binding domain or an adaptor (e.g., a linker of the adaptor) and immobilized on the substrate.
[0156] In some embodiments, the first adaptor is separated from the second adaptor on the substrate surface so that each adaptor can only interact with one target nucleic acid (i.e., the target nucleic acid immobilized by the binding domain). In some embodiments, the binding domain and the adaptor are arranged on the substrate surface to ensure the interaction of the adaptor with the target nucleic acid bound to the binding domain. In some embodiments, the adaptor is at least 1 nm away from the binding domain, and at most 30 nm away. For example, in some embodiments, the adaptor and the binding domain are about 15 nm away.
[0157] In some embodiments, the multiple copies of the adaptor are about 1 adaptor / 5 nm. 2 Approximately 1 adapter / 50 nm 2 , e.g. 1 adapter / 20 nm 2 In some embodiments, multiple copies of the binding domain are bound to the substrate at a density of 1000 nm 2 Approximately 1 binding domain per 2 Approximately one binding domain per 2 The antibodies are bound to the substrate at a density of one binding domain per antibody.
[0158] Generally, the purpose of binding domains to a substrate is to ensure the transfer of adapters and / or barcodes to target nucleic acids bound to the binding domains. Figures 1A-1H show non-limiting examples of methods for binding domains and adapters to a substrate and immobilizing them. These examples are described in detail below.
[0159] Binding of the binding domain to the substrate In some embodiments, the binding domain is directly or indirectly bound to the substrate. In some embodiments, multiple binding domains can be immobilized on the substrate using site-specific chemistry. For example, in some embodiments, the binding domain may contain a site that can be immobilized on the substrate. Binding of the binding domain to the substrate surface is facilitated by fusing an autocatalytic protein tag (e.g., Spycatcher, Sortase A, SNAP tag, Halo tag, and CLIP tag) to the end of the binding domain. These protein tags on the binding domain can be covalently reacted with cognate reactive sites on the substrate surface. For example, Spycatcher protein can be artificially incorporated into the binding domain. Spytag forms a covalent bond with Spytag protein (13aa peptide). When Spytag is bound to the substrate surface, the binding domain bound to Spycatcher and the reaction of Spytag serves to covalently bind the binding domain to the substrate. Similarly, the binding domain can be fused to a Sortase A tag and reacted with pentaglycine bound to the substrate surface. As another example, the binding domain can be fused to a SNAP tag and reacted with O6-benzylguanine bound to a substrate surface. In some embodiments, the binding domain can be fused to a CLIP tag and used to react with O2-benzylcytosine bound to a substrate surface. In some embodiments, the binding domain can be fused to a Halo tag and used to react with alkyl halides present on a substrate surface.
[0160] In some embodiments, the binding domain may include a biotin moiety, and such binding molecules can be immobilized on a substrate surface by a capture molecule that binds biotin (e.g., streptavidin).
[0161] The binding domain may be attached to the substrate via a Spytag-Spycatcher interaction. This can be achieved by functionalizing the substrate with the Spytag peptide at an appropriate surface density using standard NHS chemistry. Spytag is a short peptide of 13 aa (AHIVMVDAYKPTK; SEQ ID NO: 11). Spycatcher is a 139 amino acid protein that can be engineered to most binding domains: msyyhhhhhh dydipttenl yfqgamvdtl sglsseqgqs gdmtieedsa thikfskrde dgkelagatm elrdssgkti stwisdgqvk dfylypgkyt fvetaapdgy evataitftv neqgqvtvng katkgdahi (SEQ ID NO: 10). When a binding domain modified with SpyCatcher is exposed to a surface coated with SpyTag, the C-terminus of SpyTag and the N-terminus of SpyCatcher spontaneously react to form an isopeptide bond, immobilizing the binding domain.
[0162] Commercially available streptavidin and protein G beads are useful substrates for immobilizing binding domains. In some embodiments, streptavidin beads are functionalized with a mixture of biotinylated adapters and biotinylated protein G. In a second step, protein G is further conjugated to the antibody binding domain by affinity binding (Figure 1D). The surface density of biotinylated adapters and protein G can be adjusted to achieve high yield and specific barcode transfer. In some embodiments, transposase beads can be prepared by conjugating 5' biotinylated ME adapters to streptavidin beads and then loading the ME adapters with Tn5 transposase (Figure 1G). In some embodiments, protein G beads are functionalized with adapters using chemical conjugation of lysines of proteins with amino-modified adapters. In a second step, antibody binding domains are conjugated to protein G (Figure 1B). Here, the labeling stoichiometry of protein G and adapters must be controlled to maintain the ability of protein G to bind to antibodies. In some embodiments, transposase beads can be prepared from Protein G beads by first loading them with an antibody binding domain and then binding the Tn5 transpoase-Protein A fusion protein to the antibody (Figure 1H).
[0163] substrate In some embodiments, the compositions herein include one substrate. In some embodiments, the compositions herein include two or more substrates. In some embodiments, the compositions include multiple substrates, each substrate being formed from the same material. In some embodiments, the compositions include multiple substrates, each substrate being formed from a different material. In some embodiments, the substrate is a bead, chip, plate, tube, slide, dish, gel, or three-dimensional polymer matrix. The substrate can be formed from a variety of materials. In some embodiments, the substrate is a resin, membrane, fiber, polymer. In some embodiments, the substrate includes sepharose, agarose, cellulose, polystyrene, polymethacrylate, and / or polyacrylamide. In some embodiments, the substrate is a polymer, such as a synthetic polymer. A non-limiting list of synthetic polymers includes poly(ethylene) glycol, polyisocyanopeptide polymer, polylactic-co-glycolic acid, poly(ε-caprolactone) (PCL), polylactic acid, poly(3-hydroxybutyrate-co-3-hydroxyvalerate) (PHBV), chitosan, and cellulose.
[0164] In some embodiments, the target nucleic acid can be indirectly attached to the substrate via the binding domain. In some embodiments, the adapter is attached to a surface-activated bead that contains the binding domain. The surface-activated bead can present epoxy, tosyl, carboxylate, or amine groups for covalent attachment. Carboxylic acid beads usually need to be reacted with carbodiimide to facilitate peptide bond formation, and amine beads usually require bifunctional NHS-linkers. In some embodiments, the surface of the beads is passivated to prevent non-specific binding. Passivation can be achieved in some embodiments by co-grafting polyethylene glycol (PEG) molecules with the same linking chemistry. For example, binding domains and amino-terminated polyethylene glycol (PEG) are used such that, on average, most substrate sites are occupied by PEG molecules that serve to spatially separate the binding domains. When an excess of PEG is used, the binding domains are, on average, spatially separated from each other. The surface density of the binding domains can be adjusted by varying the ratio of binding domains to PEG molecules.
[0165] In some embodiments, the beads are sepharose beads made of mTet (tetrazine) and carboxy-PEG. Lowering the ratio of mTet to carboxy-PEG reduces cross-linking between target nucleic acids. In some embodiments, the mTet:carboxy-PEG ratio is 1:500, 1:600, 1:700, 1:800, 1:900, 1:1000, 1:1100, 1:1200, 1:1300, 1:1400, 1:500 or 1:2000. In some embodiments, the mTet:carboxy-PEG ratio is 1:1000.
[0166] In some embodiments, the substrate comprises multiple identical binding domains, hi some embodiments, the substrate comprises multiple identical adaptors.
[0167] Nucleic acid analysis method The compositions described herein (e.g., compositions comprising a binding domain, an adaptor, and a substrate) can be used in various methods of analyzing nucleic acids, particularly to recognize non-canonical features on target nucleic acids. Thus, the present invention provides methods of analyzing non-canonical features on target nucleic acids, including methods of multiplex profiling of RNA and DNA modifications across transcriptomes and genomes. In these methods, the non-canonical features of RNA or DNA are recognized by the binding domain. The adaptor or a portion thereof (e.g., a barcode) is then transferred from the substrate to the target nucleic acid (i.e., to generate a labeled / barcoded target nucleic acid) or to a copy of the target nucleic acid. This step serves to write information from the recognition event into the nucleic acid sequence of the target nucleic acid, since the barcode is unique to the particular non-canonical feature bound by the target nucleic acid. The barcoded target nucleic acid is converted into a sequencing library and read by DNA / RNA sequencing. This step reveals the sequence of the barcode, which correlates with the non-canonical feature of the target nucleic acid. Sequencing can also confirm the localization of non-canonical features in the target nucleic acid. The rapid profiling methods described herein allow for the identification of some or all DNA / RNA modifications in parallel.
[0168] The methods described herein include a series of steps as set forth below. As will be appreciated by one of skill in the art, in some embodiments, various steps may be omitted and / or performed in a different order.
[0169] Contact between the binding domain and the target nucleic acid In some embodiments, the methods described herein include contacting the compositions described herein (e.g., substrates, binding domains, and adapters) with one or more target nucleic acids. The target nucleic acid includes DNA, RNA, or a combination of DNA and RNA. The target nucleic acid can be isolated, for example, from cells or tissues of an organism. In some embodiments, the target nucleic acid can be fragmented.
[0170] The contact between the composition described herein and the target nucleic acid can occur in solution. For example, the composition comprising one or more target nucleic acids can be contacted with one or more compositions comprising substrate, binding domain, and adaptor. In some embodiments, the contact occurs in a dilute solution, and only one binding domain can interact with each target nucleic acid.
[0171] In some embodiments, one or more binding domains can be bound to a substrate, and one or more target nucleic acids can be contacted with the binding domains bound to the substrate.
[0172] The target nucleic acid may be contacted with only one type of binding domain (i.e., to detect one non-canonical feature), or in some embodiments, the target nucleic acid may be contacted with two or more binding domains to detect multiple non-canonical features. For example, the target nucleic acid may be contacted with at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least twenty, at least thirty, at least forty, at least fifty, at least sixty, at least seventy, at least eighty, at least ninety, at least one hundred, at least one fifty, or at least two hundred or more different types of binding domains. In some embodiments, the target nucleic acid may be contacted with 1-5, 5-10, 10-25, 25-50, 50-100, 100-150, 150-175, 175-200, or more different types of binding domains. When multiple binding domains are used, contacting can be simultaneous (i.e., the target nucleic acid is contacted simultaneously with multiple binding domains that recognize different non-canonical features) or sequential (i.e., the target nucleic acid is contacted with a first binding domain that recognizes a first non-canonical feature, followed by a second binding domain that recognizes a second non-canonical feature).
[0173] Barcode Transfer Each binding domain specifically binds to a non-canonical feature of the target nucleic acid, and the adapter bound proximate to the binding domain allows the adapter to interact with either the 3' or 5' end of the target nucleic acid. The adapter (e.g., an adapter that includes or consists of a barcode) can then be transferred to the target nucleic acid. In some embodiments, the adapter is attached to the substrate by a cleavable linker. In some embodiments, when the adapter binds to the target nucleic acid, the adapter is released at the cleavage site. In some embodiments, the transcription is performed in an environment that substantially prevents off-target generation of barcoded nucleic acids. Such an environment is, for example, an environment in which the adapters and binding domains are at a defined density, and each binding domain and its cognate adapter occupies a defined space separate from a second binding domain and its cognate adapter (e.g., each binding domain and adapter pair is on a separate bead, spot, or array that cannot interact with the second binding domain and adapter pair). In some embodiments, the transfer is performed by copying the target nucleic acid, generating a labeled / barcoded copy of the target nucleic acid. For example, when an adapter containing at least a barcode and a universal primer site is transferred to a target nucleic acid, polymerase chain reaction (PCR) can be used to generate a barcoded copy of the target nucleic acid.
[0174] Modification of the target nucleic acid (or a copy thereof) In some embodiments, the method may include modifying the barcoded target nucleic acid(s) or the barcoded copy(s) thereof, which may occur after the binding domain binds to the non-canonical feature, and in some embodiments, may occur after the barcode is transferred to the target nucleic acid (or the barcoded copy of the target nucleic acid is generated).
[0175] The modification is carried out so that the position of the non-canonical feature can be distinguished based on the primary nucleic acid sequence of the barcoded target nucleic acid or its barcoded copy, and can therefore be detected in a downstream sequencing step.For this purpose, many different types of modification are carried out.For example, in some embodiments, the modification can prevent polymerase bypass during copying of the target nucleic acid (or its barcoded copy).
[0176] In some embodiments, modification is accomplished, in part, by chemically modifying the binding domain, which in some embodiments can induce cleavage during copying of the target nucleic acid while the binding domain is bound.
[0177] In some embodiments, the modification comprises photochemically linking the binding domain (or a fragment thereof, such as the binding domain) to the target nucleic acid (or a barcoded copy thereof). Methods for photochemically linking nucleic acids and proteins are known to those skilled in the art. For example, photochemical linking can be induced by exposing a complex comprising the binding domain and the target nucleic acid to ultraviolet (UV) light.
[0178] In some embodiments, the modification includes editing, for example, within 1-20 bases, at or near the site where the binding domain binds to the target nucleic acid. For example, cytosine deaminase or adenosine deaminase can be used to edit bases. The base editing molecule can be attached to the binding domain via a second recognition element. In some embodiments, cytosine deaminase can be genetically fused to protein A and attached to the Fc region of an antibody binding domain. In some embodiments, cytosine deaminase can be genetically fused to Spycatcher and attached to a Spytag-tagged binding domain. Adenosine deaminase converts adenosine (A) to inosine (I), and an amplification enzyme base pairs with cytosine (C) to introduce a thymine (T) to cytosine (C) mutation. Cytosine deaminase converts cytosine (C) near the modification site to uracil (U) and introduces a guanine (G) to adenosine (A) mutation. Another method to localize non-canonical features is the NEB method. (登録商標) USER (商標) The first step is the cleavage of uracil (U) by the enzyme combination uracil deglycosylase and endonuclease VIII.
[0179] Amplification and sequencing After the target nucleic acid (or a barcoded copy thereof) has been modified, it can be amplified and sequenced. This process reveals the sequence of the barcode, which correlates with the non-canonical feature to which the binding domain in the target nucleic acid was originally bound. Sequencing also reveals the length of the cleavage fragment, which allows localization of the non-canonical feature in the target nucleic acid. Sequencing may also reveal mutations near the non-canonical feature, which can informatively derive the location of the non-canonical feature. Mutations may be the result of base editing by deaminase enzymes, or may be the result of increased base insertion error rates of the enzymes (DNA polymerases if the target is DNA, reverse transcriptases if the target is RNA) used to copy and paste the non-canonical feature in the nucleic acid target. The non-canonical feature may naturally increase the enzymatic bypass error rate, or the effect may be amplified by chemically modifying the non-canonical feature.
[0180] Thus, in some embodiments, the methods described herein may include sequencing the barcoded target nucleic acid or a copy thereof. The sequencing step can be performed using any suitable method known in the art. For example, the sequencing can be performed using next generation sequencing (NGS), massively parallel sequencing, or deep sequencing. There are many NGS platforms that can be used in the methods of the present invention. For example, Illumina (登録商標) (Solexa (登録商標) Roche sequencing works by identifying and adding DNA bases to a nucleic acid strand, with each base emitting a fluorescent signal. (登録商標) 454 sequencing is based on pyrosequencing, a technique that uses fluorescence to detect the release of pyrophosphate after a nucleotide is incorporated into a new strand of DNA by a polymerase. Ion Torrent (Proton / PGM Sequencing) measures the direct release of a proton (H+) from the incorporation of individual nucleotides by DNA polymerase.
[0181] In some embodiments, sequencing is not necessary to detect target nucleic acid. For example, PCR can be used to detect target nucleic acid. For example, PCR can be used to detect whether target nucleic acid (e.g., barcode) is present. In some embodiments, target nucleic acid is detected using fluorescent probe (e.g., fluorescently labeled hybridization probe). In some embodiments, microarray or other nucleic acid array is used to detect target nucleic acid.
[0182] In some embodiments, sequencing is not required to detect the addition of barcodes by a reaction mediated by a nucleic acid binding molecule. For example, the presence of DNA / RNA modifications can be confirmed by detecting the associated barcode using nucleic acid electrophoresis, fluorescent hybridization probes, PCR, rolling circle amplification, LAMP, or other nucleic acid amplification methods that can be triggered by the barcode.
[0183] Exemplary methods for identifying and quantifying non-canonical features on a target nucleic acid In some embodiments, the methods described herein can be used not only to identify modifications (i.e., non-canonical features) on a target nucleic acid, but also to quantify the number of modifications present. In some embodiments, the methods described herein are used to identify multiple modifications (i.e., non-canonical features) on multiple target nucleic acids and to quantify the number of each modification present.
[0184] In some embodiments, a method for detecting a non-canonical feature in a target nucleic acid includes: (i) contacting the target nucleic acid with a composition described herein; (ii) (a) transferring a nucleic acid barcode to the target nucleic acid to generate a barcoded target nucleic acid or (b) generating a barcoded copy of the target nucleic acid; and (iii) detecting the presence of the barcode in the target nucleic acid or a copy thereof.
[0185] In some embodiments, a method for detecting and / or quantifying two or more non-canonical features in a plurality of target nucleic acids includes: (i) contacting the target nucleic acid with at least two compositions, each composition comprising a binding domain and an adaptor, wherein the binding domain of each nucleic acid binding molecule binds to a different non-canonical feature of DNA or RNA, and the adaptor comprises a nucleic acid barcode sequence specific to the non-canonical feature specifically bound by each binding domain; (ii) (a) transferring the nucleic acid barcode to the target nucleic acid to generate a barcoded target nucleic acid or (b) generating a barcoded copy of the target nucleic acid; (iii) modifying the barcoded target nucleic acid or the barcoded copy thereof such that the location of the non-canonical feature can be identified based on the primary nucleic acid sequence of the barcoded target nucleic acid or the barcoded copy thereof; and (iv) determining the sequence of the barcoded target nucleic acid. In some embodiments, the method includes amplifying the barcoded target nucleic acid or the copy thereof prior to sequencing.
[0186] In some embodiments, a method for analyzing multiple target nucleic acids includes the steps of: (i) contacting the target nucleic acid with a composition described herein; (ii) (a) transferring a nucleic acid barcode to the target nucleic acid to generate a barcoded target nucleic acid or (b) generating a barcoded copy of the target nucleic acid; (iii) modifying the barcoded target nucleic acid or barcoded copy thereof such that the location of a non-canonical feature can be identified based on the primary nucleic acid sequence of the barcoded target nucleic acid or barcoded copy thereof; and (iv) sequencing the barcoded target nucleic acid.
[0187] In some embodiments, any one or more of the above steps are repeated at least once (e.g., at least two times, at least three times, at least four times, at least five times, at least six times, at least seven times, at least eight times, at least nine times, at least ten times, or more). In certain aspects, one or more of the above steps may be performed simultaneously or sequentially. In some embodiments, the same or different binding domains are used each time steps (i)-(iii) are repeated. In some embodiments, the method includes amplifying the barcoded target nucleic acid or a copy thereof prior to sequencing.
[0188] In one embodiment, an RNA sample is provided that includes modified and unmodified RNA transcripts. Each transcript of the RNA sample may or may not include a non-canonical feature. The RNA transcripts are then contacted with beads, and the beads are directly or indirectly bound to a binding domain specific for the non-canonical feature (i.e., type 1, type 2, and type III beads in FIG. 4A). The modified RNA molecules bind to the beads, while the unmodified RNA remains in the supernatant. To quantify the level of RNA modification, both fractions (substrate-bound and supernatant fractions) can be processed and converted into a sequencing library. The unmodified RNA molecules are capped at both ends with adapters that include UFPs and URPs, while the modified RNA molecules receive a barcode that indicates its modification (i.e., it is transferred from the adapters bound to the beads).
[0189] In some embodiments, the methods described herein include a substrate, where the substrate is a bead. In some embodiments, the substrate is a pool of beads. In some embodiments, each bead comprises a different binding domain. In some embodiments, each bead comprises a different adaptor. In some embodiments, each bead comprises a different binding domain and adaptor, where the adaptor comprises a nucleic acid barcode sequence unique to the non-canonical feature that is specifically bound by the binding domain.
[0190] The present invention provides a method for measuring target genes, comprising contacting a plurality of target genes with a substrate, wherein the substrate is immobilized on a microarray. In some embodiments, the microarray is a spotted microarray. In some embodiments, the microarray is a printed microarray. An example of a microarray is shown in FIG. 4B. In some embodiments, each spot on the microarray comprises a different binding domain and an adaptor, and the adaptor comprises a nucleic acid barcode sequence specific for the non-canonical feature that is specifically bound by the binding domain. In some embodiments, each spot on the microarray comprises a different composition as described herein.
[0191] The present invention provides a method for measuring target genes, comprising contacting a plurality of target genes with a substrate, wherein the substrate is immobilized in a channel of a microfluidic device. In some embodiments, the microfluidic device comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, or 50 channels. An example of a microfluidic device is shown in FIG. 4C. In some embodiments, each channel of the microfluidic device comprises a different binding domain and an adaptor, and the adaptor comprises a nucleic acid barcode sequence unique to the non-canonical feature that is specifically bound by the binding domain. In some embodiments, each channel of the microfluidic device comprises a different composition as described herein.
[0192] In some embodiments, the methods herein comprise analyzing a plurality of target nucleic acids. In some embodiments, the methods comprise contacting a plurality of target nucleic acids with any of the compositions described herein.
[0193] In one aspect, the invention includes a method for analyzing a plurality of target nucleic acids, the method comprising: (i) contacting a solution containing a plurality of target nucleic acids with a composition of the invention, wherein the target nucleic acids containing non-canonical features bind to the binding domain; (ii) do one of the following: (a) transferring a nucleic acid barcode to a target nucleic acid that includes a non-canonical feature to generate a barcoded target nucleic acid; or (b) generating a barcoded copy of the target nucleic acid that includes the non-canonical feature; (iii) amplifying the barcoded target nucleic acid; and (iv) determining the sequence of the barcoded target nucleic acid; wherein steps (i) and (ii) are carried out sequentially or simultaneously.
[0194] In one aspect, the invention includes a method for analyzing a plurality of target nucleic acids, the method comprising: (i) contacting a solution containing a plurality of target nucleic acids with a composition of the invention, wherein the target nucleic acids containing non-canonical features bind to the binding domain; (ii) Do any one of the following: (a) transferring a nucleic acid barcode to a target nucleic acid that includes a non-canonical feature to generate a barcoded target nucleic acid; or (b) generating a barcoded copy of the target nucleic acid that includes the non-canonical feature; (iii) amplifying the barcoded target nucleic acid; and (iv) determining the base sequence of the barcoded target nucleic acid; wherein steps (i) and (ii) are carried out sequentially or simultaneously.
[0195] In one aspect, the invention includes a method for analyzing a plurality of target nucleic acids, the method comprising: (i) providing a plurality of target nucleic acids by reverse transcribing target RNA molecules to form DNA-RNA heteroduplex molecules or by providing target double-stranded DNA molecules; (ii) contacting a solution containing a plurality of target nucleic acids with a composition of the invention, wherein the target nucleic acids containing non-canonical features bind to the binding domain; (iii) using a transposase to transfer two adapters, at least one of which comprises a nucleic acid barcode, to a double-stranded target nucleic acid that comprises a non-canonical feature to generate a barcoded target nucleic acid; (iv) amplifying the barcoded target nucleic acid; and (v) determining the sequence of the barcoded target nucleic acid; wherein steps (ii) and (iii) are carried out sequentially or simultaneously;
[0196] In one aspect, the invention includes a method for detecting a plurality of non-canonical features in a plurality of target nucleic acids, the method comprising: (i) contacting a solution containing a plurality of target nucleic acids with a plurality of compositions of the invention; wherein the number of the plurality of compositions contacted in step (i) is equal to or greater than the number of non-canonical features; wherein the binding domains of the plurality of compositions each bind to a different non-canonical feature of DNA or RNA, or the binding domains bind to the same non-canonical feature of DNA or RNA; and wherein the adaptors of the plurality of compositions each comprise a nucleic acid barcode sequence unique to a non-canonical feature specifically bound by a binding domain of that composition, or unique to a binding domain; (ii) performing one of the following: (a) transferring the nucleic acid barcode sequence of each of the plurality of compositions to a plurality of target nucleic acids, or (b) generating barcoded copies of the plurality of target nucleic acids; (iii) amplifying the barcoded target nucleic acid; and (iv) determining the base sequence of the barcoded target nucleic acid Includes. In some embodiments, the transfer is adapter transfer by a transposase.
[0197] In one aspect, the invention includes a method for detecting a plurality of non-canonical features in a plurality of target nucleic acids, the method comprising: (i) providing a microarray, bead, and / or fluidic device comprising a plurality of compositions of the invention; wherein the number of the plurality of compositions provided in step (i) is equal to or greater than the number of non-canonical features; wherein the binding domains of the plurality of compositions each bind to a different non-canonical feature of DNA or RNA, or the binding domains bind to the same non-canonical feature of DNA or RNA; and wherein the adaptors of the plurality of compositions each comprise a nucleic acid barcode sequence unique to a non-canonical feature specifically bound by a binding domain of that composition, or unique to a binding domain; (ii) contacting a plurality of target nucleic acids with a plurality of compositions, and performing one of the following: (a) transferring a nucleic acid barcode sequence of each of the plurality of compositions to the plurality of target nucleic acids, or (b) generating barcoded copies of the plurality of target nucleic acids; (iii) amplifying the barcoded target nucleic acid; and (iv) determining the base sequence of the barcoded target nucleic acid; Includes.
[0198] In some embodiments, a method for analyzing multiple target nucleic acids includes contacting a solution containing multiple target nucleic acids with multiple compositions described herein, wherein the substrate of each composition is a bead, such as that shown in FIG. 4A.
[0199] In some embodiments, the method of analyzing a plurality of target nucleic acids comprises: (i) contacting a microfluidic device with a solution comprising a plurality of target nucleic acids, wherein the microfluidic device comprises a plurality of channels, each channel comprising a composition described herein, wherein the adapter comprises a nucleic acid barcode sequence unique to a non-canonical feature specifically bound by a binding domain, and each of the compositions binds a different non-canonical feature, thereby binding a plurality of non-canonical features on the target nucleic acids; (ii) either (a) transferring a nucleic acid barcode to a target nucleic acid to generate a barcoded target nucleic acid, or (b) generating a barcoded copy of the target nucleic acid; (iii) amplifying the barcoded target nucleic acid; and (iv) determining the base sequence of the barcoded target nucleic acid Includes.
[0200] In some embodiments, a method for analyzing multiple target nucleic acids includes contacting a solution containing multiple target nucleic acids with multiple compositions described herein, where a substrate for each composition is immobilized on a microarray as shown in FIG. 4B.
[0201] In some embodiments, the method of analyzing a plurality of target nucleic acids comprises: (i) contacting a solution comprising a plurality of target nucleic acids with a plurality of compositions described herein, where each composition is immobilized on a microarray in which an adaptor comprises a nucleic acid barcode sequence specific for a non-canonical feature specifically bound by a binding domain, and where each composition binds a different non-canonical feature, thereby binding a plurality of non-canonical features on the target nucleic acids; (ii) either (a) transferring a nucleic acid barcode to a target nucleic acid to generate a barcoded target nucleic acid, or (b) generating a barcoded copy of the target nucleic acid; (iii) amplifying the barcoded target nucleic acid; and (iv) determining the base sequence of the barcoded target nucleic acid Includes.
[0202] In some embodiments, a method for analyzing multiple target nucleic acids includes contacting a solution containing multiple target nucleic acids with multiple compositions described herein, where a substrate for each composition is immobilized in a channel of a microfluidic device, as shown in FIG. 4C.
[0203] 1. A method for detecting a plurality of non-canonical features in a plurality of target nucleic acids, the method comprising: (i) contacting a solution containing a plurality of target nucleic acids with a plurality of compositions described herein; wherein the binding domains of the plurality of compositions each bind to a different non-canonical feature of DNA or RNA, or the binding domains bind to the same non-canonical feature of DNA or RNA; wherein the number of the plurality of compositions contacted in step (i) is equal to or greater than the number of non-canonical features. wherein each of the adaptors of the plurality of compositions comprises a nucleic acid barcode sequence specific to a non-canonical feature specifically bound by a binding domain of that composition, or specific to a binding domain; (ii) performing any one of the following: (a) transferring the nucleic acid barcode sequence of each of a plurality of compositions to a plurality of target nucleic acids; or (b) generating multiple barcoded copies of the target nucleic acid; (iii) amplifying the barcoded target nucleic acid; and (iv) determining the base sequence of the barcoded target nucleic acid.
[0204] Therefore, this method allows the detection of the same modification at multiple binding domains, each displaying a unique barcode.
[0205] In some embodiments, normalization probes (controls) are spiked into the solution (surface-bound, supernatant) containing the target nucleic acid to allow relative quantification. Furthermore, absolute quantification can be achieved by counting the unique molecular barcodes that may be present in the adapters. Many RNA modifications occur at low copy numbers. Therefore, modified and unmodified fractions of the target nucleic acid can be combined in a ratio that provides optimal sensitivity for low copy number transcripts at a given sequencing depth. This approach allows the stoichiometry and abundance of RNA modifications to be measured. "Stoichiometry" is a relative number, calculated as the copy number of a particular locus containing a non-canonical feature divided by the total copy number of that locus. "Abundance" refers to the absolute occurrence of a non-canonical feature of a nucleic acid at a locus.
[0206] In some embodiments, the method of analyzing multiple target nucleic acids may include RNA profiling by barcode transfer by ligation and localization of non-canonical features by cDNA cleavage. One or more compositions described herein can then be added to the RNA sample. A binding domain of the composition recognizes the RNA modification, and an adapter (e.g., an adapter containing a DNA barcode) is attached to the end of the RNA target. In some embodiments, the target RNA and the binding domain can be crosslinked (e.g., photochemically crosslinked) to generate a mark (i.e., a modification) that prevents reverse transcriptase from copying past the recognition element. In some embodiments, a stop point can be created without crosslinking by selecting and engineering a recognition element that disrupts the polymerase-RNA interaction and / or presents an additional reactive group that can be engaged for the same purpose. Then, single-stranded adapter ligation can be used to provide a primer binding site for reverse transcription, and cDNA can be synthesized by primer extension. The cDNA is synthesized such that the end of the transcript indicates the location of the RNA modification. The resolution at which the modification is localized depends on the nature of the cleavage mechanism.
[0207] The cDNA molecules may be circularized. For example, cDNA molecules with type B adapters can be circularized by cycligase. Cleavage of the circularized cDNA results in strand-specific linear cDNA fragments that can be easily converted into a sequencing library using PCR amplification. Primers can be used to introduce adapter pieces useful for downstream processes such as sequencing.
[0208] In some embodiments, the method of analyzing multiple target nucleic acids can be used to detect / quantify one DNA or RNA modification per reaction. In some embodiments, the method of analyzing multiple target nucleic acids can be adapted to detect multiple DNA or RNA modifications by sample division.
[0209] In some embodiments, the transposase is bound to the substrate as described herein. In some embodiments, tagmentation is used for barcoding. In some embodiments, tagmentation is used for barcoding as shown in FIG. 9. Transposases are found in both prokaryotes and eukaryotes and catalyze the movement of defined DNA elements (transposons) to different locations in the genome in a 'cut and paste' mechanism. The transposase molecule is loaded with a double-stranded DNA adapter that displays a specific RNA modification. The transposase binds to the double-stranded DNA adapter and ligates to the 5' end of the double-stranded DNA substrate, cleaving and inserting the adapter. The 3' end is untagged and the resulting gap can be filled by a polymerase reaction. In some embodiments, the transposase can use a DNA / RNA heteroduplex as a substrate. Tagmentation reactions generally generate fragments of 30-200 nt in length and can be optimized depending on the sample input. In some embodiments, the binding domain-transposase complex is added to unfragmented total RNA or enriched / deleted RNA. Upon recognition of modified RNA bases, the transposase inserts a specific barcode into the RNA / DNA duplex and adds a universal primer site and a reverse primer site. Library preparation is completed by filling in the gaps with an appropriate polymerase. Tagmentation frames the RNA modification site by the specific barcode, and positional information is obtained by engineering the transposase-linker to a length that optimizes positional resolution. In some embodiments, the transposase is a Tn5 transposase.
[0210] Transposases are widely used in many biomedical applications. For example, the hyperactive Tn5 transposase from E. coli can bind to a double-stranded synthetic 19 bp mosaic end (ME) recognition sequence, which can be added to any sequencing adapter. In some embodiments, the ME-adapter comprises CTGTCTCTTATACACATCT (SEQ ID NO: 16). In some embodiments, the ME-adapter comprises AGATGTGTATAAGAGACAG (SEQ ID NO: 24). In some embodiments, the ME-adapter comprises TTTGTGAUGCGATGAACTCAGAGTGCTTNNNNNNNNNNNNAGATGTGTATAAGAGACAG (SEQ ID NO: 52), where N is a barcode. In some embodiments, the mosaic end comprising SEQ ID NO: 16 is hybridized to the ME-adapter comprising SEQ ID NO: 52. Each transposase molecule simultaneously adds two ME-tagged adapters. Tn5 transposase has been utilized in in vitro tagmentation reactions (where target sequences are fragmented and simultaneously tagged with sequencing adapters) with double-stranded DNA or RNA / DNA heteroduplexes as substrates. A major advantage of tagmentation is that it reduces the amount of nucleic acid used, greatly simplifying the assay workflow. Tagmentation is typically performed with picogram amounts of DNA or RNA and has been successfully used in single-cell approaches.
[0211] In some embodiments, the binding domain-enzyme conjugate comprises a binding domain that specifically binds to RNA modification, DNA modification, or both RNA and DNA modification, and guides transposase to target nucleic acid. The transposase that binds to the modification-specific binding domain inserts a specific barcode into the RNA / DNA duplex and adds a universal primer site and a reverse primer site. Tagmentation is dependent on magnesium ions, and can be induced by the addition of magnesium ions. The length of the tagged duplex varies depending on reaction conditions, and can be optimized to be as short as 30 base pairs. Thus, targeted tagmentation can detect DNA or RNA modifications with a base resolution of up to 30 base pairs.
[0212] In some embodiments, the transposase may not be directly tethered or fused to the binding domain that recognizes the DNA / RNA modification. In some embodiments, the transposase may be tethered or fused to a peptide or protein domain that covalently or non-covalently binds to the structural elements of the binding domain that recognizes the DNA / RNA modification. In some embodiments, the binding domain, e.g., an antibody, is genetically fused to a Spy-tag peptide, while the transposase is genetically fused to a SpyCatcher protein. The Spy-tag and SpyCatcher spontaneously form a covalent bond, targeting the transposase to the modification site. In some embodiments, the transposase is genetically fused to Protein A, G, or L. In some embodiments, the transposase is genetically fused to Protein G. In some embodiments, the transposase is genetically fused to Protein L. Protein A, G, or L binds to a specific region of an IgG antibody and directs the transposase activity to the DNA or RNA modification binding antibody.
[0213] In some embodiments, the transposase can bind to ME-tagged adapters covalently linked to the binding domain. The adapters are present as ME-tagged single strands, and hybridization of ME complements triggers in situ loading of the transposase. The binding domain can present two or more ME-adapter molecules, allowing the transposase to add two adapters. In some embodiments, the ME-adapter molecules have the same sequence. In some embodiments, the ME-adapter molecules have different sequences. In some embodiments, the ME-adapter comprises a barcode specific for a DNA or RNA modification.
[0214] The methods described herein can be used to diagnose a disease, disorder, or condition. For example, in some embodiments, the methods can be used to diagnose cancer in a subject in need thereof. In some embodiments, the kits can be used to monitor a disease, disorder, or condition over time, such as response to one or more treatments. For example, the kits can be used to monitor epigenetic and / or epitranscriptomic changes over time in a subject undergoing cancer treatment (i.e., chemotherapy, radiation therapy, etc.). In some embodiments, the methods can be used to analyze cells or tissues from a subject in need thereof. For example, the methods can be used to detect non-canonical features in cells or tissues isolated from blood samples, biopsy samples, autopsy samples, etc.
[0215] In some embodiments, the methods may be used to detect and / or monitor epigenetic changes in cells that are used commercially to produce one or more products, such as cells used in industrial fermentation. In some embodiments, the methods may be used to detect and / or monitor epigenetic changes in plant cells or tissues.
[0216] Nucleic Acid Analysis Kits The compositions described herein can be provided in a kit (e.g., as a component of a kit). For example, the kit can include the composition, or one or more components thereof, and informational material. In some embodiments, the kit includes two or more compositions described herein. The informational material can be, for example, explanatory, instructional, marketing, or other material regarding the use of the methods and / or compositions described herein. The format of the informational material of the kit is not limited. In some embodiments, the informational material can include information regarding the manufacture of the composition, molecular weight, concentration, expiration date, batch or origin information, etc. In some embodiments, the informational material can include a list of disorders and / or conditions that can be diagnosed or evaluated using the kit.
[0217] In some embodiments, the compositions may be provided in a manner suitable for use in the methods described herein (e.g., easy-to-use tubes, appropriate concentrations, etc.). In some embodiments, the kits may require some preparation or manipulation of the composition prior to use. In some embodiments, the compositions are provided in liquid, dry, or lyophilized form. In some embodiments, the compositions are provided in aqueous solutions. In some embodiments, the compositions are provided in sterile, nuclease-free solutions. In some embodiments, the compositions are substantially free of nucleic acids other than those that may constitute the molecule itself.
[0218] In some embodiments, the kit may include one or more syringes, tubes, ampoules, foil packages, or blister packs. The containers of the kit may be airtight, waterproof (i.e., to prevent changes in moisture or evaporation), and / or light-tight.
[0219] In some embodiments, the kit can be used to perform one or more of the methods described herein, such as the method of analyzing a population of target nucleic acids. In some embodiments, the kit can be used to diagnose a disease, disorder, or condition. For example, in some embodiments, the kit can be used to diagnose cancer. In some embodiments, the kit can be used to monitor a disease, disorder, or condition over time, such as response to one or more treatments. For example, the kit can be used to monitor epigenetic and / or epitranscriptomic changes over time in a subject undergoing treatment for cancer. EXAMPLES
[0220] The following non-limiting examples further illustrate aspects of the compositions and methods of the present invention.
[0221] Example 1: Selection of binding domains Binding domains specific for pseudouridine, inosine, m5C, and m6A are selected based on their binding rates (on-rates) and dissociation rates (off-rates) measured by biolayer interferometry (BLI). First, commercially available antibodies are screened. The goal is to identify antibodies with minimal off-rates and high specificity.
[0222] The BLI instrument (Gator Prime) was loaded with a Protein G probe (Gator Bio, Cat. No. 160006). The Protein G probe has the capacity to bind 0.02-2000 μg / mL of IgG antibodies of most isoforms. The IgG antibodies are immobilized on the Protein G probe (5 μg / μL antibody in phosphate-buffered saline (PBS)) at a density corresponding to a 1 nm shift in the BLI signal. The real-time on-rate of the antigen is obtained by immersing the BLI probe in a 1-250 nM solution of the RNA target exhibiting one or more modifications. The off-rate is generated by transferring the probe into PBS buffer without antigen. The same procedure is repeated with the unmodified RNA strand. Depending on the molecular weight of the test RNA analyte, it may be necessary to amplify the signal by binding a high molecular weight reporter molecule to the RNA, for example by using biotin-labeled RNA conjugated with streptavidin. The antibody with the lowest off-rate and the highest off-rate selectivity for a particular target (off-rate 特異的 / Off Rate 非特異的 ) are selected for further characterization.
[0223] Example 2: Preparation of beads with covalently attached antibodies and DNA adapter molecules This example outlines the preparation of bead surfaces with covalently attached antibodies and DNA adapters (Figure 1A). Antibodies are site-specifically linked to maintain activity, and the densities of antibody and DNA adapters are independently tunable. A 10-fold excess of adapters over antibodies is used to achieve efficient barcoding yields while minimizing side products.
[0224] Standard 1-ethyl-3-(-3-dimethylaminopropyl)carbodiimide hydrochloride (EDC) was used to bind carboxylated magnetic beads (Thermo Fisher, Dynabeads (登録商標)M-270 Carboxylic Acid) is activated for amine coupling. The EDC-activated surface is functionalized with a ternary mixture of a passivating molecule (COOH-PEG4-amine, Broadpharm Cat. No. BP-20423), an antibody-reactive linker (DBCO-PEG10-amine, Broadpharm Cat. No. BP-24181), and a DNA-reactive linker (mTET-PEG3-amine, Broadpharm Cat. No. BP-26276). The antibody is activated for DBCO coupling using site-click chemistry (Thermo Fisher, Cat. No. S20026). Site-click chemistry introduces azide groups at the glycosylation sites of the Fc region of an IgG antibody. The amino-modified DNA adapter is functionalized with TCO-PEG4-NHS ester (Broadpharm, Cat. No. BP-22418).
[0225] To generate a surface with 3' immobilized ligation barcodes, the following general construction technology adapters are used (SEQ ID NO:1): TIFF2024541478000013.tif20164The 5' end is phosphorylated to enable enzymatic ligation, followed by a 7b barcode (underlined) indicating RNA modification, a unique molecular barcode of at least 3 bases (NNN, where N is any nucleotide), an Illumina adapter (bold), an 18-atom hexaethylene glycol spacer (iSp18), one uracil surrounded by filler AT repeats for release from the surface by cleavage with the USER enzyme (NEB), and a 3' amino moiety (3AmMO).
[0226] The surface with the 5' immobilized primer extension barcode is SEQ ID NO:2 It was prepared using the general assembly technique of TIFF2024541478000014.tif19160, where 5AmMC6 is the 5'-amine moiety and the CACTCAGT sequence is a spacer for barcoding by primer extension.
[0227] The final functionalization of the beads is performed in a stepwise manner: first, an azide-activated antibody is immobilized at the DBCO site, followed by filling of the mTet site with a TCO adaptor.
[0228] Example 3: Preparation of adapter-loaded Protein G and antibody-displaying beads This example describes an alternative to Example 2. Instead of immobilizing the DNA adapters directly on the bead surface, they are attached to Protein G (Figure 1B), which also serves to immobilize IgG antibodies by affinity binding.
[0229] The lysine residues of protein G on the magnetic beads (Thermo Fisher) are labeled with S-HyNic linkers (Vector Labs, Cat. No. 50-204-5741). The size of full-length protein G isolated from Streptococcus pyogenes is 63 kDa, but commercial versions have been engineered to be smaller while maintaining subnanomolar affinity for IgG antibodies (e.g., Abcam, Uniprot ID: P19909). To protect the IgG-binding sites of protein G from functional damage, the HyNic reaction is performed in the presence of a sacrificial IgG antibody that is eluted with 0.2 M glycine pH 2 after labeling. The HyNic modification reacts readily with DNA adapters (e.g., SEQ ID NO: 1 or 2) whose amine groups have been activated with S-4FB linkers (Vector Labs, Cat. No. 50-204-5743).
[0230] The bead preparation is complete once the sacrificial antibody is removed and loaded with the RNA modification-specific antibody of interest.
[0231] Example 4: Preparation of planar arrays of antibodies In this example, we use DNA microarray technology to immobilize antibodies on a flat surface via DNA hybridization probes (Figure 1F and Figure 4B). After patterning, 48 spots are formed on the surface, with one RNA modification-specific antibody on each spot, along with an RNA modification-specific barcoded i7 adapter and a universal i5 Illumina adapter. In this example, the i7 adapter contains one uracil for cleavage with the USER enzyme mix, and the i5 adapter contains one 8-oxoG for cleavage with the FpG enzyme. The goal is to incorporate the patterned surface into a flow cell to allow for clonal amplification of the captured nucleic acid sequences and subsequent in situ sequencing. Selective cleavage of the forward or reverse strand is an essential step for linearization of the strand preceding read 1 and read 2, respectively. The flow cell is mounted on a Peltier element and connected to a pump-driven fluidics system, allowing for automation of temperature control and liquid exchange. These features are exploited to build a fully automated library preparation workflow, as outlined in Example 6.
[0232] Microscope slides are patterned by inkjet printing of synthetic DNA probes and assembled into flow cells by a common adhesion procedure. There are 48 spots on the microscope slide, with each spot containing a mixture of three different oligonucleotides. The i7 adapters exhibit an 8b spacer region at their 3' ends to allow barcoding by primer extension according to Figure 2D. The density of the DNA probes is experimentally optimized to facilitate barcoding. RNA modification-specific antibodies are site-specifically labeled at DNA addresses using site-click chemistry (Thermo Fisher). Antibodies are loaded onto the array by hybridization to capture probes via the DNA addresses.
[0233] Example 5: RNA modification-specific barcoding by ligation using bead pools In this example, we describe a workflow for profiling RNA modifications using a bead pool prepared according to Example 2. Each bead type is presented with an antibody targeting one RNA modification and a DNA adapter to which a barcode is transcribed by ligation into the target RNA (Figure 5). The identity of the barcode is determined by next-generation sequencing, revealing the properties of the RNA modification.
[0234] Four types of beads are prepared: bead type 1 presents m6A antibody and DNA adapter for barcoding by ligation (SEQ ID NO:1). Three more types of beads are generated using antibodies against m5C, pseudouridine, and inosine, as well as DNA adapters with different barcodes (SEQ ID NO:3-5). The beads are pooled and incubated with 100 ng of RNA sample that has been chemically fragmented to an average size of 100 b and dephosphorylated. After washing the unmodified RNA, the 3' ends of the modified RNA are ligated to the surface-bound adapters by the action of T4 RNA ligase 1. The DNA adapters are primed and first- and second-strand synthesis is carried out in a single reaction containing dNTPs, DTT, template switching oligonucleotide (AGACGTGTGCTCTTCCGrGrGrG, where r represents ribonucleotide; SEQ ID NO:6), SuperScript IV reverse transcriptase, and appropriate enzyme buffer. The resulting cDNA library is PCR amplified to introduce the complete Illumina adapter and sequenced.
[0235] Example 6: RNA modification-specific barcoding by on-flow cell primer extension followed by amplification In this example, patterned arrays prepared according to Example 4 are used for profiling RNA modifications. The advantage of patterned arrays is that they can be integrated into a fluidics system, allowing a fully automated library preparation workflow. In this example, we detect all eight RNA modifications present in mRNA (m5C, m6A, m7G, m1A, m3C, ac4C, inosine, pseudouridine). For each modification, the array displays at least three cognate antibody spots.
[0236] RNA samples are chemically fragmented by treatment with magnesium chloride at 95°C. RNA fragments are dephosphorylated with shrimp alkaline phosphatase and T4 polynucleotide kinase. The RNA solution is contacted with an antibody array, where modified RNA strands are specifically captured by antibodies, and RNA fragments are separated into spots according to the degree of modification. The 3' spacer of the RNA strand hybridizes to an Illumina i7 adapter, and the adapter is extended by Superscript IV reverse transcriptase to generate barcoded cDNA strands (Figure 6). An i5 adapter complement is added to the 3' end of the first strand by including a template switching oligonucleotide in the reverse transcription reaction, as described in Example 5. Treating the surface with 0.1 M sodium hydroxide hydrolyzes the RNA and strips off the antibody. DNA is amplified by temperature cycling in the presence of a thermostable DNA polymerase (e.g., Bst polymerase) (Figure 7). The temperature protocol involves three steps: (1) annealing of DNA and surface-bound adapters at 37 °C, (2) extension of the adapters at 60 °C, and (3) denaturation at 60–95 °C in the presence of denaturing agents such as formamide, ethylene glycol, betaine, or propanediol to lower the melting temperature. This step creates clonal copies of the barcoded cDNA. When this process is performed in the presence of low antibody density, it produces spatially separated monoclonal clusters that are suitable for direct sequencing by synthesis (SBS) (Figure 8).
[0237] Example 7: Barcoding of m6A modified RNA using immobilized transposase Here, we utilize antibody-mediated pull-down of RNA modifications followed by transferase to introduce barcodes into modified RNA fragments in a rapid, one-step reaction (Figure 9 and Figure 22A).
[0238] Tagmentation is a well-established process for NGS library preparation and refers to the Mg ion-dependent "cut-and-ligate" activity of the Tn5 transposase, which selectively binds short 19 bp "mosaic end (ME)" duplexes that can be appended to DNA adapters used for tagmentation.
[0239] In this example, the surface of the beads is loaded with transposomes (FIG. 1G). Transposomes are composed of transposase dimers carrying two mosaic end (ME)-containing adapter molecules. As used herein, ME and ME' (mosaic end and mosaic end prime, respectively) are used to represent the double-stranded sequence 5'-CTG TCT CTT ATA CAC ATC T-3' (SEQ ID NO: 7) to which Tn5 transposase spontaneously binds. This sequence can be fused to any DNA sequence, such as a universal primer site or an Illumina adapter fragment (see, e.g., FIG. 13).
[0240] Streptavidin beads were loaded with an equimolar ratio of Illumina i5 and i7 ME adapters along with m6A antibody at 5%, 10%, 20% or 40% of the total loading capacity (Figure 1G). The sequence of the i7 ME adapter was as follows: The sequence of the ME adapter was as follows: TIFF2024541478000016.tif29160ME sequence is shown in bold, NNNNNNNN indicates barcode.
[0241] After bead preparation, a mixture of unmodified IVT RNA and m6A modified IVT RNA was reverse transcribed using Superscript IV reverse transcriptase. The reverse transcribed RNA samples were then immunoprecipitated using streptavidin beads co-immobilized with ME adapters and m6A antibody. After washing the beads, Tn5 transposase (Diagenode, Cat. No. C01070010-10) was loaded onto the ME adapters in binding buffer (50 mM HEPES pH 7.5, 300 mM NaCl, 0.1 mM EDTA, 0.05% Tween® (polysorbate)-20). Mg 2+ Addition of the tagmentation buffer (10 mM Tris-HCl pH 8.5, 5 mM MgCl2, 10% DMF) inserts adapters into the captured DNA-RNA duplex. In this format, the tagmentation products were reliably captured on the beads and subjected to gap-fill PCR (0.5 uM forward primer, 0.5 uM reverse primer, NEBNext Ultra II Q5 (catalog no. M0544X, New England Biolabs) for 17-19 cycles (5 min at 72°C, 2 min at 98°C, n cycles of 10 s at 98°C-75 s at 65°C, and a final extension of 5 min at 65°C). The DNA libraries were sequenced and barcodes were deconvoluted and sequence aligned.
[0242] The coverage plot (Figure 22B) shows significant enrichment of m6A-containing fragments, proving selective tagmentation of m6A-modified RNAs. Loading the beads with ME adapters at 5% or 10% of the total binding capacity slightly improved the enrichment signal-to-noise and also resulted in higher library prep yields than higher ME densities (data not shown). This experiment demonstrates that it is possible to detect m6A-modified RNAs using beads with co-immobilized ME adapters and m6A antibodies. To detect multiple RNA modifications, multiple bead types, each representing one antibody, and ME adapters with barcodes encoding the antibodies can be mixed.
[0243] Example 8: Read phasing of long RNAs with multiple m6A modifications This example extends Example 7 by introducing a base editing step to mark the positions of multiple modifications of the same type (Figure 10).
[0244] The full-length RNA strand is reverse transcribed and captured with beads displaying m6A antibody and biotin-labeled ME adapter. After washing, the ADAR-Protein L conjugate is introduced. Protein L binds specifically and with high affinity to the light chain of the IgG antibody. The ADAR enzyme edits the double-stranded RNA and the DNA strand in the DNA / RNA heteroduplex with an A>I (inosine) mutation. The ligation structure of the ADAR-Protein L conjugate is such that it restricts ADAR activity to the direct vicinity of the m6A modification. The adenine to inosine mutation (A-to-I) introduced by the ADAR marks the position of the m6A. After base editing, the transposome is assembled by binding of Tn5 transposase to the ME adapter ligated to the surface. Translocation tag sequencing allows the identification of reads originating from the same molecule with the same barcode and allows the reconstruction of long transcripts from short sequence reads (Figure 11).
[0245] Example 9: Binding Domain Selection Binding domains specific for pseudouridine, inosine, m5C, and m6A were selected based on their on-rate and off-rate measured by BioLayer Interferometry (BLI). First, commercially available antibodies were screened to measure the on-rate and off-rate of the antibodies and to correlate their properties with their performance in the barcode assay.
[0246] A streptavidin probe (catalog no. 160002, Gator Bio) was added to the BLI instrument (Gator Prime). 5' biotinylated RNA oligos with m5C, inosine, m6A or pseudouridine bases in the center were immobilized with a sparse surface coverage to ensure the formation of 1:1 antibody:RNA complexes. Oligos without base modifications served as negative controls. Real-time on-rates of the antigen were obtained by immersing the BLI probe in 1-250 nM solutions of the antibody. Off-rates were obtained by transferring the probe to PBS buffer without antibody. The same procedure was repeated with the unmodified RNA strand.
[0247] Figures 14A-14G show the on-rates and off-rates of several commercially available antibodies against m6A, m5C, inosine and pseudouridine (Ab02(m6A)=MA5-33030, Thermo Fisher; Ab05(m6A)=345E11, Synaptic Systems; Ab08(m6A)=Rb212B11, Synaptic Systems; Ab09(m6A)=C15200082-50, Diagenode; Ab10(inosine)=C15200251, Diagenode; Ab16(m5C)=MA5-24694, Thermo Fisher; Ab19(pseudouridine)=D347-3, MBL). The on-rates for specific antigen binding are 10 4 From 10 5 M -1 s -1 whereas the off-rate is 10 -4 From 10 -2 s -1 The corresponding dissociation constant K D The K ranges from 3.5 to 150 nM. In general, negligible binding was observed with the negative oligo controls, confirming target specificity. Based on the ELISA data, most antibodies bind weakly to unmodified RNA and their K D is 100-500 times greater than that for a specific target.
[0248] All of the antibodies shown in Figures 14A-14G are useful in antibody-mediated barcoding assays (see Example 6), demonstrating the suitability of this method for a variety of antibody characteristics. Antibodies with nanomolar affinities are readily available through hybridoma technology, demonstrating the versatility of this method. D Values above 150 nM result in low capture efficiency of RNA targets.
[0249] Example 10: Preparation of beads for immunoprecipitation and barcoding of modified nucleic acids In this example, two IgG antibodies were loaded onto protein G magnetic beads by affinity binding. One antibody specifically binds to one type of nucleic acid modification. The other antibody has no nucleic acid binding activity but is labeled with a DNA adapter (reporter antibody) (Figure 1E). Part of the adapter design is a modification barcode (MBC). In the following example, the nucleic acid modification was detected by transferring the MBC to the target nucleic acid in an antibody-mediated reaction. The purpose of this bead structure, and especially the reporter antibody, is to present the DNA adapter in a spatial orientation that significantly facilitates the transfer of the barcode from the adapter to the target nucleic acid without directly labeling the antibody (see Figure 1E). Each bead was "monoclonal" and loaded to contain a single type of REPA with a single modification-specific antibody and a unique MBC (Figure 15C).
[0250] The reporter antibody was mTET-PEG5-NHS ester (catalog number BP-22945, Broadpharm). Any IgG antibody without nucleic acid binding activity could be used, for example, monoclonal anti-bovine serum albumin antibody (catalog number MA1-82941, Thermo Fisher). Coupling of mTET-NHS ester and reporter antibody was performed in phosphate buffered saline (PBS) containing up to 1 mg of antibody and 30 mol equivalents of linker. The reaction was allowed to proceed for 12 hours at 25°C, and the resulting antibody-linker complex was purified on a 7 kDA MWCO Zeba desalting column (catalog number 89882, Thermo Fisher) to remove excess linker. In a separate reaction, the adapter DNA oligo (e.g., / 5AmMC6 / T / iSp18 / / iSp18 / / AGACGTGTGCTCTTCCGATCTNNNCAGCTTTCACTCAGT, where 5AmMC6 is the 5' amino modification and iSp18 is the PEG spacer (SEQ ID NO: 23), Integrated DNA Technologies) was incubated with trans-cyclooctene (TCO)-PEG4-NHS ester (catalog no. BP-22418, Broadpharm) in PBS buffer at 25°C for 12 hours. The final product was purified by acetone precipitation. The iSp18 linker unit provides both spatial flexibility and reach, and is necessary for barcoding in the described format. The final adapter-labeled reporter antibody was prepared by incubating the mTET antibody with a stoichiometric equivalent of TCO. Because the antibody was highly labeled with mTET, the final labeling rate was determined by the molar equivalent of TCO-labeled adapter reacting in quantitative yield. Analysis of the size of the resulting antibody-oligo complexes by denaturing SDS gel electrophoresis shows that the labeling stoichiometry titrates proportionally to the TCO-oligo excess (Figure 15A). A 3.5-fold molar excess of TCO-oligo is ideal to eliminate the generation of unmodified antibodies while preventing overlabeling that can inhibit Protein G binding. This procedure generated reporter antibodies that display an average of 2-3 adapters, regardless of IgG subtype and adapter sequence.
[0251] In a standard barcoding reaction, 2 μL of Protein G Dynabeads (catalog no. 10004D, Thermo Fisher) were loaded with a total of 0.5ug of mixture containing modification-specific and reporter antibodies. Antibodies were loaded in PBST for 30 minutes at room temperature, and excess antibodies were removed by washing three times with PBST. Typically, a 50:50 mixture of nucleic acid-specific and reporter antibodies was used. Varying the ratio, ranging from 20% to 80% reporter antibody, does not significantly affect barcoding specificity but does alter barcoding yield. Barcoding yield (ratio of barcoded RNA molecules divided by captured RNA molecules) increases with increasing surface density of reporter antibody, as measured by capture, barcoding, elution, and denaturing gel electrophoresis of dye-labeled modified RNA (Figure 15B).
[0252] Example 11: Preparation of in vitro transcribed RNA using modified bases as a truth model This example describes the preparation of RNA targets with known modification content. The resulting modified RNA targets were used as truth sets in the barcoding experiments described below.
[0253] In vitro transcribed (IVT) RNA was analyzed using HiScribe (商標) The manual for the T7 High Yield RNA Synthesis Kit (catalog number E2040S, New England Biolabs) was followed. The template DNA amplicon used in the IVT reaction was prepared by amplifying a region of genomic phage or bacterial DNA using a primer with a T7 promoter sequence and ligating it with PureLink RNA. (商標)Amplicons were purified using a PCR Purification Kit (Cat. No. K310001, Thermo Fisher Scientific). The following genomes were used for T7 tag amplicon generation (New England Biolabs): ΦX174 viral DNA (Cat. No. N3023S), M13mp18 single-stranded DNA (Cat. No. N4040S), lambda DNA (Cat. No. N3011S), and FLuc control plasmid (Cat. No. E2040S). IVT reactions were performed using a T7 promoter representing the PCR amplicon as input and 10–50% of the natural NTPs were replaced with modified NTPs such as methyl adenosine-5'-triphosphate (m6ATP, Catalog No. N-1013-5, TriLink), inosine-5'-triphosphate (ITP, Catalog No. N-1020, TriLink), 5-methylcytidine-5'-triphosphate (m5CTP, Catalog No. N-1014, TriLink) or pseudouridine-5'-triphosphate (YTP, Catalog No. N-1019, TriLink). IVT reactions were treated with DNAse I (Cat. No. M0303S, New England Biolabs) to remove the DNA template and Monarch (登録商標) Purification was performed using RNA Cleanup Columns (catalog number T2047L, New England Biolabs).
[0254] This procedure generated model target pools consisting of IVT RNAs derived from different genomes: for example, PhiX RNA was unmodified, Fluc RNA contained m6A, M13mp18 RNA contained m5C, and Lambda RNA contained inosine.
[0255] Model RNA pools with known modifications were used in barcoding experiments and sequenced. Barcoding specificity was determined by aligning reads from immunoprecipitated and barcoded samples, counting the number of RNA fragments that display the correct modification barcode (MBC), and normalizing the results to the input sample.
[0256] Example 12: Preparation of RNA samples for downstream modification analysis by addition of universal spacer sequences This example provides a protocol to add spacer sequences to a pool of RNA molecules. During proximity encoding, the spacer binds to the spacer' complement of a bead-anchored adaptor and is extended by DNA polymerase (Figure 2D) or reverse transcriptase (Figure 2G).
[0257] RNA was fragmented by incubation in 1X T4 RNA Ligase I buffer (New England Biolabs) at 90°C for 8-25 minutes. This treatment resulted in peak fragment sizes of 60-150 bases. The 3' ends of the RNA were then dephosphorylated by the addition of T4 Polynucleotide Kinase (Cat. No. T4PK-200, MCLab) in the presence of RNase inhibitor (Cat. No. AM2694, Thermo Fisher) for 30 minutes at 37°C. The RNA was barcoded by ligating a spacer via primer extension with DNA polymerase or reverse transcriptase. The spacer was ligated for 1 hour at 20°C in a reaction containing 0.3 units / uL T4 RNA Ligase I, 10uM spacer ( / 5Phos / NNACTGAGTG), 1X T4 RNA Ligase I buffer, 1mM ATP, 1mM DTT, 15% PEG-8000, 0.2 units / uL RNase inhibitor. The spacer-ligated RNA was purified with 1X RNAClean XP beads (catalog no. A63987, Beckman Coulter) and then prepared for barcoding assays. Figure 16 shows the fragment sizes obtained after fragmenting a mixture of 1.5 kb IVT RNA fragments and then ligating a spacer. Addition of the spacer increased the apparent fragment size from 104 nt to 109 nt.
[0258] Example 13: Multiplex detection of m6A, inosine and m5C using reverse transcription encoding, template switching and protein G beads In this example, we describe an end-to-end library preparation workflow that integrates a barcoding step for detecting RNA modifications. Barcoding is achieved by bidirectional extension of the RNA target and adapters using reverse transcriptase (Figure 17A).
[0259] To detect m5C, m6A and inosine in an RNA sample, a minimum of three beads are required and are prepared according to Example 10. The first bead type is anti-m6A (catalog no. 345E11, Synaptic Systems) and MBC-3 The second bead type is anti-inosine (Cat. no. C15200251, Diagenode) and MBC-4 The third bead type is a reporter antibody conjugated to an adaptor containing anti-m5C antibody (catalog no. MA5-24694, Thermo Fisher Scientific) and MBC-5 TIFF2024541478000019.tif20160 is a reporter antibody with an adaptor.
[0260] The adaptor included a spacer' sequence (bold at the 3' end), MBC (underlined), UMI (NNN) and an i7 Illumina adaptor (5' sequence of UMI).
[0261] Equal volumes of each loaded bead type were combined for each sample. The first assay step was immunoprecipitation (IP) of the spacer-linked RNA prepared according to Example 4. The bead pool, 0.5–50 ng of RNA, and 10 units / uL of RNase inhibitor were incubated in 1XPBST. After incubation, the beads were washed with PBST buffer and resuspended in 1X Superscript IV reverse transcription buffer (catalog no. 18090050, Thermo Fisher). The washes removed nonspecifically bound RNA, while the specific RNA modification-antibody complexes were preserved. In the next step, MBC, which contains the i7 adapter and the universal i5 adapter, was added to the target RNA. In this step, reverse transcriptase extends the 3' end of the RNA target to copy the MBC and i7 adapter, while simultaneously extending the 3' end of the adapter to synthesize cDNA. Template switching required a reverse transcriptase with terminal deoxynucleotidyl transferase (TdT) activity, such as M-MLV variants Superscript II or IV (catalog no. 18064014 or 18090200, Thermo Fisher), Maxima H-minus (catalog no. EP0751, Thermo Fisher) or Smartscribe reverse transcriptase (catalog no. 18064014, Takara Bio). The TdT activity adds a C-tail to the end of the DNA / RNA heteroduplex, allowing binding and copying of a template switching oligo (TSO) containing an Illumina i5 adapter and ending with three G bases.
[0262] The IP beads were added to the reverse transcription reaction (1X SSIV buffer, 0.5u / uL Superase-In, 5u / uL SSIV reverse transcriptase, 1mM dNTPs, 2uM template switching oligo "TSO") and incubated at 23°C for 15 minutes, then 50°C for 60 minutes. Several versions of TSO were used, for example, CTACACGACGCTCTTCCGATCTrGrG+G (rG is riboG and +G is LNA-G) (SEQ ID NO:28), CTACACGACGCTCTTCCGATCTrGrGrG (SEQ ID NO:29), or It worked well, as shown in TIFF2024541478000020.tif10160. After the reaction was completed, the supernatant was subjected to 10-13 cycles of standard Illumina index primers (0.5uM forward primer, 0.5uM reverse primer, NEBNext Ultra II Q5 (catalog no. M0544X, New England Biolabs) (n cycles of 98°C for 30 seconds, 98°C for 10 seconds, 65°C for 75 seconds, and 65°C for 5 minutes).
[0263] The libraries were sequenced, and by bioinformatic deconvolution of the MBCs attached to each RNA fragment, RNA modifications were identified and localized to specific loci. Figure 17B shows the sequencing results obtained with the described barcoding method using 5 ng of pooled IVT as input and SuperScript IV reverse transcriptase. Figure 17C shows the same results using Maxima Minus reverse transcriptase for encoding. In this example, the IVT RNA pool consisted of 70% unmodified PhiX RNA, 10% m6A modified FLuc-RNA, 10% inosine modified Lambda RNA, and 10% m5C modified M13 RNA. The plot shows that most of each MBC is bound to the correct genome, and the SuperScript IV dataset showed a better signal-to-noise ratio.
[0264] Example 14: Multiplex detection of m6A, inosine, and m5C using encoding DNA polymerase In this example, we describe a different version of barcoding by primer extension, providing an alternative to library preparation by template switching. Similar to barcoding by reverse transcription, this workflow requires ligation of a spacer sequence to the upstream RNA pool. After immunoprecipitation of the spacer-extended RNA, a DNA polymerase (Klenow fragment exo-) was used to add the barcode to the target RNA by top-strand primer extension (Figure 19A).
[0265] To detect m5C, m6A and inosine in RNA samples, three types of beads were prepared as described in Example 10. However, in this example, the 3' end of the adapter sequence was blocked from extension, e.g., by / 3SpC3 / (see nomenclature by Integrated DNA Technologies). Bead loading and IP followed the same protocol as described in Example 13. After IP washing, beads representing captured RNA were washed with 1X Klenow buffer (50 mM Tris pH 7.9, 2 mM MgCl2, 50 mM NaCl, 0.1% Tween 10 ... (登録商標) -20) and resuspended in an equal volume of barcoding mix (200uM dNTPs, 0.5 units / uL Klenow fragment exo- (Cat. no. KPIM-200, MCLAB), 50mM Tris pH 7.9, 2mM MgCl2, 50mM NaCl, 0.1% Tween (登録商標)-20) was added. The Klenow reaction was carried out for 5 minutes at room temperature. Barcoded RNA was eluted from the beads by incubation in water with 5 mM DTT and 1 mM EDTA at 37°C for 5 minutes. The eluted RNA was added i5 adapter (2uM I5 RNA adapter ( / 5SpC3 / rCrUrArCrArCrGrArCrGrCrUrCrUrUrCrCrGrArUrCrU) (SEQ ID NO: 31), 1XT4 RNA ligase buffer, 1mM ATP, 10% PEG-8000, 0.5u / uL Superase-in, 1u / uL T4 polynucleotide kinase, 1u / uL T4 RNA ligase 1) and incubated for 1 hour at room temperature. After washing with 3X Ampure beads, the adaptor-ligated RNA was reverse transcribed at 55°C for 10 minutes (1 uM cDNA primer (AGACGTGTGCTCTTCCG) (SEQ ID NO: 32), 0.5 mM dNTPs, 1X SSIV buffer, 5 mM DTT, 2 u / uL RNAseOUT, 10 u / uL SuperScript IV reverse transcriptase). Optionally, the cDNA can be NaOH treated at this point, neutralized, and washed with 3X Ampure beads, or used directly as input for index PCR (cDNA, 0.5 uM forward primer, 0.5 uM reverse primer, NEBNext Ultra II Q5) for 10-13 cycles (98°C for 30 seconds, followed by n cycles of 98°C for 10 seconds, 65°C for 75 seconds, and 65°C for 5 minutes).
[0266] Using this workflow, BLI-characterized antibodies were screened in singleplex experiments (Example 9 and Figures 14A-14G). That is, one bead loaded with modification-specific and reporter antibodies is exposed to an IVT RNA pool containing m6A-modified Fluc RNA, inosine-modified Lambda RNA and 5C-labeled M13 RNA, and optionally unmodified PhiX RNA. For each antibody, at least 80% of the MBCs were associated with the correct genome based on sequencing analysis (Figures 18A-18G). Combining three beads in a 3plex reaction (Figure 19A) gave similar results, although the background was slightly increased compared to the corresponding singleplex reaction (Figure 19B).
[0267] Example 15: Simultaneous detection of m6A and m5C using splint ligation for encoding This example introduces modification-specific barcodes by enzymatic ligation rather than primer extension. Specifically, this example uses DNA splint ligation catalyzed by T4 DNA ligase (Figure 20A).
[0268] In this example, the adapter was attached to the reporter antibody via the 3'-amine group and presented with a 5'-phosphate for ligation (see Example 10). In addition, a uracil base was introduced to allow cleavage of the adapter strand if necessary. (MBC3: / 5Phos / CAGCTTTNNNAGATCGGAAGAGCACACGTCT / ideoxyU / ATATATA / iSp18 / / iSp18 / / iSp18 / / iSp18 / T / 3AmMO / (SEQ ID NO: 33); and, MBC4: / 5Phos / CCTATATNNNAGATCGGAAGAGCACACGTCTTAATATTTAATAT / ideoxyU / ATATAT / iSp18 / / iSp18 / / iSp18 / / iSp18 / T / 3AmMO / ) (SEQ ID NO: 34).
[0269] Two types of beads were prepared, one with a reporter antibody (m6A) bearing MBC3 and Ab05, and the other with a reporter antibody (m5C) bearing MBC4 and Ab16. IP of spacer-modified RNA samples was performed as described above. RNA was loaded into a ligation mix containing a mixture of splint oligonucleotides, and washed beads were added to induce barcoding. The splints were designed to hybridize to the spacer region of the target RNA on one side and to be complementary to a 7-nt-long MBC3 or MBC4 of the adapter on the other side. One set of splints hybridizes to 6 bases of the spacer region (AAAGCTGCACTCA / 3SpC3 / (7-6 MBC3) (SEQ ID NO:18) and ATATAGGCACTCA / 3SpC3 / (7-6 MBC4) (SEQ ID NO:19)) and the other set binds to 3 bases of the spacer region (AAAGCTGCAC / 3SpC3 / (7-3 MBC3) (SEQ ID NO:20) and ATATAGGCAC / 3SpC3 / (7-3 MBC4) (SEQ ID NO:21)). The length and sequence of both sides of the splints, the universal spacer and the adaptor complement, were adjusted to prevent binding stabilization by mechanisms other than modification recognition by antibodies to ensure encoding by proximity ligation. In workflows relying on primer extension for encoding, the spacer and spacer complement were present during the IP step (i.e., as shown in Figure 17A and Figure 19A), but in this protocol, splints were added after IP to distinguish between nucleic acid hybridization and IP. To simultaneously detect m6A and m5C, the ligation mix contained 0.5uM MBC3 and MBC4 splints, 10 units / uL T4 DNA ligase, 50mM Tris-HCl, 10mM MgCl2, 1mM ATP, 10% PEG 8000. After adapter ligation was completed, the i7 adapters were primed, followed by reverse transcription with template switching and PCR amplification as described in Example 13.By sequencing the libraries and bioinformatically deconvoluting the MBCs appended to each RNA fragment, the RNA modifications were identified and localized to specific loci (Figures 20B and 20C). This workflow was able to detect m6A and m5C with similar specificity as reported for reverse transcription-encoding (Example 13).
[0270] Example 16: A-tailing and encoding examples In this example, a universal sequence for encoding by primer extension was introduced by A-tailing the 3' end of the RNA (Figure 21A), thus eliminating the need for single-stranded spacer ligation (Example 12). A-tailing reactions are known to be higher yielding and more unbiased than single-stranded ligation, and they offer advantages to assay sensitivity in terms of improving transcriptome coverage.
[0271] The 1.5 kb IVT RNA was fragmented to 150 bases by incubation in 1X T4 RNA Ligase I Buffer (New England Biolabs) at 90°C for 20 min. The 3' end of the RNA was dephosphorylated by adding RNase inhibitor (catalog no. AM2694, Thermo Fisher) in the presence of T4 Polynucleotide Kinase (catalog no. T4PK-200, MCLab) for 30 min at 37°C. The reaction was supplemented with 5 units of E. coli poly(A) polymerase (catalog no. M0276L, New England Biolabs), 0.95 mM ATP, 0.05 mM dATP, and 1X E. coli poly(A) polymerase buffer and incubated at 37°C for 10 min. The A-tailed RNA was purified with 1.8 volumes of RNAClean XP beads.
[0272] To detect m6A in RNA samples, Ab05(m6A) and a reporter antibody conjugated to an adapter containing a barcode identifying m6A(MBC000) Beads displaying TIFF2024541478000021.tif20160 were prepared. The adapter structure includes a poly(dT) sequence that hybridizes to the A-tail RNA, an MBC (underlined), a UMI (NNNNNNNN), and an i7 Illumina adapter (the 5' sequence of the UMI).
[0273] For each sample, beads were loaded and IP of A-tailed RNA fragments was performed in the same manner as in Example 13. Briefly, beads, 0.05-50ng RNA, and 10 units / uL RNase inhibitor were incubated in 1X PBST. After incubation, beads were washed and reverse transcribed by extension of immobilized Illumina i7 adapters. Template switching with TSO introduced Illumina i5 adapters required for PCR amplification and sequencing. After the reaction was completed, the supernatant was subjected to PCR using standard Illumina index primers (1 uM forward primer, 1 uM reverse primer, NEBNext Ultra II Q5 (catalog no. M0544X, New England Biolabs) for 10-13 cycles (98°C for 30 s, followed by n cycles of 98°C for 10 s, 65°C for 75 s, and a final extension at 65°C for 5 min). The libraries were sequenced, and by bioinformatic deconvolution of the MBCs appended to each RNA fragment, RNA modifications were identified and localized to specific loci.
[0274] Figure 21B shows the sequencing results obtained with the described barcoding method using 0.5 ng of pooled IVT as input. In this example, the IVT RNA pool consisted of 70% unmodified PhiX RNA, 10% m6A modified FLuc-RNA, 10% inosine modified Lambda RNA, 10% m5C modified M13 RNA. The plot shows that the MBCs are associated with the correct genome.
[0275] Example 17: Multiplex detection of m6A, inosine and m5C in mRNA using reverse transcription encoding, template switching and streptavidin beads This example describes an end-to-end library preparation workflow with an integrated barcoding step for detecting RNA modifications in mRNA-enriched samples derived from a human lung cancer immortalized cell line (A549, Cat. No. 636141, Takara). Barcoding is achieved by bidirectional extension of the RNA target and adapter using reverse transcriptase (Figure 17A).
[0276] To detect m5C, m6A and inosine in RNA samples, a minimum of three types of beads are required: biotinylated adapters and protein G were bound to streptavidin-coated beads (catalog no. 65305, Thermo Fisher) followed by affinity binding of modification-specific antibodies (Figure 1D). The first bead type was anti-m6A (catalog no. MA5-3303, Thermo Fisher) and MBC-111 TIFF2024541478000022.tif20160 The second bead type displays an adapter containing anti-inosine (catalog no. PM098, MBL) and MBC-112 TIFF2024541478000023.tif20160 The third bead type was an adapter containing anti-m5C (catalog no. MA5-24694, Thermo Fisher Scientific) and MBC-113 TIFF2024541478000024.tif20160 The adaptor containing
[0277] The adapter contained a spacer sequence (bold at the 3' end), MBC (underlined), UMI (NNNNNNNNNNN), and an i5 Illumina adapter (sequence 5' to the UMI).
[0278] For each sample, an equal amount of each bead was combined to form the substrate for IP. The first assay step was the IP of the spacer-ligated RNA prepared according to Example 12. The bead pool was mixed with 10uL of 50ng RNA and 10 units / uL of RNase inhibitor in 1XTBST and incubated for 30 minutes. After incubation, the beads were washed with 1XTBST buffer and resuspended in 1X Superscript IV reverse transcription buffer (Cat. No. 18090050, Thermo Fisher). The wash removed non-specifically bound RNA and retained the specific RNA modification-antibody complex. The reverse transcriptase extended the 3' end of the RNA target, thereby copying the MBC and i5 adapters, and simultaneously extended the 3' end of the adapters to synthesize cDNA.
[0279] The IP beads were added to a reverse transcription reaction (1X Superscript IV buffer, 0.5u / uL Superase-In, 5u / uL Superscript IV reverse transcriptase, 1mM dNTPs, 2uM template switching oligo, "TSO" (AGACGTGTGCTCTTCCGATCTrGrGrG) (SEQ ID NO: 9) and incubated at 23°C for 15 minutes, then 50°C for 60 minutes. After the reaction was complete, the beads were washed with 1XTBST, denatured with 0.1N NaOH to remove RNA, and neutralized by washing again with 1XTBST. The cDNA attached to the beads was PCR amplified for 17-19 cycles (98°C for 30 seconds, followed by n cycles of 98°C for 10 seconds, 65°C for 75 seconds, and 65°C for 5 minutes) by adding the beads directly to a reaction mixture containing standard Illumina index primers (0.5uM forward primer, 0.5uM reverse primer, NEBNext Ultra II Q5 (Cat. no. M0544X, New England Biolabs)).
[0280] The libraries were sequenced on an Illumina sequencer, and RNA modifications were identified and localized to specific gene loci by bioinformatic deconvolution of the MBCs appended to each RNA fragment. Figure 23A shows the global barcode representation of technical triplicates of IP RNA and non-enriched (input) samples. As expected, IP samples show enrichment of MBC111 reads, since the m6A modification of mRNA is known to occur 5-10 times more frequently than inosine or m5C. Reads were aligned, and the stack-up of reads for each barcode was compared between IP and input samples, and peaks were called with MACS2. Figure 23B shows the location of the called peaks within the gene. The shift in the peak calls for MBC111-m6A towards the 3' end is consistent with the known bias of the m6A modification towards the 3' UTR (untranslated region). Figure 23C shows a Venn diagram of the number of peaks called for each modification and each replicate sample. The numbers of high-confidence peaks, i.e., peaks that appeared in all three replicates, were 6,805 for m6A, 773 for inosine, and 2741 for m5C, which is consistent with the numbers of modification sites for these modifications reported by other methods.
[0281] Example 18: Detection of m6A using antibody-protein A-Tn5 complex In this example, we describe the use of an immobilized conjugate containing an antibody and a Protein A-Tn5 fusion protein for tagmentation of DNA-RNA heteroduplexes specific to m6A modification sites (Figure 1H). The tagmentation reaction introduces a barcode that identifies the RNA modification. The advantage of this format is that the transposase is directly conjugated to the antibody, thus restricting tagmentation activity to the RNA modification site.
[0282] m6A-specific beads were prepared by forming a conjugate containing m6A antibody and Protein A-Tn5 molecules (Diagnode, Cat. No. C01070002) dissolved in solution and immobilized on Protein G beads (Figures 1H and 24A). Like Protein G, Protein A also binds strongly to the Fc region of antibodies, so Tn5 is immobilized on the beads in direct proximity to the m6A antibody binding pocket. Each Tn5 dimer was loaded with a pair of mosaic end (ME) adapters containing a barcode representing m6A (I7 ME adapter): 5'Phos- CTGTCTCTTATACACATCT (SEQ ID NO: 16) hybridized to CAAGCAGAAGACGGCATACGAGAT-NNNNNNNN-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO: 39); i5 ME adapter: 5'Phos- CTGTCTCTTATACACATCT (SEQ ID NO: 16) hybridized to AATGATACGGGCGACCACCGAGATCTACAC-NNNNNNNN-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG (SEQ ID NO: 40).
[0283] First, RNA containing a mixture of unmodified IVT RNA and m6A modified A-tailed IVT RNA (see Example 11) was reverse transcribed using Superscript IV reverse transcriptase and poly-dT oligo primers. The DNA-RNA heteroduplex was then reverse transcribed using IP buffer (50 mM HEPES pH 7.5, 300 mM NaCl, 0.1 M EDTA, 0.05% Tween (登録商標) The beads were added to the 50-mL PBS containing 10% PBS-200 mM NaCl and immunoprecipitated (IP) for 30 minutes. During this process, the m6A antibody selectively bound to the m6A-modified RNA. The beads were washed and incubated with Mg 2+The tagmentation reaction was initiated by adding the tagmentation buffer (10 mM Tris-HCl pH 8.5, 5 mM MgCl2, 10% DMF). The tagmented DNA-RNA heteroduplexes were gap-filled and PCR amplified for 17–19 cycles (5 min at 72 °C, 2 min at 98 °C, followed by n cycles of 10 s at 98 °C, 75 s at 65 °C, final extension 5 min at 65 °C) using a reaction mixture containing standard Illumina index primers or library amplification primers (0.5 uM forward primer, 0.5 uM reverse primer, NEBNext Ultra II Q5 (catalog no. M0544X, New England Biolabs). Libraries were sequenced on an Illumina sequencer, and RNA modifications were identified and localized to specific loci by bioinformatic deconvolution of barcodes added to each RNA fragment.
[0284] Figure 24B compares the read coverage plots of the input (control) and immunoprecipitated samples. The m6A modified region shows significant read enrichment in the immunoprecipitated samples, while other regions are depleted compared to the RNA input. To determine the optimal Protein A-Tn5 loading ratio, experiments were performed using 2-, 4-, and 8-fold excess of Protein A-Tn5 over antibody. Although specific enrichment of m6A was observed in all conditions, higher Protein A-Tn5 ratios negatively impacted library yield without improving specificity, and we concluded that a 2-4-fold excess of Protein A-Tn5 was ideal. From these experiments, the combination of IP and tagmentation with antibody-pA-Tn5 complexes proved to be effective in detecting m6A in complexed RNA pools.
[0285] Although the subject matter of the present disclosure has been described and illustrated in some detail with reference to certain exemplary embodiments including various combinations and subcombinations of features, those skilled in the art will readily recognize other embodiments and variations and modifications thereof as encompassed within the scope of the present disclosure. Moreover, the description of such embodiments, combinations, and subcombinations is not intended to convey that the claimed subject matter requires features or combinations of features other than those expressly recited in the claims. Accordingly, the scope of the present disclosure is intended to include all modifications and variations encompassed within the spirit and scope of the following appended claims.
Claims
1. 1. A method for analyzing a plurality of non-canonical features in a plurality of target nucleic acids, comprising: (i) contacting a solution containing a plurality of target nucleic acids with one or more compositions, wherein the target nucleic acids containing the non-canonical feature bind to the binding domain; wherein the one or more compositions each independently comprise: (a) substrate; a binding domain that binds to a substrate via a first linker or a second recognition element; and an adaptor that binds to a second recognition element or substrate via a second linker; wherein the binding domain specifically binds to a non-canonical feature of DNA or RNA and the adapter comprises a nucleic acid barcode sequence unique to the non-canonical feature; (b) a substrate; a binding domain that binds to a substrate via a first linker or multiple recognition elements; and an adaptor that binds to one or more of the plurality of recognition elements or the substrate via a second linker; wherein the binding domain specifically binds to a non-canonical feature of DNA or RNA, and the adapter comprises a nucleic acid barcode sequence unique to the non-canonical feature; or (c) a combination of (a) and (b); Selected from: (ii) performing one of (a1) and (b1) below: (a1) transferring a nucleic acid barcode to a target nucleic acid comprising a non-canonical feature to generate a barcoded target nucleic acid; or (b1) generating a barcoded copy of a target nucleic acid that includes a non-canonical feature; (iii) amplifying the barcoded target nucleic acid; and (iv) determining the sequence of the barcoded target nucleic acid wherein steps (i) and (ii) are carried out sequentially or simultaneously; Optionally, the amplification step (iii) is carried out on the surface of the substrate, Optionally, the nucleic acid barcode is transferred to the target nucleic acid by transferase or chemical ligation; Optionally, prior to step (i), the method further comprises ligating a universal nucleic acid sequence to the 3' end or 5' end or both ends of the target nucleic acid; Optionally, prior to step (i), the method further comprises enzymatically tailing the 3' end of the target nucleic acid with a plurality of single-type nucleotides; Optionally, step (iii) further comprises forming clusters of identical copies of the target nucleic acid; Optionally, in step (iv), the method further comprises in situ sequencing of clusters of identical copies of the target nucleic acid on the substrate. The method.
2. 10. The method of claim 1, wherein the nucleic acid barcode is enzymatically transferred to the target nucleic acid by single-strand ligation, splint ligation, primer extension, reverse transcription, or double-strand ligation.
3. the 3' end of the adapter is attached to the 3' end of the target nucleic acid, and step (ii) further comprises introducing a modification-specific barcode, wherein the 3' end of the adapter is extended by a reverse transcriptase or a DNA polymerase, or the adapter comprises a 3' spacer sequence and site-specifically attaches to a synthetic spacer sequence exhibited by the target nucleic acid, and one or both of the 3' end of the adapter and the 3' end of the target nucleic acid are extended by a reverse transcriptase or a DNA polymerase; or an adaptor having a 3' degenerate base randomly primes the target nucleic acid, and step (ii) further comprises introducing a modification-specific barcode, wherein the 3' end of the adaptor is extended by reverse transcriptase or DNA polymerase; The method according to claim 1 or 2.
4. 3. The method of claim 1 or 2, wherein the barcoded target nucleic acid is immobilized on the substrate via its 5' end or via its 3' end.
5. further comprising modifying the barcoded target nucleic acid or a copy thereof such that the location of the non-canonical feature is identifiable based on the primary nucleic acid sequence of the barcoded target nucleic acid or a copy thereof, wherein optionally, said modifying comprises editing bases within 1 to 20 bases of the binding site where the binding domain binds to the target nucleic acid; The method according to claim 1 or 2.
6. i) substrate; ii) a binding domain that binds to the substrate via a first linker or a second recognition element; and iii) an adaptor that binds to a second recognition element or substrate via a second linker, wherein the binding domain is configured to specifically bind to a non-canonical feature of DNA or RNA, and the adaptor comprises a nucleic acid barcode sequence unique to the non-canonical feature; Optionally, further comprising a base editing enzyme bound to the binding domain. The composition.
7. the composition comprises a plurality of second recognition elements, wherein the plurality of second recognition elements comprises second recognition elements that are different from one another, and the adapter is attached to one of the plurality of second recognition elements, and the binding domain is attached to a different second recognition element; the composition comprises a plurality of second recognition elements, wherein the adapter is attached to one of the plurality of second recognition elements and the binding domain is attached to another instance of the same second recognition element; a second recognition element attached to the substrate, a binding domain that binds to the substrate via the second recognition element, and an adaptor that binds to the substrate via a second linker; or a second recognition element is attached to the substrate, a binding domain that binds to the substrate via a first linker, and an adaptor that binds to the substrate via a second recognition element; The composition of claim 6.
8. the second recognition element is Protein G, Protein L, Protein A, Protein AG, Protein AL, Protein LG, a peptide tag, a covalent peptide tag, a protein tag, a nucleic acid, streptavidin, avidin, or neutravidin; and / or the binding domain comprises an antibody, scFv, Fab fragment, antibody light chain (VL), antibody heavy chain (VH), variable fragment (Fv), F(ab')2 fragment, diabody, VHH domain, nanobody, aptamer, reader protein, writer protein, eraser protein, endonuclease V, artificial polymeric scaffold, artificial protein scaffold, or selective covalent capture reagent, or a fragment or derivative thereof; and / or The non-canonical feature is a modified nucleoside, nucleic acid damage or structural element; and / or Modified nucleosides include 3-methylcytidine (m3C), 5-methylcytidine (m5C), N4-acetylcytidine (ac4C), pseudouridine (Ψ), 1-methyladenosine (m1A), N6-methyladenosine (m6A), inosine (I), 7-methylguanosine (m7G), 7-methylguanosine (m7G)-Cap, dihydrouridine (D), 3-methyluridine (m3U), 5-methyluridine (m5U), 1-methylguanosine (m1G), N2-methylguanosine (m2G), 5-methyldeoxycytidine (m4C), and 5-methyl-1 ... cytidine (m5dC), N4-methyldeoxycytidine (4mdC), 5-hydroxymethylcytidine (5-hmC), 5-hydroxymethyldeoxycytidine (5hmdC), 5-carboxydeoxycytidine (5cadC), 5-carboxycytidine (5caC), 5-formylcytidine (5fC), 5-formyldeoxycytidine (5fdC), 6-methyldeoxyadenosine, N7-methylguanosine (m7G), 2,7,2'-methylguanosine, ribose methylation (Nm), N2,N2-dimethylguanosine (m22) G), 5-carbamoylmethyl-2'-O-methyluridine (ncm5Um), 5-methoxycarbonylmethyluridine (ncm5mU), 5-methoxycarbonylmethyl-2-thiouridine (mcm5s2U), queuosine (Q), 2-thiouridine (s2U), 5-taurinomethyluridine (τm5U), 5-taurinomethyl-2-thiouridine (τm5s2U), N6-isopentenyladenosine (I6A), 2-methylthio-N6-threonylcarbamoyladenosine (ms2t6A), and the nucleic acid damage is 8-oxoguanine (8-oxoG), one or more abasic site, a cis-platin bridge, a benzo(a)pyrene diol epoxide (BPDE) adduct, a cyclobutene pyrimidine dimer (CPD), a pyrimidine-pyrimidone (6-4) photoproduct (6-4PP), 6-O-methylguanine (O6-MedG), or O6-(carboxymethyl)-2'-deoxyguanosine (O6-CMdG), and the structural element is a hairpin, a loop, a Z-DNA structure, a G-quadruplex, a triplex, an I-motif, a bulge, a three-way junction, a cruciform structure, a tetraloop, a ribose zipper, or a pseudoknot; and / or the adapter comprises at least one of a universal forward primer (UFP) and a universal reverse primer (URP), a molecular barcode (UMI), one or more unnatural nucleobases, or two or more random bases at the 3' end; and / or the binding domain and the adapter are spatially separated on the substrate to allow transfer of the adapter or a copy of the adapter sequence to the target nucleic acid, optionally with the binding domain attached to the substrate at a distance of less than 200 nm from the adapter; and / or the base editing enzyme is covalently attached to the binding domain or is linked to the binding domain via a targeting moiety selected from a peptide tag, a protein tag, a secondary antibody, a nucleic acid sequence, or a bioorthogonal reactive group; and optionally, the base editing enzyme is an adenosine deaminase, cytosine deaminase, glycosylase, methylase, demethylase, or dioxygenase. The composition according to claim 6 or 7.
9. 1. A method for analyzing a plurality of target nucleic acids, comprising: (i) contacting a solution containing a plurality of target nucleic acids with a composition comprising: substrate, a binding domain that binds to a substrate via a first linker or a second recognition element; an adaptor that binds to a second recognition element or substrate via a second linker; wherein the binding domain specifically binds to a non-canonical feature of DNA or RNA, and the adapter comprises a nucleic acid barcode sequence unique to the non-canonical feature, wherein a target nucleic acid comprising the non-canonical feature is bound to the binding domain; (ii) doing one of the following: (a) transferring a nucleic acid barcode to a target nucleic acid comprising a non-canonical feature to generate a barcoded target nucleic acid; or (b) generating a barcoded copy of the target nucleic acid that includes the non-canonical feature; (iii) amplifying the barcoded target nucleic acid; and (iv) determining the sequence of the barcoded target nucleic acid wherein steps (i) and (ii) are carried out sequentially or simultaneously; Optionally, the amplification step (iii) is carried out on the surface of the substrate; The method.
10. 10. The method of claim 9, wherein the nucleic acid barcode is transferred to the target nucleic acid by enzymatic transfer, chemical transfer, or chemical ligation, Optionally, wherein the nucleic acid barcode is enzymatically transferred to the target nucleic acid by single-strand ligation, splint ligation, primer extension, reverse transcription, or double-strand ligation; or or, prior to step (i), further comprising ligating a universal nucleic acid sequence to the 3' end or 5' end or both ends of the target nucleic acid, or, prior to step (i), further comprising enzymatically tailing the 3' end of the target nucleic acid with a plurality of mono-type nucleotides, optionally wherein the enzymatic tailing is carried out using poly(A) polymerase or poly(U) polymerase; or the 3' end of the adapter is attached to the 3' end of the target nucleic acid, and step (ii) further comprises introducing a modification-specific barcode, wherein the 3' end of the adapter is extended by a reverse transcriptase or a DNA polymerase, or the adapter comprises a 3' spacer sequence and site-specifically attaches to a synthetic spacer sequence exhibited by the target nucleic acid, and one or both of the 3' end of the adapter and the 3' end of the target nucleic acid are extended by a reverse transcriptase or a DNA polymerase; or an adaptor having a 3' degenerate base randomly primes the target nucleic acid, and step (ii) further comprises introducing a modification-specific barcode, wherein the 3' end of the adaptor is extended by reverse transcriptase or DNA polymerase; or the barcoded target nucleic acid is immobilized on the substrate via its 5' end, or the barcoded target nucleic acid is immobilized on the substrate via its 3' end; The method.
11. the method comprising modifying a barcoded target nucleic acid or a copy thereof such that the location of the non-canonical feature is distinguishable based on a primary nucleic acid sequence of the barcoded target nucleic acid or the copy thereof; Optionally, the modifying comprises editing bases within 1 to 20 bases of the binding site, wherein the binding domain is bound to the target nucleic acid.
11. The method according to claim 9 or 10.
12. 11. The method of claim 9 or 10, further comprising forming clusters of identical copies of the target nucleic acid in step (iii), and optionally further comprising in situ sequencing the clusters of identical copies of the target nucleic acid on the substrate in step (iv).
13. A composition comprising: i) substrate; ii) a binding domain that binds to the substrate via a first linker or a second recognition element; iii) a mosaic end (ME) adaptor that binds to the substrate via a second linker or second recognition element; and iv) a transposase; wherein the transposase is bound to immobilized ME adapters, the binding domain specifically binds to a non-canonical feature in DNA or RNA, and at least one of the ME adapters comprises a nucleic acid barcode sequence unique to the non-canonical feature; or i) substrate; ii) a binding domain that binds to the substrate via a linker or second recognition element; and iii) a transposase linked to a binding domain; wherein the transposase is bound to ME adapters, the binding domain specifically binds to a non-canonical feature of DNA or RNA, and at least one of the ME adapters comprises a nucleic acid barcode sequence unique to the non-canonical feature; Optionally, the composition wherein the transposase is a Tn5 transposase.
14. 14. The composition of claim 13, wherein the transposase is linked to the binding domain via Protein L, Protein G, Protein A, Protein GA, Protein AL, Protein GL, a peptide tag, a covalently linked peptide tag, a protein tag, or a nucleic acid.
15. 14. The composition of claim 6 or 13, wherein the substrate is a bead, a chip, a plate, a slide, a dish, a gel, or a three-dimensional polymer matrix.
16. 14. The composition of claim 13, wherein the binding domain comprises an antibody, scFv, Fab fragment, antibody light chain (VL), antibody heavy chain (VH), variable fragment (Fv), F(ab')2 fragment, diabody, VHH domain, nanobody, aptamer, reader protein, writer protein, eraser protein, artificial polymeric scaffold, endonuclease V, artificial protein scaffold, or selective covalent capture reagent, or a fragment or derivative thereof.
17. The composition of claim 9 or 16, wherein the reader protein is a NUDT16, ALYREF or YTH domain protein, or a fragment or derivative thereof, the writer protein is a DNMT protein, a NAT10 protein, a METTL protein, a TRM protein, a BMT protein, a DUS protein, a PUS protein, an ADAR protein or an NSUN protein, or a fragment or derivative thereof, and the eraser protein is an FTO protein, an ALKBH protein, a TET protein, or a fragment or derivative thereof.
18. The composition of claim 13 , wherein the non-canonical feature is a modified nucleoside, a nucleic acid lesion, or a structural element.
19. Modified nucleosides include 3-methylcytidine (m3C), 5-methylcytidine (m5C), N4-acetylcytidine (ac4C), pseudouridine (Ψ), 1-methyladenosine (m1A), N6-methyladenosine (m6A), inosine (I), 7-methylguanosine (m7G), 7-methylguanosine (m7G)-Cap, dihydrouridine (D), 3-methyluridine (m3U), 5-methyluridine (m5U), 1-methylguanosine (m1G), N2-methylguanosine (m2G), 5-methyldeoxycytidine (m4C), and 5-methyl-1 ... cytidine (m5dC), N4-methyldeoxycytidine (4mdC), 5-hydroxymethylcytidine (5-hmC), 5-hydroxymethyldeoxycytidine (5hmdC), 5-carboxydeoxycytidine (5cadC), 5-carboxycytidine (5caC), 5-formylcytidine (5fC), 5-formyldeoxycytidine (5fdC), 6-methyldeoxyadenosine, N7-methylguanosine (m7G), 2,7,2'-methylguanosine, ribose methylation (Nm), N2,N2-dimethylguanosine (m 2 2 G), 5-carbamoylmethyl-2'-O-methyluridine (ncm5Um), 5-methoxycarbonylmethyluridine (ncm5mU), 5-methoxycarbonylmethyl-2-thiouridine (mcm5s2U), queuosine (Q), 2-thiouridine (s2U), 5-taurinomethyluridine (τm5U), 5-taurinomethyl-2-thiouridine (τm5s2U), N6-isopentenyladenosine (I6A), 2-methylthio-N6-threonylcarbamoyladenosine (ms2t6A), the nucleic acid damage results from bulky adduct formation or base alkylation by an exogenous agent, and optionally the damage is 8-oxoguanine (8-oxoG), one or more abasic sites, cis-platin crosslinks, benzo(a)pyrene diol epoxide (BPDE) adducts, cyclobutene pyrimidine dimers (CPDs), pyrimidine-pyrimidone (6-4) photoproducts (6-4PPs), 6-O-methylguanine (O6-MedG), or O6-(carboxymethyl)-2'-deoxyguanosine (O6-CMdG); and 19. The composition of claim 18, wherein the structural element is a hairpin, a loop, a Z-DNA structure, a G-quadruplex, a triplex, an I-motif, a bulge, a three-way junction, a cruciform structure, a tetraloop, a ribose zipper, or a pseudoknot.
20. the adapter comprises at least one of a universal forward primer (UFP) and a universal reverse primer (URP); the adapter comprises a molecular barcode (UMI); the adapter comprises one or more unnatural nucleobases; the adapter comprises one or more 3' or 5' blocking groups; or the adapter comprises a 19 bp mosaic end (ME); Optionally, one or more 3' or 5' blocking groups are independently selected from dideoxyribose, phosphate, reverse base, or linker; The composition of claim 13.
21. 14. The composition of claim 13, wherein the binding domain and adapter are spatially separated on the substrate to allow transfer of the adapter to the target nucleic acid by the transposase, and optionally the binding domain is bound to the substrate at a distance of less than 200 nm from the adapter.
22. 1. A method for analyzing a plurality of target nucleic acids, comprising: (i) providing a plurality of target nucleic acids by reverse transcribing target RNA molecules to form DNA-RNA heteroduplex molecules or by providing target double-stranded DNA molecules; (ii) contacting a solution containing a plurality of target nucleic acids with a composition comprising: substrate, a binding domain that binds to a substrate via a first linker or a second recognition element; a mosaic end (ME) adaptor that binds to the substrate via a second linker or second recognition element; and a composition comprising a transposase, wherein the transposase is loaded onto immobilized ME adapters, the binding domain specifically binds to a non-canonical feature of DNA or RNA, and at least one of the ME adapters comprises a nucleic acid barcode sequence unique to the non-canonical feature; or substrate, a binding domain that binds to the substrate via a linker or a second recognition element; and a composition comprising a transposase bound to a binding domain, wherein the transposase is loaded onto ME adapters, the binding domain specifically binds to a non-canonical feature of DNA or RNA, and at least one of the ME adapters comprises a nucleic acid barcode sequence unique to the non-canonical feature. wherein a target nucleic acid containing a non-canonical feature is bound to the binding domain; (iii) using a transposase to transfer two adapters, at least one of which comprises a nucleic acid barcode, to a double-stranded target nucleic acid comprising a non-canonical feature to generate a barcoded target nucleic acid; (iv) amplifying the barcoded target nucleic acid; and (v) determining the sequence of the barcoded target nucleic acid wherein steps (ii) and (iii) are carried out sequentially or simultaneously; The method.
23. Step (ii) is Mg 2+ step (ii) is carried out before step (iii); and / or 23. The method of claim 22, wherein step (iii) further comprises adding Mg 2+ ions, and wherein step (iii) is performed after step (ii), or wherein the nucleic acid barcode is transferred to the target nucleic acid by a transferase, or wherein the ME adapter strand containing the barcode is immobilized on a substrate via its 5' end.
24. 24. The method of claim 22 or 23, wherein the method comprises modifying the barcoded target nucleic acid or a copy thereof such that the location of the non-canonical feature is distinguishable based on the primary nucleic acid sequence of the barcoded target nucleic acid or the copy thereof.
25. 25. The method of claim 24, wherein the modifying step comprises editing bases within 1 to 20 bases of the binding site where the binding domain binds to the target nucleic acid.
26. 1. A method for detecting a plurality of non-canonical features in a plurality of target nucleic acids, comprising: (i) contacting a solution containing a plurality of target nucleic acids with a plurality of compositions; wherein each of the plurality of compositions is independently selected from: (a) substrate; a binding domain that binds to a substrate via a first linker or a second recognition element; and an adaptor that binds to a second recognition element or substrate via a second linker; wherein the binding domain specifically binds to a non-canonical feature of DNA or RNA and the adapter comprises a nucleic acid barcode sequence unique to the non-canonical feature; (b) a substrate; a second recognition element that binds to the substrate via a first linker; an adaptor that binds to a second recognition element or substrate via a second linker; and Binding domain wherein the binding domain specifically binds to a non-canonical feature of DNA or RNA, the binding domain is immobilized by a second recognition element, and the adapter comprises a nucleic acid barcode sequence unique to the non-canonical feature; (c) a substrate; a second recognition element; a binding domain that binds to the substrate via a second recognition element; and an adaptor that binds to a substrate via a linker; wherein the binding domain is configured to specifically bind to a non-canonical feature of DNA or RNA, and the adapter comprises a nucleic acid barcode sequence unique to the non-canonical feature; (d) a substrate; a second recognition element that binds to the substrate; a binding domain that binds to a substrate via a linker; Adapters that bind to the substrate via a second recognition element wherein the binding domain is configured to specifically bind to a non-canonical feature of DNA or RNA, and the adapter comprises a nucleic acid barcode sequence unique to the non-canonical feature; (e) substrate; a binding domain that binds to a substrate via a first linker or a plurality of second recognition elements; and an adaptor that binds to one or more of the plurality of second recognition elements or the substrate via a second linker; wherein the binding domain is configured to specifically bind to a non-canonical feature of DNA or RNA, and the adapter comprises a nucleic acid barcode sequence unique to the non-canonical feature; wherein the number of the plurality of compositions contacted in step (i) is equal to or greater than the number of non-canonical features, the binding domains of said plurality of compositions each bind to a different non-canonical feature of DNA or RNA, or multiple binding domains bind to the same non-canonical feature of DNA or RNA, and the adapters of said plurality of compositions each comprise a nucleic acid barcode sequence unique to the non-canonical feature or unique to the binding domain specifically bound by the binding domain of that composition; (ii) doing one of the following: (a) transferring the nucleic acid barcode sequence of each of the plurality of compositions to a plurality of target nucleic acids; or (b) generating multiple barcoded copies of the target nucleic acid; (iii) amplifying the barcoded target nucleic acid; and (iv) determining the base sequence of the barcoded target nucleic acid; The method comprising: