Modification of pseudouridine

By employing 2-bromoacrylamide-assisted circularization sequencing (BACS), the selectivity and sensitivity deficiencies of existing pseudouridine detection methods have been addressed, enabling high-resolution and high-precision quantitative analysis of pseudouridine, particularly in densely modified regions and continuous uridine sequences.

CN121368599APending Publication Date: 2026-01-20LUDWIG INSTITUTE FOR CANCER RESEARCH LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480042101.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-04-26
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Existing methods for detecting pseudouridine suffer from low selectivity and insufficient sensitivity, making it difficult to accurately distinguish between genuine pseudouridine signals and RNA secondary structure background noise. Furthermore, they require a large amount of starting material, leading to RNA degradation and low detection efficiency.

Method used

The 2-bromoacrylamide-assisted circularization sequencing (BACS) method is used to induce pseudouridine to react with Michael addition receptors. Through the induction of Ψ to C mutations during reverse transcription, it provides quantitative analysis at single-base resolution, avoids truncated or missing features, and improves detection accuracy.

Benefits of technology

This technology enables precise identification of pseudouridine locations in densely modified regions and continuous uridine sequences, improving the resolution and chemometric quantification capabilities of pseudouridine detection and reducing the risk of RNA degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121368599A_ABST
    Figure CN121368599A_ABST
Patent Text Reader

Abstract

A 2-bromoacrylamide assisted cyclization sequencing (BACS) method allows quantitative profiling of pseudouridine (psi) at single base resolution. Based on bromoacrylamide cyclization chemistry, BACS induces psi-to-C mutations during reverse transcription (RT) instead of truncation or deletion of features, thus providing higher resolution and enabling more accurate quantification of psi stoichiometry compared to CMC and BS based methods. The BACS of the invention allows for precise identification of psi positions, in particular in densely modified psi regions and continuous uridine sequences, compared to known methods. Accordingly, methods, compositions, and kits for detecting psi are possible. In the case of RNA molecules, this allows psi to be detected and sequenced in such sequences.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to the field of molecular biology, and more specifically, to methods, compositions, and kits for modifying, detecting, localizing, and otherwise determining the presence of pseudouridine in ribonucleic acid sequences. BACKGROUND

[0002] Pseudouridine (Ψ), sometimes referred to as pseudouracil, is a C-C glycoside isomer of uridine (U) and is the most abundant post-transcriptional modification in cellular RNA. Ψ is ubiquitous in almost all classes of non-coding RNA (ncRNA), including ribosomal RNA (rRNA), transfer RNA (tRNA), and small nuclear RNA (snRNA). It is also known to be present in messenger RNA (mRNA).

[0003] Ψ has been found to play an important role in splicing, translation, RNA stability, and RNA-protein interactions. In eukaryotes, Ψ is installed by various pseudouridine synthases (PUS), which have been shown to be associated with many diseases, including cancer. Therefore, there is a great need to establish a precise and sensitive method to detect Ψ.

[0004] In ribosomes, Ψ residues cluster together, having a role in stabilizing RNA-RNA and / or RNA-protein interactions. This stability can help in the folding of rRNA and the assembly of ribosomes. The presence of Ψ in rRNA affects the stability of the nearby structures, thus affecting the speed and accuracy of decoding and proofreading during translation.

[0005] In snRNA, Ψ residues ensure the correct folding and assembly of the spliceosome, which is required for pre-mRNA processing.

[0006] In tRNA, the presence of Ψ stabilizes stem-loop structures in a way that a normal uracil residue cannot.

[0007] In mRNA, Ψ is the second most abundant internal modification (0.1-0.4% Ψ / U ratio measured by mass spectrometry). The presence of Ψ residues in mRNA is known to, for example, affect the coding specificity of stop codons UAA, UGA, and UAG, and this modification of U to Ψ leads to nonsense suppression.

[0008] Ψ is formed in post-transcriptional RNA structures, which is achieved by various Ψ synthases (PUS) in eukaryotes. These enzymes can play an important role in mRNA processing, stability, and translation.

[0009] Certain genetic mutants that lack Ψ residues in tRNA or rRNA have been found to have difficulties in translation, resulting in slower growth rates of cells compared to wild type. Ψ modification defects have been found to correspond to certain diseases, such as keratosis palmoplantar, mitochondrial myopathy and sideroblastic anemia (MLASA).

[0010] Ψ is also known to have some functions in the regulation of latency in human immunodeficiency virus (HIV) infection.

[0011] Pseudouridylation has also been known to be associated with maternally inherited diabetes and deafness (MIDD). Point mutations in mitochondrial tRNA can abolish pseudouridylation of nucleotides, resulting in changes in tRNA structure that lead to instability, causing poor mitochondrial translation and respiration.

[0012] Ψ in mRNA can also be associated with various types of cancer and other diseases, and can serve as a biomarker for early cancer detection.

[0013] Therefore, in many areas of biological and medical research, there is a need for precise and sensitive methods to detect Ψ in RNA molecules. Traditionally, detection of Ψ has relied mainly on N-cyclohexyl-N’-(2-morpholinoethyl) carbodiimide methy1-p-toluenesulfonate (CMC) chemistry (Bakin, A. V. & Ofengand, J. Methods Mol. Biol. 77, 297-309 (1998)). CMC can readily react with amide or imide functionalities in nucleobases (e.g., amide for guanosine, imide for uridine - Gilham, P. T. J. Am. Chem. Soc. 84, 687-688 (1962) and Ho, N. W. Y. & Gilham, P. T. Biochemistry 6, 3632-3639 (1967)), while it is unable to react with the N 3 forms a more stable adduct, thus enabling differentiation of Ψ from U by subsequent base treatment (pH ~ 10.4) (Ho, N. W. Y. & Gilham, P. T. Biochemistry 10, 3651-3657 (1971)). Since Ψ’s N 3- CMC adducts significantly interfere with Watson-Crick side base pairing, which would lead to a truncation signature in reverse transcription (RT). With the help of these RT terminators, CMC chemistry has been widely applied to sequencing within the transcriptional range of Ψ, as shown by Ψ-seq (Carlile, T. M. et al., Nature 515, 143-146 (2014)), pseudouridine-seq (Schwartz, S. et al., Cell 159, 148-162 (2014)), and PSI-seq (Lovejoy, A. F. et al., PLoS One 9, e110799 (2014)).

[0014] However, CMC-based methods have low labeling efficiency and selectivity for Ψ, which makes it inherently difficult to distinguish the true Ψ signal from background noise from other bases and RNA secondary structures (see Incarnato, D. et al., Genome Biol. 15, 491 (2014) and Wang, P. Y. et al., RNA 25, 135-146 (2019)). Only about 100-400 and about 50-100 Ψ sites were detected on human and yeast mRNA, respectively.

[0015] CeU-seq uses azide-CMC to enrich the truncated signal, which can detect about 1000-2000 Ψ sites in the human transcriptome. However, according to the results of mass spectrometry, there seem to be much more Ψ sites than actually detected. Therefore, although CeU-seq uses azide-labeled CMC to enrich the truncated signal and thus improve sensitivity, this method still has the problem of partial reactivity of CMC chemistry and harsh alkaline treatment (Li, X. et al., Nat. Chem. Biol. 11, 592-597 (2015)). Thus, a fundamental disadvantage of CMC-based methods can arise because of the low selectivity of CMC for Ψ, which makes it inherently more difficult to distinguish the true Ψ signal from background. In addition, a relatively large amount of starting material (about 5-10 μg) is required when using CMC-based methods, which can be due to inevitable RNA degradation caused by harsh alkaline treatment. Basically, all CMC-based methods lack stoichiometric information for Ψ.

[0016] Recently, bisulfite (BS) treatment has been used for cytosine modification detection (Singhal, R. P. Biochemistry 13, 2924-2932 (1974) and Everett, D. W. Part I: Reaction of pseudouridine with bisulfite. Part II: Reaction of glyoxal with guanine derivatives: A spectrophotometric probe of molecular structure. (New York University, 1980)). Surprisingly, it was found that BS treatment converts Ψ to a Ψ-BS adduct, which ultimately leads to a deletion signature in RT (RBS-seq), providing an improvement over the use of CMC (see Khoddami, V. et al., Proc. Natl. Acad. Sci. U. S. A. 116, 6784-6789 (2019); see also Fleming, A. M. et al., J. Am. Chem. Soc. 141, 16450-16460 (2019)). Because unmodified cytosine (C) is also deaminated to U in the regular BS reaction (see Shapiro, R. et al., J. Am. Chem. Soc. 92, 422-424 (1970) and Hayatsu, H. et al., J. Am. Chem. Soc. 92, 724-726 (1970)), BID-seq and pseudouridine assessment by 19-sulfite / sulfite treatment (PRAISE) further optimize the BS treatment to near neutral pH to eliminate most side reactions on C, enabling quantitative detection of Ψ across the transcriptome (see Dai, Q. et al., Nat. Biotechnol. 41, 344-354 (2023) and Zhang, M. et al., Nat. Chem. Biol. (2023)).

[0017] WO2022 / 232795 of the University of Chicago discloses methods to modify Ψ including modified bisulfite treatment. As explained therein, careful examination of the reactivity of Ψ with bisulfite allowed modification reactions of DNA or DNA with bisulfite in the pH 6.8-7.2 range, thereby enabling quantitative Ψ-BS formation, and no C to U conversion. Related to this is the corresponding scientific publication by Dai Q. et al. (2022) Nature Biotechnology 27 October 2022 DOI https: / / doi.org / 10.1038 / s41587-022-01505-w.

[0018] Zhang M. et al. (2022) bioRxiv doi: https: / / doi.org / 10.1101 / 2022.10.25.513650 describe a PRAISE method that relies on bisulfite-induced deletion features during reverse transcription, thus enabling quantitative pseudouridine assessment by bisulfite / sulfite treatment. PRAISE is based on fourfold read alignment, thus enabling accurate measurement of Ψ stoichiometry in spike-in RNA and rRNA.

[0019] While the above BS-based methods can detect about 1000-2000 Ψ sites on human mRNA, they still suffer from low deletion rates and high false positive rates in some sequence contexts. Moreover, due to the deletion feature, it is difficult to determine the exact Ψ site when it is adjacent to one or more U or consecutive Ψ sites are detected. Furthermore, it can be laborious to detect low modified Ψ sites and Ψ sites in lowly expressed RNAs, as sufficient coverage can be required to be generated.

[0020] There is a need for an improved pseudouridine detection and sequencing method that can be used in all areas of cell biology. SUMMARY

[0021] According to the present invention, methods, compositions and kits for modifying and detecting pseudouridine are provided; and in case pseudouridine forms part of an RNA molecule, then the pseudouridine residues in the RNA molecule are modified, detected and sequenced. Thus, the present invention provides a 2-bromopropenamide assisted circularization sequencing (BACS) method for quantitative profiling of Ψ at single base resolution. Based on this bromopropenamide circularization chemistry, BACS induces Ψ to C mutations during reverse transcription (RT), rather than truncation or deletion features, thus providing higher resolution and enabling more accurate quantification of Ψ stoichiometry compared to CMC- and BS-based methods. Thus, BACS of the present invention allows precise identification of Ψ positions compared to known methods, particularly in densely modified Ψ regions and consecutive uridine sequences.

[0022] In a first aspect, the present application provides a method of modifying pseudouridine comprising reacting pseudouridine with a Michael addition acceptor, wherein the Michael addition acceptor is a compound according to Formula (I):

[0023] (I)

[0024] wherein:

[0025] X is independently selected from halogen, tosyl, mesyl and triflyl;

[0026] R 1 is independently selected from H, C1-C4-alkyl, C1-C4-haloalkyl and Co-C4-alkylene-R 1a , wherein R 1a is independently selected from C3-C6-cycloalkyl, phenyl and 4- to 6-membered heterocyclyl; wherein when R 1a is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 1a is optionally substituted with 1 to 4 R 1b , and when R 1a is phenyl, R 1a is optionally substituted with 1 to 4 R 1c ;

[0027] R 2 is independently selected from H, C1-C4-alkyl, C1-C4-haloalkyl and Co-C4-alkylene-R 2a , wherein R 2a is independently selected from C3-C6-cycloalkyl, phenyl and 4- to 6-membered heterocyclyl; wherein when R 2a is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 2a is optionally substituted with 1 to 4 R 2b , and when R 2a is phenyl, R 2a is optionally substituted with 1 to 4 R 2c ;

[0028] R 3 is independently selected from -C(O)OR 3a , -C(O)R 3a , -C(O)NR 3b R 3b , -CN, -NO2, -S(O)2OR 3a , -S(O)2R 3a , -S(O)2NR 3b R 3b and 5- to 10-membered heteroaryl, wherein when R 3 is 5- to 10-membered heteroaryl, R3 optionally substituted by 1 to 4 R 3e substituents;

[0029] R 3a is independently selected from the group consisting of H, C1-C6-alkyl, C1-C6-haloalkyl and Co-C6-alkylene-R 3c ; wherein when R 3a is alkyl or haloalkyl, R 3a is optionally substituted by a group selected from -N3, -CºCH, dibenzocyclooctyne alcohol (DIBO), azido-dibenzocyclooctyne (DBCO), bicyclononyne (BCN), trans-cyclooctene (TCO) and tetrazine;

[0030] R 3b is independently selected from the group consisting of H, C1-C6-alkyl, C1-C6-haloalkyl and Co-C6-alkylene-R 3c ; wherein when R 3b is alkyl or haloalkyl, R 3b is optionally substituted by a group selected from -N3, -CºCH, dibenzocyclooctyne alcohol (DIBO), azido-dibenzocyclooctyne (DBCO), bicyclononyne (BCN), trans-cyclooctene (TCO) and tetrazine;

[0031] or wherein two R 3b groups together with the nitrogen atom to which they are attached form a 5- or 6-membered heterocycloalkyl group, which is optionally substituted by 1 to 4 R 3d substituents;

[0032] R 3c is independently selected from the group consisting of C3-C6-cycloalkyl, phenyl and 4- to 6-membered heterocyclyl; wherein when R 3c is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 3c is optionally substituted by 1 to 4 R 3d substituents, and when R 3c is phenyl, R 3c is optionally substituted by 1 to 4 R 3e substituents;

[0033] R 1b and R 2b are each independently at each occurrence selected from the group consisting of =0, =S, halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl;

[0034] R 1c and R 2chalogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl;

[0035] R 3d is at each occurrence independently selected from =O, =S, halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl, C1-C4-haloalkyl, -N3, -C≡CH, dibenzooctalynol (DIBO), azido-dibenzooctalynol (DBCO), bicyclononyne (BCN), trans-cyclooctene (TCO) and tetrazine;

[0036] R 3e is at each occurrence independently selected from halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl, -N3, -C≡CH, dibenzooctalynol (DIBO), azido-dibenzooctalynol (DBCO), bicyclononyne (BCN), trans-cyclooctene (TCO) and tetrazine;

[0037] R 4 is at each occurrence independently selected from H and C1-C4-alkyl; or when two R 4 groups are attached to the same nitrogen, the two R 4 groups together with the nitrogen atom to which they are attached optionally form a 5- to 6- membered heterocycloalkyl group, which is optionally substituted with 1 to 4 R 5 ; and

[0038] R 5 is at each occurrence independently selected from =O, =S, halogen, nitro, cyano, C1-C4-alkyl and C1-C4-haloalkyl.

[0039] The reaction according to the present application advantageously modifies pseudouridine in a substantially single step, thus allowing for the facile and efficient modification of pseudouridine into a cyclized derivative, particularly when contained in an RNA or any other molecule. The cyclized derivative, which has a different molecular weight and chemical properties, can then be used as a proxy for the identification of the presence of pseudouridine prior to the reaction taking place. In certain aspects, the modified pseudouridine comprises a reactive moiety capable of reacting with other molecules, as described below, such as a linker, an affinity tag, an imaging probe, or a fluorophore. Such reactive moiety can be of the type used in click chemistry, such as azide, alkyne, dibenzocyclooctyne alcohol (DIBO), azido-dibenzocyclooctyne (DBCO), bicyclononyne (BCN), trans-cyclooctene (TNO), or tetrazine. It can be appreciated that a second reaction step can be required to link the modified pseudouridine to the additional molecule, such as a linker, an affinity tag, an imaging probe, or a fluorophore.

[0040] In another aspect, the present application provides a method of tagging or labeling pseudouridine in a sample, comprising reacting at least a portion of the sample with a Michael acceptor, wherein the Michael acceptor is a compound according to Formula (I):

[0041] (I)

[0042] wherein:

[0043] X is independently selected from the group consisting of halogen, tosyl, mesyl, and triflate;

[0044] R 1 is independently selected from the group consisting of H, C1-C4-alkyl, C1-C4-haloalkyl, and Co-C4-alkylene-R 1a , wherein R 1a is independently selected from the group consisting of C3-C6-cycloalkyl, phenyl, and 4- to 6-membered heterocyclyl; wherein when R 1a is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 1a is optionally substituted with 1 to 4 R 1b , and when R 1a is phenyl, R 1a is optionally substituted with 1 to 4 R 1c ;

[0045] R 2 is independently selected from the group consisting of H, C1-C4-alkyl, C1-C4-haloalkyl, and Co-C4-alkylene-R 2a , wherein R 2a is independently selected from the group consisting of C3-C6-cycloalkyl, phenyl, and 4- to 6-membered heterocyclyl; wherein when R 2a is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 2aoptionally substituted by 1 to 4 R 2b substituents; and when R 2a is phenyl, R 2a is optionally substituted by 1 to 4 R 2c substituents;

[0046] R 3 is independently selected from the group consisting of -C(O)OR 3f , -C(O)R 3f , -C(O)NR 3g R 3f , -S(O)2OR 3f , -S(O)2R 3f , -S(O)2NR 3g R 3f ;

[0047] R 3f is a linker covalently attached to an affinity tag or an imaging probe;

[0048] R 3g is independently selected from the group consisting of H, C1-C6-alkyl and C1-C6-haloalkyl; R 1b and R 2b are each independently at each occurrence selected from the group consisting of =O, =S, halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl;

[0049] R 1c and R 2c are each independently at each occurrence selected from the group consisting of halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl;

[0050] R 4 is independently at each occurrence selected from the group consisting of H and C1-C4-alkyl; or when two R 4 groups are attached to the same nitrogen, the two R 4 groups together with the nitrogen atom to which they are attached optionally form a 5- to 6-membered heterocycloalkyl group, which is optionally substituted by 1 to 4 R 5 substituents; and

[0051] R 5each independently at each occurrence is selected from =0, =S, halogen, nitro, cyano, C1-C4-alkyl and C1-C4-haloalkyl.

[0052] The reaction according to this aspect of the application advantageously modifies pseudouridine in a substantially single step, thus allowing for the convenient and efficient modification of pseudouridine to a cyclized derivative, particularly when comprised in an RNA or any other molecule. Linking the resulting cyclized derivative to an affinity tag or imaging probe allows for the convenient isolation and / or identification of the derivative molecule using a variety of possible techniques, either of qualitative or quantitative nature.

[0053] The affinity tag can be one selected from the group comprising biotin, FLAG tag, His tag, HA tag, Strep tag, Avi tag, GST, c-myc tag, V5 tag, E tag, S tag, SBP tag, poly(Glu) tag, calmodulin tag.

[0054] The linker can be a flexible linker, a cleavable linker; optionally wherein the cleavable linker is a photocleavable linker. Particular linkers of the aforementioned types can be selected from (a) polyethylene glycol (PEG); (b) a peptide; (c) a nucleic acid; or (d) an oligosaccharide.

[0055] When a PEG linker is used, the number of PEG units can range from n = 1-12.

[0056] When a peptide is used as a linker, its molecular weight can range from 100 to 5000 g / mol.

[0057] When a nucleic acid is used as a linker, it can contain a number of nucleotides ranging from 1 to 40 nucleotides.

[0058] When an oligosaccharide is used as a linker, it is preferably an oligosaccharide containing 1 to 40 monosaccharides.

[0059] When an imaging probe is used, this can be selected from the group comprising a fluorescent moiety, a radionuclide and a metal complex.

[0060] When the imaging probe is a fluorophore, it can be selected from the group comprising fluorescein, rhodamine, BIODIPY, Alexa fluor, Cy dyes or ATTO dyes. Examples of fluorescein fluorophores include FAM, HEX or VIC. Examples of rhodamine fluorophores include ROX, TAMRA, TEX 615. Examples of Alexa Fluor include Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 594, Alexa Fluor 647, Alexa Fluor 660, Alexa Fluor 750. Examples of Cy dyes include Cy 3, Cy 5, Cy 5.5. Examples of ATTO Dye include ATTO 488, ATTO 532, ATTO 550, ATTO 565, ATTO Rho101, ATTO 590, ATTO 633, ATTO 647N. A particularly preferred fluorophore has a maximum excitation in the range 350 to 850 nm; preferably it is ATTO 488 or DY676.

[0061] In certain cases, a fluorophore such as FAM can be used in conjunction with an anti-FAM antibody in a pulldown enrichment similar to that described herein in relation to biotin-avidin.

[0062] When the pseudouridine is modified and the derivative comprises a fluorophore, then for certain sample types the presence, location and (optionally) amount or concentration of the fluorescent derivative can be determined in vitro. Suitable methods can include, for example, those described in Knutson, K. D. et al. (2018) Bioconjugate Chem. Vol 29(9): 2899 - 2903.

[0063] When the imaging probe is a lanthanide complex, it can be Gd, Mn, Dy or Eu.

[0064] When the imaging probe is a radionuclide complex, it can comprise 64 Cu, 68 Ga, 18 F, 99 mTc, 123 I, 125 I, 131 I, 57 Co, 51 Cr, 67 Ga, 64 Cu, 90 Y.

[0065] The application also includes a method of isolating RNA comprising pseudouridine from a sample, comprising reacting the sample with a Michael addition acceptor according to the method defined above, wherein the pseudouridine is tagged with an affinity tag, and then contacting the sample with a matrix, said matrix comprising a binding partner for the affinity tag. The Michael addition acceptor can already comprise the affinity tag as described herein, or the Michael addition acceptor can comprise a reactive moiety with which the affinity tag can react, either before, after or simultaneously with the reaction with the pseudouridine. In this way, isolation of the pseudouridine-containing RNA from the sample can be achieved. The matrix is preferably a solid phase matrix, such as agarose, cellulose, dextran, polyacrylamide, latex or controlled pore glass; and desirably, the solid phase matrix is porous.

[0066] As will be familiar to the skilled person, RNA affinity isolation using tagged pseudouridine can be performed in batch or column format. Typically, a series of steps involved include incubating the sample with the matrix under conditions which allow for maximum possible binding of the tagged RNA to the matrix. Unbound sample components are then washed from the matrix using a suitable buffer or buffers, during which time the tagged RNA remains bound to the matrix. The tagged RNA is then eluted from the matrix using altered buffer conditions which dissociate the tagged RNA from the matrix. In this way the tagged RNA can be efficiently collected.

[0067] Advantageously, the tagged RNA can be subjected to a pulldown enrichment process prior to further processing, for example including sequencing or reverse transcription. This can be useful when dealing with RNA having low abundance pseudouridine sites.

[0068] In another aspect, the solid phase affinity matrix can be used for the purposes of a binding assay, for detecting and optionally quantifying the presence of the corresponding affinity-tagged pseudouridine in a sample.

[0069] In a particularly preferred method of affinity isolation or quantification of pseudouridine-tagged RNA, the affinity tag is biotin and the binding partner of the solid phase is avidin or streptavidin. Alternatively, the affinity tag can be avidin or streptavidin and the binding partner of the solid phase is biotin.

[0070] In another aspect, the application provides a method of visualising pseudouridine in a sample comprising RNA, comprising reacting the sample with a Michael addition acceptor according to the method defined above, wherein the pseudouridine is tagged with an imaging probe, and subjecting the sample to a visualisation procedure selected from optical observation, microscopic observation and image capture.

[0071] The sample in relation to the visualization method can comprise tissue or cells. In such samples where individual cells can be distinguished, this allows the location and frequency of pseudouridines to be observed, and thereby correlated to cell type or cell location in the sample. The visualization method of pseudouridines can also be combined with probes for specific RNA sequences of interest.

[0072] In another aspect, the present application provides a method of determining the presence and sequence position of pseudouridines comprised in an RNA sample, comprising:

[0073] (a) reacting at least a portion of the sample with a Michael acceptor,

[0074] wherein the Michael acceptor is a compound according to formula (I):

[0075] (I)

[0076] wherein:

[0077] X is independently selected from halogen, tosyl, mesyl and triflyl;

[0078] R 1 is independently selected from H, Ci-C4-alkyl, Ci-C4-haloalkyl and Co-C4-alkylene-R 1a wherein R 1a is independently selected from C3-C6-cycloalkyl, phenyl and 4- to 6-membered heterocyclyl; wherein when R 1a is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 1a is optionally substituted with 1 to 4 R 1b and when R 1a is phenyl, R 1a is optionally substituted with 1 to 4 R 1c ;

[0079] R 2 is independently selected from H, Ci-C4-alkyl, Ci-C4-haloalkyl, Co-C4-alkylene-R 2a wherein R 2a is independently selected from C3-C6-cycloalkyl, phenyl and 4- to 6-membered heterocyclyl; wherein when R 2a is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 2a is optionally substituted with 1 to 4 R 2b and when R 2a is phenyl, R 2a is optionally substituted with 1 to 4 R 2c ;

[0080] R 3 is independently selected from -C(O)OR 3a , -C(O)R3a -C(O)NR 3b R 3b -CN, -NO2, -S(O)2OR 3a -S(O)2R 3a -S(O)2NR 3b R 3b and 5- to 10-membered heteroaryl, wherein when R 3 is 5- to 10-membered heteroaryl, R 3 is optionally substituted with 1 to 4 R 3e ;

[0081] R 3a is independently selected from the group consisting of H, C1-C6-alkyl, C1-C6-haloalkyl and C0-C6-alkylene-R 3c ;

[0082] R 3b is independently selected from the group consisting of H, C1-C6-alkyl, C1-C6-haloalkyl and C0-C6-alkylene-R 3c ; or wherein two R 3b groups together with the nitrogen atom to which they are attached form a 5- or 6-membered heterocycloalkyl group, which is optionally substituted with 1 to 4 R 3d ;

[0083] R 3c is independently selected from the group consisting of C3-C6-cycloalkyl, phenyl and 4- to 6-membered heterocyclyl; wherein when R 3c is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 3c is optionally substituted with 1 to 4 R 3d , and when R 3c is phenyl, R 3c is optionally substituted with 1 to 4 R 3e ;

[0084] R 1b , R 2b and R 3d are each independently at each occurrence selected from the group consisting of =O, =S, halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl;

[0085] R 1c , R 2c and R 3e are each independently at each occurrence selected from the group consisting of halogen, nitro, cyano, C(O)OR 4 , C(O)R4 C(O)NR 4 R 4 C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl;

[0086] R 4 are independently at each occurrence selected from H and C1-C4-alkyl; or when two R 4 groups are attached to the same nitrogen, the two R 4 groups together with the nitrogen atom to which they are attached optionally form a 5- to 6- membered heterocycloalkyl group, which is optionally substituted with 1 to 4 R 5 ; and

[0087] R 5 are each independently at each occurrence selected from =O, =S, halogen, nitro, cyano, C1-C4-alkyl and C1-C4-haloalkyl;

[0088] (b) (i) sequencing the RNA; or

[0089] (ii) reverse transcribing the RNA produced in step (a) to provide cDNA and amplifying the cDNA;

[0090] (c) sequencing the DNA of step (b)(ii);

[0091] (d) comparing the DNA sequence of step (c) with a reference DNA sequence to identify the sequence position of a guanine (G) in the sequence which is an adenine (A) in the reference sequence, the position of the G in the DNA sequence being the position of a pseudouridine in the corresponding RNA sequence.

[0092] The reference sequence can be obtained from a separate portion of the RNA sample which is not subjected to the Michael addition reaction of step (a) but is sequenced according to step (b)(i); or is reverse transcribed according to step (b)(ii) and sequenced according to step (c).

[0093] The amplification of the cDNA can employ an isothermal amplification method; wherein the isothermal amplification method is selected from the group consisting of polymerase chain reaction (PCR), strand displacement amplification (SDA), rolling circle amplification (RCA), whole genome amplification (WGA), loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA) and multiple displacement amplification (MDA); optionally wherein a one-step RT-PCT is used.

[0094] The reverse transcription of step (b) can use a reverse transcriptase (i.e. an "RNA-directed DNA polymerase"; EC 2.7.7.49); optionally a reverse transcriptase selected from the group consisting of Maxima H-, SuperScript II, SuperScript III, SuperScript IV, ProtoScript II, SMARTScribe, PrimeScript II, HiScript III, MMLV, AMV, TGIRT-III or recombinant HIV; preferably Maxima H-, SuperScript IV and TGIRT-III; more preferably Maxima H-. The Marathon reverse transcriptase or Induro reverse transcriptase can also be used, especially in methods involving tRNA, as they are Group II intron-encoded RT enzymes like TGIRT-III.

[0095] In the methods of the application, in which the modified RNA is sequenced directly, the preferred sequencing method is ion torrent sequencing, for example by using the Oxford nanopore system. In such direct sequencing of RNA, the modified pseudouridine residues can be detected directly. Optionally, a control or reference sample of RNA can be used, which has not been modified with pseudouridine, and which is also sequenced using the ion torrent sequencing method, allowing a side-by-side comparison of the modified and unmodified RNA.

[0096] When the methods of the application comprise a sequencing method, any suitable sequencing method familiar to the person skilled in the art can be used. Examples of such sequencing methods are explained briefly below.

[0097] In single molecule real-time (SMRT) sequencing (Pacific Biosciences), DNA is synthesized in zero-mode waveguides (ZMWs). These are pores with a capture sequence, including unmodified polymerase and fluorescently labeled nucleotides in solution. Only fluorescence occurring at the bottom of the pore is detected. SMRT sequencing allows for 20,000 nucleotides or more in a read, with an average read length of 5 kilobases.

[0098] Ion Torrent sequencing (Thermo Fisher Scientific) uses conventional sequencing chemistry but is equipped with a special semiconductor device. The basis of this technology relies on the detection of hydrogen ions released during DNA polymerization by their charge. Microwells containing the template DNA strand to be sequenced are exposed to only one type of nucleotide at a time. If the nucleotide is complementary to the template at the position being synthesized, it is incorporated and releases a hydrogen ion, which is measured to confirm that the reaction (of a particular base) has occurred. This sequencing method can provide individual read lengths of about 800 bp. As mentioned above, another form of this sequencing is the Oxford nanopore, which is useful in embodiments of the method of the application that require direct sequencing of the RNA molecules modified according to the method of the application.

[0099] Pyrosequencing (454) (Roche Diagnostics) uses water droplets in oil solution (emulsion PCR) as the environment for DNA synthesis. Each droplet contains a single DNA template attached to a primer-coated single bead, forming a clonal colony. The sequencer has picoliter volume wells, each with a single bead and sequencing enzymes. Luciferase is used to generate light to detect the incorporation event of a single nucleotide attached to the nascent DNA strand. This method provides read lengths of about 700 bp.

[0100] Synthetic sequencing (Illumina) has multiple versions such as MiniSeq, NextSeq, MiSeq, HiSeq 2500, HiSeq3 / 4000 and HiSeq X. In this method, DNA molecules and primers are first attached to a glass slide or flow cell and amplified with a polymerase, forming local clonal DNA colonies, which are called “DNA clusters”. To determine the sequence, four types of reversible terminator bases (RT-bases) are added and unincorporated nucleotides are washed away. A camera takes pictures of the fluorescently labeled nucleotides. Then, the dye and the 3’ blocking agent are chemically removed from the DNA, allowing the next cycle to begin. These methods can provide read lengths in the range of 50-600 bp.

[0101] Sequencing by ligation (SOLiD sequencing) (Thermo Fisher) employs sequencing by ligation. All possible oligonucleotide pools of a fixed length are labeled according to the sequencing position. The oligonucleotides are annealed and ligated; the signal information for the nucleotide at that position is generated by the preferential ligation of a DNA ligase for the matching sequence. Before sequencing, the DNA is amplified by emulsion PCR and the resulting beads (each containing a single copy of the same DNA molecule) are deposited on a glass slide. Read lengths of 50 bp are obtained.

[0102] Other sequencing methods will be apparent to those skilled in the art and are readily incorporated into the methods of the present application.

[0103] Preferably, in each aspect of the present application, X is selected from Br, CI, and I.

[0104] Also, preferably, in each aspect of the present application, R 1 and R 2 are each independently selected from H, -OR 1 , -NR 2 R 3 , -C(O)OR 3a , -C(O)R 3a , -C(O)NR 3b R 3b , and -CN; more preferably, wherein R 3 is -C(O)NR 3b R 3b , optionally, wherein R 3 is -C(O)NH2.

[0105] Preferably, R 3 is independently selected from -C(O)OR 3a , -C(O)R 3a , -C(O)NR 3b R 3b , and -CN; more preferably, wherein R 3 is -C(O)NR 3b R 3b , optionally, wherein R 3 is -C(O)NH2.

[0106] In each aspect of the present application, the compound of Formula (I) is preferably selected from:

[0107] , , , , , , , and .

[0108] In the methods of the present application, suitable reaction conditions are used, for example, the Michael addition reaction is carried out at a pH in the range of about 7.0 to about 9.5.

[0109] The reaction of the RNA molecule with the Michael addition acceptor can be performed at a pH ranging from about pH 7.0 to about pH 9.5. The term "within a range" herein includes the upper and lower limits of the range. The pH can also be within any of the following ranges: about pH 7.1 to about pH 9.5, about pH 7.2 to about pH 9.5, about pH 7.3 to about pH 9.5, about pH 7.4 to about pH 9.5, about pH 7.5 to about pH 9.5, about pH 7.6 to about pH 9.5, about pH 7.7 to about pH 9.5, about pH 7.8 to about pH 9.5, about pH 7.9 to about pH 9.5, about pH 8.0 to about pH 9.5, about pH 8.1 to about pH 9.5, about pH 8.2 to about pH 9.5, about pH 8.3 to about pH 9.5, about pH 8.4 to about pH 9.5, or about pH 8.5 to about pH 9.5. Alternatively, the pH can be within any of the following ranges: about pH 7.0 to about pH 9.4, about pH 7.0 to about pH 9.3, about pH 7.0 to about pH 9.2, about pH 7.0 to about pH 9.1, about pH 7.0 to about pH 9.0, about pH 7.0 to about pH 8.9, about pH 7.0 to about pH 8.8, about pH 7.0 to about pH 8.7, about pH 7.0 to about pH 8.6, or about pH 7.0 to about pH 8.5. Alternatively, the pH can be within any of the following ranges: about pH 7.5 to about pH 9.4, about pH 7.5 to about pH 9.3, about pH 7.5 to about pH 9.2, about pH 7.5 to about pH 9.1, about pH 7.5 to about pH 9.0, about pH 7.5 to about pH 8.9, about pH 7.5 to about pH 8.8, about pH 7.5 to about pH 8.7, about pH 7.5 to about pH 8.6, or about pH 7.5 to about pH 8.5. Alternatively, the pH can be within any of the following ranges: about pH 8.0 to about pH 9.4, about pH 8.0 to about pH 9.3, about pH 8.0 to about pH 9.2, about pH 8.0 to about pH 9.1, about pH 8.0 to about pH 9.0, about pH 8.0 to about pH 8.9, about pH 8.0 to about pH 8.8, about pH 8.0 to about pH 8.7, about pH 8.0 to about pH 8.6, or about pH 8.0 to about pH 8.5.Alternatively, the pH can be in any range from about pH 7.1 to about pH 9.0, about pH 7.2 to about pH 9.0, about pH 7.3 to about pH 9.0, about pH 7.4 to about pH 9.0, about pH 7.5 to about pH 9.0, about pH 7.6 to about pH 9.0, about pH 7.7 to about pH 9.0, about pH 7.8 to about pH 9.0, about pH 7.9 to about pH 9.0, about pH 8.0 to about pH 9.0, about pH 8.1 to about pH 9.0, about pH 8.2 to about pH 9.0, about pH 8.3 to about pH 9.0, about pH 8.4 to about pH 9.0, or about pH 8.5 to about pH 9.0.

[0110] The reaction can preferably be carried out at any pH selected from about pH 7.0, about pH 7.1, about pH 7.2, about pH 7.3, about pH 7.4, about pH 7.5, about pH 7.6, about pH 7.7, about pH 7.8, about pH 7.9, about pH 8.0, about pH 8.1, about pH 8.2, about pH 8.3, about pH 8.4, about pH 8.5, about pH 8.6, about pH 8.7, about pH 8.8, about pH 8.9, about pH 9.0, about pH 9.1, about pH 9.2, about pH 9.3, about pH 9.4, or about pH 9.5.

[0111] Suitable concentrations of reagents are used in the methods of the application, for example, wherein the Michael addition acceptor is present in a concentration ranging from about 10 mM to about 2 M; preferably in the range from about 100 mM to about 500 mM; more preferably about 250 mM.

[0112] Further, suitable reaction temperature conditions are used, desirably at ambient pressure, wherein the Michael addition reaction is carried out at a temperature in the range of about 25 °C to about 95 °C; preferably about 65 °C to about 95 °C; more preferably about 85 °C. The reaction of the RNA molecule with the Michael addition acceptor can be carried out at a temperature in the range of about 40 °C to about 85 °C. As with pH, the term "in a range" herein includes the upper and lower limits of the range of temperatures. For reactions according to the present application, the upper threshold temperature can be selected from about 84 °C, about 83 °C, about 82 °C, about 81 °C, about 80 °C, about 79 °C, about 78 °C, about 77 °C, or about 76 °C. This can be combined with a lower threshold temperature of about 40 °C, about 41 °C, about 42 °C, about 43 °C, about 44 °C, about 45 °C, about 46 °C, about 47 °C, about 48 °C, about 49 °C, about 50 °C, about 51 °C, about 52 °C, about 53 °C, about 54 °C, about 55 °C, about 56 °C, about 57 °C, about 58 °C, about 59 °C, about 60 °C, about 61 °C, about 62 °C, about 63 °C, about 64 °C, about 65 °C, about 66 °C, about 67 °C, about 68 °C, about 69 °C, about 70 °C, about 71 °C, about 72 °C, about 73 °C, or about 74 °C. The reaction according to the present application can be carried out at any temperature selected from about 40 °C, about 41 °C, about 42 °C, about 43 °C, about 44 °C, about 45 °C, about 46 °C, about 47 °C, about 48 °C, about 49 °C, about 50 °C, about 51 °C, about 52 °C, about 53 °C, about 54 °C, about 55 °C, about 56 °C, about 57 °C, about 58 °C, about 59 °C, about 60 °C, about 61 °C, about 62 °C, about 63 °C, about 64 °C, about 65 °C, about 66 °C, about 67 °C, about 68 °C, about 69 °C, about 70 °C, about 71 °C, about 72 °C, about 73 °C, about 74 °C, about 75 °C, about 76 °C, about 77 °C, about 78 °C, about 79 °C, about 80 °C, about 81 °C, about 82 °C, about 83 °C, about 84 °C, or about 85 °C.

[0113] Further, suitable reaction times are used according to temperature, for example wherein the Michael addition reaction is carried out for a time in the range of about 5 minutes to about 2 hours at a temperature of about 70 °C or higher; or for a time in the range of about 2 hours to about 16 hours at a temperature of about 70 °C or lower.

[0114] The Michael addition reaction according to the present application can be carried out at a temperature of about 70 °C or higher for a time period in the range of 10 minutes to 60 minutes. As with pH and temperature, the term "in a range" herein includes both the upper and lower limits of the specified range of time. For the reaction according to the present application, the upper threshold time can be selected from about 60 minutes, about 59 minutes, about 58 minutes, about 57 minutes, about 56 minutes, about 55 minutes, about 54 minutes, about 53 minutes, about 52 minutes, about 51 minutes, about 50 minutes, about 49 minutes, about 48 minutes, about 47 minutes, about 46 minutes, about 45 minutes, about 44 minutes, about 43 minutes, about 42 minutes, about 41 minutes, about 40 minutes, about 39 minutes, about 38 minutes, about 37 minutes, about 36 minutes, about 35 minutes. The foregoing can be combined with the following lower threshold times: about 10 minutes, about 11 minutes, about 12 minutes, about 13 minutes, about 14 minutes, about 15 minutes, about 16 minutes, about 17 minutes, about 18 minutes, about 19 minutes, about 20 minutes, about 21 minutes, about 22 minutes, about 23 minutes, about 24 minutes, about 25 minutes, about 26 minutes, about 27 minutes, about 28 minutes, about 29 minutes, or about 30 minutes.

[0115] The Michael addition reaction according to the present application can be carried out for any time period selected from about 10 minutes, about 11 minutes, about 12 minutes, about 13 minutes, about 14 minutes, about 15 minutes, about 16 minutes, about 17 minutes, about 18 minutes, about 19 minutes, about 20 minutes, about 21 minutes, about 22 minutes, about 23 minutes, about 24 minutes, about 25 minutes, about 26 minutes, about 27 minutes, about 28 minutes, about 29 minutes, about 30 minutes, about 31 minutes, about 32 minutes, about 33 minutes, about 34 minutes, about 35 minutes, about 36 minutes, about 37 minutes, about 38 minutes, about 39 minutes, about 40 minutes, about 41 minutes, about 42 minutes, about 43 minutes, about 44 minutes, about 45 minutes, about 46 minutes, about 47 minutes, about 48 minutes, about 49 minutes, about 50 minutes, about 51 minutes, about 52 minutes, about 53 minutes, about 54 minutes, about 55 minutes, about 56 minutes, about 57 minutes, about 58 minutes, about 59 minutes, or about 60 minutes. The present application does not exclude the possibility that the reaction time exceeds 60 minutes, e.g. about 70 minutes, about 80 minutes, about 90 minutes, or about 100 minutes.

[0116] In any aspect of the present application, the RNA molecule can be selected from one or more of mRNA, tRNA, rRNA, snRNA, miRNA, IncRNA, or circRNA. In particular, the RNA molecule can be from a biological sample.

[0117] The present application also includes a composition comprising a Michael addition acceptor according to formula (I) as defined above and a RNA molecule.

[0118] The present application also includes an RNA molecule comprising at least one carbamido-1, 02-ethano Ψ, nce1,2Ψ. Such a modified RNA molecule is the reaction product of an RNA molecule comprising one or more Ψ residues. The reaction of an RNA molecule comprising one or more Ψ residues results in an RNA molecule wherein at least one of said Ψ residues is modified to a carbamido-1, 02-ethano Ψ, nce1,2Ψ.

[0119] In some aspects, a proportion of Ψ residues are modified and the proportion can be any proportion in the range of 1% to 100% of all Ψ residues in the RNA molecule. The proportion can be selected from any proportion of at least 1%, at least 2%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%.

[0120] The present application includes a composition comprising a modified RNA molecule as defined herein.

[0121] The method of the present application can be used to determine the presence and sequence position of pseudouridines in any RNA molecule. The RNA molecule can be selected from any one of an mRNA molecule, a tRNA molecule, a rRNA molecule, a snRNA molecule, a miRNA molecule, or an IncRNA molecule. The RNA molecule can be from a cfRNA sample. Furthermore, the RNA molecule can be one RNA molecule of a plurality of RNA molecules, wherein the method of the present application further comprises quantifying the number of pseudouridines in the plurality of RNA molecules.

[0122] The RNA used in the method of the present application can be obtained from a sample, more specifically from an environmental or biological sample. Biological samples include samples taken from any organism. This includes medical or veterinary samples and the skilled person is well aware of the range of possible ways of obtaining such samples. For example biopsy methods performed on the body such as fine needle aspiration, needle core biopsy, vacuum assisted biopsy, incisional biopsy, excisional biopsy, punch biopsy, shave biopsy or skin biopsy. Normal tissue and corresponding tumour or cancerous tissue can be sampled and compared. Pseudouridine residues in RNA can be compared between normal and cancerous samples in the study of cancer development and metastasis. Other samples from which RNA can be obtained include blood, sweat, hair follicles, oral tissue, tears, menses, faeces or saliva.

[0123] Particularly preferred samples include blood, cell-free RNA, and liquid biopsy samples.

[0124] The sample can include, but is not limited to, tissue, cells, or biological material from or derived from cells. These cells can be selected from any of bone marrow mononuclear cells, buffy coat, dissociated tumor cells, epithelial cells, fibroblasts, hepatocytes, mesenchymal stem cells, myoblasts, PBMCS, purified immune cells (e.g., T cells, B cells, NK cells, etc.), or red blood cells (RBCs).

[0125] The biological sample can be a homogenous or heterogeneous population of cells or tissue. In some cases, the sample can be cell-free, and thus, for example, is serum or plasma.

[0126] Biological fluids can provide suitable samples for RNA for examination, such as blood, bile, bone marrow aspirate, breast milk, plasma, saliva, cerebrospinal fluid (CSF), serum, stool, sputum, oral or nasal fluid, urine, or synovial fluid.

[0127] According to the methods of the present application, tissue can provide a suitable sample containing RNA for examination, and such tissue can be collected by biopsy or surgical procedure. Some tissue can have been fixed, frozen, or otherwise treated for analysis.

[0128] Cell-free nucleic acid samples can be interrogated according to the methods of the present application. Such samples include cell-free DNA (cfDNA) and / or cell-free RNA (cfRNA). Nucleic acids in such samples can be isolated from the sample and optionally purified to varying degrees.

[0129] According to the present application, RNA can be extracted from the sample prior to reaction with the Michael acceptor. Commercial RNA extraction kits can be used, with specific examples of such kits including those manufactured and sold under the following names: (a) New England Biolabs Monarch® RNA MiniPrep Kit, (b) Qiagen AllPrep® PowerViral® DNA / RNA Kit, (c) Zymo Quick RNA™-Viral Fecal / Soil Microbe Microprep Kit, or (d) Zymo Quick RNA™-Viral with Inhibitor Removal. Other suitable commercial kits can be used, including the Quick-RNA Miniprep Kit (Zymo) or the Invitrogen™ TRIzol™ Plus RNA Purification Kit.

[0130] A suitable sample for effective application of the present invention can comprise at least, at most, or about 1 ng, 2 ng, 3 ng, 4 ng, 5 ng, 6 ng, 7 ng, 8 ng, 9 ng, or 10 ng of nucleic acid. In addition, about 20 ng, about 30 ng, 40 ng, 50 ng, 60 ng, 70 ng, 80 ng, 90 ng, 100 ng, 200 ng, 300 ng, 400 ng, 500 ng, 600 ng, 700 ng, 800 ng, 900 ng, 1 pg, 2 pg, 3 pg, 4 pg, 5 pg, 6 pg, 7 pg, 8 pg, 9 pg, or 10 pg. A sample can comprise an amount of nucleic acid within a range of nucleic acid weights defined by any of the above as a lower limit and any of the above as an upper limit.

[0131] The methods of the present invention can be used in clinical and biological research and diagnostics. The kinds of RNA molecules that can be analyzed using the methods of the present invention can include messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), long non-coding RNA (IncRNA), short non-coding RNA (sncRNA), microRNA (miRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), small interfering RNA (siRNA), short hairpin RNA (shRNA), and circular RNA (circRNA). The assessment is used to determine pseudouridine residues in the sequences of these RNA molecules.

[0132] The presence of pseudouridine residues can serve as a new biomarker for a disease or condition. Accordingly, the methods of the present invention can be used to find this biomarker in patient samples. Similarly, this biomarker can be used to monitor disease progression and / or patient responsiveness to drugs or other treatments. The methods of the present invention are particularly useful in oncology and cancer research. Examples of cancers that can be studied with the methods of the present invention can include common types of cancer such as bladder cancer, breast cancer, colon cancer, rectal cancer, endometrial cancer, kidney cancer, leukemia, liver cancer, lung cancer, melanoma, non-Hodgkin’s lymphoma, pancreatic cancer, prostate cancer, and thyroid cancer. Any type of other less common cancer can be studied or monitored according to the methods of the present invention.

[0133] Increasing knowledge about the relationship between pseudouridine in tRNA and cancer. For example, in relation to glioblastoma (Cui, Q. et al. (2021) “Targeting PUS7 suppresses tRNA pseudouridylation and glioblastoma tumorigenesis” Nat Cancer 2(9): 932 - 949) and acute myeloid leukemia (Guzzi, N. et al. (2022) “Pseudouridine-modified tRNA fragments repress aberrant protein synthesis and predict leukemic progression in myelodysplastic syndrome” Nat. Cell Biol. 24: 299 - 306). The methods of the present invention have the potential to deepen the understanding of these and other cancer types.

[0134] The present invention also provides a kit for modifying pseudouridine comprising: (a) a solution comprising a Michael addition acceptor as defined above; (b) instructions for reacting a sample comprising pseudouridine with the solution.

[0135] The present invention further comprises a kit for tagging or labelling pseudouridine comprised in RNA comprising: (a) a solution comprising a Michael addition acceptor as defined above; (b) instructions for reacting RNA with the solution.

[0136] The present invention further comprises a kit for determining the presence of a sequence position of pseudouridine comprised in RNA comprising: (a) a solution comprising a Michael addition acceptor as defined above; (b) instructions for reacting RNA with the solution.

[0137] In any of the foregoing cases, the kit of the present invention can comprise one or more of: (a) one or more buffers, (b) a reverse transcriptase, or (c) a DNA polymerase. The kit can also comprise one or more containers, wherein each container holds one of the elements of the kit.

[0138] When present, the reverse transcriptase can be selected from the group consisting of: Maxima H-, SuperScript II, SuperScript III, SuperScript IV, ProtoScript II, SMARTScribe, PrimeScript II, HiScript III, MMLV, AMV, TGIRT-III, recombinant HIV, Marathon reverse transcriptase, or Induro reverse transcriptase; preferably Maxima H-, SuperScript IV, and TGIRT-III; more preferably Maxima H-.

[0139] When present, the DNA polymerase can be selected from the group consisting of Taq DNA polymerase, Bst DNA polymerase, or Bsu DNA polymerase.

[0140] The kit according to the application can comprise instructions for its use. The instructions can be about how to incubate a nucleic acid molecule sample with the included solutions, e.g. the Michael acceptor as defined herein. The instructions can state the conditions necessary to modify at least a portion of the pseudouridines in the nucleic acid molecule. These conditions can include, for example, pH conditions, temperature conditions, incubation time, as described elsewhere herein. Examples of these conditions necessary to modify pseudouridines are disclosed herein. The instructions can include a statement about incubating the sample with the Michael acceptor for a given period of time, as disclosed herein.

[0141] In a further aspect of the application, the kit can comprise a polynucleotide kinase, such as T4 polynucleotide kinase.

[0142] The kit can optionally provide additional components useful in the procedure. These optional components include buffers, capture reagents, colorimetric reagents, labels, reaction surfaces, devices for detecting or controlling the sample. In addition to instructions, the kit can also include explanatory information.

[0143] Certain kits can comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 100, 500, 750, 1,000, or more probes, primers or primer sets, synthetic molecules, or inhibitors.

[0144] The kits can include components that can be individually packaged or placed in a container, such as a tube, bottle, vial, syringe, or other kind of container. Individual components can be present in the kit in concentrated form and diluted prior to use.

[0145] Some kits can include control nucleic acids, for example, an RNA molecule that does not contain pseudouridine as a negative control, and an RNA molecule that contains pseudouridine as a positive control.

[0146] In one embodiment, X is halogen. In one embodiment, X is selected from Cl (chlorine), Br (bromine), and I (iodine). In one embodiment, X is Br.

[0147] In one embodiment, R 1 is independently selected from H, C1-C4-alkyl, and C1-C4-haloalkyl. In one embodiment, R 1 is independently selected from H, C1-C2-alkyl, and C1-C2-haloalkyl. In one embodiment, R 1 is H.

[0148] In one embodiment, R 2 is independently selected from H, C1-C4-alkyl, and C1-C4-haloalkyl. In one embodiment, R 2 is independently selected from H, C1-C2-alkyl, and C1-C2-haloalkyl. In one embodiment, R 2 is H.

[0149] In one embodiment, R 1 and R 2 are both H. 1 In one embodiment, R 2 and R 3 are both H.

[0150] In one embodiment, R 3a is independently selected from -C(O)OR 3a , -C(O)R 3b , and -C(O)NR 3b R 3 . In one embodiment, R 3 is independently selected from -CN and -NO2. In one embodiment, R 3a is independently selected from -S(O)2OR 3a , -S(O)2R 3b , and -S(O)2NR 3b R 3 . In one embodiment, R

[0151] In one embodiment, R3 -C(O)OR 3a In one embodiment, R 3 -C(O)R 3a In one embodiment, R 3 -C(O)NR 3b R 3b .

[0152] In one embodiment, R 3a is H. In one embodiment, R 3a is C1-C6-alkyl. In one embodiment, R 3a is C1-C6-haloalkyl. In one embodiment, R 3a is C0-C6-alkylene-R 3c .

[0153] In one embodiment, R 3b is H. In one embodiment, R 3b is C1-C6-alkyl. In one embodiment, R 3b is C1-C6-haloalkyl. In one embodiment, R 3b is C0-C6-alkylene-R 3c .

[0154] In embodiments where R 3 -C(O)NR 3b R 3b is H and the other R 3b is H. In one embodiment, R 3b is as defined herein.

[0155] In one embodiment, R 3c is C3-C6cycloalkyl. In one embodiment, R 3c is phenyl. In one embodiment, R 3c is 4- to 6-membered heterocyclyl.

[0156] In one embodiment, R 3 -C(O)NH2. In one embodiment, R 3 -C(O)OMe. In one embodiment, R 3 -CN.

[0157] In one embodiment, R 1b , R 2b and R 3d are each independently at each occurrence selected from =O, =S, halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4R 4 C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl.

[0158] In one embodiment, R 1b R 2b and R 3d are each independently selected from =0, =S, halogen, nitro and cyano.

[0159] In one embodiment, R 1c R 2c and R 3e are each independently selected from halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl.

[0160] In one embodiment, R 1c R 2c and R 3e are each independently selected from halogen, nitro and cyano.

[0161] In one embodiment, R 4 is H. In one embodiment, R 4 is C1-C4-alkyl.

[0162] In embodiments relating to methods of tagging or labeling pseudouridines in a sample, R 3 is independently selected from -C(O)OR 3f , -C(O)R 3f and -C(O)NR 3g R 3f . In embodiments relating to methods of tagging or labeling pseudouridines in a sample, R 3 is -C(O)NR 3g R 3f .

[0163] In embodiments, R 3g is H. In embodiments, R 3g is C1-C6-alkyl, optionally C1-C3-alkyl. In embodiments, R 3g is C1-C6-haloalkyl, optionally C1-C3-haloalkyl.

[0164] In embodiments, R 3fis a linker covalently attached to an affinity tag. In embodiments, R 3f is a linker covalently attached to an imaging probe.

[0165] In embodiments, the affinity tag is selected from the group comprising biotin, FLAG tag, His tag, HA tag, Strep tag, Avi tag, GST, c-myc tag, V5 tag, E tag, S tag, SBP tag, poly(Glu) tag, calmodulin tag.

[0166] In embodiments, the affinity tag is biotin.

[0167] In embodiments, the imaging probe is selected from the group comprising a fluorescent moiety, a radionuclide and a metal complex.

[0168] In embodiments, the imaging probe is selected from the group comprising (a) a fluorophore with maximum excitation in the range of 350-850 nm; preferably ATT0488 or DY676; (b) a lanthanide complex; preferably Gd, Mn, Dy, Eu; (c) a radionuclide complex 64 Cu, 68 Ga, 18 F, 99 mTc, 123 I, 125 I, 131 I, 57 Co, 51 Cr, 67 Ga, 64 Cu, 90 Y.

[0169] In embodiments, the imaging probe is selected from the group comprising fluorescein (FAM, HEX, VIC), rhodamine (ROX, TAMRA, TEX615), BODIPY, Alexa Fluor (Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 594, Alexa Fluor 647, Alexa Fluor 660, Alexa Fluor 750), Cy dyes (Cy 3, Cy 5, Cy 5.5) and ATTO dyes (ATTO 488, ATTO 532, ATTO 550, ATTO 565, ATTO Rho101, ATTO 590, ATTO 633, ATTO 647N).

[0170] In one embodiment, the linker is a flexible linker; optionally selected from the group comprising (a) a polyethylene glycol, (b) a peptide; preferably a peptide having a molecular weight in the range of 100 to 5000 g / mol, (c) a nucleic acid; preferably a nucleic acid containing 1 to 40 nucleotides, or (d) an oligosaccharide; preferably an oligosaccharide containing 1 to 40 monosaccharides.

[0171] In one embodiment, R 3f may have the following structure:

[0172]

[0173] wherein:

[0174] n is an integer selected from 1, 2, 3, 4, 5 and 6; m is an integer selected from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 1 1 and 12;

[0175] Z is an affinity tag or an imaging probe as defined herein; and

[0176] Y is a diazene (-N=N-) group, a disulfide (-S-S-) group, a Dde group or a photocleavable group. BRIEF DESCRIPTION OF DRAWINGS

[0177] Embodiments of the application are further described below with reference to the accompanying drawings, in which:

[0178] Figure 1 is a schematic overview of the BACS reaction. After BACS labeling, nce 1,2 O of Ψ 2 cannot act as a hydrogen bond acceptor and is therefore read as C during RT.

[0179] Figure 2a is the BACS labeled MALDI characterization of the modified (Ψ) 10mer product and the unmodified (U) 10mer.

[0180] Figure 2b shows the Ψ depletion rate of HeLa total RNA and 1.8 kb 10% Ψ modified RNA after BACS treatment, quantified by UHPLC-MS / MS. Data given in two independent experiments.

[0181] Figure 3 shows the mutation rate of Ψ sites in a 72mer model RNA after BACS treatment. Data shown as mean ± s.d. of ten independent experiments (n = 10). Figure 4Cumulative (upper bar graph) and motif-dependent (lower table) results showing BACS conversion rates and false positive rates for synthetic 30mer NNΨNN and NNUNN spike-ins. Data are shown as the mean ± s.d. of three and eight independent experiments for NNΨNN (n = 3) and NNUNN (n = 8) spike-ins, respectively.

[0182] Figure 5a BACS calibration curve for Ψ stoichiometric quantification in the NNUNN motif is shown. Experiment was performed once. Figure 5b BACS calibration curve for Ψ stoichiometric quantification in the UGUAG motif is shown. Experiment was performed once.

[0183] Figure 6 is a flowchart of BACS library construction.

[0184] Figure 7 is a graph showing the number of Ψ sites identified in human cy-rRNA and mt-rRNA.

[0185] Figure 8a is a graph showing the ratio of BACS (dark) and control (light) samples in 28S rRNA. Figure 8b is a graph showing the ratio of BACS (dark) and control (light) samples in 18S rRNA. Figure 8c is a graph showing the ratio of BACS (dark) and control (light) samples in 5.8S rRNA.

[0186] Figure 9a shows BACS conversion rates for Ψ sites identified in 18S rRNA - data shown as the mean ± s.d. of four independent experiments (n = 4). Figure 9b shows modification levels for Ψ sites detected in 18S rRNA - data shown as the mean ± s.d. of four independent experiments (n = 4). Figure 9c shows BACS conversion rates for Ψ sites identified in 28S rRNA - data shown as the mean ± s.d. of four independent experiments (n = 4). Figure 9d shows modification levels for Ψ sites detected in 28S rRNA - data shown as the mean ± s.d. of four independent experiments (n = 4).

[0187] Figure 10 Correlation density plots between two biological replicates of BACS are shown. The extent of shading represents density.

[0188] Figure 11is a venn diagram showing the overlap of Ψ sites detected in human cy-rRNA between BACS and SILNAS MS.

[0189] Figure 12 is a comparison of conversion rates of cy-rRNA between BACS and control samples.

[0190] Figure 13 is a comparison of conversion rates of BACS with deletion rates of BID-seq and PRAISE for selected Ψ sites in 18S rRNA. Because BID-seq and PRAISE cannot quantify multiple Ψ sites (≥ 2) located in the same contiguous uridine environment, these sites were excluded.

[0191] Figure 14 is a comparison of conversion rates of BACS with deletion rates of BID-seq and PRAISE for selected Ψ sites in 28S rRNA. Because BID-seq and PRAISE cannot quantify multiple Ψ sites (≥ 2) located in the same contiguous uridine environment, these sites were excluded.

[0192] Figure 15 shows a comparison of Ψ site modification levels in cy-rRNA and mt-rRNA. Box plots represent median, quartiles, extremes, and outliers.

[0193] Figure 16 shows how BACS detects Ψ sites adjacent to one or more U and densely modified Ψ sites. Blue and red represent C and T bases, respectively. Ψ sites are identified by U to C mutations following BACS.

[0194] Figure 17 shows how BACS results are not affected by other modified uridine bases. Blue and red represent C and T bases, respectively. m1acp3Ψ and m3U sites can be easily filtered out by comparing BACS results with untreated controls.

[0195] Figure 18 shows the number of Ψ sites identified in human spliceosomal snRNAs.

[0196] Figure 19 is a venn diagram showing the overlap of Ψ sites detected in human spliceosomal snRNAs between BACS and SILNAS MS.

[0197] Figure 20 shows conversion rates of BACS (dark) and control (light) samples in U2 snRNA. Data are given in four independent experiments.

[0198] Figure 21 Conversion rates of BACS (dark) and control (light) samples for U4atac snRNA showing the new Ψ 11 and known Ψ 12 sites. Data given in four independent experiments.

[0199] Figure 22 Base pairing interactions between U4atac and U6atac snRNA in the stem II region are shown. The new Ψ 11 site is to the right of the known Ψ 12 site.

[0200] Figure 23 Number of Ψ sites identified in human snoRNAs and TERC.

[0201] Figure 24 a is a violin plot showing the overlap of Ψ sites detected in human snoRNAs between BACS and Ψ-seq. Figure 24 b is a violin plot showing the overlap of Ψ sites detected in human snoRNAs between BACS and BID-seq.

[0202] Figure 25 Number of Ψ sites identified in box C / D snoRNAs, box H / ACA snoRNAs and scaRNAs with high (50-100%, top), medium (20-50%, middle) and low (5-20%, bottom) level of modification.

[0203] Figure 26 Distribution of the level of modification of Ψ sites in box C / D snoRNAs, box H / ACA snoRNAs and scaRNAs. Boxplots visualize all Ψ sites in each class of snoRNAs showing the median, quartiles and extrema.

[0204] Figure 27 Macrogenomic plot of Ψ sites in box C / D snoRNAs. Each shade represents box C / C’ and D / D’.

[0205] Figure 28 Macrogenomic plot of Ψ sites in box H / ACA snoRNAs. Each shade represents box H and ACA.

[0206] Figure 29 Potential base pairing interactions between snoRNAs (dark) and their targets in rRNAs (light) for box C / D snoRNAs. Identified snoRNA Ψ sites are highlighted with position numbers. 2’-O-methylation targets in rRNAs are underlined.

[0207] Figure 30 Potential base-pairing interactions between snoRNAs (dark) and their targets in rRNAs (light) are shown for the box H / ACA snoRNAs. Identified snoRNA Ψ sites are highlighted with position numbers. 2'-O-methylation targets in rRNAs are underlined.

[0208] Figure 31 Ψ modification levels in TERC are shown, with each identified Ψ site labeled accordingly. Data is given in four independent experiments.

[0209] Figure 32 Distribution of identified Ψ sites in each human cy-tRNA paralog family (left) and mt-tRNA (right) is shown.

[0210] Figure 33 Median number of identified Ψ sites per tRNA in each cy-tRNA (left) and mt-tRNA (right) paralog family is shown. (n / a = not applicable).

[0211] Figure 34 Heatmap showing modification levels of high-confidence Ψ sites in human cy-tRNA. Only one representative tRNA paralog per paralog family is shown.

[0212] Figure 35 Overall view of Ψ profiles of human cy-tRNA (left) and mt-tRNA (right) is shown.

[0213] Figure 36 Comparison of modification levels of Ψ sites at selected positions in human cy-tRNA. Boxplots visualize all Ψ sites per position, showing median, quartiles and extremes.

[0214] Figure 37 Venn diagram showing overlap of Ψ sites in human mt-RNA reported by BACS and previously published datasets.

[0215] Figure 38 Heatmap of Ψ modification levels in human mt-tRNA.

[0216] Figure 39 Comparison of modification levels of Ψ sites at selected positions in human mt-tRNA. Boxplots visualize all Ψ sites per position, showing median, quartiles and extremes.

[0217] Figure 40 is a m 1 A to m 6Schematic overview of the Dimroth rearrangement of A.

[0218] Figure 41 is a heatmap showing the change in mutation rate of all adenosine sites in human mt-tRNA after BACS treatment. The degree of shading indicates an increase or decrease in mutation rate. 1 Comparison of mutation rates at A sites. Boxplot representation of median, quartiles and extremes.

[0219] Figure 42 is a heatmap showing the change in mutation rate of all adenosine sites in human mt-tRNA after BACS treatment. The degree of shading indicates an increase or decrease in mutation rate.

[0220] Figure 43 Distribution of mapped reads of polyA tail RNA in control library is shown. Data represent two independent experiments.

[0221] Figure 44 Number of Ψ sites with high (50-100%, top), medium (20-50%, middle) and low (5-20%, bottom) level of modification identified in HeLa polyA tail RNA is shown.

[0222] Figure 45 Distribution of modification levels of Ψ sites in HeLa polyA tail RNA is shown.

[0223] Figure 46 Example of genome reader view showing highly modified Ψ sites is shown. Top shading indicates T count. Bottom shading indicates C count.

[0224] Figure 47 Comparison of Ψ modification levels in different RNA classes is shown. Boxplot representation of median, quartiles and extremes.

[0225] Figure 48 Distribution of Ψ sites in different features of HeLa mRNA and ncRNA is shown.

[0226] Figure 49 Metagenomic plot of Ψ sites in HeLa mRNA is shown.

[0227] Figure 50 Gene ontology enrichment analysis (biological processes) of HeLa mRNA Ψ sites is shown.

[0228] Figure 51 Correlation density plot of RNA expression levels between BACS and control samples is shown. The degree of shading represents density. Data represent two independent experiments.

[0229] Figure 52Correlation of HeLa RNA expression levels between BACS and BID-seq input libraries. Pearson r values are shown.

[0230] Figure 53 Distribution of mRNA Ψ sites in single and contiguous uridine contexts is shown.

[0231] Figure 54 Motif frequencies of Ψ sites in HeLa mRNA are shown.

[0232] Figure 55 Distribution of modification levels of HeLa mRNA Ψ sites within selected motifs, median is shown in each plot.

[0233] Figure 56 Comparison of Ψ modification levels of selected motifs between tRNA (top panel) and mRNA (bottom panel). Boxplots indicate median, quartiles and extreme values.

[0234] Figure 57 Number of mRNA Ψ sites located in different codons is shown.

[0235] Figure 58 Codons encoding different amino acids are shown.

[0236] Figure 59 Number of mRNA Ψ sites located in different codon positions is shown.

[0237] Figure 60 Venn diagram illustrating the overlap of mRNA Ψ sites between the "highest confidence" list and the BACS and CMC-based integrated dataset is shown.

[0238] Figure 61 Venn diagram illustrating the overlap of mRNA Ψ sites between the "high confidence" list and the BACS and CMC-based integrated dataset is shown.

[0239] Figure 62 Venn diagram illustrating the overlap of mRNA Ψ sites between the "high confidence" list and the BACS and CMC-based integrated dataset is shown. Only Ψ sites consistently detected in more than 8 samples in the "high confidence" list are considered.

[0240] Figure 63 Venn diagram illustrating the overlap of mRNA Ψ sites between BACS and BID-seq is shown.

[0241] Figure 64 Distribution of Ψ sites in the BACS dataset only for BID-seq is shown.

[0242] Figure 65A Venn diagram showing the overlap of mRNA Ψ sites between BACS and PRAISE is shown.

[0243] Figure 66 A distribution of PRAISE-only Ψ sites in the BACS data set is shown. DETAILED DESCRIPTION

[0244] The inventors, seeking to improve existing pseudouridine detection and sequencing methods, developed a 2-bromoacrylamide-assisted circularization sequencing (BACS) method for direct and base-resolution sequencing of Ψ. B romo A crylamide-assisted C yclization S equencing, BACS) method for direct and base-resolution sequencing of Ψ.

[0245] BACS provides quantitative and base-resolution sequencing of Ψ. BACS uses a bromoacrylamide circularization chemistry to induce Ψ to C mutation signatures, rather than truncation or deletion signatures, allowing for more accurate quantification of Ψ stoichiometry. Importantly, BACS provides higher resolution compared to CMC-based methods, especially in highly structured regions. Furthermore, BACS overcomes the inherent limitations of BS-based methods in two key aspects: it enhances detection of densely modified Ψ sites with higher accuracy and sensitivity, and it facilitates precise determination of the exact location of Ψ sites positioned adjacent to one or more uridines. These advances make BACS a valuable tool for studying and understanding Ψ modifications in cellular RNAs, as it can provide a more comprehensive and accurate description of the Ψ landscape across various RNA species. Using BACS, the inventors successfully generated the first quantitative Ψ maps of human snoRNAs and tRNAs, illuminating the distribution of Ψ modifications in small RNAs. The combination of BACS with state-of-the-art technologies has the potential to further improve its performance on small RNAs. When applied to mRNAs, the reliability and robustness of BACS were again demonstrated, as it consistently produced results consistent with reported data sets. Thus, BACS is a powerful method for studying Ψ modifications.

[0246] BACS will allow the exploration of changes in Ψ across different cell types. BACS with minimal RNA degradation provides an excellent opportunity to combine with single-cell technologies, which can enable the study of Ψ dynamics across different cell populations. The BACS method of the present invention can enable the quantification and base resolution sequencing of Ψ, which can be used to study pseudouridylation of nascent RNA. Furthermore, there are 13 putative PUS enzymes in the human genome, and it remains challenging to identify the PUS enzymes responsible for many Ψ sites. Furthermore, their substrate specificity and potential redundancy are not fully clear. BACS can serve as a valuable tool to study PUS knockout cells to elucidate the properties and functions of these enzymes.

[0247] As used herein, the term C m -C n means a group with m to n carbon atoms. For the avoidance of doubt, the term "C0" means a group with 0 carbon atoms.

[0248] The term "alkyl" means a monovalent linear or branched saturated hydrocarbon chain. For example, C1-C6-alkyl can mean methyl, ethyl, n-propyl, i-propyl, n-butyl, s-butyl, t-butyl, n-pentyl and n-hexyl. Alkyl groups can be unsubstituted or substituted by one or more substituents.

[0249] The term "alkylene" means a divalent linear saturated hydrocarbon chain. For example, C1-C3-alkylene can mean methylene, ethylene or propylene. Alkylene groups can be unsubstituted or substituted by one or more substituents. For the avoidance of doubt, the term "C0-alkylene" means a group in which the alkylene chain is not present. For example, "C0-alkylene-R a " means R a .

[0250] The term "haloalkyl" means a hydrocarbon chain substituted with at least one halogen atom, which is independently selected at each occurrence from: fluorine, chlorine, bromine and iodine. The halogen atom can be present at any position on the hydrocarbon chain. For example, C1-C6-haloalkyl can mean chloromethyl, fluoromethyl, trifluoromethyl, chloroethyl (e.g. 1-chloromethyl and 2-chloroethyl), trichloroethyl (e.g. 1,2,2-trichloroethyl, 2,2,2-trichloroethyl), fluoroethyl (e.g. 1-fluoromethyl and 2-fluoroethyl), trifluoroethyl (e.g. 1,2,2-trifluoroethyl and 2,2,2-trifluoroethyl), chloropropyl, trichloropropyl, fluoropropyl, trifluoropropyl. Haloalkyl can be fluoroalkyl, i.e. a hydrocarbon chain substituted with at least one fluorine atom. Thus, haloalkyl can have any number of halogen substituents. The group can contain a single halogen substituent, it can have two or three halogen substituents, or it can be saturated with halogen substituents.

[0251] The term "alkenyl" refers to a branched or straight chain hydrocarbon chain containing at least one double bond. The double bond can be present as an E or Z isomer. The double bond can be at any possible location in the hydrocarbon chain. For example, "C2-C6-alkenyl" can refer to ethenyl, propenyl, butenyl, butadienyl, pentenyl, pentadienyl, hexenyl, and hexadienyl. The alkenyl group can be unsubstituted or substituted by one or more substituents.

[0252] The term "alkynyl" refers to a branched or straight chain hydrocarbon chain containing at least one triple bond. The triple bond can be at any possible location in the hydrocarbon chain. For example, "C2-C6-alkynyl" can refer to ethynyl, propynyl, butynyl, pentynyl, and hexynyl. The alkynyl group can be unsubstituted or substituted by one or more substituents.

[0253] The term "cycloalkyl" refers to a saturated hydrocarbon ring system containing 3, 4, 5, or 6 carbon atoms. For example, "C3-C6-cycloalkyl" can refer to cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl. The cycloalkyl group can be unsubstituted or substituted by one or more substituents.

[0254] The term "y- to z-membered heterocycloalkyl" refers to a y- to z-membered heterocycloalkyl group. Thus, it can refer to a monocyclic or bicyclic saturated or partially saturated group having y to z atoms in the ring system and comprising 1 or 2 heteroatoms in the ring system independently selected from O, S and N (in other words, 1 or 2 of the atoms forming the ring system are selected from O, S and N). Partially saturated means that the ring can contain one or two double bonds. This applies in particular to monocyclic rings having 5 to 6 members. The double bond is usually between two carbon atoms, but can also be between a carbon atom and a nitrogen atom. Examples of heterocycloalkyl groups include: piperidine, piperazine, morpholine, thiomorpholine, pyrrolidine, tetrahydrofuran, tetrahydrothiophene, dihydrofuran, tetrahydropyran, dihydropyran, dioxane, azepine. The heterocycloalkyl group can be unsubstituted or substituted by one or more substituents.

[0255] The aryl group can be any aromatic carbocyclic ring system (i.e. a ring system containing 2(2n + 1)π electrons). The aryl group can have 6 to 10 carbon atoms in the ring system. The aryl group is usually phenyl. The aryl group can be naphthyl or biphenyl.

[0256] The term "heterocyclyl" refers to a ring comprising 1 to 4 heteroatoms independently selected from O, S and N. The rings can be heterocycloalkyl rings (including saturated and partially saturated rings) or heteroaryl rings. The term "heterocyclyl" also includes groups which are tautomers of hydroxyheteroaryl groups, such as pyridinones, and tautomers of hydroxyheteroaryl groups which are substituted on the nitrogen, such as N-alkylpyridinones.

[0257] The term "heterocycloalkenyl" refers to a partially saturated ring comprising 1 to 2 heteroatoms independently selected from O, S and N.

[0258] The term "heteroaryl" refers to any aromatic (i.e., a ring system comprising 2(2n + 1) pi electrons) ring system comprising 1 to 4 heteroatoms independently selected from O, S, and N (in other words, 1 to 4 of the atoms forming the ring system are selected from O, S, and N). Thus, any heteroaryl can be independently selected from: 5-membered heteroaryl, wherein the heteroaromatic ring is substituted with 1-4 heteroatoms independently selected from O, S, and N; and 6-membered heteroaryl, wherein the heteroaromatic ring is substituted with 1-3 (e.g., 1-2) nitrogen atoms. In particular, the heteroaryl can be independently selected from: pyrrole, furan, thiophene, pyrazole, imidazole, oxazole, isoxazole, triazole, oxadiazole, thiadiazole, tetrazole; pyridine, pyridazine, pyrimidine, pyrazine, triazine, quinoline, isoquinoline, indole, benzofuran, benzopyrazole, benzimidazole.

[0259] Examples

[0260] Method

[0261] Model RNA preparation. Regular and Ψ-labeled 10mer RNA oligonucleotides and 30mer spike-in were purchased from Integrated DNA Technologies (IDT). Model RNA containing 72mer Ψ for mutation analysis and 1.8 kb 10% Ψ modified RNA for UHPLC-MS / MS were prepared by T7 in vitro transcription using HiScribe T7 High Yield RNA Synthesis Kit (New England Biolabs (NEB)) and pseudo-UTP (Jena Bioscience) according to the manufacturer’s protocol. DNA template was removed by adding 2 μΐ of Turbo DNase (Invitrogen) to the reaction and incubating at 37°C for 30 min. The product was finally purified with Monarch RNA Cleanup Kit (NEB). The sequences of the RNA oligonucleotides can be found in Table 1 below:

[0262] Table 1

[0263]

[0264] Mass spectrometric analysis of short oligonucleotides. MALDI was performed on a Voyager-DE Biospectroscopic Workstation (Applied Biosystems) with 2’,4’,6’- trihydroxyacetophenone (THAP) as matrix. All oligonucleotides were analyzed in the positive mode.

[0265] Ψ levels were quantified by UHPLC-MS / MS. Untreated and treated RNA were digested into nucleosides by nucleoside digestion mix (NEB) in 50 mΐ solution according to the manufacturer’s protocol. After filtration with Amicon Ultra-0.5 mL 3K centrifugal filters (Millipore), the digested samples were analyzed by UHPLC-MS / MS as described in Muller, C. A. et al. Nat. Methods 16, 429-436 (2019). A 1290 Infinity LC system (Agilent) equipped with a ZORBAX RRHD SB-C18 column (2.1 x 150 mm, 1.8 pm, Agilent) was connected to a 6495B triple quadrupole mass spectrometer (Agilent). Ions were monitored in positive mode with mass transitions m / z 245 to 125 (Ψ+H) and m / z 245 to 113 (rU+H) according to the compound-dependent UHPLC-MS / MS parameters for nucleoside quantification shown in Table 2 below.

[0266] Table 2

[0267]

[0268] RT: retention time; CE: collision energy.

[0269] The concentration of nucleosides in the RNA samples was inferred by fitting the signal peak area into a standard curve.

[0270] Cell culture. HeLa cells were cultured in DMEM medium (Gibco) at 37 °C with 5% CO2, supplemented with 10% (v / v) FBS (Gibco) and 1% penicillin / streptomycin (Gibco). For RNA isolation, cells were harvested by centrifugation at 1,000 x g and room temperature for 5 min.

[0271] RNA isolation. Total RNA was isolated using TRIzol (Invitrogen) and Direct-zol RNA mini prep Plus (Zymo Research) according to the manufacturer’s protocol. Ribo-RNA was isolated using RiboMinus Eukaryotic System v2 (Invitrogen) according to the manufacturer’s protocol. PolyA +RNA. To remove genomic DNA contamination, RNA was then treated with Turbo DNase and purified by Zymo-IC columns with RNA binding buffer.

[0272] BACS for Ψ detection. 50-100 ng ribo – or polyA + RNA was fragmented for 4 min at 94°C according to the manufacturer’s protocol by NEBNext Magnesium RNA Fragmentation Module and purified by Zymo-IC columns with RNA binding buffer. Fragmented RNA was mixed with 5 μl 10x T4 PNK reaction buffer (NEB), 5 μl T4 PNK (NEB), and 2.5 μl SUPERase•In RNase Inhibitor (Invitrogen) in 50 μL final solution and incubated at 37°C for 1 h. 3’-repaired RNA was purified by Zymo-IC columns with RNA binding buffer and eluted with 10 μl nuclease-free H2O. Eluted RNA was then mixed with 1 μl of synthetic 30mer spike-in (2%) and 1 μl 20 μM RNA adapter (5’- / 5rApp / AGATCGGAAGAGCGTCGTG / 3SpC3 / -3’), incubated at 70°C for 2 min, and immediately placed on ice. Next, 2.5 μl 10x T4 RNA ligase reaction buffer (NEB), 1 μl SUPERase•In RNase Inhibitor, 7.5 μl 50% PEG 8000 (NEB), and 2 μl T4 RNA ligase 2, truncated KQ (NEB) were added to the mixture and the reaction was incubated at 25°C for 2 h and then at 16°C for 14 h. To digest excess adapter, the solution was further diluted to 47 μl with nuclease-free H2O, treated with 2 μl 5’-Deadenylase (NEB) at 30°C for 1 h, and then 1 μl RecJ f(NEB) and incubated at 37°C for 1 h. The 3 '-ligated RNA was purified by Zymo-IC column with RNA binding buffer and eluted with 10 μΐ nuclease-free H20. A 7 μΐ aliquot was subjected to BACS library construction, while the remaining 3 μΐ was saved as a control sample and diluted to 12.5 μΐ with nuclease-free H20. For BACS, 1 m 2-bromopropenamide (Enamine) was prepared by dissolving the solid in DMSO. The 7 μΐ 3 '-ligated RNA was added to a 20 μΐ solution containing 250 mM 2-bromopropenamide and 625 mM phosphate buffer (pH 8.5) and incubated at 85°C for 30 min. The treated RNA was double purified by Micro Bio-Spin P-6 Tris column (Bio-Rad) and Zymo-IC column with RNA binding buffer, and finally eluted with 12.5 μΐ nuclease-free H20.

[0273] The treated and control RNA samples were mixed with 1 μΐ of 2 μΜ RT primer (5'-ACACGACGCTCTTCCGATCT-3') and 1 μΐ of 10 mM dNTP mix (NEB), incubated at 70°C for 2 min, and immediately placed on ice. Next, 4 μΐ 5x Maxima H-RT buffer (Thermo), 0.5 μΐ RiboLock RNase inhibitor (Thermo), and 1 μΐ Maxima H –Reverse Transcriptase (Thermo) and incubate the reaction at 50 °C for 1 h. To digest excess RT primer, treat the solution with 1 μΐ Exo I (NEB) and incubate at 37 °C for 30 min, followed by the addition of 1 μΐ 0.5 M EDTA (Sigma) to quench the reaction. To hydrolyze the RNA, add 2.5 μΐ 1 M NaOH (Sigma) and then incubate the solution at 70 °C for 12 min, followed by the addition of 2.5 μΐ 1 M HC1 (Sigma) to neutralize the NaOH. Finally, purify the cDNA with Dynabeads MyOne Silane (Invitrogen) and elute with 13 μΐ nuclease-free H2O. The eluted cDNA is then mixed with 2 μΐ 25 μΜ cDNA adaptor (5’- / 5Phos / NNNNNNAGATCGGAAGAGCACACGTCTG / 3SpC3 / -3’), incubate at 70 °C for 2 min and immediately place on ice. Next, add 5 μΐ 10x T4 RNA Ligase Reaction Buffer, 25 μΐ 50% PEG 8000, 0.5 μΐ 100 mM ATP (NEB), 3.5 μΐ DMSO (Thermo) and 1 μΐ high-concentration T4 RNA Ligase 1 (NEB) to the mixture and incubate the reaction at 25 °C for 16 h. The ligated cDNA is purified with Dynabeads MyOne Silane and eluted with 15 μΐ nuclease-free H2O. The eluted DNA is amplified with NEBNext Multiplex Oligos for Illumina (96 unique dual index primer pairs) and NEBNext Ultra II Q5 Master Mix according to the manufacturer’s protocol for 10 cycles. The PCR product is purified with 0.8x AMPure XP beads and quantified with the Qubit dsDNA HS Assay Kit (Thermo) according to the manufacturer’s protocol. The BACS and control libraries are sequenced on a NextSeq 2000 (60 bp paired-end) without the addition of PhiX.

[0274] Data preprocessing. Raw sequencing reads were processed by Cutadapt v.4.2 (Martin, M. EMBnet.journal 17, 10-12 (2011)) to remove low-quality bases (-q 20) and short reads (-m 18), and to trim adapters. 6mer UMIs were extracted by UMI tool extract v.1.0.1 (Smith, T. et al., Genome Res. 27, 491-499 (2017)) and used for deduplication. Paired reads were then merged into single reads using fastp v.1.0.1 (Chen, S. F. et al., Bioinformatics 34, 884-890 (2018)).

[0275] Read alignment. Purified reads were first mapped to synthetic spike-in and rRNA references using bowtie2 v.2.4.4 (Langmead, B. & Salzberg, S. L. Nat. Methods 9, 357-359 (2012)) with the following key parameters: bowtie2 -p 2 --no-unal --local -L 16 -N 1 --mp 4. Unmapped reads were then mapped to snoRNA and tRNA references using the same parameters. Human snoRNA sequences were downloaded from RefSeq belonging to HGNC “small nucleolar RNA” genomic class. Repeated snoRNA sequences were removed. High-confidence human tRNA sequences (hg38) were downloaded from GtRNAdb (Chan, P. P. & Lowe, T. M. Nucleic Acids Res. 44, D184-D189 (2016)). Only non-redundant tRNA sequences were kept and appended with a “3’-CCA” end. Finally, unmapped reads were aligned to human genome (hg38) with STAR v.2.7.9a GENCODE v.43 annotation (Dobin, A. et al. Bioinformatics 29, 15-21 (2013)). Aligned reads were then filtered and classified using samtools v.1.16.1 (Li, H. et al. Bioinformatics 25, 2078-2079 (2009)). For synthetic spike-in and rRNA, only reads with MAPQ > 10 were kept. For snoRNA and tRNA, only reads with MAPQ > 1 were kept. For mRNA, only uniquely mapped reads with at most 3 mutation counts were kept (-q 30). Deduplication was performed using UMI tool dedup v.1.0.1 (Smith, T. et al. Genome Res. 27, 491-499 (2017)). In addition, GATK ClipReads (v.4.1.7.0) 81 Patching of polyC counts (more than 3 cytosines) at the start and end of STAR aligned reads was performed to avoid potential false positive signals. Finally, mutations were counted by samtools mpileup v.1.16.1 (McKenna, A. et al. Genome Res. 20, 1297-1303 (2010)) and cpup (v.0.1.0) (https: / / github.com / y9c / cpup).

[0276] Ψ site calling. BACS raw conversion rate was calculated as C / (T+C). Ψ modification level was calculated using the following linear equation: Ψ modification level = (R-F) / (C-F), where R, F and C represent raw conversion rate, motif-specific false positive rate (from NNUNN spike-in) and motif-specific conversion rate (from NNΨNN spike-in), respectively. P-value for each site was calculated using motif-specific false positive rate and then adjusted following Benjamini-Hochberg (BH) procedure. The following criteria were used to call Ψ sites: (1) coverage in both BACS and control library is higher than 20; (2) in control library, background conversion rate is lower than 0.01 or T to C mutation count is lower than 2; (3) Ψ modification level is higher than 0.05; (4) adjusted p-value is lower than 0.001; (5) consistently detected in all replicates. To call cy-tRNA Ψ sites, criteria (3) and (5) were modified to require Ψ modification level higher than 0.10 in at least two out of three replicates. Only Ψ sites identified in expressed cy-tRNA isoacceptors were reported.

[0277] RNA structure visualization. RNA-RNA interactions were visualized using r2r v.1.0.6 (Weinberg, Z. & Breaker, R. R. BMC Bioinformatics 12, 3 (2011)). snoRNA-rRNA interactions were adapted from snoRNA Atlas.

[0278] Downstream analysis. snoRNA boxes and guide sequences were downloaded from snoDB 2.0 (Bergeron, D. et al., Nucleic Acids Res. 51 (2022)). In metagenomic analysis, snoRNA sequences showing considerable similarity were condensed, leaving only one representative snoRNA. Ψ sites identified in poly-A tail RNAs were annotated using bedtools intersect v.2.30.0 (Quinlan, A.R. & Hall, I. M. Bioinformatics 26, 841-842 (2010)) and GENCODE v.43 annotation. Ψ sites in regions of interest were visualized by Integrative Genomics Viewer (IGV) (Robinson, J. T. et al., Nat. Biotechnol. 29, 24-26 (2011)).

[0279] Read counts obtained from featureCounts v.1.6.4 (Liao, Y. et al., Bioinformatics 30, 923-930 (2014)) were normalized based on sequencing depth and gene length using the transcripts per million (TPM) method. GO analysis of mRNA Ψ sites was performed using enrichR (Kuleshov, M. V. et al., Nucleic Acids Res. 44, W90-W97 (2016)).

[0280] Publicly available data. Relevant publicly available data were downloaded from the Gene Expression Omnibus (GEO) database: BID-seq of HeLa cells (GSE179798) (Dai, Q. et al., Nat. Biotechnol. 41, 344-354 (2023)).

[0281] Example 1: Reaction of oligonucleotides with 2-bromopropenamide

[0282] 10mer short Ψ-tagged and U-tagged RNA oligos for MALDI were purchased from IDT. The oligo containing pseudouridine, 5’-UACUGΨAGCU-3’ [SEQ ID NO: 1], was reacted with 2-bromopropenamide. The control oligo lacking pseudouridine, 5’-UACUGUAGCU-3’ [SEQ ID NO: 2], was reacted with 2-bromopropenamide alone under the same time and same conditions.

[0283] The reaction products of the experiment and control were then analyzed by MALDI. As shown in Figure 2a , the observed mass of the Ψ oligo increased from 3122.5 Da to 3191.9 Da after the reaction, thus a mass increase of 69 Da was found. In contrast, the mass of the control oligo did not increase significantly. The calculated mass of each reaction product is shown directly below its respective sequence. The increase in mass value of the oligo containing pseudouridine was found to be evidence of the formation of the cyclization product (ureido-1, O 2 -ethanol Ψ, nce 1,2 Ψ). This reaction was further confirmed by ultra-high performance liquid chromatography-tandem mass spectrometry (UHPLC-MS / MS) (see Figure 2b ).

[0284] Ψ was compared to U, which contains a free N1 atom, and it was found to be highly reactive towards Michael acceptors such as acrylonitrile, acrylamide, and other acrylate compounds. Reference is made to Figure 1, the presence of an alpha-halo group on the acceptor induces the O 2 - Intramolecular alkylation induces tandem cyclization of the N1-acrylic acid adduct of Ψ, thereby inducing the desired Ψ to C mutation.

[0285] Example 2: nce 1,2 Determination of U to C mutation profile of Ψ

[0286] 72mer in vitro transcribed Ψ / U containing RNA was used to validate nce 1,2 Ψ to C mutation profile. To prepare model RNA and spike-in, 72mer Ψ and U containing RNA oligos were synthesized by T7 in vitro transcription using HiScribe® T7 High Yield RNA Synthesis Kit (NEB) and pseudo UTP (Jena) or UTP, and ATP, CTP and GTP according to the manufacturer’s protocol. Template DNA was removed by adding 2 μΐ Turbo DNase (Thermo) to the reaction and incubating at 37°C for 30 min. The template DNA sequence used was: TM

[0287] 5’-GTTGTCTTTGCCTTCGCTTCGGTCCTCGATTTCTGTTGTTGTACCGTTGGTTTCGTTGTGGTGTGTTCTC CCTATAGTGAGTCGTATTA -3’ [SEQ ID NO:3]

[0288] The final RNA sequence is:

[0289] 5’-GGGAGAACACACCACAACGAAACCAACGG(Ψ / U)ACAACAACAGAAA(Ψ / U)CGAGGACCGAAGCGAAGGCAAAGACAAC-3’ [SEQ ID NO:4]

[0290] The product was finally purified with Monarch® RNA Cleanup Kit (NEB). By sequencing, 80% U to C mutation rate was observed on both sites, while U to R (R = A or G) mutation rate was below 1% (see Figure 3 ). Therefore, U to C mutation rate can be used as the conversion rate of BACS.

[0291] Example 3: Sequence preference of BACS

[0292] To further demonstrate the sequence preference of BACS, libraries were generated with synthetic 30mer RNA spike-in containing NNΨNN and NNUNN (N = A, C, G or U) respectively. As Figure 4 ​As shown, after BACS, the conversion rate of Ψ was 82.7% and the false positive rate of uridine was 0.7% when all motifs were accumulated. Out of all 256 motifs, 224 showed conversion rates higher than 80% and 255 showed conversion rates higher than 70%, indicating the high efficiency of BACS chemistry. Low false positive rates (< 1%) were observed in most motifs (214 out of 256 motifs). Certain motifs, particularly those with one or more cytidines flanking the 5’- or 3’-side of the uridine site (e.g. GCUCC and ACUCC), showed slightly higher false positive rates (3%), which can be due to the bias of RT. However, BACS clearly showed higher conversion rates and lower false positive rates than BID-seq both in general and on specific motifs. Figure 5a and 5b As shown, by mixing NNΨNN and NNUNN spike-ins at different ratios, excellent calibration curves were generated for accurate quantification of Ψ modification levels (r 2 = 1.00).

[0293] Example 4: Validation of BACS on human rRNA

[0294] BACS was applied to cytosolic rRNA (cy-rRNA) of HeLa cells, which is known to have a series of highly conserved Ψ sites. Figure 6 The workflow of library generation is shown. Figure 7 Based on the BACS-induced U to C mutation signals, 2, 40 and 62 Ψ sites were detected in 5.8S, 18S and 28S rRNAs, respectively. Most of the detected sites showed high modification levels (> 80%), consistent with the fact that Ψ sites in human cy-rRNA43are highly modified (see Figure 8a , 8b , 8c, 9a, 9b, 9c and 9d). The raw signals of BACS from two biological replicates were examined. As shown in Figure 10 , this revealed strong correlation between them (Pearson r = 1.00 for two biological replicates). Compared with the reported SILNAS mass spectrometry (SILNAS MS) results, 103 out of 105 known Ψ sites in human cy-rRNA were identified with high confidence (see Figure 11 ). However, Ψ1136in 18S rRNA was not detected, which can be due to the low modification level of 3.8% for it (see Fig. 9a and 9b). Interestingly, the known Ψ36site in 18S rRNA was found to have a 20% U to C mutation rate in the control library, although the mutation rate increased to 75% after BACS treatment (see Figure 8b and Figure 12). Similar results were obtained from the BID-seq control library, suggesting that there might be uncharacterized single nucleotide polymorphism (SNP) sites. In addition, a new Ψ4938 site was detected in 28S rRNA, located near the previously known Ψ4937 site. The presence of Ψ4938 was supported by two public databases, which both predicted that the small nucleolar RNA (snoRNA) SNORA17B would be responsible for catalyzing this modification (see Jorjani, H. et al., Nucleic Acids Res. 44, 5068-5082 (2016) and Tan, K. T. et al., Sci. Adv. 7, eabd2605 (2021)). Thus, BACS further confirmed the presence of the Ψ4938 site. Notably, while some Ψ or uridine modifications can induce intrinsic mutation signals (such as the U-to-C mutation of m1acp3Ψ1248 in 18S rRNA and the U-to-A mutation of m3U4500 in 28S rRNA), these signals can be easily filtered out by comparing the results of BACS libraries and control libraries (see Figure 12 ).

[0295] As expected, BACS is clearly superior to BS-based methods in several aspects. First, the U-to-C mutation signature enables BACS to determine the accurate positions of Ψ sites in consecutive uridine sequences (adjacent to one or more uridines (e.g., Ψ801 / Ψ814 / Ψ815 / Ψ822 in 18S rRNA and Ψ1847 / Ψ1849 in 28S rRNA)) and dense Ψ sites in narrow regions (e.g., Ψ3737 / Ψ3741 / Ψ3743 / Ψ3747 / Ψ3749 in 28S rRNA and Ψ4263 / Ψ4266 / Ψ4269 in 28S rRNA), which remain challenging for BS-based methods (see Figure 8a , 8b , 13 and 14). Compared to BS-based methods, more uniform Ψ site conversion rates were obtained using BACS in different regions of rRNA, indicating that BACS results are not significantly affected by the density of pseudouridylation, thus enabling more accurate quantification of Ψ stoichiometry (see Figure 13 and 14 ). In addition, it was found that BACS was able to achieve a higher conversion rate (85%) at the 28S rRNA Ψm3797 site than BS-based methods (10-20%), as BACS relies only on the availability of the N1 atom of Ψ (see Figure 14 ).

[0296] In addition to cy-rRNA, BACS was applied to mitochondrial rRNA (mt-rRNA), and 6 and 1 Ψ sites were detected in 12S and 16S rRNA, respectively. Four of them were also detected by pseudouridine-seq4. In general, the modification level of Ψ sites in mt-rRNA was significantly lower than its cytosolic counterpart (see Figure 15 ).

[0297] As shown in Figure 16 and 17 , the mutation and deletion profile induced by BACS treatment was examined for each base (A, C, G, U, and Ψ). In general, BACS exhibited a high mutation rate at known Ψ sites, while exhibiting a low background at unmodified A, C, G, and U sites. In addition, U to C mutations were confirmed to be the major type of Ψ mutation, and thus can be used to calculate the conversion rate of Ψ. BACS did not cause significant deletion features at Ψ or other bases, which addressed a fundamental problem in BS-based methods. As can be seen from Figure 16 , the mutation profile can accurately detect one or more Ψ sites near U (e.g., 18S rRNA Ψ801, Ψ822 and 28S rRNA Ψ1847, Ψ1849), or dense Ψ sites in a narrow region (18S rRNA Ψ801, Ψ814, Ψ815, Ψ822; 28S rRNA Ψ3737, Ψ3741, Ψ3743, Ψ3747, Ψ3749; Ψ4263, Ψ4266, Ψ4269, Ψ4282). Although some Ψ or U modifications would induce intrinsic mutation profiles (U to C mutation of 18S rRNA m1acp3Ψ1248 and U to A mutation of 28S rRNA m3U4500), Figure 17 shows that these sites can be easily excluded by comparing BACS results with untreated RNA-seq data.

[0298] Example 5: BACS identifies highly conserved Ψ sites in human spliceosomal snRNAs

[0299] BACS was validated by applying it to the spliceosomal snRNAs from HeLa cells, which are known to contain multiple consecutive Ψ sites. The initial focus was on the major spliceosomal snRNA species. 2, 14, 3, 4, and 4 Ψ sites were detected in U1, U2, U4, U5, and U6 snRNAs, respectively (see Figure 18 and 19 ), which is highly consistent with the latest SILNAS MS results by Yamaki, Y. et al., Anal. Chem. 92, 11349-11356 (2020). Only Ψ 59was not detected by BACS, as it can be hypo-modified in HeLa cells. It is worth noting that BACS successfully mapped all 14 Ψ sites in human U2 snRNA, which is unattainable by any other high-throughput sequencing method, further demonstrating the advantage of BACS in detecting dense and contiguous Ψ sites (see Figure 20 ). Unlike the snRNA components of the human major spliceosome, the Ψ profiles of the minor spliceosomal snRNA species were only revealed using the CMC-based primer extension assay, mainly because of their low abundance. However, considering the potential difficulties of the CMC-based method in partial labeling efficiency and "stuttering" phenomenon, we believe that BACS can be a better method to study the pseudouridylation of these snRNA species. In fact, 2, 2, and 1 Ψ sites were consistently detected in U12, U4atac, and U6atac snRNAs, respectively, while no Ψ site was detected in U11 snRNA (see Figure 18 ). It is worth noting that two consecutive Ψ sites (Ψ11 / Ψ12) instead of one Ψ12 site in U4atac snRNA were confirmed, which provides new insights into its interaction with U6atac snRNA (see Figure 21 and 22 ).

[0300] In addition, a conserved Ψ 247 and Ψ 250 site was detected in 7SK RNA, and no high-confidence Ψ site was revealed in U7 snRNA, RNase P RNA, RNase MRP RNA, vault RNA, and Y RNA. However, the known Ψ 211 site was not detected in 7SL RNA, possibly due to the difference in cell lines.

[0301] Example 6: BACS reveals Ψ profiles of human snoRNAs

[0302] The Ψ profile of yeast snoRNAs has been revealed by pseudouridine-seq and Ψ-seq, however it remains relatively unexplored in human snoRNAs. Using BACS, 282 Ψ sites were detected in the snoRNAs of HeLa cells, including those previously identified by Ψ-seq and BID-seq (see Figure 23 、 24Ψ sites in the box C / D snoRNAs, box H / ACA snoRNAs, and small Cajal body- specific RNAs (scaRNAs) were identified (Figures 24a and 24b). Analysis revealed 192, 62, and 28 Ψ sites in the box C / D snoRNAs, box H / ACA snoRNAs, and scaRNAs, respectively. Notably, all three types of snoRNAs exhibited a large number of highly modified Ψ sites (see Figure 25 and 26 ). Moreover, Ψ sites observed in the box C / D snoRNAs were enriched in the 5 '-upstream region of the box D' and the 3 '-downstream region of the box C', while Ψ sites in the box H / ACA snoRNAs were enriched in the 5 '-upstream region of the box H and the ACA (see Figure 27 and 28 ). These patterns imply a potential role of Ψ in mediating the interaction between snoRNAs and their targets. In fact, a subset of Ψ sites identified in the box C / D and box H / ACA snoRNAs also located in the predicted guide regions, consistent with the Ψ-seq results (see Figure 29 and 30 ).

[0303] In addition, the human telomerase RNA component (TERC) shares similar features with snoRNAs, as it contains a conserved box H / ACA scaRNA domain at the 3 '-end. Upon BACS treatment, 7 Ψ sites in TERC from HeLa cells, of which 4 are previously discovered putative Ψ sites by CMC-based primer extension method (see Figure 23 and 31 ). In particular, all 3 new Ψ 38 / Ψ 100 / Ψ 155 sites, together with the known Ψ 161 and Ψ 179 sites, were found in the core domain of TERC. This observation suggests a potential involvement of Ψ in stabilizing the TERC structure, similar to what has been demonstrated for Ψ 306 and Ψ 307 in the P6.1 loop of TERC.

[0304] Example 7: Comprehensive Ψ profile of human tRNAs

[0305] Ψ is one of the most fundamental and ubiquitous modifications in human tRNA. However, given that most tRNA species are extensively modified and highly structured, quantitative analysis of Ψ in tRNA remains challenging for CMC- or BS-based methods. BACS offers a better solution to this problem because BACS-induced mutational signaling is not significantly affected by RT blockade or other intrinsic mismatches. Applying BACS to tRNA from HeLa cells, 625 high-confidence Ψ sites in cytoplasmic tRNA (cy-tRNA) were successfully detected (see...). Figure 32 The number of Ψ sites identified by each cy-tRNA varies across different isoreceptor families (see [link to relevant documentation]). Figure 33 In cy-tRNAs, the Ψ site is primarily located in highly conserved positions, including positions 13, 27–28, 38–40, and 55, while Ψ sites at other positions are limited to specific types of cy-tRNAs (see [link to relevant documentation]). Figure 34 Then, based on the standardized tRNA numbering system, a comprehensive view of the Ψ spectrum of human cy-tRNAs was summarized (see...). Figure 35 Subsequently, the Ψ modification level at each tRNA site was compared, providing valuable insights into the nature of the corresponding PUS enzyme (see [link to article]). Figure 36 Notably, position 55 became the most frequently and highly modified Ψ site in cy-tRNA, primarily installed by TRUB1. Furthermore, position 13, designated as the PUS7 target, also showed high levels of Ψ modification. Conversely, the modification levels of the PUS1 target (positions 27-28) and the PUS3 target (positions 38-40) exhibited considerable differences. Further identification of the PUS enzymes responsible for other sites is important for a full understanding of the diversity and specificity of Ψ modification patterns in human cy-tRNA.

[0306] Using BACS, 50 Ψ sites were also detected in human mitochondrial tRNA (mt-tRNA), which is highly consistent with the published dataset (see BACS). Figure 32 , 33 And 37). Three reported Ψ sites (including mt-tRNA) Ala Ψ in 38 mt-tRNA Met Ψ in 55 and mt-tRNA Pro Ψ in 38 Due to its low modification level (< 5%, see...) Figure 38 However, it was not identified as a high-confidence site. Although BACS clearly showed higher resolution than CMC and BS-based tRNA profiling methods, the mt-tRNA detected by BACS was not classified as a high-confidence site. Leu(CNN) Ψ in 20 and mt-tRNAAsn Ψ in 25 They are not located in the same place and further verification is still needed (see...). Figure 35 Overall, human mt-tRNA exhibits lower levels of pseudouridine acidification compared to cy-tRNA (see [link to relevant documentation]). Figure 34 , 37 38 and 39).

[0307] Similar to RBS-seq, BACS will also induce N1-methyladenosine (m 1 A) to N 6 -Methyladenosine (m 6 A) Dimroth rearrangement, therefore m may be detected. 1 A and Ψ (see supplement) Figure 6 a). As expected, m was observed at tRNA positions 58 (for cy-tRNA) and 9 (for mt-tRNA) after BACS treatment. 1 A significant reduction in A mutation signal, comparable to the efficiency of demethylases (see...). Figure 41 and 42 These results suggest that BACS may be a powerful tool for simultaneously studying multiple modifications in tRNA.

[0308] Example 8: Profiling and quantification of Ψ in HeLa mRNA

[0309] After successfully applying BACS to various types of ncRNAs, it was used to map and quantify Ψ modifications in HeLa mRNA. Given the high abundance of Ψ in rRNA, snRNA, snoRNA, and tRNA, the proportion of reads mapped to these ncRNAs in polyadenylated tail RNA samples was detected to assess enrichment efficiency. Only a small fraction of reads (3.7%) were mapped to these ncRNAs (see [link to BACS]). Figure 43 Using the remaining reads, a total of 1381 Ψ sites were mapped in the HeLa polyadenylated tail RNA (see...). Figure 44 Most of these Ψ sites showed low modification levels (< 20%), while only a limited number of Ψ sites showed high modification levels (> 50%) (see [link to relevant documentation]). Figure 44 and 45 The study revealed representative highly modified Ψ sites in DKC1, as evidenced by high U-to-C mutation signals in the BACS library and low background in the control library (see [link to study]). Figure 46 Compared to the aforementioned ncRNAs, the Ψ modification level in polyadenylated tail RNAs was significantly lower (see...). Figure 47 Of the 1381 Ψ sites, 1167 and 214 are located in mRNA and ncRNA (excluding rRNA, snRNA, snoRNA, and tRNA), respectively (see [link to relevant documentation]).Figure 48 ). In mRNA, Ψ is enriched in the coding sequence (CDS) and 3'-untranslated region (3'-UTR), but relatively depleted in the 5'-untranslated region (5'-UTR), consistent with previous findings (see Figure 48 and 49 ). Gene ontology (GO) analysis shows that Ψ-modified mRNAs are enriched in functions such as translation and regulation of apoptotic process (see Figure 50 ). Importantly, BACS can provide mRNA expression levels while mapping Ψ, which shows strong correlation with control library (Pearson r = 1.00) and BID-seq input library (Pearson r = 0.95-0.96), indicating minimal RNA degradation induced by BACS (see Figure 51 and 52 ).

[0310] Next, the sequence context of Ψ in HeLa mRNA was analyzed. First, the analysis shows that most Ψ sites (59.6%) are located in consecutive uridine sequences (see Figure 53 ). These positions cannot be determined precisely by BS-based methods, which further highlights the advantage of BACS. With the benefit of high-resolution signals from BACS, it is found that Ψ is mainly enriched in USΨAG (S = C or G) and GUΨCN (N = A, C, G or U) motifs, which correspond to previously identified PUS7 and TRUB1 motifs, respectively (see Figure 54 ). In addition, it is also observed that Ψ tends to be enriched in motifs containing multiple consecutive uridines, such as CUΨUG, ACΨUU, and even UUΨUU. The stoichiometry of Ψ in these motifs is also compared, which demonstrates that GUΨCN exhibits relatively higher modification levels (see Figure 55 ). However, these potential TRUB1 targets are significantly less modified in mRNA than their counterparts in cy-tRNA, a similar situation is also observed for the putative PUS7 targets (see Figure 56 ). These results suggest that mRNA might not be the major substrate for these independent PUS enzymes. In addition, from the analysis of codon preference of Ψ in mRNA, it is seen that Ψ is enriched in codons containing consecutive uridines, such as UUY (Y = C or U), UUG, AUU, and GUU, which encode phenylalanine (Phe), leucine (Leu), isoleucine (Ile), and valine (Val), respectively (see Figure 57 and 58 ). In codons, Ψ is mainly located at the second position (see Figure 59 ). Single Ψ sites located at the start codon (AUG) are observed, while 2 sites are found in the stop codon (UAG). These might facilitate stop codon readthrough.

[0311] To further assess the performance of BACS, the identified mRNA sites were thoroughly compared to published datasets. First, BACS was compared to the most recent dataset (Safra, M. et al., Genome Res. 27, 393-406 (2017)) that integrated three CMC-based methods. Notably, BACS accurately identified 63 out of 70 (90.0%) of the Ψ sites listed in the “highest confidence” category (see Figure 60 ). However, there was a substantial overlap between BACS and the “high confidence” list (186 out of 321, 57.9%) only when considering Ψ sites consistently detected across multiple samples (> 8) (see Figure 61 and 62 ). BACS was further compared to two recently developed BS-based methods. In comparison to CMC-based methods, BACS demonstrated a better overlap with BID-seq results as expected (236 out of 575, 41.0%) (see Figure 63 ). Most of the sites specific to the BID-seq dataset in our BACS library showed low levels of modification (see Figure 64 ). In comparison to PRAISE, 626 out of 1995 (31.4%) Ψ sites showed overlap with BACS results (see Figure 65 ). Similarly, most of the PRAISE-only Ψ sites in our dataset were lightly modified (see Figure 66 ). Regarding the 6 Ψ sites identified by BACS in mitochondrial mRNA (mt-mRNA), 4, 2, and 3 of them were also detected by pseudouridine-seq, BID-seq, and PRAISE, respectively. Potentially, the extent of overlap between different methods can be affected by differences in sequencing depth and different bioinformatics pipelines used for analysis (e.g., mapping to the genome or mapping directly to the transcriptome).

[0312] Throughout the description and claims of this specification, the words “comprise” and “contain” and variations of the words, for example “comprising” and “comprises”, means “including but not limited to”, and they are not intended to (and they do not) exclude other moieties, additives, components, integers or steps. Throughout the description and claims of this specification, the singular encompasses the plural unless the context requires otherwise. Specifically, unless explicitly stated otherwise, the description should be understood to accommodate plurals as well as singles.

[0313] Features, integers, characteristics, compounds, chemical moieties or groups described in conjunction with a particular aspect, embodiment or example of the application are to be understood to be applicable to any other aspect, embodiment or example described herein unless incompatible therewith. All of the features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, can be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. The application is not restricted to the details of any foregoing embodiments. The application extends to any novel one, or any novel combination, of the features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel one, or any novel combination, of the steps of any method or process so disclosed.

[0314] The reader's attention is directed to all papers and documents submitted herewith or concurrently with the present application in connection with this application as part of the disclosure, and the contents of all such papers and documents are incorporated herein by reference.

Claims

1. A method of modifying pseudouridine comprising reacting pseudouridine with a Michael addition acceptor, wherein the Michael addition acceptor is a compound according to Formula (I): (I) wherein: X is independently selected from the group consisting of halogen, tosyl, mesyl and triflate; R 1 is independently selected from H, C1-C4-alkyl, C1-C4-haloalkyl and C0-C4-alkylene-R 1a wherein R 1a is independently selected from C3-C6-cycloalkyl, phenyl and 4- to 6-membered heterocyclyl; wherein when R 1a is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 1a is optionally substituted with 1 to 4 R 1b and when R 1a is phenyl, R 1a is optionally substituted with 1 to 4 R 1c ; R 2 is independently selected from H, C1-C4-alkyl, C1-C4-haloalkyl and C0-C4-alkylene-R 2a wherein R 2a is independently selected from C3-C6-cycloalkyl, phenyl and 4- to 6-membered heterocyclyl; wherein when R 2a is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 2a is optionally substituted with 1 to 4 R 2b and when R 2a is phenyl, R 2a is optionally substituted with 1 to 4 R 2c ; R 3 independently selected from -C(O)OR 3a , -C(O)R 3a , -C(O)NR 3b R 3b , -CN, -NO2, -S(O)2OR 3a , -S(O)2R 3a , -S(O)2NR 3b R 3b , and 5- to 10-membered heteroaryl, wherein when R 3 is 5- to 10-membered heteroaryl, R 3 is optionally substituted with 1 to 4 R 3e ; R 3a is independently selected from H, C1-C6-alkyl, C1-C6-haloalkyl and C0-C6-alkylene-R 3c ; wherein when R 3a is alkyl or haloalkyl, R 3a is optionally substituted with a group selected from -N3, -CºCH, dibenzo cyclooctyne alcohol (DIBO), azido-dibenzo cyclooctyne (DBCO), bicyclononyne (BCN), trans-cyclooctene (TCO) and tetrazine; R 3b is independently selected from H, C1-C6-alkyl, C1-C6-haloalkyl and C0-C6-alkylene-R 3c ; wherein when R 3b is alkyl or haloalkyl, R 3b is optionally substituted with a group selected from -N3, -CºCH, dibenzo cyclooctyne alcohol (DIBO), azido-dibenzo cyclooctyne (DBCO), bicyclononyne (BCN), trans-cyclooctene (TCO) and tetrazine; or wherein two R 3b groups together with the nitrogen atom to which they are attached form a 5- or 6- membered heterocycloalkyl group, which is optionally substituted with 1 to 4 R 3d substituents; R 3c is independently selected from C3-C6cycloalkyl, phenyl, and 4- to 6- membered heterocyclyl; wherein when R 3c is C3-C6cycloalkyl or 4- to 6- membered heterocyclyl, R 3c is optionally substituted with 1 to 4 R 3d , and when R 3c is phenyl, R 3c is optionally substituted with 1 to 4 R 3e ; R 1b and R 2b are each, independently at each occurrence, selected from the group consisting of =0, =S, halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl; R 1c and R 2c are each, independently at each occurrence, selected from the group consisting of halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl; R 3d independently at each occurrence selected from =0, =S, halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl, C1-C4-haloalkyl, -N3, -CºCH, dibenzo cyclooctyne alcohol (DIBO), azido-dibenzo cyclooctyne (DBCO), bicyclononyne (BCN), trans-cyclooctene (TCO), and tetrazine; R 3e independently at each occurrence selected from halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl, -N3, -CºCH, dibenzo cyclooctyne alcohol (DIBO), azido-dibenzo cyclooctyne (DBCO), bicyclononyne (BCN), trans-cyclooctene (TCO) and tetrazine; R 4 independently at each occurrence selected from H and C1-C4-alkyl; or when two R 4 groups are attached to the same nitrogen, the two R 4 groups together with the nitrogen atom to which they are attached optionally form a 5- to 6- membered heterocycloalkyl group, optionally substituted with 1 to 4 R 5 ; and R 5 are each independently at each occurrence selected from the group consisting of =0, =S, halogen, nitro, cyano, Ci-C4-alkyl and Ci-C4-haloalkyl.

2. A method of tagging or labelling pseudouridine in a sample comprising reacting at least a portion of the sample with a Michael addition acceptor, wherein the Michael addition acceptor is a compound according to Formula (I): (I) wherein: X is independently selected from the group consisting of halogen, tosyl, mesyl and triflate; R 1 is independently selected from H, C1-C4-alkyl, C1-C4-haloalkyl and C0-C4-alkylene-R 1a wherein R 1a is independently selected from C3-C6-cycloalkyl, phenyl and 4- to 6-membered heterocyclyl; wherein when R 1a is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 1a is optionally substituted with 1 to 4 R 1b and when R 1a is phenyl, R 1a is optionally substituted with 1 to 4 R 1c ; R 2 is independently selected from H, C1-C4-alkyl, C1-C4-haloalkyl and C0-C4-alkylene-R 2a wherein R 2a is independently selected from C3-C6-cycloalkyl, phenyl and 4- to 6-membered heterocyclyl; wherein when R 2a is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 2a is optionally substituted with 1 to 4 R 2b and when R 2a is phenyl, R 2a is optionally substituted with 1 to 4 R 2c ; R 3 independently selected from -C(O)OR 3f , -C(O)R 3f , -C(O)NR 3g R 3f , -S(O)2OR 3f , -S(O)2R 3f , -S(O)2NR 3g R 3f ; R 3f is a linker covalently attached to an affinity tag or imaging probe; R 3g independently selected from H, Ci-C6-alkyl and Ci-C6-haloalkyl; R 1b and R 2b are each independently at each occurrence selected from =0, =S, halogen, nitro, cyano, C(0)OR 4 , C(0)R 4 , C(0)NR 4 R 4 , Ci-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and Ci-C4-haloalkyl; R 1c and R 2c each, independently from each other in each occurrence, is selected from the group consisting of halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl; R 4 independently at each occurrence selected from H and C1-C4-alkyl; or when two R 4 groups are attached to the same nitrogen, the two R 4 groups together with the nitrogen atom to which they are attached optionally form a 5- to 6- membered heterocycloalkyl group, optionally substituted with 1 to 4 R 5 ; and R 5 are each, independently at each occurrence, selected from the group consisting of =0, =S, halogen, nitro, cyano, Ci-C4-alkyl and Ci-C4-haloalkyl.

3. The method of claim 2, wherein the affinity tag is selected from the group comprising biotin, FLAG tag, His tag, HA tag, Strep tag, Avi tag, GST, c-myc tag, V5 tag, E tag, S tag, SBP tag, poly(Glu) tag, calmodulin tag.

4. The method of claim 2 or 3, wherein the linker is a flexible linker, a cleavable linker; optionally a photocleavable linker.

5. The method of claim 4, wherein the linker is selected from the group consisting of (a) polyethylene glycol; (b) a peptide; (c) a nucleic acid; or (d) an oligosaccharide.

6. The method of claim 2, wherein the imaging probe is selected from the group comprising a fluorescent moiety, a radionuclide and a metal complex.

7. The method of claim 6, wherein the imaging probe is (a) a fluorophore selected from the group comprising fluorescein, rhodamine, BIODIPY, Alexa fluor, Cy dye or ATTO dye; (b) a lanthanide complex; or (c) a radionuclide complex.

8. A method of isolating RNA comprising pseudouridine from a sample comprising reacting the sample with a Michael addition acceptor according to the method of any one of claims 2 to 5, and then contacting the sample with a matrix comprising a binding partner for the affinity tag.

9. The method of claim 8, wherein the affinity tag is biotin and the binding partner is avidin or streptavidin.

10. A method of visualising pseudouridine in a sample comprising RNA comprising reacting the sample with a Michael addition acceptor according to the method of any one of claims 2, 6 or 7, and subjecting the sample to a visualisation procedure selected from optical observation, microscopic observation and image capture.

11. A method of determining the presence and sequence position of pseudouridine comprised in a sample of RNA comprising: (a) reacting at least a portion of the sample with a Michael addition acceptor, wherein the Michael addition acceptor is a compound according to Formula (I): (I) wherein: X is independently selected from the group consisting of halogen, tosyl, mesyl and triflate; R 1 is independently selected from H, C1-C4-alkyl, C1-C4-haloalkyl and C0-C4-alkylene-R 1a wherein R 1a is independently selected from C3-C6-cycloalkyl, phenyl and 4- to 6-membered heterocyclyl; wherein when R 1a is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 1a is optionally substituted with 1 to 4 R 1b and when R 1a is phenyl, R 1a is optionally substituted with 1 to 4 R 1c ; R 2 is independently selected from H, C1-C4-alkyl, C1-C4-haloalkyl, C0-C4-alkylene-R 2a wherein R 2a is independently selected from C3-C6-cycloalkyl, phenyl and 4- to 6-membered heterocyclyl; wherein when R 2a is C3-C6-cycloalkyl or 4- to 6-membered heterocyclyl, R 2a is optionally substituted with 1 to 4 R 2b and when R 2a is phenyl, R 2a is optionally substituted with 1 to 4 R 2c ; R 3 independently selected from -C(O)OR 3a , -C(O)R 3a , -C(O)NR 3b R 3b , -CN, -NO2, -S(O)2OR 3a , -S(O)2R 3a , -S(O)2NR 3b R 3b , and 5- to 10-membered heteroaryl, wherein when R 3 is 5- to 10-membered heteroaryl, R 3 is optionally substituted with 1 to 4 R 3e ; R 3a independently selected from the group consisting of H, C1-C6-alkyl, C1-C6-haloalkyl and C0-C6-alkylene-R 3c ; R 3b independently selected from the group consisting of H, C1-C6-alkyl, C1-C6-haloalkyl and C0-C6-alkylene-R 3c ; or two R 3b groups together with the nitrogen atom to which they are attached form a 5- or 6- membered heterocycloalkyl group, which is optionally substituted with 1 to 4 R 3d groups; R 3c is independently selected from C3-C6cycloalkyl, phenyl, and 4- to 6- membered heterocyclyl; wherein when R 3c is C3-C6cycloalkyl or 4- to 6- membered heterocyclyl, R 3c is optionally substituted with 1 to 4 R 3d , and when R 3c is phenyl, R 3c is optionally substituted with 1 to 4 R 3e ; R 1b , R 2b and R 3d are each independently at each occurrence selected from the group consisting of =0, =S, halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl; R 1c , R 2c and R 3e are each independently at each occurrence selected from the group consisting of halogen, nitro, cyano, C(O)OR 4 , C(O)R 4 , C(O)NR 4 R 4 , C1-C4-alkyl, C2-C4-alkenyl, C2-C4-alkynyl and C1-C4-haloalkyl; R 4 independently at each occurrence selected from H and C1-C4-alkyl; or when two R 4 groups are attached to the same nitrogen, the two R 4 groups together with the nitrogen atom to which they are attached optionally form a 5- to 6- membered heterocycloalkyl group, optionally substituted with 1 to 4 R 5 ; and R 5 are each independently at each occurrence selected from the group consisting of =0, =S, halogen, nitro, cyano, Ci-C4-alkyl and Ci-C4-haloalkyl; (b) (i) sequencing the RNA; or (ii) reverse transcribing the RNA produced in step (a) to provide cDNA and amplifying the cDNA; (c) sequencing the DNA of step (b)(ii); (d) comparing the DNA sequence of step (c) to a reference DNA sequence to identify the sequence position of guanine (G) in the sequence which is adenine (A) in the reference sequence, the position of G in the DNA sequence being the position of pseudouridine in the corresponding RNA sequence.

12. The method of claim 11, wherein the reference sequence can be obtained from a separate portion of the RNA sample, which portion is not subjected to the Michael addition reaction of step (a), but is sequenced according to step (b)(i); or is reverse transcribed according to step (b)(ii) and sequenced according to step (c).

13. The method of claim 11 or 12, wherein amplification of the cDNA employs an isothermal amplification method; wherein the isothermal amplification method is selected from the group consisting of polymerase chain reaction (PCR), strand displacement amplification (SDA), rolling circle amplification (RCA), whole genome amplification (WGA), loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA), and multiple displacement amplification (MDA); optionally, wherein one-step RT-PCT is used.

14. The method of any one of claims 11 to 13, wherein reverse transcription of step (b) uses a reverse transcriptase; the reverse transcriptase is optionally selected from the group consisting of: Maxima H-, SuperScript II, SuperScript III, SuperScrupt IV, ProtoScript II, SMARTScribe, PrimeScript II, HiScript III, MMLV, AMV, TGIRT-III, or recombinant HIV; preferably Maxima H-, SuperScript IV, and TGIRT-III; more preferably Maxima H-.

15. The method of any one of the preceding claims, wherein X is selected from the group consisting of Br, CI, and I.

16. The method of any one of the preceding claims, wherein at least one of R 1 and R 2 is H; optionally, wherein R 1 and R 2 are both H.

17. The method of any one of the preceding claims, wherein R 3 is independently selected from -C(O)OR 3a , -C(O)R 3a , -C(O)NR 3b R 3b , and -CN.

18. The method of any one of the preceding claims, wherein R 3 is -C(O)NR 3b R 3b , optionally wherein R 3 is -C(O)NH2.

19. The method of any one of claims 1 to 18, wherein the compound according to formula (I) is selected from the group consisting of: 、 、 、 、 、 、 、 and .

20. The method of any one of the preceding claims, wherein the Michael addition reaction is performed at a pH in the range of about 7.0 to about 9.

5.

21. The method of any one of the preceding claims, wherein the Michael addition acceptor is present at a concentration in the range of about 10 mM to about 2 M; preferably in the range of from about 100 mM to about 500 mM; more preferably about 250 mM.

22. The method of any one of the preceding claims, wherein the Michael addition reaction is performed at a temperature in the range of about 25 °C to about 95 °C; preferably about 65 °C to about 95 °C; more preferably about 85 °C.

23. The method of any one of the preceding claims, wherein the Michael addition reaction is performed at a temperature of about 70 °C or greater for a time in the range of about 5 minutes to about 2 hours; or at a temperature of about 70 °C or less for a time in the range of about 2 hours to about 16 hours.

24. The method of any one of the preceding claims, wherein the RNA molecule is selected from one or more of mRNA, tRNA, rRNA, snRNA, miRNA, IncRNA, or circRNA.

25. The method of any one of the preceding claims, wherein the RNA molecule is from a biological sample.

26. A kit for modifying pseudouridines, comprising: (a) a solution comprising the Michael acceptor of any one of claims 1 to 7 or 15 to 19; (b) instructions for reacting a sample comprising pseudouridines with the solution.

27. A kit for tagging or labeling pseudouridines comprised in an RNA, comprising: (a) a solution comprising the Michael acceptor of any one of claims 2 or 15 to 19; (b) instructions for reacting an RNA with the solution.

28. A kit for determining the presence of a sequence position of pseudouridines comprised in an RNA, comprising: (a) a solution comprising the Michael acceptor of any one of claims 11 or 15 to 19; (b) instructions for reacting an RNA with the solution.

29. The kit of any one of claims 26 to 28, further comprising one or more buffers.

30. The kit of claim 28 or 29, further comprising a reverse transcriptase; optionally, a reverse transcriptase selected from Maxima H-, SuperScript II, SuperScript III, SuperScrupt IV, ProtoScript II, SMARTScribe, PrimeScript II, HiScript III, MMLV, AMV, TGIRT-III, recombinant HIV, Marathon reverse transcriptase, or Induro reverse transcriptase; preferably Maxima H-, SuperScript IV, and TGIRT-III; more preferably Maxima H-.

31. The kit of any one of claims 28 to 30, further comprising a DNA polymerase; optionally, a DNA polymerase selected from Taq DNA polymerase, Bst DNA polymerase, or Bsu DNA polymerase.

Citation Information

Patent Citations

  • Compositions and methods related to modification and detection of pseudouridine and 5-hydroxymethylcytosine

    WO2022232795A1