Methods for modifying and identifying nucleic acids
Nucleotide analog derivatization chemistry facilitates the detection of nucleic acid modifications with single-nucleotide resolution, addressing the inefficiencies of current RNA analysis methods by enabling scalable and cost-effective whole-transcriptome analysis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-08-29
- Publication Date
- 2026-03-16
AI Technical Summary
Current methods for analyzing RNA dynamics in cells are time-consuming, expensive, and impractical due to small signal-to-noise ratios and off-target reactivity, especially when dealing with short RNA species, and they cannot analyze labeled RNA in the context of whole RNA.
A method for nucleotide analog derivatization chemistry that allows for the detection of modified nucleic acids with single-nucleotide resolution, enabling scalable, quantifiable, and time-efficient whole-transcriptome analysis of polynucleotide modifications by altering the base-pairing ability of nucleic acid bases through hydrogen bonding partner addition or removal.
Enables rapid and automated detection of nucleic acid modifications without the need for purification, providing a cost-effective and efficient method for high-throughput sequencing of RNA dynamics within cells.
Smart Images

Figure 0007830393000004 
Figure 0007830393000005 
Figure 0007830393000006
Abstract
Description
[Technical Field]
[0001] This invention relates to the field of nucleic acid processing and sequencing. [Background technology]
[0002] 4-thiouridine (s) is a nucleotide analog. 4 U) and 6-thioguanosine (s 6 G), for example, is readily incorporated into nascent RNA by natural enzymes (Tani et al., Genome Res. Vol. 22, pp. 947-956 (2012)). Well-known analogs include 5-bromouridine (5BrU), 5-ethinyluridine (5-EU), and 6-thioguanosine (s 6 G), 4-thiouridine (s 4 There are 4-thiouridines (s), which are readily incorporated by cells and further provide unique physiological and chemical properties for antibody detection, cyclization reactions, and thiol-specific reactivity and affinity, respectively (Eidinoff et al., Science. Vol. 129, pp. 1550-1551 (1959); Jao et al., PNAS Vol. 105, pp. 15779-15784 (2008); Melvin et al., Eur. J. Biochem. Vol. 92, pp. 373-379 (1978); Woodford et al., Anal. Biochem. Vol. 171, pp. 166-172 (1988); Dolken et al., RNA Vol. 14, pp. 1959-1972 (2008); Rabani et al., Nat Biotechnol. Vol. 29, pp. 436-442 (2011)). 4 U) is the most widely used nucleotide for studying RNA expression dynamics. 4 Like other nucleotides, U is rapidly taken up by cells without the need for electroporation or lipofection. In cells, phosphorylated by cellular uridine kinases, s 4A pool of accumulated U is generated, which is effectively incorporated into newly synthesized RNA in various types of cells (including fly, mouse, and human cells) (Dolken, 2008, supra). Further, in flies and mice, 4-thiouridine can be used in combination with the cell-type specific expression of uracil phosphoribosyltransferase (UPRT) of Toxoplasma gondii to label in vivo transcripts with a cell-type specific label. This UPRT couples ribose 5-phosphate to the N1 nitrogen of uracil (or 4-thiouridine) to generate (4-thio)uridine monophosphate, which is incorporated into RNA (Cleary et al., Nat Biotechnol. Vol. 23, pp. 232-237 (2005)). The current protocols that characterize the dynamics of RNA biosynthesis, processing, and metabolic turnover in cells using 4-thiouridine (s 4 U) metabolic RNA labeling utilize biochemical separation through reversible biotinylation of the thiol group in s 4 U [e.g., through N-[6-(biotinamido)hexyl]-3'-(2'-pyridyldithio)propionamide (HPDP-biotin) or methanethiosulfonate (MTS-biotin) coupled with biotin] (Cleary et al., 2005, supra). However, like all biochemical separation methods, the underlying protocols are time-consuming and typically encounter the problem of a small signal-to-noise ratio due to biotinylation efficiency (especially when applied to short RNA species) and off-target reactivity (Duffy et al., Mol Cell. Vol. 59, pp. 858-866 (2015); Neymotin et al., RNA Vol. 20:1645-1652 (2014)).
[0003] WO 2006 / 125808 A1 describes a method for analyzing newly transcribed RNA containing thiolated RNA based on a microarray.
[0004] WO 2004 / 101825 A1 and WO 2016 / 154040 A2 concern methods for separating RNA by biosynthesis labeling.
[0005] Miller et al., Nature Methods Vol. 6(6), 2009: pp. 439-441, describe the labeling of RNA in Drosophila via a 4-thioluracil food source.
[0006] Schwalb et al., Science, Vol. 352 (6290), 2016: pp. 1225-1228, describes a transient transcriptome sequencing method that can evaluate the synthesis and degradation of total mRNA.
[0007] Hartmann et al. (Handbook of RNA Biochemistry, Vol. 2, 2014, Chapter 8.3.3, pp. 164-166) describe post-synthesis labeling of 4-thiouridine-modified RNA by modifying the 4-thiouridine residue with iodoacetamide or a sulfur-based compound.
[0008] Testa et al., Biochemistry, Vol. 38 (50), 1999: pp. 16655-16662 contains information on thiouracil (s 2 U and s 4 It has been disclosed that the base-pairing ability of U) is different compared to uracil.
[0009] Hara et al., Biochemical and Biophysical Research Communications, Vol. 38(2), 1970: pp. 305-311, discloses 4-thiouridine-specific spin labeling of tRNA.
[0010] Fuchs et al., Genome Biology, Vol. 15(5), 2014: pp. 1465-6906, describes determining genome-wide transcription elongation rates by examining 4-thiouridine tags on RNA. This method requires biotinylation and purification of the thus labeled RNA.
[0011] Furthermore, in reversible biotinylation strategies, labeled RNA can only be analyzed in isolation; that is, it cannot be analyzed in the context of whole RNA. Therefore, accurately measuring RNA dynamics within cells by high-throughput sequencing requires analyzing three RNA subsets (labeled RNA, whole RNA, and unlabeled RNA) at each time point, making such approaches expensive and impractical for downstream analysis. [Overview of the project] [Problems that the invention aims to solve]
[0012] Therefore, one objective of the present invention is to simplify the method for detecting modified nucleic acids to the extent that automated detection is possible. [Means for solving the problem]
[0013] This invention is based on nucleotide analog derivatization chemistry that enables the detection of modifications within polynucleotide (PNA) species with single-nucleotide resolution. The method of this invention provides a scalable, highly quantifiable, cost-effective, and time-efficient method for rapid, whole-transcriptome analysis of PNA modifications.
[0014] In the first aspect, the present invention provides a method for identifying polynucleic acid (PNA), comprising the steps of: preparing PNA; modifying one or more nucleic acid bases of the PNA by adding or removing hydrogen bonding partners to change the base-pairing ability of one or more nucleic acid bases; forming base pairs with complementary nucleic acid with the PNA, including base pairing with at least one modified nucleic acid base; and identifying the sequence of the complementary nucleic acid at a position complementary to at least one modified nucleic acid base.
[0015] In preferred embodiments, PNA is synthesized intracellularly, in particular, already having modifications that alter its base-pairing ability, or can be further modified to alter its base-pairing ability. Thus, the present invention can also be defined as a method for identifying polynucleic acid (PNA), the method comprising: expressing PNA in a cell; isolating the PNA from a cell; modifying one or more nucleic acid bases of the PNA in the cell and / or after isolation, thereby altering the base-pairing ability of the one or more nucleic acid bases by adding or removing hydrogen bonding partners of the one or more nucleic acid bases through the modification in the cell, or after isolation, or both in the cell and after isolation; base-pairing a complementary nucleic acid with the PNA, the base-pairing of which includes base-pairing with at least one modified nucleic acid base; and identifying the sequence of the complementary nucleic acid at at least one position complementary to the modified at least one nucleic acid base.
[0016] The present invention further provides a kit for carrying out the method of the present invention, which includes a nucleic acid base modified with a thiol in particular, and an alkylating agent suitable for alkylating the thiol-modified nucleic acid base at the position of the thiol group, wherein the alkylating agent includes a hydrogen bond donor or a hydrogen bond acceptor.
[0017] All embodiments of the present invention are described collectively in the following detailed description, and all preferred embodiments relate equally to all embodiments, aspects, methods, and kits. For example, a kit or its components may be used in the methods of the present invention, or may be suitable for use in the methods of the present invention. Any component used in the described methods may be included in a kit. The preferred and detailed descriptions of the methods of the present invention should be read similarly to the components of a kit for a given step of the method, or their suitability, or combinations of components of the kit. All embodiments can be combined with one another, unless otherwise stated. [Modes for carrying out the invention]
[0018] This invention relates to a method for producing synthetic PNA (also called modified PNA) by modifying polynucleic acid (abbreviated as PNA). The presence of synthetic PNA in a PNA sample can be detected in the PNA sequencing read results of that sample, thereby identifying the modified PNA. One advantage of this invention is that this identification can be performed without purification / separation from unmodified PNA.
[0019] More specifically, the method of the present invention includes the steps of: modifying one or more nucleic acid bases of PNA by adding or removing hydrogen bonding partners to change the base-pairing ability (or behavior) of the one or more nucleic acid bases; and causing a complementary nucleic acid to form a base pair with the PNA, the base pairing of which includes base pairing with at least one modified nucleic acid base.
[0020] The natural nucleic acid bases are A (adenine), G (guanine), C (cytosine), T (thymine) / U (uracil). The modifications of this invention result in non-natural nucleic acid bases compared to A, G, C, and U nucleotides in RNA, and non-natural nucleic acid bases compared to A, G, C, and T nucleotides in DNA. This modification alters the base-pairing behavior, changing the preferential base-pairing (hydrogen bonding) between A and T / U, and between C and G. This means that base-pairing with a complementary nucleic acid changes from one natural nucleic acid to another. Preferably, the complementary nucleic acid is DNA, and T is used instead of U. The modified A can bind to C or G; the modified T or U can bind to A or T / U; the modified C can bind to A or T / U; and the modified G can bind to A or T / U. Such modifications are known in the art. Modifications are usually minor, and the changes are kept to a minimum so that only the behavior of base pairing is altered. For example, A and G maintain the purine ring system to which they belong, and C and T / U maintain the pyrimidine ring to which they belong. For example, Harcourt et al. (Nature 2017, Vol. 541: pp. 339-346) present an overview and summary of such modifications. An example of modification is from A to m 6 To A, m 1 Modifications to A, inosine, and 2-aminoadenine; from C to m 5 To C(5-methylcytosine), hm 5 Modifications to C(5-hydroxymethylcytosine), pseudouridine, 2-thiocytosine, 5-halocytosine, 5-propynyl(-C=C-CH3)cytosine, and 5-alkynylcytosine; from T or U to 2-thiouracil, s 4Modifications include: U (4-thiouracil) to 2-thiothymine, 4-pyrimidinone, pseudouracil, 5-halouracil, e.g., 5-bromouracil (also called 5-bromouridine (5BrU)), 5-propynyl(-C=C-CH3)uracil, 5-alkynyluracil, e.g., 5-ethinyluracil; modification of G to hypoxanthine, xanthine, and isoguanine; and modification of A or G to 6-methyl derivatives of adenine and guanine and other 6-alkyl derivatives, and to 2-propyl derivatives of adenine and guanine and other 2-alkyl derivatives. Further modifications include to 6-azo-uracil, 6-azo-cytosine, 6-azo-thymine, 8-halo-adenine and 8-halo-guanine, 8-amino-adenine and 8-amino-guanine, 8-thiol-adenine and 8-thiol-guanine, 8-thioalkyl-adenine and 8-thioalkylguanine, 8-hydroxyl-adenine and 8-hydroxyl-guanine, other 8-substituted adenines and guanines, and 5-halo (especially 5-halo-guanine). Modifications include those to rom-uracil and 5-halo(especially 5-bromo)-cytosine, 5-trifluoromethyl-uracil and 5-trifluoromethyl-cytosine, other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine, 8-azaadenine, 7-deazaguanine, 7-deazaadenine, 3-deazaguanine, and 3-deazaadenine. While it is preferable for natural nucleic acid bases to be modified to the closest modified nucleic acid base as described above, in principle, nucleic acid bases can be modified to any of the above modified nucleic acid bases. The important factor is the change in the hydrogen bonding pattern, which results in modified nucleic acid bases having different base-pairing partners compared to unmodified nucleic acid bases. The change in the binding partner does not need to be absolutely certain; it is sufficient that the certainty of binding to the natural binding partner changes by, for example, at least 10%, or at least 20%, or at least 30%, or at least 40%, or at least 50%, or at least 60%, or at least 70%, or at least 80%, or at least 90%, or 100%.Special nucleic acid bases can bind to two or more complementary nucleic acid bases (especially fluctuating bases). Reference conditions for determining the changes are the standard conditions for reverse transcriptase in isotonic physiological aqueous solutions, preferably at atmospheric pressure and 37°C. Examples of conditions include 50 mM Tris-HCl, 75 mM KCl, 3 mM MgCl2, and 10 mM DTT at pH 7.5–8.5. Any such changes can be monitored by current detection methods (sequencing, sequence comparison, etc.). Furthermore, two or more modifications can be contained within a single PNA molecule, and only one modification per molecule or one modification per molecule needs to be detected. Of course, the greater the rate at which hydrogen bonds change from one native nucleic acid base to another (within complementary nucleic acids), the greater the certainty of detection. Therefore, a large change in the rate of base pairing is preferable (e.g., at least 50%, at least 80%).
[0021] "Halo" means a halogen, particularly one of F, Cl, Br, or I, with Br being particularly preferred, as it is found in 5BrU, for example. "Alkyl" means an alkyl residue, preferably one with a length of C1 to C2. 12 These are branched or unbranched, substituted or unsubstituted alkyl residues. Preferred are alkyl residues of length C1-C4 with optional O substituents and / or optional N substituents, as in acetamide, or other alkylcarbonyl, carboxylic acid, or amide.
[0022] Particularly preferred modified nucleic acid bases of PNA are 5-bromouridine (5BrU), 5-ethinyluridine (5-EU), and 6-thioguanosine (s 6 G), 4-thiouridine (s 4 U), 5-vinyluridine, 5-azidomethyluridine, N 6 -Allyladenosine (a 6 A) is correct.
[0023] The behavior of base pairing is known in this field or can be derived from changes in hydrogen bond donors or acceptors (including the inhibition of their pairing by interference). For example, 4-pyrimidinone (modified U or T) preferentially forms base pairs with G instead of A (Sochacka et al., Nucleic Acids Res. March 11, 2015; Vol. 43(5): pp. 2499-2512).
[0024] Modification of the nucleic acid bases of PNA is possible by replacing a hydrogen (H) on the oxygen (O) or nitrogen (N) atom with a substituent such as carbon (e.g., as in a methyl group or other alkyl group), thereby removing H as a hydrogen bond donor. Modification is also possible by replacing a free electron pair on the oxygen (O) or nitrogen (N) atom with a substituent such as carbon (e.g., as in a methyl group or other alkyl group), thereby removing the electron pair that acts as a hydrogen bond acceptor. Modification may include replacing O with sulfur (S) or SH, followed by one of the above modifications, particularly alkylation of S or SH. One preferred method for replacing O with S or SH is by biosynthesis, where an enzyme (e.g., transcriptase) is prepared to produce a nucleotide modified with S or SH (s 4 (e.g., U). Transcriptases can exist within cells.
[0025] The modification of the present invention can be a one-step modification or a modification of two or more steps (e.g., two, three, or more steps). For example, the first part of the modification may be carried out in one reaction environment (e.g., cells), and the second modification may be carried out in a different reaction environment after, for example, the PNA has been isolated from those cells. Preferably, such a second or further modification depends on the first modification and is carried out, for example, on atoms that have been altered by the first modification. Particularly preferred is a multi-step modification, in which case the first modification is an enzymatic modification, for example, a modification by incorporating a modified nucleotide / nucleic acid base into the PNA by an enzyme (e.g., RNA polymerase or DNA polymerase). This step involves only minor modifications so as not to impair or only to an acceptable degree of enzyme activity with respect to the enzyme's processing capacity. Minor modifications are, for example, changes of only one or two atoms (hydrogen not counted) compared to the corresponding native nucleic acid base. In subsequent steps, the incorporated modified nucleic acid base can be further modified by any means (such as wet chemical processes, including alkylation) to obtain the modified nucleic acid bases described herein, for example. Such further modifications can be carried out extracellularly, enzymatically, or non-enzymatically. This modification preferably targets the modification introduced in the initial step. As an intracellular (initial) modification, induced or enhanced modification is possible by supplying the modified nucleic acid base (e.g., modified nucleotide) to the cell, which then incorporates the modified nucleic acid base into the biosynthesized PNA. "Enhanced" means exceeding the rate of modification occurrence in nature.
[0026] As a (first) modification, intrinsic intracellular processes that do not supply modified nucleic acid bases to the cell are also possible. Such intrinsic processes include, for example, tRNA thiolation (Thomas et al., Eur J Biochem. 1980, Vol. 113(1): pp. 67-74; Emilsson et al., Nucleic Acids Res 1992, Vol. 20(17): pp. 4499-4505; Kramer et al., J. Bacteriol. 1988, Vol. 170(5): pp. 2344-2351). Such naturally occurring modifications can also be detected by the method of the present invention, for example, by directly detecting base mismatches with these modified nucleic acid bases or altered base pairing behavior, or by further (second) modifications of these naturally modified nucleic acid bases. Some intrinsic modifications may be the result of stress responses or other environmental influences. Therefore, such cellular responses and intracellular influences can be detected using the method of the present invention. One example is the reaction to UV light (especially near-UV irradiation), particularly within tRNA. 4 It is a U modification (Kramer et al., see above reference). In particular, s within tRNA 4 U modification can also be used to measure the rate of cell proliferation (Emilsson et al., cited above). This modification is used as an indicator of proliferation and can be detected according to the method of the present invention. Preferably, eubacteria or archaea are used for such natural modifications.
[0027] In preferred embodiments of the present invention, the modification step is carried out by incorporating a thiol-modified nucleic acid base into PNA (first part of modification) and alkylating the thiol nucleic acid base with an alkylating agent (second part of modification). Thiol-reactive alkylating agents include iodoacetamide, maleimide, benzyl halide, and bromomethyl ketone. The alkylating agent may also contain the above alkyl groups and release groups (halogens such as Br and Cl). The alkylating agent reacts to S-alkylate the thiol, thereby producing a stable thioether product. Arylation reagents (NBD (4-nitrobenzo-2-oxa-1,3-diazole) halides) react with thiols or amines, resulting in nucleophilic substitution of aromatic halides. Thiosulfates can also be used for reversible thiol modification. Thiosulfates react stoichiometrically with thiols to form mixed disulfides. Thiols also react with isothiocyanates and succinimidyl esters. Isothiocyanates and succinimidyl esters can also be used in reactions with amines.
[0028] The modification of thiols may also include the step of converting the thiol to a thioketone. The thioketone group can then be further modified by the addition or removal of a hydrogen bonding partner. Conversion to thioketones may include the removal of hydrogen on a transition metal cluster as a catalyst, as described, for example, by Kohler et al. (Angew. Chem. Int. Ed. Engl. 1996, Vol. 35(9): pp. 993-995). Conversion to thioketones allows for additional reaction chemistry options for carrying out the modifications of the present invention. Kohler et al. also describe the introduction of thiols or thioketones into aryls. This is also one option for generating thio modifications (thiols, thioketones) to the modified nucleic acid bases of the present invention.
[0029] Thiol alkylation is also called thiol (SH) bond alkylation. The advantage of thiol alkylation is that it is selective for "soft" thiols, while non-thiolated nucleic acid bases remain unchanged (HSAB theory - "hard, soft (Lewis) acids and bases", Pearson et al., JACS 1963, Vol. 85(22): pp. 3533-3539). Iodoacetamide readily reacts with any thiol to form a thioether. Iodoacetamide is somewhat more reactive than bromoacetamide, but bromoacetamide can also be used. Maleimide is an excellent reagent for thiol-selective modification, quantification, and analysis. In this reaction, the thiol is added across the double bond of maleimide to produce a thioether. Alkylation is also possible with the thioketones mentioned above.
[0030] Preferably, the modification involves alkylation at the 4-position of uridine. Interfering with the innate hydrogen bonding behavior of uridine at this position is highly effective. Such modification can be carried out using an alkylating agent, for example, an alkylating agent containing a hydrogen bonding partner (a hydrogen bonding acceptor is preferred), or an alkylating agent that does not contain a hydrogen bonding partner and therefore inhibits the hydrogen bonding that would normally occur at the 4-position of uridine. Such alkylation can be carried out in a two-step modification via 4-thiouridine as described above.
[0031] Another preferred alkylation is at the 6-position of guanosine. Such alkylation increases the rate of mispair formation from standard GC base pairs to G*A fluctuation base pairs, which have only two effective hydrogen bonds (instead of three in GC). In a particularly preferred embodiment, introducing alkylation at the 6-position of guanosine leads to 6-thioguanosine (s) from guanidine. 6 This involves modification to G) and alkylation of the thio position. Therefore, this is another preferred example of such alkylation in the two-step modification via 6-thioguanosine described above. 6-thioguanosine can be incorporated into PNA by biosynthesis in the presence of 6-thioguanosine nucleotides.
[0032] A preferred alkylating agent is the formula Hal-(C) x O y N z (Hydrogen is not shown) is present. Hereinafter, Hal means halogen, C is a branched or unbranched carbon chain consisting of x C atoms, where x is 1 to 8, O is y oxygen substituents on a C atom, where y is 0 to 3, and N is z nitrogen substituents on a C atom, where z is 0 to 3. N is preferably at least one -NH2 or double bond=NH, and O is preferably -OH or double bond=O. Hal is preferably selected from Br and I.
[0033] Particularly preferred is that the PNA contains one or more 4-thiouridine or 6-thioguanosine molecules, and more preferably two, three, four, five, six, seven, eight, nine, ten, or more 4-thiouridine or 6-thioguanosine molecules. Modification of one or more nucleic acid bases may include attaching hydrogen bonding partners (hydrogen bonding acceptors or hydrogen bonding donors) to the thiol-modified nucleic acid bases. Such attachment can be achieved, for example, by any chemical modification (alkylation is preferred) with the above-mentioned halogen-containing alkylating agents.
[0034] One alternative to alkylation is modification by oxidation. Such modifications are disclosed, for example, in Burton, Biochem J Vol. 204 (1967): p. 686 and Riml et al., Angew Chem Int Ed Engl. 2017; Vol. 56 (43): pp. 13479-13483. For example, nucleic acid bases, especially thiolated nucleic acid bases, can be modified by oxidation to become hydrogen bond donors or hydrogen bond acceptors. In the above two-step method using thiolated nucleic acid bases, the sulfur of the thiol group can be oxidized by, for example, OsO4, NaIO3, NaIO4, or by peroxides (such as chloroperoxybenzoic acid or H2O2). 4 U can be oxidized to C (Schofield et al., Nature Methods, doi:10.1038 / nMeth.4582). This changes the base pairing / hybridization behavior from UA to CG. As Burton (see above) shows, this oxidation does not require a thiol intermediate, although such an intermediate is preferable, especially in the case of biosynthetic modification (see below). Such a C analog is, for example, trifluoroethylated cytidine (e.g., the product of oxidation in the presence of 2,2,2-trifluoroethylamine). The C analog can retain the base pairing behavior of cytosine and / or the pyrimidine-2-one ring. The 4-position on the pyrimidine-2-one ring can be substituted with, for example, an amino group (in C) or other substituents (R-NH group (where R is selected from alkyl groups, aromatic groups, alkane groups, NH2, trifluoroethylene, MeO, etc.)) (see Schofield et al., in particular Additional Figure 1 of the above; incorporated herein by reference).
[0035] In a preferred embodiment, modified nucleic acid bases, such as thiol-modified bases, are incorporated into PNA through intracellular biosynthesis or by cellular enzymes (e.g., by in vitro transcription). Alternatively, modified nucleic acid bases can be chemically introduced, for example, by (biological) chemical PNA synthesis (organic or semi-synthetic synthesis). Biosynthesis is the synthesis of PNA based on template PNA (usually DNA, particularly genomic DNA) and template-dependent synthesis (transcription, reverse transcription). Suitable enzymes for such transcription include RNA polymerase, DNA polymerase, and reverse transcriptase. These enzymes can incorporate native nucleotides and modified nucleotides (containing modified nucleic acid bases) into the biosynthesized PNA molecule. Multiple nucleotide monomer units are linked together when forming PNA. Such monomers are provided in a modified form and can be incorporated into PNA. Preferably, only one type of native nucleotide (A, G, C, T / U) is modified, i.e., only one type of native nucleotide has a modified (non-natural) partner to incorporate into the PNA. It is preferable that all types of natural nucleotides are present, and that the number of modified nucleic acid bases is smaller than the corresponding natural (unmodified) nucleic acid bases. "Corresponding" means the natural nucleic acid base for which the change in atoms (hydrogen is not counted) required to reproduce the natural nucleic acid base is minimal. For example, A, G, C, T / U are provided in addition to modified U (or any other type of modified nucleotide selected from A, G, C, T). Preferably, the ratio of modified nucleotides to unmodified (natural) nucleotides of a given type is 20% or less, e.g., 15% or less, or 10% or less, or even 5% or less (all in molar percent). The modified nucleotides are incorporated in place of the corresponding natural nucleotides, but then undergo atypical base pairing in the method of the present invention (the behavior of base pairing changes as described in detail above). Then, another complementary nucleotide forms a base pair with a modified nucleotide that is different from the natural counterpart nucleotide that would have formed the base pair.Therefore, a change will occur in the sequence of the hybridized complementary chain (for example, a newly synthesized complementary chain). As a result, base pairing with at least one modified nucleic acid base enables base pairing with a different nucleotide than base pairing with an unmodified but otherwise identical nucleic acid base.
[0036] It is also possible to incorporate alkylated nucleic acid bases into PNA through biosynthesis (for example, without using thiol intermediates). For example, alkylated nucleotides can be incorporated into cells and made available to those cells during PNA synthesis. Such methods are described in Jao et al., Proc. Nat. Acad. Sci. USA Vol. 105(41), 2008: p. 15779 and Darzynkiewicz et al., Cytometry A Vol. 79A, 2011: p. 328. In particular, a modified nucleotide that is effective for use in this invention is 5-ethynyluridine (5-EU). Ethynyl-labeled uridine can penetrate cells and be incorporated into nascent RNA in place of the natural analog uridine. In preferred embodiments, the resulting ethynyl functional PNA is further modified through click chemistry catalyzed by, for example, Cu(I) (as described, e.g., Presolski et al., Current Protocols in Chemical Biology, Vol. 3, 2011: p. 153; or Hong et al., Angew. Chem. Int. Ed., Vol. 48, 2011: p. 9879) to introduce additional functional groups via azide functional molecules (e.g., NHS esters, maleimides, azido acids, azidoamines) that affect the ability of the ortho ketone to hydrogen bond with the ethynyl group.
[0037] In another embodiment, such azide functional molecules can be introduced into the cells themselves for biosynthesis and then introduced into PNA as modified nucleic acid bases. The resulting azide functional molecules can then be detected by Cu(I)-catalyzed (CuAAC) or Cu(I)-free azide-alkyne cycloaddition (SPAAC) click chemistry to introduce functional groups that alter the hydrogen bonding ability of the nucleic acid bases compared to unmodified nucleic acid bases (C, T / U, A, G).
[0038] Another example of modifying one or more nucleic acid bases in PNA involves incorporating vinyl functional nucleic acid bases (such as 5-vinyluridine) into PNA. The vinyl group can be further modified to alter the hydrogen bonding ability of the nucleic acid base if it were unmodified (see Rieder et al., Angew. Chem. Int. Ed. Vol. 53, 2014: p. 9168).
[0039] In particularly preferred embodiments, modification of one or more nucleic acid bases of PNA includes cyclization of an allyl group and / or halogenation (particularly iodization) of the nucleic acid bases of PNA. The modified nucleic acid bases are allyl nucleic acid bases (N 6 -Allyladenosine ("a 6 It is preferable that the allyl nucleic acid base is N) (e.g., A). The allyl nucleic acid base can be further modified by cyclization involving the allyl group. Such allyl nucleic acid bases can be incorporated into PNA during PNA synthesis, particularly into cells as described in other embodiments herein. Halogenation and / or cyclization can be carried out according to the principle described in Shu et al., J. Am. Chem. Soc., 2017, Vol. 139(48): pp. 17213-17216. Preferably, this method is used to introduce N into cells. 6- This involves iodation using, for example, elemental iodine (I2) after the incorporation of allyl adenosine. The iodized pre-allyl group is then cyclized with nitrogen on the purine group of the nucleic acid base (in the case of modified A or G) or the pyrimidine group (in the case of modified C or T / U). This modification alters base pairing, which can be read during sequencing or hybridization. For example, a 6 Because A behaves like A, it can be metabolically incorporated into newly synthesized RNA within mammalian cells. under mild buffer conditions 6 A's N 6 -N is formed by iodation of the allyl group 1 ,N 6 - Cycladized adenosine is spontaneously induced, leading to the generation of mutations in the opposite region during complementary DNA synthesis in reverse transcription.
[0040] In yet another preferred embodiment, modification of one or more nucleic acid bases in PNA includes the introduction of a 5-bromouridine (5-BrU) nucleic acid base into the PNA. 5-BrU is a mutagenic substance that exists as a tautomer. This means that 5-BrU exists in keto or enol forms that form base pairs with adenine or guanine (see Figure 37a), which in turn leads to increased T>C conversion (compared to unmodified U) in amplification reactions such as PCR. Thus, in this embodiment of the invention and in generally preferred embodiments, modification of one or more nucleic acid bases in PNA introduces a tautomer nucleic acid base that can form base pairs with both purine bases (A and G) in the case of modified T / U and modified C, and can form base pairs with both pyrimidine bases (T and C) in the case of modified A and G. Base pairing with both purine and pyrimidine bases means, as used herein, that the base pairing behavior is more equitable (but not necessarily equitable) than that of unmodified A, G, U / T, and C (which rarely pair with non-complementary bases). In other words, tautomer bases exhibit increased pairing with non-complementary bases of the same base core structure (purine or pyrimidine) compared to unmodified bases (fluctuating behavior). This increase occurs, in particular, under standard conditions for PCR.
[0041] Such fluctuation behavior can be determined from the resulting increase in mixed bases at specific positions corresponding to modified nucleic acid bases. Detection of fluctuation bases is a preferred read in all embodiments of the present invention (compare Figures 5B, 24B, and 37C).
[0042] In yet another related embodiment, further modification of 5-BrU or any other halogenated nucleic acid base is possible by substituting the halogen with an amino group. For example, 5-BrU can be converted to 5-aminouridine by heating it with ammonia. Such amino-modified nucleic acid bases alter base pairing during reverse transcription and / or introduce additional fluctuation behavior.
[0043] PNAs (containing modified nucleic acid bases) can contain or consist of RNA or DNA. Examples of RNA include mRNA, microRNA (miRNA or miR), short hairpin RNA (shRNA), small interfering RNA (siRNA), PIWI-interacting RNA (piRNA), ribosomal RNA (rRNA), small RNA derived from tRNA (tsRNA), transfer RNA (tRNA), small nucleolar RNA (snoRNA), small nuclear RNA (snRNA), long non-coding RNA (lncRNA), or their precursor RNA molecules. DNA can be, for example, genomic DNA, cDNA, plasmid DNA, or a DNA vector. PNAs can be double-stranded or single-stranded.
[0044] "Contains" relates to a term that has no limit on the number of items that can be listed, so a molecule can also contain other members, such as other types of nucleotides (RNA or DNA, which can include LNA, an artificially modified nucleotide). "Consists of" is considered a closed definition that requires the member to meet the condition (i.e., complete RNA or complete DNA).
[0045] Preferably, for each type of nucleotide selected from A, G, C, U, or T, the modified PNA contains more native nucleotides than the modified nucleotides. In this specification, PNA refers to the final PNA having any modifications according to the present invention. The PNA preferably contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more, and up to 30 modified nucleotides. Preferably, a small number of nucleic acid bases (e.g., 20% or less) in the PNA molecule are modified, for example, 15% or less, or 10% or less, or even 5% or less (all in molar percent) of the nucleotides are modified.
[0046] PNA molecules can have any length. Preferably, PNA molecules have a length of at least 10 nt (nucleotides). Particularly preferred are lengths of 10 nt, 20 nt, 30 nt, 40 nt, 50 nt, 75 nt, 100 nt, 250 nt, 500 nt, 1000 nt, 2500 nt, 5000 nt, 10000 nt, 25000 nt, 50000 nt, 100000 nt, or any range between these values. Preferred ranges are lengths from 10 nt to 100000 nt, or from 50 nt to 50000 nt.
[0047] Preferably, the PNAs originate from a specific cellular fraction of nucleotides (total RNA fraction, mRNA fraction, DNA fraction (such as plasmid DNA or genomic DNA)). The fractions can be selected by isolating PNAs that share common characteristics (length, nucleotide type, or sequence (e.g., poly(A) tail or 5' cap in mRNA)).
[0048] The method of the present invention includes the step of causing PNA to form a base pair with a complementary nucleic acid. In this base pairing, at least one modified nucleic acid base must form a base pair (usually by the base pairing of several nucleic acid bases of the PNA). Base pairing with a complementary nucleic acid can be easily achieved by hybridizing the PNA to a nucleic acid chain. This may occur during an extension reaction (e.g., PCR) or by hybridizing a probe nucleic acid. The complementary nucleic acid can have any length, for example, the lengths disclosed above with respect to PNA.
[0049] The complementary nucleic acid sequence is identified at a position complementary to at least one modified nucleic acid base. Sequencing can be achieved by any common procedure known in the art. Such methods include those based on generating the whole or part of the complementary strand by PCR, for example, as in next-generation sequencing (NGS), a fragment-based sequencing method. If desired, the fragment reads can be assembled into a combined sequence. However, in the use of the present invention, this is not necessary as long as the nucleic acid base complementary to the modified nucleic acid base is identified, in particular, along with its adjacent sequences (±5 nt, or ±10 nt, or ±15 nt, or ±20 nt). Another method for determining the sequence is binding to a probe. In this way, the sequence of the PNA is determined as a complementary sequence through the sequence of a known probe that hybridizes.
[0050] Another option is sequencing of small nucleic acids, especially when the complementary nucleic acid is small, such as in the case of nucleic acids complementary to miRNA, shRNA, and siRNA. Small nucleic acids can be in the range of, for example, 10 nt to 200 nt in length, preferably in the range of 12 nt to 100 nt, or 14 nt to 50 nt. They can also be longer than 200 nt or shorter than 10 nt. Fragments of complementary nucleic acids can have such lengths on average. Fragments can be generated by physical or chemical means known in the art with respect to NGS. In the case of small nucleic acids (including fragments obtained during NGS), it is preferable to conjugate an adapter to that nucleic acid which can be used as a hybridization sequence for primers or probes. Such adapters may also contain a characteristic sequence (such as a barcode) for identifying the small nucleic acid by labeling. The barcode can provide labeling for the source of the sample from which the PNA was obtained, or the source of the PNA molecule, or the source of the nucleic acid that was complementary to that PNA molecule and was a fragment (source of the fragment). Such barcodes may be useful in multiplex sequencing where a large number of nucleic acids with different sequences (e.g., multiple different complementary nucleic acids and / or multiple fragments of one or more complementary nucleic acids) are sequenced. Such a large number could be, for example, 2 to 1000 or more nucleic acids. Another possibility is to hybridize a primer or probe to a complementary nucleic acid sequence corresponding to the PNA, without necessarily requiring an adapter. Such primers or probes can be hybridized to known sequences or randomly using, for example, random primers. Random primers are described below in relation to the kit of the present invention. Any such random primer can be used in the method of the present invention.
[0051] In a preferred embodiment of the present invention, the PNA of a single cell is identified according to the present invention. Thus, the cell's PNA is isolated and maintained separate from the PNA of other cells. "Maintaining the PNA in isolation" means that the PNA sequencing information of the cell under investigation does not mix with the sequencing information of other cells, and the PNA of the cell under investigation remains identifiable. This can be achieved by physically separating the PNA or by labeling the PNA or complementary nucleic acid with a label (e.g., a barcode) to identify the cell of interest. This makes it possible to analyze the PNA metabolism of a single cell. Single-cell analysis can be performed by single-cell sequencing (Eberwine et al., Nat. Methods. Vol. 11(1): pp. 25-27). Alternatively, complementary nucleic acids or fragments thereof may be prepared in a library, preferably (but not necessarily) with an adapter. This library can then be independently sequenced or provided for other uses.
[0052] The modification of the present invention (e.g., thiol-specific alkylation) facilitates a quantifiable degree of "wrong" incorporation of complementary nucleotides, which now form different hydrogen bonding partners as described above. Guanosine may be incorporated, for example, during transcription or reverse transcription, through complementary nucleic acid bonding with a modified nucleic acid base (e.g., alkylated 4-thiouridine) instead of adenosine. Nevertheless, the processing performance of the (reverse) transcriptase is usually not affected, because the nucleotides that instead formed base pairs can be amplified with the corresponding PNA without further interference. Preferably, the second modification is combined with the first modification (by incorporation of a modified nucleic acid base) using an enzyme. For example, the well-established and non-toxic s 4 This is a combination with a U metabolic labeling protocol.
[0053] The sequencing method of the present invention, in which the sequence of complementary nucleic acids is altered due to modified nucleic acid bases, can be combined with available high-throughput sequencing methods (e.g., NGS). Sequence changes that differ between individual different molecules of PNA / complementary nucleic acids can be identified by available computerized methods, especially when they are incomplete or partial. For example, T>C conversion (resulting from U modification leading to increased G base pairing) can be tracked in next-generation sequencing datasets. By combining such highly automated methods with computerized analysis, the present invention provides rapid access to the dynamics of intracellular RNA processing. This is one preferred application of the present invention. The present invention can accurately inform RNA polymerase II-dependent transcriptional output resulting from complementary base pairing. Insights into the intracellular dynamics of RNA biosynthesis, processing, and turnover are essential for elucidating the molecular basis of changes in gene expression patterns that affect virtually every biological process in vivo.
[0054] Therefore, in one preferred embodiment, the method of the present invention can be used to reveal modifications or easily modifiable changes of intracellular PNA. Such "easily modifiable changes" relate, for example, to the multi-step method described above, in which the first modification (also called a change) is carried out intracellularly, and subsequent modifications are carried out extracellularly, usually in a second or subsequent step, after the PNA has been isolated.
[0055] Preferably, the method of the present invention is used to carry out at least the first modification / change within a cell (especially a living cell) to modify RNA (such as PNA). Since the expressed RNA is modified, this makes it possible to track changes in RNA expression.
[0056] Regulated expression of genetic information is essential for maintaining cellular homeostasis, giving cells the flexibility to respond to changing environmental conditions, while abnormal regulation can lead to human diseases such as cancer. Behind these critical biological processes are tightly regulated molecular events that control the relative dynamics of RNA transcription, processing, and degradation in a transcript-specific manner.
[0057] The cellular RNA pool (which includes numerous RNA species, such as mRNA and non-coding RNAs (microRNAs)) is determined by the transcription of selected gene loci within the genome and can be qualitatively and quantitatively evaluated using RNA profiling techniques (such as high-throughput sequencing). However, measuring RNA content at steady-state levels does not accurately reflect transcriptional activity itself. In reality, RNA stability plays a major role in determining the relative content of RNA molecules. Therefore, approaches to measure transcription rates and RNA degradation rates at the genome scale are useful for gaining insight into the dynamics of RNA expression and the underlying regulatory mechanisms. According to the present invention, it is possible to elucidate the intracellular dynamics of RNA biosynthesis and metabolic turnover.
[0058] RNA can be altered or modified by the cell's own metabolism by incorporating altered or modified nucleotides into naturally processed RNA. Such alterations can be used to selectively introduce the modifications of the present invention, altering the behavior of hydrogen bonding (either alone or after further modification). When the modified nucleotides are subsequently sequenced due to metabolic influence, such a method is called "metabolic sequencing." The sequencing process, or any base-pairing process with complementary (poly)nucleotides in general, can be automated and handled in the high-throughput sequencing methods described above. The present invention provides a high-throughput metabolic labeling protocol suitable for elucidating the intracellular dynamics of RNA biosynthesis and turnover. This protocol solves the problem of providing high temporal resolution of intracellular RNA expression dynamics (including biosynthesis and turnover) by accurately measuring the polyadenylated transcription output dependent on RNA polymerase II and reproducing a comprehensive post-transcriptional gene regulatory signature.
[0059] Possible cells include any type of cell, such as bacterial cells (including eukaryotic and prokaryotic cells), Gram-negative cells, Gram-positive cells, fungal cells, algal cells, plant cells, animal cells, mammalian cells (rodent cells, primate cells, human cells, non-human cells, etc.), archaeal cells, avian cells, amphibian cells (frog cells, etc.), reptile cells, and marsupial cells.
[0060] By controlling the modification over time, it is possible to monitor the changes, for example, by comparing a stage of cellular RNA expression without modification with a stage of cellular RNA expression with modification. It is preferable that such stages be compared within the same cells or cell culture, for example, a stage with modification followed by a stage without modification, or vice versa. Thus, in a preferred embodiment of the present invention, one or more cells are cultured in at least two culture stages. In this case, one culture stage involves incorporating a modified nucleotide into biosynthesized RNA, where the RNA is modified by the addition or removal of a hydrogen bonding partner, while another culture stage lacks such incorporation of the modified nucleotide into the biosynthesized RNA. The “another culture stage” may also involve incorporating the modified nucleotide into the biosynthesized RNA at a different (e.g., lower) concentration than in the other culture stage. The different or lower concentration must be sufficient to observe a difference (particularly a difference in concentration) in the incorporation of the modified nucleotide into the biosynthesized RNA. Therefore, the method of the present invention can be defined as a method for identifying polynucleic acid (PNA), which includes the steps of: expressing PNA in a cell; modifying one or more nucleic acid bases of the PNA; isolating the PNA from the cell; optionally, further modifying the PNA before, after, or simultaneously with isolation, thereby altering the base-pairing ability of one or more nucleic acid bases by adding or removing hydrogen bonding partners of one or more nucleic acid bases; base-pairing a complementary nucleic acid with the PNA, including base-pairing with at least one modified nucleic acid base; and identifying the sequence of the complementary nucleic acid at a position complementary to at least one modified nucleic acid base. Particularly preferred metabolic labeling (i.e., modification by cellular metabolism (e.g., enzymes such as RNA polymerase)) is by a 4-thiouridine integration event. This can be used to alter the base-pairing behavior of U.
[0061] Particularly preferred is a method involving at least two stages of cell culture, which facilitates ensuring that the levels of PNA modification, particularly RNA modification, differ in at least two stages of culture. This can be achieved by supplying cells with modified nucleic acid bases at different concentrations, thereby allowing the cells to incorporate the modified nucleic acid bases into the PNA (particularly RNA) at different levels or concentrations. As described above, the modified nucleic acid bases are preferably thiol-modified. The level of PNA modification in one stage can be no modification at all. Each stage, particularly the stage in which PNA modification occurs, needs to have a predetermined period for that PNA modification. By comparing the incorporation at different stages with each other, it is possible to calculate the turnover rate during that predetermined period. In one particularly preferred embodiment, the rate of turnover or degradation is calculated based on comparing the incorporation of modified nucleic acid bases into the PNA in at least one stage with that of other stages. The stages are preferably consecutive culture stages.
[0062] Further comparisons can be performed between different cell culture stages. Such comparisons make it possible to estimate expression differences and PNA turnover between these cells. One cell or cell group can be used as a control, and the other cell or cell group can be a candidate cell or candidate cell group to be examined. Both cells or cell groups may have a stage in which modified nucleic acid bases are incorporated into PNA, and their PNAs are compared. Preferably, such an incorporation stage is controlled by supplying the cells with modified nucleic acid bases so that they are incorporated into PNA. Preferably, equal amounts of modified nucleic acid bases suitable for comparison of cellular metabolism are supplied to each cell or cell group. Preferably, the incorporation stage is followed by a stage of no further incorporation, which is, for example, by stopping further supply of modified nucleic acid bases to the cells or cell group. It is also possible that the incorporation stage is followed by a stage of reduced incorporation or a stage of incorporation at different levels. After some change in the level at which modified nucleic acid bases are incorporated into PNA, cellular metabolic adaptation follows. This adaptation can be monitored by the method of the present invention. For example, if an embedded stage is followed by a stage with reduced or no embedded stages, it is possible to monitor the decomposition of the modified PNA. If a stage with no embedded stages or limited embedded stages is followed by a stage with more embedded stages than the limited embedded stages, it is possible to monitor the increase in the modified PNA.
[0063] Therefore, one application of the present invention is to compare identified sequences of complementary nucleic acids in at least two cells or at at least two different growth stages within one cell at a position complementary to at least one modified nucleic acid base (as described above). These at least two cells or at least two growth stages exhibit different expression (usually gene expression, but including mRNA or regulatory RNA expression) between the at least two cells or between the at least two growth stages. This difference in (gene) expression can be induced by repressing or stimulating at least one gene within the cell. Using such a method, it is possible to screen the effect of certain disturbances in cellular metabolism on differences in expression. Since this difference in expression may be due to an unknown gene in the screening method, for example, it is possible to investigate whether a regulatory inhibitor or regulatory activator, or any other substance that has an effect on phenotype, has a particular gene effect within the cell. In another embodiment of this method, the target gene can be known so that further secondary effects of other genes on gene expression can be investigated. For example, known regulatory genes (such as oncogenes or tumor suppressor genes) can be known genes.
[0064] PNA can be administered as cells or cell populations in an in vitro culture, or in vivo from living organisms (plants, bacterial cells, fungal cells, algal cells, non-human animals, humans). In the case of in vivo cells, modified nucleic acid bases can be supplied to the cells by systemic administration to the organism, for example, to the vascular system, or by local administration to an organ of interest. Thus, it is possible to monitor PNA metabolism in vivo or in a specific organ of interest. PNA is then isolated from the organism, for example, by autopsy, or from a fluid sample in the case of secreted PNA, or by euthanasia of a non-human organism. Preferably, after isolating single-cell PNA from the organism, it is labeled, for example as described above, and / or a library is generated and / or analyzed by single-cell sequencing according to the method of the present invention. Any description of the culture stage also applies to in vivo processing and is called the “growth stage.” The “growth stage” does not require cell proliferation or cell doubling and refers to PNA metabolism or “growth” that is identified and analyzed.
[0065] Comparing the levels of PNA and PNA turnover is important for revealing differences in cellular metabolism between cells in different states during organism growth or disease. Measuring the rate of PNA turnover helps to identify which pathways are active and which are less active or inactive. In this regard, the turnover rate provides an additional indicator to steady-state concentration measurements of PNA (e.g., mRNA present in cells, tissues, or organs), which only measure the concentration of PNA (e.g., mRNA present in cells, tissues, or organs).
[0066] Preferably, the biosynthesized PNA (preferably RNA) from the two culture stages is recovered from the cells and preferably mixed, and the formation of base pairs with the PNA on the complementary nucleic acid includes generating a complementary polynucleotide chain (preferably a DNA chain) by transcription (reverse transcription if the PNA is RNA).
[0067] Particularly beneficial to the present invention is the elimination of the need to separate the modified PNA produced from the unmodified or less modified equivalent PNA, or their respective complementary nucleic acids. The base pairing of PNA with complementary nucleic acids can be performed in a mixture of both modified and unmodified PNA. The PNA / complementary nucleic acid sequences can then be determined in combination because the sequence / characteristics of the complementary nucleic acid can be determined in both modified and unmodified cases, and the modification event can be inferred by comparison. Such comparisons are preferably computerized sequence comparisons. A particularly preferred embodiment of the present invention, which involves forming a base pair with at least one modified nucleic acid base, thereby leading to the formation of a base pair with a different nucleotide than with an unmodified nucleic acid base, further includes determining the sequence of a complementary polynucleotide chain and comparing its chain sequence, so that the complementary nucleic acid altered as a result of modification by the addition or removal of hydrogen bond partners can be identified by comparison with that complementary nucleic acid in the unmodified case. Preferably, the nucleotide sequence is obtained as a fragment, such as those used in NGS and high-throughput sequencing. (The desired sequence may contain at least one modified nucleic acid base and a complementary position.) The desired sequence can be 10 nt to 500 nt, preferably 12 nt to 250 nt, or 15 nt to 100 nt in length.
[0068] Computer identification of a complementary nucleic acid sequence at a position complementary to at least one modified nucleic acid base can include comparison with an unmodified PNA sequence. Such comparison sequences can be obtained from sequence databases (EBI or NCBI) or by generating PNA without introducing modifications (e.g., by base-pairing a natural base with a naturally complementary base). A computer program product for such comparison, or a computer-readable medium for the method, can be included in the kit of the present invention.
[0069] The present invention further provides a kit suitable for carrying out the method of the present invention, comprising a thiol-modified nucleic acid base and an alkylating agent suitable for alkylating the thiol group of the thiol-modified nucleic acid base, wherein the alkylating agent comprises a hydrogen bond donor or a hydrogen bond acceptor. The alkylating agent is preferably one of those mentioned above, particularly iodoacetamide. However, any of the above alkylating agents, or an alkylating agent suitable for modified nucleotides (such as thiol-modified nucleotides) having any of the above modifications (particularly modified bases), can be included in the kit of the present invention.
[0070] The kit preferably further includes primers, nucleotides selected from A, G, C, and T, and reverse transcriptase, or a combination thereof, and preferably all of these components. An example of a primer is a random primer. A random primer is a mixture of several randomly selected primers. Such a random primer mixture may contain at least 50, or at least 100, or at least 500 different primers. A random primer may contain random hexamers, random pentamers, random pentamers, random octamers, and so on.
[0071] The kit may further include PNA polymerase, preferably a buffer for polymerizing that polymerase. The polymerase can be either DNA polymerase or RNA polymerase.
[0072] The kit of the present invention may also include an adapter nucleic acid. Such an adapter can be bound to a nucleic acid to produce a complementary nucleic acid to which the adapter is bound. The adapter may include one or more barcodes as described above. The kit may also include a ligase (such as a DNA ligase).
[0073] The components of the kit can be provided in appropriate containers (such as vials or flasks).
[0074] The kit may also include instructions or a manual for carrying out any method of the present invention.
[0075] The present invention will be further illustrated by the following drawings and embodiments, but the present invention is not limited to these aspects. [Brief explanation of the drawing]
[0076] [Figure 1] A schematic diagram of thiol (SH) linkage alkylation for performing metabolic sequencing of RNA. When cells are treated with 4-thiouridine (s4U), the s4U is incorporated into newly transcribed RNA upon uptake by the cells. When total RNA is prepared at a given time, s4U residues present in the newly generated RNA species are carboxyamidomethylated by treatment with iodoacetamide (IAA), resulting in bulky groups at the base-pairing interface. When the presence of bulky groups at the s4U integration site is combined with well-established RNA library preparation protocols, G is mistakenly incorporated into the alkylated s4U to a specific and quantifiable degree during reverse transcription (RT). The s4U-containing site can be identified at single-nucleotide resolution by bioinformatics in a high-throughput sequencing library by calling the conversion from T to C. [Figure 2A-E]Derivatization of 4-thiouracil by thiol-linked alkylation. (A) 4-thiouracil (s4U) reacts with the thiol-reactive compound iodoacetamide (IAA) in a nucleophilic substitution (SN2) reaction, resulting in the addition of a carboxyamide methyl group to the thiol group in s4U. The wavelengths at which absorption is maximal are shown for the extract (4-thiouracil; s4U; λmax ≈ 335 nm) and the product (carboxyamide methylated 4-thiouracil; *s4U; λmax ≈ 297 nm). (B) Absorption spectra of 4-thiouracil (s4U) in the absence and presence of iodoacetamide (IAA) at the stated concentrations. 1 mM s4U was incubated with IAA at the stated concentrations in 50 mM sodium phosphate buffer (pH 8.0) and 10% DMSO at 37°C for 1 hour. Data represent the mean ± SD of at least three independent replicates. (C) Quantitative results of absorption at 335 nm shown in (B). The p-value (Student's t-test) is shown. (D) Absorption spectra after incubation of 1 mM 4-thiouracil (s4U) for 5 minutes at the specified temperature in the presence of 50 mM sodium phosphate buffer (pH 8.0) and 10% DMSO in the absence and presence of 10 mM iodoacetamide (IAA). The data represent the mean ± SD of at least three independent replicates. (E) Quantitative results of absorption at 335 nm shown in (D). The p-value (Student's t-test) is shown. [Figure 2F-M]Derivatization of 4-thiouracil by thiol-linked alkylation. (F) Absorption spectra after incubation of 1 mM 4-thiouracil (s4U) in the presence of 50 mM sodium phosphate buffer (pH 8.0) and 10% DMSO at 37°C for the stated time, in the absence and presence of 10 mM iodoacetamide (IAA). Data represent the mean ± SD of at least three independent replicates. (G) Quantification of absorption at 335 nm shown in (F). p-value (Student's t-test) is shown. (H) Absorption spectra after incubation of 1 mM 4-thiouracil (s4U) in the presence of 50 mM sodium phosphate buffer (pH 8.0) and the stated amount of DMSO at 50°C for 2 minutes, in the absence and presence of 10 mM iodoacetamide (IAA). Data represent the mean ± SD of at least three independent replicates. (I) Quantification results of absorption at 335 nm shown in (H). The p-value (Student's t-test) is shown. (J) Absorption spectra after incubation of 1 mM 4-thiouracil (s4U) in the absence and presence of 10 mM iodoacetamide (IAA) at 50°C for 5 minutes in the presence of 50 mM sodium phosphate buffer (pH 8.0) and 10% DMSO (optimal reaction [rxn] conditions) in the absence and presence of 10 mM iodoacetamide (IAA). The data represent the mean ± SD of at least three independent replicates. (K) Quantitative results of absorption at 335 nm shown in (J). The p-value (Student's t-test) is shown. (L) Absorption spectra after incubation of 1 mM 4-thiouracil (s4U) in the absence and presence of 10 mM iodoacetamide (IAA) at 50°C for 15 minutes in the presence of 50 mM sodium phosphate buffer (pH 8.0) and 50% DMSO (optimal reaction [rxn] conditions). The data represent the mean ± SD of at least three independent replicates. The quantitative results of absorption at 335 nm are shown in (M) and (L). The p-value (Student's t-test) is also shown. [Figure 3A]Derivatization of 4-thiouridine by thiol bond alkylation. (A) 4-thiouridine (s4U) reacts with the thiol-reactive compound iodoacetamide (IAA) to undergo a nucleophilic substitution (SN2) reaction, resulting in the addition of a carboxyamide methyl group to the thiol group in s4U. [Figure 3B-C] (B) Derivatization of 4-thiouridine by thiol-linked alkylation. Analysis of s4U alkylation by mass spectrometry. 40 nanomoles of 4-thiouridine were incubated with the specified concentration of iodoacetamide at 50°C for 15 minutes in standard reaction buffer (50 mM NaPO4 (pH 8), 50% DMSO). The reaction was stopped with 1% acetic acid. The acidified sample was separated at a flow rate of 100 μl / min using a Kinetex F5 pentafluorophenyl column (150 mm × 2.1 mm; 2.6 μm, 100 Å; Phenomenex) on an Ultimate U300 BioRSLC HPLC system (Dionex; Thermo Fisher Scientific). The nucleoside was electrospray ionized and then analyzed online using a TSQ Quantiva mass spectrometer (Thermo Fisher Scientific). The SRMs at that time were as follows: 4-thiouridine m / z 260 → 129, alkylated 4-thiouridine m / z 318 → 186. Data were interpreted using the Trace Finder software suite (Thermo Fisher Scientific) and manually verified. Quantitative results from two independent experiments with two technical replicates shown in (C) and (B). The proportion of alkylated s4U with the stated IAA concentration represents the normalized relative signal intensity at peak retention time for s4U and alkylated s4U. Data are shown as mean ± SD. [Figure 4A-C]Alkylation of 4-thiouridine-containing RNA does not affect the processing performance of reverse transcriptase. (A) To clarify the effect of s4U alkylation on the processing performance of reverse transcriptase, a 76 nt synthetic RNA containing 4-thiouracil (s4U) incorporated at a single position (p9) within the sequence of dme-let-7, a small Drosophila RNA with adjacent 5' and 3' adapter sequences, was used. Reverse transcription was performed using commercially available reverse transcriptase, following the extension of a 5' 32P-labeled DNA oligonucleotide (with a sequence opposite and complementary to the 3' adapter sequence), both before and after treatment with iodoacetamide (IAA). (B) The reaction products prepared as in (A) were analyzed by polyacrylamide gel electrophoresis, followed by fluorescence imaging. The results of primer extension of s4U-containing RNA and s4U-free RNA with and without IAA treatment are shown, using either the reverse transcriptase Superscript II (SSII), Superscript III (SSIII), or Quant-seq RT (QS). The sequences of RNA components excluding the adapter sequence are shown. The position of the s4U residue is shown in red. RNA sequencing was performed by adding the described ddNTP to the reverse transcript reaction product. PR: DNA primer labeled at 5' 32P; bg: background stop signal; *p9: stop signal at position 9; FL: full-length product. Quantitative results of three independent replicates from the experiments shown in (C) and (B). The ratio of elimination signals at p9 (with IAA treatment vs. without IAA treatment), normalized to the preceding background elimination signal, was determined for control RNA and s4U-containing RNA using the described reverse transcriptase. Data represent mean ± SD. Statistical analysis was performed using Student's t-test. [Figure 5A] Alkylation allows for the quantitative identification of s4U integration into RNA with single-nucleotide resolution. (A) RNA with or without 4-thiouridine (s4U) at a single position (p9) was treated with iodoacetamide (IAA), followed by reverse transcription and gel extraction of the full-length product, and then PCR amplification and high-throughput (HTP) sequencing. [Figure 5B-1] (B) The mutation rates at each position of control RNA (left figure) and s4U-containing RNA (right figure) are shown, with and without iodoacetamide (IAA) treatment using the reverse transcriptase described. The bars represent the mean mutation rate ± SD of the three independent replicates. The number of reads sequenced for each replicate (r1~r3) is shown. The nucleotides appearing at p9 are indicated. [Figure 5B-2] (B) The mutation rates at each position of control RNA (left figure) and s4U-containing RNA (right figure) are shown, with and without iodoacetamide (IAA) treatment using the reverse transcriptase described. The bars represent the mean mutation rate ± SD of the three independent replicates. The number of reads sequenced for each replicate (r1~r3) is shown. The nucleotides appearing at p9 are indicated. [Figure 5C] (C) Mutation rates for the described mutations, with and without IAA treatment, using either reverse transcriptase Superscript II (SSII), Superscript III (SSIII), or Quant-seq RT (QS). Mutation rates were averaged for the same nucleotide positions in both s4U-containing and s4U-non-containing RNA oligonucleotides. The p-value (determined by Student's t-test) is shown. ns indicates non-significant (p>0.05). [Figure 6]Effects of s4U treatment on mES cell viability and metabolic RNA labeling. (A) Vitalization of mES cells cultured for 12 hours (left) or 24 hours (right) in the presence of 4-thiouridine (s4U) at the stated concentrations compared to untreated cells. The final concentration (100 μM) used in subsequent experiments is indicated by triangles and dotted lines. (B) Quantification of s4U incorporated into total RNA after pulsed s4U metabolic labeling over the stated time, or after medium changes in uridine tracking. s4U incorporation was determined to single-nucleoside precision by HPLC analysis after digestion and dephosphorylation of total RNA. The signal intensity of s4U at 330 nm, normalized to 24 hours after pulsed labeling, with background subtracted, and the absorbance of the signal intensity of uridine at 260 nm are shown. (C) The substitution rate of s4U compared to unmodified uridine was determined by HPLC. s4U integration into total RNA at all time points in experiments with s4U metabolic pulse and tracking labeling in mES cells. Numerical values represent the mean ± SD of three independent replicates. Maximum integration rate is shown 24 hours after labeling. [Figure 7] This protocol prepares a Quant-seq mRNA 3' end sequencing library. Because Quant-seq uses total RNA as input, there is no need to pre-enrich poly(A) or deplete rRNA. Library generation is initiated by oligo(dT) priming. Primers already contain Illumina-compatible linker sequences (shown in green, top: "adapter", next step: final bend). After synthesizing the first strand, RNA is removed, and the synthesis of the second strand is initiated by random priming and DNA polymerase. Random primers also contain Illumina-compatible linker sequences (shown in blue). Purification is not required between the synthesis of the first and second strands. Insertion size is optimized for shorter reads (SR50 or SR100). After synthesizing the second strand, a purification step based on magnetic beads follows. The library is then amplified to introduce the sequences necessary for cluster generation (shown in red and purple). For redundancy, an external barcode (BC) is introduced between PCR-amplified antibodies. [Figure 8A] s4U integration events in metabolically labeled mES cell mRNA. (A) Representative genome browser screenshots for three independent mRNA libraries generated from total RNA of mES cells. These mRNA libraries were prepared using standard mRNA sequencing (top three figures), Cap analysis gene expression (CAGE; middle three figures), and mRNA 3' end sequencing (bottom three figures). Representative regions encoding Trim28 in the mouse genome are shown. [Figure 8B-1] (B) Enlarged view of the genomic region encoding the mRNA 3' end of Trim28, including the 3' untranslated region (UTR). Coverage plots are shown of modified mRNA 3' end sequencing libraries prepared from total RNA of untreated mES cells or mES cells that have been s4U metabolically labeled for 24 hours, followed by sequencing. A subset of individual reads that forms the basis of the coverage plot is shown. Red bars within individual reads represent T>C conversions, and black bars represent any mutation other than T>C. [Figure 8B-2] (B) Enlarged view of the genomic region encoding the mRNA 3' end of Trim28, including the 3' untranslated region (UTR). Coverage plots are shown of modified mRNA 3' end sequencing libraries prepared from total RNA of untreated mES cells or mES cells that have been s4U metabolically labeled for 24 hours, followed by sequencing. A subset of individual reads that forms the basis of the coverage plot is shown. Red bars within individual reads represent T>C conversions, and black bars represent any mutation other than T>C. [Figure 9]Overall analysis of mutation rates in mRNA 3'-terminal sequencing after s4U metabolic labeling of RNA in mES cells. mRNA 3'-terminal sequencing libraries generated from total RNA in mES cells were mapped to annotated 3' non-truth regions (UTRs) before and after 24 hours of s4U metabolic labeling, and mutation rates were determined for all expressed genes. Tukey box plots show mutations per UTR as percentages. Outliers are not shown. The median frequency observed for each individual mutation is shown. Statistical analysis of increased T>C conversion was performed using the Mann-Whitney test. [Figure 10A] Measurement of the transcriptional output of polyadenylated mRNA in mES cells. (A) Experimental setup for determining the transcriptional output of polyadenylated mRNA in mES cells according to the present invention. [Figure 10B] (B) The relative content of transcripts containing T>C conversion ("SLAM-seq") and transcripts not containing T>C conversion ("steady state") was determined in parts per million counts (cpm) by mRNA 3' end sequencing as described in (A). Transcripts that are excessively expressed in SLAM-seq compared to the steady state are shown in red (high transcription output; n=828). Most transcripts that are abundant in the steady state are shown in yellow (high steady-state expression; n=825). Transcripts corresponding to the major mES cell-specific miRNA clusters miR-290 and miR-182 are shown. [Figure 10C] (C) The top 825 genes detected in the steady state by conventional mRNA 3' end sequencing (steady state) and 828 genes that were over-reproduced from newly transcribed RNA (SLAM-seq) were compared in terms of prediction of underlying transcription factors (using Ingenuity Pathway Analysis; www.Ingenuity.com) and prediction of molecular pathways (using Enrichr). [Figure 11]Overall analysis of mRNA stability in mES cells. (A) Experimental setup for investigating the stability of polyadenylated mRNA in mES cells according to the present invention. (B) Overall analysis of mRNA half-life in mES cells. The relative proportion of T>C conversion-containing reads mapped to the annotated 3' UTR of 9430 genes abundantly expressed in mES cells was normalized to the point of 24-hour pulse labeling. When the time course of the median, upper quartile, and lower quartile was fitted using a single exponential decay kinetics, it was found that the median mRNA half-life (~t1 / 2) was 4.0 hours. (C) Half-life calculations for examples of individual transcripts. For Junb, Id1, Eif5a, and Ndufa7, the average proportion of T>C conversion-containing reads in three independent replicates is shown compared to the point of 24-hour pulse, fitting a single exponential decay kinetics. The mean half-life (t1 / 2) of each transcript, determined by curve fitting, is shown. (D) Tukey box plots of mRNA half-lives, determined by mRNA 3'-end sequencing, for transcripts classified according to the relevant GO term (i.e., transcriptional regulation, signal transduction, cell cycle, development) or housekeeping (i.e., extracellular matrix, metabolic processes, protein synthesis). The number of transcripts in each category is shown. The p-value (determined by the Mann-Whitney test) is shown. [Figure 12A] Thiol-linked alkylation for metabolic labeling of small RNAs. (A) Screenshot of a representative genome browser for a small RNA library generated from size-selected total RNA in Drosophila S2 cells. Representative regions encoding miR-184 in the Drosophila genome are shown. The relative positions of nucleotides encoding thymine (T, red) are shown relative to the 5' end of each small RNA species. For miR-184-3p and miR-184-5p, 99% of all 5' isoforms are shown, along with the number of each read in parts per million (ppm). [Figure 12B](B) Small RNA sequencing libraries generated from total RNA of Drosophila S2 cells before and after 24 hours of s4U metabolic labeling were mapped to annotated miRNAs, and mutation rates were determined for miRNAs that were abundantly expressed (over 100 ppm). Tukey box plots show the mutations for each miRNA as percentages. Outliers are not shown. The median frequency of observed mutations is shown for each individual mutation. The p-values obtained by the Mann-Whitney test are shown. [Figure 13A-B] Intracellular dynamics of microRNA biosynthesis. (A) MicroRNAs originate from hairpin-containing PNA polymerase II transcripts (primary microRNAs, pre-miRNAs) and are processed sequentially by the RNase III enzymes Drosha in the nucleus and Dicer in the cytoplasm to form approximately 22 nt double-stranded microRNAs. Hairpin processing intermediates (precursor microRNAs, pre-miRNAs) are transported from the nucleus to the cytoplasm by Ranbp21 in a RanGTP-dependent manner. (B) Cumulative distribution plots show the median T>C mutations for 42 abundantly expressed miRNAs (left) or 20 miR* (right) in a small RNA library generated from total RNA of Drosophila ago2ko S2 cells treated with s4U over the indicated time. p-values were determined by the Kolmogorov-Smirnov test (****=p<10⁻⁴). Bgmax indicates the maximum background error rate. [Figure 13C] (C) The average content of the listed miRNA (left) or miR* (right) is shown for the steady state or for T>C mutation-containing reads at the time point after metabolic labeling with s4U. The number of reads normalized to all small RNAs is shown in units of ppm (parts per million). Cases where the mutation rate (Mu) is greater than the background maximum (Bgmax) are shown. [Figure 13D-E](D) Miltron hairpins are generated through splicing of the protein encoding the transcript. Post-transcriptional uridylation is performed on the miltron hairpins in the cytoplasm after debranching of the lasso intron. This alters the 2 nt 3' overhang of pre-miltron, and miRNA biosynthesis is inhibited by Dicer. (E) T>C mutation rate (top) for classical miRNA (gray) or miltron (red) at the time point described in the s4U labeling experiment, or T>C reads normalized to small RNA (units are parts per million (ppm)). Median and interquartile range are shown. p-values (Mann-Whitney test) are shown (*, p<0.05; **, p<0.01; ***, p<0.001). nd, not detected. [Figure 14] Intracellular dynamics of microRNA loading. (A) MicroRNA double strands, upon generation, are loaded into the Argonaut protein Ago1. In this process, one strand (miR strand) is selectively retained within Ago1, while the other strand (miR* strand) is excluded and degraded. Single-stranded miRNAs bound to Ago1 form the mature miRNA-induced silencing complex (miRISC). (B) Median accumulation (in ppm) of T>C conversion-containing reads of 20 richly expressed miR and miR* pairs during s4U metabolic labeling experiments in ago2ko S2 cells. Median and interquartile ranges are shown for miR (red) and miR* (blue). The values are from two independent experiments. The p-values indicate significant segregation of miR and miR* as determined by the Mann-Whitney test (*, p<0.05; ****, p<0.0001). (C) Enlarged view of the time-dependent changes shown in (B) for miR (red, top) and miR* (blue, bottom). [Figure 15]Intracellular dynamics of exonucleolytic miRNA trimming. (A) A model for exonucleolytic miRNA maturation in Drosophila. MicroRNAs (e.g., miR-34) are aggregated by Dicer to generate a longer miRNA double strand of approximately 24 nt. This miRNA double strand is loaded into Ago1, and when the miR* strand is removed, it undergoes exonucleolytic maturation by the 3'→5' exoribonuclease Nibbler, forming a mature gene regulatory miRNA-induced silencing complex. (B) Steady-state length distribution of miR-34-5p in Drosophila ago2ko S2 cells was determined by high-throughput sequencing of small RNAs (left, bars represent the mean ± standard deviation of 18 measurements; mean count of cloning is shown as a percentage per million (ppm)) or by Northern-Hybridization experiments (right). (C) Length distribution of miR-34-5p in a library prepared from s4U-metabolized Drosophila ago2ko S2 cells over the time period indicated. The length distribution of T>C conversion-containing reads (labeled, red, top) and all reads (steady state, black, bottom) is shown. The number of underlying reads is indicated. The data shows the mean ± standard deviation of two independent replicates. (D) Weighted mean length of miR-34-5p in a library prepared from s4U-metabolized Drosophila ago2ko S2 cells over the time period indicated. The data shows the mean ± standard deviation of T>C conversion-containing reads (labeled, red) and all reads (steady state, black). The decrease in the weighted mean length of T>C conversion-containing reads indicates that exonucleic acid degradation trimming occurred (highlighted by the gray area). (E) Loading of miR-34-5p determined by the relative content of T>C conversion-containing reads in miR-34-5p (miR chain, red) and miR-34-3p (miR* chain, blue) after s4U metabolic labeling of Drosophila S2 cells. The mean ± standard deviation of two independent experiments is shown. Loading is represented by the separation of miR from miR* and is highlighted by the gray area. [Figure 16A-C]Differences in miRNA stability. (A) When a microRNA double-strand is loaded into Ago1, a pre-miRISC is formed, and as a result of the degradation of the miR* strand (blue), it becomes a mature miRNA-induced silencing complex (miRISC). The stability of miRNAs within the miRISC remains precisely unknown. (B) Increase in mutation rate per T-site for 41 miRNAs (red, left) and 20 miR* (blue, right) that were highly expressed during the s4U metabolic labeling period in Drosophila ago2ko S2 cells. For each small RNA, the median mutation rate at all T-sites was calculated and normalized to 24 hours. The median and interquartile range are shown. The values represent the mean of two independent replicates. The median half-life (t1 / 2) and 95% confidence interval were obtained by fitting a single exponential curve. (C) Tukey box plot showing the half-lives of 41 miR (red) and 20 miR* (blue). The p-value was determined by the Mann-Whitney test. [Figure 16D] (D) Steady-state content and mean half-life for the miR (red) and miR* (blue) listed. The mean half-life represents the average of two independent replicates. Individual half-life measurements from two independent experiments (r1 and r2) are shown. Half-life data exceeding the total measurement time are shown as greater than 24 hours. [Figure 16E-F] (E) Comparison of half-life values obtained from two independent biological replicates for 41 miR strands (red) and 20 miR* strands (blue). Pearson correlation coefficient (rp) and associated p-values are shown. (F) MicroRNA stability makes different contributions to the steady-state miRNA content. The half-life values for 40 miRs are shown compared to the steady-state content. The data represent the mean values from two independent biological replicates. [Figure 17A-B]The Argonaut protein profile determines the stability of small RNAs. (A) In Drosophila, miRNAs are preferentially loaded onto Ago1 to form miRISCs. In parallel, a subset of miR* is loaded onto Ago2 to form siRISCs. SiRISC formation involves specific methylation of the 2' position of the 3' terminal ribose of the bound small RNA by Ago2. It is unknown whether the Argonaut protein profile has different effects on the stability of small RNAs. (B) The pie charts show the relative content of miR-NAs and different end-siRNA classes in a small RNA library derived from wild-type Drosophila S2 cells. Results from a standard cloning protocol (unoxidized, top figure) and results from a cloning strategy that enriches the 3' ends of small RNAs with modifications (oxidized, bottom figure) are shown. The proportions of miR and miR* are shown for both libraries. The mean distribution of the seven datasets is shown. The mean library depth is shown. [Figure 17C] (C) The heatmap shows the relative content of miR (red) and miR* (blue) in the library (shown on a gray scale). The relative occurrence ratio in the library indicates that smaller RNAs preferentially bind to AGO1 (green) or AGO2 (red). [Figure 17D-G](D) Western blot analysis of wild-type (wt) S2 cells or S2 cells deficient in Ago2 by CRISPR / Cas9 genomic manipulation (ago2ko). Actin represents the loading control. (E) Relative content of Ago2-enriched miR and miR* in wild-type (wt) S2 cells and ago2ko S2 cells of Drosophila. Median and interquartile range are shown. p-values were determined by Wilcoxon's signed-rank test for paired data. (F) Degradation kinetics of Ago2-enriched small RNA (left) and Ago1-enriched small RNA (right) in standard libraries prepared from s4U-metabolized wild-type S2 cells (wt, black) or ago2ko S2 cells (red), or from wild-type S2 cells (wt oxidized, blue) using a cloning strategy that enriches small RNAs with modified 3' ends. The median and interquartile range of the biphasic or monophasic exponential fit (as specifically described in the text) are shown. The half-life (t1 / 2) obtained by curve fitting is shown. In the case of biphasic pharmacokinetics, the relative contributions of fast and slow pharmacokinetics are shown. (G)ago2ko Half-lives of the 30 most abundant miRNAs (red, Ago1) in S2 cells, or the 30 most abundant miRs and miR* (blue, Ago2) in a small RNA library using a cloning strategy to enrich small RNAs with modified 3' ends. The median and interquartile range are shown. The p-value was obtained by the Mann-Whitney test. [Figure 18] Metabolic labeling of 4-thiouridine in Drosophila S2 cells. Quantitative results of s4U incorporation into total RNA after s4U metabolic labeling over the indicated time in pulse labeling experiments in Drosophila S2 cells. The substitution rate of s4U compared to unmodified uridine, determined by HPLC, is shown. This substitution rate was determined as previously described (Spitzer et al. (2014), Meth Enzymol, Vol. 539, pp. 113-161). The values represent the mean ± SD of three independent replicates. The maximum incorporation rate after 24 hours of labeling is shown. [Figure 19]Iodoacetamide treatment does not affect the quality of small RNA libraries. Small RNA sequencing libraries generated from total RNA of Drosophila S2 cells before and after treatment were mapped to annotated miRNAs, and miRNAs expressed in large quantities (over 100 ppm) were analyzed. (A) Mutation rates were determined for each miRNA from small RNA libraries of total RNA treated with or untreated with iodoacetamide. Tukey box plots show the mutations for each miRNA in percentages. Outliers are not shown. The median frequency of observed mutations is shown for each individual mutation. (B) Content of miRNAs in small RNA libraries prepared from total RNA treated with or untreated with iodoacetamide. Pearson correlation coefficients and associated p-values are shown. (C) Multiplicative changes in the expression of individual miRNAs in small RNA libraries prepared from total RNA treated with or untreated with iodoacetamide. [Figure 20] Frequency of s4U incorporation in metabolically labeled small RNAs. Tukey box plots show the percentage of T>C-converted reads containing one, two, or three T>C mutations for each of 71 miRNAs expressed in large quantities (over 100 ppm) in a small RNA library prepared from size-selected total RNA of Drosophila S2 cells that underwent 24-hour s4U metabolizing. The median percentage of T>C-converted reads is shown. [Figure 21]The s4U metabolic labeling does not affect microRNA biosynthesis or loading. Excessive or under-excessive T>C conversion at specific locations of a given small RNA (left) derived from the 5p or 3p arm of a microRNA precursor, or a given small RNA (right) that constitutes a miR or mirR* strand as defined by selective algonaut loading. Results are from 71 miRNAs expressed in large quantities (over 100 ppm) (corresponding to 35 5p-miRNAs and 36 3p-miRNAs, or 44 miRs and 27 miR*s). Statistical significance of relative occurrence was compared to the population as a whole by the Mann-Whitney test for the described locations. ns, p>0.05; nd, not possible due to limited data points. [Figure 22] Precursor miRNA tailing hinders efficient miRNA biosynthesis. Correlation between pre-miRNA uridylation and T>C mutation rate in a small RNA library prepared by SLAM-seq from ago2ko S2 cells after treatment with s4U over the described time. Pearson's correlation coefficient (rp) and associated p-values are shown. [Figure 23] Chemical modification of 6-thioguanosine (s6G) using iodoacetamide. The chemical reaction between 6-thioguanosine and iodoacetamide (A) and the alkylation efficiency determined by mass spectrometry when 6-thioguanosine is treated with iodoacetamide (B) are shown. [Figure 24A] Alkylation allows for the identification of s6G integration into RNA at single-nucleotide resolution through G-to-A conversion. (A) RNA with or without single-position (p8) 6-thioguanosine (s6G) integration was treated with iodoacetamide (IAA), followed by reverse transcription and gel extraction of the full-length product, and then PCR amplification and high-throughput (HTP) sequencing. [Figure 24B](B) The mutation rates at each position of control RNA (left figure) and s6G-containing RNA (right figure) are shown, with and without iodoacetamide (IAA) treatment using the reverse transcriptase described. The values represent the mean mutation rate ± SD of the three independent replicates. The number of reads sequenced for each replicate (r1~r3) is shown. The nucleotides appearing at p8 are indicated. [Figure 24C] (C) Mutation rates for the described mutations, with and without IAA treatment, using either reverse transcriptase Superscript II (SSII), Superscript III (SSIII), or Quant-seq RT (QS). Mutation rates were averaged for the same nucleotide positions for both s6U-containing and s6U-non-containing RNA oligonucleotides. The p-value (determined by Student's t-test) is shown. ns indicates non-significant (p>0.05). [Figure 25A] Time-resolved mapping of transcriptional responses using SLAM-seq. (A) Example workflow for a SLAM-seq experiment mapping responses after treatment with flavopyridol (300 nM) for 15-60 minutes in K562 cells. [Figure 25B-D] (B) Mapping of all reads and converted (two or more T>C) reads for the low turnover gene GAPDH, with or without flavopyridol treatment. (C) Box plots of flavopyridol-induced expression changes for all reads, reads with one or more T>C conversions, and reads with two or more T>C conversions. Whiskers indicate a range of 5-95%. (D) Schematic diagram of the BCR / ABL effector pathway and kinase inhibitors investigated using SLAM-seq in K562 cells (pretreatment for 30 minutes, s4U labeling for 60 minutes). [Figure 25E] (E) Heatmaps and hierarchical clusterings of the 50 genes that showed the greatest upregulation and downregulation by SLAM-seq in K562 cells treated with the listed inhibitors, and the behavior of these genes at the total mRNA level. [Figure 25F-G](F) Estimated half-life of genes detected to be downregulated to less than half by at least one of the inhibitors listed in (D) in whole mRNA or SLAM-seq. (G) Principal component analysis of SLAM-seq reads from K562 treated with the inhibitors listed in (D). [Figure 26A-D] BETi hypersensitivity is distinctly different from comprehensive transcriptional regulation by BRD. (A) Schematic diagram of the AID-BRD4 knock-in allele and the Tir1 delivery vector SOP. (B) Immunoblotting of BRD4 in K562AID-BRD4 cells transduced using SOP and treated with auxin (100 μM IAA). (C) SLAM-seq response of K562AID-BRD4 cells treated with auxin for 30 minutes before s4U labeling for 60 minutes. (D) SLAM-seq response of K562 cells treated with JQ1 (200 nM) for 30 minutes before s4U labeling for 60 minutes. [Figure 26E-F] (E)(C) Comparison of SLAM-seq responses to JQ1 in K562 cells and MV4-11 cells treated in the same way. R is the Pearson correlation coefficient. (F) Comparison of mean SLAM-seq response and mean CRISPR score, showing the importance of genes in K562, MOLM-13, and MV4-11 cells (14, 15). All genes that were significantly downregulated in all three cell lines are shown (FDR ≤ 0.1). [Figure 26G-H] (G) Principal component analysis of s4U-labeled SLAM-seq reads from MOLM-13 cells treated with JQ1 or the CDK9 inhibitor NVP-2 as performed in (C). (H) Heatmap and hierarchical clustering of Spearman rank correlations between SLAM-seq responses to JQ1 and CDK9 suppression in the described cell lines. [Figure 27A-E]Chromatin context determines BETi hypersensitivity. (A) ROC curves for identifying BETi hypersensitivity genes from a matched control set based on described predictors in K562 cells. (B) Venn diagram showing overlap between BETi hypersensitivity genes and published super-enhancer targets in K562 cells. (C) Sample tracks and super-enhancer annotations from H3K27ac ChIP-seq for selected genes that exemplify category (B). (D) Simplified modeling workflow for classifying BETi hypersensitivity genes based on 214 chromatin signatures within 500 bp or 2000 bp from TSS. (E) Similar ROC curves to (A) for two independent models of BETi hypersensitivity evaluated in a holdout test set, these models based on chromatin signatures. [Figure 27F-G] (F) A figure showing the relative contributions of the five strongest positive and negative predictors to the GLM shown in (E), based on the normalized model coefficients of these predictors. (G) Heatmap and hierarchical clustering of the relative ChIP-seq density of the predictors shown in (F) in the TSS of 125 BETi hypersensitive genes. [Figure 28A-D] MYC is a selective direct activator of cellular metabolism in various types of cancer. (A) Schematic diagram of the MYC-AID allele and Tir1 vector. (B) MYC immunoblotting in K562MYC-AID cells after auxin treatment over the indicated time. (C) SLAM-seq profile after MYC degradation in K562MYC-AID cells (auxin pretreatment for 30 minutes, s4U labeling for 60 minutes). (D) SLAM-seq response of total mRNA and significantly enriched gene sets. [Figure 28E-L](E) MYC immunoblotting in HCT116MYC-AID cells similar to (B). (F) Comparison of SLAM-seq responses in K562MYC-AID cells and HCT116MYC-AID cells. (G) ROC curves of various predictors identifying MYC-dependent genes (FDR≦0.1, log2FC≦-1) from a matched control set in (C). MYC / MAX ChIP, genes ranked by ChIP-seq signals within 2 kbp of TSS; GLM, elastic net GLM based on 214 chromatin profiles. (H) Relative contribution of the strongest positive and negative predictors to GLM in (G). (I) Expression of MYC and SLAM-seq-based MYC-targeted signatures in 672 cancer cell lines. Samples with MYCN or MYCL levels exceeding MYC levels are highlighted. (J, K)GSEA of MYC-target signatures in cell lines from (I) or from AML patients with high or low MYC expression. (L)Expression of MYC-target signatures in 5583 patient samples isolated based on MYC expression levels and cancer type. ****, p<0.0001 (Wilcoxon rank-sum test). [Figure 29A-D]Experimental setup for SLAM-seq to map differences in gene expression. (A) Schematic diagram of the primary and secondary effects of a single target disturbance on mRNA levels in genes with different turnover rates, illustrating the indirect repression of MYC target genes by BETi treatment. (B) Alkylation of 4-thiouridine residues in mRNA with iodoacetamide. (C) Contribution of background error rate to reads with one or more or two T>C conversions in a SLAM-seq experiment with a labeling time of 60 minutes. Background signal was measured by performing alkylation and sequencing of mRNA from s4U-treated and untreated cells in parallel. The box plot shows the error rates of genes with high mRNA turnover (top 10%), moderate turnover (mid 45-55%), and low mRNA turnover (bottom 10%) estimated from the proportion of labeled reads in the s4U-treated library. Whiskers represent a range of 5-95%. (D) Recovery of newly synthesized mRNA by SLAM-seq is shown as a function of T content per read. The estimation is based on a labeling efficiency of 11.4% per T in newly synthesized mRNA. The histogram shown above shows the T content of 3'UTR reads for different read lengths. [Figure 29E] (E) Confirmation of MAPK·AKT pathway suppression in K562 cells treated with the described kinase inhibitors for the SLAM-seq samples shown in Figure 25E. [Figure 30A-C] Direct JQ1 response in myeloid leukemia cell lines. (A) Box plot summarizing the overall changes in gene expression measured by SLAM-seq for K562 cells treated with JQ1 in Figure 26D. (B, C) Primary response to JQ1 treatment mapped by SLAM-seq for cell lines MOLM-13 and MV4-11, as shown in Figure 2D. [Figure 30D-G](D)(A) A summary of the overall changes in gene expression when MOLM-13 cells and MV4-11 cells were treated with JQ1. (E, F) Pairwise comparisons of changes in gene expression when the described cell lines were treated with JQ1. (G) Comparison of SLAM-seq responses to JQ1 and rapid BRD4 degradation in K562 cells, shown in Figures 26B and 26C. Genes induced by JQ1 treatment and genes induced by BRD4 degradation in MOLM-13, MV4-11, and K562 cells are highlighted in blue. **** indicates that p<0.0001 when the median is deviated from 0, calculated by Wilcoxon's signed-rank test. [Figure 31A-B] Synergistic effects of CDK9 inhibition and BET bromodomain inhibition at the cellular and transcriptional levels. (A) Results of the inhibition of MOLM-13 and OCI-AML3 cell proliferation by the stated doses of NVP-2 and JQ1 measured for 3 days using the CellTiter-Glo luminescent cell survival assay. (B) Figure showing the synergistic effect of NVP-2 and JQ1 in (A) as an excess due to Bliss additivity. [Figure 31C] Primary response to JQ1 treatment, NVP-2 treatment, and combination treatment at the described doses, performed for 30 minutes in MOLM-13 cells before (C)s4U labeling for 60 minutes. [Figure 31D-F] (D) Pairwise comparison of the response to JQ1 and the response to the intermediate dose of NVP-2 (6 nM) shown in (C). R represents the Pearson correlation coefficient. (E) Overall change in mRNA output when OCI-AML3 cells were treated with NVP-2 and JQ1 as in (C). The dots indicate the proportion of s4U-labeled reads to total reads in SLAM-seq for three independent replicates. (F) Principal component analysis of SLAM-seq reads from OCI / AML3 cells treated with NVP-2 and JQ1 in (E). % indicates the proportion of total variance explained by each principal component. [Figure 32A-B]Derivation and validation of factors predicting JQ1 supersensitivity. (A) Workflow for selecting a balanced set of JQ1 supersensitive genes and control genes based on SLAM-seq response after 90 minutes of treatment with JQ1 (30 minutes of pretreatment + 60 minutes of s4U labeling). Control genes were selected by repeatedly subsampling and validating that baseline expression was the same using the Kolmogorov-Smirnov test (KS). (B) ROC curves for predicting JQ1 supersensitive genes by comparing them with control genes using superenhancers (SEs). Each SE was assigned to the gene with the closest TSS, and the genes were sorted by the rank of the superenhancer. [Figure 32C-D] (C) Workflow diagram for deriving factors for classifying BETi hypersensitivity in K562 cells based on TSS, as shown in Figure 27E. SVM - Support Vector Machine, GLM - Generalized Linear Model derived by Elastic Net Regularization, GBM - Gradient Boosting Model. [Figure 33A] Characterization of the GLM predicting JQ1 hypersensitivity in K562 cells. (A) Bar graph of coefficients for all predictors contributing to the GLM shown in Figure 27E. The adjacent table lists each predictor, the corresponding publicly available ChIP-seq track discriminant (see also Table S1), and whether the signal was measured within 500 bp or 2000 bp from the transcription start site. [Figure 33B] (B) Heatmaps and hierarchical clustering of JQ1 supersensitive genes and control genes based on the relative ChIP-seq signals of five unique positive and negative strongest predictors from (A). [Figure 34A-B]Primary response to MYC degradation in endogenously tagged MYC-AID cell lines. (A) Volcano plot of SLAM-seq response of auxin-treated K562MYC-AID cells shown in Figure 4C. The boxes indicate the number of downregulated and upregulated genes (FDR ≤ 0.1) and the total number of genes with more than twofold deregulation. The values on the vertical axis are limited to 20 for readability. (B) Figure showing the SLAM-seq response of HCT116MYC-AID cells in (A) for genes related to the listed GO terms. [Figure 34C-D] (C) A Volcano plot showing the results of SLAM-seq of HCT116MYC-AID cells treated with auxin as shown in (A). (D) The SLAM-seq response to rapid MYC degradation in K562MYC-AID cells is shown, grouped by GO term for genes encoding subunits of different RNA polymerase complexes. [Figure 35A-B] Chromatin-based prediction of MYC-dependent gene expression. (A) ROC curves measuring the ability to distinguish MYC-dependent genes (FDR ≤ 1, log2FC ≤ -1) from genes that do not respond to MYC degradation (FDR ≤ 0.1, -0.2 ≤ log2FC ≤ 0.2) for the five classification factors in Figure 28C. The model was derived as shown in Figure 32, trained on a test set of 802 genes, and evaluated on a test set of 268 genes. (B) ROC curves evaluating the performance of the GLM shown in (A) on an additional set of holdout genes for validation. [Figure 35C-E](C) Coefficients of all GLM predictors shown in (A). (D) ROC for predicting MYC-dependent genes based on the presence of a MYC ChIP-seq peak within 2000 bp of the gene's TSS. As shown in Figure 32, for each gene, the TSS that contributes most to the level of cellular mRNA was selected based on the CAGE-seq signal. SPP peak - peak called by SPP; IDR peak - peak exceeding 2%, which is the detection irreproducibility threshold. (E) Venn diagram of the 7135 genes classified as MYC-binding genes and MYC-dependent genes as described above. [Figure 36] Expression of MYC orthologues in human cancer cell lines. (A) Comparison of MYC expression and MYCL expression in RNA-seq profiles from 672 cancer cell lines. (B) Comparison of MYC expression and MYCN expression in the same cancer cell lines as (A). [Figure 37] The pH-dependent frequency of base pairing during reverse transcription of 5-bromouridine (5BrU)-containing RNA oligonucleotides was measured by sequencing. (a) Tautomer morphs of 5BrU exhibit different pH-dependent base pairing properties with respect to adenine or guanine. (b) An experimental outline for detecting the pH-dependent effect on the misintegration of nucleotides at 5BrU-modified positions in the context of RNA. (c) Conversion rates for specific conversions observed at 5BrU-modified nucleotide positions after reverse transcription under the described pH conditions. Increased conversion rates are shown compared to pH 7 conditions. The number of independent measurements was n=3, and the bars represent the mean convergence rate ± SD. P-values (unpaired, parametric Student's t-test) are shown. ns, not significant (p>0.05); *, p<0.05; **, p<0.01; ***, p<0.001; ****, p<0.0001. [Examples]
[0077] Example 1: Materials and Method
[0078] s 4Carboxamide methylation of U
[0079] Unless otherwise specified, carboxyamidomethylation is performed using 1 mM 4-thiouracil (SIGMA), 800 μM 4-thiouridine (SIGMA), s 4 The reaction was carried out under standard conditions (50% DMSO, 10 mM iodoacetamide, 50 mM sodium phosphate buffer, pH 8, 50°C for 15 minutes) using 5 to 50 μg of total RNA prepared from metabolic labeling experiments. The reaction was stopped by adding excess DTT.
[0080] Adsorption measurement
[0081] Unless otherwise specified, 1 mM 4-thiouracil was incubated under optimal conditions (10 mM iodoacetamide, 50% DMSO, 50 mM sodium phosphate buffer, pH 8, at 50°C for 15 minutes). The reaction was stopped by adding 100 mM DTT, and the adsorption spectrum was measured using a Nanodrop 2000 instrument (Thermo Fisher Scientific). The baseline was then subtracted from the adsorption at 400 nm.
[0082] mass spectrometry
[0083] 40 nanomoles of 4-thiouracil or 6-thioguanosine were reacted at 50°C for 15 minutes under standard reaction conditions (50 mM sodium phosphate buffer, pH 8; 50% DMSO) in the absence or presence of 0.05, 0.25, 0.5, or 5 micromoles of iodoacetamide. The reaction was stopped with 1% acetic acid. The acidified samples were separated on an Ultimate U300 BioRSLC HPLC system (Dionex; Thermo Fisher Scientific) using a Kinetex F5 pentafluorophenyl column (150 mm × 2.1 mm; 2.6 μm, 100 Å; Phenomenex) at a flow rate of 100 μl / min. Nucleosides were electrospray ionized and then analyzed online using a TSQ Quantiva mass spectrometer (Thermo Fisher Scientific). The SRM results at that time were as follows: 4-thiouridine m / z 260 → 129, alkylated 4-thiouridine m / z 318 → 186, 6-thio-guanosine m / z 300 → 168, alkylated 6-thio-guanosine m / z 357 → 225. The data were interpreted using the Trace Finder software suite (Thermo Fisher Scientific) and manually verified.
[0084] Primer extension assay
[0085] The primer extension assay was performed essentially as previously described by Nilsen et al. (Cold Spring Harb Protoc. 2013, pp. 1182-1185). Briefly, the template RNA oligonucleotide (5L-let-7-3L or 5L-let-7-s) was used. 4Up9-3L (Dharmacon; see table for sequence) was deprotected according to the manufacturer's instructions and purified by denatured polyacrylamide gel elution. 100 μM of purified RNA oligonucleotide was treated with 10 mM iodoacetamide (+IAA) or EtOH (-IAA) at 50°C for 15 minutes under standard reaction conditions (50% DMSO, 50 mM sodium phosphate buffer, pH 8). After stopping the reaction with the addition of 20 mM DTT, ethanol precipitation was performed. γ- 32 After radiolabeling the 5' end of RT primers (see table for sequences) using P-ATP (Perkin-Elmer) and T4-polynucleotide kinase (NEB), denaturation polyacrylamide gel purification was performed. In 2× annealing buffer (500 mM KCl, 50 mM Tris, pH 8.3) in a PCR instrument, 640 nM γ-nucleotides were added. 32 P-RT primer in 400 nM 5L-let-7-3L or 5L-let-7-s 4The samples were annealed in Up9-3L (3 minutes at 95°C, 30 seconds at 85°C, transfer rate 0.5°C / sec, 5 minutes at 25°C, transfer rate 0.1°C / sec). Reverse transcription was performed using either Superscript II (Invitrogen), Superscript III (Invitrogen), or Quant-seq RT (Lexogen) as recommended by the manufacturer. For the dideoxynucleotide reaction, ddNTP at a final concentration of 500 μM was added to the RT reaction product (as instructed). After the reaction was complete, the RT reaction product was resuspended in formaldehyde hydroloading buffer (Gel Loading Buffer II, Thermo Fisher Scientific) and subjected to 12.5% denatured polyacrylamide gel electrophoresis. The gel was dried, exposed to a storage fluorescence screen (Perkin-Elmer), images were acquired using a Typhoon TRIO variable-mode imaging system (Amersham Biosciences), and quantification was performed using ImageQuant TL v7.0 (GE Healthcare). To analyze the elimination, the signal intensity at the p9 position was normalized for each reaction to the preceding elimination signal intensity (bg, Figure 4B). 4 U-containing RNA oligonucleotides and s 4 The values representing the change in elimination signal (+IAA / -IAA) for U-free RNA oligonucleotides were compared for the described reverse transcriptases. 6-thioguanosine modification, replacing 4-thiouridine, was analyzed using the corresponding settings.
[0086] [Table 1]
[0087] [Table 2]
[0088] [Table 3]
[0089] s 4 U or s 6 HPLC analysis of RNA labeled with G
[0090] s to total RNA after metabolic labeling 4 U or s 6 The incorporation of G was analyzed in the manner previously described by Spitzer et al. (Meth Enzymol. Vol. 539, pp. 113-161 (2014)).
[0091] Cell survival assay
[0092] The day before the experiment, 5000 mES cells were seeded into 96 wells. (The different concentrations of s) (as described) were used. 4 Cells were treated with a culture medium containing u for 12 or 24 hours. Cell viability was assessed using the CellTiter-Glo® Luminescent Cell Viability Assay (Promega) according to the manufacturer's instructions. Fluorescence signals were measured using Synergy Gen5 software (v2.09.1).
[0093] Cell culture
[0094] Mouse embryonic stem (mES) cells (clone AN3-12) were obtained from Haplobank (Elling et al., WO 2013 / 079670) and cultured in 15% FBS (Gibco), 1× penicillin-streptomycin solution (100 U / ml penicillin, 0.1 mg / ml streptomycin, SIGMA), 2 mM L-glutamine (SIGMA), 1× MEM non-essential amino acid solution (SIGMA), 1 mM sodium pyruvate (SIGMA), 50 μM 2-mercaptoethanol (Gibco), and 20 ng / ml LIF (homemade). The cells were maintained at 37°C with 5% CO2 and subcultured every two days.
[0095] RNA modification to alter sequencing ("SLAM-seq")
[0096] The day before the experiment, 10 mES cells were placed in a 10 cm dish. 5 Cells were seeded at a density of cells / ml. mES cells were cultured in standard medium, but from a 500 mM storage solution in water. 4 U or s 6 By incubating points in different culture media with 100 μM of G (SIGMA Corporation) added, s 4 Metabolic labeling was performed. During metabolic labeling, s 4 U or s 6 The culture medium containing G was replaced every 3 hours. For the uridine tracking experiment, s 4 U or s 6 The medium containing G was discarded, the cells were washed twice with 1×PBS, and incubated with standard medium supplemented with 10 mM uridine (SIGMA). The cells were directly lysed in TRIzol® (Ambion), and RNA was extracted according to the manufacturer's instructions, with the exception of adding 0.1 mM DTT during isopropanol precipitation. The RNA was resuspended in 1 mM DTT. 5 μg of total RNA was treated with 10 mM iodoacetamide under optimal reaction conditions, followed by ethanol precipitation to prepare a Quant-seq 3'-terminal mRNA library (Moll et al., Nat Methods Vol. 11 (2014); WO 2015 / 140307).
[0097] Preparation of RNA libraries
[0098] A standard RNA sequence library was prepared using the NEBNext® Ultra® Directional RNA Library Prep Kit for Illumina® (NEB) according to the manufacturer's instructions. A Cap-seq library was prepared as previously described by Mohn et al. (Cell. Vol. 157, pp. 1364-1379 (2014)), but with the difference that ribosomal RNA was depleted using the Magnetic RiboZero Kit (Epicenter) before fragmentation. Messenger RNA 3' ends were sequenced using the Quant-seq mRNA 3' End Library Preparation Kit (Lexogen) according to the manufacturer's instructions.
[0099] Data Analysis
[0100] Gel images were quantified using ImageQuant v7.0a (GE Healthcare). Curve fitting was performed using Prism v7.0 (GraphPad) or R (v2.15.3) according to the integrated rate law for first-order reactions. Statistical analysis was performed using either Prism v7.0a (GraphPad), Excel v15.22 (Microsoft), or R (v2.15.3).
[0101] Bioinformatics
[0102] For sequencing analysis of synthetic RNA samples (Figure 5), the barcoded library was demultiplexed using Picard Tools BamIndexDecoder v1.13, which allows zero mismatches in the barcode. The resulting files were converted to fastq format using picard-tools SamToFastq v1.82. The adapter was trimmed using Cutadapt v1.7.1 (allowing 10% mismatches in the adapter sequence by default), and filtered to search for sequences of length 21 nt. The obtained sequences were aligned to the mature dme-let-7 sequence (5'-TGAGGTAGTAGGTTGTATAGT-3', sequence ID number 7) using bowtie v0.12.9, which allows three mismatches, and converted to bam format using samtools v0.1.18. Sequences containing ambiguous nucleotides (N) were removed. The remaining aligned reads were converted to pile-up format. Finally, the percentage of each mutation at each position was extracted from the pile-up. The output table was analyzed and plotted using Excel v15.22 (Microsoft) and Prism v7.0a (GraphPad).
[0103] To analyze mRNA 3' end sequencing data, the barcoded library was demultiplexed using Picard Tools BamIndexDecoder v1.13, which allows one mismatch within the barcode. Adapters were clipped using Cutadapt v1.5, and reads were filtered by size to find those with 15 nucleotides or more. Reads were aligned to the mouse genome mm10 using STAR aligner v2.5.2b. Alignments were filtered to find those with an alignment score of 0.3 or higher, and alignments with an alignment identity of 0.3 or higher were normalized to read length. Only alignments with 30 or more matches were reported. Only chimeric alignments with 15 bp or more overlap were accepted. Two-pass mapping was used. Introns shorter than 200 kb were filtered, and alignments containing non-classical junctions were filtered. Alignments with a mismatch ratio of 0.1 or higher to the mapped base, or alignments with a maximum of 10 mismatches, were filtered. The maximum number of gaps allowed at the splice joint for 1, 2, 3, and N reads was set to 10 kb, 20 kb, 30 kb, and 50 kb, respectively. For (1) non-classical motifs, (2) GT / AG motifs and CT / AC motifs, (3) GC / AG motifs and CT / GC motifs, and (4) AT / AC motifs and GT / AT motifs, the minimum overhang lengths for the splice joints on both sides were set to 20, 12, 12, and 12, respectively. Using "pseudo" splice filtering, the maximum number of multiple alignments allowed in a single read was set to 1. Exon reads (Gencode) were quantified using FeatureCounts.
[0104] For capped gene expression analysis (CAGE), the barcoded library was demultiplexed using Picard Tools BamIndexDecoder v1.13, which allows one mismatch within the barcode. The first 4 nt of each read was trimmed using seqtk. Reads were screened for ribosomal RNA by alignment with known rRNA sequences (Ref-Seq) using BWA (v0.6.1). Reads with rRNA subtracted were aligned to the mouse genome (mm10) using TopHat (v1.4.1). The maximum multi-hit was set to 1, the segment length to 18, and the segment-mismatch to 1. In addition, the gene model was provided as a GTF (Gencode VM4).
[0105] To analyze mRNA 3' end sequencing (Quant-seq) datasets, reads were demultiplexed using Picard Tools BamIndexDecoder v1.13, which allows one mismatch within the barcode. The five nucleotides at the 5' end of the demultiplexed reads were trimmed. Reads were aligned to the mouse genome mm10 using SLAMdunk (Neumann and Rescheneder, t-neumann.github.io / slamdunk / ; Herzog et al., Nature Methods Vol. 14, pp. 1198-1204 (2017)), and the alignments in the annotated 3' UTR (Gencode) were counted. Briefly, SLAMdunk is a flexible and fast read mapping program that relies on NextGenMap and is custom-made using a fit scoring scheme that eliminates the T>C mismatch penalty for the mapping process. The T>C-containing and T>C-free reads aligned to the 3'UTR were quantified, and s 4 The content of U-labeled transcripts and unlabeled transcripts was derived, respectively.
[0106] To analyze transcriptional output, high-throughput sequencing data was aligned to the mouse genome mm10 using SLAMdunk. For all genes, the number of normalized reads (in cpm; "steady-state expression") and the number of normalized reads containing one or more T / C mutations (in cpm; "transcriptional output") were obtained. Mitochondrial (Mt-) genes and predicted (GM-) genes were excluded from the analysis. Background T / C reads (s 4 The number of T / C reads observed without U labeling was subtracted from the T / C reads at 1 hour, and an expression threshold of over 5 cpm was set for the average value of "steady-state expression". To identify genes with high transcriptional output, log 10 (steady state occurrence) vs log 10 After plotting the transcription output (number of genes: 6766) and fitting it using linear regression, the equation Y = 0.6378 × X - 1.676 was obtained. For each gene, the distance to the fitted line (ΔY) was calculated as ΔY = transcription output (cpm) - (0.6378 × steady-state expression (cpm) - 1.676). Genes with "high transcription output" were defined as ΔY > 0.5 (number of genes: 828). Genes with "high expression" were defined as those with steady-state CPM > log 10 (2.15) was defined (number of genes: 825). To predict the transcription factor network that governs each class of genes, Ingenuity Pathway Analysis (Qiagen) v27821452 was used for "high transcription output" or "high expression" genes. Ingenuity Pathway Analysis is a web-based application that enables biologists to discover, visualize, and explore therapeutic networks that are important to their experimental results. For a detailed explanation of Ingenuity Pathway Analysis, please visit www.Ingenuity.com. The top 5 predicted upstream regulators are shown.
[0107] To predict pathways with "high transcription output" or "highly expressed" genes, the online tool Enrichr was used with two gene classes as input. The top five predicted pathways are displayed.
[0108] Example 2: Thiol-linked alkylation for RNA metabolic sequencing
[0109] As a way to verify the principle, we used s to elucidate the RNA expression dynamics in cultured cells. 4 U or s 6 As an example of a derivatization strategy to avoid the need to biochemically isolate G-labeled or unlabeled RNA species, we selected thiol nucleotides (Figure 1). This strategy is based on a well-established metabolic labeling approach but avoids the ineffective and time-consuming biotinylation step. This strategy is s 4 It includes a short chemical treatment protocol that involves modifying U-containing RNA with iodoacetamide. Iodoacetamide is a sulfhydryl-reactive compound, and s 4 U or s 6 When reacted with G, it produces bulky groups at the base-pairing interface (Figure 1). 4 If a bulky group is present at the U integration site, when combined with a well-established RNA library preparation protocol, G may be mistakenly incorporated during reverse transcription (RT) to a degree that is specific and quantifiable, but this does not affect the processing performance of RT. Therefore, s 4 U or s 6Sequences containing G can be identified at single-nucleotide resolution by calling the transition from T to C, by sequence comparison, i.e., by bioinformatics in a high-throughput sequencing library. Importantly, no enzyme is known to convert uridine to cytosine, and similar error rates (i.e., T>C conversion) in datasets for high-throughput RNA sequencing using the Illumina HiSeq 2500 platform are extremely low, occurring at a frequency of less than 1 in 10,000. This approach is called "SLAM-seq," an abbreviation for thiol(SH)-linked alkylation of RNA for the metabolic sequencing, which is its most preferred embodiment.
[0110] SLAM-seq is based on nucleotide analog derivatization chemistry, enabling the detection of 4-thiouridine incorporation events derived from metabolism and labeling in RNA species with single-nucleotide resolution via high-throughput sequencing. We demonstrate that this novel method accurately measures RNA polymerase II-dependent polyadenylated transcription output in mouse embryonic stem cells and reproduces comprehensive post-transcriptional gene regulatory signatures. The present invention provides a scalable, highly quantifiable, cost-effective, and time-efficient method for rapid, whole-transcriptome analysis of RNA expression dynamics with high temporal resolution.
[0111] s 4 For U derivatization, an example of an effective first thiol-reactive compound is nucleophilic substitution (S N 2) Iodoacetamide (IAA) was used to add a carboxyamide methyl group to the thiol group as a result of the reaction (Figure 2A). A similar reaction can be performed with other thiolated nucleic acid bases (s 6 This occurs in G, etc. (Figure 23A). 4To quantitatively monitor the efficiency of U derivatization as a function of various parameters (i.e., time, temperature, pH, IAA concentration, DMSO), the characteristic absorption spectrum of 4-thiouracil (approximately 335 nm, which shifts to approximately 297 nm upon reaction with iodoacetamide (IAA)) was monitored (Figures 2B-2K) (45). Under optimal reaction conditions (10 mM IAA; 50 mM NaPO4, pH 8; 50% DMSO; 50°C; 15 minutes), the absorption at 335 nm was reduced to 1 / 50th compared to untreated 4-thiouracil, resulting in an alkylation rate of at least 98% (Figures 2L and 2M). (Note: Since 4-thiouracil and its alkylated derivatives partially overlap, the conversion rate may be underestimated.) Mass spectrometry analysis of thiol-specific alkylation in the context of ribose (i.e., 4-thiouridine or 6-thioguanosine) confirmed near-perfect derivatization efficiency (Figures 3 and 23B).
[0112] s 4 U or s 6 In the quantitative recovery of the G integration event, reverse transcriptase is used to remove alkylated s without elimination. 4 It is assumed that the U residue group is passed through. 4 U alkylation or s 6 To clarify the effect of G alkylation on the processing performance of reverse transcriptase, 4 U or s 6 Synthetic RNA containing only one G integration (see Table in Example 1 for sequence information) was used to test three commercially available reverse transcriptases (RTs) (Superscript II, Superscript III, Quant-seq RT) using a primer extension assay (Figure 4A). When normalized to the background elimination signal, the same sequence was observed, but s 4 U or s 6 When compared to oligosaccharides that do not contain G, s 4 U alkylation or s 6No significant effect of G alkylation on the processing performance of RT was observed (Figure 4B, Figure 4C, Figure 24B). We concluded that alkylation does not cause premature termination of reverse transcription.
[0113] s 4 U alkylation and s 6 To evaluate the effect of G alkylation on nucleotide incorporation directed by reverse transcriptase, the full-length products of the primer extension reaction were isolated, the cDNA was amplified by PCR, and high-throughput sequencing was performed on the library using an Illumina HiSeq 2500 instrument (Figure 5A and Figure 24A). As expected, s 4 All three RTs in the U-free control RNA or s 6 G-free control RNA accurately reverse transcribed uridine, and the average mutation rate was less than 10 -2 even in the absence of treatment with iodoacetamide (Figure 5B and Figure 24B, left panel). In contrast, s 4 the presence of U promoted a constant 10% - 11% conversion from T to C even without alkylation, probably due to s 4 changes in the base pairing of the U tautomer (Figure 5B, upper right panel). s 6 In the case of G, a conversion from G to A was observed (Figure 24B, right panel). Notably, alkylation of s 4 U by iodoacetamide treatment increased the conversion from T to C by 8.5-fold, and the mutation rate exceeded 0.94 in all RTs examined (lower right panel of Figure 5B). A signal-to-noise ratio of over 940:1 was obtained compared to the sequencing errors (less than 10 -3 reported in the Illumina high-throughput sequencing dataset. Importantly, no significant effect of iodoacetamide treatment on the mutation rate of thiol-free nucleotides was observed (Figure 5C and Figure 24C). We found that reverse transcription following iodoacetamide treatment introduced s 4 U or s 6We concluded that this method allows for the quantitative identification of G incorporation at single-nucleotide resolution, while remaining unaffected by the sequence information of thiol-free nucleotides.
[0114] Example 3: Incorporation of modified nucleotides into metabolically labeled mRNA in mES cells
[0115] Mouse embryonic stem cells at various concentrations 4 U after 12 hours or 24 hours 4 We investigated the ability to tolerate U metabolic RNA labeling (Figure 6A). As previously reported, high concentrations of s 4 U impairs cell survival, and EC 12 or 24 hours after labeling 50 These concentrations are 3.1 mM or 380 μM, respectively (Figure 6A). Therefore, the labeling conditions are 100 μM s 4 U was used. At this concentration, there was no serious effect on cell survival. Under these conditions, s was added to the total RNA preparation 3 hours, 6 hours, 12 hours, and 24 hours after labeling. 4 A steady increase in U inclusion was detected, along with a steady decrease at 3, 6, 12, and 24 hours of uridine tracking (Figure 6B). As expected, the dynamics of inclusion followed a single exponential function, s 4 The maximum average inclusion rate of U was 1.78%, which is one s for every 56 uridines in the total RNA. 4 This corresponds to the incorporation of U (Figure 6C). These experiments showed that s in mES cells 4 U-labeling conditions are established. These conditions can then be used to measure RNA biosynthesis and turnover under undisturbed conditions.
[0116] This method is suitable for high-throughput sequencing datasets. 4 To verify the ability to identify U embedded events, a 24-hour s 4An mRNA 3'-end library was generated using total RNA prepared from cells cultured after U-metabolism RNA labeling (using Lexogen's Quant-seq 3' mRNA sequencing library preparation kit) (Figure 7) (Moll et al., cited above). The Quant-seq 3' mRNA-Seq Library Prep Kit generates a library of sequences close to the 3' end of polyadenylated RNA that is compatible with Illumina, as exemplified for the gene Trim28 (Figure 8A). Unlike other mRNA sequencing protocols, only one fragment is generated per transcript, eliminating the need to normalize reads to gene length. As a result, accurate gene expression values with high strand specificity are obtained.
[0117] Furthermore, a library that is easy to sequence can be generated in just 4.5 hours, with an execution time of approximately 2 hours. Combining Quant-seq with the present invention makes it easy to accurately determine the mutation rate in transcript-specific regions because the library has a low degree of sequence heterogeneity. In fact, s 4 When a library of U-modified RNA was generated from total RNA of mES cells using the Quant-seq protocol 24 hours after U-metabolism RNA labeling, a stronger accumulation of T>C conversion was observed compared to a library prepared from total RNA of unlabeled mES cells (Figure 8B). To confirm this observation across the entire transcriptome, reads were aligned to annotated 3' UTRs, and the occurrence of mutations in each UTR was examined (Figure 9). 4 In the absence of U metabolite RNA labeling, a median mutation rate of less than 0.1% was observed for any given mutation. This mutation rate is consistent with the sequencing error rate reported in Illumina. 4 24 hours after U metabolite RNA labeling, the T>C mutation rate was statistically significant (p<10). -4A 25-fold increase was observed (Mann-Whitney test), while all other mutation rates remained below the expected sequencing error rate (Figure 9). More specifically, 2.56% s after 24 hours of labeling. 4 A median U was observed. This is because every 39 uridines 4 This corresponds to the incorporation of one U. (Note: This median incorporation frequency for mRNA is higher than the value estimated by HPLC for total RNA [Figure 6C]. This is because stable non-coding RNA species (such as rRNA) are likely to be strongly over-represented in total RNA.) From these analyses, it was found that, using a novel method, in cultured cells, s 4 U-metabolic RNA labeling followed by s of mRNA 4 It is confirmed that the U embedded event will become apparent.
[0118] As previously reported (Eidinoff et al., Science. Vol. 129, pp. 1550-1551 (1959); Jao et al., PNAS Vol. 105, pp. 15779-15784 (2008); Melvin et al., Eur. J. Biochem. Vol. 92, pp. 373-379 (1978); Woodford et al., Anal. Biochem. Vol. 171, pp. 166-172 (1988)), other modified nucleotides (s 6 The same incorporation results are expected with (G or 5-ethinyluridine).
[0119] Example 4: Use of 5-bromouridine and pH-dependent base pairing frequency during reverse transcription
[0120] Yu et al. (The Journal of Biological Chemistry, Vol. 268:21, pp. 15935-15943, 1993) showed that the base analog bromouracil mistakenly pairs with G (guanine) during polymerization as a function of pH. Furthermore, it has been shown that 5BrU is taken up by cells, phosphorylated, and incorporated into nascent RNA (Larsen et al., Current Protocols in Cytometry, Vol. 12 (7.12): pp. 7.12.1-7.12.11, 2001). We demonstrate that 5BrU labeling can be identified by sequencing of pH variant NGS library preparations using both of these methods. 100 picomoles of synthetic RNA oligonucleotide containing 5BrU modification at only one central position were used. RNA sequence 5'- ACACUCUUUCCCUACACGACGC UCUUCCGAUCU- UGAGGUAGU[5BrU]AGGUUGUAUAGUA GAUCGGAAGAGCACACGUCUC -3' (sequence ID number 8) has two underlined linker sequences (used for reverse transcription and amplification) and [5BrU] (5-bromouridine label at the central position). Reverse transcription was performed using Superscript II (Thermo Fisher Scientific) according to the manufacturer's instructions, with RT DNA oligonucleotide primers (5'-GTGACTGGAGTTCAGACGTGTGCTCTTCCGAT-CT-3', sequence ID number 9) and 5× RT buffer (their pH adjusted to pH 7, pH 8, and pH 9, respectively). After reverse transcription, DNA oligonucleotide Solexa PCR Fwd (5'- AATGATACGGCG ACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTCCGATCT-3', sequence ID number 10) and Solexa IDX rev(5'-CAAGCAGAAGACGGCATACGAGATNNNNNN-GTGACTGGAGTTCA Using the sequence GACGTGTGCTCTTCCGATCT-3', sequence ID number 11 (NNNNNN indicates the barcode - nucleotide position), a 1 picomole reverse transcript was amplified by PCR using the KAPA Real-Time Library Amplification Kit (KAPA Biosystems) according to the manufacturer's instructions. The amplified library was sequenced by high-throughput sequencing using the Illumina MiSeq platform. The conversion rate of 5BrU nucleotides was determined by counting the frequencies of nucleotides A (adenine), G (guanine), and C (cytosine), rather than using T (thymine) reads, which are expected to be the majority at the 5BrU position.
[0121] The pH-dependent conversion rate of 5BrU or T to A (T>A conversion) in the final read increased from pH 8 to pH 9 during reverse transcription, resulting in a rate of 3 × 10⁻¹⁰ at the background pH of 7. -4 Compared to the conversion rates shown, these rates increase by 1.1 times and 1.4 times, respectively (Figure 37C). The T>G conversion rate increases by 1.2 times and 1.9 times, respectively, compared to the lower conversion rate of 1 × 10⁻⁴ at the background pH of 7. In contrast, the T>C conversion rate increases by 3 × 10⁻⁴ at the background pH of 7. -4 The conversion rate increases by 2.2 times and 4.3 times compared to the previous rate. The signature that the conversion rate changes depending on pH is characteristic of 5BrU in RNA and can be used to identify the number of 5BrU incorporated into nascent RNA during transcription.
[0122] Example 5: Measurement of polyadenylated transcription output in mES cells
[0123] Short s after mRNA 3' end sequencing 4 To verify whether U-pulse labeling accurately indicates polyadenylated transcription output, mES cells were subjected to a short 1-hour s pulse. 4After applying a U pulse, total RNA was extracted, and an mRNA 3'-terminal library was prepared (Figure 10A). The library generated by Quant-seq was mapped to annotated 3' UTRs in the mouse genome, and the presence of T>C conversions representing newly transcribed RNA was analyzed (Figure 10A). Comparing the relative content of newly transcribed (i.e., containing T>C conversions) polyadenylated transcripts with that of polyadenylated transcripts in the steady state (i.e., containing both T>C and non-T>C conversions), a subset of transcripts consisting of 828 over-existing genes was observed among the newly transcribed RNA (Figure 10B). The top over-existing transcripts included clusters of mES cell-specific microRNAs (miR-290 cluster and miR182 cluster). These are thought to be particularly short-lived because they are rapidly degraded when microRNA hairpins are excised by Drosha in the nucleus, and therefore do not accumulate at high levels in the steady state (Figure 10B). To more systematically characterize the de novo transcription output measured by this novel method, gene list enrichment analysis was performed on 828 overexpressing genes among the newly transcribed RNAs, and on 825 genes detected in steady state by conventional mRNA 3'-end sequencing for predicted underlying transcription factors (using Ingenuity Pathway Analysis, www.Ingenuity.com) and related molecular pathways (using Enrichr) (Figure 10C). As expected, high levels of expression in steady state did not predict the transcription factor network related to pluripotency, with the majority being related to housekeeping pathways (ribosomal proteins, mRNA processing, electron transport, etc.). In contrast, de novo transcript analysis using the present invention successfully predicted the pluripotency network, as well as key mES cell-specific transcription factors (including Oct4 (POU5F1), NANOG, and SOX2) (Figure 10C). We found that short s 4We concluded that combining U-pulse labeling with RNA modification and mRNA 3' end sequencing allows for the separation of immediate transcriptional output from transcript stabilization effects, thus providing a rapid and scalable method for studying the regulation of transcriptional genes.
[0124] Example 6: Measurement of mRNA Transcript Stability
[0125] To investigate whether the stability of mRNA transcripts can be measured using the method of the present invention, we added s to RNA in mES cells. 4 After labeling with U for 24 hours, total RNA was prepared at various time points (0 min, 15 min, 30 min, 1 hour, 3 hours, 6 hours, 12 hours, 24 hours) using excess thiol-free uridine, followed by U modification and mRNA 3' end sequencing (Figure 11A). The library was again mapped to annotated 3' UTR in the mouse genome and analyzed to search for the presence of T>C conversion representing older transcripts (Figure 11A). For 9430 genes, the transcripts containing T>C conversion were normalized to their steady-state content (which remained constant during analysis), and their changes over time were analyzed as a whole, revealing that the median mRNA half-life was 4 hours (Figure 11B). As expected, the half-lives of individual transcripts varied by more than an order of magnitude (Figure 11C).
[0126] Controlling mRNA stability is crucial for the temporal order of gene expression. The novel method replicates key fundamental principles because regulatory transcripts related to GO terms such as "transcriptional regulation," "signal transduction," "cell cycle," and "development" exhibited significantly shorter half-lives compared to housekeeping transcripts related to GO terms such as "extracellular matrix," "metabolic processes," and "protein synthesis" (Figure 11D). In conclusion, the present invention enables accurate assessment of mRNA stability, providing a convenient method for studying post-transcriptional gene regulation. As demonstrated here, this method replicates the entire post-transcriptional gene regulation signature in mouse embryonic cells.
[0127] Example 7: Thiol-linked alkylation for metabolic sequencing of small RNAs
[0128] To gain insight into the intracellular dynamics of small RNA silencing pathways, we applied a nucleotide analog derivatization strategy that allows us to elucidate the dynamics of RNA biosynthesis and turnover in the context of whole RNA, while avoiding the need for biochemical isolation of labeled RNA species (Figure 12). This strategy is based on the well-established 4-thiouridine(s) 4 U) Based on a metabolic RNA labeling approach, but instead of the biotinylation step which is ineffective and experimentally difficult, s 4 A short chemical treatment is performed by reacting U-containing RNA with iodoacetamide, a sulfhydryl-reactive compound (iodoacetamide is s 4 (When reacted with U, a covalently bonded amide methyl group is formed at the base pairing interface) (Figure 1). 4 The bulky group present at the U integration site, when combined with conventional RNA library preparation protocols (such as small RNA sequencing), can lead to the specific and quantifiable misintegration of G during reverse transcription (RT), but this does not interfere with the processing performance of RT. 4 U inclusion events can usually be identified at single-nucleotide resolution by bioinformatics in high-throughput sequencing libraries through sequence comparison, thus eliminating the need for spike-in solutions to measure absolute labeling efficiency in the context of unlabeled RNA by calling for T-to-C conversion. Importantly, no enzyme is known to convert uridine to cytosine, and T>C errors are rare, occurring at a frequency of less than 1 in 10,000, in datasets from high-throughput RNA sequencing using the Illumina HiSeq 2500 platform, for example. We call this approach thiol (SH)-linked alkylation for metabolic sequencing of small RNAs.
[0129] To validate the performance of this method for identifying metabolically labeled small RNAs, Drosophila S2 cells were subjected to conditions that do not interfere with cell survival (i.e., 500 μM) 4 After incubation with U for 24 hours, total RNA was extracted and small RNA sequences were performed. Metabolic labeling was confirmed by HPLC analysis of total RNA (Figure 18), revealing a labeling efficiency of 2.3%. This indicates that one of the 43 uridine molecules was s 4 This corresponds to the incorporation of U. 4 When a library was generated from size-selected total RNA in Drosophila S2 cells after 24 hours of U metabolic labeling, a stronger accumulation of T>C conversion was observed compared to a library prepared from unlabeled cells (for example, the highly expressed microRNA mir-184-3p and the less abundant miR * This is the miR-184 locus that produces miR-184-5p (Figure 12A). To confirm this observation on a genome scale, the reads were aligned to annotated microRNA loci in the Drosophila genome, and the frequency of mutations normalized to the total T content for each miRNA was examined (Figure 12B). 4 In the absence of U metabolic labeling, the median mutation rate for any mutation was observed to be less than 0.1%. This is consistent with the error rate reported in Illumina. (Note that treatment of unlabeled total RNA with iodoacetamide had no detectable effect on the small RNA content and mutation rate determined by high-throughput sequencing. See Figure 19.) Over a 24-hour period... 4 After U metabolic labeling, T>C mutations were statistically significant (p<10). -4 A 74-fold increase was observed (Mann-Whitney test), while all other mutations remained below the expected sequencing error rate (Figure 12B). More specifically, after 24 hours of labeling, 4 The median U inclusion rate was observed to be 2.22%, which corresponds to 1 s per 45 uridines. 4This corresponds to U incorporation and is consistent with the incorporation rate determined by HPLC measurement of total RNA (Figure 18). Therefore, metabolic labeling under undisturbed conditions results in most (over 95%) of the small RNAs being infused with at most one s 4 This shows a U integration event (Figure 20). This is insufficient for quantifiable recovery even with an improved biotinylation strategy (Duffy et al., 2015, see above reference), but chemically... 4 It can be easily identified by U-derivativeization (Figure 12).
[0130] The method of the present invention can achieve a single nucleotide resolution. 4 The ability to replicate U integration allows for s for microRNA processing and loading. 4 We were able to systematically analyze the effects of U metabolic labeling. For this purpose, we used a given small RNA derived from the 5p arm or 3p arm of a microRNA precursor, or the miR strand or miR * We examined the over- and under-occurrence of T>C conversion at individual positions of a given small RNA constituting the chain. This can be determined by selective algonaut loading. (35 5p-miRNAs and 36 3p-miRNAs, or 44 miRs and 27 miRs) * When we examined 71 abundantly expressed (over 100 ppm) microRNAs (corresponding to ), no significant phylogenetic changes in relative T>C mutation rates were observed at any given position (Figure 21). We found that s 4 We concluded that U metabolic labeling does not affect microRNA biosynthesis or loading.
[0131] When combined, the s in cultured Drosophila S2 cells 4 Performing SLAM-seq after U metabolic labeling allows for the detection of small RNAs. 4 Because U integration events are recovered to a degree that allows for quantification with single-nucleotide resolution, s 4 It becomes clear that U labeling does not have a significant, location-dependent effect on microRNA biosynthesis or loading.
[0132] Example 8: Intracellular dynamics of microRNA biosynthesis
[0133] MicroRNAs are derived from hairpin-containing RNA polymerase II transcripts and are sequentially processed by nuclear Drosha and cytoplasmic Dicer to produce mature miRNA double strands (Figure 13). To investigate the intracellular dynamics of miRNA biosynthesis, we examined the increase in the T>C mutation rate of miRNAs expressed abundantly (over 100 ppm in steady state) in a small RNA library prepared from size-selected total RNA of Drosophila S2 cells. 4 Metabolic labeling using U was determined at 5 minutes, 15 minutes, 30 minutes, and 60 minutes later, s 4 The error rate was compared with the error rate detected without U labeling (Figure 13B). A significant increase in the T>C conversion rate was already detected after a short labeling time of 5 minutes, representing 17% of the total miR. * 90% of the samples showed an error rate exceeding the maximum background level. This percentage increased over time, reaching 74%, 93%, and 100% at 15 minutes, 30 minutes, and 1 hour, respectively. Based on conventional measurements, more than 50% of all miRNAs (22 out of 42) were generated to a detectable level within a short period (i.e., 5 minutes). (i.e., even miRs, and their miRs) * Even with our partner, we exceeded the maximum background T>C conversion rate. This reveals that the miRNA processing mechanism in cells is extremely efficient.
[0134] As an alternative indicator of miRNA molecule production over time, the number of reads containing T>C conversion was also determined (Figure 13C). On average, miRNAs that are abundant in the steady state (such as miR-184 and miR-14) also show high production rates, but others (such as miR-276b and miR-190) have high biosynthesis rates but low steady-state content, or vice versa (i.e., miR-980 or miR-11). This indicates that the production rate of miRNAs is not the sole determinant of their intracellular content.
[0135] Our overall analysis revealed that total miRNA biosynthesis was unexpectedly efficient, but selected smaller RNAs were produced at significantly lower rates. An example of such poorly produced miRNAs is miltron (i.e., miR-1003, miR-1006, miR-1008; Figure 13C). Miltron is a class of microRNA that is produced by splicing instead of processing directed by Drosha. Miltron biosynthesis can be selectively reduced because the precursor hairpin derived from splicing is a specific target of uridylation by the terminal nucleotidyltransferase Tailor, which recognizes the 3' terminal splice acceptor site, thereby blocking efficient processing mediated by Dicer and initiating exonucleic degradation directed by dmDis312 (Figure 13D). Our data provide experimental evidence for the selective suppression of miltron biosynthesis. This is because Miltron showed a significantly lower T>C mutation rate compared to classical miRNAs, and reads containing T>C accumulated less over time (Figure 13D). Regarding the hypothesis that suppression of Miltron biosynthesis is behind the uridylation of pre-miRNA hairpins, a significant correlation was also detected between pre-miRNA tailing and the T>C conversion rate (Figure 22).
[0136] In summary, the novel method revealed a significantly higher rate of miRNA production within cells and reproduced the selective inhibitory effect of precursor hairpin uridylation on miRNA biosynthesis.
[0137] Example 9: Monitoring of small RNA loading into ribonucleoprotein complexes
[0138] MicroRNA biosynthesis generates a double-stranded miRNA. However, only one of the two strands of the miRNA double-stranded miRNA (the miR strand) is preferentially loaded onto the surface of Ago1 and selectively stabilized, while the other strand (miR *The miRNA strand is eliminated and degraded in the miRNA loading process (Figure 14A). To verify whether the method of the present invention replicates the processes of miRNA double-strand formation and miRNA loading in living cells, the relative accumulation of reads containing T>C conversion was analyzed for 20 miRNAs, and it was found that miR and its partner miRNA * Both were detected at sufficiently high levels (Figure 14B). 4 At early time points after U metabolic labeling (i.e., 5 minutes, 15 minutes, 30 minutes; Mann-Whitney test p>0.05), miR pairs and miR * Since no significant difference in content was observed between the pairs, it was confirmed that miRNAs initially accumulate as double strands in a process unrelated to loading in terms of kinetics. It was only after 1 hour that the miR strands began to form miR * Significantly greater accumulation was detected compared to the chain (Mann-Whitney test, p>0.05, Figure 14B). This indicates that, on average, the loading of miRNA into Ago1 occurs at a much slower rate than miRNA biosynthesis.
[0139] More detailed analysis revealed a two-step process underlying miRNA accumulation. The first step involves miR and miR * They are the same (k miR =0.35±0.03 and k miR* =0.32±0.03). Therefore, this reflected that miRNAs accumulate as double strands. Of note is that the second, slower stage involves miR and miR * In both cases, the process differed from the biosynthetic stage. * A rapid decrease in the accumulation rate (k miR* (=0.32±0.03) represents the majority (i.e., about 81%) of miR * The chain showed that it was rapidly degraded as a consequence of miR loading. In contrast, the miR chain was miR * chain(k miR* The accumulation rate (k = 0.32 ± 0.03) is much larger compared to this. miRThe initial biosynthesis rate (k) was 0.26±0.03. This suggests that the miR chain is selectively stabilized, likely due to the formation of miRISC. However, the initial biosynthesis rate (k) was also observed. miR Compared to (=0.35±0.03), miR also shows a decrease in reaction rate in the second stage (k miR The result was 0.26±0.03). This indicates that only about 74% of the miR strands were effectively loaded, while nearly a quarter of the miRNAs were likely degraded as double-stranded cells. This is consistent with empty Argonaut availability and is one important limiting factor regarding miRNA accumulation, whose overexpression increases the overall intracellular miRNA content.
[0140] Furthermore, individual miR:miR * Pair studies revealed differences in loading efficiency between miRNA double strands. Strand separation, and therefore loading, was detectable within minutes for miR-184, while bantam showed a slightly delayed loading kinetics, with strand separation not occurring until approximately 30 minutes later. miR-282 ranked among the miRNAs with the lowest effective loading. Notably, the differences in loading kinetics followed thermodynamic rules for Ago1 loading, suggesting that mismatches in the species or 3' support region of the miRNA double strand facilitated the efficient formation of miRISCs.
[0141] In summary, this novel method has revealed a detailed understanding of the dynamics of miRNA biosynthesis and loading.
[0142] Example 10: Production of isomiR
[0143] Accumulated evidence suggests that while diverse intracellular processes diversify the sequences and functions of microRNAs, a poor understanding of the underlying mechanisms makes analysis from small, steady-state RNA sequencing libraries challenging. One well-established example regarding isomiR production is the exonucleic maturation of miRNAs in flies. Most Drosophila miRNAs are produced as small RNAs of approximately 22 nt, but selected miRNAs are produced as longer, approximately 24-mers. These 24-mers require further exonucleic maturation mediated by the 3'→5' exoribonuclease Nibbler to form gene regulatory miRISCs. In standard small RNA sequencing libraries and high-resolution Northern hybridization experiments, miR-34-5p exhibits a range of length profiles, resulting in expression-rich isoforms ranging from 24-mers to 21-mers, arising from 3'-end cleavage of the same 5' isoform (Figure 15B). To verify the ability of the present invention to elucidate the intracellular sequence of events that give rise to diverse miR-34-5p isoforms, 4 We analyzed a small RNA library prepared from total RNA of S2 cells after U metabolic labeling. Unlike the stationary small RNA, which showed a very similar length profile throughout the entire time course, the reads of miR-34-5p containing T>C conversion initially accumulated as a complete 24-mer isoform, consistent with previous in vitro processing experiments using recombinant Dcr-1 i.e., fly lysate and synthetic pre-miR-34 (Figure 15C, bottom). Only after 3 hours was the emergence of the shorter T>C conversion-containing miR-34-5p 3' isoform detected (Figure 15C, top). From this point onward, the weighted average length of miR-34-5p continuously decreased over time, approaching the average length profile of miR-34-5p observed in the stationary state (Figure 15D). We concluded that the emergence of isomiR is revealed by the method of our invention, as exemplified by the 3'→5' exonucleic acid degradation trimming directed by Nibbler.
[0144] Exonucleic degradation and maturation of miRNAs require loading the miRNA into Ago1, and biochemical evidence suggests that trimming is necessary for miRNA development. * This suggested that it occurred only after strand removal. This is likely because Nibbler provided function as a single-strand specific 3'→5' exoribonuclease. Since our method made it possible to simultaneously measure miRNA loading and isomiR production, we tested this hypothesis by comparing the miR-34-5p trimming signal (Figure 15D) with the miR-34 double-strand loading dynamics (Figure 15E). Trimming loads miR-34 and miR * What actually happens after chain removal was observed (Figure 15E). This is determined from the shift in the accumulation of miR-34-5p and miR-34-3p that begins 1 hour after metabolic labeling in our library. In conclusion, the method of the present invention reveals the intracellular sequence of miRNA isoform production and thus provides a powerful tool for shedding light on the processes that diversify the sequence and function of miRNAs in living cells.
[0145] Example 11: Stability of microRNA
[0146] Various miRNAs are assembled, otherwise forming indistinguishable protein complexes, but cumulative evidence suggests that the stabilization of these protein complexes may differ dramatically from one another (Figure 16). However, currently available technologies only measure the relative half-life, not the absolute half-life, for individual miRNAs, preventing a detailed understanding of miRNA stability. Furthermore, metabolic labeling under undisturbed conditions reveals that most (over 95%) of the labeled small RNAs have up to one s 4This facilitates the demonstration of the U incorporation event (Figure 20). However, even with improved biotinylation strategies (Duffy et al., 2015, see above), quantifiable recovery remains insufficient, and bias is introduced due to the wide range of U content differences among miRNAs with different miRNA sequences. In contrast, the method of the present invention provides rapid access to absolute, sequence-content-normalized miRNA stability in the context of a standard small RNA library (Figure 16). 4 By analyzing the T>C conversion rate normalized to miRNA-U content in a small RNA library prepared from total RNA of Drosophila S2 cells that had been labeled with U metabolically for up to 24 hours using SLAM-seq, the median half-life of 41 abundantly expressed miR chains was determined to be 12.13 hours (Figure 16B). This indicates that the average half-life of miRNA is significantly longer than that of mRNA, which exhibits an average half-life of approximately 4-6 hours. * Unlike miR, it showed a much shorter half-life of 0.44 hours (Figure 16B). This is a consequence of miR loading. * This is consistent with the enhanced metabolic turnover (Figure 14). In short, miR * Its stability is significantly lower compared to miR (Figure 16C). However, even among miRs, they exhibit fundamentally different stabilities, differing in magnitude by more than an order of magnitude. An example of this is the unstable miR-12-5p(t 1 / 2 (=1.7 hours) and stable bantam-5p(t 1 / 2 (>24 hours). Importantly, the present invention provides a robust outlook on the half-lives of individual small RNAs, thus providing highly reproducible results regarding the stability of a variety of small RNAs (Figure 16E).
[0147] MicroRNA stability is a major factor contributing to establishing the profile of small RNAs within S2 cells, as exemplified by the two miRNAs that accumulated at the highest levels in the steady state. Bantam exhibited relatively slow biosynthesis (Figure 13) and a moderate loading rate (Figure 14), but exceptionally high stability (t 1 / 2It accumulated to the highest steady-state level because (>24 hours). In contrast, the second most abundant miRNA, miR-184-3p, was bantam-5p(t 1 / 2 Although its stability was only one-third compared to the 6-hour period, it still accumulated at a high level due to its unusually large biosynthesis and loading dynamics (Figures 13 and 14). Therefore, metabolic sequencing of small RNAs by SLAM-seq reveals that miRNA biosynthesis, loading, and turnover relatively contribute to the establishment of small steady-state RNA profiles, which are major determinants of miRNA-mediated gene regulation.
[0148] Example 12: The characteristics of the Argonaut protein determine the stability of small RNAs.
[0149] Both mammalian and insect genomes encode several proteins of the Argonaut protein family. Some of these proteins selectively load small RNAs to regulate distinct subsets of transcripts through various mechanisms. miRNA double-strands are inherently asymmetric, meaning the miR strand is preferentially loaded into Ago1, but each miRNA precursor can potentially produce two mature small RNA strands, which can be classified into two distinct Argonaut proteins universally expressed in flies. * Unlike most miRs, this protein is often loaded as a functional species into Ago2, an effector protein in the RNAi pathway, and in the final step of the Ago2-RISC assembly, the 2' position of the 3' terminal ribose undergoes selective methylation by the methyltransferase Hen1 (Figure 17A). Since Ago2 deficiency does not impair cell survival, we were able to investigate the role of Ago proteins in the selective stabilization of small RNAs. Furthermore, Ago2 deficiency provided an experimental framework to understand whether the half-life of small RNAs is inherently determined by the characteristics of the associated Ago protein.
[0150] First, we established a set of miRNAs specifically assembled to become Ago2 in wild-type Drosophila S2 cells by comparing a small RNA library prepared from total RNA by conventional small RNA cloning (which preferentially reflects small RNAs bound to Ago1) with a library prepared from total RNA but rich in small RNAs that have been methylated by oxidation (i.e., bound to Ago2). While most small RNAs in conventional cloning approaches consisted of miRNAs (particularly miR chains), Ago2-binding endogenous small RNAs derived from transposons, genes (primarily derived from duplicated mRNA transcripts), and loci that produce long folded transcripts (structured loci) were selectively enriched by oxidation. As previously described, methylated (i.e., bound to Ago2) miRNAs are miR * Selective enrichment was achieved by comparing an unoxidized small RNA library with an oxidized small RNA library, and by comparing miR strands and miR * The strands could be classified according to their accumulation within Ago1 and Ago2 (Figure 17C). By comparing the content of Ago2-enriched small RNAs in a conventional small RNA library generated from wild-type S2 cells and S2 cells deficient in Ago2 by CRISPR / Cas9 genome manipulation (Figure 17D), it was confirmed that the content of Ago2-enriched small RNAs was significantly reduced when Ago2 was deficient (p<0.002, Wilcoxon's signed-rank test for paired data, Figure 17E).
[0151] Next, s 4 U metabolic labeling and subsequent SLAM-seq differentiate wild-type cells from Ago2-deficient cells (ago2 ko The stability of Ago2-enriched small RNAs was determined by measuring the stability of small RNAs in all small RNA libraries prepared from cells. In wild-type cells, Ago2-enriched small RNAs followed a biphase degradation kinetics. Here, the majority of the population (i.e., 94%) showed high stability (t 1 / 2While it showed a half-life of >24 hours, only a very small amount certainly decomposed rapidly, with a half-life of miR * It was the same as (t 1 / 2 (=0.2 hours). We investigated whether the population associated with the long half-life represented the Ago2-binding fraction by examining the stability of the same small RNA species in a methylated small RNA library. Indeed, the Ago2-enriched small RNAs followed a single exponential degradation kinetics and had a half-life of over 24 hours (Figure 17F). Conversely, in Ago2-deficient S2 cells, the Ago2-enriched small RNAs also followed a two-phase degradation kinetics, but this time the majority (63%) of the population were miR * It showed similar stability to (t 1 / 2 (=0.4 hours). This indicates that in the absence of Ago2, these small RNAs are primarily degraded when their partner strand is loaded onto Ago1. In contrast, Ago1-enriched small RNAs exhibited the same stability in the presence and absence of Ago2 (Figure 17F). Thus, our data reveal that miRNAs possess population-specific stability, and this stability is determined by the loading of the miRNA onto a specific Ago protein.
[0152] Finally, to analyze whether the half-life of small RNAs is intrinsically determined by the characteristics of the Ago protein, the stability of 30 small RNAs most abundant in Ago1 and Ago2 was compared (Figure 17G). This analysis revealed that small RNAs bound to Ago2 showed significantly greater stability compared to small RNAs bound to Ago1 (p<10). -4 (Mann-Whitney test). This is likely because the methylation of the small RNA bound to Ago2 contributed to the stabilization of the small RNA within Ago2, rather than Ago1.
[0153] In summary, this provides an experimental framework for analyzing the molecular mechanisms underlying the establishment and maintenance of small RNA profiles that influence gene expression in healthy and diseased states.
[0154] Example 13: SLAM-seq determines the direct gene regulatory function of the BRD4-MYC axis.
[0155] Identifying the genes directly targeted by transcription factors such as BRD4 and MYC is crucial for both understanding the fundamental function of those genes in cells and developing therapeutic methods. However, elucidating direct regulatory relationships is difficult for various reasons. While genome binding sites can be mapped, for example, by chromatin immunoprecipitation and sequencing (ChIP-seq), the binding of a single factor alone cannot predict the regulatory function of neighboring genes. Another approach involves experimentally disrupting a given regulator and then examining the resulting differences in expression.
[0156] For example, to further investigate whether SLAM-seq can capture more specific transcriptional responses caused by disruption of signaling pathways, K562 cells were treated with BCR / ABL, the oncogene that drives these cells, as well as small molecule inhibitors of MEK and AKT, which act as mediators in another signaling cascade downstream of BCR / ABL (Figures 25D, 30A, and 30B).
[0157] Cell culture
[0158] Leukemia cell lines K562, MOLM-13, and MV4-11 were cultured in RMPI 1640 and 10% fetal bovine serum (FCS). OCI / AML-3 cells were grown in MEM-α-containing 10% FCS. HCT116 and Lenti-X lentiviral packaging cells (Clontech) were cultured in DMEM and 10% FCS. L-glutamine (4 mM) was supplemented with all growth media. To determine the growth curve, cells were cultured at an initial density of 2 × 10⁶ cells in the presence or absence of 100 μM IAA (indole-3-sodium acetate, Sigma-Aldrich). 6Cells were seeded at a rate of cells / ml, and the culture medium and IAA were refreshed every 24 hours in a 1:2.6 ratio to maintain a non-crowded cell environment. Cell density was measured every 24 hours using a Guava EasyCyte flow cytometer (Merck Millipore).
[0159] Cell viability assays were performed using the CellTiter-Glo Luminescent Cell Viability Assay (Promega) after treating cells with a combination of JQ1 and NVP-2 for 72 hours. Relative luminescence signals (RLU) were recorded using the EnSpire Multimode Plate Reader (Perkin Elmer). The partial response to drug treatment was α = 1 - (RLU). treated / RLU untreated ) was defined as such, and the synergistic effect was calculated as an excess (eob) relative to Bliss additivity. However, eob = α NVP-2,JQ1 - α JQ1 - ( α NVP-2 · 1 - α JQ1 )
[0160] Plasmid and vector used in this example
[0161] SpCas9 and sgRNA were expressed from the plasmid pLCG (hU6-sgRNA-EFS-SpCas9-P2A-GFP). The pLCG was cloned based on a publicly available Cas9 expression vector (lentiCRISPR v2, Addgene plasmid #52961). This plasmid contains a modified chiRNA context. The sgRNA sequence was cloned into the pLCG. As a donor for homologous recombination repair of the target locus, an AID knock-in cassette was synthesized by gene synthesis (Integrated DNA Technologies), and approximately 500 bp homology arms (HAs) from the genomic DNA of the target cell line were amplified by PCR. When all the components were assembled and placed into the lentiviral plasmid backbone (Addgene plasmid #14748), constitutive GFP expression for monitoring transfection was added, yielding the final vectors pLPG-AID-BRD4 (5'HA-Blast(registered trademark)-P2A-V5-AID-spacer-3'HA-hPGK-eGFP) and pLPG-MYC-AID (5'HA-spacer-AID-P2A-Blast(registered trademark)-3'HA-hPGK-eGFP). For acute protein deficiency experiments, Tirl from Asian rice (Oryza sativa) was introduced using the publicly available lentiviral vector SOP (pRRL-SFFV-Tir1-3xMYC-tag-T2A-Puro). For competitive growth assays, Tir1 was introduced using the vector SO-Blue (pRRL-SFFV-Tir1-3xMYC-tag-T2A-EBFP2). RNAi was performed using the LT3GEN vector, which delivers shRNAmir inserts.
[0162] Genome editing and lentiviral transduction
[0163] To derive AID knock-in cell lines, plasmid pLCG and pLPG were simultaneously delivered by electroporation using a MaxCyte STX electroporation system (K562) or by transfection using FuGENE HD Transfection Reagent (Promega) (HCT116). After selecting successful knock-ins using brassicidine (10 μg / ml, Invitrogen), GFP was selected using a BD FACSAria III cell sorter (BD Biosciences). - Single-cell clones were isolated. Clones were characterized by PCR genotyping of crude cell lysates. Knock-in was further confirmed by immunoblotting of tagged proteins, and for K562, clones were characterized by flow cytometry to best match the immunophenotype of wild-type cells.
[0164] For acute protein deficiency experiments, a validated homozygous AID knock-in clone was transduced using the Tir1 expression vector SOP. Following standard procedures, the helper plasmids pCMVR8.74 (Addgene plasmid #22036) and pCMV-VSV-G (Addgene plasmid #8454) were transduced into polyethyleneimine (PEI, M). W Lentiviral particle packaging was performed in Lenti-X cells by transfection (25000, Polysciences). Target cells were infected at limiting dilution and selected on puromycin (2 μg / ml, Sigma-Aldrich). To avoid potential silencing of the transgene, all deficiency experiments were performed on newly transfected and selected cells.
[0165] Immunoblotting and immunophenotypic testing
[0166] The chemiluminescence of the primary antibody was detected using HRP-labeled secondary antibodies (Cell Signaling Technology, catalog numbers #7074, #7076, #7077). Alternatively, the fluorescence of rabbit and mouse primary antibodies was detected with the Odyssey CLx Imaging System (LI-COR Biosciences) using the secondary antibodies IRDye 680RD goat anti-rabbit IgG and IRDye 800CW goat anti-mouse IgG (LI-COR Biosciences).
[0167] For immunophenotyping, cells were washed with FACS-buffer (5% FCS in PBS) and pre-incubated for 10 minutes at room temperature with FCS-receptor blocking peptide (Human TruStain FcX, Biolegend, diluted 1:20 in FACS-buffer). Fluorophore-labeled antibodies were added at a final dilution rate of 1:400 and the cells were incubated for 20 minutes at 4°C. The stained cells were washed twice, resuspended in FACS-buffer, and then analyzed with a BD LSRFortessa flow cytometer (BD Biosciences).
[0168] Fragmentation of chromatin
[0169] To fragment chromatin, cells were washed in ice-cold PBS and resuspended in chromatin extraction buffer (20 mM Tris-HCl, 100 mM NaCl, 5 mM MgCl2, 10% glycerol, 0.2% IGEPAL CA-630, 20 mM β-glycerophosphate, 2 mM NaF, 2 mM Na3VO4, protease inhibitor cocktail (without EDTA, Roche), pH 7.5). The insoluble fraction was precipitated by centrifugation (16,000×g, 5 min, 4 °C), washed three times in chromatin extraction buffer, and then resuspended therein. Whole cell fractions and supernatants were sampled before and after the first precipitation. All fractions supplemented with SDS (sodium dodecyl sulfate, 0.1% (w / v)) were digested with Benzonase (Merck Millipore, 30 min, 4 °C) and redissolved by sonication in a Bioruptor sonicator (Diagenode).
[0170] SLAM-seq
[0171] All SLAM-seq assays were performed at a confluency of 60–70% for adherent cells and at a maximum cell density of 60% counted with a hemocytometer for suspension cells. The growth medium was aspirated and replaced 5–7 h before each assay. Unless otherwise indicated, cells were pretreated with the indicated small molecule inhibitor or 100 μM IAA for 30 min to pre-establish a state where the target was completely inhibited or degraded. Labeling of newly synthesized RNA was performed with a final concentration of 100 μM 4-thiouridine (s 4The procedure was performed using a 3' mRNA sequencing kit (Carbosynth) for the instructed time (45 or 60 minutes). Adherent cells were recovered by rapidly freezing the plate directly on dry ice. Suspended cells were rapidly frozen immediately after centrifugation. RNA extraction was performed using the RNeasy Plus Mini Kit (Qiagen). Total RNA was alkylated with iodoacetamide (Sigma, 10 mM) for 15 minutes, and the RNA was purified again by ethanol precipitation. Using 500 ng of alkylated RNA as input, a 3' mRNA sequencing library was generated using a commercially available kit (Quant-seq 3' mRNA-Seq Library Prep Kit FWD for Illumina and PCR Add-on Kit for Illumina, Lexogen). Deep sequencing was performed using the HiSeq1500 platform and HiSeq2500 platform (Illumina).
[0172] Differential gene expression analysis, PCA, GO terminology enrichment
[0173] To analyze at the gene level, raw reads mapped to different UTR annotations of the same gene were summed using Entrez Gene ID. A preliminary study of K562 cells with kinase inhibitors was performed as a single experimental group. Analysis of gene expression differences was limited to genes read at least 10 times under at least one condition (flavopyridol and DMASO) for 50 bp sequencing runs, and to genes read at least 20 times under at least one condition (mk2206, trametinib, nilotinib, trametinib + mk2206, DMSO) for 100 bp sequencing runs. To assess expression differences, one pseudocount per raw read was added to all genes.
[0174] All other SLAM-seq experiments were performed in triplicate and analyzed as follows: Differential gene expression calling was performed using DESeq2 (version 1.14.1) with default settings for raw read counts with T>C conversion of 2 or greater, and using the size factor estimated for the corresponding total mRNA for overall normalization. Downstream analysis was limited to genes that passed through all internal filters for FDR estimation by DESeq2. After variance stabilization conversion of the 500 genes with the greatest variability under all experimental conditions, principal component analysis was performed. K562 MYC-AID For genes that were significantly and strongly downregulated by SLAN-seq during IAA treatment with +Tir1 (FDR ≤ 0.1, log2FC ≤ -1), GO term enrichment analysis was performed using the PANTHER over-recurrence test (Fisher's exact test with FDR multiple test correction, pantherdb.org).
[0175] Evaluation of mRNA metabolic turnover
[0176] To roughly estimate mRNA turnover in K562 cells without disturbance, we assume that mRNA biosynthesis is in equilibrium in a steady state, and s 4 We assumed that the gene degrades at a first-order rate after prolonged exposure to U, approaching the fully labeled state. Therefore, for any gene i, s 4 β is the percentage of converted reads (2 or more T>C conversions) out of the total read count after 60 minutes of U-labeling. i Using this method, the half-life of cellular mRNA is: t 1 / 2,i = -60 × ln (2) / ln (β i ) We were able to calculate it as follows.
[0177] Deep sequencing after chromatin immunoprecipitation (ChIP-Seq)
[0178] For ChIP-Seq, 1 × 10 8 ~2×10 8 K562 AID-BRD4+Tir1 cells were treated with 100 μM IAA or DMSO for 1 hour, crosslinked with 1% formaldehyde over 10 minutes at room temperature, quenched with 500 mM glycine for 5 minutes, and then washed twice with ice-cold PBS. After isolating the nuclei, the pellet was dissolved in a lysis buffer containing a protease inhibitor (Complete, Roche) (10 mM Tris-HCl, 100 mM NaCl, 1 mM EDTA, 0.5 mM EGTA, 0.1% sodium deoxycholate, 0.5% N-lauroyl sarcosine, pH 8.0). Chromatin was sheared using a Bioruptor sonicator (Diagenode). Cell residue was pelleted by centrifugation at 16000 × g for 10 minutes at 4°C. To allow direct comparison between ChIP-Seq samples treated with DMSO and those treated with IAA, chromatin from mouse AML cell lineage (RN2) was added in a ratio of RN2:K562 ≈ 1:10 as a spike-in control for internal normalization. Immunoprecipitation was performed by adding Triton X-100 (final concentration 1%) and incubating the chromatin lysate with 5–10 μg of antibody overnight at 4°C on a rotating wheel. The antibody-chromatin complexes were captured on magnetic Sepharose beads (G&E Healthcare; blocked at room temperature for 2 hours with TE containing 1 mg / ml BSA) on a rotating wheel at 4°C for 2 hours. The beads were washed once with RIPA buffer (150 mM NaCl, 50 mM Tris-HCl, 0.1% SDS, 1% IGEPAL CA-630, 0.5% sodium deoxycholate, pH 8.0), Hi-Salt buffer (500 mM NaCl, 50 mM Tris-HCl, 0.1% SDS, 1% IGEPAL CA-630, pH 8.0), and LiCl buffer (250 mM LiCl, 50 mM Tris-HCl, 1% IGEPAL CA-630, 0.5% sodium deoxycholate, pH 8.0), and then washed twice with TE. The immunocomplexes were eluted in 1% SDS and 100 mM NaHCO3.The samples were treated with RNase A (100 μg / ml) at 37°C for 30 minutes, followed by the addition of NaCl (200 mM) and crosslinking inversion at 65°C for 6 hours. The samples were then digested with 200 μg / ml proteinase K at 45°C for 2 hours. Genomic DNA was recovered from both the precipitated material and the sheared chromatin input (1% of the material used for ChIP) by phenol-chloroform extraction and ethanol precipitation. Libraries for Illumina sequencing were prepared using the NEBNext Ultra II DNA Library Preparation Kit for Illumina (New England Biolabs, #7645).
[0179] Analysis of spike-in controlled ChIP-seq data
[0180] To analyze spike-in controlled ChIP-seq samples, a hybrid reference genome was prepared by mixing human and mouse genome sequences (GRCh38 and mm10). Reads were first aligned to this hybrid genome using bowtie 2 v2.2.9 (high sensitivity), and then separated into human and mouse bins. Read coverage for each track was calculated using deeptool v2.5.0.1 and rescaled using a spike-in normalization factor. After further subtracting the respective input signals from the resulting normalized coverage tracks, the ratio of DMSO-treated samples to IAA-treated samples was calculated.
[0181] Reanalysis of ChIP-Seq and Click-Seq data and super-enhancer calling
[0182] Previously published Click-seq data, H3K27ac ChIP-seq data, and matching input samples were realigned to GRCh38 using bowtie (version 1.1.2) after removing the adapter sequence using cutadapt. For K562 cells, proximal super-enhancer genes were used. For MV4-11 and MOLM-13, H3K27ac peak calling was performed using MACS2 (v2.1.0.20140616) with default parameters. Super-enhancer calling was performed using ROSE v0.1 with default parameters. Super-enhancers were assigned to genes based on the nearest TSS within 100 kb. Subsequent comparisons were limited to proximal super-enhancer transcripts that were assigned an Entrez GeneID and whose expression could be detected in SLAM-seq.
[0183] Modeling for predicting transcriptional responses
[0184] The TSS locations of all Refseq transcripts in GRCh38.p9 were downloaded from www.ensembl.org / biomart. The density of CAGE-seq reads within 300 bp of each TSS on each strand was extracted from publicly available CAGE-seq data for K562. The TSS with the highest mean signal of two replicates was retained and further analyzed. 213 publicly available, pre-analyzed ChIP-Seq tracks and one whole-genome bisulfite sequencing experiment were obtained from the ENCODE project (www.encodeproject.org / ) or the Cistrome Data Browser (cistrome.org / db / ). ChIP-Seq signals within 500 bp and 2000 bp around each TSS were used as input for classification models.
[0185] To model the prediction of JQ1 hypersensitivity, the response of K562 cells to 200 nM JQ1 was measured by SLAM-seq (FDR ≤ 0.1, log2FC ≤ -0.7), and genes were classified as downregulated based on these results. Unaffected genes (FDR > 0.1, -0.1 ≤ log2FC ≤ 0.1) were subsampled, and a matched control set with the same size and baseline mRNA expression was obtained by repeated resampling and comparison with the target distribution using the Kolmogorov-Smirnov test. The question genes and control genes were cross-referenced with a TSS-ChIP-seq signal matrix and split into a training set (75%) and a test set (25%). Using scaled and centered ChIP signals, five independent classifiers with cross-validation of 5-fold or more (elastic net GLM, gradient boosting machine, SVM with linear kernel, SVM with multinomial kernel, SVM with radial kernel) were trained using the CARET package during parameter tuning. All four final models were compared using a holdout test set.
[0186] To model the prediction of transcription dependent on MYC, K562 MYC-AID Based on the SLAM-seq response when +Tir1 cells were treated with IAA, genes were classified as either downregulated (FDR ≤ 0.1, log2FC ≤ -1) or unaffected (FDR ≤ 0.1, -0.2 ≤ log2FC ≤ 0.2). Unaffected genes were further subsampled to obtain a control set of the same size and matching expression, as described for JQ1 response modeling. Due to the large sample size, the genes were split into a training set (60%), a test set (20%), and an additional validation set (20%), and processed as described for JQ1 response modeling.
[0187] Analysis of direct MYC target signatures in cell lines and RNA-seq data from cancer patients
[0188] To compare MYC expression with empirical MYC response signatures, FPKM-normalized gene expression data from 672 human cancer cell lines were obtained from Klijn et al. (Nat. Biotechnol. Vol. 33, pp. 306-312 (2015)). Cell lines expressing MYCN or MYCL at higher levels than MYC were excluded, and the remaining samples were classified as MYC-high (top 20% of MYC expression) or MYC-low (bottom 20% of MYC expression). All annotated genes from the cell line expression datasets using Entrez GeneID and K562 were used. MYC-AID +Tir1 and HCT116 MYC-AID Among all genes significantly downregulated (FDR ≤ 0.1) within +Tir1, the 100 genes with the highest mean downregulation in both cell lines were defined as common MYC response signatures. To obtain balanced estimates of the expression of all signature genes, the FPKM values of each gene were scaled across all cell lines, and the scaled expression values of all signature genes were averaged across each cell line. Normalized upper quartile gene expression data from 5583 cancer patients from 11 TCGA projects were downloaded from portal.gdc.cancer.gov and processed independently for each cancer type as described for the cell line datasets. Gene set enrichment analysis was performed using GSEA Desktop v3.0 beta.
[0189] Sample preparation for proteomics
[0190] K562 AID-BRD4 +Tir1 cells were treated with 100 μM IAA or DMSO for 60 minutes in three independent experiments, washed three times with ice-cold PBS, pelletized by centrifugation, and rapidly frozen. The pellets were resuspended in lysis buffer (10 M urea, 50 mM HCl), incubated at room temperature for 10 minutes, and then fused in 1 M Tris-buffer (Tris-HCl, c finalThe pH was adjusted using 100 mM Tris buffer (pH 8). Nucleic acids were digested with benzonase (Merck Millipore, 250 U per pellet, 1 hour, 37 °C), alkylated by adding iodoacetamide (15 mM, 30 minutes, room temperature), and quenched with DTT (4 mM, 30 minutes, 37 °C). For proteolysis, 200 μg of protein per sample was diluted with 100 mM Tris buffer to a urea concentration of 6 M and digested with Lys-C (Wako) at an enzyme:protein ratio of 1:50 (3 hours, 37 °C). The sample was further diluted with 100 mM Tris buffer to a final urea concentration of 2 M and digested with trypsin (Trypsin Gold, Promega) at an enzyme:protein ratio of 1:50 (37 °C, overnight). The pH was adjusted to less than 2 using 10% trifluoroacetic acid (TFA, Pierce), and desalted using a C18 cartridge (Sep-Pak Vac (50 mg), Waters). Peptides were eluted using 70% acetonitrile (ACN, Chromasolv, gradient grade, Sigma-Aldrich) and 0.1% TFA, and then lyophilized. Isobaric labeling was performed using the TMTsixplex Isobaric Label Reagent Set (Thermo Fisher Scientific), samples were mixed in the same molar amount, and lyophilized. After re-purifying the peptides using a C18 cartridge, the peptides were eluted using 70% ACN and 0.1% formic acid (FA, Suprapur, Merck), and then lyophilized.
[0191] Fractionation of proteomics samples by strong cation exchange chromatography (SCX)
[0192] The dried sample was dissolved in SCX buffer A (5 mM NaH2PO4, 15% ACN, pH 2.7). SCX was performed with 200 μg of peptide. A Thermo Fisher Scientific UltiMate 3000 Rapid Separation system with a flow rate of 35 μl / min and a custom-made TOSOH TSKgel SP-2PW SCX column (5 μm particles, pore size 12.5 nm, inner diameter 1 mm × 250 mm) were used. A ternary gradient was used for separation, starting with 100% buffer A for 10 minutes, and linearly increasing over 80 minutes to 10% buffer B (5 mM NaH2PO4, 1 M NaCl, 15% ACN, pH 2.7) and 50% buffer C (5 mM Na2HPO4, 15% ACN, pH 6). This was then reduced to 25% buffer B and 50% buffer C over 10 minutes, and then to 50% buffer B and 50% buffer C over 10 minutes, followed by isocratic elution for a further 15 minutes. The flow-through was collected as a single fraction, and the gradient fraction was collected minute by minute over 140 minutes, pooled to form 110 fractions, and stored at -80°C.
[0193] Quantitative analysis of peptides using LC-MS / MS.
[0194] LC-MS / MS was performed using a Thermo Fisher RSLC nanosystem (Thermo Fisher Scientific) connected to a Q Exactive HF mass spectrometer (Thermo Fisher Scientific) equipped with a Proxeon nanospray source (Thermo Fisher Scientific). Peptides were loaded at 25 μl / min into a trap column (Thermo Fisher Scientific, PepMap C18, 5 mm × 300 μm inner diameter, 5 μm particles, 100 Å pore size) using 0.1% TFA as the mobile phase. After 10 minutes, the trap column was switched to align with the analysis column (Thermo Fisher Scientific, PepMap C18, 500 mm × 75 μm inner diameter, 2 μm, 100 Å). The gradient was started with mobile phases of 98%A (H2O / FA, 99.9 / 0.1, v / v) and 2%B (H2O / ACN / FA, 19.92 / 80 / 0.08, v / v / v), increased to 35%B over 60 minutes, then increased to 90%B over 5 minutes, maintained for 5 minutes, and then decreased to 98%A and 2%B over 5 minutes to reach equilibrium at 30°C.
[0195] The Q Exactive HF mass spectrometer was operated in data-dependent mode, and an MS / MS scan was performed after a full scan of the 10 most abundant ions (m / z range 350–1650, nominal resolution 120000, target value 3E6). MS / MS spectra were acquired using a normalized collision energy of 35%, isolation width 1.2 m / z, resolution 60,000, target value 1E5, and a first fixed mass set of 115 m / z. Precursor ions selected for fragmentation (those with unassigned charge states 1, >8 excluded) were placed on a dynamic exclusion list for 30 seconds. In addition, the minimum AGC target was set to 1E4 and the intensity threshold was 4E4. Peptide matching features were set to preferred, and isotopic features to exclude were enabled.
[0196] Analysis of proteomics data
[0197] Raw data were processed using Proteome Discoverer (version 1.4.1.14, Thermo Fisher Scientific). Database searches were performed using MS Amanda (version 1.4.14.8240) in a database consisting of the human SwissProt database and attached contaminants (a total of 20,508 protein sequences). Methionine oxidation was defined as a dynamic modification, while N-terminal cysteine and TMT carbanide methylation and lysine carbanide methylation were designated as fixed modifications. Trypsin was defined as a protease that cleaves after lysine or arginine (except when followed by proline), with up to two incorrect cleavages allowed. Precursor and fragment ion tolerances were set to 5 ppm and 0.03 Da, respectively. Identified spectra were rescored using Percolator and filtered to a peptide spectral agreement level of 0.5% FDR. Protein grouping was performed in Proteome Discoverer using a strict parsing principle. The intensity of the reporter ion was extracted from the most reliable center of gravity mass with an integral tolerance of 10 ppm. For all proteins detected to have at least two unique peptides, the protein level quantification results were calculated based on all unique peptides contained in a given protein group. The statistical reliability of proteins with different amounts was calculated using limma.
[0198] Preparation of cellular metabolites for mass spectrometry
[0199] In a preheated growth medium, cells were raised in the presence of 100 μM IAA or DMSO (1:5000 (v / v)) for 2 × 10⁶ cells. 5 Cells were seeded at a rate of cells / ml. The culture medium was changed after 24 hours, and cells were counted and harvested after 48 hours. The cells were washed twice with PBS and rapidly frozen. 4 × 10⁶ cells per sample. 6Cell pellets were dissolved in a mixture of MeOH, ACN, and H2O (ratio 2:2:1 (v / v)), vortexed, and rapidly frozen. To ensure complete lysis, the cells underwent three cycles of rapid freezing, thawed, and subjected to sonication (10 minutes, 4°C, maximum intensity in a Bioruptor sonicator (Diagenode)). Proteins were allowed to settle at -20°C for 1 hour, then centrifuged (15 minutes, 18000×g, 4°C). The supernatant was collected, evaporated in a SpeedVac concentrate, and the pellets were dissolved again in a 1:1 mixture of ACN and H2O (v / v) by sonication (10 minutes, 4°C). The residue was removed by centrifugation (4°C, 15 minutes, 18000×g).
[0200] Targeted LC-MS / MS of cellular metabolites
[0201] Prior to analysis, 50 μl of ACN was added to each 60 μl sample, and 3 μl was injected into an UltiMate 3000 XRS HPLC system (Dionex, Thermo Scientific). Metabolites were separated using a ZIC-HILIC column (100 × 2.1 mm, 3.5 μm, 200 Å, Merck) and a flow rate of 100 μl / min, utilizing a gradient that increased from 5% mobile phase A (water containing 10 μm of ammonium acetate, pH 7.5) to 50% A in phase B (ACN) for 14 minutes. MS / MS was performed using a TSQ Quantiva triple quadrupole mass spectrometer (Thermo Scientific) with selective reaction monitoring (SRM) in negative ion mode. Samples from three independent experiments were analyzed separately in technical triplicates, and MS data were analyzed using TraceFinder (Thermo Scientific).
[0202] result
[0203] Pre-treatment for 30 minutes, s 4SLAM-seq after 60 minutes of U labeling revealed a significant immediate response to small molecule inhibitors (Figures 25E and 30). This response was not biased by mRNA half-life, while changes at the total mRNA level were limited to a few short-lived mRNAs (Figure 25F). Monotherapy with all three inhibitors initiated specific and clear transcriptional responses (Figure 30C), whereas combined treatment with MEK and AKT inhibitors was similar to the effect observed after BCR / ABL suppression (Figures 25E and 25G). This is consistent with the function of these inhibitors in the major effector pathway of BCR / ABL. Taken together, these experimental studies establish SLAM-seq as a rapid, accessible, and scalable approach for investigating specific and overall transcriptional responses at the mature mRNA level, regardless of mRNA half-life, on a timescale that eliminates indirect effects. Therefore, by combining SLAM-seq with rapid disturbances of specific regulatory factors, it becomes possible to identify direct transcriptional target genes without ambiguity.
[0204] To generalize this approach and enable the investigation of numerous regulators for which selective inhibitors are not available, as in the case of BRD4, we attempted to combine SLAM-seq with chemogenetic proteolysis. To achieve a sufficiently fast reaction rate for the purpose of unambiguous target identification, we utilized an auxin-inducible degron (AID) system. This system degrades AID-tagged proteins within one hour. Specifically, the BRD4 locus in K562 cells was modified to affix a very small AID tag (Figure 26A), and homozygous tagged clones were transformed using a lentiviral vector expressing high levels of the rice F-box protein Tir1. Tir1 mediates the ubiquitination of AID-tagged proteins when treated with auxin (indole-3-acetic acid, IAA). Indeed, treatment of AID-tagged proteins with auxin initiated highly specific (Figures 31A, 31C) and nearly complete degradation of BRD4 within 30 minutes (Figures 26B, 31B). Tagging or Tir1 expression and auxin treatment were well tolerated, but prolonged degradation of BRD4 strongly inhibited proliferation (Figure 31D, Figure 31E). This is consistent with the important functions reported for BRD4.
[0205] Next, to accurately describe the direct transcriptional consequences of BRD4 degradation, cells were treated with auxin for 30 minutes, and newly synthesized RNA was transcribed. 4The mRNA was labeled with U for the following 60 minutes. Subsequent quantification of the labeled mRNA by SLAM-seq revealed overall downregulation of transcription, similar to the effect of CDK9 repression (Figure 26C, Figure 31F). To investigate the regulatory events behind this phenomenon, we measured the levels of the chromatin-bound core transcription mechanism when BRD4 was degraded. While components of the pre-initiation complex, as well as DSIF, NELF, and P-TEFb, did not impair overall recruitment, we noticed a significant decrease in Pol2 phosphorylation at the S2 position rather than S5 of its C-terminal base repeat (Figure 26D). This indicates a defect in promoter-proximal rest release. Indeed, from spike-in controlled ChIP sequencing of Pol2 when BRD4 was degraded (60 minutes after auxin treatment), Pol2 occupying the active transcription start site (TSS) increased significantly, while Pol2 density decreased throughout the transcriptional region of the gene (Figure 26E, Figure 26F, Figure 32A). Similarly, levels of S5-phosphorylated Pol2 increased at the promoter site, while S2-phosphorylated Pol2 (involved in the late elongation process) decreased significantly throughout the gene's transcriptional region (Figures 26E, 26F, 32B, and 32C). These findings are consistent with the widespread transcriptional decrease that occurs when pan-BET proteins are degraded independently of CDK9 recruitment to chromatin, indicating that the loss of BRD4 alone is sufficient to mediate these effects. Together, these results establish a central role for BRD4 in the licensing release of polymerases that are arrested at most active promoter sites.
[0206] These findings are consistent with BRD4 randomly binding to active TSSs and physically interacting with the core transcription mechanism, but contradict the selective effects observed after BETi treatment in conventional expression analyses. To clarify the immediate transcriptional effect of BETi and compare it to that of BRD4 degradation, SLAM-seq was performed after treatment with various doses of BETi JQ1 in K562 cells and the acute myeloid leukemia (AML) cell line MV4-11. In both cell types, treatment with high doses (1 μM or 5 μM) of JQ1 broadly repressed transcription (Figure 26G, Figure 33A) and reduced Pol2-S2 phosphorylation overall (Figure 33B). This is similar to the effect observed after BRD4 degradation and indicates that the comprehensive transcriptional function of BRD4 is dependent on the BET bromodomain. Importantly, the effect of high-dose BETi treatment on Pol2-S2 phosphorylation was also replicated by BRD4 knockdown at multiple time points prior to the antiproliferative effect (Figure 33C, Figure 33D), whereas repression of BRD2 or BRD3 (two other universally expressed BETi targets) did not initiate such a phenomenon. These results indicate that the comprehensive transcriptional effect of BETi is primarily mediated by BRD4 repression and cannot be compensated for by other BET-bromodomain-containing proteins.
[0207] Since doses of JQ1 exceeding 1 μM significantly exceed the inhibitory concentration in AML and other JQ1-sensitive cancer cell lines, we attempted to investigate the direct transcriptional response to a more carefully selected dose of 200 nM, which elicits a strong antileukemic effect in a wide range of AML models. In K562 cells, one of the few BETi-insensitive leukemia cell lines, 200 nM of JQ1 induced selective degradation of a small number of transcripts (Figure 26H). Surprisingly, treating two highly sensitive AML cell lines with the same dose initiated transcriptional responses on a comparable scale (Figures 26H, 34A, and 34B), affecting a similar set of BETi-hypersensitive transcripts (including MYC and other panmyelo-dependent leukemias) (Figures 26I, 34C, and 34D). These findings suggest that BETi resistance in leukemia is determined not by a lack of primary transcriptional response, but by secondary adaptation. We also noticed a small group of genes that were commonly upregulated after BRT repression or BRD4 degradation (Figure 34E). Interestingly, this included ERG1, a tumor suppressor in AML that may contribute to the potent effect of BETi in this context. Taken together, our results reveal that the primary transcriptional response to BETi is strongly dose-dependent, and that a therapeutically effective dose initiates the antileukemic effect by deregulating a small group of hypersensitive genes. Furthermore, this indicates that partial repression of the underlying transcriptional mechanism can induce a highly specific response that selectively targets cancer-dependent mechanisms.
[0208] To identify factors that make certain transcripts hypersensitive to BETi, we investigated whether this phenomenon is simply a reflection of a clear sensitivity to interference with the general Pol2 resting release mechanism. To verify this, we used SLAM-seq to compare the transcriptional response to BET suppression (200 nM JQ1) with the effects initiated by various doses of the selective CDK9 inhibitor NVP-2. High doses (60 nM NVP-2) suppressed CDK9, resulting in overall transcriptional repression, while intermediate doses (6 nM NVP-2) resulted in a selective transcriptional response distinctly different from the conventional response to BETi (Figures 27A, 27B, and 35B) (Figure 35A). Since CDK9 inhibitors and BET inhibitors have shown strong synergistic effects in previous reports and in our study in AML (Figures 35C and 35D), we attempted to investigate the transcriptional response behind this phenomenon. In contrast to selective monotherapy effects, the combination of intermediate doses of JQ1 and NVP-2 initiated a total loss of transcription, similar to CDK9 repression at high doses (Figures 27A, 27B, and 35A). These observations apply to genetically distinct AML cell lines (Figures 35E and 35F). This raises concerns about the tolerability of this combination, as it suggests that the therapeutic synergy between BETi and CDK9i is primarily based on synergistic repression of comprehensive transcription. Overall, our results reveal that therapeutically effective doses of CDK9 inhibitors and BET inhibitors, despite the common role of their targets in Pol2 rest release, exploit different bottlenecks in this process to initiate a selective transcriptional response.
[0209] To investigate whether the phenomenon of BETi hypersensitivity is determined by specific chromatin characteristics, we first examined whether BRD4 occupancy levels in TSS, or the accessibility of BRD4 to BETi, could distinguish direct BETi targets (FDR ≤ 0.1, log2FC ≤ -0.7) from cohorts of the same size consisting of non-responsive genes with the same baseline expression (FDR ≤ 0.1, -0.1 ≤ log2FC ≤ 0.1; Figure 36A). While BRD4 occupancy is unlikely to outweigh random gene selection (AUC 0.52, Figure 31C), recent reports on chromatin binding levels of BETi measured by Click-seq may partially explain the BETi response (AUC 0.63; Figure 36B). This suggests that differences in drug accessibility contribute to selective BETi effects. Another widely adopted model attributes the transcriptional and therapeutic effects of BETi to its ability to selectively suppress superenhancers, a model challenged by recent studies that identified H3K27ac-based regulatoryity as an excellent predictor of BETi targets. Because these studies relied on conventional RNA-seq after long-term drug treatment, we re-evaluated both models using SLAM-seq. In addition to H3K27ac-based regulatoryity of genes, the association of superenhancers and their potential predicted hypersensitivity to BETi with moderate accuracy (AUC 0.66 and 0.64, respectively, Figure 27C). However, two-thirds of BETi-sensitive genes could not be assigned to superenhancers, and the majority of expressed superenhancer-related genes did not respond to BETi treatment (Figures 27D, 27E). These observations apply to other leukemia cell lines (Figure 36C), indicating that sensitivity to BET suppression is associated with the presence of superenhancers, but not determined solely by their presence. This suggests that more complex factors are behind this phenomenon.
[0210] To explore this, we leveraged comprehensive profiling data available for K562 cells and devised an unbiased approach to model the combinational modes of gene regulation. Specifically, we extracted signals from 214 ChIP sequencing and methylome sequencing experiments within 500 bp and 2000 bp of the TSS of BETi-sensitive and non-responsive genes. This data was then used to train various classification models, which were later evaluated based on holdout test genes (Figure 27F, Figure 36D). This approach yielded numerous classifiers that predicted BETi sensitivity with high fidelity (AUC > 0.8, Figure 27G, Figure 36D), including a generalized linear model (GLM) derived by elastic network regression. Re-analysis of the coefficients in this model revealed that several factors (including high levels of proximal TSS REST and H3K27ac) are associated with BETi hypersensitivity, while high occupancy of SUPT5H (SUPT5H itself is a regulator of elongation) was the strongest negative predictor (Figures 27H and 37A). Unsupervised clustering revealed that most predictive factors and cofactors are abundant only in separate subclusters of BETi-sensitive or non-responsive genes (Figures 27I and 37B). For example, high loading of the positive predictors NFRKB and HMBOX1 was found in separate groups of BETi-sensitive genes, while high binding of CREM and SUPT5H was observed in separate subclusters of BETi-insensitive genes. Taken together, these findings suggest that the transcriptional response to BETi is determined by locus-specific regulators and cannot be predicted based on a single unified chromatin factor.
[0211] Similar to the complex determinants of BETi sensitivity, the therapeutic effects of BETi are likely to arise through the deregulation of numerous hypersensitive target genes. After confirming that MYC is a prominent BETi hypersensitive gene in leukemia, transcriptional and cellular responses to MYC repression should be considered as important effector mechanisms of BETi in this context. However, the direct gene regulatory function of MYC remains controversial, with studies describing its activating, repressive, and dose-dependent effects on specific targets, as well as studies describing its role as a general transcriptional amplification factor. To validate these models, we attempted to measure the direct changes in mRNA output after rapid loss of endogenous MYC. For this purpose, we manipulated the MYC locus in K562 cells and tagged it with an AID (Figure 28A). This tag induced rapid MYC degradation within 30 minutes in homozygous Tir1-expressing clones (Figure 28B). Next, we used SLAM-seq to quantify the output of newly synthesized mRNA 60 minutes after MYC degradation. Compared to the degradation of BRD4 and the drug-induced suppression of CDK9 and BET, the rapid loss of MYC resulted in highly specific changes rather than overall changes in mRNA production (Figure 28C). These changes were dominated by repressive effects on 712 genes, while only 15 mRNAs were strongly upregulated. Therefore, in K562 cells, MYC functions primarily as a transcriptional activator for specific target genes, rather than acting as a direct transcriptional repressor or general amplification factor.
[0212] Since MYC is known to occupy virtually all active promoters, we then investigated how MYC exerts selective transcriptional activation despite its universal binding. To this end, we trained a classification model to predict MYC-dependent transcripts (FDR ≤ 0.1, log2FC ≤ -1) based on different ChIP-seq signals at promoter locations. From elastic net regression, a simple GLM was derived that predicts MYC-dependent gene regulation well (AUC 0.91). The strongest contributing factor in this model was the content of MYC itself. Indeed, the presence of MYC at promoter locations determined by conventional peak calling fails to identify MYC-sensitive transcripts, while the binding levels of MYC or its cofactor MAX predict MYC-dependent gene regulation with moderate accuracy (AUCs 0.76 and 0.74, respectively). Taken together, these results suggest that transcripts directly dependent on MYC are determined by strong MYC binding and further alteration or compensation by additional factors (such as MNT, NKRF, TBL1XR1, EP300, and YY1).
[0213] To investigate the cellular function of MYC-dependent gene regulation, we analyzed the richness of biological processes among direct MYC target genes. Surprisingly, rapid loss of MYC primarily downregulated genes involved in protein and nucleotide biosynthesis (Figure 28D). These genes include 36% of all ribosomal biosynthesis factors, key regulators in AMP metabolism, and all six enzymes in the denovopurine synthesis pathway (Figures 28C, 28D). Indeed, MYC degradation gradually impairs protein synthesis (Figure 28E), leading to significant decreases in cellular AMP and GMP levels, as well as their upstream intermediate AICAR, before proliferation defects began to occur (Figure 28C, 28D). 8F). The role of MYC in directly regulating several subunits of polymerases I, II, and III, as well as being an important enzyme in protein and nucleotide biosynthesis, explains the reported increase in total cellular RNA when MYC is overexpressed. This supports the idea that these effects are secondary and not attributable to comprehensive transcriptional effects.
[0214] To verify whether the direct transcriptional function of MYC is conserved in other contexts, a homozygous AID-tag was inserted into the MYC locus of HCT116 colorectal cancer cells. These cells express MYC at particularly high levels. Regarding K562, HCT116 expresses Tir1. MYC-AID When cells were treated with auxin, complete degradation of MYC began within 30 minutes (Figure 28G). SLAM-seq profiling revealed a highly selective transcriptional effect (Figure 28H) that affected the same cellular process and correlated with the response in K562 cells (R=0.64, Figure 28H). To investigate whether the conservation of MYC targets between the two unrelated cell lines extends to other types of cancer, we derived signatures of the 100 most strongly downregulated genes in SLAM-seq and compared their expression to MYC levels in a panel of 672 cancer cell lines. Indeed, MYC expression levels correlated well with our signatures (Figure 28I), except for a small percentage of outliers that expressed MYC at low levels without losing their signatures. Notably, all of these outliers expressed MYCN or MYCL at high levels. This indicates that MYC paralogs have redundant functions in regulating core MYC targets. Our signature for a direct MYC target strongly correlated with MYC levels in TCGA RNA-seq profiles derived from 5583 primary patient samples across 11 major human cancers (Figure 28J). Together, these findings reveal that MYC drives the expression of a conserved set of transcriptional targets across various human cancers. This should be considered a starting point for inhibiting the oncogenic function of MYC.
[0215] In summary, combining rapid chemical-genetic disturbances with SLAM-seq establishes a simple yet powerful strategy for investigating the specific and comprehensive direct functions of transcription factors and cofactors. Using this approach, we characterized the function of BRD4, a factor widely studied as a regulator of cell lineage and disease-related expression programs, as a comprehensive cofactor in transcriptional pause-release. Meanwhile, MYC, previously considered a comprehensive transcriptional amplification factor, was found to activate a limited and conserved set of target genes, driving fundamental anabolic processes (particularly protein and nucleotide biosynthesis). More generally, SLAM-seq provides a simple, robust, and scalable method for defining direct transcriptional responses to any disturbance by directly quantifying changes in mRNA output, thereby exploring cellular regulatory networks.
Claims
1. A method for identifying polynucleic acid (PNA), A step of preparing PNA; a step of modifying one or more nucleic acid bases of the PNA by incorporating thiol-modified nucleic acid bases into the PNA through intracellular biosynthesis, and further modifying the thiol nucleic acid bases by alkylating them with an alkylating agent containing a hydrogen bonding partner, thereby changing the base pairing ability of the one or more nucleic acid bases; A step of base-pairing a complementary nucleic acid with its PNA, wherein this base-pairing includes base-pairing with at least one modified nucleic acid base; A method comprising the step of identifying the sequence of a complementary nucleic acid at a position complementary to at least one modified nucleic acid base.
2. The method according to claim 1, wherein the modification alters the behavior of base pairing, thereby altering the preferential base pairing between A and T / U and between C and G compared to natural nucleic acid bases selected from A, T / U, C, and G.
3. The method according to claim 1 or 2, wherein the modification comprises alkylating the 4-position of uridine using an alkylating agent containing the hydrogen bonding partner.
4. The method according to any one of claims 1 to 3, wherein the polynucleic acid comprises one or more 4-thiouridines or 6-thioguanosines.
5. The method according to any one of claims 1 to 4, wherein the polynucleic acid is synthesized in a cell by modification that alters the base pairing ability.
6. The method according to any one of claims 1 to 5, wherein base pairing with at least one modified nucleic acid base results in base pairing with another nucleotide, rather than base pairing with an unmodified nucleic acid base.
7. The method according to any one of claims 1 to 6, wherein the PNA includes or consists of RNA or DNA.
8. The method according to any one of claims 1 to 7, wherein, for each type of nucleotide selected from A, G, C, U, or T, the modified PNA contains more natural nucleotides than the modified nucleotide.
9. The method according to any one of claims 1 to 8, wherein the PNA comprises one, two, three, four, five, six, seven, eight, nine, ten, or more modified nucleotides, up to 30.
10. The PNA is PNA expressed in cells, and the PNA is further isolated from a sample containing the cells containing the PNA; After isolation, one or more nucleic acid bases of the PNA are modified; The method according to any one of claims 1 to 9, wherein the modification after isolation includes changing the base-pairing ability of one or more nucleic acid bases by adding or removing hydrogen bonding partners of one or more nucleic acid bases.
11. Culture or grow one or more cells in at least two culture or growth stages. Here, one culture or growth step comprises incorporating a modified nucleotide into a biosynthesized polynucleic acid, where the polynucleic acid is modified by the addition or removal of a hydrogen bonding partner. Another culture or growth stage lacks such incorporation of modified nucleotides into the biosynthesized polynucleic acid, or the modified nucleotides are incorporated into the biosynthesized polynucleic acid at a different concentration than in the other culture or growth stage; or, The method according to any one of claims 1 to 10, wherein the method comprises incorporating a biosynthesized polynucleic acid from at least two different cells, or a modified nucleotide from at least two different cell populations.
12. The method according to claim 11, wherein the method compares the incorporation of at least two different cells or at least two different cell populations.
13. The method according to claim 11 or 12, wherein the biosynthesized PNA from the two culture or proliferation steps is recovered from the cells, or the biosynthesized PNA from at least two different cells or two different cell populations is recovered from the cells, and the base pairing of a nucleic acid complementary to the PNA includes the generation of a complementary PNA chain by transcription.
14. The method according to claim 13, wherein the biosynthesized PNA is mixed.
15. The method according to claim 13 or 14, comprising labeling the PNA in accordance with the cell from which the PNA originates.
16. The method according to any one of claims 13 to 15, wherein the complementary PNA strand is a DNA strand.
17. The method according to any one of claims 13 to 16, wherein the transfer is a reverse transfer.
18. The method according to any one of claims 13 to 17, further comprising determining the sequence of the complementary PNA chain and comparing the sequences of the chains, wherein the complementary nucleic acid altered as a result of modification by addition or removal of hydrogen bonding partners can be identified by comparison with the complementary nucleic acid without modification.
19. The method according to any one of claims 1 to 18, comprising comparing the identified sequence of a complementary nucleic acid at a position complementary to at least one modified nucleic acid base in at least two cells, or at a position complementary to at least one modified nucleic acid base in at least two different growth stages within a cell, wherein the at least two cells or at least two growth stages have differences in gene expression between the at least two cells or at least two growth stages.
20. The method according to claim 19, wherein the difference in gene expression is caused by the suppression or stimulation of at least one gene in the cell.
Citation Information
Patent Citations
Classification of nucleic acid templates
US20110183320A1