RNA higher-order structure analysis method
The use of reactive alkylating agents for RNA modification and mutation profiling addresses inefficiencies in detecting RNA higher-order structures, particularly G4 structures, by providing precise and comprehensive structural analysis.
Patent Information
- Application Number
- JP2023510651
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-25
- Filing Date
- 2022-02-22
- Publication Date
- 2025-07-31
- Estimated Expiration
- 2042-02-22
AI Technical Summary
Existing methods for detecting RNA higher-order structures, particularly those involving non-Watson-Crick base pairs like G4 structures, are inefficient and cannot accurately reflect multiple equilibrium states, leading to incomplete or inaccurate structural analysis.
A method using reactive OFF-ON type alkylating agents covalently bonded to low-molecular-weight compounds for RNA modification, followed by mutation profiling to determine the base sequence and interaction positions, allowing for precise detection of RNA higher-order structures.
Enables efficient detection of a wider variety of RNA higher-order structures, including non-Watson-Crick base pair types, with improved accuracy and retention of sequence information, overcoming limitations of previous techniques.
Smart Images

Figure 0007716055000021 
Figure 0007716055000022 
Figure 0007716055000023
Abstract
Description
Cross-reference
[0001] This application claims priority based on Japanese Patent Application No. 2021-054713 filed on March 29, 2021, and Japanese Patent Application No. 2021-105526 filed on June 25, 2021, in Japan, and all the contents described in the applications are incorporated herein by reference as they are. In addition, all the contents described in all the patents, patent applications, and documents cited in this application are incorporated herein by reference as they are.
Technical Field
[0002] The present invention relates to a method for analyzing the higher-order structure of RNA and the like.
Background Art
[0003] RNA is a biomolecule that functions as a template for protein synthesis. On the other hand, RNA itself forms a highly folded higher-order structure and controls gene expression, intracellular localization of transcripts, splicing mechanisms, and the like. Many of these functional RNAs are defined by the fact that the bases as the primary sequence take a specific three-dimensional arrangement in structure formation. This RNA higher-order structure is formed from combinations of various structural motifs such as stem (STEM), stem-loop (STEM-LOOP), kissing-loop (KISSING-LOOP), multi-junction (MULTI-JUNCTION), kink-turn (KINK-TURN), pseudoknot (PSEUDOKNOT), and quadruplex (QUADRUPLEX). For example, a guanine quadruplex (hereinafter sometimes referred to as "G4") is a higher-order structure formed by a guanine (G)-rich sequence. The core structure of G4 is formed by four guanines through Hoogsteen hydrogen bonds. The monovalent metal cation coordinated to the O of guanine (Na 6 or K + or K +) enhances the stability of the G4 structure. A single-stranded RNA containing consecutive guanines can form a four-stranded helix structure in which G4s stack on top of each other within the folded structure. The types and combinations of these structural motifs, including G4, are enormous, and it is difficult to predict because they can take multiple equilibrium states. Therefore, in the research of RNA biology to understand the function of RNA, the development of technologies for actually measuring the higher-order structure of RNA is strongly demanded.
[0004] In recent years, technologies for determining the higher-order structure of RNA have been developed by combining chemical modification reactions for specific bases and sequence data obtained by parallel sequencers. For example, as methods using modification reactions for bases that do not form Watson-Crick base pairs, DMS-MaPseq using dimethyl sulfate (DMS) (Non-Patent Document 1), SHAPE-MaP that selectively modifies the 2-position carbon of the sugar of nucleic acids (Non-Patent Document 2), and Chem-CLIP-Map-Seq (Chemical Cross-Linking and Isolation by Pull-down to Map Small Molecule-RNA Binding Sites) using a cross-linking reaction at the binding position with low- and medium-molecular compounds (Non-Patent Document 3) are known. In Chem-CLIP-Map-Seq, it is possible to detect a specific RNA higher-order structure by using an RNA higher-order structure-specific binding molecule. Furthermore, as a technology for identifying the binding site between a low-molecular compound and RNA, a method using a modification reaction specific to the binding site of the low-molecular compound has been developed (Non-Patent Document 4, Patent Document 1).
[0005] On the other hand, reactive OFF-ON type alkylating agents have also been developed (Non-Patent Document 5) that exist as stable precursors until the low-molecular compound approaches the target DNA or RNA and are activated at the target site.
Prior Art Documents
Non-Patent Documents
[0006]
Non-Patent Document 1
[0007] [Patent Document 1] Japanese Patent Application Laid-Open No. 2019-511562 [Summary of the Invention] [Problems to be Solved by the Invention]
[0008] However, in the detection of RNA higher-order structures using the modification reactions disclosed in Non-Patent Document 1 and Non-Patent Document 2, a method is adopted in which the mutation information obtained by mutation profiling is given to RNA secondary structure prediction software, for example, RNAstructure. At this time, mainly the presence or absence of Watson-Crick base pairs is estimated, and the entire RNA higher-order structure is constructed. However, there are also higher-order structures of RNA that are difficult to identify only with the information of Watson-Crick base pairs. For example, the above-described G4 structure is a higher-order structure formed by guanine being arranged planar and layered by Hoogsteen hydrogen bonds. As functions of G4 in RNA, translational control and control of mRNA localization have been reported so far. Therefore, identifying G4 from intracellular transcripts is significant in RNA biology and nucleic acid chemistry. However, since the formation of G4 composed of Hoogsteen base pairs competes with Watson-Crick base pairs, their formation is antagonistic. Therefore, it is difficult to detect a structure such as G4 by the mutation profiling that discriminates the presence or absence of Watson-Crick base pairs used in the above-described SHAPE-MaP and DMS-MaPSeq. As an example, when SHAPE-MaP is used, the G4 possessed by the RNA of HIV-1 is presented as a stem structure composed of Watson-Crick base pairs.
[0009] In addition, in the structural detection using existing low-molecular-weight compounds (for example, the methods disclosed in Non-Patent Document 3 and Non-Patent Document 4), the modification position is identified by regarding the stop position of cDNA synthesis during reverse transcription as a modified base. Therefore, there is a problem that only single information can be obtained from one RNA molecule. For example, when there are two higher-order structures to be detected in one RNA molecule, only one of the information can be obtained. It is less efficient than mutation profiling in that the information on the structure after the reverse transcription termination position is lost. In addition, there is a drawback that the modification patterns co-occurring at multiple positions cannot be measured and the true structure cannot be reflected. Therefore, an object of the present invention is to establish a technique for efficiently detecting a wider variety of RNA higher-order structures including non-Watson-Crick base pair-type higher-order structures.
Means for Solving the Problems
[0010] The present invention has been made to solve the above problems, and provides a structural detection technique by mutation profiling using a reactive OFF-ON type alkylating agent covalently bonded to a low-molecular-weight compound as a modifying molecule.
[0011] That is, a method for analyzing the higher-order structure of RNA according to one embodiment of the present invention is represented by the following formula (I), (II), (III) or (IV):
Chemical formula
[0012] Preferred embodiments and other embodiments of the above method will be described in detail in the embodiments for carrying out the following invention.
Advantages of the Invention
[0013] According to the method of the present invention, a wider variety of RNA higher-order structures, including non-Watson-Crick base pair type higher-order structures, can be efficiently detected.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Mode for Carrying Out the Invention
[0015] Next, each embodiment of the present invention will be described with reference to the drawings. Note that each embodiment described below does not limit the invention according to the claims, and not all of the elements and combinations thereof described in each embodiment are essential for the solution means of the present invention.
[0016] (Definition) As used herein, the higher-order structure of RNA refers to, in solution, mainly a secondary structure such as a stem-loop including partial double-strand formation based on base pairing within the molecule, a single-stranded structure or a circular single-stranded structure in a portion without such base pairing, a tertiary structure such as a junction or a pseudoknot, and further a quaternary structure composed of a complex of these. Also included in the higher-order structure of RNA are a triple strand formed by inserting a nucleoside not involved in base pairing into the minor groove of an RNA double helix, and a guanine quartet in which four guanine bases form a planar structure by Hoogsteen-type hydrogen bonds and are stacked. Furthermore, motifs called coaxial stacking include kissing loops and pseudoknots. In a kissing loop, the single-stranded loop regions of two hairpins interact by base pairing, and a helix is synthesized by coaxial stacking. The pseudoknot motif is generated when the single-stranded region of a hairpin loop forms a base pair with a sequence upstream or downstream of the same RNA strand. Such a structure is in a specific equilibrium state depending on the state of the solution (temperature, salt concentration, etc.) and fluctuates with the movement of the RNA molecule.
[0017] The term "motif" or "motif region" means a functional structural unit that includes the above-described higher-order structure of RNA and through which the RNA interacts with a target substance. The motif region contained in the RNA to be analyzed for its higher-order structure in the present invention may consist of a single stem-loop structure (hairpin loop structure), or may include a plurality of stem-loop structures (multi-branched loop structure) or other higher-order structures.
[0018] "Target" or "target RNA" refers to an RNA that contains such an RNA motif and can be a target for gene expression control in cells and therapeutic intervention by small molecules. It is understood that various RNA molecules play important regulatory roles in both normal and diseased cells. Non-coding transcripts (non-coding transcriptome) represent a large group of new therapeutic targets. Non-coding RNAs such as microRNA (miRNA) and long non-coding RNA (lncRNA) regulate transcription, splicing, mRNA stability / degradation, and translation. In addition, non-coding regions of mRNA such as the 5' untranslated region (5'UTR), 3'UTR, and introns can play regulatory roles in the expression level of mRNA, alternative splicing, translation efficiency, and the intracellular localization of mRNA and protein. The higher-order structure of RNA is extremely important for these regulatory activities.
[0019] (Compound Design and Its Embodiments) The compounds used in the present invention have the following structure in which the target-binding moiety Sm and the RNA-modifying moiety Y are linked via a linker L. [Chemical formula]
[0020] [Target-binding moiety] The target-binding portion is a portion that interacts with the higher-order structure formed by RNA, preferably a specific RNA structural motif. Novel compounds that interact with RNA forming higher-order structures in vivo have great therapeutic potential. For example, branaplam is known to recognize the bulge structure in the stem of exon 7 of SMN2 (Campagne, S., Boigner, S., Rudisser, S. et al. Structural basis of a small molecule targeting RNA for a specific splicing correction. Nat Chem Biol 15, 1191-1198 (2019). https: / / doi.org / 10.1038 / s41589-019-0384-5), and ribocil is known to recognize the multi-branched loop structure of the FMN riboswitch (Howe, J., Wang, H., Fischmann, T. et al. Selective small-molecule inhibition of an RNA structural element. Nature 526, 672-677 (2015). https: / / doi.org / 10.1038 / nature15542).
[0021] [[ID=
[0022] To date, approximately 1000 small molecules targeting G4 structures have been reported in the G-Quadruplex Ligands Database (http: / / www.g4ldb.org / ). Small molecule G4 binders generally have an aromatic surface for π-π stacking with G-tetrads, a positive or basic group that binds to the loop or groove of G4, and steric bulk to prevent intercalation with double-stranded DNA.
[0023] Therefore, in certain embodiments, the target binding moiety is selected to have a structure that binds to RNA from any compound or a portion thereof. In one embodiment, the G4 binders described above are included. Specific G4 binders include, but are not limited to, acridine, berberine, pyridostatin, porphyrin derivatives such as TMPyP4, and macrocyclic compounds such as telomestatin. As other embodiments, triphenylene-based scaffold structures that stabilize RNA 3-way Junctions have been reported (S.A. Barros and D.M. Chenoweth, Recognition of Nucleic Acid Junctions Using Triptycene Based Molecules, Angew Chem Int Ed Engl. 2014, 53 (50), pp. 13746-50). In still other embodiments, some small molecule compounds in clinical or preclinical trials that act on various RNAs, as shown in FIG. 8, are included.
[0024] <RNA modification moiety> The RNA modification moiety in this embodiment has a structure that is activated by contacting RNA from an inactive precursor and is part of a compound represented by the following formula (I), (II), (III), or (IV).
Chemical formula
[0025] In the formula, Sm represents the target binding moiety described above. L represents a linker that connects the target binding moiety and the RNA modification moiety, and X is -S-R 4 , -S(O)-R 4 , -O-R 5 or -N(R 6 )-R 7 wherein R 1 , R 2 and R 3 are each independently a hydrogen atom, a halogen, an alkyl which may have a substituent, an alkenyl which may have a substituent, an alkynyl which may have a substituent, an alkoxy which may have a substituent, an aryl which may have a substituent, an aralkyl which may have a substituent, a cycloalkyl which may have a substituent, or a heteroaryl which may have a substituent, or R 1 and R 2 or R 2 and R 3 together form a ring which may have a substituent, R 4 represents an alkyl which may have a substituent, an aryl which may have a substituent or a heteroarylalkyl which may have a substituent, R 5 represents a hydrogen atom or an alkyl which may have a substituent, R 6 and R 7 are each independently a hydrogen atom, an alkyl which may have a substituent or an aryl which may have a substituent, or R 6 and R 7 together form a ring which may have a substituent.
[0026] Here, the "alkyl" in the "alkyl which may have a substituent" represented by R 1 ~R 7 is usually a linear or branched alkyl having 1 to 15 carbon atoms (C 1-15means "(alkyl)", and examples thereof include methyl, ethyl, propyl, isopropyl, butyl, isobutyl, sec-butyl, tert-butyl, pentyl, isopentyl, neopentyl, hexyl, heptyl, octyl, nonyl, decyl, and the like. Preferably, C such as methyl, ethyl, propyl, isopropyl, butyl, isobutyl, sec-butyl, tert-butyl or pentyl 1-6 alkyl is mentioned, more preferably methyl or ethyl is mentioned, and most preferably methyl.
[0027] R 1 ~R 3 The "alkenyl" in the "optionally substituted alkenyl" represented by is, for example, a linear or branched alkenyl having 2 to 10 carbon atoms (C 2-10 alkenyl). Specifically, vinyl, allyl, 1-propenyl, isopropenyl, methacryl, butenyl, crotyl, pentenyl, hexenyl, heptenyl, octenyl, nonenyl, decenyl, and the like.
[0028] Similarly, the "alkynyl" in the "optionally substituted alkynyl" represented by R 1 ~R 3 is, for example, a linear or branched alkynyl having 2 to 10 carbon atoms (C 2-10 alkynyl). Specifically, ethynyl, propargyl, butynyl, pentynyl, hexynyl, heptynyl, octynyl, nonynyl, decynyl, and the like.
[0029] R 1 ~R 3 The "alkoxy" in the "optionally substituted alkoxy" represented by is, for example, a linear or branched alkoxy having 1 to 15 carbon atoms (C 1-15 alkoxy). Specifically, methoxy, ethoxy. Also in this specification, "halo-C 1-15 alkoxy" includes the above C 1-15 alkoxy substituted with one or more halogen atoms.
[0030] R1 ~R 7 The "aryl" in the "aryl which may have a substituent" represented by 7 means aryl (C 6-14 aryl) having 6 to 14 carbon atoms, for example, phenyl, naphthyl, or an ortho-fused bicyclic group having 8 to 10 ring atoms and at least one ring being an aromatic ring (such as indenyl, etc.).
[0031] R 1 ~R 3 The "aralkyl" in the "aralkyl which may have a substituent" represented by 3 is "arylalkyl" having an alkyl group having 1 to 8 carbon atoms which may be linear or branched. For example, benzyl, benzhydryl, 1-phenylethyl, 2-phenylethyl, phenylpropyl, phenylbutyl, phenylpentyl, phenylhexyl, naphthylmethyl, naphthylethyl, etc., C 6-14 aryl-C 1-8 alkyl can be mentioned, but benzyl or naphthylmethyl is preferred.
[0032] R 1 ~R 3 The "cycloalkyl" in the "cycloalkyl which may have a substituent" represented by 3 means cycloalkyl (C 3-7 cycloalkyl) having 3 to 7 carbon atoms. Specifically, for example, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, etc. can be mentioned. Preferably, cyclopropyl, cyclobutyl, cyclopentyl or cyclohexyl, and more preferably cyclopropyl or cyclobutyl can be mentioned.
[0033] R 1 ~R 4The "heteroaryl" in the "optionally substituted heteroaryl" represented by means a 5- to 7-membered aromatic heterocyclic (monocyclic) group containing 1 to 4 heteroatoms selected from nitrogen, sulfur, and oxygen atoms in addition to carbon atoms as ring atoms. Examples include furyl, thienyl, pyrrolyl, thiazolyl, pyrazolyl, oxazolyl, isoxazolyl, isothiazolyl, imidazolyl, 1,2,4-oxadiazolyl, 1,3,4-oxadiazolyl, 1,2,3-triazolyl, 1,2,4-triazolyl, 1,2,4-thiadiazolyl, 1,3,4-thiadiazolyl, tetrazolyl, pyridyl, pyrimidinyl, pyrazinyl, pyridazinyl, 1,3,5-triazinyl, azepinyl, diazepinyl, etc. Further, the "heteroaryl" includes a group derived from an aromatic heterocyclic ring (bicyclic or higher) formed by condensing a 5- to 7-membered aromatic heterocyclic ring containing 1 to 4 heteroatoms selected from nitrogen, sulfur, and oxygen atoms in addition to carbon atoms as ring atoms with a benzene ring or the above aromatic heterocyclic (monocyclic) group. Examples include indolyl, isoindolyl, benzo[b]furyl, benzo[b]thienyl, benzimidazolyl, benzoxazolyl, benzoisoxazolyl, benzothiazolyl, benzoisothiazolyl, quinolinyl, isoquinolinyl, etc.
[0034] Examples of the substituents in the optionally substituted alkyl, optionally substituted alkenyl, optionally substituted alkynyl, and optionally substituted alkoxy include, the same or different, for example, a halogen atom, C 1-15 alkyl (preferably C 1-6 alkyl), halo-C 1-15 alkyl, C 1-15 alkoxy, halo-C 1-15 alkoxy, hydroxy, nitro, cyano, and amino, etc. Further, in this specification, examples of the "halogen atom" include a fluorine atom, a bromine atom, a chlorine atom, and an iodine atom. Preferably, they are a bromine atom and a chlorine atom.
[0035] Aryl which may have a substituent, aralkyl which may have a substituent, and substituents in a ring which may have a substituent may be the same or different and include, for example, halogen with 1 to 3 substituents, hydroxy, sulfanyl, nitro, cyano, carboxy, carbamoyl, C 1-10 alkyl, trifluoromethyl, C 3-8 cycloalkyl, C 6-14 aryl, aliphatic heterocyclic group, aromatic heterocyclic group, C 1-10 alkoxy, C 3-8 cycloalkoxy, C 6-14 aryloxy, C 7-16 aralkyloxy, C 1-8 alkanoyloxy, C 7-15 aroyloxy, C 1-10 alkylsulfanyl, C 1-8 alkanoyl, C 7-15 aroyl, C 1-10 alkoxycarbonyl, C 6-14 aryloxycarbonyl, C 1-10 alkylcarbamoyl and diC 1-10 Substituents such as those selected from the group consisting of alkylcarbamoyl are exemplified, and preferably halogen with 1 substituent, hydroxy, sulfanyl, nitro, cyano, carboxy, C 1-3 alkyl, trifluoromethyl, C 1-3 alkoxy and the like are exemplified.
[0036] The RNA modification moiety of this embodiment promotes activation from an inactive precursor by interacting with the target RNA. For example, the RNA modification moiety contained in the compound of formula (I) is considered to promote activation only in the presence of the target RNA by an elimination reaction of a conjugate base molecule (E1cB reaction) as shown in the following scheme.
Chemical formula
[0037] Since the vinyl group in the active compound is bonded to an electron-withdrawing carbonyl group, it is expected to be highly reactive. Therefore, the inactive compound of formula (I) was converted into a precursor compound by protecting this highly reactive vinyl group with several functional groups (X) shown below. Scheme 1 shows the reaction mechanism in which the leaving group X is eliminated when the target binding moiety Sm reaches and interacts with the target RNA. The promotion of activation is thought to occur by the abstraction of a hydrogen atom by a neighboring available nucleobase and phosphate backbone (denoted as :B in Scheme 1) to which the target binding moiety Sm is attached. The resulting reactive RNA modification site (vinyl group) then efficiently alkylates the target base.
[0038] To meet this objective, various thiol groups and sulfoxide groups can be used as the leaving group X. For example, X is -S-R 4 , -S(O)-R 4 , -O-R 5 or -N(R 6 )-R 7 , where R 4 represents alkyl which may have a substituent, aryl which may have a substituent or heteroarylalkyl which may have a substituent, R 5 represents a hydrogen atom or alkyl which may have a substituent, and R 6 and R 7 each independently represent a hydrogen atom, alkyl which may have a substituent or aryl which may have a substituent, or R 6 and R 7 together form a ring which may have a substituent.
[0039] Preferred examples of X include -S-C 1-6 alkyl, -S-aryl, -S(O)-C 1-6 alkyl, -S(O)-aryl, -O-H or -N(C 1-6is (alkyl)2, more preferably -S-CH3, -S-phenyl, -S(O)-CH3, -S(O)-phenyl, -O-H or -N(CH3)2. The above phenyl may be substituted at the para position, meta position or para position with methoxy, methyl, fluorine atom, chlorine atom or bromine atom, etc.
[0040] Regarding the compound represented by the above formula (II), (III) or (IV), similar to the compound of formula (I), an ethylene group having a leaving group X that enables an elimination reaction (E1cB reaction) of one molecule of the conjugate base in the 6-membered ring containing a nitrogen atom is bonded. For this reason, an active vinyl form can be generated by the same mechanism as the compound of formula (I), and it is considered that it can be an OFF-ON type RNA modifier.
[0041] In a preferred embodiment of the present invention, the RNA modification moiety (Y) is a vinylquinazolinone precursor (VQ) represented by the following formula (V). [Chemical formula]
[0042] In the formula, Sm, L and X have the same meanings as above, and R 8 , R 9 , R 10 and R 11 independently of each other represent a hydrogen atom, a halogen, alkyl which may have a substituent, alkenyl which may have a substituent, alkynyl which may have a substituent, alkoxy which may have a substituent, aryl which may have a substituent, aralkyl which may have a substituent, cycloalkyl which may have a substituent, or heteroaryl which may have a substituent.
[0043] R 8 A preferred example of is a hydrogen atom, a halogen or C 1-15 alkyl, more preferably a hydrogen atom or C 1-6 alkyl, and most preferably a hydrogen atom. R 9Suitable examples thereof include a hydrogen atom, a C 1-15 alkyl which may have a substituent, a C 1-15 alkynyl or a heteroaryl which may have a substituent, more preferably a hydrogen atom or the following formula (VI) or (VII):
[0044]
Chemical formula
[0045] R 10 Suitable examples thereof include a hydrogen atom, a halogen or a C 1-15 alkyl, more preferably a hydrogen atom or a C 1-6 alkyl, and most preferably a hydrogen atom.
[0046] R 11 Suitable examples thereof include a hydrogen atom, a halogen or a C 1-15 alkyl, more preferably a hydrogen atom or a C 1-6 alkyl, and most preferably a hydrogen atom.
[0047] Suitable examples of X include -S-R 4 or -S(O)-R 4 wherein R 4 is methyl, hydroxyethyl, 2-pyridylmethyl or a phenyl which may have a substituent. In another embodiment, X is -N(R 6 )-R 7 wherein R 6 and R 7 are each independently a hydrogen atom, methyl or a phenyl which may have a substituent, or R 6 and R 7 together may form a cycloalkyl ring which may have a substituent, a morpholine ring which may have a substituent or a piperazine ring which may have a substituent.
[0048] <Linker> The present invention can link a target-binding moiety Sm and an RNA-modifying moiety Y using various divalent or trivalent linkers, providing optimal binding and reactivity to bases proximal to the binding site of the target RNA. For example, in one embodiment, the linker is a polyethylene glycol (PEG) group of, for example, 1 to 20 ethylene glycol subunits. In other embodiments, the linker is an optionally substituted C1-12 aliphatic group or a peptide containing 1 to 8 amino acids.
[0049] Preferred examples of the linker L include -(C2H4-O) n -C2H4- (n is an integer from 1 to 5, preferably 2 or 3.) and -CONH-(C2H4-O-C2H4) m -NHCO- (m is an integer from 1 to 5, preferably 1 or 2.) and the like can be mentioned.
[0050] <Synthesis method of the compound of this embodiment> The compounds of the present invention can generally be prepared or isolated by synthetic and / or semi-synthetic methods known to those skilled in the art for similar compounds, as well as by the methods described in detail in the examples and drawings herein. For example, various compounds of the present invention can be synthesized with reference to Schemes 2 to 9 described below.
[0051] In specific protecting groups ("PG"), leaving groups ("LG"), or detailed descriptions, schemes, and chemical reactions indicating conversion conditions in the examples, other protecting groups, leaving groups, and conversion conditions can be easily used according to the common general knowledge of those skilled in the art. As used herein, the expression "leaving group" (LG) includes, but is not limited to, halogens (e.g., fluoride, chloride, bromide, iodide), sulfonates (e.g., mesylate, tosylate, benzenesulfonate, brosylate, nosylate, triflate), diazonium, and the like.
[0052] As used herein, the expression "oxygen protecting group" includes, for example, carbonyl protecting groups, hydroxyl protecting groups, etc. Hydroxyl protecting groups are well known in the art. Suitable hydroxyl protecting groups include, for example, esters, allyl ethers, ethers, silyl ethers, alkyl ethers, arylalkyl ethers, and alkoxyalkyl ethers, but are not limited thereto. Such esters include, for example, formates, acetates, carbonates, and sulfonates.
[0053] Amino protecting groups are also well known in the art. Suitable amino protecting groups include, but are not limited to, aralkylamines, carbamates, cyclic imides, allylamines, amides, etc. Such groups include, for example, t-butyloxycarbonyl (BOC), ethyloxycarbonyl, methyloxycarbonyl, trichloroethyloxycarbonyl, allyloxycarbonyl (Alloc), benzyloxycarbonyl (CBZ), allyl, phthalimide, benzyl (Bn), fluorenylmethylcarbonyl (Fmoc), formyl, acetyl, chloroacetyl, dichloroacetyl, trichloroacetyl, phenylacetyl, trifluoroacetyl, benzoyl, etc.
[0054] Those skilled in the art will understand that various functional groups present in the compounds of the present invention, such as, for example, aliphatic groups, alcohols, carboxylic acids, esters, amides, aldehydes, halogens, and nitriles, can be interconverted by techniques well known in the art (including, but not limited to, reduction, oxidation, esterification, hydrolysis, partial oxidation, partial reduction, halogenation, dehydration, partial hydration, and hydration).
[0055] (Method for Analyzing Higher-Order Structure of RNA) FIG. 1 is a flowchart showing a method for analyzing the higher-order structure of RNA according to an embodiment of the present invention. This method includes a step (S10) of preparing a compound represented by the above-described formula (I), (II), (III), or (IV), a step (S20) of preparing a target RNA to be analyzed, a step (S30) of contacting these compounds with one or more target RNAs to modify this RNA, a step (S40) of detecting modified bases by determining the base sequence of the RNA modified in step S30, and a step (S50) of analyzing the higher-order structure of RNA by determining the position and / or region on the RNA that interacts with the target-binding portion of the above compound based on the determined base sequence. Note that the preparation step (S10) of the above compound has already been described.
[0056] <Step (S20) of Preparing Target RNA> The target RNA is the RNA to be analyzed. A mixture of one or more types of RNAs can be used, and it may be extracted from a living body or artificially synthesized. This target RNA preferably contains a motif region for exerting its function in a living body. The motif region may consist of a single stem-loop structure (hairpin loop structure) or may include a plurality of stem-loop structures (multibranched loop structure). In the present embodiment, it is possible to include a motif region extracted based on the stem structure (see, for example, the specification of WO2018 / 003809). Thereby, a target RNA reflecting the functional structural unit existing in the RNA can be prepared without fragmenting the motif region. The motif region may have an arbitrary sequence length as long as its function is maintained, for example, 1000 bases or less, 900 bases or less, 800 bases or less, 700 bases or less, 600 bases or less, 500 bases or less, 400 bases or less, 300 bases or less, 200 bases or less, 150 bases or less, 100 bases or less, 50 bases or less.
[0057] The target RNA of the present embodiment can be synthesized by any conventionally known genetic engineering method. Preferably, the target RNA can be produced by transcribing template DNA synthesized by commissioning a contract manufacturer for synthesis. In order to perform transcription from DNA to RNA, the DNA containing the sequence of the target RNA may have a promoter sequence. Although not particularly limited, as a preferable promoter sequence, a T7 promoter sequence is exemplified. When the T7 promoter sequence is used, for example, RNA can be transcribed from DNA having a desired target RNA sequence using the MEGAshortscript TM T7 Transcription Kit provided by Life Technologies. In the present embodiment, the RNA may be modified RNA as well as adenine, guanine, cytosine, and uracil. Modified RNAs include, for example, pseudouridine, 5-methylcytosine, 5-methyluridine, 2'-O-methyluridine, 2-thiouridine, and N 6 -methyladenosine.
[0058] In one embodiment, the target RNA may be used as a target RNA library containing a plurality of target RNAs having different sequences. In this embodiment, it is preferable to synthesize a variety of target RNAs simultaneously, and it can be carried out using oligonucleotide library synthesis technology. This uses an inkjet technique that prints individual bases at defined positions on a slide to synthesize one base at a time and extend a template DNA of a specified length. Next, the constructed oligos are cleaved from the slide, pooled, dried, and stored in a single tube. The oligo library can then be redissolved, amplified, and a target RNA library can be prepared by an in vitro transcription reaction. Although not particularly limited in the present invention, oligonucleotide library synthesis can be produced by commissioning Agilent Technologies or Twist Bioscience.
[0059] <Modification Step of Target RNA (S30)> The compound synthesized in Step S10 is added to a solution containing the target RNA prepared in Step S20 to bring the compound into contact with the target RNA. This solution may be a solution containing the compound at different concentrations and amounts. It may also contain various surfactants, polymers, and osmotic agents. It may also be a biological solution containing proteins, cells, viruses, lipids, monosaccharides and polysaccharides, amino acids, nucleotides, DNA, as well as various salts and metabolites at different concentrations and amounts. The concentration of the compound can be adjusted to specifically bind to a specific motif of the target RNA.
[0060] Furthermore, when the reactivity of the RNA-modifying moiety of the compound is pH-dependent, the pH may be maintained, for example, in the range of 6.5 to 8.0, but not limited thereto. The RNA can be replaced by any procedure that folds it into the desired conformation at the desired pH (e.g., about pH 7). First, this RNA is heated to eliminate the multimeric form and then quickly cooled in a low ionic strength buffer. Subsequently, a folding solution is added so that the RNA can achieve the correct conformation and react with the compound of the present embodiment.
[0061] <Detection Step of Modified Bases (S40)> This step is to detect modified bases by determining the nucleotide sequence of the RNA obtained in the above-described modification step (S30). There is no particular limitation as long as it is a method for reading modified bases in the RNA sequence. For example, it may be a pull-down method using an antibody specific to the modified base or a nanopore sequencing method for directly reading the potential of RNA. This direct RNA nanopore sequencing method is a technique for detecting modified sites of RNA at the single-molecule level. Currently, in the direct RNA sequencing platform developed and commercially available from Oxford Nanopore Technologies, RNA bound to a motor protein moves through a biological nanopore suspended in a membrane. When RNA passes through the pore under a voltage bias, changes in picoampere ion current are observed depending on the chemical identity (i.e., sequence) of a short sequence (5 nucleotides) passing through the pore constriction (see Garalde, D. R., et al. (2018) Highly parallel direct RNA sequencing on an array of nanopores. Nat. Methods, and Workman, R. E., et al. (2019) Nanopore native RNA sequencing of a human poly(A) transcriptome. Nat. Methods, 16, 1297-1305).
[0062] In a preferred embodiment, the step of detecting modified bases (S30) is mutation profiling (MaP) including conversion of RNA to complementary DNA (cDNA). In this embodiment, first, cDNA is synthesized using one or more target RNAs obtained in step S30 as a template by a reverse transcriptase or other polymerase. A reverse transcriptase is an enzyme that synthesizes cDNA from RNA, and examples include, but are not limited to, thermostable enzymes such as mouse or avian reverse transcriptases. Alternatively, it may be a thermostable group II intron reverse transcriptase (TGIRT) present in retrotransposons of prokaryotes or fungi, etc.
[0063] As shown in Fig. 2, these enzymes terminate the reverse transcription reaction at or near the position alkylated by the RNA modification moiety (Y) on the target RNA, or skip the alkylated nucleotide, incorporating an inaccurate (non-complementary) nucleotide at this modified site on the cDNA. This step includes the step of detecting chemical modifications in RNA by such a method. As used herein, "inaccurate" with respect to nucleotide incorporation means incorporating a non-complementary nucleotide (a nucleotide that violates the Watson-Crick rule) into the nucleotide present in the original sequence. This includes deletions within the sequence. Also, as disclosed in Non-Patent Document 3 and Non-Patent Document 4, it is also possible to detect this RNA modification site by terminating the reverse transcription reaction.
[0064] Subsequently, the nucleotide sequence of the cDNA is determined and the plurality of read nucleotide sequences are aligned. The cDNA can efficiently detect chemical modifications in nucleic acids such as RNA using the massively parallel sequencing method (MPS) by using a library derived from a mixture of various types of target RNAs. As an example, in the next-generation sequencer of Illumina, the 5'-end side is fixed on the flow cell via adapters at both ends of tens of millions to hundreds of millions of DNA fragments. Next, the 5'-end side adapter pre-fixed on the flow cell is annealed to the adapter sequence on the 3'-end side of the DNA fragment to form a bridge-shaped DNA fragment. In this state, by performing a nucleic acid amplification reaction with DNA polymerase, a large number of single-stranded DNA fragments can be locally amplified and fixed. Then, in the next-generation sequencer, by performing sequencing using the obtained single-stranded DNA as a template, as of 2020, extremely large sequence information of about 3 Tb can be obtained in one analysis.
[0065] In one embodiment, the sequence data (reads) obtained by a next-generation sequencer are aligned in a form including barcode sequences. This is because by aligning the sequence data for each individual barcode sequence, it is possible to sequence samples containing multiple types of target RNAs simultaneously. Also, even when the RNAs to be analyzed have similar sequences, such as those belonging to a gene family or single nucleotide polymorphisms, it becomes possible to identify and analyze them. A "barcode sequence" is a tag having a unique sequence added to each type or each molecule of a nucleic acid molecule. If barcode sequences having different unique sequences are added to each of the multiple RNAs to be analyzed, after simultaneously modifying and amplifying the multiple RNAs, each RNA can be identified and analyzed based on the types of the added barcodes.
[0066] Alternatively, after aligning all the cDNAs together, the alignment of the low-confidence alignments may be evaluated by taking into account the mutation information of the barcodes. In any method, the accuracy of the sequence information can be improved by aligning the RNA sequence to be analyzed together with the barcode sequence.
[0067] Based on the nucleotide sequence aligned in this way, the positions and frequencies of the generated mutations are detected. The mutation rate at a given nucleotide is simply the number of mutations (mismatches, deletions, and insertions) at that location divided by the number of reads. The data obtained by calculating the raw reactivity for each nucleotide can be normalized using various criteria. Quality control of the data is possible by considering the sequencing read depth and standard error.
[0068] <Step for analyzing the higher-order structure of RNA (S50)> Based on the position and frequency of the mutation on the target RNA detected in the above step S40, the higher-order structure formed by the target RNA can be analyzed. For example, if it is known that the target-binding moiety Sm in a compound interacts with a specific RNA structural motif, the higher-order structure formed by the target RNA can be estimated based on that information. For example, it is a method of predicting the G4 structure of RNA using a G4 binder. Alternatively, when a specific compound without such information is used as the target-binding moiety Sm, it is presumed that the RNA region interacting with this is the binding site with the compound. Therefore, in one embodiment, any compound or a part thereof can be used as the target-binding moiety to identify the RNA that interacts with the any compound from among a plurality of target RNAs.
[0069] And based on the three-dimensional structure of the RNA to which an arbitrary compound binds, it is also possible to estimate the three-dimensional structure formed by the RNA region, for example, the structure of the binding pocket of the target-binding moiety (this is also referred to as the "ligand-binding pocket") and the pharmacophore complementary to this. The structure of such a binding pocket or the pharmacophore is also included in one of the higher-order structures of RNA. The binding pocket refers to a pore or depression having a size sufficient for a ligand molecule to bind, observed on the molecular surface of RNA forming a higher-order structure. The pharmacophore refers to an aggregate of steric and electronic features necessary to ensure an optimal supramolecular interaction with a specific biological target and to cause (or block) a biological reaction. For example, using a compound that recognizes a complex RNA structural motif with high drug discovery potential, such as a 3-way junction structure, etc., leads to a comprehensive discovery of RNA structures with high drug discovery potential.
[0070] (Method for identifying the structure of a target-binding moiety that regulates the function of a target RNA) In another embodiment of the present invention, a plurality of compounds represented by the above-described formulas (I), (II), (III), or (IV) are prepared, and a step of contacting these plurality of compounds with one or more target RNAs, and a step of determining the nucleotide sequence of the target RNA contacted with these compounds, and a step of selecting a compound that interacts with each target RNA based on the determined nucleotide sequence, and a method for identifying the structure of a target-binding moiety that regulates the function of the target RNA is provided.
[0071] The structure of this target-binding moiety is important for the development of small-molecule compounds with beneficial pharmacological activity. Small-molecule compounds can be optimized to exhibit excellent absorption from the intestine, excellent distribution to target organs, and excellent cell permeability. Small-molecule compounds can be used to modulate pre-mRNA splicing. One example is spinal muscular atrophy (SMA), which is also associated with some of the compounds shown in FIG. 8. SMA is the result of insufficient levels of the survival motor neuron (SMN) protein. Humans have two SMN genes, SMN1 and SMN2. Since SMA patients have a mutated SMN1 gene, the SMN protein in these patients depends only on SMN2. The SMN2 gene has a silent mutation in exon 7 that causes inefficient splicing, so exon 7 is skipped in most of the SMN2 transcripts, leading to the production of defective proteins that are rapidly degraded within the cell. As a result, the amount of SMN protein produced from this locus is limited. Small-molecule compounds that promote the efficient inclusion of exon 7 during the splicing of SMN2 transcripts would be an effective treatment for SMA. Thus, in one aspect, the present invention provides a method for identifying the structure of a target-binding moiety that modulates the splicing of a target pre-mRNA to treat a disease or disorder, the method comprising contacting the target pre-mRNA with one or more compounds represented by formula (I), (II), (III), or (IV), and selecting a compound that interacts with the target RNA by analyzing the results of the analysis of the higher-order structure of the RNA disclosed herein. In some embodiments, the pre-mRNA is an SMN2 transcript. In some embodiments, the disease or disorder is spinal muscular atrophy (SMA).
[0072] As an example of defective splicing causing disease, there is the dystrophin gene in Duchenne muscular dystrophy (DMD). Various different mutations leading to premature stop codons in DMD patients can be removed by exon skipping promoted by oligonucleotides. Small molecules that bind to RNA structures and affect splicing are predicted to have a similar effect. Thus, in one aspect, the present invention is a method for identifying the structure of a target-binding moiety that modulates the splicing pattern of a target pre-mRNA to treat a disease or disorder, comprising contacting the target pre-mRNA with a compound represented by one or more of formulas (I), (II), (III), or (IV) for binding to the target pre-mRNA, and selecting a compound that interacts with the target RNA by analyzing the results of the analysis of the higher-order structure of the RNA disclosed herein.
[0073] Next, examples are given to further illustrate the present invention, but the present invention is not limited to these examples in any way.
Examples
[0074] For G4, acridine-VQ(SPh) and berberine-VQ(SPh), which specifically bind to and alkylate G4, were used as modifying molecules (hereinafter, they may be collectively referred to as Sm-VQ), and mutation profiling (MaP) was performed on target RNA1 containing G4. Acridine-VQ(SPh) and berberine-VQ(SPh) are low molecular weight compounds prepared by covalently bonding acridine and berberine, which selectively bind to the G4 structure, to a VQ precursor having a thiophenyl (SPh) group, respectively (Figs. 3(a) to (d)). In addition, in order to confirm that the modification reaction occurs by the modification (alkylation) reaction between Sm-VQ(SPh) and the target base, a control experiment and analysis using Sm-VQ(SMe), which reduces only the modification activity while retaining the binding, as a modifying molecule were performed for each of the modifying molecules of acridine-VQ and berberine-VQ. In addition, in order to confirm that the mutations confirmed by MaP occur by a time-dependent chemical reaction, experiments with different reaction times were performed using acridine-VQ(SPh).
[0075] (Synthesis Example 1) Synthesis of Acridine-VQ [Chemical formula] To a solution of 2-aminobenzamide (301 mg, 2.21 mmol) in DMF (4.0 mL), K2CO3 (919 mg, 6.65 mmol) and tert-butyl bromoacetate (485 μL, 3.31 mmol) were added, and the mixture was stirred at 90 °C. After stirring for 40 hours, the mixture was cooled to room temperature and diluted with CH2Cl2 (30 mL) and water (10 mL). The organic layer was separated, dried over anhydrous Na2SO4, filtered, and evaporated under reduced pressure. The residue was purified by column chromatography (CHCl3 / MeOH = 99 / 1) to obtain Compound 5 (265.7 mg, 48%) as a pale yellow solid.
[0076] To a solution of compound 5 (100.3 mg, 0.40 mmol) in CH2Cl2 (3.5 mL) was added 3-(methylthio)propionyl chloride (140 μL, 1.21 mmol), and the mixture was stirred at room temperature. After stirring for 3 hours, the reaction mixture was diluted with CH2Cl2 (10 mL) and washed with saturated aqueous NaHCO3 (15 mL × 4), water (15 mL), and brine (15 mL). The organic layer was dried over anhydrous Na2SO4, filtered, and concentrated under reduced pressure. The crude product was suspended in Et2O / hexane = 1 / 2 (10 mL). The solid was collected by filtration and then washed with Et2O / hexane = 1 / 2 (20 mL) to obtain the desired compound 6 (97.1 mg, 73%) as a pale yellow solid.
[0077] To a solution of compound 6 (41 mg, 0.13 mmol) in DCM (0.2 mL) were added triisopropylsilane (40 μL, 0.19 mmol) and TFA (0.82 mL), and the reaction mixture was stirred at room temperature. After stirring for 4 hours, the reaction mixture was concentrated under reduced pressure and co-evaporated three times with acetonitrile. The residue was purified by column chromatography (EtOAc only → EtOAc:MeOH = 4:1) to obtain compound 7 as a white solid (25 mg, 72%).
[0078] For compound 7 1 1H NMR (DMSO-d6, 400 MHz) δ (ppm) 8.14 (1H, d, J = 7.6 Hz), 7.87 (1H, dd, J = 7.2, 8.0 Hz), 7.64 (1H, d, J = 8.4 Hz), 7.57 (1H, dd, J = 7.2, 7.6 Hz), 5.24 (2H, s), 3.16 (2H, brs), 2.87 (2H, t, J = 7.2 Hz), 2.49 (2H, br), 2.12 (3H, s). 13 13C NMR (DMSO-d6, 125 MHz) δ (ppm) 169.0, 164.4, 163.4, 140.7, 135.3, 127.7, 127.2, 119.6, 116.8, 49.0, 34.0, 30.2, 15.1; ESI-HRMS (m / z): C 13 H 15 N2O3S + as [M + H] + Calculated 279.0798, found 279.0795.
[0079]
Chem.
[0080] 9-Chloroacridine (Compound 8) (230 mg, 1.08 mmol) and amine linker (Compound 9) (321 mg, 1.29 mmol) were dissolved in phenol (1.1 g), and the reaction mixture was stirred at 100 °C for 3 hours. The reaction mixture was cooled to room temperature, and 1N aqueous NaOH solution (10 mL) was poured in. This solution was extracted with CH2Cl2 (30 mL × 2), washed with brine (20 mL), dried over anhydrous Na2SO4, filtered, and evaporated. The residue was purified by column chromatography (CHCl3:MeOH = 9:1 → 7:1 → 5:1 → 3:1) to obtain Compound 10 as a yellow oil (442 mg, 96%).
[0081] TFA (0.95 mL) was added to a solution of Compound 10 (14 mg, 0.03 mmol) in DCM (0.2 mL), and the reaction mixture was stirred at room temperature for 2 hours. The reaction mixture was concentrated and co-evaporated three times with acetonitrile. The residue was passed through amino silica, concentrated, and dissolved in DMF (0.5 mL). The reaction solution was added to a new flask containing Compound 7 (11 mg, 0.04 mmol) in DMF (0.1 mL). To the reaction mixture, HBTU (15 mg, 0.04 mmol), HOBt (5.3 mg, 0.04 mmol), and DIPEA (58 μL, 0.33 mmol) were added, and the mixture was stirred at room temperature. After stirring for 2 hours, the reaction mixture was diluted with DCM and washed with saturated aqueous NaHCO3 and brine. The organic layer was separated, dried over Na2SO4, filtered, and evaporated. The residue was purified by column chromatography (EtOAc:MeOH = 49:1 → 29:1 → 19:1 → 9:1) to obtain Compound 3-SMe as a yellow solid (10 mg, 52%). A portion of this solid was further purified by reverse-phase HPLC using a C-18 column (Nacalai tesque: COSMOSIL 5C18-AR-II, 10 × 250 mm) with a linear gradient of 0 - 45% / 30 min acetonitrile in 0.1% TFA buffer at a flow rate of 4 mL / min at 40 °C, with UV detection at λ = 254 nm and fluorescence detection (λex = 266 nm, λ em = 450 nm) was monitored, and the desired product was obtained as a pale yellow solid. The concentration of compound 3-SMe was determined by quantitative 1 1H NMR using maleic acid as an internal standard (ε 260 = 48,750 M -1 ·cm -1 ).
[0082] For compound 3-SMe, 1 1H NMR ((DMSO-d6, 600 MHz) δ (ppm) 13.48 (1H, s), 9.64 (1H, dd, J = 5.4, 6.0 Hz), 8.59 (2H, d, J = 9.0 Hz), 8.56 (1H, dd, J = 5.4, 6.0 Hz), 8.04 (1H, dd, J = 5.4, 6.0 Hz), 8.04 (1H, dd, J = 1.2, 7.8 Hz), 7.98 (2H, dd, J = 1.2, 8.4 Hz), 7.83 (2H, dd, J1.2, 8.4 Hz), 7.72 (1H, dd, J = 1.2, 8.4 Hz), 7.55 (2H, dd, J = 7.2, 7.8 Hz), 7.41 (2H, dd, J = 7.2, 8.4 Hz), 4.92 (2H, s), 4.27 (2H, q, J = 5.4 Hz), 3.92 (2H, t, J = 5.4 Hz), 3.57 - 3.58 (2H, m), 3.47 - 3.50 (2H, m), 3.36 (2H, t, J = 5.4 Hz), 3.19 (2H, dd, J = 5.4, 11.4 Hz), 3.04 (2H, br-s), 2.85 (2H, t, J = 7.8 Hz), 2.09 (3H, s). 13 13C NMR ((DMSO-d6, 150 MHz) δ (ppm) 167.2, 166.2, 163.1, 158.3, 158.1, 157.8, 141.2, 135.3, 133.9, 127.3, 125.6, 123.4, 119.3, 118.6, 115.5, 69.9, 69.4, 68.8, 68.2, 49.0, 48.7, 40.1, 38.8, 34.3, 29.9, 14.9. ESI-HRMS (m / z). C 32 1H 36 N5O4S + as [M + H] + Calculated value 586.2483; Measured value 586.2484.
[0083] Synthesis of Aminoacridine-VQ Conjugated Thiophenol (3-SPh) [Chemical formula]
[0084] A solution of MMPP (1.2 nmol) in water (1.2 μL) was added to a solution of compound 3-SMe (2 nmol) in DMSO (2 μL), and the mixture was allowed to stand at room temperature for 1 minute to obtain compound 3-S(O)Me. A carbonate buffer solution at pH 10 (50 mM, 0.4 μL), thiophenol (100 nmol) in DMSO (0.2 μL), and DMSO (1.2 μL) were added, and these mixtures were incubated at 37 °C for 3 hours. The mixed solution was purified by HPLC to obtain compound 3-SPh.
[0085] Large-scale synthesis: A solution of MMPP (10.8 μmol) in water (708 μL) was added to a solution of compound 3-SMe (11.8 μmol) in DMSO (250 μL) and water (930 μL), and the mixture was allowed to rest at room temperature for 1 minute to obtain compound 3-S(O)Me. A carbonate buffer at pH 10 (50 mM, 232 μL), thiophenol (5.9 mmol) in DMSO (116 μL), and DMSO (690 μL) were added, and the mixture was incubated at 37 °C for 3 hours. 2,2'-Dipyridyl disulfide (2.9 mmol) in DMSO (58 μL) was added to this solution, and the solution was purified by HPLC to obtain compound 3-SPh.
[0086] of 3-SPh 11H NMR (600 MHz, DMSO-d6): δ (ppm) = 13.42 (1H, s), 9.60 (1H, t, J = 5.4 Hz), 8.59 (2H, d, J = 8.4 Hz), 8.48 (1H, t, J = 5.4 Hz), 8.05 (1H, dd, J = 7.8, 1.8 Hz), 7.97 (2H, dd, J = 8.4, 7.2 Hz), 7.82 (2H, d, J = 8.4 Hz), 7.71 (1H, dd, J = 7.8, 7.2, 1.8 Hz), 7.54 (2H, t, J = 8.4 Hz), 7.43 - 7.39 (2H, m), 7.33 (2H, d, J = 7.2 Hz), 7.29 (2H, t, J = 7.2 Hz), 7.16 (1H, t, J = 7.2 Hz), 4.9 (2H, s), 4.27 (2H, q, J = 5.4 Hz), 3.91 (2H, t, J = 5.4 Hz), 3.57 (2H, t, J = 5.4 Hz), 3.47 - 3.45 (2H, m), 3.36 - 3.31 (4H, m), 3.15 (2H, t, J = 5. Hz), 3.07 (2H, br).
[0087] of 3-SPh 13 13C NMR (150 MHz, DMSO-d6): δ (ppm) = 167.26, 166.07, 162.60, 158.23, 157.70, 141.18, 135.97, 135.21, 133.78, 129.09, 128.11, 127.26, 125.77, 125.50, 119.28, 118.52, 115.23, 69.84, 69.40, 68.73, 68.16, 48.91, 48.64, 38.71, 34.18, 28.97. ESI-HRMS (m / z): C 32 H 36 N5O4S + as [M + H] + Calculated value 586.2483; Measured value 586.2484.
[0088] (Synthesis Example 2) Synthesis of Berberine-VQ
Chemical Structure
[0089] To a solution of Compound 1 (5 mg, 35.93 μmol) in DMF (0.4 mL), DIPEA (9.5 μL), HBTU (12.8 mg, 33.75 μmol), and HOBt (3.4 mg, 25.16 μmol) were added. After stirring at room temperature for 30 minutes, N-(tert-butoxycarbonyl)-2-(2-aminoethoxy)ethylamine (4.5 μL, 22.54 μmol) was added and the reaction was carried out for 24 hours. The reaction solution was evaporated with an oil pump to remove DMF, then extracted with CHCl3 (15 mL) and washed with NaHCO3 (10 mL × 2) and brine (10 mL). The organic solution was dried over Na2SO4 and concentrated. The crude compound was purified by the following method. Silica gel column chromatography (Pasteur pipette, CHCl3:MeOH = 50:1 → 30:1 → 20:1 → 10:1) was performed to obtain a white solid of Compound 2 (1.4 mg, 3.01 μmol, 16.8%).
[0090]
Chem.
[0091] To a solution of Compound 4 (15 mg, 41.92 μmol) in DMF (1.5 mL), K2CO3 (11.3 mg, 81.75 μmol) and t-butyl-2-bromoacetate (12.5 μL, 85.21 μmol) were added, and the reaction mixture changed from yellow to brown. After stirring at room temperature for 21 hours, the reaction mixture was filtered and turned yellow. A solid precipitated on the cotton. The precipitate was dissolved in MeOH and evaporated to obtain a yellow solid (5.5 mg, 12.60 μmol). The residue filtrate was recrystallized with EA:MeOH:hexane = 1.7 mL:1 mL:6 mL to obtain a yellow fine powder (7 mg, 16.04 μmol, the total yield of the obtained Compound 5 was 68.3%).
[0092]
Chem.
[0093] To a solution of compound 5 (7 mg, 16.04 μmol) in DCM (105 μL), triethylsilane (3.85 μL, 24.06 μmol) and TFA (420 μL) were added. The reaction mixture was stirred at room temperature for 1 h, evaporated and co-evaporated three times with MeCN, and the crude compound was purified by silica gel column chromatography (EA:MeOH = 0:1 → 8:1 → 5:1 → 1:1 → 1:1:10) to give a yellow solid of compound 6. (3.5 mg, 9.20 μmol, 57.4%)
[0094]
Chem.
[0095] To a solution of compound 2 (2.8 mg, 6.03 μmol) in DCM (40 μL), triethylsilane (1.45 μL) and TFA (150 μL) were added. After stirring at room temperature for 30 min, the reaction mixture was evaporated and co-evaporated three times with MeCN. The crude compound was quickly passed through a silica gel column to remove TFA (CHCl3:MeOH = 10:1 → 1:1), and the resulting solution was concentrated (washed with DMF 100 μL × 2) and added to a DMF solution (150 μL) of compound 6 (2.3 mg, 6.05 μmol), DIPEA (3.15 μL, 18.14 μmol), HOBt (1.9 mg, 14.01 μmol), and HBTU (5.6 mg, 14.76 μmol). After stirring at room temperature for 1 h, HBTU (2.6 mg, 6.86 μmol) was replenished. After 30 min, the reaction mixture was evaporated, dissolved in DMSO, and filtered through a membrane (Advantec 13 HPO45AN 0.45 μm). The filtrate was purified by HPLC to give a yellow solution of compound 7. (3.26 μmol, 53.9%)
[0096] of compound 7 11H NMR (600 MHz, DMSO) δ (ppm) 9.96 (1H, s), 8.91 (1H, s), 8.59 (1H, d, J = 5.4 Hz), 8.22 (1H, t, J = 5.4 Hz), 8.18 (1H, d, J = 9.6 Hz) was used to measure 1H-NMR (600 MHz, DMSO) δ (ppm) 9.96 (1H, s). 8.06 (1H, d, J = 7.8 Hz), 7.99 (1H, d, J = 9 Hz), 7.78 (1H, s), 7.76 (1H, t, J = 7.2, 8.4 Hz), 7.44 (2H, m), 7.09 (1H, s), 6.18 (2H, s), 4.97 (2H.s), 4.90 (2H, d, J = 6 Hz), 4.79 (2H, s), 4.04 (3H, s), 3.48 (4H, m), 3.35 (2H, t, J = 6 Hz), 3.19 (2H, t, J = 6 Hz), 3.06 (2H, s), 2.86 (2H, t, J =.7.2 Hz), 2.10 (3H, s).
[0097] Of compound 7 13 13C NMR (600 MHz, DMSO) δ (ppm) 167.90, 167.35, 166.32, 163.00, 149.89, 149.83, 147.72, 145.82, 141.92, 141.28, 137.53, 133.73.132.89, 130.63, 127.26, 126.62, 125.44, 123.76, 121.27, 120.42, 120.14, 119.23, 115.32, 108.44, 105.45, 102.11, 71.62, 68.70, 57.12, 55.44, 48.64, 38.76, 38.28, 34.37, 29.81, 26.37, 14.84. HRMS (ESI-TOF): C 38 H 40 N5O8S + As [M] + Calculated value 726.2592, measured value 726.2567; C 38 H 41 N5O8S + As [M + H] 2+ Calculated value 363.6333, measured value 363.6345.
[0098]
Chemical formula
[0099] To a solution of compound 7 (5 μmol) in DMSO (269 μL) was added a solution of MMPP (25 μmol) in water (1.25 mL), and the mixture was stirred at room temperature for 1 minute to obtain compound 8. Carbonate buffer pH 10 (50 mM, 1 mL), thiophenol (400 μmol) in DMSO (800 μL), and DMSO (2.7 mL) were added, and the mixture was incubated at 37 °C for 3 hours. This solution was purified by HPLC to obtain compound 9 (3 μmol, 60%).
[0100] For compound 9 1 H NMR (600 MHz, DMSO) δ (ppm) 9.96 (1H, s), 8.90 (1H, s), 8.55 (1H, s), 8.23 (1H, s), 8.17 (1H, d, J = 9 Hz), 8.06 (1H, d, J = 7.8 Hz), 7.98 (1H, d, J = 9 Hz), 7.78 (1H, s), 7.74 (1H, t, J = 8.4 Hz), 7.44 (2H, t, J = 7.8 Hz), 7.35 (2H, d, J = 7.8 Hz), 7.31 (2H, t, J = 7.8 Hz), 7.18 (1H, t, J = 7.2 Hz), 7.08 (1H, s), 6.18 (2H, s), 4.90 (4H, m), 4.79 (2H, s), 4.03 (3H, s), 3.48 (2H, t, J = 6.0 Hz), 3.43 (2H, t, J = 6.0 Hz), 3.36 (2H, m), 3.35 (2H, m), 3.27 (2H, t, J = 5.4 Hz), 3.19 (2H, t, J = 6.0 Hz), 3.08 (2H, s).
[0101] For compound 9 1313C NMR (150 MHz, DMSO) δ (ppm) 167.92, 167.33, 166.20, 162.61, 149.89, 149.83, 147.72, 145.78, 141.93, 141.23, 137.49, 136.00, 133.78, 132.90, 130.61, 129.11, 1218.14, 127.29, 126.60, 125.80, 125.51, 123.76, 121.26, 120.41, 120.12, 119.26, 118.41, 116.42, 115.31, 108.44, 105.45, 102.11, 71.63, 68.71, 68.57, 57.12, 55.45, 48.63, 38.74, 38.29, 34.20, 28.98, 26.38.
[0102] HRMS (ESI-TOF) of Compound 9 43 H 42 N5O8S + as [M] + : Calcd 788.2749, Found 788.2708, 43 H 43 N5O8S + as [M+H] 2+ : Calcd 394.6411, Found 394.6415.
[0103] (Example 1) Mutation Profiling (MaP) Using Sm-VQ as a Modifying Molecule <Sequence of Target RNA1> To demonstrate the utility of the acridine-VQ synthesized in Synthesis Example 1 and berberine-VQ synthesized in Synthesis Example 2, the following sequence was used as the RNA to be analyzed: 5’-[Cassette sequence]-GUCUCGCGAGAGUGAGGCAAGCAUACCGGGGCGGGCCUUGGGCGGGGUGUAUGCAAUGGUGCUGAGAGGCACCACAAAU-[Cassette sequence]-3’ (SEQ ID NO: 1) was used. This sequence contains a partially artificially modified sequence: 5’-AGCAUACCGGGGCGGGCCUUGGGCGGGG-3’ (SEQ ID NO: 2) of the G4 sequence present within the promoter sequence of human vascular endothelial growth factor, and forms a stable G4 structure. The 5’ end of RNA1 contains an arbitrary sequence (5’ cassette sequence) necessary for the DNA amplification reaction, and the 3’ end contains an arbitrary sequence (3’ cassette sequence) necessary for the reverse transcription reaction and the DNA amplification reaction.
[0104] <Alkylation reaction against target RNA1> First, target RNA1 was incubated at 95°C for 5 minutes and then cooled to 4°C in a 20 mM phosphate buffer (pH 7.0), 80 mM KCl, 20 mM NaCl solution (PKN Buffer) to perform RNA folding. Next, each Sm-VQ was reacted with target RNA1. The scale of the reaction solution was 20 μL, and the composition was 4 μM target RNA1, 1×PKN Buffer, 20 μM of each Sm-VQ precursor. For the negative control sample, dimethyl sulfoxide (DMSO) and 20 mM EDTA (diluted with 1×PKN Buffer) were added instead of 20 μM Sm-VQ precursor. After the reaction, target RNA1 was purified. For purification, RNA Clean & Concentrator-5 from Zymo Research or AMPure XP (manufactured by Beckman Coulter) was used.
[0105] <Reverse transcription reaction for mutation profiling> For the RNA sample after the alkylation reaction, a reverse transcription reaction was performed using a reverse primer having a sequence complementary to the 3' cassette sequence. First, annealing of the reverse transcription primer was performed on the RNA after the alkylation reaction. The scale of the reaction solution was 10 μL, and the composition was 7 μL of the RNA solution after the alkylation reaction, 1 μL of 2 μM reverse primer, and 2 μL of 10 mM dNTP. Here, 2.22×RT Buffer required for the reverse transcription reaction was prepared. The composition was 2.22×MaP pre-buffer, 2.22 M betaine, and 11.1 mM MgCl2. 2.22×MaP pre-Buffer was prepared in advance. The composition of 5×MaP pre-buffer was 250 mM Tris (pH 8.0), 375 mM KCl, and 50 mM DTT. Next, the reverse transcription reaction was performed according to the protocol of holding at 25 °C for 10 minutes → 60 °C for 90 minutes → 90 °C for 10 minutes → 4 °C. The scale of the reaction solution was 20 μL, and the composition was 1 μL of TGIRT-III, 9 μL of 2.22×RT Buffer, and 10 μL of the reaction solution after annealing. Next, 1 μL of RNaseH was added to the solution after the reverse transcription reaction and reacted at 37 °C for 20 minutes to decompose the remaining RNA. Finally, the cDNA was purified. For purification, RNA Clean & Concentrator-5 from Zymo Research or AMPure XP (manufactured by Beckman Coulter) was used.
[0106] <Preparation of Illumina sequencing library> Amplicon PCR and index PCR were performed as DNA amplification reactions for library preparation. 0.5 ng of the reverse transcription product, 1× Platinum TM SuperFi TM Amplicon PCR was performed in a 25 μL reaction volume using a PCR Master Mix and 1× SuperFi GC Enhancer (both from Thermo Fisher Scientific), 500 nM forward and reverse primers. First, it was heated at 98 °C for 30 seconds, and then three-step PCR was performed at 98 °C for 10 seconds, 64 °C for 10 seconds, and 72 °C for 20 seconds. After the last cycle, the temperature was held at 72 °C for 5 minutes and then cooled to 4 °C. After PCR, 2.5 μL of Exonuclease I (from NEW ENGLAND Biolabs) was added to degrade the remaining primers and reacted at 37 °C for 15 minutes. For purification, the DNA clean-up and concentration protocol of the Monarch PCR & DNA Cleanup Kit (5 μg) (from New England Biolabs) was used. 8 μL of DNA elution buffer was used for the final elution. This prepared the amplicon for indexing for Illumina sequencing. Next, index PCR was performed in a 25 μL reaction volume using 1 ng of the amplicon PCR product. The other reaction components were 1× Platinum TM SuperFi TM PCR Master Mix and 1 μM index primers of the Nextera XT Index Kit v2 (Illumina). First, it was heated at 98 °C for 30 seconds, and then seven 3-cycle PCRs were performed at 98 °C for 10 seconds, 55 °C for 10 seconds, and 72 °C for 20 seconds. After the last cycle, the temperature was held at 72 °C for 5 minutes and then cooled to 4 °C. For purification, it was cleaned up using AMPure XP (from Beckman Coulter). For elution, 14 μL of water was added to the dried beads, mixed well, incubated at room temperature for 10 minutes, and the supernatant was recovered. Then, samples with different indexes were mixed into the same solution for sequencing.
[0107] <Illumina sequencing> For sequencing, NextSeq500 / 550 Mid Output Kit v2.5 (150 cycles) using paired-end reads and standard read primers or Miseq Micro Kit v2, Miseq Nano Kit v3 were used.
[0108] <Alignment and Data Analysis> After removing the adapter region from the FASTQ files, alignment was performed against the reference using BWA. The deletion rate was calculated by summing the number of deletions for each nucleotide and dividing by the total number of reads at a certain base position. To reduce noise due to sequence-specific mutations, the deletion rate of the unmodified sample was subtracted from the deletion rate of the Sm-VQ modified sample to obtain the delta deletion rate (ΔDeletion rate) of the following formula (1).
[0109] Delta deletion rate (ΔDeletion rate) = Deletion rate 被修飾 - Deletion rate 未修飾
[0110] Results and Discussion <Detection of G4 Structure of Target RNA1> For the target RNA1 containing the G4 structure, the above-described experiments and analyses were performed, and the G4 structure was detected by identifying the binding sites of small molecule compounds that bind to G4. As modification molecules, acridine-VQ(SPh) and berberine-VQ(SPh) were used. From the sequence data, the deletion probabilities at each nucleotide position were calculated for the sample containing Sm-VQ (Sm-VQ) and the control sample without Sm-VQ (DMSO) (Figure 4(a)). Furthermore, the difference obtained by subtracting the deletion probability of the control sample without Sm-VQ from the deletion probability of the sample containing Sm-VQ was taken, and the deletion probability that occurred only in the sample containing Sm-VQ (ΔDeletion rate = deletion probability in Sm-VQ - deletion probability in DMSO) was evaluated (Figure 4(b)). In Figure 4(b), for both modification molecules used, in target RNA1, high deletion probability peaks were observed at cytosine and uracil, which are the bases modified by Sm-VQ in the G4 region. To show the statistical significance of these deletion probability peaks, the ΔSHAPE framework cited from a previous study on SHAPE-MaP (Matthew J Smola & Kevin M Weeks, In-cell RNA structure probing with SHAPE-MaP. Nature Protocols 13, 1181-1195 (2018)) was used as a statistical filter. In the ΔSHAPE framework, Z-factor and Standard Score are used to test the statistical significant difference in the mutation probability at each base of the sequence. In Figure 4b, the bases determined to be statistically significant by the ΔSHAPE framework (Z-factor > 0, Standard Score > 1) are shown with a gray background color. From Figure 4b, when acridine-VQ was used, a statistically significant deletion probability peak was observed at uracil in the G4 region, and when berberine-VQ was used, statistically significant deletion probability peaks were observed at uracil and cytosine in the G4 region. From this, it is considered that Motif-MaP using Sm-VQ as a modification molecule can quantitatively detect the target structure by the statistical significant difference test of the deletion probability calculated from the sequence data.
[0111] <Evaluation of the length of deletions in target RNA1> To evaluate how much sequence information is lost due to deletions in sequencing, the length of deletions for each nucleotide of target RNA1 was calculated using the same sequence data as in Example 1 above. The length of deletions for each nucleotide was calculated from the sequencing data of the sample containing Sm-VQ and the control sample without Sm-VQ, the difference was taken, and the number of deletions that occurred only in the sample containing Sm-VQ was calculated for each length of deletion, and the ratio to the total number of deletions of any base was evaluated (Figure 5). For any base, most deletions are of length 1, and it is considered that only the bases where the modification reaction occurred are deleted. From this result, it was found that the mutation profiling using the deletion probability used in this technology does not lose the sequence information and thus the structural information of the modified RNA1 molecule due to deletions. This feature has more binding sites and thus higher-order structure sites obtained from a single molecule compared to conventional structure detection techniques using RT-Stop where sequence information after transcription termination is largely lost. Therefore, it is useful for the discovery of multiple binding sites, co-occurring binding patterns, and the detection of fluctuations in RNA higher-order structures.
[0112] <Verification of the alkylation reaction time dependence of the deletion probability> To verify whether the deletions observed in MaP using target RNA1 are due to chemical reactions by modifying molecules, the time-dependent change in the deletion probability was confirmed. Specifically, in Figure 4(a), the reaction time with 18 hours as the standard was replaced with eight different conditions of 0, 1, 2, 4, 8, 16, 18, 24, and 32 hours. Actually, this experiment and analysis were carried out using acridine-VQ as the modifying molecule, and it was verified how the deletion probability changes over time (Figure 6). From Figure 6, the number of deletions at the positions of uracil and cytosine in the G4 region in target RNA1 increased over time. This is considered to mean that the number of uracil and cytosine in the G4 region modified by acridine-VQ increased with each passing reaction time. That is, the deletions in MaP using acridine-VQ as the modifying molecule are considered to be derived from the chemical reaction of the modifying molecule with target RNA1.
[0113] <Control experiment and analysis using the control modifier molecule Sm-VQ-SMe> In the mutation profiling using target RNA1, the deletions confirmed were derived from the modification reaction of VQ and not caused by the specific binding of small molecules (acridine and berberine) to G4. To show this, a control experiment was carried out using the negative control molecule Sm-VQ(SMe) of Sm-VQ as the modifier molecule. In Sm-VQ(SMe), the SPh group of the VQ precursor is replaced by an SMe group. The SMe group is less likely to undergo an elimination reaction compared to the SPh group, and the conversion efficiency of the VQ precursor to VQ is low. That is, similar to Sm-VQ(SPh), Sm-VQ(SMe) binds to the target higher-order structure, but the modification efficiency is lower than that of Sm-VQ(SPh) (Non-Patent Document 5). For each of acridine and berberine, the ΔDeletion rate when using Sm-VQ(SPh) and Sm-VQ(SMe) as modifier molecules was compared (Figure 7). In both acridine and berberine, the significantly high deletion probability peak observed with SPh was not detected at any base of target RNA1, including the G4 region, with SMe. From this, it was confirmed that the deletions in MaP using Sm-VQ as the modifier molecule are derived from the modification reaction by VQ and do not depend on the binding of small molecules to RNA.
[0114] (Example 2) Cluster analysis of target RNA having single nucleotide polymorphisms <Preparation of target RNA sequences> As the sequences to be analyzed, two sequences, wild-type or SNP-type, derived from the microRNA precursor (pre-miRNA-1229) were used. The wild-type pre-miRNA-1229 sequence is It contains 5’-GGGUAGGGUUUGGGGGAGAGCGUGGGCUGGGGUUCAGGGACA-3’ (SEQ ID NO: 3). The SNP-type pre-miRNA-1229 sequence contains the sequence where the 21st cytosine of pre-miRNA-1229 is substituted with uracil: 5’-GGGUAGGGUUUGGGGGAGAGUGUGGGCUGGGGUUCAGGGACA-3’ (SEQ ID NO: 4). This single nucleotide substitution is known as rs2291418. At the 5’ end of each RNA sequence, an arbitrary sequence necessary for DNA amplification reaction and mapping (5’ cassette sequence) and an arbitrary sequence necessary for sequence discrimination (5’ barcode sequence) were added, and at the 3’ end, an arbitrary sequence necessary for reverse transcription reaction and DNA amplification reaction (3’ cassette sequence) and an arbitrary sequence necessary for sequence discrimination (3’ barcode sequence) were added. Analyzed RNAs as follows containing different barcode sequences for each target RNA sequence were constructed.
[0115] 5’-[Cassette sequence]-[Barcode sequence]-[SEQ ID NO: 3 or SEQ ID NO: 4]-[Barcode sequence]-[Cassette sequence]-3’
[0116] Hereinafter, the analyzed RNA containing wild-type pre-miRNA-1229 is denoted as WT, and the analyzed RNA containing SNP-type pre-miRNA-1229 is denoted as SNP.
[0117] <Association between rs2291418 and Alzheimer's disease> rs2291418 is an SNP within pre-miRNA-1229 for which an association with Alzheimer's disease (AD) has been reported. AD is known as a disease caused by protein misfolding, and the accumulation of tau protein and β-amyloid (Aβ) protein triggers the symptoms. The processing and trafficking of Aβ involve various proteins including sortilin-related receptor 1 (SORL1). miRNA-1229-3p is known to control the translation of SORL1, and the expression level of miRNA-1229-3p is known to increase in the rs2291418 pre-miRNA-1229 variant.
[0118] Pre-miRNA-1229 has been reported to be in an equilibrium state between G4 and hairpin structures. Also, rs2291418 has been reported to change the balance between these structures. (Joshua A. Imperatore., et al. (2020) Characterization of a G-Quadruplex Structure in Pre-miRNA-1229 and in Its Alzheimer’s Disease-Associated Variant rs2291418: Implications for miRNA-1229 Maturation. Int. J. Mol. Sci, see reference)
[0119] <Mutation Profiling by Berberine-VQ> Using the two types of RNA to be analyzed, WT and SNP, prepared above, an alkylation reaction was performed with berberine-VQ. The conditions of the alkylation reaction were basically the same as those in Example 1, but the concentration of the target RNA was different. In Example 1, 4 μM of target RNA1 was used, whereas in this example, an alkylation reaction was performed on a library containing 22 types of RNA sequences including 4 μM of the two types of RNA to be analyzed, WT and SNP. Subsequently, under the same conditions as in Example 1, reverse transcription reaction, preparation of cDNA library, and mutation profiling by sequencing were performed.
[0120] <Implementation of Clustering of Defect Patterns> (1) First, the missing information of the sequence to be analyzed was extracted from the SAM format file obtained by mapping the reads in the berberine-VQ modification group sample to the reference sequence. Specifically, 2000 reads with at least one length-1 deletion in the sequence to be analyzed were randomly selected from the SAM format file of the berberine-VQ modification group sample, and for each read, an array containing the information on the length of the deletion of each base in the sequence to be analyzed was generated. The length of each array was equal to the length of the sequence to be analyzed, and as components of the array, numbers 0 or 1 were included based on the presence or absence of a deletion. 1 corresponded to the base where a deletion occurred, and 0 corresponded to the base where no deletion occurred. This process was carried out for each of WT and SNP.
[0121] (2) Next, using UMAP, the arrays containing the deletion information of each read extracted in (1) were compressed two-dimensionally. This compression was performed collectively on the 4000 arrays extracted from the WT and SNP data.
[0122] (3) Next, using k-means, the deletion information compressed two-dimensionally in (2) was clustered. First, the elbow method was used to estimate the appropriate number of clusters. The elbow method is a technique for estimating the optimal number of clusters by calculating the sum of squared residuals at each number of clusters while changing the number of clusters and presenting it graphically. Here, the number of clusters used for this clustering was set to 4. Next, in (2), the generated two-dimensional list was clustered. From the cluster information obtained here, the two-dimensional list in (2) was color-coded for each cluster and plotted (Figure 9).
[0123] (4) A graph of the ΔDeletion rate corresponding to the four clusters obtained in (3) was generated. First, information on the array of the length of the analysis target array before dimensional compression was extracted from the two-dimensional array of each cluster. Next, the total number of deletions for each base in each cluster was calculated. Then, the ratio of the deletions for each base occupied by each cluster in the whole was calculated. Here, an array including information on the ratio of the deletions occupied by each base in each cluster was generated. The ΔDeletion rate for each cluster was calculated by multiplying each component of this array by the ΔDeletion rate of the corresponding base, and is shown (Figure 10).
[0124] <Results and Discussion> The deletion information of each sequence of WT and SNP was two-dimensionally compressed and classified into four clusters as shown in Figure 9. In WT and SNP, the ratio occupied by each cluster in the whole was different. Specifically, the ratios of Cluster 1 and Cluster 2 were higher in WT, and the ratio of Cluster 3 was higher in SNP (Figure 10(a), Figure 10(b)). This difference is considered to have occurred because the modification pattern of berberine-VQ was different between WT and SNP. Also, the difference in the modification pattern of berberine-VQ between clusters is considered to be the difference in the higher-order structure formed by the target RNA sequence. Specifically, it is considered that multiple RNA structures of pre-miRNA-1229 were in an equilibrium state, and due to the SNP, the balance between the structures changed, resulting in the difference in the modification pattern of berberine-VQ and thus the difference in the deletion pattern.
[0125] RNA can form multiple structures from one sequence, and multiple bases corresponding to each structure react with small molecules. Thus, Motif-MaP has been shown to be able to not only detect the target RNA higher-order structure, but also distinguish and detect the binding pattern of co-occurring small molecules and the fluctuations (structural equilibrium state) between multiple RNA higher-order structures. These results indicate that by combining mutation profiling (MaP) and cluster analysis, the higher-order structure of the target RNA can be analyzed more precisely and in detail.
[0126] (Example 3) Concentration of Modified RNA Using RNA Pull-Down 1. First In Example 1, mutation profiling was performed using the molecule Sm-VQ that modifies the binding site of a low-molecular-weight compound. That is, the deletion probability at each base of the RNA was determined from the sequence data, and bases with a significantly high deletion probability according to the binding-modification reaction were regarded as low-molecular-weight binding positions, and the target RNA higher-order structure was detected. Therefore, in order to efficiently detect the target RNA higher-order structure from limited sequence data, it was necessary to extract more deletion information or information on modified RNA.
[0127] When the RNA to be analyzed contains unmodified RNA, there is a problem that the sequencing cost increases when reverse transcription and amplification are uniformly performed on the modified RNA and the unmodified RNA in the same solution. Therefore, steps for selectively concentrating the modified RNA from the mixed solution of the modified RNA and the unmodified RNA were added, and the Motif-MaP method was carried out.
[0128] The concentration of the modified RNA mainly consists of three types of steps. First, a specific modification reaction induced by RNA-low molecular weight interaction is performed using a low molecular weight binding alkylating agent having an azide group. As a result, an azide group is added to the modified RNA. Next, the azide group added in the modification reaction is converted to biotin by a click reaction. Finally, an RNA pull-down assay using biotin-avidin interaction is carried out. In this pull-down assay, since the RNA to which biotin is added, and thus the modified RNA, preferentially binds to avidin beads, the modified RNA can be concentrated.
[0129] 2. Experimental Method <Sequences Used> A target RNA library consisting of 10 types of sequences shown in Table 1 below was used. As for the RNA to be analyzed included in the library, for SEQ ID NOs: 5 to 13, RNAs each consisting of 5'-[cassette sequence]-[target sequence]-[cassette sequence]-3' were used. Among the 10 types of sequences, the modification efficiency of Sm-VQ has been examined for 9 types.
[0130]
Table 1
[0131] <Modification reaction using a modified molecule having an azide group> First, a target RNA library 1 containing 10 sequences was incubated at 95 °C for 5 minutes and then cooled to 4 °C in a 20 mM phosphate buffer (pH 7.0), 80 mM KCl, 20 mM NaCl solution (PKN Buffer) to perform RNA folding. Next, acridine-VQ(NMe2) having a covalently bonded azide group (its structural formula is shown below) was reacted with the target RNA library 1.
[0132]
Chemical formula
[0133] The scale of the reaction solution was 20 μL, and the composition was 4 μM target RNA library 1, 1×PKN Buffer, and 20 μM of each Sm-VQ precursor. For the negative control sample, dimethyl sulfoxide (DMSO) and 20 mM EDTA (diluted with 1×PKN Buffer) were added instead of 20 μM of the acridine-VQ(NMe2) precursor. After the reaction, the target RNA library 1 was purified. For purification, RNA Clean & Concentrator-5 from Zymo Research or AMPure XP (manufactured by Beckman Coulter) was used.
[0134] <Click reaction> For 1500 ng of the RNA sample after the modification reaction, 2 μL of 2 mM Click-iT TMAfter adding Biotin sDIBO Alkyne (manufactured by Thermo Fisher Scientific) and 1 μL of RiboLock RNase Inhibitor (manufactured by Thermo Fisher Scientific), each sample was adjusted to a final volume of 30 μL using ultrapure water. Next, all reaction solutions were stirred at 37 °C and 1000 rpm for 2.5 hours using an Eppendorf Thermomixer. After the reaction, Target RNA Library 1 was purified using RNA Clean & Concentrator-5 from Zymo Research.
[0135] <RNA Pull-down> In a 1.5 mL tube, 20 μL of SpeedBeads TM Magnetic Neutravidin Coated particles (manufactured by Cytiva) were dispensed, and after placing the tube on a magnetic rack, the supernatant was removed. Next, 500 μL of 1x PKN Buffer was added and mixed by inversion. Then, the tube was placed on the magnetic rack and the supernatant was removed. Next, the RNA sample after the click reaction was added, and 1x PKN Buffer was added until the total volume reached 1000 μL. Then, it was stirred at 25 °C and 1200 rpm for 1 hour using an Eppendorf Thermomixer, after which the tube was placed on the magnetic rack and the supernatant was removed. As a washing operation, 1000 μL of 1x PKN Buffer was added and mixed by inversion. After spin-down, the tube was placed on the magnetic rack and the supernatant was removed. This series of washing operations was performed 3 times in total. After the washing operation, 50 μL of Elution Buffer (95% formamide, 10 mM EDTA, pH 8.2) was added and heat-treated at 80 °C for 5 minutes. Then, after standing at room temperature for 5 minutes, the tube was placed on the magnetic rack and the supernatant was transferred to a new DNA Lobind tube, and Target RNA Library 1 was purified using RNA Clean & Concentrator-5 from Zymo Research.
[0136] <Reverse Transcription Reaction of Target RNA Library 1 and Preparation of Illumina Sequencing Library> In the same manner as in Example 1, reverse transcription reaction and preparation of an Illumina sequencing library were performed.
[0137] <Illumina sequencing> For sequencing, iSeq 100 i1 Reagent v2 (300-cycle) using paired-end reads and standard read primers was used.
[0138] <Experimental results> Among the target RNA libraries 1, graphs of deletion profiling for four target sequences that have been found to have high modification efficiency in other assays and high binding affinity with small molecules are shown in FIGS. 11(A) to 11(D), and graphs of deletion profiling for five target sequences that have been found to have low modification efficiency and low binding affinity with small molecules are shown in FIGS. 12(A) to 12(E). In each graph, the sequence is shown on the horizontal axis and the Δ deletion probability is shown on the vertical axis. Also, the dark gray graph is a sample in which the modified RNA has been concentrated using RNA pull-down, and the light gray graph is the result of the control sample without this treatment.
[0139] From FIG. 11, it was found that for the four sequences with high binding affinity with small molecules, the deletion probability increased by concentration. Also, many bases with increased deletion probability were bases of U or bases around it, which are considered to be the most easily modified by Sm-VQ. On the other hand, from FIG. 12, for the five sequences with low binding affinity with small molecules, no significant increase in deletion probability as seen in the results of FIG. 11 was observed.
[0140] From these results, it was shown that the probability of deletion increases depending on the strength of the binding affinity with the small molecule due to the enrichment of the modified RNA. In Motif-MaP, the target RNA higher-order structure is identified using the information on the deletions induced and generated by the modification reaction at each base of each sequence. Therefore, the higher the probability of deletion of a base, the more likely it is to be recognized as the small molecule binding position with a limited number of sequencer reads. That is, it is considered that the selective enrichment of the modified RNA described in this paper enables the identification of the target RNA higher-order structure with higher detection efficiency than the existing Motif-MaP method.
Claims
1. The following formula (I), (II), (III) or (IV): 【Chemical 1】 (wherein Sm represents a target binding moiety, L represents a linker, X is -S-R 4 , -S(O)-R 4 , -O-R 5 or -N(R 6 )-R 7 wherein R 1 , R 2 and R 3 are each independently a hydrogen atom, a halogen, an alkyl which may have a substituent, an alkenyl which may have a substituent, an alkynyl which may have a substituent, an alkoxy which may have a substituent, an aryl which may have a substituent, an aralkyl which may have a substituent, a cycloalkyl which may have a substituent, or a heteroaryl which may have a substituent, or R 1 and R 2 or R 2 and R 3 together form a ring which may have a substituent, R 4 represents an alkyl group which may have a substituent, an aryl group which may have a substituent, or a heteroarylalkyl group which may have a substituent, R 5 represents a hydrogen atom or an alkyl group which may have a substituent, R 6 and R 7 each independently represents a hydrogen atom, an alkyl group which may have a substituent or an aryl group which may have a substituent, or R 6 and R 7 together form a ring which may have a substituent.) preparing a compound represented by, contacting the compound with one or more RNAs, determining the nucleotide sequence of complementary DNA synthesized by reverse transcriptase using the RNA contacted with the compound as a template, based on the nucleotide sequence of the complementary DNA in which the base modified by the compound on the RNA is skipped and one or several bases are deleted, determining the position and / or region on the RNA that interacts with the target binding moiety of the compound, A method for analyzing the higher-order structure of RNA, comprising:
2. The compound represented by formula I is the following formula (V): [Chemical Formula 2] (wherein Sm, L and X have the same meanings as in Claim 1, R 8 , R 9 , R 10 , and R 11 are each independently a hydrogen atom, a halogen, an alkyl which may have a substituent, an alkenyl which may have a substituent, an alkynyl which may have a substituent, an alkoxy which may have a substituent, an aryl which may have a substituent, an aralkyl which may have a substituent, a cycloalkyl which may have a substituent, or a heteroaryl which may have a substituent. ) The method according to Claim 1, which is a compound represented by.
3. In the compound represented by the formula V, R 8 , R 10 and R 11 are hydrogen atoms, and R 9 is the following formula (VI) or (VII): 【Chemical 3】 represents a substituent represented by, The method according to Claim 2, further comprising a step of concentrating the RNA modified by the compound after contacting the compound with a plurality of RNAs.
4. wherein X is -S-R 4 or -S(O)-R 4 and R 4 is methyl, hydroxyethyl, 2-pyridylmethyl or phenyl which may have a substituent, according to any one of claims 1 to 3.
5. wherein X is -N(R 6 )-R 7 , R 6 and R 7 are each independently a hydrogen atom, methyl or phenyl which may have a substituent, or R 6 and R 7 together form a cycloalkyl ring which may have a substituent, a morpholine ring which may have a substituent or a piperazine ring which may have a substituent, the method according to claim 1 or 2.
6. The method according to any one of Claims 1 to 5, wherein the RNA contains a structural motif that forms a higher-order structure.
7. The method according to Claim 6, wherein the structural motif is a stem-loop, multi-branched loop, junction, bulge, kink turn, pseudoknot, triple-stranded structure or quadruple-stranded structure, or a combination thereof.
8. The linker is a polyethylene glycol (PEG) group having 1 to 20 ethylene glycol subunits, an alkyl group having 1 to 12 carbon atoms which may have a substituent, an alkenyl group which may have a substituent, an alkynyl group which may have a substituent, and a cycloalkyl group which may have a substituent, and a divalent group selected from the group consisting of peptides containing 1 to 8 amino acids. The method according to any one of Claims 1 to 7.
9. The method according to any one of Claims 1 to 8, wherein the target binding moiety is an arbitrary compound or a part thereof, and the method includes a step of identifying the RNA that interacts with the arbitrary compound from the one or more RNAs.
10. The method according to any one of Claims 1 to 9, further comprising a step of estimating the higher-order structure of the RNA region that interacts with the target binding moiety.
11. The method according to any one of Claims 1 to 10, wherein the target binding moiety is a compound that binds to a guanine quadruplex or a part thereof.
12. The following formula (I), (II), (III) or (IV): 【Chemical Formula 4】 (wherein, Sm represents a target binding moiety, L represents a linker, X is -S-R 4 , -S(O)-R 4 , -O-R 5 or -N(R 6 )-R 7 wherein R 1 、 R 2 and R 3 are each independently a hydrogen atom, a halogen, an alkyl which may have a substituent, an alkenyl which may have a substituent, an alkynyl which may have a substituent, an alkoxy which may have a substituent, an aryl which may have a substituent, an aralkyl which may have a substituent, a cycloalkyl which may have a substituent, or a heteroaryl which may have a substituent, or R 1 and R 2 or R 2 and R 3 together form a ring which may have a substituent, R 4 represents alkyl which may have a substituent, aryl which may have a substituent, or heteroarylalkyl which may have a substituent, R 5 represents a hydrogen atom or an alkyl group which may have a substituent, R 6 and R 7 each independently represents a hydrogen atom, an alkyl group which may have a substituent, or an aryl group which may have a substituent, or R 6 and R 7 together form a ring which may have a substituent.) Prepare a plurality of compounds represented by, Contacting the plurality of compounds with one or more target RNAs; Determining the nucleotide sequence of complementary DNA synthesized by reverse transcriptase using the target RNA contacted with the compound as a template; Selecting a compound that interacts with each of the target RNAs based on the nucleotide sequence of the complementary DNA in which the bases modified by the compound are skipped and one or several bases are deleted on each of the target RNAs; A method for identifying the structure of a target binding moiety that regulates the function of the target RNA, comprising:
13. The method according to claim 12, wherein the target RNA is an RNA that forms a guanine quadruplex.
14. The method according to claim 12 or 13, further comprising a step of concentrating the RNA modified by contacting with the compound.
Citation Information
Patent Citations
Medium-small-length RNA high-throughput sequencing library building method and application thereof
CN110952148A
Compounds and methods for treating RNA-mediated diseases
JP2019511562A