RNA-targeting ligand, composition thereof, and method of preparation and use thereof

A fragment-based screening method using SHAPE and SHAPE-MaP RNA structure detection identifies high-affinity ligands for RNA molecules, addressing inefficiencies in current RNA-targeting technologies and enabling regulation of cellular states and diseases.

KR102993514B1Active Publication Date: 2026-07-21더유니버시티오브노쓰캐롤라이나엣채플힐
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
더유니버시티오브노쓰캐롤라이나엣채플힐
Filing Date
2020-08-05
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Current methods for identifying small molecule ligands that target RNA molecules are inefficient and lag behind in developing high-affinity inhibitors that can regulate RNA function, despite RNA's crucial role in biological processes and disease states.

Method used

A fragment-based screening strategy using SHAPE and SHAPE-MaP RNA structure detection to identify small molecule fragments that bind to TPP riboswitches with high affinity, followed by structure-activity-relationship studies to design linked fragment ligands with nanomolar affinity.

Benefits of technology

Enables the rapid and efficient identification of small molecule ligands that bind to RNA molecules with high affinity, potentially regulating cellular states and diseases, and can be applied to various RNA targets beyond TPP riboswitches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 112022029354843-PCT00082_ABST
    Figure 112022029354843-PCT00082_ABST
Patent Text Reader

Abstract

The disclosure of the present invention relates to a compound that binds to a target RNA molecule, e.g., a TPP riboswitch, a composition comprising said compound, and a method for preparing and using said compound. said compound contains two structurally different fragments that enable binding to the target RNA at two different binding sites, thereby producing a binding ligand with higher affinity compared to a compound that binds to only a single RNA binding site.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The disclosure of the present invention relates to a compound that binds to a target RNA molecule, e.g., a TPP riboswitch, a composition comprising said compound, and a method for preparing and using said compound. said compound contains two structurally different fragments that enable binding to the target RNA at two different binding sites, thereby producing a binding ligand with higher affinity compared to a compound that binds to only a single RNA binding site.

[0002] Reference citation in sequence list

[0003] The data in the attached sequence list is incorporated herein by reference in its entirety. The attached file named sequence list 39397600002_ST25 was created on August 5, 2020, and is 4 KB.

[0004] Government support

[0005] This invention was made with government support under approval numbers GM098662 and AI068462 granted by the National Institutes of Health (NIH). The government holds specific rights to this invention. Background Technology

[0006] The majority of small molecule ligands are developed to manipulate biological systems primarily by targeting proteins. Proteins possess highly complex three-dimensional structures, which are important for their proper functioning and contain gaps and pockets where small molecule ligands can bind. 1,2 Transcripts—the set of all RNA molecules produced in an organism—also contain promising targets for studying and manipulating biological systems. For example, RNA transcripts not only play an important role in mammalian systems, but they are also present in both bacteria and viruses and thus regulate gene expression by providing targets for small molecules.

[0007] It is worth noting the key features required for the development of highly selective ligands. 4 A complex three-dimensional structure comparable to the structure of a protein 3 It can be adopted, and RNA plays a pervasive role in governing the behavior of biological systems. 5 Originally, RNA was viewed solely as a carrier of genetic information existing to transmit messages that guide protein encoding and protein biosynthesis processes; however, the modern perspective on RNA has evolved to encompass expanded roles, where a diverse range of RNA molecules are now understood to play a wide-ranging role in regulating gene expression and other biological processes through various mechanisms. Furthermore, the majority of newly discovered non-coding RNAs have been found to be associated with diseases such as cancer and non-tumorous disorders. Therefore, the reality that RNA contributes to disease states, distinct from encoding pathogenic proteins, offers a wealth of previously unrecognized therapeutic targets.

[0008] However, although small molecule ligands can bind to mRNA and have been shown to have the potential to regulate protein expression in cells by up- or down-regulating decoding efficiency, 6,7 There are problems related to the identification of small molecule RNA that were not encountered when targeting proteins. 4,11,12 It also includes the development of small molecules directed to non-coding RNAs that exhibit a rich target pool. 8-10 Unfortunately, while various techniques have been developed for the analysis of RNA structures and the discovery of new functions, the ability to efficiently and rapidly identify or design inhibitors that bind to RNA and interfere with its function is lagging far behind. Therefore, there is a great demand in the industry for the development of new methods and techniques that enable the rapid and efficient identification of small-molecule ligands that target RNA molecules.

[0009] outline

[0010] As previously mentioned above, the transcript represents an attractive but underutilized set of targets for small molecule ligands. Small molecule ligands (and ultimately drugs) targeted to messenger RNA and non-coding RNA have the potential to regulate cellular states and diseases. In the disclosure herein, small molecule fragments that bind to target RNA structures were discovered using a fragment-based screening strategy utilizing selective 2'-hydroxyl acylation (SHAPE) and SHAPE-mutation profiling (MaP) RNA structures analyzed by primer extension. In particular, fragments binding to TPP riboswitches with millimolar to micromolar affinities and cooperative binding fragment pairs were identified. Structure-activity-relationship (SAR) studies were performed to obtain information for efficiently designing linked fragment ligands that bind to TPP riboswitches with high nanomolar affinity. The principles of the present disclosure are not limited to TPP riboswitches, but may be broadly applied to other target RNA structures by utilizing cooperative and multisite binding to develop high-quality ligands for various RNA targets.

[0011] As such, one aspect of the main content disclosed herein is a compound having the structure of Formula I or a pharmaceutically acceptable salt thereof:

[0012] Chemical Formula I

[0013]

[0014] In the above formula,

[0015] X1, X2, and X3 are CR1, CHR1, respectively. Independently selected from N, NH, O, and S, wherein adjacent X1, X2, and X3 are not simultaneously selected as being O or S;

[0016] The dashed line indicates any double combination;

[0017] Y1, Y2, and Y3 are each independently selected from CR2 and N in each case;

[0018] n is 1 or 2, where, if n is 1, only one of the dashed lines is a double bond;

[0019] L is

[0020] , , , and Selected from,

[0021] Here, p, q, r, and v are independently selected from integers 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10, and z is selected from 1, 2, 3, 4, and 5;

[0022] A is

[0023] , , and Selected from,

[0024] Here, X4, X5, X6, and X7 are independently selected from CR3 and N;

[0025] Here, R1, R2, and R3 are independently selected from -H, -Cl, -Br, -I, -F, -CF3, -OH, -CN, -NO2, -NH2, -NH(C1-C6alkyl), -N(C1-C6alkyl)2, -COOH, -COO(C1-C6alkyl), -CO(C1-C6alkyl), -O(C1-C6alkyl), -OCO(C1-C6alkyl), -NCO(C1-C6alkyl), -CONH(C1-C6alkyl), and substituted or unsubstituted C1-C6alkyls;

[0026] m is 1 or 2 and;

[0027] W is -O or -NR4, where R4 is selected from -H, -CO(C1-C6alkyl), substituted or unsubstituted C1-C6alkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted cycloalkyl, -CO(aryl), -CO(heteroaryl), and -CO(cycloalkyl);

[0028] However, at least two of X1, X2, X3, X4, X5, X6, and X7 are N.

[0029] Further aspects of the main content described herein include compounds as described herein that bind to regions of RNA molecules.

[0030] Further aspects of the main content described herein include compositions comprising a therapeutically effective amount of the compound described herein in a pharmaceutically acceptable carrier, diluent, or excipient.

[0031] Further aspects of the main content described herein include a method for treating a disease or disorder associated with dysfunction in RNA expression, said method comprising the step of administering a therapeutically effective dose of a compound described herein to a subject in need thereof.

[0032] Further aspects of the main content disclosed herein include a method for producing the compounds described herein.

[0033] Further aspects of the main points described herein are provided below. Brief explanation of the drawing

[0034] Fig. 1 Figure 1 shows a schematic of the RNA screening construct and fragment screening workflow. It depicts RNA motifs 1 and 2, the barcode helix; and the structure cassette helix. RNA is detected using SHAPE in the presence or absence of small molecule fragments, and chemical modifications corresponding to ligand-dependent structural information are read by multiplexed MaP sequencing. Fig. 2Figure 1 shows representative mutation rates for fragment hits and non-hits. Normalized mutation rates for fragment-exposed samples are labeled as +ligand, +2, or +4 and compared to ligand-free traces labeled as ligand-free. Statistically significant changes in mutation rates are indicated by triangles (refer to Figure 8 for SHAPE verification data). (Top) Comparison of mutation rates for representative compounds that do not bind to the test construct. (Center) Fragment hits on the TPP riboswitch region of RNA. (Bottom) Non-specific hits inducing reactivity changes across the test construct. Motif 1 and 2 landmarks are shown below the SHAPE profile. Figures 3a and 3b (Fig. 3a) fragment 17 versus (Fig. 3b) intrinsic TPP ligand (2HOJ 28 This shows a comparison of the structures of TPP riboswitches coupled by ). RNA structures are shown with similar orientations in each image. Hydrogen bonds between the ligand and RNA are indicated by dashed lines. Figures 4a and 4b shows the stepwise ligand binding affinities for the thermodynamic cycle and fragments 2 and 31. Fig. 4a shows compound 2 (dark brown, K 1) and compound 31 (light brown, K 2) Shows an overview of the combination by fragments. K D The value is determined by ITC. Figure 4b shows ITC data illustrating cooperative binding by the single compound and fragments 2 and 31. The linkage of the two fragments exhibits an additive effect in binding energy, and the sub-micromolar ligand compound 37 ( K L Induces ). The ITC trace is shown as light brown as the baseline trace (ligand is titrated with buffer) and as dark brown as the experimental trace. The curve pit is shown as a 95% confidence interval in the brown shade. Fig. 5Figure 17 shows the covalent linkage of fragments 17 and 31 as a function of linker type and length, terminal group chemistry type, and terminal group orientation. Modifications increasing RNA binding affinity are present in compounds 36 and 37 (dark brown); negative modifications are present in compounds 35, 39, and 40 (dark brown); and a neutral modification is present in compound 38 (dark brown). Dissociation constants determined by ITC. Fig. 6 Figure 37 shows a comparison of fragment-linker-fragment ligands developed by the fragment-based method and aligned by their linkage coefficients (E). Values ​​are plotted on the logarithmic axis. Cooperative linkage corresponds to lower E values ​​(top of the vertical axis). Fragment 37 exhibits an E value of 2.5 and an LE value of 0.34. Dissociation constants for individual fragments (left, center) and linked ligands (right) are shown below the component fragments; E values ​​(top) and ligand efficiencies (bottom) are indicated. Covalent linkages introduced between fragments are highlighted in light brown. The structures for the component fragments are shown in detail in Table 7. Figures 7a and 7b shows the screening construction design. Fig. 7a is an RNA sequence having the following components ( Sequence number 6 Shows ): GGUCGCGAGUAAUCGCGACC ( Sequence number 7 ) is a structural cassette and; G CU G CA AGAGAU UG U AG C ( Sequence number 8 ) is an RNA barcode (barcode NT underlined); GUGGGCACUUCGGUGUCCAC ( Sequence number 9 ) is a structure cassette; ACGCGA AG GAAACCGCGUGUC A ACUGUGCAACAGCUGACAAAGAGAUUC CU ( Sequence number 10 ) is a DENV pseudoknot (mutation in bold); AAAACU is a linker; CAGUACUCGGGGUGCCCUUCUGCGUGAAGGCUGAGAAAUACCCGUAUCACCUGAUCUGGGAAUAAUGCCAGCGUAGGGAAGU G C UG ( Sequence number 11 ) is a TPP riboswitch (mutant in bold); GAUCCGGUUCGCCGGAUCAAUCGGGCUUCGGUCCGGUUC ( Sequence No. 12 ) is a structural cassette. Fig. 7b shows the secondary structure of the RNA-sequence barcode in relation to its self-folding hairpin. Fig. 8 Figure 1 shows the SHAPE profiles for non-hit, hit, and non-specific hit fragments. Mutation rate traces corresponding to fragment-exposed and ligand-absent control traces are indicated by solid brown shades and black outlines, respectively. Nucleotides determined to be statistically significantly different between fragment-exposed and fragment-absent samples are indicated by triangles. Mutation rate traces for the same fragment are schematically shown in Figure 2. Specific details for implementing the invention

[0035] The main points disclosed herein are described more fully from now on. However, many variations and other embodiments of the main points disclosed herein will be recalled to those skilled in the art to which the main points disclosed herein pertain and benefit from the teachings provided in the prior description. Accordingly, it should be understood that the main points disclosed herein are not limited to the specific embodiments described, and that variations and other embodiments are intended to be included within the scope of the appended claims. In other words, the main points disclosed herein include all alternatives, variations, and equivalents. In the event that one or more of the integrated literature, patents, and similar materials, including but not limited to defined terms, usage of terms, and described techniques, differ from or contradict the present application, the present application shall prevail. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. All publications, patent applications, patents, and other reference literature mentioned herein are incorporated by reference in their entirety.

[0036] definition

[0037] The term "alkyl group" as used herein refers to a saturated hydrocarbon radical containing 1 to 8, 1 to 6, 1 to 4, or 5 to 8 carbons. In some embodiments, the saturated radical contains more than 8 carbons. The alkyl group is structurally similar to an acyclic alkane compound modified by the removal of one hydrogen from the acyclic alkane and thus the substitution of a non-hydrogen group or radical. The alkyl group radical may be branched or unbranched. Lower alkyl group radicals have 1 to 4 carbon atoms. Higher alkyl group radicals have 5 to 8 carbon atoms. Examples of alkyl, lower alkyl, and higher alkyl group radicals include, but are not limited to, radicals such as methyl, ethyl, n-propyl, isopropyl, n-butyl, secondary butyl, t-butyl, amyl, t-amyl, n-pentyl, n-hexyl, i-octyl, etc.

[0038] As used herein, the names "(CO)" and "C(O)" are used to denote carbonyl moiety. Examples of suitable carbonyl moiety include, but are not limited to, ketone and aldehyde moiety.

[0039] The term "cycloalkyl" refers to a hydrocarbon having 3-8, 3-7, 3-6, 3-5, or 3-4 members and may be monocyclic or bicyclic. The ring may be saturated or may have some degree of unsaturation. The cycloalkyl group may optionally be substituted with one or more substituents. In one embodiment, 0, 1, 2, 3, or 4 atoms of each ring of the cycloalkyl group may be substituted by one substituent. Representative examples of cycloalkyl groups include cyclopropyl, cyclopentyl, cyclohexyl, cyclobutyl, cycloheptyl, cyclopentenyl, cyclopentadienyl, cyclohexenyl, cyclohexadienyl, etc.

[0040] The term "aryl" refers to a hydrocarbon monocyclic, bicyclic, or tricyclic aromatic ring system. An aryl group may optionally be substituted with one or more substituents. In one embodiment, 0, 1, 2, 3, 4, 5, or 6 atoms of each ring of the aryl group may be substituted by a single substituent. Examples of aryl groups include phenyl, naphthyl, anthracenyl, fluorenyl, indenyl, azulenyl, etc.

[0041] The term "heteroaryl" refers to an aromatic 5- to 10-membered ring system, wherein the heteroatom is selected from O, N, or S, and the remaining ring atom is a carbon (having an appropriate hydrogen atom unless otherwise indicated). The heteroaryl group may optionally be substituted with one or more substituents. In one embodiment, 0, 1, 2, 3, or 4 atoms of each ring of the heteroaryl group may be substituted by a single substituent. Examples of heteroaryl groups include pyridyl, furanyl, thienyl, pyrrolyl, oxazolyl, oxadiazolyl, imidazolyl, thiazolyl, isoxazolyl, quinolinyl, pyrazolyl, isothiazolyl, pyridazinyl, pyrimidinyl, pyrazinyl, triazinyl, isoquinolinyl, indazolyl, etc.

[0042] The term “substituted” as used herein refers to a moiety (e.g., heteroaryl, aryl, alkyl and / or alkenyl), said moiety being bonded to one or more additional organic or inorganic substituent radicals. In some embodiments, the substituted moiety comprises one, two, three, four, or five additional substituent groups or radicals. Suitable organic and inorganic acid substituent radicals include, but are not limited to, hydroxyl, cycloalkyl, aryl, substituted aryl, heteroaryl, heterocyclic ring, substituted heterocyclic ring, amino, monosubstituted amino, disubstituted amino, acyloxy, nitro, cyano, carboxy, carboalkoxy, alkyl carboxamide, substituted alkyl carboxamide, dialkyl carboxamide, substituted dialkyl carboxamide, alkylsulfonyl, alkylsulfinyl, thioalkyl, alkoxy, substituted alkoxy, or haloalkoxy radicals as defined herein. Unless otherwise indicated herein, the organic substituent may comprise 1 to 4 or 5 to 8 carbon atoms. Where a substituted moiety is bonded to more than one substituent radical, the substituent radicals may be the same or different.

[0043] The term “unsubstituted” as used herein refers to a moiety (e.g., heteroaryl, aryl, alkenyl, and / or alkyl) that is not bonded to one or more additional organic or inorganic acid substituent radicals as described above, and means that said moiety is substituted only with hydrogen.

[0044] It is understood that the structure and "substitution" or "substituted with" provided herein include an absolute proviso that said structure and substitution depend on the allowed valence of the substituted atom and the substituent, and said substitution induces a stable compound that does not undergo spontaneous transformation, such as by rearrangement, ring closure, removal, etc.

[0045] The term "RNA," as used herein, refers to ribonucleic acid, a polymeric molecule essential for various biological roles in encoding, decoding, regulating, and expressing genes. RNA and DNA are nucleic acids and, along with lipids, proteins, and carbohydrates, constitute the four major macromolecules essential to all known forms of life. Like DNA, RNA is assembled as a chain of nucleotides; however, unlike DNA, RNA is found naturally as a single strand that folds itself rather than as a paired double strand. Cellular organisms use messenger RNA (mRNA) to convey genetic information (using the nitrogenous bases guanine, uracil, adenine, and cytosine, designated by the letters G, U, A, and C) that directs the synthesis of specific proteins. Many viruses use RNA genomes to encode their genetic information. Some RNA molecules perform active roles within cells by regulating gene expression or catalyzing biological responses that detect and communicate responses to cellular signals. One of these active processes is protein synthesis, a universal function in which RNA molecules direct the synthesis of proteins on ribosomes. The above process uses a transfer RNA (tRNA) molecule that delivers amino acids to a ribosome, where ribosomal RNA (rRNA) then links the amino acids together to form an encoded protein.

[0046] The term "non-coding RNA (ncRNA)" as used herein refers to RNA molecules that are not decoded into proteins. The DNA sequence from which functional non-coding RNA is transcribed is commonly referred to as an RNA gene. Abundant and functionally important types of non-coding RNA include transfer RNA (tRNA) and ribosomal RNA (rRNA), small RNAs, e.g., microRNA, siRNA, piRNA, snoRNA, snRNA, exRNA, scaRNA, and long ncRNAs, e.g., Xist and HOTAIR.

[0047] The term "coding RNA" as used herein refers to RNA that encodes proteins, namely messenger RNA (mRNA). Such RNA includes a transcriptome.

[0048] The term "riboswitch" as used herein refers to a regulatory segment of a messenger RNA molecule that binds to a small RNA molecule and induces a change in the production of a protein encoded by the mRNA. Thus, mRNA containing a riboswitch is directly involved in regulating its own activity in response to the concentration of its effector molecule.

[0049] The term “TPP riboswitch,” as used herein and also known as the THI element and Thi-box riboswitch, refers to a highly conserved RNA secondary structure. It acts as a riboswitch that directly binds to thiamine pyrophosphate (TPP) and regulates gene expression through various mechanisms in archaea, bacteria, and eukaryotic cells. TPP is the active form of thiamine (vitamin B1), an essential coenzyme synthesized in bacteria by the coupling of a pyrimidine and a thiazole moiety.

[0050] The term "pseudoknot" as used herein refers to a nucleic acid secondary structure containing at least two stem-loop structures, wherein half of one stem is inserted between two halves of another stem. Pseudoknots were first recognized in the turnip yellow mosaic virus in 1982. Pseudoknots fold into a three-dimensional knot-like shape but are not true morphological knots.

[0051] "Aptamers" refer to nucleic acid molecules capable of binding to specific molecules of interest with high affinity and specificity (Tuerk and Gold, 1990; Ellington and Szostak, 1990), and may be of human processed or natural origin. The binding of a ligand to an aptamer, which is typically RNA, alters the conformation of the aptamer and the nucleic acid within it on which the aptamer is located. In some cases, this conformational alteration inhibits the decoding of the mRNA on which the aptamer is located or interferes with the normal activity of the nucleic acid. Aptamers may also consist of DNA or include non-natural nucleotides and nucleotide analogs. Aptamers are mostly obtained by in vitro selection for binding to target molecules. However, in vivo selection of aptamers is also possible. Aptamers are also the ligand-binding domains of riboswitches. Aptamers are typically about 10 to about 300 nucleotides long. More typically, aptamers are about 30 to about 100 nucleotide long. For example, refer to U.S. Patent No. 6,949,379, which is incorporated herein by reference. Examples of aptamers useful for the present invention include, but are not limited to, PSMA aptamers (McNamara et al., 2006), CTLA4 aptamers (Santulli-Marotto et al., 2003), and 4-1BB aptamers (McNamara et al., 2007).

[0052] The term "PCR" as used herein refers to the polymerase chain reaction and describes a widely used method in molecular biology that rapidly produces millions to billions of copies of a specific DNA sample, enabling scientists to acquire DNA from a very small sample and amplify it to a quantity large enough to study it in detail.

[0053] The term "pharmaceuticalally acceptable" indicates that a substance or composition is chemically and / or toxicologically compatible with other components constituting the formulation and / or subjects treated with them.

[0054] The term "pharmaceutically acceptable salt" as used herein refers to an acceptable organic acid or inorganic acid salt of a compound of the present invention. Exemplary salts include sulfates, citrates, acetates, oxalates, chlorides, bromides, iodides, nitrates, bisulfates, phosphates, acid phosphates, isonicotinates, lactates, salicylates, acid citrates, tartrates, oleates, tannates, pantothenates, bitartrates, ascorbates, succinates, maleates, gentisinates, fumarates, gluconates, glucuronates, saccharates, formates, benzoates, glutamates, methanesulfonates "mesylates", ethanesulfonates, benzenesulfonates, p-toluenesulfonates, pamoate (i.e., 1,1'-methylene-bis-(2-hydroxy-3-naphthoate)) salts, alkali metal (e.g., sodium and potassium) salts, and alkaline earth metal (e.g., magnesium) salts. Salts, and ammonium salts, are included but not limited thereto. Pharmaceutically acceptable salts may contain inclusions of other molecules, such as acetate ions, succinate ions, or other counterions. The counterions may be any organic or inorganic acid moiety that stabilizes the charge on the parent compound. Additionally, pharmaceutically acceptable salts may have more than one charged atom in their structure. Where multiple charged atoms are part of the pharmaceutically acceptable salt, the salt may have multiple counterions. Thus, pharmaceutically acceptable salts may have one or more charged atoms and / or one or more counterions.

[0055] "Carriers" as used herein comprise pharmaceutically acceptable carriers, excipients, or stabilizers that are non-toxic to cells or mammals exposed thereto at the doses and concentrations used. Commonly, pharmaceutically acceptable carriers are aqueous pH buffer solutions. Non-limiting examples of physiologically acceptable carriers include buffers such as phosphates, citrates, and other organic acids; antioxidants including ascorbic acid; low molecular weight (less than about 10 residues) polypeptides; proteins such as serum albumin, gelatin, or immunoglobulin; hydrophilic polymers such as polyvinylpyrrolidone; amino acids such as glycine, glutamine, asparagine, arginine, or lysine; monosaccharides, disaccharides, and other carbohydrates including glucose, mannose, or dextrin; chelating agents such as EDTA; sugar alcohols such as mannitol or sorbitol; salt-forming counterions such as sodium; and / or include nonionic surfactants such as TWEEN™, polyethylene glycol (PEG), and PLURONICS™. In certain embodiments, the pharmaceutically acceptable carrier is a non-naturally occurring pharmaceutically acceptable carrier.

[0056] The terms “treat” and “treatment” refer to both therapeutic treatment and prophylactic or preventive measures, wherein the objective is to prevent or slow down unintended physiological changes or disorders, such as the onset or spread of cancer. For the purposes of the present invention, beneficial or desired clinical outcomes include alleviation of symptoms, reduction in the severity of the disease, a stabilized (i.e., non-deterioration) state of the disease, delay or deceleration of disease progression, improvement or alleviation of the disease state, and remission (partial or complete), whether or not they are detected. “Treatment” may also mean extending survival rates compared to the expected survival rate in the absence of treatment. Those requiring treatment include conditions or disorders, or those prone to having conditions or disorders that must be prevented, as well as those already having conditions or disorders.

[0057] The terms "administration" or "administering" include a route of introducing the compound(s) to a subject to perform their intended function. Examples of administration routes that may be used include injection (including, but not limited to, subcutaneous, intravenous, parenteral, intraperitoneal, and intrathecal), topical, oral, inhalation, rectal, and transdermal.

[0058] The term "effective dose" includes an amount effective for the dose and duration required to achieve the desired result. The effective dose of a compound may vary depending on factors such as the subject's disease status, age and body weight, and the compound's ability to induce the desired response in the subject. The dosage regimen may be adjusted to provide an optimal therapeutic response.

[0059] As used herein, the terms “systemic administration,” “systemically administered,” “peripheral administration,” and “peripherally administered” refer to the administration of compound(s), drugs, or other substances so as to be applied to metabolism and other similar processes as they enter the patient’s system.

[0060] The term “therapeutic effective dose” means an amount of a compound of the present invention that (i) treats or prevents a specific disease, condition, or disorder; (ii) alleviates, improves, or eliminates one or more symptoms of a specific disease, condition, or disorder; or (iii) prevents or delays the onset of one or more symptoms of a specific disease, condition, or disorder described herein. In the case of cancer, the therapeutic effective dose of the drug may reduce the number of cancer cells; reduce tumor size; inhibit cancer cell invasion into peripheral organs (i.e., to some extent slow and preferably stop); inhibit tumor metastasis (i.e., to some extent slow and preferably stop); inhibit tumor growth to some extent; and alleviate one or more symptoms associated with cancer to some extent. There may be inhibition of cell proliferation and / or cytotoxicity to the extent that the drug prevents growth and / or kills existing cancer cells. In the case of cancer therapy, efficacy may be measured, for example, by evaluating the time to disease progression (TTP) and / or determining the response rate (RR).

[0061] The term "object" refers to animals such as mammals, including but not limited to primates (e.g., humans), cattle, sheep, goats, horses, dogs, cats, rabbits, rats, mice, etc. In certain embodiments, the object is a person.

[0062] The present disclosure relates to a fragment-based ligand discovery strategy suitable for the identification of small molecules that bind to specific RNA regions with high affinity. Generally, fragment-based ligand discovery enables the identification of one or more small molecule "fragments" with low to intermediate affinity that bind to a target of interest. These fragments are then refined or linked to generate more potent ligands. 13,14Typically, these fragments exhibit a molecular weight of less than 300 Da and make significant, high-quality contact with the target of interest to bind detectably.

[0063] Fragment-based ligand discovery has been successfully used to identify early hit compounds, which are single-fragment hit bindings to specific RNAs. 15-19 The identification of multiple fragments binding to the same RNA will enable the utilization of potential additive and cooperative interactions between fragments within the binding pocket. 20,21 However, it has recently been shown that many RNAs bind to ligands through multiple "subsites," which are binding pocket regions that contact the ligand in an independent or cooperative manner. 22 Additionally, it was shown that high-affinity RNA binding can occur even when subsite binding exhibits only a slight cooperative effect. These characteristics are a good indication for the effectiveness of fragment-based ligand discovery as applied to RNA targets.

[0064] Accordingly, based on the foregoing, the present disclosure relates to a method for identifying a fragment binding to RNA of interest, such as, for example, a TPP riboswitch. Secondly, the method disclosed herein relates to establishing the localization of fragment binding within RNA approximately in nucleotide separation. Thirdly, the method disclosed herein relates to identifying a second-site fragment bound near the site of an initial fragment hit. The method disclosed herein combines a fragment-based ligand discovery approach with SHAPE-MaP RNA structure detection. 23,24 This was used to identify both RNA binding fragments and establish individual sites of fragment binding. The ligand ultimately produced by linking the two fragments is not similar to natural riboswitch ligands and binds with high affinity to structurally complex TPP riboswitch RNA.

[0065] The identification of the method and ligand disclosed herein will be described in more detail below.

[0066] A. Compounds

[0067] A first aspect of the main content disclosed herein is a compound having the structure of Formula I or a pharmaceutically acceptable salt thereof:

[0068] Chemical Formula I

[0069]

[0070] In the above formula,

[0071] X1, X2, and X3 are CR1, CHR1, respectively. Independently selected from N, NH, O, and S, wherein adjacent X1, X2, and X3 are not simultaneously selected as being O or S;

[0072] The dashed line indicates any double combination;

[0073] Y1, Y2, and Y3 are each independently selected from CR2 and N in each case;

[0074] n is 1 or 2, where, if n is 1, only one of the dashed lines is a double bond;

[0075] L is

[0076] , , , , and Selected from, where k, p, q, r, and v are independently selected from integers 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10, and z is selected from integers 1, 2, 3, 4, and 5;

[0077] A is

[0078] , , and Selected from,

[0079] Here, X4, X5, X6, and X7 are independently selected from CR3 and N;

[0080] Here, R1, R2, and R3 are independently selected from -H, -Cl, -Br, -I, -F, -CF3, -OH, -CN, -NO2, -NH2, -NH(C1-C6alkyl), -N(C1-C6alkyl)2, -COOH, -COO(C1-C6alkyl), -CO(C1-C6alkyl), -O(C1-C6alkyl), -OCO(C1-C6alkyl), -NCO(C1-C6alkyl), -CONH(C1-C6alkyl), and substituted or unsubstituted C1-C6alkyls;

[0081] m is 1 or 2 and;

[0082] W is -O or -NR4, where R4 is selected from -H, -CO(C1-C6alkyl), substituted or unsubstituted C1-C6alkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted cycloalkyl, -CO(aryl), -CO(heteroaryl), and -CO(cycloalkyl);

[0083] However, at least two of X1, X2, X3, X4, X5, X6, and X7 are N.

[0084] As in any of the above embodiments, the compound is at least one of X1, X2, or X3 being N.

[0085] As in any of the above embodiments, the compound is X1 is N.

[0086] As in any of the above embodiments, the compound is X2 is N.

[0087] As in any of the above embodiments, the compound is X3 is N.

[0088] As in any of the above embodiments, in each case, two of X1, X2, and X3 of the compound are N.

[0089] As in any of the above embodiments, the compound is X1 and X3 being N.

[0090] As in any of the above embodiments, the compound is N, at least one of Y1, Y2, and Y3.

[0091] As in any of the above embodiments, the compound is Y1 is N.

[0092] As in any of the above embodiments, the compound is Y2 N.

[0093] As in any of the above embodiments, the compound is Y3 N.

[0094] As in any of the above embodiments, the compounds are at least one of Y1, Y2, and Y3 being CR2.

[0095] As in any of the above embodiments, the compound Y1 is CR2.

[0096] As in any of the above embodiments, the compound is Y2 is CR2.

[0097] As in any of the above embodiments, the compound Y3 is CR2.

[0098] As in any of the above embodiments, the compound has n of 2.

[0099] As in any of the above embodiments, the compound has the structure of Formula II:

[0100] Chemical Formula II

[0101]

[0102] In the above formula,

[0103] X 2a and X 2b is independently selected from CR1 and N;

[0104] X1 and X3 are independently selected from CR1 and N;

[0105] L and A are as provided for Chemical Formula I;

[0106] X1, X 2a , X 2b , and 2 of X3 are N.

[0107] As in any of the above embodiments, the compound has the structure of Formula III:

[0108] Chemical Formula III

[0109]

[0110] In the above formula,

[0111] L and A are as provided for Chemical Formula I.

[0112] As in any of the above embodiments, the compounds are selected independently from integers p, q, r, and v of 0, 1, 2, and 3.

[0113] As in any of the above embodiments, the compound L is selected from the following:

[0114] , , and .

[0115] As in any of the above embodiments, the compound is L

[0116] am.

[0117] As in any of the above embodiments, the compound has q and r 0 or 1.

[0118] As in any of the above embodiments, the compound has q of 1.

[0119] As in any of the above embodiments, the compound has r of 1.

[0120] As in any of the above embodiments, the compound has r of 0.

[0121] As in any of the above embodiments, the compound has q and r as 1.

[0122] As in any of the above embodiments, the compound has q of 1 and r of 0.

[0123] As in any of the above embodiments, the compound has m of 1.

[0124] As in any of the above embodiments, the compound W is selected from -NH, -O, and -N(C1-C6alkyl)2.

[0125] As in any of the above embodiments, the compound has W as -NH.

[0126] As in any of the above embodiments, the compound is at least one of X4, X5, X6 and X7 being N.

[0127] As in any of the above embodiments, the compound is X4 N.

[0128] As in any of the above embodiments, the compound is X5 is N.

[0129] As in any of the above embodiments, the compound is X6 is N.

[0130] As in any of the above embodiments, the compound is X7 is N.

[0131] As in any of the above embodiments, the compound is X4 and X6 being N.

[0132] As in any of the above embodiments, the compound is X5 and X7 is N.

[0133] As in any of the above embodiments, the compound is such that X5 or X6 is N, and X4 and X7 are both independently CR2.

[0134] As in any of the above embodiments, the compound is A

[0135] am.

[0136] As in any of the above embodiments, the compound has the following structure:

[0137] .

[0138] As in any of the above embodiments, the compound is L

[0139] or am.

[0140] As in any of the above embodiments, the compounds Y1, Y2, and Y3 are each independently selected from CR2 and N, where R1 is selected from -H, -Cl, -Br, -I, -F, -OH, and -NH2.

[0141] As in any of the above embodiments, the compound has z of 2.

[0142] As in any of the above embodiments, the compound is Y2 N.

[0143] As in any of the above embodiments, the compound is such that Y2 is CR2 and R1 is selected from -H, -F, -OH, and -NH2.

[0144] As in any of the above embodiments, the compound is A

[0145] am.

[0146] As in any of the above embodiments, the compound has the following structure:

[0147] or

[0148] .

[0149] As in any of the above embodiments, the compound has the following structure:

[0150] , , or .

[0151] B. Screening Method

[0152] The disclosure of this application relates to the development and validation of flexible-selective 2'-hydroxyl acylations analyzed by a primer extension (SHAPE)-based fragment screening method. Fragment-based ligand discovery has proven to be an effective approach for identifying compounds that form significant intimate contacts with macromolecules containing RNA. 13,14,17A prerequisite for the success of the above discovery strategy is a high-quality biophysical assay that can be adopted to detect ligand binding. Accordingly, in some embodiments, ligand binding was detected using SHAPE RNA structure detection, and 23-25 , this measures local nucleotide flexibility as the relative reactivity of the ribose 2'-hydroxyl group toward an electrophilic reagent. SHAPE can be used on all RNA and provides data for almost all nucleotides of RNA in a single experiment, generating structural information per nucleotide in addition to simply detecting binding, which is described in more detail below. Additionally, the disclosure of this invention also covers SHAPE-mutation profiling (MaP) 23,24 This relates to the application of SHAPE, which combines SHAPE with readings by high-speed sequencing to enable multiplexing of thousands of samples and efficient high-speed analysis.

[0153] Accordingly, in some embodiments, the present disclosure relates to a screening method using SHAPE and / or SHAPE-MaP to identify small molecule fragments and / or compounds that bind to and / or associate with an RNA molecule of interest. The method described herein further comprises using SHAPE and / or SHAPE-MaP to identify a small molecule fragment (e.g., fragment 2) that binds to and / or associates with an RNA molecule that has been pre-incubated with another small molecule fragment (e.g., fragment 1). Although not bound by theory, it is believed that fragment 1 binds to a first binding site on the same RNA molecule and fragment 2 binds to a second binding site (e.g., a subsite). Accordingly, it is believed that the combination of structural characteristics of fragment 1 and fragment 2 to produce a compound as described herein (e.g., linking the two fragments with linker L) confers a linked fragment ligand with increased RNA binding affinity compared to fragment 1 and / or fragment 2 alone.

[0154] The screening methods SHAPE and SHAPE-MaP are described in more detail below.

[0155] I. SHAPE Chemistry

[0156] SHAPE chemistry is based, at least in part, on the observation that the nucleophilicity of the RNA ribose 2'-position is sensitive to the electronic influence of the adjacent 3'-phosphodiester group. Unbound nucleotides sample more forms that form base pairs or otherwise enhance the nucleophilicity of the 2'-hydroxyl group compared to bound nucleotides. Thus, hydroxyl-selective electrophiles, such as N-methylisatosan anhydride (NMIA), form stable 2'-O adducts more rapidly using flexible RNA nucleotides, though they are not limited to this. Since all RNA nucleotides (with the exception of a small number of cellular RNAs undergoing post-transcriptional modification) possess a 2'-hydroxyl group, local nucleotide flexibility at all positions of an RNA molecule can be investigated simultaneously in a single experiment. Because 2'-hydroxyl reactivity is insensitive to base identity, absolute SHAPE reactivity can be compared at all positions of RNA. In addition, nucleotides may be reactive because they are constrained in a form that enhances the nucleophilicity of a specific 2'-hydroxyl. Nucleotides of this class are expected to be rare, contain non-canonical local geometry, and are correctly scored as unpaired positions.

[0157] The main point disclosed herein is to provide a method for detecting structural data in RNA molecules by investigating structural constraints in RNA molecules of any length and structural complexity in some embodiments. In some embodiments, the method comprises the steps of: annealing an RNA molecule containing a 2'-O-adduct with a (labeled) primer; annealing an RNA molecule not containing a 2'-O-adduct with a (labeled) primer as a negative control; extending the primer to generate a cDNA library; analyzing the cDNA; and generating an output file containing structural data for the RNA.

[0158] RNA molecules may be present in biological samples. In some embodiments, RNA molecules may be modified in the presence of proteins or other small and large biological ligands and / or compounds. Primers may optionally be labeled with radioisotopes, fluorescent labels, heavy atoms, enzyme labels, chemiluminescent groups, biotinyl groups, predetermined polypeptide epitopes recognized by secondary reporters, or combinations thereof. The analysis may include separation, quantification, sizing, or a combination thereof. The analysis may include extracting fluorescence or dye quantity data as a function of elution time data referred to as trace. For example, cDNA may be analyzed in a single column of a capillary electrophoresis instrument or in a microflow device.

[0159] In some embodiments, the maximum area under trace for RNA molecules containing a 2'-O-adduct and for RNA molecules not containing a 2'-O-adduct can be calculated. The trace can be compared and aligned with the sequence of RNA. The trace describing the cDNA generated by sequencing is a single nucleotide longer than the corresponding position under trace for RNA molecules containing a 2'-O-adduct and for RNA molecules not containing a 2'-O-adduct. The area under each peak can be determined by performing full trace Gaussian-fit integration.

[0160] Accordingly, the present invention provides a method for forming a covalent ribose 2'-O-adduct having an RNA molecule in a complex biological solution in some embodiments. In some embodiments, the method comprises contacting an electrophile with an RNA molecule, wherein the electrophile selectively modifies an unconstrained nucleotide within the RNA molecule to form a covalent ribose 1'-O-adduct.

[0161] In some embodiments, an electrophile, such as N-methylisatosan anhydride (MMIA), is dissolved in an anhydrous polar aprotic solvent such as DMSO. The reagent-solvent solution is added to a complex biological solution containing RNA molecules. The solution may contain proteins, cells, viruses, lipids, mono- and polysaccharides, amino acids, nucleotides, DNA, and different salts and metabolites at different concentrations and amounts. The concentration of the electrophile can be adjusted to achieve a desired degree of modification within the RNA molecule. The electrophile has the potential to react with any free hydroxyl groups in the solution to form a ribose 2'-O-adduct on the RNA molecule. Additionally, the electrophile can selectively modify nucleotides that do not form pairs or are otherwise unconstrained in the RNA molecule.

[0162] RNA molecules can be exposed to electrophiles at concentrations that produce a rare RNA modification to form a 2'-O-adduct, which can be detected by the ability to inhibit primer extension by reverse transcriptase. Because the chemical targets the general reactivity of the 2'-hydroxyl group, all RNA sites can be investigated in a single experiment. In some embodiments, a control extension reaction with the electrophile excluded and a dideoxy sequencing extension to assign nucleotide positions may be performed in parallel to evaluate the background. These combined steps are referred to as primer extension or selective 2'-hydroxyl acylation analyzed by SHAPE.

[0163] In some embodiments, the method further comprises the steps of: contacting an RNA molecule containing a 1'-O-adduct with a (labeled) primer; contacting an RNA that does not contain a 2'-O-adduct as a negative control with a (labeled) primer; extending the primer to generate an array of linear cDNA; analyzing the cDNA; and generating an output file containing structural data for the RNA.

[0164] The number of nucleotides investigated in a single SHAPE experiment varies depending on the characteristics of the RNA modification as well as the detection and resolution capabilities of the separation technique used. Given specific reaction conditions, almost all RNA molecules have a length with at least one modification. As primer extensions reach this length, the amount of extended cDNA decreases, causing the experimental signal to attenuate. Adjusting conditions to reduce modification yield can increase the read length. However, a decrease in reagent yield can also reduce the signal measured for each cDNA length. Taking these considerations into account, the preferred maximum length of a single SHAPE read is preferably about 1 kilobase of RNA, but should not be limited thereto.

[0165] II. SHAPE-MaP

[0166] In SHAPE-MaP, SHAPE adducts are detected by mutation profiling (MaP), which utilizes the ability of reverse transcriptase to insert non-complementary nucleotides or create deletions at sites of SHAPE chemical adducts. In some embodiments, SHAPE-MaP can be used for library construction and sequencing. In some embodiments, multiplexing technology can be used in SHAPE-MaP.

[0167] Generally, RNA is treated with a SHAPE reagent that reacts morphologically with kinetic nucleotides. During reverse transcription, polymerase reads through chemical adducts in RNA and incorporates nucleotides non-complementary to the original sequence into cDNA. The obtained cDNA is sequenced using any large-scale parallel approach to generate mutation profiles (MaPs). Sequence reads are aligned to a reference sequence, nucleotide-separated mutation rates are calculated, corrected against the baseline, and normalized to generate standard SHAPE reactivity profiles. SHAPE reactivity can then be used to model secondary structures, visualize competing and alternative structures, or quantify any processes or functions regulating local nucleotide RNA dynamics. After SHAPE modification of the RNA molecule, mutation profiles are generated using reverse transcriptase. This step encodes the location and relative frequency of SHAPE adducts as mutations within the cDNA. cDNA is converted to dsDNA using a method known in the art (e.g., PCR reaction), and the dsDNA is further amplified in a second PCR reaction to add sequencing analysis for multiplexing. After purification, the sequencing library has a uniform size, and each DNA molecule contains the entire sequence of interest.

[0168] Accordingly, according to some embodiments of the main gist disclosed herein, a method for detecting one or more chemical modifications in nucleic acids is provided. In some embodiments, the method comprises the steps of: providing a nucleic acid suspected of having a chemical modification; synthesizing a nucleic acid using a polymerase and the provided nucleic acid as a template, wherein the synthesis is performed under conditions in which the polymerase reads through the chemical modification of the provided nucleic acid to produce an incorrect nucleotide in the nucleic acid obtained at the site of the chemical modification; and detecting the incorrect nucleotide.

[0169] According to some embodiments of the main gist disclosed herein, a method for detecting structural data within a nucleic acid is provided. In some embodiments, the method comprises the steps of: providing a nucleic acid suspected of having a chemical modification; synthesizing a nucleic acid using a polymerase and the provided nucleic acid as a template, wherein the synthesis is performed under conditions in which the polymerase reads through the chemical modification of the provided nucleic acid to produce an incorrect nucleotide in the nucleic acid obtained at the site of the chemical modification; detecting the incorrect nucleotide; and generating an output file containing structural data for the provided nucleic acid.

[0170] In some embodiments of the main gist disclosed herein, the nucleic acid provided is an RNA molecule (e.g., coding RNA and / or non-coding RNA molecule). In some embodiments, the method comprises the step of detecting two or more chemical modifications. In some embodiments, the polymerase is read through multiple chemical modifications to produce a number of incorrect nucleotides, and the method comprises the step of detecting each of the incorrect nucleotides.

[0171] In some embodiments, the nucleic acid (e.g., RNA molecule) is exposed to a reagent that provides a chemical modification, or said chemical modification is already present in the nucleic acid (e.g., RNA molecule). In some embodiments, the already present modification is a 2'-O-methyl group and / or is / is not limited to an epigenetic modification generated by the cell from which the nucleic acid is derived, such as an epigenetic modification, or said modification is 1-methyladenosine, 3-methylcytosine, 6-methyladenosine, 3-methyluridine, and / or 2-methylguanosine. In some embodiments, the nucleic acid, such as an RNA molecule, may be modified in the presence of proteins or other small and large biological ligands and / or compounds.

[0172] In some embodiments, the reagent comprises an electrophile. In some embodiments, the electrophile optionally modifies an unconstrained nucleotide within an RNA molecule to form a covalent ribose 2'-O-adduct. In some embodiments, the reagent is 1 M7, 1 M6, NMIA, DMS, or a combination thereof. In some embodiments, the nucleic acid is present in or derived from a biological sample.

[0173] In some embodiments, the polymerase is a reverse transcriptase. In some embodiments, the polymerase is a native polymerase or a mutant polymerase. In some embodiments, the synthesized nucleic acid is cDNA.

[0174] In some embodiments, detection of incorrect nucleotides includes the step of sequencing the nucleic acid. In some embodiments, the sequence information is aligned with the sequence of the provided nucleic acid. In some embodiments, detection of incorrect nucleotides includes the step of using bulk parallel sequencing of the nucleic acid. In some embodiments, the method includes the step of amplifying the nucleic acid. In some embodiments, the method includes the step of amplifying the nucleic acid using a site-directed approach using specific primers, a whole-genome approach using random priming, a whole-transcriptome approach using random priming, or a combination thereof.

[0175] According to some embodiments of the principal claims disclosed herein, a computer program product comprising computer-executable instructions implemented on a computer-readable medium is provided when performing a step comprising any method step of any embodiment of the principal claims disclosed herein. According to some embodiments of the principal claims disclosed herein, a nucleic acid library generated by any method of the principal claims disclosed herein is provided.

[0176] III. SHAPE Electrophile

[0177] As disclosed herein, SHAPE chemistry utilizes the finding that the nucleophilic reactivity of the ribose 2'-hydroxyl group is gated by local nucleotide flexibility. In nucleotides constrained by base pairing or ternary interactions, 3'-phosphodiester anions and other interactions reduce the reactivity of the 2'-hydroxyl. In contrast, the flexible position preferentially adopts a form that reacts with electrophiles, including but not limited to MMIA, to form a 2'-O-adduct. For example, NMIA generally reacts with all four nucleotides, and the reagent proceeds with a parallel self-inactivating hydrolysis reaction. Indeed, the main point disclosed herein provides that any molecule capable of reacting with nucleic acids as described herein may be used according to some embodiments of the main point disclosed herein. In some embodiments, the electrophile (also referred to as the SHAPE reagent) may be selected from, but is not limited to, isotosan anhydride derivatives, benzoyl cyanide derivatives, benzoyl chloride derivatives, phthalic anhydride derivatives, benzyl isocyanate derivatives, and combinations thereof. The isotosan anhydride derivative may include 1-methyl-7-nitroisatosan anhydride (1M7). The benzoyl cyanide derivative may be selected from the group including, but not limited to, benzoyl cyanide (BC), 3-carboxybenzoyl cyanide (3-CBC), 4-carboxybenzoyl cyanide (4-CBC), 3-aminomethylbenzoyl cyanide (3-AMBC), 4-aminomethylbenzoyl cyanide, and combinations thereof. The benzoyl chloride derivative may include benzoyl chloride (BCl). Phthalic anhydride derivatives may include 4-nitrophthalic anhydride (4NPA). Benzyl isocyanate derivatives may include benzyl isocyanate (BIC).

[0178] IV. RNA Molecular Design

[0179] SHAPE reactivity can be evaluated in one or more primer extension reactions, and information may be lost at both 5' ends and near the primer binding site of the RNA molecule. Typically, adduct formation at 10 to 20 nucleotides adjacent to the primer binding site is difficult to quantify due to interruption by reverse transcriptase (RT) enzymes during the initial stage of primer extension or the presence of cDNA fragments reflecting non-template extension. It may be difficult to visualize the presence of abundant full-length extension products at 8-10 positions at the 5' end of the RNA.

[0180] To monitor SHAPE reactivity at the 5' and 3' ends of the sequence of interest, the RNA molecule may be embedded within a larger fragment of the original sequence or positioned between strongly folded RNA sequences containing unique primer binding sites. In some embodiments, a structural cassette containing 5' and 3' flanking sequences of nucleotides is designed so that any location within the RNA molecule of interest can be evaluated in any separation technique that imparts nucleotide separation capabilities, such as sequence analysis gels or capillary electrophoresis, but is not limited to this. In some embodiments, both the 5' and 3' extensions may be folded into a stable hairpin structure that does not interfere with the folding of various internal RNAs. The primer binding sites of the cassette can efficiently bind to cDNA primers. The sequences of the 5' and 3' structural cassette elements can be examined to ensure that they do not tend to form stable base-pairing interactions with the internal sequence.

[0181] In some embodiments, the RNA molecule of interest comprises two different target motifs linked to a nucleotide linker. The target motif may be any nucleotide sequence of the molecule of interest. Exemplary target motifs include, but are not limited to, riboswitches, viral regulatory elements, and structural regions in mRNA, multi-helical junctions, pseudoknots and / or aptamers. In some embodiments, the first target motif is a pseudoknot, for example, a pseudoknot from the 5' UTR of the dengue fever virus genome. In some embodiments, the second target motif is an aptamer domain, such as a TPP riboswitch aptamer domain. For the nucleotide linker, the number of nucleotides may vary. For example, in some embodiments, the number of nucleotides in the linker is in the range of about 1 to about 20 nucleotides, about 1 to about 15 nucleotides, about 1 to about 10 nucleotides, or about 5 to about 10 nucleotides (or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides).

[0182] In some embodiments, the RNA molecule additionally comprises an RNA barcode region. The RNA barcode region is a distinctive barcode that enables the identification of a specific RNA molecule in a mixture of RNA molecules (e.g., during multiplexing). The location of the RNA barcode region may vary, but is typically found adjacent to one of the cassettes present in the RNA molecule. In some embodiments, the RNA barcode is designed to fold into a self-contained structure that does not interact with any other part of the RNA molecule. The structure of the RNA barcode region may vary. In some embodiments, the structure of the RNA barcode region comprises a base pair helix containing about 1 to about 10 base pairs (or about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 base pairs). In some embodiments, the RNA barcode region comprises 7 base pairs. In some embodiments, the base pairs are capped with tetraloops fixed to the terminal base pairs of the base pair helix. The capping of the base pair helix maintains the overall hairpin stability of the RNA barcode region. In some embodiments, the tetraloop includes the nucleotide sequence GNRA, but is not meant to be limited thereto. In some embodiments, the RNA barcode region is designed to undergo at least two mutations so that any individual barcode is misinterpreted as another barcode.

[0183] V. Folding of RNA molecules

[0184] The main points disclosed herein may be performed with RNA molecules produced by methods including, but not limited to, in vitro transcription and RNA molecules generated in cells and viruses. In some embodiments, RNA molecules may be purified and regenerated by denaturing gel electrophoresis to achieve a biologically relevant form. Additionally, any procedure for folding RNA molecules into a desired form at a desired pH (e.g., about pH 8) may be replaced. RNA molecules may first be heated and rapidly cooled in a low ionic strength buffer to remove the multimeric form. A folding solution is then added to allow RNA molecules to achieve a suitable form and be prepared for structure-sensitive probing with electrophiles. In some embodiments, RNA may be folded in a single reaction and subsequently separated by (+) and (-) electrophilic reactions. In some embodiments, RNA molecules do not naturally fold prior to modification. Modification may occur while RNA molecules are denatured by heating and / or low-salt conditions.

[0185] VI. RNA Molecular Modification

[0186] Electrophiles can be added to RNA to form 2'-O adducts at flexible nucleotide positions. Subsequently, the reaction can be incubated until virtually all electrophiles react with the RNA or are degraded by hydrolysis with water. No specific quenching step is required. Modification can be performed in the presence of various salts as well as complex ligands and biomolecules. RNA can also be modified within cells and viruses. These salts and complex ligands may include salts of magnesium, sodium, manganese, iron, and / or cobalt. Complex ligands may include, but are not limited to, proteins, lipids, other RNA molecules, DNA, or organic acid small molecules. In some embodiments, the complex ligand is a small molecule fragment as described herein. In some embodiments, the complex ligand is a compound as described herein. The modified RNA may be purified from reaction products and buffer components that may be detrimental to the primer extension reaction, for example, by ethanol precipitation.

[0187] VII. Primer Extension and Polymerization

[0188] According to the main points disclosed herein, the analysis of RNA adducts by primer extension may, in various embodiments, involve the use of optimized primer binding sites, heat-stable reverse transcriptases, low MgCl2 concentrations, elevated temperatures, short extension times, and any combination thereof. Intact, undissolved RNA free from reaction byproducts and other small molecule contaminants may also be used as a template for reverse transcription. The RNA component of the obtained RNA-cDNA hybrid may be degraded by treatment with a base. Subsequently, the cDNA fragment may be separated using, for example, polyacrylamide sequencing gels, capillary electrophoresis, or other separation techniques, as will be apparent to those skilled in the art after reviewing the disclosures herein.

[0189] Deoxyribonucleotide triphosphates dATP, dCTP, dGTP, and dTTP and / or deoxyribonucleotide triphosphate (dNTP) may be added to the synthesis mixture in appropriate amounts, either together with or separately from the primer, and the resulting solution may be heated to about 50-100°C for about 1 to 10 minutes. After the heating period, the solution may be cooled. In some embodiments, a suitable agent for carrying out the primer extension reaction may be added to the cooled mixture, and the reaction is allowed to take place under conditions known in the art. In some embodiments, the polymerization agent may be added together with other reagents if it is heat-stable. In some embodiments, the synthesis (or amplification) reaction may take place at room temperature. In some embodiments, the synthesis (or amplification) reaction may take place up to a temperature at which the polymerization agent is no longer functional.

[0190] The polymerization agent may be any compound or system capable of achieving the synthesis of a primer extension product, for example, including an enzyme. Suitable enzymes for this purpose include, but are not limited to, E. coli DNA polymerase I, the Klenow fragment of E. coli DNA polymerase, polymerase mutain, reverse transcriptase, and other enzymes that are heat-stable (i.e., enzymes that perform primer extension after exposure to a temperature raised sufficiently to cause denaturation), for example, murine or algal reverse transcriptase enzymes. Suitable enzymes may facilitate the combination of nucleotides in an appropriate manner to form a primer extension product complementary to each polymorphic locus nucleic acid strand. In some embodiments, synthesis may be initiated at the 5' end of each primer and proceed in the 3' direction until synthesis is terminated at the template end by the incorporation of dideoxynucleotide triphosphate or a 2'-O-adduct, thereby producing molecules of various lengths.

[0191] The newly synthesized strand and its complementary nucleic acid strand may form a double-stranded molecule under the hybridization conditions described herein, said hybrid being used in a subsequent step as disclosed in the methods described in U.S. Patent No. 10,240,188 and U.S. Patent No. 8,318,424, the full text of which is incorporated herein by reference. In some embodiments, the newly synthesized double-stranded molecule may also be subjected to denaturation conditions using any procedure known in the art to provide a single-stranded molecule.

[0192] VII. Processing of Raw Data

[0193] The principal content described herein regarding nucleic acids, such as RNA molecules, chemical modification analysis, and / or nucleic acid structure analysis may be implemented using a computer program product comprising computer-executable instructions implemented on a computer-readable medium. Exemplary computer-readable media suitable for implementing the principal content described herein include chip memory devices, disk memory devices, programmable logic devices, and application-specific integrated circuits. Additionally, the computer program product implementing the principal content described herein may be located on a single device or computing platform or distributed across multiple devices or computing platforms. Accordingly, the principal content described herein may include a set of computer instructions that perform specific functions for nucleic acids, such as RNA structure analysis, when executed by a computer.

[0194] Considering the aforementioned items I–VII, the modular RNA screening construct was designed to implement SHAPE as a fast-processing assay for reading ligand bindings (Fig. 1, top). The construct targets two motifs, for example, the 5' UTR of the dengue virus genome, which reduces viral fitness when its structure is disrupted. 26 and designed to contain pseudoknots from TPP riboswitch aptamer domains. 27-29Including two distinct structural motifs within a single construct allows each to serve as an internal specificity control for the other. Fragments bound to both RNA structures were readily identified as non-specific complexes. These two structures are connected by a six-nucleotide linker designed to be single-stranded, ensuring that the two RNA structures remain structurally independent. Flanking the structural core of the construct is a structural cassette. 25 These stem-loop-forming regions are used as primer-bonding sites for steps required in the screening workflow and are designed not to interact with other structures within the material (Fig. 7).

[0195] Another component of the screening construct is the RNA barcode; barcoding enables multiplexing, which significantly reduces downstream workload. In the 96-well plate used to screen the fragment library, each well contains RNA with a unique barcode within the context of another identical construct, and thus the barcode sequence identifies the well location and the fragment (or fragments) present after multiplexing (Fig. 1). The RNA barcode region is designed to fold into a self-contained structure that does not interact with any other part of the construct. The barcode structure is a 7-base-pair helical structure capped with a GNRA tetraloop and is anchored with GC base pairs to maintain hairpin stability (Fig. 7). Each set of 96 barcodes is designed so that any individual barcode undergoes two or more mutations to be misinterpreted as another barcode.

[0196] The above construct provides flexibility in selecting RNA structures for screening for ligand binding and supports simple screening experiments (Fig. 1). In a 96-well plate containing the same RNA construct with a unique RNA barcode, each well is incubated with one or several small molecule fragments or a control (solvent) lacking fragments, and then exposed to the SHAPE reagent. The resulting SHAPE adduct chemically encodes structural information per nucleotide. After SHAPE detection, the information required to determine the fragment entity (RNA barcode) and fragment binding (SHAPE adduct pattern) is permanently encoded in each RNA strand; thus, RNA from the 96 wells of the plate can be pooled into a single sample. Fragment screening experiments are processed in a manner very similar to the standard MaP structure-detection workflow. 24 For example, in some embodiments, a specialized relaxed fidelity reverse transcription reaction is used to produce cDNA containing a non-template-encoding sequence change at any location of a shape adduct on RNA. 30 These cDNAs are subsequently used to prepare DNA libraries for high-speed sequencing. The experimental multiple plates are barcoded at the DNA library level. 24 It is possible to collect data for thousands of compounds in a single sequencing operation (Fig. 1). The obtained sequencing data contains hundreds of individual reads, each corresponding to a specific RNA strand. These reads are classified by barcode to enable the analysis of data for each small molecule fragment or combination of fragments. The determination and identification of small molecule fragments (e.g., fragment 1 and / or fragment 2) using the methods described above, such as SHAPE and / or SHAPE-MaP, are described in more detail in the following section.

[0197] C. Ligand Identification and Selection

[0198] As mentioned above, SHAPE and SHAPE-MaP were used to identify small molecule fragments that bind to or are associated with RNA molecules of interest. In particular, when testing small molecule fragments using SHAPE-MaP, the detection of the associated fragment signature with a SHAPE-MaP mutation rate per nucleotide involves several steps to normalize data across large-scale experimental screens and ensure statistical rigor. Key characteristics of the SHAPE-based hit analysis strategy include: (i) comparison of each fragment-exposed RNA or "experimental sample" with five negative fragment-absent control samples to account for plate-to-plate and well-to-well variability; (ii) hit detection performed independently for each of the two structural motifs within the composition in the present disclosure, namely the pseudoknot and the TPP riboswitch; and (iii) shielding of individual nucleotides with low reactivity across all samples, as these nucleotides are unlikely to exhibit fragment-induced changes. and (iv) calculation of the difference per nucleotide in mutation rates between fragment-exposed experimental samples and fragment-absent negative control samples. These nucleotides having a difference of 20% or more in mutation rates between one of the motifs and the fragment-absent control were selected for Z-score analysis. However, those skilled in the art recognize that the difference in mutation rates may vary and may adjust the difference in mutation rates accordingly. For example, in some embodiments, the difference in mutation rates may be 25%, 30%, 35%, 45%, or 50% or more. In some embodiments, the difference in mutation rates may be 15%, 10%, or 5% or more. A fragment is defined as having three or more nucleotides with a Z-value greater than 2.7 in one of the two motifs (determined by comparison of the Poisson number for the two motifs). 31(Refer to Example 2) It was determined that the SHAPE reactivity pattern is significantly altered. However, Z-values ​​may vary, and those skilled in the art may adjust them accordingly. For example, in some embodiments, the Z-value is 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, or greater than 3.9. In some embodiments, the Z-value is 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, or greater than 2.6.

[0199] A series of steps is performed to identify small molecule fragments having a SHAPE and / or SHAPE-MaP that are subsequently linked together to produce a compound as disclosed herein. First, a primary screening is performed to screen a number of compounds, e.g., at least 100 compounds, to identify any initial lead or hit compound exhibiting suitable binding activity for a target RNA molecule. In Step 2, these hit compounds are further investigated in a Structure-Activity-Relationship (SAR) study, in which changes in target RNA binding affinity are determined as the structure of the hit compound is modified. If a number of small molecule fragments are identified as suitable binding ligands for the target RNA molecule, additional binding studies may be performed to further investigate the binding sites for each small molecule fragment (i.e., Step 3). For example, in some embodiments, the target RNA is pre-incubated with the first fragment (identified as the target RNA binding ligand according to the SAR study in Step 2) before exposure to the second fragment of the target RNA (identified as the RNA binding ligand in the SAR study of Step 2) to determine whether the second fragment can bind to the target RNA if the first fragment is already bound. If the second fragment having suitable binding activity to the RNA of interest is identified, the first fragment can be linked with a linker to produce a compound as disclosed herein (i.e., Step 4). Each of the aforementioned steps is described in more detail below.

[0200] Step 1: Primary Screening

[0201] In the first screening, 1,500 fragments were tested, and 41 fragments were detected as hits, resulting in an initial hit rate of 2.7%. Hit confirmation was performed via triple SHAPE analysis (Figs. 2 and 8), and compounds were accepted as true hits only if they were detected as binders in all three replicates. Subsequently, these replicated hit compounds were analyzed by isothermal titration calorimetry (ITC) to determine their binding affinity to RNA corresponding only to the target motif (excluding flanking sequences from the screening constructs). Of these initial hits, eight hits were confirmed by replication analysis and ITC (Table 1). Seven of the hits bound to the TPP riboswitch based on their mutation signatures located mostly or throughout the TPP riboswitch region of the test constructs. The remaining hits were non-specific because the fragments affected nucleotides across all parts of the RNA constructs. No compound that specifically binds to the dengue fever-like knot region of the test specimen was detected.

[0202] [Table 1] Fragment binding to the TPP riboswitch as detected by SHAPE detection.

[0203]

[0204] Hits were detected by SHAPE structure detection and verified by replicate analysis and ITC. Dissociation constants were determined by ITC; error values ​​marked with ‡ represent standard errors derived from ≥3 replicates, and other error estimates are calculated based on 95% confidence intervals for least squares regression of the coupling curve. Intrinsic TPP ligands are included for comparison.

[0205] The seven fragments binding to the TPP riboswitch, as confirmed by the ITC, have diverse chemical types; most have little to no similarity to the native TPP ligand (Table 1). Overall, heteroaromatic nitrogen-containing rings are predominant; these are likely to participate in hydrogen bonding interactions. Three compounds have pyridine rings, and two have pyrazine rings. Azole ring moiety is present in three compounds: two thiadiazoles and one imidazole. The native TPP ligand has a thiazole ring, but this moiety does not participate in binding interactions with RNA. 28,29,33 Additionally, many of the identified fragments contain primary amines, esters and ethers, and fluorine groups that can act as hydrogen bond acceptors or donors.

[0206] Step 2: Structure-Activity Relationship (SAR) of Riboswitch-Coupling Fragments

[0207] Next, analogs of some of the initial hits were investigated with the goal of identifying sites where fragment hits could be modified into linkers to increase binding affinity without interfering with binding. In particular, analogs of compounds 2 and 5 were considered because these two fragments are structurally distinct and analogs are commercially available. Analog-RNA binding was evaluated by ITC. Sixteen analogs of 2 were tested. Modification of the core-quinoxaline structure of 2 by removing one or both of the ring nitrogens induced changes in binding activity (Table 2A).

[0208] [Table 2A] SAR for Fragment 2 analogs.

[0209]

[0210] The modification of the quinoxaline core was investigated, and the dissociation constant was obtained by ITC.

[0211] Improvements in binding affinity resulted from the introduction of methylene-linked hydrogen bond donors or acceptors (Table 2B, compounds 16 and 17). Various substituents at different positions on the quinoxaline ring core reduced binding activity. Compound 2 was a good candidate for further development based on the high degree of flexibility and even binding improvement observed when a substituent was modified at the C-6 position.

[0212] [Table 2B] Structure-activity relationships for analogs of fragment 2 binding to TPP riboswitch RNA. Modifications to the pendant group of the quinoxaline core. Dissociation constants were obtained by ITC.

[0213]

[0214] Next, an examination of 18 analogs of fragment 5 revealed that the corepyridine function of the molecule is important for binding, as changing the ring nitrogen position or adding or removing ring nitrogen reduces or eliminates all bindings (Table 3).

[0215] [Table 3] Structure-activity relationships for analogs of fragment 5 binding to TPP riboswitch RNA. Modifications for the pyridine core and dissociation constant were obtained by ITC.

[0216]

[0217] Modifications to the ring substituents generally resulted in a significant loss of binding activity (Table 4). The only affinity-increasing analog had a chloride at the C-4 position, S12, producing a compound with about 3 times higher affinity for the TPP riboswitch than fragment 5.

[0218] [Table 4] Structure-activity relationships for analogs of fragment 5 binding to TPP riboswitch RNA. Modifications for the pendant group of the pyridine core. Dissociation constants were obtained by ITC.

[0219]

[0220] Step 3: Identification of the fragment binding to the second site on the TPP riboswitch

[0221] Using a second round of screening, fragments bound to the TPP riboswitch region of screening compositions pre-bound to compound 2 or S12 were identified. This screening identified fragments that preferentially interact with the TPP riboswitch when 2 or S12 is already bound, as new binding modes can be solubilized due to cooperative effects or structural changes occurring during primary ligand binding (Fig. 3). Of the 1,500 screened fragments, five were confirmed to bind simultaneously with 2 or S12 (Table 5).

[0222] [Table 5] Fragments binding to the TPP riboswitch in the presence of pre-bound fragment partners upon detection by SHAPE. Hits were confirmed by replicate SHAPE analysis. Primary binding partners (2,6) are shown in Table 1.

[0223]

[0224] One second screening hit (29) induced a very strong change in the SHAPE reactive signal and was shown to cause significant changes in RNA structure, including the unfolding of the P1 helix. As this fragment caused changes in other regions of RNA consistent with non-specific interactions, the fragment was not considered as a candidate for further fragment linkage. Fragment 28 was insoluble at the concentration required for ITC analysis; therefore, related analogs containing pyridine instead of the quinoline ring were investigated by ITC (Table 6). These compounds bound with weak affinity, and nevertheless, 31 and 32 showed clear but slight binding cooperation with 2.

[0225] [Table 6] Structure-activity relationships for analogs of Fragment 28 binding to TPP riboswitch RNA in the presence and absence of pre-bound Fragment 2.*

[0226]

[0227] *und (undetermined) is due to the inability to fit the ITC binding curve; insoluble, compounds insoluble at the concentration required for ITC.

[0228] [Table 7] Detailed comparison of representative proteins and RNA fragment-linker-fragment ligands developed by fragment-based methods. RNA examples are highlighted with an asterisk. Each entry details the two-component fragments and individual Kd values, the linked compound and corresponding Kd values, and the ligand efficiency (LE) and linkage factor (E) for the linked compound. 22,38,53,54,45-52

[0229]

[0230]

[0231] Step 4: Collaboration and Fragment Connection

[0232] The cooperative binding interaction between 2 and 31 was quantified by ITC. Individually, 2 samples were given 25 μM of K d It binds, and 31 are at a much higher K of 10 mM. d It binds. As in the secondary screening, the affinity of fragment 31 was investigated when it pre-binds to divalent TPP riboswitch RNA to form a 2-RNA complex. Under these conditions, fragment 31 binds to approximately 3 mM K d 2 was bound to the 2-TPP RNA complex (Fig. 4). The experiment also showed that 31 binds to the TPP RNA when binding by 2 is saturated, which means that the two fragments do not bind at the same location. By binding to different regions of the TPP RNA with excellent and reasonable affinity, 2 and 31 were linked for the purpose of creating a high-affinity ligand.

[0233] Based on the SAR analysis of fragment hits 2 (Table 2B) and 28 (Table 4), linked analogs of the most promising SAR fragments were prepared by focusing on the aminomethyl position at position 17 and two sites in the pyridine ring of fragment 31 (Fig. 5). First, the affinities of the fragments conjugated with amide or amine linkers were compared. The compound with a flexible amine linker (Compound 36) had a binding affinity five times higher than the amide-linked version (Compound 35, Fig. 5). These linkages are formed by magnesium ions, as they occur in the pyrophosphate moiety of the intrinsic TPP ligand. 35 It was introduced in relation to hydroxamic acid capable of chelating. 27,28 However, amine-linked hydroxyamic acid compound 36 binds with an affinity similar to that of parent fragment 17, suggesting that the hydroxyamic acid moiety does not confer additional binding affinity by chelating ions. Linked compound 37 binds with an affinity of 625 nM—a reasonable approximation—demonstrating that the linkage of two fragments with moderate affinity can achieve a high nanomolar linkage. Substitution of the fragment 31 entity with a tertiary amine (compound 38) reduced the affinity compared to compound 37, suggesting that the interaction between fragment 31 and RNA is mediated solely by the excess effect of the charged base. Finally, altering the linkage between the 17 and 31 moieties by length (compound 39) or the pyridine ring linkage site (compound 40) reduces the affinity for compound 37 (Fig. 5). Ultimately, by linking compounds that bind individually to the TPP riboswitch with affinities of 5.0 μM (Compound 19) and ≥10 mM (Compound 31), 625 nM K d A compound (37) that binds to RNA was produced.

[0234] Those skilled in the art will understand that steps I-IV above are not limiting and serve merely as exemplary embodiments. It will be well understood that steps I-IV above can be applied to identify alternative fragments that can be linked together to provide a compound as disclosed herein having suitable binding affinity for a TPP riboswitch. Additionally, it will be well understood that steps I-IV above can be applied to identify fragments that can be linked together to provide a compound as disclosed herein that binds to another RNA molecule of interest.

[0235] D. Overview and Additional Considerations

[0236] Because both coding (mRNA) and non-coding RNA can potentially be manipulated to alter cellular regulation and disease processes, efforts were made to develop efficient strategies for identifying small molecule ligands on structured RNA. The study disclosed herein demonstrates the prospect of using SHAPE screening readouts to detect ligand binding to RNA combined with a fragment-based strategy. Here, the strategy was used to generate ligands binding to TPP riboswitches that are structurally unrelated to natural ligands, using 625 nM Kd. The combined SHAPE and fragment-based screening approach is general regarding both RNA structures that can be targeted and ligand chemical types that can be developed. Specifically, the strategy is highly suitable for discovering ligands on RNA possessing complex structures essential for identifying RNA motifs that bind within three-dimensional pockets. 4 Additionally, the use of the MaP approach and the application of multiplexing through both RNA and DNA barcoding allow for the efficient screening of many structurally different targets with moderate effort required to screen a library of over a thousand member fragments.

[0237] Many of the obtained ligands were similar to those previously reported for single-round screening also performed on TPP riboswitches. 15,17 Hits within the primary screening were found to be slightly biased toward higher affinities, with most ligands detected by SHAPE binding in the 10–300 μM range. The hit detection assay used may be biased toward the detection of the tightest fragment binders and binders that induce the most substantial changes in SHAPE reactivity. Fragments with lower affinities may be missed. The bias toward tighter binding fragments is considered to be advantageous overall. No fragments bound to dengue-like knots were identified that reached the affinity and specificity required to meet the screening criteria. Dengue-like knot RNA is highly structured, and it may be unlikely that a fragment could disturb said structure. Another possibility is that said specific knot structures may not contain ligand-capable pockets.

[0238] A fragment pair identification strategy, in which fragment hits from primary screening are pre-bound to RNA and screened for additional fragment binding partners, was specifically used to utilize nucleotide-based information obtainable by SHAPE and to discover successfully derived fit pairs (Fig. 4). A key caveat in fragment-based ligand development is that cooperation between two fragments can be achieved through proximal binding, and that additional binding can be utilized by linking cooperative fragments with a minimally invasive covalent linker. 20,21,36,37The development of compound 37 linked from primary and secondary fragment hits demonstrates that fragment-based ligand discovery can be efficiently applied to RNA targets. There is a moderate degree of cooperation between 2 and 31. Binding by the compound was 3 to 10 times stronger than when 2 was pre-bound to RNA. Immediately upon linking these two fragments, a moderate additiveness of their binding energies was observed. 37 had an affinity of 625 nM. Since perfect localization of the fragments was not achieved, no super-addition effects were observed when linking fragments 2 and 31. 36 Small changes in the length or geometry of the linker induced large changes in the affinity of the linked ligand (Fig. 5), which implies that the precise orientation of the linker is important for the optimal orientation of the two fragments. The successful development of compound 37 reveals that it is not necessary to achieve perfection in the degree of cooperation between the fragments or in the configuration of the covalent linker connecting them to efficiently develop sub-micromolar ligands.

[0239] Although many efforts have been designed to utilize fragment-to-fragment cooperation to obtain tight-binding ligands that target proteins, targeting RNA is still in its early stages. The extent to which the SHAPE-based screening strategy disclosed herein couples with fragment linkages compared to previous (protein-centric) efforts was investigated. Compounds previously discovered using fragment-based strategies were ranked according to the linkage coefficient (E), which measured the degree to which the entire system functions together when linked (Fig. 6, Table 7(Extended from ). In the absence of positive or negative contributing factors, the binding energies of the two fragments are exactly additive, the linker is inactive, and E is equal to 1.0. Cooperative effects or favored linker interactions decrease E, while anti-cooperative efforts or negative linker interactions increase E. Crucially, E values ​​can vary on an order of magnitude within the protein system. The linkage coefficient for 37 was 2.5, which was slightly higher than the average for linked (protein-targeted) ligands in the academic literature. 37 has a ligand efficiency (LE) calculated by dividing the binding free energy by the number of non-hydrogen atoms, which compares favorably with examples of linked fragment ligands targeting proteins (Fig. 6). Based on these metrics, 37 functions nearly identically to TPPc, a ligand closely related to the intrinsic TPP riboswitch ligand. 22 Therefore, fragment-based ligand discovery, particularly efficiently implemented by SHAPE-supported multiplexed screening, holds significant prospects for enabling the rapid development of unique ligands that target a vast array of RNA structures.

[0240] E. Manufacturing method

[0241] The disclosure of this application also relates to any method for preparing the compounds disclosed herein. Those skilled in the art will understand that such preparation methods may vary. For example, in some embodiments, a method for preparing the compounds disclosed herein is:

[0242] The step of contacting a fragment of Formula IV with Formula V-1 or V-2 in the presence of a Pd catalyst:

[0243] Chemical Formula IV

[0244]

[0245] In the above formula,

[0246] X1, X2, and X3 are CHR 1,CR1, and heteroatoms N, NH, O, and S are independently selected, wherein adjacent X1, X2, and X3 are not simultaneously selected as being O or S;

[0247] The dashed line indicates any double combination;

[0248] Y1, Y2, and Y3 are independently selected from CR2 and N;

[0249] R1 and R2 are independently selected from -H, -Cl, -Br, -I, -F, -CF3, -OH, -CN, -NO2, -NH2, -NH(C1-C6alkyl), -N(C1-C6alkyl)2, -COOH, -COO(C1-C6alkyl), -CO(C1-C6alkyl), -O(C1-C6alkyl), -OCO(C1-C6alkyl), -NCO(C1-C6alkyl), -CONH(C1-C6alkyl), and substituted or unsubstituted C1-C6alkyls;

[0250] n is selected from integers 1 and 2, where, if n is 1, only one of the dashed lines is a double combination;

[0251] Chemical formula V-1

[0252]

[0253] Chemical formula V2

[0254]

[0255] In the above formula,

[0256] X is a halogen selected from F, Br, Cl, and I;

[0257] X4, X5, X6, and X7 are independently selected from CR3 and N;

[0258] R3 is selected from -H, -Cl, -Br, -I, -F, -CF3, -OH, -CN, -NO2, -NH2, -NH(C1-C6alkyl), -N(C1-C6alkyl)2, -COOH, -COO(C1-C6alkyl), -CO(C1-C6alkyl), -O(C1-C6alkyl), -OCO(C1-C6alkyl), -NCO(C1-C6alkyl), -CONH C1-C6(alkyl), and substituted or unsubstituted C1-C6alkyl;

[0259] m is 1 or 2 and;

[0260] W is -O or -NR4, where R4 is selected from -H, -CO(C1-C6 alkyl), substituted or unsubstituted C1-C6 alkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted cycloalkyl, -CO(aryl), -CO(heteroaryl), and -CO(cycloalkyl).

[0261] In some embodiments, the Pd catalyst is selected from (DPPF)PdCl2, Pd2(dba)3, PdCl2[P(o-tolyl)3]2, Pd(dba)2, and Pd(OAc)2. In some embodiments, the contact step further comprises a phosphine ligand. In some embodiments, the phosphine ligand is monodentate. In some embodiments, the phosphine ligand is bitentate. Exemplary phosphine ligands include, but are not limited to, DPPF, BINAP, and rac-BINAP. In some embodiments, the contact step further comprises a base. In some embodiments, the base is an inorganic acid. In some embodiments, the base is NaOtBu. In some embodiments, the contact step is performed purely (i.e., without a solvent). In some embodiments, the contact step is performed in the presence of a solvent. In some embodiments, the solvent is a nonpolar solvent. Exemplary solvents include, but are not limited to, toluene, benzene, dioxane, and tetrahydrofuran. In some embodiments, the contact step is performed at an elevated temperature. In some embodiments, the contact step is performed at 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, or 100°C.

[0262] In some embodiments, a method for preparing the compound disclosed herein is:

[0263] The step of contacting the fragment of Formula IV with the fragment of Formula VI-1 or VI-2 in the presence of a reducing agent:

[0264] Chemical Formula IV

[0265]

[0266] Chemical formula VI-1

[0267]

[0268] Chemical formula VI-2

[0269]

[0270] In the above formula,

[0271] X1, X2, X3, X4, X5, X6, X7, Y1, Y2, Y3, n, m, and W are as defined above.

[0272] In some embodiments, the reducing agent may be any reducing agent applicable for reductive amination chemistry. Exemplary reducing agents include, but are not limited to, boron hydride and / or aluminum hydride. In some embodiments, the reducing agent is boron hydride. In some embodiments, the reducing agent is sodium borohydride. In some embodiments, the contact step is performed in a pure form. In some embodiments, the contact step is performed in a solvent. Exemplary solvents include, but are not limited to, alcohol solvents (e.g., methanol, ethanol, isopropanol), chlorinated solvents (e.g., dichloromethane), and / or ether solvents (e.g., tetrahydrofuran). In some embodiments, the contact step is performed at a temperature below room temperature. In some embodiments, the contact step is performed at an elevated temperature.

[0273] In some embodiments, a method for preparing the compound disclosed herein is:

[0274] The step of contacting a fragment of Formula IV with a fragment of Formula V-1 or V-2 in the presence of a base:

[0275] Chemical Formula IV

[0276]

[0277] Chemical formula VII-1

[0278]

[0279] Chemical formula VII-2

[0280]

[0281] In the above formula,

[0282] X1, X2, X3, X4, X5, X6, X7, Y1, Y2, Y3, n, m, and W are as defined above;

[0283] G is -F, -Cl, -Br, -OH, -OCH3, or -OCH2CH3.

[0284] In some embodiments, the base is an organic acid (pyridine and / or trimethylamine). In some embodiments, the base is an inorganic acid (e.g., potassium / sodium carbonate and / or potassium / sodium bicarbonate). In some embodiments, the method further comprises, but is not limited to, a coupling agent such as DCC and / or EDCI. In some embodiments, the contact step is performed in the presence of a solvent. In some embodiments, the contact step is performed in the presence of a solvent. Exemplary solvents include, but are not limited to, THF, DCM, ACN, and / or DMSO. In some embodiments, the contact step is performed at room temperature. In some embodiments, the contact step is performed at an elevated temperature.

[0285] F. Composition

[0286] The compounds disclosed herein may be formulated into pharmaceutical compositions with pharmaceutically acceptable carriers.

[0287] The compounds disclosed herein may be formulated as pharmaceutical compositions according to standard pharmaceutical practices. According to the above aspects, a pharmaceutical composition comprising a compound as disclosed herein is provided in combination with a pharmaceutically acceptable diluent or carrier.

[0288] Typical formulations are prepared by mixing a compound as disclosed herein with a carrier, diluent, or excipient. Suitable carriers, diluents, and excipients are widely known to those skilled in the art and include materials such as carbohydrates, waxes, water-soluble and / or swelling polymers, hydrophilic or hydrophobic materials, gelatin, oils, solvents, water, etc. The specific carrier, diluent, or excipient used depends on the means and purpose for which the compound is applied. Solvents are generally selected based on solvents recognized by those skilled in the art as safe for administration to mammals (GRAS). Generally, safe solvents are non-toxic aqueous solvents, e.g., water, and other non-toxic solvents that are water-soluble or water-miscible. Suitable aqueous solvents include water, ethanol, polypropylene glycol, polyethylene glycol (e.g., PEG 400, PEG 300), etc., and mixtures thereof. The formulation may also include one or more buffers, stabilizers, surfactants, wetting agents, lubricants, emulsifiers, suspending agents, preservatives, antioxidants, opacifiers, lubricants, processing aids, coloring agents, sweeteners, flavoring agents, and other known additives to delicately provide aids in the manufacture of drugs (i.e., compounds or pharmaceutical compositions thereof as described herein) or pharmaceutical products (i.e., drugs).

[0289] Formulations may be prepared using conventional dissolution and mixing procedures. For example, a bulk drug substance (i.e., a compound as disclosed herein or a stabilized form of a compound (e.g., a complex with a cyclodextrin derivative or other known complexing agent)) is dissolved in a suitable solvent in the presence of one or more of the excipients mentioned above. The compound is generally formulated into a pharmaceutical dosage form to provide an easily controllable dose of the drug and to enable the patient to comply with the prescribed dosage. Pharmaceutical compositions (or formulations) for application may be packaged in various ways depending on the method used for administering the drug. Generally, articles for distribution include a container in which the pharmaceutical formulation is deposited inside in a suitable form. Suitable containers are widely known to those skilled in the art and include materials such as bottles (plastic and glass), sachets, ampoules, plastic bags, metal cylinders, etc. The container may also include a tamper-proof assembly to prevent inadvertent access to the contents of the package. Additionally, the container has a label attached thereto stating the contents of the container. The label may also include appropriate warnings.

[0290] Pharmaceutical formulations may be prepared for various routes and types of administration. For example, compounds as disclosed herein having a desired degree of purity may optionally be mixed with pharmaceutically acceptable diluents, carriers, excipients, or stabilizers (Remington's Pharmaceutical Sciences (1980) 16th edition, Osol, A. Ed.) in the form of lyophilized formulations, ground powders, or aqueous solutions. Formulation may be performed by mixing with a physiologically acceptable carrier at ambient temperature at an appropriate pH and desired purity, i.e., a carrier that is non-toxic to the recipient at the dose and concentration used. The pH of the formulation may be in the range of about 3 to about 8, although it depends primarily on the specific use and concentration of the compound. A formulation in acetate buffer at pH 5 is a suitable embodiment.

[0291] The compound may be sterile. In particular, the formulation to be used for in vivo administration must be sterile. Such sterilization is easily achieved by filtration through a sterile filter membrane.

[0292] The compound can typically be stored as a solid composition, a freeze-dried formulation, or an aqueous solution.

[0293] Pharmaceutical compositions comprising compounds as disclosed herein may be formulated, administered, and administered in a manner consistent with good medical practice, namely in amount, concentration, schedule, course, vehicle, and route of administration. Factors to be considered in this context include the specific disorder being treated, the specific mammal being treated, the clinical condition of the individual patient, the cause of the disorder, the site of delivery of the formulation, the method of administration, the administration schedule, and other factors known to the healthcare worker. The "therapeutic effective amount" of the compound to be administered will be determined by these considerations and is the minimum amount necessary to prevent, improve, or treat the coagulation factor-mediated disorder. This amount is preferably less than the amount that is toxic to the host or makes the host significantly more susceptible to bleeding.

[0294] Permitted diluents, carriers, excipients, and stabilizers are non-toxic to the recipient at the doses and concentrations used and include buffers such as phosphates, citrates, and other organic acids; antioxidants including ascorbic acid and methionine; preservatives (octadecyldimethylbenzylammonium chloride; hexamethonium chloride; benzalkonium chloride; benzethonium chloride; phenol, butyl, or benzyl alcohol; alkyl parabens such as methyl or propyl paraben; catechol; resorcinol; cyclohexanol; 3-pentanol; and m-cresol); low molecular weight (less than about 10 residues) polypeptides; proteins such as serum albumin, gelatin, or immunoglobulin; hydrophilic polymers such as polyvinylpyrrolidone; and amino acids such as glycine, glutamine, asparagine, histidine, arginine, or lysine; Monosaccharides, disaccharides, and other carbohydrates including glucose, mannose, or dextrin; chelating agents such as EDTA; sugars such as sucrose, mannitol, trehalose, or sorbitol; salt-forming counterions such as sodium; metal complexes (e.g., Zn-protein complexes); and / or nonionic surfactants such as TWEEN™, PLURONICS™, or polyethylene glycol (PEG). Active pharmaceutical ingredients may also be encapsulated in microcapsules or macroemulsions prepared in colloidal drug delivery systems (e.g., liposomes, albumin microspheres, microemulsions, nanoparticles, and nanocapsules) by coacervation technology or interfacial polymerization, for example, hydroxymethylcellulose or gelatin-microcapsules and poly-microcapsules and poly-(methyl methacrylate) microcapsules, respectively. The above technology is described in the literature (Remington's Pharmaceutical Sciences 16th edition, Osol, A. Ed. (1980)).

[0295] Delayed-release formulations of compounds may be manufactured. Suitable examples of delayed-release formulations include a semipermeable matrix of a solid hydrophobic polymer containing a compound as described herein, said matrix being in the form of a molded product, e.g., a film or a microcapsule. Examples of delayed-release matrices include polyesters, hydrogels (e.g., poly(2-hydroxyethyl-methacrylate), or poly(vinyl alcohol)), polylactide (U.S. Patent No. 3,773,919), copolymers of L-glutamic acid and gamma-ethyl-L-glutamate, non-degradable ethylene-vinyl acetate, degradable lactic acid-glycolic acid copolymers, e.g., LUPRON DEPOT™ (injectable microspheres composed of a lactic acid-glycolic acid copolymer and leuprolide acetate), and poly-D-(-)-3-hydroxybutyric acid.

[0296] Formulations include those suitable for the administration routes described in detail herein. Formulations can be readily provided in unit dosage forms and can be prepared by any method widely known in the pharmaceutical field. Techniques and formulations are generally found in the literature (Remington's Pharmaceutical Sciences (Mack Publishing Co., Easton, Pa.)). The method comprises the step of combining an active ingredient with a carrier constituting one or more accessory components. Generally, formulations are prepared by uniformly and intimately combining the active ingredient with a liquid carrier or a finely divided solid carrier, or both, and then, if necessary, forming the product.

[0297] Formulations of compounds as disclosed herein suitable for oral administration may be manufactured as separate units, such as pills, capsules, capsules, or tablets, each containing a predetermined amount of the compound.

[0298] Compressed tablets may be manufactured by compressing an active ingredient in a free-flowing form, e.g., powder or granules, mixed optionally with a binder, lubricant, inert diluent, preservative, surfactant, or dispersant in a suitable machine. Molded tablets may be manufactured by molding a mixture of powdered active ingredients soaked in an inert liquid diluent in a suitable machine. Tablets may optionally be coated or scored and optionally formulated to provide slow release or controlled release of the active ingredient.

[0299] Tablets, lozenges, aqueous or oily suspensions, dispersible powders or granules, emulsions, hard or soft capsules, e.g., gelatin capsules, syrups, or elixirs may be manufactured for oral use. Formulations of compounds as disclosed herein intended for oral use may be manufactured according to any method known in the art for the manufacture of pharmaceutical compositions, and such compositions may contain one or more agents including sweeteners, flavorings, colorings, and preservatives to provide a palatable formulation. Tablets containing the active ingredient are permitted by mixing with non-toxic, pharmaceutically acceptable excipients suitable for tablet manufacturing. These excipients may be, for example, inert diluents such as calcium or sodium carbonate, lactose, calcium or sodium phosphate; granulizing and disintegrating agents such as corn starch or alginate; binders such as starch, gelatin, or acacia; and lubricants such as magnesium stearate, stearic acid, or talc. The tablet may not be coated or may be coated by known techniques including microencapsulation that delays disintegration and adsorption in the gastrointestinal tract to provide sustained action over a long period. For example, time-delaying substances such as glyceryl monostearate or glyceryl distearate may be used alone or in combination with wax.

[0300] For the treatment of the eyes or other external tissues, e.g., the mouth and skin, the formulation may be applied as a topical ointment or cream containing the active ingredient(s) in an amount of, for example, 0.075 to 20% w / w. When formulated as an ointment, the active ingredient may be used with a paraffinic or water-miscible ointment base. Alternatively, the active ingredient may be formulated as a cream with an oil-in-water cream base.

[0301] In some cases, the aqueous phase of the cream base may comprise polyhydric alcohols, namely alcohols having two or more hydroxyl groups such as propylene glycol, butane 1,3-diol, mannitol, sorbitol, glycerol, and polyethylene glycol (including PEG 400), and mixtures thereof. The topical formulation may preferably comprise a compound that enhances the absorption or penetration of the active ingredient through the skin or other affected area. Examples of the skin penetration enhancer include dimethyl sulfoxide and related analogs.

[0302] The oil phase of the emulsion may be composed of known ingredients in a known manner. The phase may comprise only an emulsifier, but may also comprise at least one emulsifier and a mixture of fat or oil, or both fat and oil. A lipophilic emulsifier included together with a hydrophilic emulsifier may act as a stabilizer. Combined, in the presence or absence of the stabilizer(s), the emulsifier(s) constitute a so-called emulsifying wax, and the wax, together with oil and fat, constitutes a so-called emulsifying ointment base that forms an oil dispersion phase of a cream formulation. Emulsifiers and emulsifying stabilizers suitable for use in formulations include Tween® 60, Span® 80, cetostearyl alcohol, benzyl alcohol, myristyl alcohol, glyceryl monostearate, and sodium lauryl sulfate.

[0303] An aqueous suspension of a compound contains an active substance mixed with excipients suitable for the preparation of an aqueous suspension. Such excipients include suspending agents, e.g., sodium carboxymethylcellulose, croscarmellose, povidone, methylcellulose, hydroxypropyl methylcellulose, sodium alginate, polyvinylpyrrolidone, tragagant gum, and acacia gum, and dispersants or wetting agents, e.g., naturally occurring phosphatides (e.g., lexitin), condensation products of alkylene oxides and fatty acids (e.g., polyoxyethylene stearate), condensation products of ethylene oxides and long-chain aliphatic alcohols (e.g., heptadecaethyleneoxycetanol), and condensation products of ethylene oxides and partial esters derived from fatty acids and hexitol anhydrides (e.g., polyoxyethylene sorbitan monooleate). The aqueous suspension may also contain one or more preservatives such as ethyl or n-propyl p-hydroxybenzoate, one or more coloring agents, one or more flavoring agents and one or more sweeteners, for example, sucrose or saccharin.

[0304] The pharmaceutical composition of the compound may be in the form of a sterile injectable formulation, for example, in the form of a sterile injectable aqueous or oily suspension. Such suspensions may be prepared according to known methods using a suitable dispersant or wetting agent and the aforementioned suspenders. Such sterile injectable formulations may also be sterile injectable solutions or suspensions in non-toxic, parenterally acceptable diluents or solvents such as 1,3-butanediol. Sterile injectable formulations may also be prepared as lyophilized powders. Acceptable vehicles and solvents that may be used include water, Ringer's solution, and isotonic sodium chloride. Additionally, sterile fixatives may generally be used as solvents or suspension media. For this purpose, any conventional fixative, including synthetic mono- or di-glycerides, may be used. Additionally, fatty acids such as oleic acid may likewise be used in the preparation of the injectable formulation.

[0305] The amount of active ingredient that can be combined with a carrier material to create a single-dose formulation varies depending on the host to be treated and the specific mode of administration. For example, a time-release formulation intended for oral administration to humans may contain about 1 to 1000 mg of active ingredient, which is mixed in appropriate and convenient amounts that may vary from about 5 to about 95% (weight:weight) of the total composition. Pharmaceutical compositions may be prepared to provide easily measurable amounts for administration. For example, an aqueous solution for intravenous infusion may contain about 1 to 500 μg of active ingredient per milliliter of solution to enable infusion of a suitable volume at a rate of about 10 mL / hour to about 50 mL / hour.

[0306] Formulations suitable for parenteral administration include aqueous and non-aqueous isotonic sterile injectable solutions that provide an isotonic formulation of the preparations with the blood of the intended recipient; and aqueous and non-aqueous sterile suspensions that may include a suspending agent and a thickening agent.

[0307] Formulations suitable for topical administration to the eye also comprise eye drops in which the active ingredient is dissolved or suspended in a suitable carrier, particularly an aqueous solvent for the active ingredient. The active ingredient is preferably present in these formulations at a concentration of about 0.5 to 20% w / w, for example, about 0.5 to 10% w / w, for example, about 1.5% w / w.

[0308] Formulations suitable for local administration in the oral cavity include lozenges containing an active ingredient in a flavor base, generally sucrose and acacia or tragacanth; pills containing an active ingredient in an inert base such as gelatin and glycerin, or sucrose and acacia; and mouthwashes containing an active ingredient in a suitable liquid carrier.

[0309] A formulation for rectal administration may be provided as a suppository with a suitable base, for example, containing cocoa butter or salicylate.

[0310] Formulations suitable for intrapulmonary or nasal administration have particle sizes ranging from 0.1 to 500 microns, for example (including particle sizes ranging from 0.1 to 500 microns in increments of micron, such as 0.5, 1, 30 microns, 35 microns, etc.), and are administered by rapid inhalation through the nasal cavity or by inhalation through the mouth to reach the alveoli. Suitable formulations comprise an aqueous or oil solution of the active ingredient. Formulations suitable for aerosol or dry powder administration may be prepared according to conventional methods and may be delivered together with other therapeutic agents, such as compounds previously used for the treatment or prevention of disorders, as described below.

[0311] Formulations suitable for vaginal administration may be provided as pessaries, tampons, creams, gels, pastes, foams, or spray formulations containing a carrier known to be suitable in the art in addition to the active ingredient.

[0312] The above formulations may be packaged in single-dose or multi-dose containers, e.g., sealed ampoules and vials, and may be stored in a lyophilized (freeze-dried) state requiring the addition of a sterile liquid carrier, e.g., water for injection, immediately before use. Immediate injectable solutions and suspensions were prepared from sterile powders, granules, and tablets of the types previously described. Preferred unit-dose formulations are those containing a daily dose or daily unit sub-dose of the active ingredient as cited herein, or suitable fractions thereof.

[0313] The above main point therefore further provides a veterinary composition comprising at least one active ingredient as defined above, together with a veterinary carrier. The veterinary carrier is a substance useful for the purpose of administering the composition and may otherwise be a solid, liquid, or gaseous substance that is inert or acceptable in the veterinary field and compatible with the active ingredient. These veterinary compositions may be administered parenterally, orally, or by any other intended route.

[0314] In certain embodiments, a pharmaceutical composition comprising the compound disclosed herein further comprises a chemotherapeutic agent. In some of these embodiments, the chemotherapeutic agent is an immunotherapeutic agent.

[0315] G. Treatment Methods

[0316] The compounds and compositions disclosed herein may also be used in methods for treating various diseases and / or disorders identified as associated with dysfunction of RNA expression and / or function, or with the expression and / or function of proteins derived from mRNA, or with a useful role of altering the shape of RNA derived from mRNA or using small molecules, or with a role of altering the intrinsic function of riboswitches in a manner that inhibits the growth of infectious organisms. As such, the methods of the disclosures herein relate to treating disorders associated with RNA expression and / or function or to creating new convertible therapeutic agents. For example, reference is made to U.S. Patent Application Publication No. 2018 / 010146, the full text of which is incorporated herein by reference. As such, in some embodiments, a method for treating a disease or disorder (e.g., associated with dysfunction of RNA expression and / or function) as disclosed herein comprises the step of administering a therapeutically effective dose of a compound and / or composition as disclosed herein to a subject in need thereof.

[0317] Dysfunction in RNA expression is characterized by the overexpression or underexpression of one or more RNA molecule(s). In some embodiments, one or more RNA molecule(s) are associated with promoting the disease and / or disorder to be treated. In some embodiments, the RNA molecule(s) are characterized as being part of a healthy cell apparatus and thus will prevent and / or improve the disease and / or disorder to be treated. In some embodiments, the disease or disorder to be treated is associated with dysfunction of RNA function related to transcription, processing, and / or decoding. In some embodiments, the disease or disorder to be treated is associated with inaccurate expression of a protein as a result of dysfunctional RNA molecule function. In some embodiments, the disease or disorder to be treated is associated with dysfunction of RNA function related to gene expression. In some embodiments, the disease or disorder is a disease or disorder intended to reduce protein expression by binding a molecule to the mRNA. In some embodiments, the disease is advantageously treated by a therapeutic regimen that can be switched on or off using a small molecule. For example, in some embodiments, the disease or disorder is a genetic disease that is required to have the ability to switch the expression of a therapeutic gene to active or inactive.

[0318] Diseases and disorders to be treated include, but are not limited to, degenerative disorders, cancer, diabetes, autoimmune disorders, cardiovascular disorders, coagulation disorders, eye diseases, infectious diseases, and diseases caused by one or more mutations within genes.

[0319] Exemplary degenerative diseases include Alzheimer's disease (AD), amyotrophic lateral sclerosis (ALS, Lou Gehrig's disease), cancer, Charcot-Marie-Tooth (CMT) disease, chronic traumatic encephalopathy, cystic fibrosis, certain cytochrome c oxidase deficiencies (often the cause of degenerative Leigh syndrome), Ehlers-Danlos syndrome, progressive ossifying fibrodysplasia, Friedreich's ataxia, frontotemporal dementia (FTD), certain cardiovascular diseases (e.g., atherosclerosis such as coronary artery disease and aortic stenosis), Huntington's disease, infantile neuroaxonal dystrophy, keratoconus (KC), spherical cornea, leukodystrophy, macular degeneration (AMD), Marfan syndrome (MFS), certain mitochondrial myopathy, and mitochondrial DNA depletion. Syndromes, multiple sclerosis (MS), multiple system atrophy, muscular dystrophy (MD), neuroceroid pilofuscinosis, Niemann-Pick disease, osteoarthritis, osteoporosis, Parkinson's disease, pulmonary hypertension, all prion diseases (Creutzfeldt-Jakob disease, fatal familial insomnia, etc.), progressive supranuclear palsy, retinitis pigmentosa (RP), rheumatoid arthritis, Sandhoff's disease, spinal muscular atrophy (SMA, motor neuron disease), subacute sclerosing panencephalitis, Tay-Sachs disease, and vascular dementia (which is not neurodegenerative in itself but often appears alongside other forms of degenerative dementia), but are not limited thereto.

[0320] Exemplary cancers include, but are not limited to, all forms of carcinoma, melanoma, blastoma, sarcoma, lymphoma and leukemia, e.g., bladder cancer, bladder carcinoma, brain tumor, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, endometrial cancer, hepatocellular carcinoma, laryngeal cancer, lung cancer, osteosarcoma, ovarian cancer, pancreatic cancer, prostate cancer, kidney cancer and thyroid cancer, acute lymphoblastic leukemia, acute myeloid leukemia, ependymoma, Ewing sarcoma, glioblastoma, medulloblastoma, neuroblastoma, osteosarcoma, rhabdomyosarcoma, rhabdomyocarcinoma and nephroblastoma (Wilms tumor).

[0321] Exemplary autoimmune disorders include adult Still's disease, agammaglobulinemia, alopecia areata, amyloidosis, ankylosing spondylitis, anti-GBM / anti-TBM nephritis, antiphospholipid syndrome, autoimmune angioedema, autoimmune dyskinesia, autoimmune encephalomyelitis, autoimmune hepatitis, autoimmune inner ear disease (AIED), autoimmune myocarditis, autoimmune oophoritis, autoimmune orchitis, autoimmune pancreatitis, autoimmune retinopathy, autoimmune urticaria, axonal and neuroneuropathy (AMAN), Valo disease, Behcet's disease, benign mucosal pemphigus, bullous pemphigoid, Castleman disease (CD), celiac disease, Chagas disease, chronic inflammatory demyelinating polyneuropathy (CIDP), chronic relapsing multifocal osteomyelitis (CRMO), Churg-Strauss syndrome (CSS) or eosinophilic granulomatosis (EGPA), and scarring. Pemphigus, Cogan syndrome, cold agglutinin disease, congenital heart block, Coxsackie's cardiomyopathy, CREST syndrome, Crohn's disease, superficial dermatitis, dermatomyositis, Devic's disease (neuromyelitis optica), discoid lupus, Dressler syndrome, endometriosis, eosinophilic esophagitis (EoE), eosinophilic fasciitis, erythema nodosum, essential mixed cryoglobulinemia, Evans syndrome, fibromyalgia, fibrous pneumonia, giant cell arthritis (temporal arteritis), giant cell cardiomyopathy, glomerulonephritis, Goodpasture syndrome, granulomatosis with polyangiitis, Graves' disease, Guillain-Barré syndrome, Hashimoto's thyroiditis, hemolytic anemia, Henoch-Schönlein purpura (HSP), gestational herpes or pemphigoid-like pregnancy (PG), hidradenitis suppurativa (HS) (inverse acne), Hypogammaglobulinemia, IgA nephropathy, IgG4-related sclerosing disease, Immune thrombocytopenic purpura (ITP), inclusion body myositis (IBM), interstitial cystitis (IC), juvenile arthritis, juvenile diabetes (Type 1 diabetes), pediatric myositis (JM), Kawasaki disease, Lambert-Eaton syndrome, leukoblastic vasculitis, lichen planus, lichen sclerosus, mungular conjunctivitis, linear IgA disease (LAD), lupus, chronic Lyme disease, Meniere's disease,Microscopic polyangiitis (MPA), Mixed Connective Tissue Disease (MCTD), Muren ulcer, Muha-Haberman disease, Multifocal Motor Neuropathy (MMN) or MMNCB, Multiple Sclerosis, Myasthenia Gravis, Myositis, Narcolepsy, Neonatal Lupus, Neuromyelitis Optic, Neutropenia, Pemphigus Ocularis, Optic Neuritis, Palindromic Rheumatism (PR), PANDAS, Subacute Cerebellar Degeneration (CD), Paroxysmal Nocturnal Hemoglobinuria (PNH), Facial Hemiplegia, Planopathitis (Peripheral Uveitis), Pasonage-Turner Syndrome, Pemphigus, Peripheral Neuropathy, Venous Encephalomyelitis, Pernicious Anemia (PA), POEMS Syndrome, Polyarteritis Nodularis Syndrome, Polymyositis Syndrome Types I, II, III, Polymyalgia Rheumatica, Polymyalgia Rheumatica, Polymyositis, Post-Myocardial Infarction Syndrome, Postpericardiotomy syndrome, primary biliary cirrhosis, primary sclerosing cholangitis, progesterone dermatitis, psoriasis, psoriatic arthritis, pure red blood cell aplasia (PRCA), pyoderma gangrene, Raynaud's phenomenon, reactive arthritis, reflex sympathetic dystrophy, recurrent polychondritis, restless legs syndrome (RLS), retroperitoneal fibrosis, rheumatic fever, rheumatoid arthritis, sarcoidosis, Schmidt syndrome, scleroderma, Sjögren's syndrome, sperm and testicular autoimmunity, rigid man syndrome (SPS), subacute bacterial endocarditis (SBE), Susac syndrome, sympathetic ophthalmitis (SO), Takaya's arteritis, temporal arteritis / giant cell arteritis, thrombocytopenic purpura (TTP), Tolosa-Hunt syndrome (THS), transverse myelitis, type 1 diabetes, Ulcerative colitis (UC), undifferentiated connective tissue disease (UCTD), uveitis, vasculitis, leukoplakia, and Vogt-Koyanagi-Harada disease, including but not limited to these.

[0322] Exemplary cardiovascular disorders include, but are not limited to, coronary artery disease (CAD), angina pectoris, myocardial infarction, stroke, heart attack, heart failure, hypertensive heart disease, thematic heart disease, cardiomyopathy, abnormal heart rhythm, congenital heart disease, valvular heart disease, carditis, aortic aneurysm, peripheral artery disease, thromboembolism, and venous thrombosis.

[0323] Exemplary coagulation disorders include, but are not limited to, hemophilia, von Willebrand disease, disseminated intravascular coagulation, liver disease, overdevelopment of circulating anticoagulants, vitamin K deficiency, platelet dysfunction, and other coagulation deficiencies.

[0324] Exemplary eye diseases include, but are not limited to, macular degeneration, protruding eye, cataract, CMV retinitis, diabetic macular edema, glaucoma, keratoconus, ocular hypertension, ocular migraine, retinoblastoma, subconjunctival hemorrhage, pterygium, keratitis, dry eye, and corneal abrasion.

[0325] Exemplary infectious diseases include acute flaccid myelitis (AFM), anaplasmosis, anthrax, babesiosis, botulism, brucellosis, campylobacteriosis, carbapenem-resistant infection (CRE / CRPA), chancroid, chikungunya virus infection (chikungunya), chlamydia, ciguatera (harmful avian broom (HAB)), Clostridium difficile infection, Clostridium perfringens (epsilon toxin), coccidioides fungal infection (valley fever), COVID-19 (coronavirus disease 2019), Creutzfeldt-transmissible Jacphon disease, cryptosporidiosis (Crypto), cytosporidiosis, dengue fever, 1, 2, 3, 4 (dengue fever), diphtheria, and lice. E. coli infection, Shiga toxin production (STEC), Eastern equine encephalitis (EEE), Ebola hemorrhagic fever (Ebola), herlichiosis, encephalitis, arbovirus or parainfectious, enterovirus infection, non-polio (non-polio enterovirus), enterovirus infection, D68 (EV-D68), Giardiasis (Giardia), Glanders, gonorrhea, inguinal granuloma, Haemophilus influenzae disease, Hepatitis B (Hib or H-flu), Hantavirus pulmonary syndrome (HPS), hemolytic uremic syndrome (HUS), Hepatitis A (Hepatitis A), Hepatitis B (Hepatitis B), Hepatitis C (Hepatitis C), Hepatitis D (Hepatitis A), Hepatitis E (Hepatitis E), herpes, shingles, VZV shingles (herpes zoster (Shingles), histoplasmosis infection (histoplasmosis), human immunodeficiency virus / AIDS (HIV / AIDS), human papillomavirus (HPV), influenza (flu), lead poisoning, Legionellosis (military disease), leprosy (Hansen's disease), leptospirosis, listeriosis (Listeria), Lyme disease, lymphogranuloma sexually transmitted infection (LGV), malaria, measles, melioidosis, meningitis, viral (meningitis, viral),Meningococcal disease, bacterial (meningitis, bacterial), Middle East Respiratory Syndrome Coronavirus (MERS-CoV), mumps, norovirus, paralytic shellfish poisoning (paralytic shellfish poisoning, ciguatera), polio (lice, head lice), pelvic inflammatory disease (PID), whooping cough, infectious disease; Lymph nodes, septic, pneumonia (infectious disease), pneumococcal disease (pneumonia), polio, Foucault, psytacosis (parrot's fever), non-coccal (crab; pubic lice infection), pustular rash disease (smallpox, monkeypox), microcephaly - fever, rabies, ricin poisoning, rickettsialism (Rocky Mountain spotted fever), congenital (including rubella), salmonellosis gastroenteritis (Salmonella), scabies infection (Scabies), scabies, septic shock (sepsis), severe acute respiratory syndrome (SARS), dysenteric gastroenteritis (Shigella), smallpox, staphylococcal infection, methicillin resistance (MRSA), staphylococcal food poisoning, enterotoxin B poisoning (staphylococcal food poisoning), staphylococcal infection, vancomycin intermediate (VISA), staphylococcal infection (VRSA), streptococcal disease, Group A (invasive) (Streptococcal disease (invasive)), Streptococcal disease, Group B (Streptococcus-B), Streptococcal Toxic Shock Syndrome, STSS, Toxic Shock (STSS, TSS), Syphilis (Primary, Secondary, Early Latent, Latent, Congenital), Tetanus Infection, Tetanus (Lock Jaw), Trichomoniasis (Trichomonas Infection), Trichomonas Infection (Trichomonas Infection), Tuberculosis (TB), Tuberculosis (Latent) (LTBI), Tularemia (Rabbit) Fever), Typhoid Fever (Group D), Typhoid Fever, Vaginitis, Bacterial (Yeast Infection), E-cigarette-Related Lung Injury (E-cigarette-Related Lung Injury), Chickenpox (Chickenpox), Vibrio Cholera (Cholera), Vibriosis (Vibrio), Viral Hemorrhagic Fever (Ebola, Lassa, Marburg), West Nile Virus, Yellow Fever, Yersenia (Yesinia), Zika Virus Infection (Zika) Includes, but is not limited to.

[0326] Examples

[0327] Example 1 : RNA construct design and preparation

[0328] The screening construct was designed to allow the integration of one or more diverse internal target RNA motifs. Two motifs are present in the said construct: of the dengue fever virus. 27 TPP Riboswitch domain and pseudoknot from 5'-UTR 26 The design for the complete construct sequence, including the structure cassette, RNA barcode helix, and two test RNA structures (separated by a 6-nucleotide linker), was evaluated using the RNA structures. 39 To reduce the likelihood of the two test structures interacting, a few sequences were modified to maintain the intrinsic folding and prevent misfolded structures predicted by the RNA structure (Fig. 7). The structure of the final construct was verified by SHAPE-MaP.

[0329] RNA barcodes were designed to fold into self-contained hairpins (Fig. 7). All possible permutations of RNA barcodes were calculated and folded in the context of the entire construct sequence, and any barcodes that were likely to interact with other parts of the RNA construct were removed from the set. The barcode constructs were detected by SHAPE-MaP using the "ligand absence" protocol and folded using RNA constructs with SHAPE reactive constraints to verify whether the barcode helix had folded into the intended self-contained hairpin.

[0330] RNA production

[0331] DNA templates for in vitro transcription (integrated DNA technology) encoded target construct sequences (including dengue-like knot sequences, single-strand linkers, and TPP riboswitch sequences) and flanking structure cassettes. 25 : 5'- GTGGG CACTT CGGTG TCCAC ACGCG AAGGA AACCG CGTGT CAACT GTGCA ACAGC TGACA AAGAG ATTCC TAAAA CTCAG TACTC GGGGT GCCCT TCTGC GTGAA GGCTG AGAAA TACCC GTATC ACCTG ATCTG GATAA TGCCA GCGTA GGGAA GTGCT GGATC CGGTT CGCCG GATCA AT CGG GCTTC GGTCC GGTTC -3' ( Sequence No. 1 ). Primer binding sites are underlined. RNA barcodes were individually added to each of the 96 constructs in individual PCR reactions using orientation PCR primers containing unique RNA barcodes and T7 promoter sequences. The sample orientation primer sequences with bolded barcode nucleotides and underlined primer binding sites are as follows: 5'- GAAAT TACGA CTCAC TATAG GTCGC GAGTA ATCGC GACCG GC G CT AGAGA T AG T GC C GTG GGCAC TTCGG TGTC -3' ( Sequence No. 2 ).

[0332] DNA was amplified by PCR using 200 μM dNTP mix (New England Biolabs), 500 nM forward-orientation primers, 500 nM reverse-orientation primers, 1 ng DNA template, 20% (v / v) Q5 reaction buffer, and 0.02 U / μL Q5 hot-start high-fidelity polymerase to generate a template for in vitro transcription. DNA was purified (PureLink Pro 96 PCR Purification Kit; Invitrogen) and quantified on a Tecan Infinite M1000 Pro microplate reader (Quant-iT dsDNA High-Sensitivity Assay Kit; Invitrogen).

[0333] In vitro transcription was performed in a 96-well plate format, with each well containing a total reaction volume of 100 μL. Each well contained 5 mM NTP (New England Biolabs), 0.02 U / μL inorganic pyrophosphatase (yeast, New England Biolabs), and 0.05 mg / mL T7 polymerase in 25 mM MgCl2, 40 mM Tris, pH 8.0, 2.5 mM spermidine, 0.01% Triton, 10 mM DTT, and 200–800 nM uniquely barcoded DNA templates (generated by PCR). The reaction mixture was incubated at 37°C for 4 hours and subsequently treated with TurboDNase (RNase-free, Invitrogen) at a final concentration of 0.04 U / μL; Incubation was performed at 37°C for 30 minutes, followed by the addition of a second DNase to a total final concentration of 0.08 U / μL and incubation at 37°C for an additional 30 minutes. The enzymatic reaction was stopped by adding EDTA to a final concentration of 50 mM, and the sample was placed on ice. RNA was purified into a 96-well format (Agencourt RNAclean XP magnetic beads; Beckman Coulter) and resuspended in 10 mM Tris pH 8.0 and 1 mM EDTA. RNA concentration was quantified on a Tecan Infinite M1000 Pro microplate reader (Quant-iT RNA Broad-spectrum assay kit; Invitrogen), and RNA from each well was individually diluted to 1 pmol / μL. The RNA was stored at -80°C.

[0334] Example 2: Chemical modification and screening of small molecule fragments

[0335] The fragments were obtained as a fragment screening library from Maybridge, comprising a subset of the Ro3 diversity fragment library and 1,500 compounds dissolved in DMSO at 50 mM. Most of these compounds follow the "Rule of 3" for fragment compounds; they have a molecular mass < 300 Da, contain ≤ 3 hydrogen bond donors and ≤ 3 hydrogen bond acceptors, and ClogP ≤ 3.0. With the exception of the compounds listed in Example 5, all compounds used in the ITC were purchased from Millipore-Sigma and used without further purification. Screening experiments were performed in 25 μL in a 96-well plate format on a Tecan Freedom Evo-150 liquid handler equipped with an 8-channel air displacement pipetting arm, disposable filter tips, a robotic manipulator arm, and an EchoTherm RIC20 remote-controlled heating / cooling drying bath (Torrey Pines Scientific). The liquid handler program used for screening is available upon request.

[0336] For the first fragment-ligand screening, 5 pmol of RNA per well was diluted to 19.6 μL in RNase-free water on a 4°C cooling block. The plates were heated at 95°C for 2 minutes and immediately quenched at 4°C for 5 minutes. 19.6 μL of 2× folding buffer (final concentration 50 mM HEPES pH 8.0, 200 mM potassium acetate, and 10 mM MgCl2) was added to each well, and the plates were incubated at 37°C for 30 minutes. For the second fragment-ligand screening, 24.3 μL of folded RNA per well was diluted to 2.7 μL of primary binding fragment in DMSO at a final concentration of 10x K dThe fragments were added to form the fragments, and the samples were incubated at 37°C for 10 minutes. To combine the target RNA with the fragments, 24.3 μL of RNA solution or RNA + primary binding fragment was added to wells containing 2.7 μL of 10x screening fragment (in DMSO to yield a final fragment concentration of 1 mM). The solutions were thoroughly mixed by pipetting and incubated at 37°C for 10 minutes. For SHAPE detection, 22.5 μL of RNA-fragment solution from each well of the screening plate was added to 2.5 μL of 10x SHAPE reagent in DMSO over a 37°C heating block, and a uniform distribution of the RNA-based SHAPE reagent was achieved by rapidly mixing by pipetting. After an appropriate reaction time, the samples were placed on ice. For the first fragment screening, 1-methyl-7-nitrosatosan anhydride (1M7) was used as the SHAPE reagent at a final concentration of 10 mM with a reaction time of 5 minutes. For the second fragment screening, 5-nitrosatosan anhydride (5NIA) 40 It was used as a SHAPE reagent at a final concentration of 25 mM with a reaction time of 15 minutes. Excess fragments, solvent, and hydrolyzed SHAPE reagent were removed using an AutoScreen-A 96-well plate (GE Healthcare Life Sciences), and 5 μL of modified RNA from each well of the 96-well plate was pooled as a single sample per plate for sequencing of the library preparations.

[0337] Each screening consisted of 19 fragment test plates, two plates containing the distribution of positive (fragment 2, final concentration 1 mM) and negative (solvent, DMSO) controls, and one negative SHAPE control plate treated with solvent (DMSO) instead of SHAPE reagent. For the hit confirmation experiment, the well location of each hit fragment was varied to control for well location and RNA barcode effects. Plate maps for both primary and secondary screenings are also available.

[0338] Once the screening of the test fragments is completed, statistical tests were performed to identify differences in the strain of a given nucleotide. Specifically, the screening analysis requires a statistical comparison of the strain of a given nucleotide in the presence of the fragment compared to its absence. For each nucleotide, the number of strains in the given reaction is a Poisson process with known variance; therefore, the statistical significance of the difference in strain observed between the two samples can be confirmed by performing two Poisson coefficient comparison tests. 31 That is, counting m1 variations of the tested nucleotide during the n1 reading of Sample 1, and m during the n2 readings of Sample 2 2 When counting n deformations, the tested null hypothesis is that the proportion of deformation in Sample 1 among all counted deformations (m1 + m2) is p1 = n1 / (n1 + n2). The Z-test for the above hypothesis is as follows:

[0339]

[0340] If the Z value exceeds a specified significance threshold, the tested nucleotide is considered to be statistically significantly affected by the presence of the test fragment.

[0341] Next, for each fragment, a Z-test must be performed on multiple nucleotides containing the RNA sequence, increasing the false-positive probability. While the number of false-positive assignments of SHAPE reactivity per nucleotide can be minimized by raising the Z significance threshold, this approach reduces the sensitivity of the screening (meaning it reduces the ability to detect weaker binding ligands). To reduce the number of Z-tests performed, these tests were applied only to nucleotides of interest in the RNA screening constructs, rather than to all nucleotides. For the dengue fever motif in RNA, the region of interest was at positions 59–110; for the TPP motif, the region of interest was at positions 100–199. The number of Z-tests was further reduced by removing nucleotides with low strain from both samples. A threshold for considering nucleotides with low strain was set to 25% of the plate average strain, which was calculated for all nucleotides in all 96 wells of a given plate. A Z-test was performed only on nucleotides that caused the strain in at least one of the two compared samples to exceed the 25% threshold.

[0342] Ideally, the only difference between the conditions in two comparison samples is that a fragment is present in one sample but not in the other. By using a test of negative control samples against each other, the prevalence of uncontrolled factors that can introduce variability in nucleotide strains among the samples can be measured. For example, if the Z-significance threshold is set to 2.7, a Z-test applied to a pair of negative control (fragment-free) samples should theoretically identify differentially reactive nucleotides with a probability of P = 0.0035 in the absence of any of the aforementioned factors. However, when a Z-test was applied to pairs of negative control samples randomly selected from the 587 negative control samples tested in the primary screening, the actual probability was 90-fold higher at P = 0.32. Therefore, there is statistically significant variability in SHAPE reactivity in individual nucleotides without fragments.

[0343] Most replicates essentially shared the same profile, but there were a significant number of replicates with different profiles; some coefficients of determination were as low as 0.85. The application of the Z-test to different negative control samples often resulted in the generated nucleotides being misclassified as differentially reactive. To avoid this result, each sample was compared to the five most correlated negative control samples. The Z-test applied to these selected pairs of negative controls, which had a Z significance threshold of 2.7, resulted in the identification of differentially reactive nucleotides with a probability of P = 0.067.

[0344] The above probability is approximately 20 times higher than the theoretical P = 0.0035, indicating variability in sample processing. Some of this variability is extended evenly across the reactivity of all nucleotides of all RNA in the sample. This variability can be eliminated by scaling down the total reactivity of the more reactive sample to match the total reactivity of the less reactive sample. This scaling was performed by (i) calculating the ratio of the strain in the more reactive sample to the strain in the less reactive sample for each nucleotide in the RNA sequence, and (ii) dividing the strain of all nucleotides in the more reactive sample by the median of the ratio obtained in step (i). This scaling of the pairs with maximized correlation in the negative control well reduced the probability of finding a nucleotide hit to P = 0.030, which is 9 times higher than the theoretical probability. Therefore, false-positive identification of fragments was performed as is actually done in all fast processing screening assays, and actual fragment hits from non-ligand variants were distinguished by replicate shape confirmation and direct ligand binding measurements using ITC.

[0345] Since effective ligands are expected to influence the modification rates of multiple nucleotides in target RNA, fragments were recognized as hits only when the number of nucleotides with reactivity different from the reactive control exceeded a defined threshold set to 2. Second, when looking for the relatively potent effect of fragments on RNA, small relative differences in nucleotide reactivity were excluded from the total number of differentially reactive nucleotides, even if they were statistically significant. In practice, the minimum allowable difference was set at 20% of the mean.

[0346]

[0347] Here, r 1 and r2 represents the nucleotide strain in the two samples. Third, the specified samples were tested against the five negative control samples showing the highest correlation. All five tests were required to identify the test samples that were modified compared to the negative control samples.

[0348] Finally, the sensitivity and specificity of the screening were controlled by the selection of Z significance thresholds. Evaluation of samples containing fragments and all negative control samples was performed at multiple Z significance threshold settings. For each of these settings, the false positive fraction (FPF) was calculated as the fraction of negative control samples identified as altered, and the ligand fraction (LF) was evaluated by subtracting the FPF from the fraction of altered samples containing fragments. The balance between LF and FPF was quantified by the LF / FPF ratio. The best balance for TPP riboswitch RNA (1.3) was achieved with Z significance thresholds in the range of 2.5 to 2.7, where 0.022 > FPF > 0.014. For dengue-like knots, the best balance (LF / FPF 4) was achieved with a Z significance threshold in the range between 2.5 and 2.65 where 0.007 > FPF > 0.005.

[0349] Example 3: Library Preparation and Sequencing

[0350] Reverse transcription was performed on 100 μL of pooled modified RNA. 6 μL of reverse transcription primers was added to 71 μL of pooled RNA to achieve a final primer concentration of 150 nM, and the sample was incubated at 65°C for 5 minutes and then placed on ice. To the solution, 6 μL of 10x first-strand buffer (500 mM Tris pH 8.0, 750 mM KCl), 4 μL of 0.4 M DTT, 8 μL of dNTP mix (10 mM each), and 15 μL of 500 mM MnCl2 were added, and the solution was incubated at 42°C for 2 minutes before adding 8 μL of SuperScript II reverse transcriptase (Invitrogen). The reaction mixture was incubated at 42°C for 3 hours, followed by heat inactivation at 70°C for 10 minutes before being placed on ice. The generated cDNA product was purified (Agencourt RNAClean magnetic beads; Beckman Coulter), eluted in RNase-free water, and stored at -20°C. The sequence of the reverse transcription primer was 5'-CGGGC TTCGG TCCGG TTC-3' (Sequence No. 3).

[0351] A DNA library for sequencing was prepared using a two-step PCR reaction to amplify DNA and add the necessary TruSeq adapter. 24DNA was amplified by PCR using 200 μM dNTP mix (New England Biolabs), 500 nM forward-oriented primers, 500 nM reverse-oriented primers, 1 ng DNA or double-stranded DNA template, 20% (v / v) Q5 reaction buffer, and 0.02 U / μL Q5 hot-start high-fidelity polymerase. Excess unintegrated dNTPs and primers were removed by affinity purification (Agencourt AmpureXP magnetic beads, Beckman Coulter, 0.7:1 sample-to-bead ratio). The DNA library was quantified on a Qubit fluorescence analyzer (Invitrogen) (Qubit dsDNA High Sensitivity Assay Kit; Invitrogen), quality checked (Bioanalyzer 2100 on-chip electrophoresis instrument, Agilent), and sequenced on an Illumina NextSeq 550 high-speed sequencer.

[0352] The amplicon-specific positive orientation primers for the SHAPE-MaP library are 5'CCCTA CACGA CGCTC TTCCG ATCTN NNNN G GCCTT CGGGC CAAGG A -3' ( Sequence No. 4 ) was. The amplicon-specific reverse orientation primers for the SHAPE-MaP library fabrication were 5'GACTG GAGTT CAGAC GTGTG CTCTT CCGAT CTNNN NNTT G AACCG GACCG AAGCC CGATT T -3' ( Sequence No. 5 It was. Sequences that overlap with RNA screening constructs are underlined.

[0353] Example 4: Isothermal Titration Calorimeter

[0354] ITC experiments were performed using a Microcal PEAQ-ITC automated instrument (Malvern Analytical) under RNase-free conditions. 41. In vitro transcribed RNA was concentrated by centrifugation (Amicon Ultra centrifuge filter, 10K MWCO, Millipore-Sigma) with 100 mM CHES, pH 8.0, 200 mM potassium acetate, and 3 mM MgC l2 It was exchanged with a folding buffer containing [the component]. The ligand was dissolved in the same buffer (to minimize mixing heat when adding the ligand to the RNA) at a concentration 10–20 times that of the target experimental RNA concentration. RNA concentration was quantified (using a Nanodrop UV-VIS spectrometer; ThermoFisher Scientific) and the expected K in the buffer was measured. d The RNA was diluted 1 to 10 times, and the diluted RNA was re-quantified to determine the final experimental RNA concentration. The RNA diluted in folding buffer was heated at 65°C for 5 minutes, placed on ice for 5 minutes, and allowed to fold at 37°C for 15 minutes. If necessary, the primary binding ligand (e.g., 2) was pre-bound to the RNA by adding 0.1 volume at 10 times the target final concentration of the bound ligand, followed by incubation at room temperature for 10 minutes.

[0355] Each ITC experiment included two runs: the ligand titrated with RNA (experimental track) and the same ligand titrated with buffer (control track). The ITC experiments were performed using the following parameters: cell temperature of 25°C, reference power of 8 μCal / sec, stirring speed of 750 RPM, high feedback mode, and 19 injections of 2 μL following an initial injection of 0.2 μL. Each injection took 4 seconds to complete, with an interval of 180 seconds between injections.

[0356] ITC data were analyzed using MicroCal PEAQ-ITC analysis software (Malvern Analytical). First, baselines for each injection peak were manually adjusted to correct inaccurately selected injection endpoints. Second, control traces were subtracted from experimental traces by point-to-point deduction. Third, least squares regression lines were fitted to the data using the Levenberg-Marquardt algorithm. N was manually set to 1.0 to enable the fitting of low c-value curves for weakly binding ligands (>500 μM).

[0357] Example 5: Chemical synthesis of test compounds 35, 36, 37, 38, 39 and 40.

[0358]

[0359] Compound 35 : 3-C linked hydroxyl acid 35 was prepared from carboxylic acid S19 via a mixed anhydride intermediate by reacting with aqueous hydroxylamine. Quinoxalin-6-amine was approached to acid S19 by treating it with the closed-ring anhydride dihydrofuran-2,5-dione.

[0360]

[0361] Compound 36 : 2-C linked analog 36 was obtained from the corresponding ester S20 by reacting it with hydroxylamine formed in situ. Ester S20 was prepared by Michael addition of quinoxalin-6-amine with ethyl acrylate.

[0362]

[0363] Compound 37 The Buchwald-Hartwig reaction was used for the synthesis of intermediates S21 and S22. The removal of the protecting group (Boc) was achieved using HCl in ether and then further treated with Na2CO3 to obtain 37.

[0364]

[0365] Compound 38 : Imine formation, and subsequent sodium borohydride reduction of quinoxalin-6-carbaldehyde and diamine produces 38.

[0366]

[0367] Compound 39 : Immigration formation, and S N Intermediate 24 was obtained by subsequent sodium borohydride reduction using quinoxaline-6-ylmethanamine hydrochloride and aldehyde S23 prepared via an Ar reaction, and then 39 was obtained by deprotecting with HCl (Boc).

[0368]

[0369] Compound 40 : Less bound analog 40 was prepared by reacting twice with 3,5-dibromopyridine in the Buchwald-Hartwigg cycle and then deprotecting with HCl (Boc).

[0370] Example 6: X-ray Crystallography

[0371] To evaluate whether the structural variant of 2 is a good binding candidate for the TPP riboswitch, compound 17 was investigated in X-ray crystallography studies. TPP riboswitch RNA was prepared by in vitro transcription as described. 27TPP riboswitch RNA (0.2 mM) and RNA-17 (2 mM) were heated at 60°C for 3 minutes in a buffer containing 50 mM potassium acetate (pH 6.8) and 5 mM MgCl2, rapidly cooled on crushed ice, and incubated at 4°C for 30 minutes prior to crystallization. For crystallization, 1.0 μL of RNA-17 complex was mixed with 1.0 μL of a storage solution containing 0.1 M sodium acetate (pH 4.8), 0.35 M ammonium acetate, and 28% (v / v) PEG4000. Crystallization was performed at 291K by applying dropwise vapor diffusion over a period of 2 weeks. The crystals were cryoprotected in a mother liquor supplemented with 15% glycerol prior to rapid freezing in liquid nitrogen. Data were collected at a wavelength of 0.9202 Å at the 17-ID-2 (FMX) beamline of NSLS-II (Brookhaven National Laboratory). Data were processed using HKL200043. The structure was elucidated through molecular replacement using Phenix44 and 2GDI riboswitch RNA structures. 27 The structure was modified from Phenix. Organic ligands, water molecules, and ions were added in the later modification stage based on the Fo-Fc and 2Fo-Fc electron density maps.

[0372] The results showed that compound 17 binds to the TPP riboswitch in a manner similar to the thiamine moiety of the TPP ligand stacking between G42 and A43 at the J3 / 2 junction (Fig. 3). 27,28. 17 forms three hydrogen bonds with RNA: one bound to the ribose of G40 and the Watson-Crick face, and one bound to the ribose of G19. Compared to RNA, there are significant changes in the local RNA structure during the complex with the intrinsic TPP ligand. In the 17-bound structure, G72 flips to the binding site containing the pyrophosphate moiety of the TPP ligand. This binding mode is consistent with previous work that visualized the orientation of G72 flipped to the fragment bound to the thiamine subsite of the riboswitch binding pocket. 17,34 Consistent with SAR analysis, the orientation of the C-6 substituent appears to be relatively undisturbed by interactions with RNA, suggesting that this vector will be a good candidate for fragment refinement.

[0373] References

[0374]

[0375]

[0376]

[0377]

Claims

Claim 1 Compound having the structure of Chemical Formula III or a pharmaceutically acceptable salt thereof: Chemical Formula III In the above equation, L is or and, q and r are independently selected from integers 0, 1, 2, and 3; z is selected from integers 1, 2, and 3; and A is , , and Selected from, X5 and X6 are independently selected from CR3 and N; X4 and X7 are CR3, and R3 is -H; m is 1 or 2; W is -O or -NR4, and R4 is (C1-C6)alkyl or -H. Claim 2 A compound according to claim 1, wherein q and r are independently selected from integers 0, 1, and 2. Claim 3 In claim 1 or 2, L Phosphorus, compound. Claim 4 A compound according to claim 3, wherein q and r are 0 or 1. Claim 5 A compound according to claim 4, wherein q and r are 1. Claim 6 A compound according to claim 4, wherein q is 1 and r is 0. Claim 7 A compound according to claim 1, wherein m is 1. Claim 8 A compound according to claim 7, wherein W is -NH or -O. Claim 9 A compound of claim 8, wherein W is -NH. Claim 10 In claim 9, A Phosphorus, compound. Claim 11 In claim 9, A Phosphorus, compound. Claim 12 A compound according to claim 3, wherein q is 0 and r is 2. Claim 13 In claim 1, z is 2 and L is Phosphorus, compound. Claim 14 In claim 12, A Phosphorus, compound. Claim 15 In claim 13, A Phosphorus, compound. Claim 16 Claim 1, wherein the compound has the following structure, a compound or a pharmaceutically acceptable salt thereof: , , , , or . Claim 17 delete Claim 18 delete Claim 19 delete Claim 20 delete Claim 21 delete Claim 22 delete Claim 23 delete Claim 24 delete Claim 25 delete Claim 26 delete Claim 27 delete Claim 28 delete Claim 29 delete Claim 30 delete Claim 31 delete Claim 32 delete Claim 33 delete Claim 34 delete Claim 35 delete Claim 36 delete