RNA targeting ligands, compositions thereof, and methods of making and using same
By employing a fragment-based screening strategy and SHAPE-MaP RNA structure detection, high nanomolar affinity RNA-targeting ligands were identified and designed, solving the problem of rapid identification of small molecule ligands binding to RNA molecules in existing technologies, and achieving efficient regulation of RNA targets such as TPP riboswitches.
Patent Information
- Application Number
- CN202511000157.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-29
- Filing Date
- 2020-08-05
- Publication Date
- 2025-12-02
AI Technical Summary
Existing technologies struggle to rapidly and efficiently identify small molecule ligands that bind with high affinity to RNA molecules, especially non-coding RNA molecules, resulting in a lack of effective means for regulating cellular states and treating diseases.
We employed a fragment-based screening strategy, using selective 2'-hydroxyacylated (SHAPE) and SHAPE-mutant profiling (MaP) RNA structure detection to identify fragments that bind to TPP riboswitches. We then designed high nanomolar affinity linker ligands through structure-activity relationship studies, and developed high-quality RNA-targeting ligands by leveraging synergistic and multi-site binding.
This technology enables the identification and design of fragment ligands that bind to TPP riboswitches with high affinity, effectively modulating RNA function and possessing broad applicability. It is suitable for the development of highly efficient targeting ligands for various RNA targets.
Smart Images

Figure CN121045142A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese patent application filed on August 5, 2020, with application number 202080069480.0 and entitled "RNA Targeting Ligand, Composition Thereof, and Method Thereof for Preparation and Use". Technical Field
[0002] This disclosure relates to compounds that bind to target RNA molecules such as TPP riboswitches, compositions comprising said compounds, and methods of preparation and use thereof. Compared to compounds that bind to only a single RNA binding site, said compounds contain two structurally distinct fragments that allow binding to the target RNA at two different binding sites, thereby producing binding ligands with higher affinity. Background Technology
[0003] The vast majority of small molecule ligands are developed primarily to manipulate biological systems by targeting proteins. Proteins possess highly complex three-dimensional structures, which are crucial for their proper functioning and contain the gaps and notches that small molecule ligands can bind to. 1,2 The transcriptome—the collection of all RNA molecules produced in an organism—also contains promising targets for studying and manipulating biological systems. For example, the RNA transcriptome plays an important role not only in mammalian systems but also in both bacteria and viruses, and therefore represents targets for small molecule regulation of gene expression.
[0004] It is worth noting that RNA can take on a three-dimensional structure with a complexity comparable to that of proteins. 3 This is a key feature required for developing highly selective ligands. 4 Furthermore, RNA plays a universal role in regulating the behavior of biological systems. 5 Initially viewed merely as a carrier of genetic information for protein coding and guiding protein biosynthesis, the modern understanding of RNA has evolved to encompass a wider range of roles. It is now known that various RNA molecules play a broad and profound role in regulating gene expression and other biological processes through various mechanisms. Even a large number of newly discovered non-coding RNAs have been found to be associated with diseases such as cancer and non-oncogenic diseases. Therefore, in addition to encoding pathogenic proteins, RNA contributes to the understanding of disease states and provides a wealth of previously unrecognized therapeutic targets.
[0005] However, although small molecule ligands have been shown to bind to mRNA and have the potential to upregulate or downregulate translation efficiency, thereby modulating protein expression in cells. 6,7 However, identifying small RNA ligands still involves challenges not encountered when targeting proteins. 4,11,12This also includes developing small molecules targeting non-coding RNAs, which represent a rich array of targets. 8 - 10 Unfortunately, despite the development of various techniques for analyzing RNA structure and the discovery of new functions, the ability to efficiently and rapidly identify or design inhibitors that bind to RNA and perturb its function remains far behind. Therefore, there is a significant need in this field to develop new methods and techniques that allow for the rapid and efficient identification of small molecule ligands targeting RNA molecules. Summary of the Invention
[0006] As mentioned above, the transcriptome represents a set of attractive but underutilized targets for small molecular ligands. Small molecular ligands (and ultimately drugs) targeting messenger RNA and non-coding RNA have the potential to modulate cellular states and diseases. In this current disclosure, a fragment-based screening strategy using selective 2'-hydroxyacylated (SHAPE) and SHAPE-mutant profiling (MaP) RNA structure probing via primer extension analysis was used to discover small molecular fragments that bind to target RNA structures. Specifically, fragments binding to TPP riboswitches with millimolar to micromolar affinities and co-binding fragment pairs were identified. Structure-activity relationship (SAR) studies were conducted to obtain information for efficiently designing linker fragment ligands that bind to TPP riboswitches with high nanomolar affinities. The principles of this current disclosure are not intended to limit us to TPP riboswitches, but can be broadly applied to other target RNA structures, thereby leveraging co-binding and multi-site binding to develop high-quality ligands for a wide range of RNA targets.
[0007] Thus, one aspect of the currently disclosed subject matter is compounds having the structure of formula (I):
[0008]
[0009] in
[0010] X1, X2, and X3 are independently selected from CR1, CHR1, N, NH, O, and S in each instance, wherein adjacent X1, X2, and X3 are not simultaneously selected as O or S;
[0011] Dashed lines represent optional double bonds;
[0012] Y1, Y2, and Y3 are independently selected from CR2 and N in each instance;
[0013] n is 1 or 2, where when n is 1, only one of the dashed lines is a double bond;
[0014] L is selected from
[0015]
[0016]
[0017] Where p, q, r, and v are independently selected from the integers 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10, and z is selected from the integers 1, 2, 3, 4, and 5; and
[0018] A is selected from
[0019]
[0020] X4, X5, X6, and X7 are independently selected from CR3 and N;
[0021] R1, R2 and R3 are independently selected from -H, -Cl, -Br, -I, -F, -CF3, -OH, -CN, -NO2, -NH2, -NH(C1-C6 alkyl), -N(C1-C6 alkyl)2, -COOH, -COO(C1-C6 alkyl), -CO(C1-C6 alkyl), -O(C1-C6 alkyl), -OCO(C1-C6 alkyl), -NCO(C1-C6 alkyl), -CONH(C1-C6 alkyl) and substituted or unsubstituted C1-C6 alkyl groups;
[0022] m is 1 or 2; and
[0023] W is -O or -NR4, where R4 is selected from -H, -CO (C1-C6 alkyl), substituted or unsubstituted C1-C6 alkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted cycloalkyl, -CO (aryl), -CO (heteroaryl), and -CO (cycloalkyl).
[0024] The condition is that at least two of X1, X2, X3, X4, X5, X6, and X7 are N;
[0025] Or its pharmaceutically acceptable salt.
[0026] Another aspect of the currently disclosed subject matter includes compounds that bind to regions of RNA molecules, as described herein.
[0027] Another aspect of the currently disclosed subject matter includes a composition comprising a therapeutically effective amount of the compound described herein in a pharmaceutically acceptable carrier, diluent, or excipient.
[0028] Another aspect of the currently disclosed subject matter includes a method for treating a disease or condition associated with dysfunction of RNA expression, the method comprising administering to a subject in need a therapeutically effective dose of a composition of the compounds described herein.
[0029] Another aspect of the subject matter disclosed herein includes methods for preparing the compounds described herein.
[0030] Another aspect of the currently publicized topic will be presented below. Attached Figure Description
[0031] Figure 1 A scheme for RNA screening constructs and fragment screening workflows is illustrated. RNA motifs 1 and 2; a barcode helix; and a structure box helix are shown. SHAPE is used to probe RNA in the presence or absence of small fragments, and chemical modifications corresponding to ligand-dependent structural information are read out by multiplexing MaP sequencing.
[0032] Figure 2 Representative mutation rates for fragment hits and misses are shown. Normalized mutation rates for fragment-exposed samples are labeled as +ligand, +2, or +4 and compared to ligandless traces labeled as ligandless. Statistically significant changes in mutation rates are indicated by triangles (see [link to SHAPE confirmation data]). Figure 8 (Top) Comparison of mutation rates of representative fragments not bound to the test construct. (Middle) Fragment hits to the TPP riboswitch region of RNA. (Bottom) Non-specific hits inducing reactive changes throughout the test construct. Motifs 1 and 2 are shown below the SHAPE spectra.
[0033] Figure 3A and 3B It shows the result of ( Figure 3A ) 17 pairs of fragments ( Figure 3B Natural TPP ligands (2HOJ) 28 Comparison of the structures of the TPP riboswitches. RNA structures are shown with similar orientations in each image. Hydrogen bonds between the ligand and RNA are shown as dashed lines.
[0034] Figure 4A and 4B The thermodynamic cycling and stepwise ligand binding affinity of fragments 2 and 31 are shown. Figure 4A A summary of the binding of fragments of compound 2 (dark gray, K1) and compound 31 (light gray, K2) is shown. D The value is determined by ITC. Figure 4B ITC data are shown, illustrating single-compound binding and co-binding by fragments 2 and 31. The connection of the two fragments exhibits an additive effect of binding energies, resulting in the submicromolar ligand compound 37 (K). L The ITC trace is shown, with the background trace (ligand titrated into buffer) shown in light gray and the experimental trace shown in dark gray. Curve fits and 95% confidence intervals are shown in gray shading.
[0035] Figure 5 The covalent linkages of fragments 17 and 31 are shown as varying with adapter type and length, terminal chemotype, and terminal orientation. Modifications increasing RNA binding affinity are shown in compounds 36 and 37 (light gray); negative modifications are shown in compounds 35, 39, and 40 (light gray), and a neutral modification is shown in compound 38 (light gray). Dissociation constants were determined by ITC.
[0036] Figure 6 A comparison of fragment-linker-fragment ligands developed using a fragment-based approach is shown, with the ligands ordered according to their linkage coefficients (E). Values are shown on a logarithmic axis. Co-linking corresponds to lower E values (top of the vertical axis). Fragment 37 exhibits an E value of 2.5 and an E value of 0.34. The dissociation constants of individual fragments (left, middle) and linked ligands (right) are indicated below the component fragments; E values (top) and ligand efficiencies (bottom) are shown. Covalent linkages introduced between fragments are highlighted in light gray. The structures of the component fragments are detailed in Table 7.
[0037] Figure 7A and 7B The design of the filter builder is shown. Figure 7A An RNA sequence (SEQ ID NO:6) containing the following components is shown: GGUCGCGAGUAAUCGCGACC (SEQ ID NO:7) is a cassette; G CU G CA AGAGAU UG U AG C (SEQ ID NO:8) is an RNA barcode (barcode NT with underline); GUGGGCACUUCGGUGUCCAC (SEQ ID NO:9) is a cassette; ACGCGAAGGAAACCGCGUGUCAACUGUGCAACAGCUGACAAAGAGAUUCCU (SEQ ID NO:10) is a DENV pseudoknot (mutant bold); AAAACU is an adapter; CAGUACUCGGGGUGCCCUUCUGCGUGAAGGCUGAGAAAUACCCGUAUCACCUGAUCUGGAUAAUGCCAGCGUAGGGAAGUGCUG (SEQ ID NO:11) is a TPP riboswitch (mutant bold); and GAUCCGGUUCGCCGGAUCAAUCGGGCUUCGGUCCGGUUC (SEQ ID NO:12) is a cassette. Figure 7B The secondary structure of the RNA sequence barcode is shown against the background of its self-folding hairpin structure.
[0038] Figure 8 The SHAPE spectra of missed, hit, and non-specific hit fragments are shown. The mutation rate traces corresponding to fragment exposure and the ligand-free control traces are shown as pure gray shading and black outlines, respectively. Nucleotides identified as statistically significantly different between the fragment sample and the unfragmented sample are marked with triangles. The mutation rate traces of the same fragment are shown in... Figure 2 The diagram is shown schematically. Detailed Implementation
[0039] The subject matter of this disclosure will now be described more fully below. However, those skilled in the art will conceive of many modifications and other embodiments of the subject matter to which this disclosure pertains, taking advantage of the teachings presented in the foregoing description. Therefore, it should be understood that the subject matter of this disclosure is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. In other words, the subject matter described herein encompasses all alternatives, modifications, and equivalents. If one or more of the cited documents, patents, and similar materials differ from or contradict this application, including but not limited to defined terms, usages of terms, described techniques, etc., this application shall prevail. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. All publications, patent applications, patents, and other references mentioned herein are incorporated herein by reference in their entirety.
[0040] definition
[0041] As used herein, the term "alkyl" refers to a saturated hydrocarbon group containing 1 to 8, 1 to 6, 1 to 4, or 5 to 8 carbons. In some embodiments, the saturated group contains more than 8 carbons. Alkyl groups are structurally similar to noncycloalkane compounds modified by removing a hydrogen atom from a noncycloalkane and replacing it with a non-hydrogen group or radical. Alkyl radicals can be branched or unbranched. Lower alkyl radicals have 1 to 4 carbon atoms. Higher alkyl radicals have 5 to 8 carbon atoms. Examples of alkyl radicals, lower alkyl radicals, and higher alkyl radicals include, but are not limited to, radicals such as methyl, ethyl, n-propyl, isopropyl, n-butyl, sec-butyl, tert-butyl, pentyl, tert-pentyl, n-pentyl, n-hexyl, and isooctyl.
[0042] As used herein, the symbols “(CO)” and “C(O)” are used to indicate carbonyl moieties. Examples of suitable carbonyl moieties include, but are not limited to, ketone and aldehyde moieties.
[0043] The term "cycloalkyl" refers to a hydrocarbon having 3-8, 3-7, 3-6, 3-5, or 3-4 members, and can be monocyclic or bicyclic. The ring may be saturated or may have a certain degree of unsaturation. The cycloalkyl group may optionally be substituted by one or more substituents. In one embodiment, 0, 1, 2, 3, or 4 atoms of each ring of the cycloalkyl group may be substituted by a substituent. Representative examples of cycloalkyl groups include cyclopropyl, cyclopentyl, cyclohexyl, cyclobutyl, cycloheptyl, cyclopentenyl, cyclopentadienyl, cyclohexenyl, cyclohexadienyl, etc.
[0044] The term "aryl" refers to an aromatic ring system of a hydrocarbon, whether monocyclic, bicyclic, or tricyclic. An aryl group may optionally be substituted with one or more substituents. In one embodiment, 0, 1, 2, 3, 4, 5, or 6 atoms of each ring of the aryl group may be substituted with substituents. Examples of aryl groups include phenyl, naphthyl, anthraceneyl, fluorenyl, indene, azulel, and the like.
[0045] The term "heteroaryl" refers to an aromatic 5-10 membered ring system in which the heteroatom is selected from O, N, or S, and the remaining ring atoms are carbon (with suitable hydrogen atoms unless otherwise specified). A heteroaryl group may optionally be substituted with one or more substituents. In one embodiment, 0, 1, 2, 3, or 4 atoms of each ring of the heteroaryl group may be substituted with substituents. Examples of heteroaryl groups include pyridinyl, furanyl, thiopheneyl, pyrroleyl, oxazolyl, oxadiazolyl, imidazolyl, thiazolyl, isoxazolyl, quinolinyl, pyrazolyl, isothiazolyl, pyridinyl, pyrazinyl, triazinyl, isoquinolinyl, indazoleyl, etc.
[0046] As used herein, the term "substituted" refers to a portion (such as heteroaryl, aryl, alkyl, and / or alkenyl) that is incorporated into one or more additional organic or inorganic substituent radicals. In some embodiments, the substituted portion comprises 1, 2, 3, 4, or 5 additional substituent groups or radicals. Suitable organic and inorganic substituent radicals include, but are not limited to, hydroxyl, cycloalkyl, aryl, substituted aryl, heteroaryl, heterocyclic, substituted heterocyclic, amino, monosubstituted amino, disubstituted amino, acyloxy, nitro, cyano, carboxyl, alkoxycarbonyl, alkylformamide, substituted alkylformamide, dialkylformamide, substituted dialkylformamide, alkylsulfonyl, alkylsulfinyl, thioalkyl, alkoxy, substituted alkoxy, or haloalkoxy radicals, wherein the terms are defined herein. Unless otherwise stated herein, organic substituents may comprise 1 to 4 or 5 to 8 carbon atoms. When more than one substituent radical is attached to the substituted portion, the substituent radicals can be the same or different.
[0047] As used herein, the term “unsubstituted” refers to a portion that is not bound to one or more other organic or inorganic substituent radicals as described above (such as heteroaryl, aryl, alkenyl, and / or alkyl), meaning that this portion is substituted only with hydrogen.
[0048] It should be understood that the structures provided herein and any statements of “substitution” or “replaced by” imply that such structures and substitutions are based on the permissible valences of the substituted atoms and substituents, and that the substitutions produce stable compounds, for example, compounds that do not spontaneously undergo transformations such as by rearrangement, cyclization, elimination, etc.
[0049] As used herein, the term "RNA" refers to ribonucleic acid, a polymeric molecule essential for various biological roles in the coding, decoding, regulation, and expression of genes. RNA and DNA are nucleic acids and, along with lipids, proteins, and carbohydrates, constitute the four major macromolecules essential for all known forms of life. Like DNA, RNA is assembled into nucleotide chains, but unlike DNA, RNA in nature is found as a self-folding single strand rather than a paired double strand. Cellular organisms use messenger RNA (mRNA) to transmit genetic information (using nitrogenous bases of guanine, uracil, adenine, and cytosine, represented by the letters G, U, A, and C) to direct the synthesis of specific proteins. Many viruses use RNA genomes to encode their genetic information. Some RNA molecules play active roles within cells by catalyzing biological reactions, controlling gene expression, or sensing and transmitting responses to cellular signals. One of these active processes is protein synthesis, a common function in which RNA molecules direct protein synthesis on ribosomes. This process uses transfer RNA (tRNA) molecules to deliver amino acids to ribosomes, where ribosomal RNA (rRNA) then links the amino acids together to form the encoded protein.
[0050] As used herein, the term "non-coding RNA (ncRNA)" refers to an RNA molecule that is not translated into protein. The DNA sequence from which functional non-coding RNAs are transcribed is usually referred to as an RNA gene. Numerous and functionally important types of non-coding RNAs include transfer RNA (tRNA) and ribosomal RNA (rRNA), as well as small RNAs such as microRNAs, siRNAs, piRNAs, snoRNAs, snRNAs, exRNAs, and scaRNAs, and long ncRNAs such as Xist and HOTAIR.
[0051] As used herein, the term "coding RNA" refers to RNA that encodes proteins, i.e., messenger RNS (mRNA). This type of RNA includes the transcriptome.
[0052] As used herein, the term "riboswitch" refers to a regulatory segment of a messenger RNA molecule that binds to small molecules, thereby altering the production of proteins encoded by the mRNA. Therefore, mRNAs containing riboswitches directly participate in regulating their own activity in response to the concentration of their effector molecules.
[0053] As used herein, the term "TPP riboswitch," also known as a THI element or Thi-box riboswitch, refers to a highly conserved RNA secondary structure. The TPP riboswitch functions as a riboswitch that binds directly to thiamine pyrophosphate (TPP) to regulate gene expression in archaea, bacteria, and eukaryotes through various mechanisms. TPP is the active form of thiamine (vitamin B1), an essential coenzyme synthesized in bacteria by coupling pyrimidine and thiazole moieties.
[0054] As used herein, the term "pseudoknot" refers to a nucleic acid secondary structure containing at least two stem-loop structures, where one half of the stem is inserted between the two halves of the other stem. Pseudoknots were first discovered in turnip yellow mosaic virus in 1982. Pseudoknots fold into a knot-like three-dimensional configuration but are not true topological knots.
[0055] An aptamer is a nucleic acid molecule capable of binding to a specific molecule of interest with high affinity and specificity (Tuerk and Gold, 1990; Ellington and Szostak, 1990), and can be either artificially engineered or of natural origin. The binding of a ligand to an aptamer (typically RNA) alters the conformation of the aptamer and the nucleic acid to which it resides. In some instances, the conformational change inhibits the translation of the mRNA to which the aptamer resides, for example, or otherwise interferes with the normal activity of the nucleic acid. Aptamers can also be composed of DNA or can include non-natural nucleotides and nucleotide analogs. Aptamers are most typically obtained through in vitro selection for binding to a target molecule. However, in vivo selection of aptamers is also possible. Aptamers are also ligand-binding domains of riboswitch molecule. The length of an aptamer typically ranges from about 10 to about 300 nucleotides. More commonly, the length of an aptamer ranges from about 30 to about 100 nucleotides. See, for example, U.S. Patent No. 6,949,379, which is incorporated herein by reference. Examples of aptamers that can be used in this invention include, but are not limited to, the PSMA aptamer (McNamara et al., 2006), the CTLA4 aptamer (Santulli-Marotto et al., 2003), and the 4-1BB aptamer (McNamara et al., 2007).
[0056] As used in this article, the term "PCR" stands for Polymerase Chain Reaction and refers to a method widely used in molecular biology to rapidly prepare millions to billions of copies of a specific DNA sample, allowing scientists to use very small DNA samples and amplify them to a sufficiently large quantity for careful study.
[0057] The phrase “pharmaceutical acceptable” indicates that the substance or composition is chemically and / or toxicologically compatible with other components constituting the formulation and / or with the subject being treated therein.
[0058] As used herein, the phrase "pharmaceutically acceptable salt" refers to a pharmaceutically acceptable organic or inorganic salt of the compounds of the present invention. Exemplary salts include, but are not limited to, sulfates, citrates, acetates, oxalates, chlorides, bromides, iodides, nitrates, bisulfates, phosphates, acid phosphates, isonicotinates, lactates, salicylates, acid citrates, tartrates, oleates, tannins, pantothenates, bitartrates, ascorbic acid salts, succinates, maleates, gentianates, fumarates, gluconates, glucurons, sucrose salts, formates, benzoates, glutamates, methanesulfonates (mesylate), ethanesulfonates, benzenesulfonates, p-toluenesulfonates, bis(hydroxynaphthyl)ates (i.e., 1,1′-methylene-bis(2-hydroxy-3-naphthylcarbamate)), alkali metal (e.g., sodium and potassium) salts, alkaline earth metal (e.g., magnesium) salts, and ammonium salts. Pharmaceutically acceptable salts may involve another molecule containing, for example, an acetate ion, a succinate ion, or other counterions. The counterion can be any organic or inorganic component that stabilizes the charge on the parent compound. Furthermore, pharmaceutically acceptable salts may have more than one charged atom in their structure. Multiple charged atoms are an example of a pharmaceutically acceptable salt that may have multiple counterions. Therefore, pharmaceutically acceptable salts may have one or more charged atoms and / or one or more counterions.
[0059] As used herein, “carrier” refers to pharmaceutically acceptable loads, excipients, or stabilizers that are non-toxic to cells or mammals exposed to them at the doses and concentrations used. Physiologically acceptable carriers are typically aqueous pH buffer solutions. Non-limiting examples of physiologically acceptable carriers include buffer solutions such as phosphates, citrates, and other organic acids; antioxidants containing ascorbic acid; low molecular weight (less than about 10 residues) peptides; proteins such as serum albumin, gelatin, or immunoglobulins; hydrophilic polymers such as polyvinylpyrrolidone; amino acids such as glycine, glutamine, asparagine, arginine, or lysine; monosaccharides, disaccharides, and other carbohydrates containing glucose, mannose, or dextrin; chelating agents such as EDTA; sugar alcohols such as mannitol or sorbitol; salt-forming counterions such as sodium; and / or TWEEN. TM Polyethylene glycol (PEG) and PLURONICS TM Nonionic surfactants are used. In some embodiments, the pharmaceutically acceptable carrier is a non-naturally occurring pharmaceutically acceptable carrier.
[0060] The terms "treat" and "treatment" refer to therapeutic treatments and preventative or preventative measures aimed at preventing or mitigating (alleviating) undesirable physiological changes or conditions, such as the development or spread of cancer. For the purposes of this invention, beneficial or desired clinical outcomes include, but are not limited to, reduction of symptoms, reduction of disease severity, stable disease (i.e., no worsening), delay or slowing of disease progression, improvement or mitigation of disease status, and parity (whether partial or overall), whether these outcomes are detectable or undetectable. "Treatment" can also mean prolonged survival compared to expected survival without treatment. Those requiring treatment include those already suffering from symptoms or conditions, those susceptible to symptoms or conditions, and those with symptoms or conditions to be prevented.
[0061] The term “administration” or “administering” refers to a route of administration in which a compound is introduced into a subject to perform its intended function. Examples of possible routes of administration include injection (including but not limited to subcutaneous, intravenous, parenteral, intraperitoneal, and intrathecal), local, oral, inhalation, rectal, and transdermal administration.
[0062] The term "effective amount" refers to the amount, measured in doses and sustained for the required period of time, that effectively achieves the desired outcome. The effective amount of a compound can vary depending on factors such as the subject's disease state, age, and weight, as well as the compound's ability to elicit the desired response in the subject. Dosing regimens can be adjusted to provide optimal therapeutic response.
[0063] As used in this article, the phrases “systemic administration” and “administered systematically” and “peripheral administration” refer to the administration of a compound, drug, or other material into the patient’s system, where it undergoes metabolism and other similar processes.
[0064] The phrase "therapeutic effective amount" refers to the amount by which the compound of the present invention (i) treats or prevents a particular disease, symptom, or condition; (ii) alleviates, improves, or eliminates one or more symptoms of a particular disease, symptom, or condition; or (iii) prevents or delays the onset of one or more symptoms of a particular disease, symptom, or condition described herein. In the case of cancer, a therapeutically effective amount of the drug may reduce the number of cancer cells; reduce tumor size; inhibit (i.e., to some extent slow down and preferably terminate) the infiltration of cancer cells into peripheral organs; inhibit (i.e., to some extent slow down and preferably terminate) tumor metastasis; inhibit tumor growth to some extent; and / or alleviate one or more of the symptoms associated with cancer to some extent. The drug may be cytoseptic and / or cytotoxic to the extent that it can stop growth and / or kill existing cancer cells. For cancer therapies, efficacy may be measured, for example, by assessing time to progression (TTP) and / or determining response rate (RR).
[0065] The term "subject" refers to animals such as mammals, including but not limited to primates (e.g., humans), cattle, sheep, goats, horses, dogs, cats, rabbits, rats, mice, etc. In some embodiments, the subject is a human.
[0066] This disclosure relates to a fragment-based ligand discovery strategy suitable for identifying small molecules that bind to specific RNA regions with high affinity. Typically, fragment-based ligand discovery allows the identification of one or more small molecule “fragments” with low to intermediate affinity binding to a target of interest. These fragments are then refined or ligated to produce more potent ligands. 13,14 Typically, these fragments exhibit molecular weights below 300 Da and establish substantial, high-quality contact with the target of interest for detectable binding.
[0067] Fragment-based ligand discovery has so far only been successfully used to identify initial-hit compounds that are single-fragment hits binding to a given RNA. 15-19 Identifying multiple fragments that bind to the same RNA will allow for the utilization of potential additive and cooperative interactions between fragments within the binding notch. 20,21However, it has recently been demonstrated that many RNAs bind to their ligands through multiple "subsites," which are regions of the binding notch that independently or cooperatively contact the ligand. 22 Furthermore, it has been demonstrated that high-affinity RNA binding can occur even when subsite binding exhibits only a moderate synergistic effect. These characteristics are a good indicator of the effectiveness of fragment-based ligand discovery in targeting RNA.
[0068] Therefore, based on the above, this disclosure relates to methods for identifying fragments that bind to RNA of interest, such as TPP riboswitches. Second, the disclosed methods involve establishing the localization of the binding fragment in the RNA at approximately nucleotide resolution. Third, the disclosed methods involve identifying a second-site fragment that binds near the initial fragment hit site. The disclosed methods combine fragment-based ligand discovery methods with SHAPE-MaP RNA structure probing for identifying RNA-binding fragments and establishing individual sites of fragment binding. 23,24 The ligand resulting from linking the two fragments bears no resemblance to the natural riboswitching ligand, and binds to the structurally complex TPP riboswitching RNA with high affinity.
[0069] The disclosed methods and ligand identification will be described in more detail below.
[0070] A. Compound
[0071] The first aspect of the subject currently disclosed is compounds having the structure of formula (I):
[0072]
[0073]
[0074] in
[0075] X1, X2, and X3 are independently selected from CR1, CHR1, N, NH, O, and S in each instance, wherein adjacent X1, X2, and X3 are not simultaneously selected as O or S;
[0076] Dashed lines represent optional double bonds;
[0077] Y1, Y2, and Y3 are independently selected from CR2 and N in each instance;
[0078] n is 1 or 2, where when n is 1, only one of the dashed lines is a double bond;
[0079] L is selected from
[0080]
[0081] Where k, p, q, r, and v are independently selected from the integers 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10, and z is selected from the integers 1, 2, 3, 4, and 5; and
[0082] A is selected from
[0083]
[0084] X4, X5, X6, and X7 are independently selected from CR3 and N;
[0085] R1, R2 and R3 are independently selected from -H, -Cl, -Br, -I, -F, -CF3, -OH, -CN, -NO2, -NH2, -NH(C1-C6 alkyl), -N(C1-C6 alkyl)2, -COOH, -COO(C1-C6 alkyl), -CO(C1-C6 alkyl), -O(C1-C6 alkyl), -OCO(C1-C6 alkyl), -NCO(C1-C6 alkyl), -CONH(C1-C6 alkyl) and substituted or unsubstituted C1-C6 alkyl groups;
[0086] m is 1 or 2; and
[0087] W is -O or -NR4, where R4 is selected from -H, -CO (C1-C6 alkyl), substituted or unsubstituted C1-C6 alkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted cycloalkyl, -CO (aryl), -CO (heteroaryl), and -CO (cycloalkyl).
[0088] The condition is that at least two of X1, X2, X3, X4, X5, X6, and X7 are N;
[0089] Or its pharmaceutically acceptable salt.
[0090] As in any of the above embodiments, a compound wherein at least one of X1, X2, or X3 is N.
[0091] As in any of the above embodiments, a compound wherein X1 is N.
[0092] As in any of the above embodiments, a compound wherein X2 is N.
[0093] As in any of the above embodiments, a compound wherein X3 is N.
[0094] As in any of the above embodiments, a compound, wherein in each instance, two of X1, X2, and X3 are N.
[0095] As in any of the above embodiments, a compound wherein X1 and X3 are N.
[0096] As in any of the above embodiments, a compound wherein at least one of Y1, Y2, and Y3 is N.
[0097] As in any of the above embodiments, a compound wherein Y1 is N.
[0098] As in any of the above embodiments, a compound wherein Y2 is N.
[0099] As in any of the above embodiments, a compound wherein Y3 is N.
[0100] As in any of the above embodiments, a compound wherein at least one of Y1, Y2 and Y3 is CR2.
[0101] As in any of the above embodiments, a compound wherein Y1 is CR2.
[0102] As in any of the above embodiments, a compound wherein Y2 is CR2.
[0103] As in any of the above embodiments, a compound wherein Y3 is CR2.
[0104] As in any of the above embodiments, a compound, wherein n is 2.
[0105] As in any of the above embodiments, a compound having the structure of formula (II):
[0106]
[0107] in
[0108] X 2a and X 2b Independently selected from CR1 and N;
[0109] X1 and X3 are independently selected from CR1 and N;
[0110] L and A are as specified for formula (I); and
[0111] X1, X 2a X 2b The two in X3 are N.
[0112] As in any of the above embodiments, a compound having the structure of formula (III):
[0113]
[0114] in
[0115] L and A are as specified for formula (I).
[0116] As in any of the above embodiments, a compound wherein p, q, r, and v are independently selected from integers 0, 1, 2, and 3.
[0117] As in any of the above embodiments, a compound wherein L is selected from...
[0118]
[0119] As in any of the above embodiments, a compound, wherein L is
[0120]
[0121] As in any of the above embodiments, a compound wherein q and r are 0 or 1.
[0122] As in any of the above embodiments, a compound, wherein q is 1.
[0123] As in any of the above embodiments, a compound, wherein r is 1.
[0124] As in any of the above embodiments, a compound, wherein r is 0.
[0125] As in any of the above embodiments, a compound wherein q and r are 1.
[0126] As in any of the above embodiments, a compound, wherein q is 1 and r is 0.
[0127] As in any of the above embodiments, a compound, wherein m is 1.
[0128] As in any of the above embodiments, a compound wherein W is selected from -NH, -O, and -N(C1-C6 alkyl)2.
[0129] As in any of the above embodiments, a compound wherein W is -NH.
[0130] As in any of the above embodiments, a compound wherein at least one of X4, X5, X6, and X7 is N.
[0131] As in any of the above embodiments, a compound wherein X4 is N.
[0132] As in any of the above embodiments, a compound wherein X5 is N.
[0133] As in any of the above embodiments, a compound wherein X6 is N.
[0134] As in any of the above embodiments, a compound wherein X7 is N.
[0135] As in any of the above embodiments, a compound wherein X4 and X6 are N.
[0136] As in any of the above embodiments, a compound wherein X5 and X7 are N.
[0137] As in any of the above embodiments, a compound wherein X5 or X6 is N, and both X4 and X7 are independently CR2.
[0138] As in any of the above embodiments, a compound, wherein A is
[0139]
[0140] As in any of the above embodiments, a compound having the following structure:
[0141]
[0142] As in any of the above embodiments, a compound, wherein L is
[0143]
[0144] As in any of the above embodiments, a compound wherein Y1, Y2 and Y3 are independently selected from CR2 and N in each instance, wherein R1 is selected from -H, -Cl, -Br, -I, -F, -OH and -NH2.
[0145] As in any of the above embodiments, a compound wherein z is 2.
[0146] As in any of the above embodiments, a compound wherein Y2 is N.
[0147] As in any of the above embodiments, a compound wherein Y2 is CR2 and R1 is selected from -H, -F, -OH and -NH2.
[0148] As in any of the above embodiments, a compound, wherein A is
[0149]
[0150] As in any of the above embodiments, a compound having the following structure:
[0151]
[0152] As in any of the above embodiments, a compound having the following structure:
[0153]
[0154] B. Screening Methods
[0155] This disclosure relates to the development and validation of a fragment screening method based on flexible selective 2'-hydroxyacylation (SHAPE) via primer extension analysis. Fragment-based ligand discovery has proven to be an effective method for identifying compounds that form substantial close contact with macromolecules, including RNA. 13,14,17 The success of this discovery strategy hinges on adaptive, high-quality biophysical assays used to detect ligand binding. Therefore, in some embodiments, SHAPE RNA structure probing is used to detect ligand binding. 23-25 The SHAPE RNA structure probe measures local nucleotide flexibility as the relative reactivity of the ribose 2'-hydroxyl group to an electrophilic reagent. SHAPE can be used on any RNA and provides data on virtually all nucleotides in the RNA in a single experiment, thus generating per-nucleotide structural information in addition to simply detecting binding, as described in detail below. Furthermore, this disclosure also relates to the application of SHAPE-mutation profiling (MaP). 23,24 It combines SHAPE with high-throughput sequencing readout, enabling the reuse of thousands of samples and efficient high-throughput analysis.
[0156] Therefore, in some embodiments, this disclosure relates to a screening method for using SHAPE and / or SHAPE-MaP to identify small molecular fragments and / or compounds that bind to and / or associate with RNA molecules of interest. The methods disclosed herein further include using SHAPE and / or SHAPE-MaP to identify small molecular fragments (e.g., fragment 2) that bind to and / or associate with RNA molecules that have been pre-incubated with another small molecular fragment (e.g., fragment 1). Not bound by theory, it is believed that fragment 1 binds to a first binding site and fragment 2 binds to a second binding site (e.g., a subsite) within the same RNA molecule. Therefore, combining the structural features of fragments 1 and 2 (e.g., connecting the two fragments with a linker L) to produce compounds as disclosed herein is believed to cause the linked fragment ligands to exhibit increased RNA binding affinity compared to fragments 1 and / or 2 alone.
[0157] The filtering methods SHAPE and SHAPE-MaP are described in more detail below.
[0158] I.SHAPE Chemistry
[0159] SHAPE chemistry is at least in part based on the observation that the nucleophilicity of the 2′-position of RNA ribose is sensitive to the electronic influence of the adjacent 3′-phosphodiester group. Unconstrained nucleotides tend to adopt configurations that enhance the nucleophilicity of the 2′-hydroxyl group compared to nucleotides that are base-paired or otherwise constrained. Therefore, hydroxyl-selective electrophiles, such as, but not limited to, N-methylindosinic anhydride (NMIA), form stable 2′-O-adducts with flexible RNA nucleotides more rapidly. Local nucleotide flexibility can be queried simultaneously at all positions in the RNA molecule in a single experiment because all RNA nucleotides (except for a few cellular RNAs carrying post-transcriptional modifications) have a 2′-hydroxyl group. Absolute SHAPE reactivity can be compared at all positions in RNA because 2′-hydroxyl reactivity is insensitive to base identity. It is also possible for a nucleotide to be reactive because it is constrained in a configuration that enhances the nucleophilicity of a particular 2′-hydroxyl group. Such nucleotides are expected to be rare, involving atypical local geometries, and will be correctly scored based on unpaired positions.
[0160] The currently disclosed subject provides, in some embodiments, a method for detecting structural data in RNA molecules by querying structural constraints in RNA molecules of arbitrary length and structural complexity. In some embodiments, the method includes: annealing RNA molecules containing 2′-O-adducts with (labeled) primers; annealing RNA molecules not containing 2′-O-adducts with (labeled) primers as a negative control; extending the primers to generate a library of cDNA; analyzing the cDNA; and generating an output file including structural data of the RNA.
[0161] RNA molecules can be present in biological samples. In some embodiments, RNA molecules can be modified in the presence of proteins or other small and large biological ligands and / or compounds. Primers can optionally be labeled with radioactive isotopes, fluorescent labels, heavy atoms, enzyme labels, chemiluminescent groups, biotinylate groups, predetermined polypeptide epitopes recognized by secondary reporter genes, or combinations thereof. Analysis can include separation, quantification, fractionation, or combinations thereof. Analysis can include extracting fluorescence or dye amount data as a function of elution time data, said fluorescence or dye amount data being referred to as traces. For example, cDNA can be analyzed in a single column or microfluidic device of a capillary electrophoresis apparatus.
[0162] In some embodiments, the peak areas in the traces of nucleotide sequences for RNA molecules containing and without the 2′-O-adduct can be calculated. The traces can be compared and aligned with the RNA sequence. It is observed that the traces of cDNA generated by sequencing are one nucleotide longer than the corresponding positions in the traces of RNA molecules containing and without the 2′-O-adduct. The area under each peak can be determined by performing a Gaussian fitting integral over the entire trace.
[0163] Therefore, in some embodiments, this document provides a method for forming covalent ribose 2′-O-adducts with RNA molecules in complex biological solutions. In some embodiments, the method includes contacting an electrophilic agent with an RNA molecule, wherein the electrophilic agent selectively modifies unconstrained nucleotides in the RNA molecule to form a covalent ribose 1′-O-adduct.
[0164] In some embodiments, an electrophilic agent, such as but not limited to N-methylindorubicin (NMIA), is dissolved in an anhydrous, polar, aprotic solvent such as DMSO. The reagent-solvent solution is added to a complex biological solution containing RNA molecules. The solution may contain varying concentrations and amounts of proteins, cells, viruses, lipids, monosaccharides and polysaccharides, amino acids, nucleotides, DNA, and various salts and metabolites. The concentration of the electrophilic agent can be adjusted to achieve the desired degree of modification in the RNA molecule. The electrophilic agent has the potential to react with any free hydroxyl groups in the solution, thereby generating a ribose 2′-O-adduct on the RNA molecule. Further, the electrophilic agent can selectively modify unpaired or otherwise unconstrained nucleotides in the RNA molecule.
[0165] RNA molecules can be exposed to electrophilic reagents at concentrations that produce small amounts of RNA modification to form 2′-O-adducts, which can be detected by the ability of reverse transcriptase to inhibit primer elongation. All RNA sites can be queried in a single experiment due to the universal reactivity of the chemical targeting the 2′-hydroxyl group. In some embodiments, control elongation reactions to remove the electrophilic reagent to assess background and dideoxy sequencing elongation for nucleotide position allocation can be performed in parallel. These combined steps are referred to as selective 2′-hydroxyacylation or SHAPE by primer elongation analysis.
[0166] In some embodiments, the method further includes: contacting an RNA molecule containing a 1′-O-adduct with (labeled) primers, contacting RNA without a 2′-O-adduct with (labeled) primers as a negative control; extending the primers to generate a linear array of cDNA, analyzing the cDNA, and generating an output file including structural data of the RNA.
[0167] The number of nucleotides queried in a single SHAPE experiment depends not only on the detection and resolution of the isolation technique used, but also on the nature of the RNA modification. Under given reaction conditions, almost all RNA molecules possess at least one modification of length. As primer extension reaches these lengths, the amount of extended cDNA decreases, which weakens the experimental signal. Adjusting conditions to reduce modification yield can increase read length. However, reducing reagent yield can also decrease the measured signal per cDNA length. Given these considerations, the preferred maximum length for a single SHAPE read is approximately 1 kilobase of RNA, but should not be limited to this.
[0168] II.SHAPE-MaP
[0169] In SHAPE-MaP, SHAPE adducts are detected using a mutation profile (MaP), which utilizes the ability of reverse transcriptase to incorporate non-complementary nucleotides or generate deletions at the sites of SHAPE chemical adducts. In some embodiments, SHAPE-MaP can be used for library construction and sequencing. In some embodiments, reuse techniques can be employed in SHAPE-MaP.
[0170] Typically, RNA is treated with SHAPE reagents that react at conformationally dynamic nucleotides. During reverse transcription, a polymerase reads the chemical adducts in the RNA and incorporates nucleotides that are not complementary to the original sequence into the cDNA. The resulting cDNA is sequenced using any massively parallel method to generate a mutation profile (MaP). The sequenced reads are aligned to a reference sequence and the nucleotide-resolution mutation rate is calculated, corrected for background, and normalized to generate a standard SHAPE reactivity profile. The SHAPE reactivity can then be used to model secondary structures, visualize competing and alternative structures, or quantify any processes or functions that regulate local nucleotide RNA dynamics. Following SHAPE modification of the RNA molecule, reverse transcriptase is used to generate a mutation profile. This step encodes the positions and relative frequencies of the SHAPE adducts as mutations in the cDNA. The cDNA is converted to dsDNA using methods known in the art (e.g., PCR reaction), and the dsDNA is further amplified in a second PCR reaction, adding sequencing for reuse. After purification, the sequencing library is uniform in size and each DNA molecule contains the entire sequence of interest.
[0171] Therefore, according to some embodiments of the currently disclosed subject matter, a method for detecting one or more chemical modifications in a nucleic acid is provided. In some embodiments, the method includes: providing a nucleic acid suspected of having a chemical modification; synthesizing the nucleic acid using a polymerase and the provided nucleic acid as a template, wherein the synthesis occurs under conditions where the polymerase reads the chemical modification of the provided nucleic acid, thereby generating an incorrect nucleotide at the site of the chemical modification in the resulting nucleic acid; and detecting the incorrect nucleotide.
[0172] According to some embodiments of the currently disclosed subject matter, methods for detecting structural data in nucleic acids are provided. In some embodiments, the method includes: providing a nucleic acid suspected of having chemical modifications; synthesizing the nucleic acid using a polymerase and the provided nucleic acid as a template, wherein the synthesis occurs under conditions where the polymerase reads the chemical modifications of the provided nucleic acid, thereby generating incorrect nucleotides at the sites of the chemical modifications in the resulting nucleic acid; detecting the incorrect nucleotides; and generating an output file including structural data of the provided nucleic acid.
[0173] In some embodiments of the currently disclosed subject matter, the provided nucleic acid is an RNA molecule (e.g., a coding RNA and / or a non-coding RNA molecule). In some embodiments, the method includes detecting two or more chemical modifications. In some embodiments, a polymerase reads multiple chemical modifications to produce multiple incorrect nucleotides and the method includes detecting each incorrect nucleotide.
[0174] In some embodiments, nucleic acids (e.g., RNA molecules) have been exposed to a reagent that provides chemical modification, or the chemical modification is pre-existing in the nucleic acid (e.g., RNA molecule). In some embodiments, the pre-existing modification is 2'-O-methyl and / or caused by the cell from which the nucleic acid originates, such as, but not limited to, epigenetic modifications, and / or the modification is 1-methyladenosine, 3-methylcytosine, 6-methyladenosine, 3-methyluridine, and / or 2-methylguanosine. In some embodiments, nucleic acids such as RNA molecules may be modified in the presence of proteins or other small and large biological ligands and / or compounds.
[0175] In some embodiments, the reagent includes an electrophilic reagent. In some embodiments, the electrophilic reagent selectively modifies unconstrained nucleotides in an RNA molecule to form a covalently ribose 2'-O-adduct. In some embodiments, the reagent is 1M7, 1M6, NMIA, DMS, or a combination thereof. In some embodiments, the nucleic acid is present in or derived from a biological sample.
[0176] In some embodiments, the polymerase is a reverse transcriptase. In some embodiments, the polymerase is a natural polymerase or a mutant polymerase. In some embodiments, the synthesized nucleic acid is cDNA.
[0177] In some embodiments, detecting incorrect nucleotides includes sequencing the nucleic acid. In some embodiments, the sequence information is compared with the sequence of the provided nucleic acid. In some embodiments, detecting incorrect nucleotides includes using massively parallel sequencing of the nucleic acid. In some embodiments, the method includes amplifying the nucleic acid. In some embodiments, the method includes amplifying the nucleic acid using a site-specific method with specific primers, a whole genome using random primers, a whole transcriptome using random primers, or a combination thereof.
[0178] According to some embodiments of the currently disclosed subject matter, a computer program product is provided comprising computer-executable instructions embodied in a computer-readable medium in the form of execution steps, said execution steps including any method steps of any embodiment of the currently disclosed subject matter. According to some embodiments of the currently disclosed subject matter, a nucleic acid library generated by any method of the currently disclosed subject matter is provided.
[0179] III. SHAPE Electrophilic Reagent
[0180] As disclosed above, SHAPE chemistry utilizes the finding that the nucleophilic reactivity of the 2′-hydroxyl group of ribose is gated by the flexibility of the local nucleotide. At nucleotides constrained by base pairing or tertiary interactions, the 3′-phosphodiester anion and other interactions reduce the reactivity of the 2′-hydroxyl group. In contrast, flexible sites preferentially adopt a configuration that reacts with electrophilic reagents, including but not limited to NMIA, to form a 2′-O-adduct. For example, NMIA generally reacts with all four nucleotides and the reagent undergoes parallel, self-inactivation, and hydrolysis reactions. In fact, the subject matter currently disclosed provides for the use of any molecule that can react with nucleic acids as disclosed herein, based on some embodiments of the subject matter currently disclosed. In some embodiments, the electrophilic reagent (also referred to as the SHAPE reagent) may be selected from, but is not limited to, indocyanine anhydride derivatives, benzoyl cyanide derivatives, benzoyl chloride derivatives, phthalic anhydride derivatives, benzyl isocyanate derivatives, and combinations thereof. Indocyanine anhydride derivatives may include 1-methyl-7-nitroindocyanine anhydride (1M7). Benzoyl cyanide derivatives may be selected from, but are not limited to, the group consisting of, but not limited to, benzoyl cyanide (BC), 3-carboxybenzoyl cyanide (3-CBC), 4-carboxybenzoyl cyanide (4-CBC), 3-aminomethylbenzoyl cyanide (3-AMBC), 4-aminomethylbenzoyl cyanide, and combinations thereof. Benzoyl chloride derivatives may include benzoyl chloride (BCl). Phthalic anhydride derivatives may include 4-nitrobenzenephthalic anhydride (4NPA). Benzyl isocyanate derivatives may include benzyl isocyanate (BIC).
[0181] IV. RNA Molecular Design
[0182] Because SHAPE reactivity can be assessed in one or more primer extension reactions, information may be lost at the 5′ end of the RNA molecule and near the primer binding site. Typically, adduct formation within 10–20 nucleotides adjacent to the primer binding site is difficult to quantify due to the presence of cDNA fragments reflecting pauses or non-modular extensions via reverse transcriptase (RT) during the initiation phase of primer extension. Due to the abundance of full-length extension products, 8–10 sites at the 5′ end of the RNA may be difficult to visualize.
[0183] To monitor the SHAPE reactivity at the 5′ and 3′ ends of the sequence of interest, RNA molecules can be embedded within large fragments of the native sequence or placed between strongly folded RNA sequences containing unique primer-binding sites. In some embodiments, the cassette can be designed to contain 5′ and 3′ side sequences of nucleotides to allow evaluation of all positions within the RNA molecule of interest in any separation technique providing nucleotide resolution, such as, but not limited to, sequencing gels, capillary electrophoresis, etc. In some embodiments, both the 5′ and 3′ extensions can fold into stable hairpin structures that do not interfere with the folding of various internal RNAs. The primer-binding sites of the cassette can bind efficiently to cDNA primers. The sequence of any 5′ and 3′ cassette element can be examined to ensure that the element does not readily react with the internal sequence to form stable base pairs.
[0184] In some embodiments, the RNA molecule of interest includes two distinct target motifs linked to a nucleotide linker. The target motif can be any nucleotide sequence of interest. Exemplary target motifs include, but are not limited to, riboswitches, viral regulatory elements, structured regions in mRNA, multi-helix knots, pseudoknots, and / or aptamers. In some embodiments, the first target motif is a pseudoknot, such as a pseudoknot from the 5′UTR of the dengue virus genome. In some embodiments, the second target motif is an aptamer domain, such as the TPP riboswitch aptamer domain. The number of nucleotides in the nucleotide linker can vary. For example, in some embodiments, the number of nucleotides in the linker ranges from about 1 to about 20 nucleotides, about 1 to about 15 nucleotides, about 1 to about 10 nucleotides, or about 5 to about 10 nucleotides (or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides).
[0185] In some embodiments, the RNA molecule further includes an RNA barcode region. An RNA barcode region is a unique barcode that allows identification of a particular RNA molecule within a mixture of RNA molecules (e.g., during reuse). The location of the RNA barcode region can vary, but it is typically found adjacent to one of the boxes present in the RNA molecule. In some embodiments, the RNA barcode is designed to fold into a self-contained structure that does not interact with any other part of the RNA molecule. The structure of the RNA barcode region can vary. In some embodiments, the structure of the RNA barcode region includes a base pair helix comprising about 1 to about 10 base pairs (or about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 base pairs). In some embodiments, the RNA barcode region comprises 7 base pairs. In some embodiments, the base pairs are capped with a quadruple loop anchored to the terminal base pairs of the base pair helix. The capping of the base pair helix maintains the overall hairpin stability of the RNA barcode region. In some embodiments, the quadruple loop comprises the nucleotide sequence GNRA, but is not intended to be limited thereto. In some embodiments, the RNA barcode region is designed such that any single barcode undergoes at least two mutations to be misinterpreted as another barcode.
[0186] V. Folding of RNA molecules
[0187] The currently disclosed subject matter can be implemented using RNA molecules generated by methods including, but not limited to, in vitro transcription and RNA molecules generated in cells and viruses. In some embodiments, RNA molecules can be purified and renatured by denaturing gel electrophoresis to achieve a biologically relevant conformation. Further, any procedure for folding RNA molecules to a desired conformation at a desired pH (e.g., about pH 8) can be replaced. RNA molecules can first be heated and rapidly cooled in a low ionic strength buffer to eliminate multimeric forms. A folding solution can then be added to achieve the appropriate conformation of the RNA molecules and make them ready for structure-sensitive probing with electrophilic reagents. In some embodiments, RNA can be folded in a single reaction and then separated for (+) and (-) electrophilic reagent reactions. In some embodiments, the RNA molecules are not natively folded prior to modification. Modification can be performed while the RNA molecules are denatured by heat and / or low-salt conditions.
[0188] VI. RNA molecular modification
[0189] Electrophilic reagents can be added to RNA to generate 2'-O-adducts at flexible nucleotide sites. The reaction can then be incubated until almost all of the electrophilic reagent has reacted with the RNA or has been degraded due to hydrolysis with water. No specific quenching step is required. Modification can be performed in the presence of complex ligands and biomolecules, as well as in the presence of various salts. RNA can also be modified within cells and viruses. These salts and complex ligands can contain salts of magnesium, sodium, manganese, iron, and / or cobalt. Complex ligands can contain, but are not limited to, proteins, lipids, other RNA molecules, DNA, or small organic molecules. In some embodiments, complex ligands are small molecular fragments as disclosed herein. In some embodiments, complex ligands are compounds as disclosed herein. Modified RNA can be purified from the reaction products and buffer components that may be detrimental to primer extension reactions, for example, by ethanol precipitation.
[0190] VII. Primer extension and polymerization
[0191] Analysis of RNA adducts via primer extension according to the currently disclosed subject matter may, in various embodiments, include the use of optimized primer binding sites, thermostable reverse transcriptases, low MgCl2 concentrations, elevated temperatures, shorter extension times, and combinations thereof. Intact, undegraded RNA, free of reaction byproducts and other small molecule contaminants, can also be used as a template for reverse transcription. The RNA component of the resulting RNA-cDNA hybrid can be degraded by treatment with an alkali. The cDNA fragment can then be resolved using, for example, polyacrylamide sequencing gels, capillary electrophoresis, or other separation techniques that will be readily apparent to those skilled in the art upon review of this disclosure.
[0192] Deoxyribonucleotide triphosphates (dATP), dCTP, dGTP, and dTTP, and / or deoxyribonucleotide triphosphates (dNTPs) can be added to the synthetic mixture in sufficient quantities, alone or together with primers, and the resulting solution can be heated to about 50-100°C for about 1 to 10 minutes. After the heating period, the solution can be cooled. In some embodiments, suitable reagents for inducing primer extension reactions can be added to the cooled mixture, and the reaction can be carried out under conditions known in the art. In some embodiments, reagents for polymerization can be added together with other reagents, provided they are thermally stable. In some embodiments, the synthetic (or amplification) reaction can be carried out at room temperature. In some embodiments, the synthetic (or amplification) reaction can be carried out at at most a temperature above which the reagents for polymerization no longer function.
[0193] The agents used for polymerization can be any compound or system for synthesizing primer extension products comprising, for example, enzymes. Suitable enzymes for this purpose include, but are not limited to, E. coli DNA polymerase I, the Klenow fragment of E. coli DNA polymerase, polymerase mutant proteins, reverse transcriptases, and other enzymes, including thermostable enzymes (i.e., those that extend primers after being subjected to temperatures raised sufficiently to cause denaturation), such as mouse or avian reverse transcriptases. Suitable enzymes can facilitate the combination of nucleotides in a suitable manner to form primer extension products complementary to the nucleic acid strands of each polymorphic locus. In some embodiments, synthesis may begin at the 5' end of each primer and proceed in the 3' direction until synthesis terminates at the end of the template by incorporation of dideoxynucleotide triphosphates or at a 2'-O-adduct, thereby producing molecules of varying lengths.
[0194] The newly synthesized strand and its complementary nucleic acid strand can form a double-stranded molecule under the hybridization conditions described herein, and this hybrid is used in subsequent steps as disclosed in U.S. Patent Nos. 10,240,188 and 8,318,424, which are incorporated herein by reference in their entirety. In some embodiments, the newly synthesized double-stranded molecule can also be subjected to denaturing conditions using any procedure known in the art to provide a single-stranded molecule.
[0195] VII. Processing of Raw Data
[0196] The topics described herein for nucleic acid, chemical modification analysis, and / or nucleic acid structure analysis, such as RNA molecules, can be implemented using a computer program product comprising computer-executable instructions embodied in a computer-readable medium. Exemplary computer-readable media suitable for implementing the topics described herein include on-chip memory devices, disk-based memory devices, programmable logic devices, and application-specific integrated circuits (ASICs). Furthermore, the computer program product for implementing the topics described herein may reside on a single device or computing platform or may be distributed across multiple devices or computing platforms. Therefore, the topics described herein may include a set of computer instructions that, when executed by a computer, perform specific functions for nucleic acids, such as RNA structure analysis.
[0197] Considering items I-VII mentioned above, a modular RNA screening construct was designed to implement SHAPE as a high-throughput assay for ligand binding readout. Figure 1 (Top). The construct is designed to contain two target motifs, such as a pseudoknot from the 5'UTR of the dengue virus genome, which reduces viral fitness when its structure is disrupted. 26 ; and the TPP riboswitching aptamer domain 27–29The inclusion of two distinct structural motifs within a single construct allows each motif to serve as an internal, specific control for the other. Fragments binding to both RNA structures are readily identified as nonspecific binders. The two structures are linked by a hexanucleotide linker and are engineered to be single-stranded to maintain structural independence between the two RNA structures. A structural cassette is attached to the structural core of the construct. 25 These stem-loop forming regions serve as primer binding sites for the steps required in the screening workflow and are designed not to interact with other structures in the construct (Figure 7).
[0198] Another component of the screening construct is the RNA barcode; barcoding supports reuse, which significantly reduces downstream workload. Each well in a 96-well plate used for screening fragment libraries contains RNA with a unique barcode against a background of otherwise identical constructs; the barcode sequence thus identifies the well location and the presence of one (or more) fragments after reuse. Figure 1 The RNA barcode regions are designed as independent structures that do not interact with any other part of the construct. The barcode structure is a seven-base-pair helix, which is capped with a GNRA quadruple ring and anchored with GC base pairs to maintain hairpin stability (Figure 7). Each set of 96 barcodes is designed such that any single barcode undergoes two or more mutations to be misinterpreted as another barcode.
[0199] This structure offers flexibility in selecting RNA structures for ligand binding screening and supports simple, direct screening experiments. Figure 1 Each well in a 96-well plate containing an otherwise identical RNA construct with a unique RNA barcode is incubated with one or more small fragments or a fragment-free control (solvent) and then exposed to SHAPE reagent. The resulting SHAPE adduct chemically encodes structural information for each nucleotide. Following SHAPE detection, the information required to determine fragment identity (RNA barcode) and fragment binding (SHAPE adduct pattern) is permanently encoded into each RNA strand, allowing RNA from the 96 wells of the plate to be pooled into a single sample. The fragment screening assay is very similar to the standard MaP structure detection workflow. 24 For example, in some embodiments, a specialized relaxation-fidelity reverse transcription reaction is used to prepare cDNA containing non-template coding sequence variations at any SHAPE adduct site on the RNA. 30 These cDNAs were then used to prepare DNA libraries for high-throughput sequencing. Multiple plates in the experiment could be bar-coded at the DNA library level. 24 To collect data on thousands of compounds in a single sequencing run. Figure 1The resulting sequencing data contains millions of individual reads, each corresponding to a specific RNA strand. These reads are sorted by barcode to allow analysis of data for each small fragment or combination of fragments. The determination and identification of small fragments (e.g., fragment 1 and / or fragment 2) using methods such as SHAPE and / or SHAPE-MaP described above will be described in more detail in the next section.
[0200] C. Ligand identification and selection
[0201] As mentioned above, SHAPE and SHAPE-MaP are used to identify small fragments that bind to or associate with RNA molecules of interest. Specifically, when testing small fragments using SHAPE-Map, detecting the binding fragment characteristics based on the SHAPE-MaP mutation rate per nucleotide involves multiple steps to normalize data and ensure statistical rigor in large-scale experimental screening. Key features of the SHAPE-based hit analysis strategy include: (i) comparison of each fragment-exposed RNA or “experimental sample” with five negative, fragment-free control samples, taking into account plate-to-plate and well-to-well variability; (ii) individual hit detection for each of the two structural motifs in the construct (pseudoknot and TPP riboswitches in this disclosure); (iii) masking of individual nucleotides with low reactivity across all samples, as these nucleotides are unlikely to show fragment-induced changes; and (iv) calculation of the difference in mutation rate per nucleotide between the fragment-exposed experimental sample and the negative control sample without fragment exposure. Nucleotides with a mutation rate difference of 20% or greater between one of the motifs and the fragment-free control are selected for Z-score analysis. However, those skilled in the art will be able to adjust the differences in mutation rates accordingly, recognizing that mutation rates can vary. For example, in some embodiments, the difference in mutation rates can be 25%, 30%, 35%, 45%, or 50% or greater. In some embodiments, the difference in mutation rates can be 15%, 10%, or 5% or greater. If the Z-value of three or more nucleotides in one of the two motifs is greater than 2.7 (as determined by comparing the Poisson counts of the two motifs), this will be beneficial. 31 (See Example 2) to determine if the fragment has a significantly altered SHAPE reactivity pattern. However, the Z-value can vary, and those skilled in the art will be able to adjust it accordingly. For example, in some embodiments, the Z-value is greater than 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, or 3.9. In some embodiments, the Z-value is greater than 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, or 2.6.
[0202] To identify the small molecule fragments subsequently linked together to produce the compounds disclosed herein using SHAPE and / or SHAPE-MaP, a series of steps are performed. First, a primary screening is conducted, screening a large number of compounds, such as at least 100 compounds, to identify any initial lead or hit compounds that exhibit suitable binding activity to the target RNA molecule. In step 2, these hit compounds are further examined in a structure-activity relationship (SAR) study, where changes in target RNA binding affinity are determined as the structure of the hit compound is modified. When multiple small molecule fragments are identified as suitable binding ligands for the target RNA molecule, additional binding studies can be performed to further investigate the binding site of each small molecule fragment (i.e., step 3). For example, in some embodiments, the target RNA can be pre-incubated with a first fragment (identified as a target RNA binding ligand according to the SAR study in step 2), and then the target RNA can be exposed to a second fragment (also identified as an RNA binding ligand in the SAR study in step 2) to determine whether the second fragment can bind to the target RNA while the first fragment has already bound. Once the second fragment has been identified as having suitable binding activity with the RNA of interest, it can be ligated to the first fragment using a adapter to produce the compound disclosed herein (i.e., step 4). Each of the steps mentioned above is described in more detail below.
[0203] Step 1: Initial Screening
[0204] In the initial screening, 1,500 segments were tested and 41 segments were detected as hits, with an initial hit rate of 2.7%. Hit verification was performed using triplicate SHAPE analysis. Figure 2 , Figure 8 Compounds were accepted as true hits only when they were detected as binding agents in all three replicates. These replicate hits were then analyzed by isothermal titration calorimetry (ITC) to determine binding affinity for RNAs corresponding only to the target motif (side sequences in the screening constructs omitted). Of these initial hits, eight were validated by replicate analysis and ITC (Table 1). Seven of these hits bound to the TPP riboswitch based on a mutation signature that was largely or entirely located within the TPP riboswitch region of the test construct. The remaining hits were nonspecific because this fragment affected nucleotides in all parts of the RNA construct. No compounds specifically binding to the dengue pseudoknot region of the test construct were detected.
[0205] Table 1: Fragments of TPP riboswitches detected by SHAPE probe.
[0206]
[0207] Hit detection was performed using SHAPE structure probing and verified through repeatability analysis and ITC. The dissociation constant was determined using ITC; marked with... The error values represent the standard error derived from ≥3 repetitions; other error estimates are calculated based on the 95% confidence interval of the least-squares regression of the binding curve. Natural TPP ligands are included for comparison.
[0208] The seven fragments of the TPP-binding riboswitches validated by the ITC have different chemical types; most bear little resemblance to the natural TPP ligand (Table 1). Overall, heteroaromatic nitrogen-containing rings are predominant; these may be involved in hydrogen bonding interactions. Three compounds have pyridine rings and two have pyrazine rings. An azole ring moiety is present in three compounds: two thiadiazoles and one imidazole. A thiazolium ring is present in the natural TPP ligand, but this moiety does not participate in RNA binding interactions. 28,29,33 In addition, many identified fragments contain primary amines, esters, and ethers, as well as fluorine groups, which can act as hydrogen bond acceptors or donors.
[0209] Step 2: Structure-activity relationship (SAR) of riboswitch binding fragments
[0210] Next, several initially hit analogs were examined to increase binding affinity and identify fragment hit sites that could be modified with adapters without hindering binding. Specifically, analogs of compounds 2 and 5 were considered because these two fragments are structurally different and the analogs are commercially available. Analog-RNA binding was evaluated by ITC. Sixteen analogs of 2 were tested. Modifying the core quinoxaline structure of 2 by removing one or two cyclic nitrogens resulted in changes in binding activity (Table 2A).
[0211] Table 2A: SAR of fragment 2 analogues.
[0212]
[0213]
[0214] The modification of the quinoxaline core was examined, and the dissociation constant was obtained by ITC.
[0215] The improved binding affinity is due to the introduction of a methylene-linked hydrogen bond donor or acceptor (Table 2B, compounds 16 and 17). Substituent changes at other positions on the quinoxaline ring core lead to decreased binding activity. Compound 2 is a promising candidate for further development based on its high flexibility, and even improved binding was observed after modification of the substituent at the C-6 position.
[0216] Table 2B: Structure-activity relationship of fragment 2 analogs that bind to TPP riboswitch RNA. Modification of side groups of the quinoxaline core. Dissociation constants were obtained via ITC.
[0217]
[0218]
[0219] Next, an examination of 18 analogues of fragment 5 showed that the core pyridine functionality of the molecule appears to be important for binding, as changing the position of the cyclic nitrogen, adding or removing the cyclic nitrogen all reduced or abolished the binding (Table 3).
[0220] Table 3: Structure-activity relationship of fragment 5 analogs that bind to TPP riboswitch RNA. Modifications to the pyridine core and dissociation constants were obtained via ITC.
[0221]
[0222]
[0223] Modification of the ring substituents generally leads to a significant loss of binding activity (Table 4). The only analogue with increased affinity, S12, is characterized by chlorine at the C-4 position, resulting in a compound with approximately three times the affinity for the TPP riboswitches compared to fragment 5.
[0224] Table 4: Structure-activity relationship of fragment 5 analogs that bind to TPP riboswitching RNA. Modification of side groups of the pyridine core. Dissociation constants were obtained via ITC.
[0225]
[0226]
[0227] Step 3: Identify the fragment that binds to the second site on the TPP riboswitch.
[0228] A second round of screening was used to identify fragments that bind to the TPP riboswitch region of the screening constructs pre-bound to compound 2 or S12. This screening identified fragments that preferentially interacted with the TPP riboswitch when 2 or S12 was already bound, either due to cooperation or due to structural changes occurring during primary ligand binding, making new binding modes available (Figure 3). Of the 1,500 fragments screened, five were verified to bind simultaneously to 2 or S12 (Table 5).
[0229] Table 5: Fragments that bind to the TPP riboswitches in the presence of pre-binding fragment couplers, as detected by SHAPE. Hit detection was verified by repeated SHAPE analysis. Primary binding couplers (2, 6) are shown in Table 1.
[0230]
[0231] A second-screen hit 29 induced a very strong change in the SHAPE reactive signal and appeared to cause a significant alteration in RNA structure, including the unwinding of the P1 helix. This fragment caused changes in other regions of the RNA consistent with nonspecific interactions, therefore this fragment was not considered further as a candidate for fragment ligation. Fragment 28 was insoluble at the concentrations required for ITC analysis; therefore, related analogues containing pyridine instead of a quinoline ring were examined by ITC (Table 6). These compounds bound with weak affinity; nevertheless, 31 and 32 showed significant but modest synergistic binding with 2.
[0232] Table 6: Structure-activity relationship of fragment 28 analogs binding to TPP riboswitching RNA in the presence and absence of pre-bound fragment 2. *
[0233]
[0234]
[0235] *und (Undetermined): Due to the inability to fit the ITC binding curve; Insoluble: The compound is insoluble at the concentration required for ITC.
[0236] Table 7: Detailed comparison of representative protein and RNA fragment-linker-fragment ligands developed using a fragment-based approach. RNA examples are highlighted with an asterisk. Each entry details the two component fragments and their individual K+. d Values, linked compounds and their corresponding K d The values, as well as the ligand efficiency (LE) and linkage coefficient (E) of the linked compounds. 22,38,53,54,45-52 .
[0237]
[0238]
[0239] Step 4: Coordination and Fragment Connections
[0240] The co-binding interaction between 2 and 31 was quantified using ITC. Individually, 2 binds at 25 μM K. d Combined, and 31 with a much higher K content of 10mM d Binding. For example, in secondary screening, the affinity of fragment 31 was also examined when fragment 2 pre-binded to TPP riboswitching RNA, thereby forming a 2-RNA complex. Under these conditions, fragment 31 bound at approximately 3 mM K. dThe fragments bind to the 2-TPP RNA complex (Figure 4). This experiment also showed that when the binding of 2 is saturated, 31 binds to the TPP RNA, meaning that the two fragments do not bind at the same site. Since 2 and 31 bind to different regions of the TPP RNA with excellent and reasonable affinity, the two fragments are associated with targets that produce high-affinity ligands.
[0241] Based on SAR analysis of fragment hits 2 (Table 2B) and 28 (Table 4), linker analogs of the most promising SAR fragments were prepared, focusing on the aminomethyl position at 17 and two sites in the pyridine ring of fragment 31. Figure 5 First, the affinity of fragments conjugated with amide or amine joints was compared. The compound with the flexible amine joint (compound 36) showed a stronger binding affinity than the amide-linked type (compound 35). Figure 5 The binding affinity is five times higher. These linkages are introduced in the context of hydroxamic acids, which can chelate magnesium ions. 35 Such as in the case of pyrophosphate moieties containing natural TPP ligands 27,28 However, the amine-linked hydroxamic acid compound 36 binds with an affinity similar to that of the parent fragment 17, indicating that the hydroxamic acid moiety does not confer additional binding affinity through chelating ions. The linked compound 37 binds with an affinity of 625 nM, suggesting that high nanomolar binding can be achieved by linking two fragments with moderate affinity under the correct approximation. Replacing the fragment 31 entity with a tertiary amine (compound 38) reduces the affinity relative to compound 37, indicating that the interaction of fragment 31 with RNA is not solely mediated by charge-based effects. Finally, altering the connection between the 17 and 31 moieties by changing the length (compound 39) or the pyridine ring linker site (compound 40) reduces the affinity relative to compound 37. Figure 5 Finally, by linking compounds that individually bind to the TPP riboswitch with an affinity of 5.0 μM (compound 19) and ≥10 mM (compound 31), a K-type riboswitch with an affinity of 625 nM was produced. d Compounds that bind to RNA (37).
[0242] Those skilled in the art will understand that steps I-IV above are not intended to be limiting, but are merely exemplary embodiments. It will be fully understood that those skilled in the art will be able to apply steps I-IV above to identify alternative fragments of compounds that can be linked together to produce the compounds disclosed herein that have suitable binding affinity for TPP riboswitches. Furthermore, it will be fully understood that those skilled in the art will be able to apply steps I-IV above to identify fragments of compounds that can be linked together to produce the compounds disclosed herein that bind to other RNA molecules of interest.
[0243] D. Summary and other considerations
[0244] Since both coding (mRNA) and non-coding RNA can potentially be manipulated to alter cellular regulatory and disease processes, there is a need to develop an efficient strategy for identifying small molecule ligands for structured RNA. The research presented in this paper demonstrates the promise of using SHAPE screening readouts for detecting RNA-bound ligands, fused with a fragment-based strategy. Here, this strategy is used to generate K+ at 625 nM. d This study identifies structurally independent ligands that bind to TPP riboswitches. Fusion-based SHAPE and fragment-based screening methods are applicable to both targetable RNA structures and exploitable ligand chemotypes. The strategy is particularly well-suited for identifying ligands for RNAs with complex structures, which may be crucial for identifying RNA motifs that bind in three-dimensional notches. 4 Furthermore, due to the use of the MaP method and the application of multiplexing via RNA and DNA barcoding, the workload required to screen libraries of more than a thousand member fragments is moderate, thus enabling efficient screening of many structurally different targets.
[0245] Many of the ligands obtained are similar to those previously reported for single-round screening, which also targeted the TPP riboswitch. 15,17 The initial screening hits appear to be moderately biased towards higher affinities, resulting in most ligands detected by SHAPE binding at 10–300 μM. The hit detection assays used may be biased towards detecting the tightest-binding fragments and those that induce the most significant changes in SHAPE reactivity. Lower-affinity fragments may be missed. This bias towards tightly bound fragments is believed to be an overall advantage. No fragments were identified that bind to dengue pseudoknots with the affinity and specificity required to meet the above screening criteria. Dengue pseudoknot RNA is highly structured, and the likelihood of fragments perturbing this structure may be low. Another possibility is that this particular pseudoknot structure may not contain ligand-compatible notches.
[0246] The fragment pair identification strategy, which involves selecting fragments from the primary screening that pre-bind to RNA and then screening for additional fragments to bind to their mates, specifically utilizes per-nucleotide information obtainable via SHAPE and was successfully used here to discover induced-fit fragment pairs (Figure 4). The core principle of fragment-based ligand development is that synergy between two fragments can be achieved through proximal binding, and this additive binding can be utilized by linking the synergistic fragments together using minimally invasive covalent linkers. 20 ,21,36 ,37Compound 37, developed from primary and secondary fragment hits, demonstrates that fragment-based ligand discovery can be efficiently applied to RNA targets. Moderate synergy exists between 2 and 31: when 2 is pre-bound to RNA, the binding of compound 31 is enhanced by 3 to 10-fold. Moderate additive binding energies were observed when connecting the two fragments: the affinity of 37 is 625 nM. No hyperadditive effect was observed when connecting fragments 2 and 31. 36 This could be because perfect fragment localization was not achieved. Small changes in the length or geometry of the joint can lead to significant variations in the affinity of the ligands. Figure 5 This means that the precise orientation of the linker is crucial for optimally orienting the two fragments. The successful development of compound 37 reveals that it is not necessary to achieve perfection in the degree of synergy between the fragments or in the construction of the covalent linker connecting them in order to efficiently develop submicromolar ligands.
[0247] Despite significant efforts to leverage synergy between fragments to obtain tight-binding ligands for targeting proteins, RNA targeting remains in its early stages. This study explored the extent to which the disclosed SHAPE-based screening strategy correlated with fragment ligation compared to previous protein-focused efforts. Compounds previously discovered using fragment-based strategies were ranked according to their linkage coefficient (E), a measure of the overall effectiveness of the system working together during ligation. 21,38 ( Figure 6 (Details are provided in Table 7). In the absence of positive or negative influencing factors, the binding energies of the two fragments are exactly additive, the linker is inert, and E equals 1.0. Cooperative effects or favorable linker interactions decrease E, while anti-cooperative effects or negative linker interactions increase E. Crucially, the E value in protein systems can vary by orders of magnitude. The linkage coefficient of 37 is 2.5, slightly higher than the average for linker (protein-targeting) ligands in the academic literature. The ligand efficiency (LE) of 37, i.e., the free energy of binding divided by the number of non-hydrogen atoms, is favorable compared to examples of linker ligands targeting proteins. Figure 6 By these metrics, 37 performs almost as well as TPPc, a ligand closely associated with the natural TPP riboswitching ligand. 22 Therefore, fragment-based ligand discovery, especially fragment-based ligand discovery that is efficiently implemented through SHAPE-enabled reuse screening, holds great promise for rapidly developing unique ligands targeting a wide range of RNA structures.
[0248] E. Preparation method
[0249] This disclosure also relates to any methods for preparing the compounds disclosed herein. Those skilled in the art will recognize that such preparation methods may vary. For example, in some embodiments, methods for preparing the disclosed compounds include:
[0250] Make Form IV fragment:
[0251] Formula (IV)
[0252] in
[0253] X1, X2, and X3 are independently selected from CHR1, CR1, and heteroatoms N, NH, O, and S, wherein adjacent X1, X2, and X3 are not simultaneously selected as O or S;
[0254] Dashed lines represent optional double bonds;
[0255] Y1, Y2, and Y3 are independently selected from CR2 and N;
[0256] R1 and R2 are independently selected from -H, -Cl, -Br, -I, -F, -CF3, -OH, -CN, -NO2, -NH2, -NH(C1-C6 alkyl), -N(C1-C6 alkyl)2, -COOH, -COO(C1-C6 alkyl), -CO(C1-C6 alkyl), -O(C1-C6 alkyl), -OCO(C1-C6 alkyl), -NCO(C1-C6 alkyl), -CONHC1-C6(alkyl), and substituted or unsubstituted C1-C6 alkyl groups; and
[0257] n is selected from integers 1 and 2, where when n is 1, only one of the dashed lines is a double bond;
[0258] With equation V-1 or V-2:
[0259]
[0260] Where X is a halogen selected from F, Br, Cl and I;
[0261] X4, X5, X6, and X7 are independently selected from CR3 and N;
[0262] R3 is selected from -H, -Cl, -Br, -I, -F, -CF3, -OH, -CN, -NO2, -NH2, -NH(C1-C6 alkyl), -N(C1-C6 alkyl)2, -COOH, -COO(C1-C6 alkyl), -CO(C1-C6 alkyl), -O(C1-C6 alkyl), -OCO(C1-C6 alkyl), -NCO(C1-C6 alkyl), -CONHC1-C6(alkyl), and substituted or unsubstituted C1-C6 alkyl groups;
[0263] m is 1 or 2; and
[0264] W is -O or -NR4, where R4 is selected from -CO (C1-C6 alkyl), substituted or unsubstituted C1-C6 alkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted cycloalkyl, -CO (aryl), -CO (heteroaryl), and -CO (cycloalkyl).
[0265] Contact is carried out in the presence of a Pd catalyst.
[0266] In some embodiments, the Pd catalyst is selected from (DPPF)PdCl2, Pd2(dba)3, PdCl2[P(o-tolyl)3]2, Pd(dba)2, and Pd(OAc)2. In some embodiments, the contacting step further includes a phosphine ligand. In some embodiments, the phosphine ligand is monodentate. In some embodiments, the phosphine ligand is bidentate. Exemplary phosphine ligands include, but are not limited to, DPPF, BINAP, and rac-BINAP. In some embodiments, the contacting step further includes a base. In some embodiments, the base is inorganic. In some embodiments, the base is NaOtBu. In some embodiments, the contacting step is carried out purely (i.e., without solvent). In some embodiments, the contacting step is carried out in the presence of a solvent. In some embodiments, the solvent is a nonpolar solvent. Exemplary solvents include, but are not limited to, toluene, benzene, dioxane, and tetrahydrofuran. In some embodiments, the contacting step is carried out at an elevated temperature. In some embodiments, the contact step is performed at 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, 80°C, 85°C, 90°C, 95°C, 100°C, 105°C, or 100°C.
[0267] In some embodiments, a method for preparing the disclosed compound includes:
[0268] Make Form IV fragment:
[0269] Formula (IV)
[0270] With formula VI-1 or VI-2 fragments:
[0271]
[0272] Where X1, X2, X3, X4, X5, X6, X7, Y1, Y2, Y3, n, m, and W are defined as above.
[0273] Contact occurs in the presence of a reducing agent.
[0274] In some embodiments, the reducing agent can be any reducing agent suitable for reducing amination chemistry. Exemplary reducing agents include, but are not limited to, borohydrides and / or aluminum hydride. In some embodiments, the reducing agent is a borohydride. In some embodiments, the reducing agent is sodium borohydride. In some embodiments, the contact step is performed purely. In some embodiments, the contact step is performed in a solvent. Exemplary solvents include, but are not limited to, alcohol solvents (e.g., methanol, ethanol, isopropanol), chlorinated solvents (e.g., dichloromethane), and / or ether solvents (e.g., tetrahydrofuran). In some embodiments, the contact step is performed below room temperature. In some embodiments, the contact step is performed at an elevated temperature.
[0275] In some embodiments, a method for preparing the disclosed compound includes:
[0276] Make Form IV fragment:
[0277] Formula (IV)
[0278] With the fragments of equation VII-1 or VII-2:
[0279]
[0280] X1, X2, X3, X4, X5, X6, X7, Y1, Y2, Y3, n, m, and W are defined as above; and
[0281] G is -F, -Cl, -Br, -OH, -OCH3, or -OCH2CH3;
[0282] Contact occurs in the presence of an alkaline substance.
[0283] In some embodiments, the base is organic (pyridine and / or trimethylamine). In some embodiments, the base is inorganic (e.g., potassium carbonate / sodium carbonate and / or potassium bicarbonate / sodium bicarbonate). In some embodiments, the method further includes, but is not limited to, coupling agents such as DCC and / or EDCI. In some embodiments, the contacting step is performed purely. In some embodiments, the contacting step is performed in the presence of a solvent. Exemplary solvents include, but are not limited to, THF, DCM, ACN, and / or DMSO. In some embodiments, the contacting step is performed at room temperature. In some embodiments, the contacting step is performed at an elevated temperature.
[0284] F. Composition
[0285] Currently disclosed compounds can be formulated into pharmaceutical compositions together with pharmaceutically acceptable carriers.
[0286] The compounds disclosed herein can be formulated into pharmaceutical compositions according to standard pharmaceutical practice. In this respect, a pharmaceutical composition is provided comprising the disclosed compounds associated with a pharmaceutically acceptable diluent or carrier.
[0287] Typical formulations are prepared by mixing the compounds disclosed herein with a carrier, diluent, or excipient. Suitable carriers, diluents, and excipients are well known to those skilled in the art and include materials such as carbohydrates, waxes, water-soluble and / or swellable polymers, hydrophilic or hydrophobic materials, gelatin, oils, solvents, water, etc. The specific carrier, diluent, or excipient used will depend on the application and intended use of the compound. Solvents are generally selected based on those who consider them to be generally safe (GRAS) for use in mammals. Generally, safe solvents are non-toxic aqueous solvents, such as water and other non-toxic solvents that are soluble or miscible with water. Suitable aqueous solvents include water, ethanol, propylene glycol, polyethylene glycol (e.g., PEG 400, PEG 300), and mixtures thereof. The formulation may also include one or more buffers, stabilizers, surfactants, wetting agents, lubricants, emulsifiers, suspending agents, preservatives, antioxidants, light-blocking agents, flow aids, processing aids, colorants, sweeteners, flavoring agents, and other known additives that provide an elegant presentation of the pharmaceutical product (i.e., the compound disclosed herein or a pharmaceutical composition thereof) or aid in the manufacture of the pharmaceutical product (i.e., the drug).
[0288] Formulations can be prepared using conventional dissolution and mixing procedures. For example, the bulk pharmaceutical substance (i.e., the compound disclosed herein or a stable form of the compound (e.g., a complex with a cyclodextrin derivative or other known complexing agent)) is dissolved in a suitable solvent in the presence of one or more of the excipients described above. The compound is typically formulated into a pharmaceutical dosage form to provide easily controlled drug dosage and enable patients to adhere to prescribed regimens.
[0289] Pharmaceutical compositions (or formulations) for application can be packaged in various ways depending on the method of administration. Typically, articles for distribution comprise containers in which the pharmaceutical formulation is already placed in a suitable form. Suitable containers are well known to those skilled in the art and include materials such as bottles (plastic and glass), capsules, ampoules, plastic bags, metal tubes, etc. Containers may also include tamper-evident components to prevent easy access to the contents. Additionally, containers have labels placed thereon describing the contents of the container. The labels may also contain appropriate warnings.
[0290] Pharmaceutical formulations can be prepared for administration via various routes and types. For example, compounds of desired purity as disclosed herein can optionally be mixed with pharmaceutically acceptable diluents, carriers, excipients, or stabilizers (Remington's Pharmaceutical Sciences (1980), 16th edition, Osol, A. ed.) in the form of lyophilized formulations, ground powders, or aqueous solutions. Formulations can be prepared by mixing with a physiologically acceptable carrier—one that is non-toxic to the recipient at the dose and concentration used—at ambient temperature, at an appropriate pH, and to the desired purity. The pH of the formulation depends primarily on the specific application and the concentration of the compound, but can range from about 3 to about 8. Formulations in an acetate buffer solution at pH 5 are a suitable example.
[0291] The compound may be sterile. Specifically, formulations intended for in vivo administration should be sterile. Such sterility can be easily achieved by means of filtration using a sterile filter membrane.
[0292] The compounds are typically stored as solid compositions, lyophilized formulations, or aqueous solutions.
[0293] The formulation, administration, and route of delivery of pharmaceutical compositions including the compounds disclosed herein—that is, the amount, concentration, schedule, duration of treatment, mediator, and route of administration—can be consistent with good medical practice. Factors considered in this context include the specific condition being treated, the specific mammal being treated, the individual patient's clinical symptoms, the cause of the condition, the site of delivery, the method of administration, the schedule of administration, and other factors known to the practicing physician. The "therapeuticly effective amount" of the compound to be administered will depend on such considerations and is the minimum amount necessary to prevent, improve, or treat clotting factor-mediated conditions. This amount is preferably below amounts that would be toxic to the host or make the host significantly more prone to bleeding.
[0294] Acceptable diluents, carriers, excipients, and stabilizers are non-toxic to recipients at the doses and concentrations used and contain buffers such as phosphates, citrates, and other organic acids; antioxidants containing ascorbic acid and methionine; preservatives (such as octadecyl dimethyl benzyl ammonium chloride; hexahydroquinone quaternary ammonium chloride; benzyl alkyl ammonium chloride, benzyl chloride; phenolic alcohols, butanol, or benzyl alcohol; alkyl parabens such as methyl or propyl parabens; catechol; resorcinol; cyclohexanol; 3-pentanol; and m-cresol); low molecular weight (less than) Polypeptides (approximately 10 residues); proteins such as serum albumin, gelatin, or immunoglobulins; hydrophilic polymers such as polyvinylpyrrolidone; amino acids such as glycine, glutamine, asparagine, histidine, arginine, or lysine; monosaccharides, disaccharides, and other carbohydrates containing glucose, mannose, or dextrin; chelating agents such as EDTA; sugars such as sucrose, mannitol, trehalose, or sorbitol; salt-forming counterions such as sodium; metal complexes (e.g., zinc-protein complexes); and / or nonionic surfactants such as TWEEN. TM PLURONICS TM Or polyethylene glycol (PEG). The active pharmaceutical ingredient can also be encapsulated in microcapsules, for example, prepared by coagulation techniques or interfacial polymerization, such as hydroxymethyl cellulose or gelatin microcapsules and poly(methyl methacrylate) microcapsules in colloidal drug delivery systems (e.g., liposomes, albumin microspheres, microemulsions, nanoparticles, and nanocapsules) or crude emulsions. Such techniques are disclosed in Remington Pharmaceutical Sciences, 16th edition, Osol, A. (1980).
[0295] Sustained-release formulations of the compounds can be prepared. Suitable examples of sustained-release formulations comprise a semi-permeable matrix containing a solid hydrophobic polymer of the compounds disclosed herein, said matrix being in the form of a molded article, such as a membrane or microcapsule. Examples of sustained-release matrices include polyesters, hydrogels (e.g., poly(2-hydroxyethyl-methacrylate), or poly(vinyl alcohol)), polylactides (US Patent No. 3,773,919), copolymers of L-glutamic acid and γ-ethyl-L-glutamate, non-degradable ethylene-vinyl acetate, and degradable lactic-glycolic acid copolymers, such as LUPRON DEPOT. TM (Injectable microspheres composed of lactic acid-glycolic acid copolymer and leuprolide acetate) and poly-D-(-)-3-hydroxybutyric acid.
[0296] The formulations comprise those suitable for the routes of administration described herein. The formulations can be conveniently presented in unit dosage forms and can be prepared by any method well-known in the pharmaceutical industry. Techniques and formulations are typically found in Remington Pharmaceutical Sciences (Mack Publishing Co., Easton, PA). Such methods involve the step of associating the active ingredient with a carrier constituting one or more excipients. Generally, the formulations are prepared by homogeneously and sufficiently associating the active ingredient with a liquid carrier or a fine solid carrier, or both, and then shaping the product if necessary.
[0297] Formulations of compounds suitable for oral administration as disclosed herein can be prepared in discrete units, such as tablets, capsules, sachets, or tablets each containing a predetermined amount of the compound.
[0298] Compressed tablets can be prepared by pressing an active ingredient (such as powder or granules) in a free-flowing form, optionally mixed with a binder, lubricant, inert diluent, preservative, surfactant, or dispersant, in a suitable machine. Molded tablets can be manufactured by molding a mixture of powdered active ingredients moistened with an inert liquid diluent in a suitable machine. Tablets can optionally be coated or scored, and optionally formulated to provide a slow or controlled release of the active ingredient.
[0299] Tablets, sugar lozenges, tablets, aqueous or oily suspensions, dispersible powders or granules, emulsions, hard capsules or soft capsules (e.g., gelatin capsules), syrups or elixirs can be prepared for oral use. Formulations of compounds intended for oral use as disclosed herein can be prepared according to any method known in the art for manufacturing pharmaceutical compositions, and such compositions can contain one or more pharmaceutical agents comprising sweeteners, flavoring agents, coloring agents, and preservatives to provide a palatable formulation. Tablets containing an active ingredient mixed with non-toxic, pharmaceutically acceptable excipients suitable for manufacturing tablets are acceptable. For example, these excipients can be inert diluents such as calcium carbonate or sodium carbonate, lactose, calcium phosphate or sodium phosphate; granulating and disintegrants such as corn starch or alginate; binding agents such as starch, gelatin, or gum arabic; and lubricants such as magnesium stearate, stearic acid, or talc. Tablets may be uncoated or can be coated using known techniques including microencapsulation to delay disintegration and adsorption in the gastrointestinal tract, and thus provide sustained action over a longer period of time. For example, delaying materials, such as glyceryl monostearate or glyceryl distearate, can be used alone or in combination with wax.
[0300] For the treatment of the eyes or other external tissues, such as the mouth and skin, the formulation can be applied as a topical ointment or cream containing, for example, 0.075 to 20% w / w of the active ingredient. When formulated into an ointment, the active ingredient can be used with a paraffin-based ointment base or a water-miscible ointment base. Alternatively, the active ingredient can be formulated into a cream using an oil-in-water emulsion base.
[0301] If desired, the aqueous phase of the cream matrix may contain polyols, i.e., alcohols having two or more hydroxyl groups, such as propylene glycol, 1,3-butanediol, mannitol, sorbitol, glycerin, and polyethylene glycol (containing PEG 400), as well as mixtures thereof. These topical formulations may desirably contain compounds that enhance the absorption or penetration of the active ingredient through the skin or other affected areas. Examples of such transdermal penetration enhancers include dimethyl sulfoxide and related analogues.
[0302] The oil phase of an emulsion can be composed of known ingredients in a known manner. While the phase may consist only of emulsifiers, it may also include at least one emulsifier with fats or oils, or a mixture of both. Hydrophilic emulsifiers included along with lipophilic emulsifiers can be used as stabilizers. Emulsifiers, with or without stabilizers, together constitute a so-called emulsified wax, and the wax, together with the oils and fats, constitutes a so-called emulsified ointment matrix, which forms the oily dispersed phase of the cream formulation. Suitable emulsifiers and emulsion stabilizers for use in formulations include... 60. 80. Cetearyl alcohol, benzyl alcohol, myristol, glyceryl monostearate, and sodium lauryl sulfate.
[0303] The aqueous suspension of the compound contains an active material mixed with excipients suitable for manufacturing the aqueous suspension. Such excipients include: suspending agents such as sodium carboxymethyl cellulose, croscarmellose, povidone, methylcellulose, hydroxypropyl methylcellulose, sodium alginate, polyvinylpyrrolidone, gum arabic, and gum arabic; and dispersing or wetting agents such as naturally occurring phospholipids (e.g., lecithin), condensation products of alkyl esters and fatty acids (e.g., polyoxyethylene stearate), condensation products of ethylene oxide and long-chain fatty alcohols (e.g., heptadecanethoxycetyl alcohol), and condensation products of ethylene oxide and metaesters derived from fatty acids and hexadiene anhydrides (e.g., polyoxyethylene sorbitan monooleate). The aqueous suspension may also contain: one or more preservatives, such as ethylparaben or n-propylparaben; one or more colorants; one or more flavoring agents; and one or more sweeteners, such as sucrose or saccharin.
[0304] The pharmaceutical composition of the compound can be in the form of a sterile injectable formulation, such as a sterile injectable aqueous or oily suspension. This suspension can be formulated using suitable dispersants or wetting agents and suspending agents mentioned above, according to known techniques. The sterile injectable formulation can also be a sterile injectable solution or suspension in a non-toxic, parenteral-acceptable diluent or solvent, such as 1,3-butanediol. The sterile injectable formulation can also be prepared as a lyophilized powder. Acceptable mediators and solvents that can be used are water, Ringer's solution, and isotonic sodium chloride solution. Additionally, sterile non-volatile oils can routinely be used as solvents or suspension media. For this purpose, any mild non-volatile oil, including synthetic monoglycerides or diglycerides, can be used. Furthermore, fatty acids such as oleic acid can also be used in the preparation of injectable formulations.
[0305] The amount of active ingredient that can be combined with a carrier material to produce a single dosage form will vary depending on the subject being treated and the specific mode of administration. For example, a sustained-release formulation for oral administration to humans may contain approximately 1 to 1000 mg of active material compounded with an appropriate and convenient amount of carrier material, said appropriate and convenient amount may vary from approximately 5% to approximately 95% (weight:weight) of the total composition. Pharmaceutical compositions can be prepared to provide an easily measurable amount for administration. For example, an aqueous solution intended for intravenous infusion may contain approximately 1 μg to 500 μg of active ingredient per milliliter of solution, with the aim of producing a suitable volume that can be infused at a rate of approximately 10 mL / h to approximately 50 mL / h.
[0306] Formulas suitable for parenteral administration include: aqueous and non-aqueous sterile injectable solutions, which may contain antioxidants, buffers, antibacterial agents, and solutes that make the formula isotonic with the blood of the intended recipient; and aqueous and non-aqueous sterile suspensions, which may contain suspending agents and thickeners.
[0307] Formulations suitable for topical application to the eyes also include eye drops, wherein the active ingredient is dissolved or suspended in a suitable carrier, particularly in an aqueous solvent of the active ingredient. The active ingredient is preferably present in such formulations at a concentration of about 0.5 to 20% w / w, for example about 0.5 to 10% w / w, for example about 1.5% w / w.
[0308] Formulations suitable for topical application in the mouth include: tablets comprising an active ingredient in a flavoring matrix (typically sucrose and gum arabic or tragacanth); soft tablets comprising an active ingredient in an inert matrix, such as gelatin and glycerin, or sucrose and gum arabic; and mouthwashes comprising the active ingredient in a suitable liquid carrier.
[0309] Formulations for rectal administration may be presented in suppository form with a suitable base (including, for example, cocoa butter or salicylates).
[0310] Formulations suitable for intrapulmonary or nasal administration have a particle size, for example, in the range of 0.1 to 500 micrometers (in increments of 0.1 to 500 micrometers, such as 0.5, 1, 30, 35 micrometers, etc.), and are administered by rapid inhalation through the nasal passage or orally to reach the alveolar sacs. Suitable formulations comprise aqueous or oily solutions of the active ingredient. Formulations suitable for aerosol or dry powder administration can be prepared according to conventional methods and can be delivered together with other therapeutic agents, such as compounds used to date for the treatment or prevention of the conditions described below.
[0311] Formulations suitable for vaginal application may be in the form of vaginal suppositories, tampons, creams, gels, pastes, foams or sprays, which, in addition to the active ingredient, contain a suitable carrier as known in the art.
[0312] The formulations may be packaged in single-dose or multi-dose containers (e.g., sealed ampoules and vials) and may be stored under lyophilized (freeze-dried) conditions where a sterile liquid carrier (e.g., water for injection) is added only just before use. Ready-to-use injectable solutions and suspensions are prepared from the types of sterile powders, granules, and tablets described above. Preferred single-dose formulations are those containing the active ingredient at the daily dose or sub-daily dose as described above, or a suitable portion thereof.
[0313] The subject matter further provides veterinary compositions comprising at least one active ingredient as defined above and a veterinary carrier therefor. The veterinary carrier is a material suitable for the purpose of administering the composition and may be a solid, liquid, or gaseous material that is originally inert or acceptable in the veterinary field and is compatible with the active ingredient. These veterinary compositions may be administered parenterally, orally, or via any other desired route.
[0314] In certain embodiments, the pharmaceutical composition comprising the currently disclosed compounds further comprises a chemotherapeutic agent. In some of these embodiments, the chemotherapeutic agent is an immunotherapeutic agent.
[0315] G. Treatment methods
[0316] The compounds and compositions disclosed herein can also be used in methods for treating various diseases and / or conditions that have been identified as being related to dysfunction of RNA expression and / or function, or to the expression and / or function of proteins produced from mRNA, or to the useful role of switching RNA conformations using small molecules, or to the natural function of altering riboswitches as a way to inhibit the growth of infectious organisms. Thus, the methods of this disclosure relate to treating diseases or conditions associated with dysfunction of RNA expression and / or function, or to creating novel switchable therapies. See, for example, U.S. Patent Application Publication No. 2018 / 010146, which is hereby incorporated by reference in its entirety. Thus, in some embodiments, a method for treating a disease or condition disclosed herein (e.g., associated with dysfunction of RNA expression and / or function) comprises administering a therapeutically effective dose of the compounds and / or compositions disclosed herein to a subject in need.
[0317] Disorders of RNA expression are characterized by overexpression or underexpression of one or more RNA molecules. In some embodiments, the one or more RNA molecules are associated with promoting a disease and / or condition to be treated. In some embodiments, the RNA molecules are characterized as part of healthy cellular mechanisms and thus can prevent and / or improve the disease and / or condition to be treated. In some embodiments, the disease or condition to be treated is associated with dysfunction of RNA function related to transcription, processing, and / or translation. In some embodiments, the disease or condition to be treated is associated with inaccurate protein expression due to dysfunction of RNA molecule function. In some embodiments, the disease or condition to be treated is associated with dysfunction of RNA function related to gene expression. In some embodiments, the disease or condition is one in which it is desired to reduce protein expression by binding a molecule to mRNA. In some embodiments, the disease is advantageously treatable by a therapy that can be used to turn small molecules on or off. For example, in some embodiments, the disease or condition is a hereditary disease in which the ability to turn the expression of a therapeutic gene on or off is desired.
[0318] The diseases and conditions to be treated include, but are not limited to, degenerative diseases, cancer, diabetes, autoimmune diseases, cardiovascular diseases, coagulation disorders, eye diseases, infectious diseases, and diseases caused by mutations in one or more genes.
[0319] Typical degenerative diseases include, but are not limited to, Alzheimer's disease (AD), amyotrophic lateral sclerosis (ALS, Lou Gehrig's disease), cancer, Charcot-Marie-Tooth disease (CMT), chronic traumatic encephalopathy, cystic fibrosis, some cytochrome c oxidase deficiencies (often the cause of degenerative Leigh syndrome), Ehlers-Danlos syndrome, progressive osteosclerotic fibrosis, Friedreich's ataxia, frontotemporal dementia (FTD), some cardiovascular diseases (e.g., atherosclerotic cardiovascular diseases such as coronary artery disease, aortic stenosis, etc.), and Huntington's disease. Diseases including: infantile axonal dystrophy, keratoconus (KC), keratoglossia, leukodystrophy, macular degeneration (AMD), Marfan syndrome (MFS), some mitochondrial myopathy, mitochondrial DNA depletion syndrome, multiple sclerosis (MS), multiple system atrophy, muscular dystrophy (MD), neuronal ceroid lipofuscin deposition disease, Niemann-Pick disease, osteoarthritis, osteoporosis, Parkinson's disease, pulmonary hypertension, all prion diseases (Creutzfeldt-Jakob disease, fatal familial insomnia, etc.), progressive supranuclear palsy, retinitis pigmentosa (RP), rheumatoid arthritis, Sandhoff disease, spinal muscular atrophy (SMA, motor neuron disease), subacute sclerosing panencephalitis, and Tay-Sachs disease. (disease) and vascular dementia (which may not be neurodegenerative dementia itself, but often occurs with other forms of degenerative dementia).
[0320] Exemplary cancers include, but are not limited to, all forms of cancer, melanoma, blastoma, sarcoma, lymphoma, and leukemia, including, but not limited to, bladder cancer, bladder carcinoma, brain tumors, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, endometrial cancer, hepatocellular carcinoma, laryngeal cancer, lung cancer, osteosarcoma, ovarian cancer, pancreatic cancer, prostate cancer, kidney cancer, and thyroid cancer, acute lymphoblastic leukemia, acute myeloid leukemia, ependymoma, Ewing's sarcoma, glioblastoma, medulloblastoma, neuroblastoma, osteosarcoma, rhabdomyosarcoma, rhabdomyosarcoma, and nephroblastoma (Wilm's tumor).
[0321] Exemplary autoimmune diseases include, but are not limited to, Adult Still's disease, immunoglobulin deficiency, alopecia areata, amyloidosis, ankylosing spondylitis, anti-GBM / anti-TBM nephritis, antiphospholipid syndrome, autoimmune angioedema, autoimmune familial autonomic dysfunction, autoimmune encephalomyelitis, autoimmune hepatitis, autoimmune inner ear disease (AIED), autoimmune myocarditis, autoimmune oophoritis, autoimmune orchitis, autoimmune pancreatitis, autoimmune retinopathy, autoimmune urticaria, axonal and neuronal neuropathy (AMAN), Baló disease, Behcet's disease, benign mucosal pemphigoid, bullous pemphigoid, and Castleman's disease. Disease (CD), celiac disease, Chagas disease, chronic inflammatory demyelinating polyneuropathy (CIDP), chronic relapsing multifocal osteomyelitis (CRMO), allergic granulomatous syndrome (CSS) or eosinophilic granulomatosis (EGPA), pemphigus cicatricialis, Cogan's syndrome, cold agglutinin disease, congenital heart block, Coxsackie myocarditis, CREST syndrome, Crohn's disease, herpetic dermatitis, dermatomyositis, Devic's disease (neuromyelitis optica), discoid lupus, Dressler's syndrome, endometriosis, eosinophilic esophagitis (EoE), eosinophilic fasciitis, erythema nodosum, mixed cryoglobulinemia, Evans syndrome Syndrome, fibromyalgia, fibrotic alveolitis, giant cell arteritis (temporal arteritis), giant cell myocarditis, glomerulonephritis, Goodpasture's syndrome, granulomatous polyangiitis, Graves' disease, Guillain-Barré syndrome, Hashimoto's thyroiditis, hemolytic anemia, Henoch-Schonlein purpura.Herpes gestationis (HSP), pemphigoid gestationis (PG), hidradenitis suppurativa (HS) (acne), hypogammaglobulinemia, IgA nephropathy, IgG4-related sclerosis, immune thrombocytopenic purpura (ITP), inclusion body myositis (IBM), interstitial cystitis (IC), juvenile arthritis, juvenile diabetes mellitus (type 1 diabetes), juvenile myositis (JM), Kawasaki disease, Lambert-Eaton syndrome, leukocytoclastic vasculitis, lichen planus, lichen sclerosus, woody conjunctivitis, linear IgA disease (LAD), lupus, chronic Lyme disease, Meniere's disease, microscopic polyangiitis (MPA), mixed connective tissue disease (MCTD), mooren's ulcer, Mucha-Habermann disease. Diseases, multifocal motor neuropathy (MMN) or MMNCB, multiple sclerosis, myasthenia gravis, myositis, narcolepsy, neonatal lupus, neuromyelitis optica, neutropenia, ocular cicatricial pemphigus, optic neuritis, relapsing rheumatism (PR), PANDAS, paraneoplastic cerebellar degeneration (PCD), paroxysmal nocturnal hemoglobinuria (PNH), Parry Romberg syndrome, periclival plaque inflammation (peripheral uveitis), Parsonage-Turner syndrome. Pemphigus, peripheral neuropathy, perivenous encephalomyelitis, pernicious anemia (PA), POEMS syndrome, polyarteritis nodosa, polyglandular syndrome (types I, II, and III), polymyalgia rheumatica, polymyositis, post-myocardial infarction syndrome, post-pericardiotomy syndrome, primary biliary cirrhosis, primary sclerosing cholangitis, progesterone dermatitis, psoriasis, psoriatic arthritis, pure red cell aplasia (PRCA), pyoderma gangrene, Raynaud's phenomenon, reactive arthritis, reflex sympathetic dystrophy, relapsing polychondritis, restless legs syndrome (RLS), retroperitoneal fibrosis, rheumatic fever, rheumatoid arthritis, sarcoidosis, Schmidt syndrome, scleroderma, Sjögren's syndrome Sperm and testicular autoimmunity, stiff-person syndrome (SPS), subacute bacterial endocarditis (SBE), Susac's syndrome, sympathetic ophthalmia (SO), Takayasu's arteritis, temporal arteritis / giant cell arteritis, thrombocytopenic purpura (TTP), painful ophthalmoplegia syndrome (THS), transverse myelitis, type 1 diabetes, ulcerative colitis (UC), undifferentiated connective tissue disease (UCTD), uveitis, vasculitis, vitiligo, and Vogt-Koyanagi-Harada disease.
[0322] Exemplary cardiovascular conditions include, but are not limited to, coronary artery disease (CAD), angina pectoris, myocardial infarction, stroke, heart attack, heart failure, hypertensive heart disease, rheumatic heart disease, cardiomyopathy, arrhythmia, congenital heart disease, valvular heart disease, carditis, aortic aneurysm, peripheral artery disease, thromboembolic disease, and venous thrombosis.
[0323] Exemplary coagulation disorders include, but are not limited to, hemophilia, von Willebrand disease, disseminated intravascular coagulation, liver disease, excessive development of circulating anticoagulants, vitamin K deficiency, platelet dysfunction, and other coagulation defects.
[0324] Exemplary eye diseases include, but are not limited to, macular degeneration, exophthalmos, cataracts, cytomegalovirus retinitis (CMV retinitis), diabetic macular edema, glaucoma, keratoconus, ocular hypertension, ocular migraine, retinoblastoma, subconjunctival hemorrhage, pterygium, keratitis, dry eye, and corneal abrasion.
[0325] Exemplary infectious diseases include, but are not limited to, acute flaccid myelitis (AFM), anaplasmosis, anthrax, babesiosis, botulism, brucellosis, campylobacteriosis, carbapenem-resistant infection (CRE / CRPA), chancroid, chikungunya virus infection (chikungunya fever), chlamydia, fish botulism (harmful algal bloom (HAB)), Clostridium difficile infection, Clostridium perfringens (ε-toxin), coccidioidomycosis (valley fever), COVID-19 (coronavirus disease 2019), and Creutzfeldt-Jacob disease. Diseases including infectious spongiform encephalopathy (CJD), cryptosporidiosis (latent), cyclosporidiosis, dengue fever, dengue 1, 2, 3, 4, diphtheria, Escherichia coli infection, Shiga toxin-producing (STEC), eastern equine encephalitis (EEE), Ebola hemorrhagic fever (Ebola), Ehrlich disease, encephalitis, arbovirus or similar infections, enterovirus infection, non-poliomyelitis (non-poliomyelitis enterovirus), enterovirus infection, D68 (EV-D68), giardiasis (giardia), glanders, gonococcal infection (gonorrhea), granuloma inguinale, Haemophilus influenzae type b (Hib or H-influenza), Hantavirus pulmonary syndrome (HPS), hemolytic uremic syndrome (HUS), hepatitis A (Hep A), hepatitis B (Hep B), hepatitis C (Hep C), hepatitis D (Hep D), hepatitis E (Hep A), hepatitis B (Hep B), hepatitis C (Hep C), hepatitis D (Hep D), hepatitis E (Hep A), hepatitis B (Hep B), hepatitis C (Hep C), hepatitis D (Hep D), hepatitis E (Hep D), hepatitis B ... E), herpes, shingles, VZV (herpes zoster), histoplasmosis, human immunodeficiency virus / AIDS (HIV / AIDS), human papillomavirus (HPV), influenza, lead poisoning, Legionnaires' disease, leprosy (Hansens disease), leptospirosis, listeriosis (Listeria), Lyme disease Diseases, lymphogranuloma venereum (LGV), malaria, measles, glanders, meningitis, viral (meningitis, viral), meningococcal disease, bacterial (meningitis, bacterial), Middle East Respiratory Syndrome Coronavirus (MERS-CoV), mumps, norovirus, paralytic shellfish poisoning (paralytic shellfish poisoning, botulism), lice infestation (lice, head lice and body lice), pelvic inflammatory disease (PID), pertussis (whooping cough), plague; inguinal gland inflammation, septicemia, pneumonia (plague), pneumococcal disease (pneumonia), poliomyelitis (poliomyelitis), Powassan, psittacosis (psittacosis), lice infestation (crabs);Pubic lice infection), impetigo (smallpox, monkeypox, cowpox), Q fever, rabies, ricin poisoning, rickettsial disease (Rocky Mountain spotted fever), rubella, including congenital (German measles), salmonellosis (Salmonella), scabies infection, scabies toxin, septic shock, severe acute respiratory syndrome (SARS), Shigella infection, smallpox, staphylococcal infection, methicillin-resistant septicemia (MRSA), staphylococcal food poisoning, enterotoxin-B poisoning (staphylococcal food poisoning), staphylococcal infection, vancomycin intermediate (VISA), staphylococcal infection, vancomycin-resistant septicemia (VRSA), group A (invasive) streptococcal infection (streptococcal A (invasive)), group B Streptococcal disease (Streptococcus-B), Streptococcal toxic shock syndrome (STSS), Toxic shock (STSS, TSS), Syphilis (primary, secondary, early latent, late latent, congenital), Tetanus infection, Tetanus (trismus), Trichomoniasis, Trichinosis, Tuberculosis (TB), Tuberculosis (latent) (LTBI), Tularemia, Typhus (Group D), Spotted typhus, Vaginal disease, Bacterial (yeast infection), Vascular pulmonary injury (e-cigarette-associated lung injury), Varicella (Chickenpox), Vibrio cholerae (Cholera), Vibrio infection, Viral hemorrhagic fever (Ebola, Lassa, Marburg), West Nile virus (West Nile virus) Yellow fever, Yersinia (Yersinia), and Zika virus infection.
[0326] Example
[0327] Example 1 RNA construct design and preparation
[0328] The screening construct was designed to allow the incorporation of various internal target RNA motifs, one or more. Two motifs are present in the construct: a TPP riboswitching domain. 27 And pseudoknots from the 5'-UTR of dengue virus 26 The design of the complete construct sequence, comprising a cassette, an RNA barcode helix, and two test RNA structures (separated by a hexanucleotide linker), was evaluated using RNA structure evaluation. 39To reduce the possibility of interaction between the two test structures, minor sequence alterations were made to prevent misfolded structures predicted by the RNA structure while preserving the natural folds (Figure 7). The structure of the final construct was confirmed by SHAPE-MaP.
[0329] RNA barcodes were designed to fold into individual hairpins (Figure 7). All possible permutations of the RNA barcodes were calculated and folded against the background of the complete construct sequence, and any barcodes that might interact with another part of the RNA construct were removed from the set. The barcode-encoded construct was probed using a ligand-free scheme via SHAPE-MaP and folded using RNA structures with SHAPE reactivity constraints to confirm that the barcodes helically folded into the desired individual hairpins.
[0330] RNA preparation
[0331] The DNA template used for in vitro transcription (Integrated DNA Technologies) encodes the target construct sequence (containing the dengue pseudoknot sequence, single linker, and TPP riboswitcher sequence) and the side-connected cassette. 25 :5'- GTGGG CACTT CGGTG TC CAC ACGCG AAGGA AACCG CGTGT CAACT GTGCA ACAGC TGACA AAGAGATTCC TAAAA CTCAG TACTC GGGGT GCCCT TCTGC GTGAA GGCTG AGAAA TACCC GTATC ACCTGATCTG GATAATGCCA GCGTA GGGAA GTGCT GGATC CGGTT CGCCG GATCA AT CGG GCTTC GGTCC GGTTC -3' (SEQ ID NO:1). The primer binding site is underlined. RNA barcodes were added to each of the 96 constructs individually using forward PCR primers containing unique RNA barcodes and T7 promoter sequences in separate PCR reactions. The sample forward primer sequences with barcode nucleotides shown in bold and primer binding sites underlined are:
[0332] DNA was amplified by PCR using a 200 μM dNTP mixture (New England Biolabs), 500 nM forward primers, 500 nM reverse primers, 1 ng DNA template, 20% (v / v) Q5 reaction buffer, and 0.02 U / μL Q5 hot-start high-fidelity polymerase (New England Biolabs) to create a template for in vitro transcription. The DNA was purified (PureLink Pro 96 PCR Purification Kit; Invitrogen) and quantified using a TecanInfinite M1000 Pro microplate reader (Quant-iT dsDNA High Sensitivity Assay Kit; Invitrogen).
[0333] In vitro transcription was performed in 96-well plates, with each well containing a total reaction volume of 100 μL. Each well contained 5 mM NTP (New England Biolabs), 0.02 U / μL inorganic pyrophosphatase (yeast, New England Biolabs), 25 mM MgCl2 containing 0.05 mg / mL T7 polymerase, 40 mM Tris, pH 8.0, 2.5 mM spermidine, 0.01% Triton, 10 mM DTT, and 200–800 nM unique bar-coding DNA template (generated by PCR). The reaction was incubated at 37 °C for 4 h; then treated with TurboDNase (RNase-free, Ingenium) at a final concentration of 0.04 U / μL; incubated at 37 °C for 30 min; then a second addition of DNase was made to a total final concentration of 0.08 U / μL, and incubated at 37 °C for an additional 30 min. The enzymatic reaction was stopped by adding EDTA to a final concentration of 50 mM and placing the plate on ice. RNA was purified in 96-well batches (Agencourt RNAclean XP magnetic beads; Beckman Coulter) and resuspended in 10 mM Tris pH 8.0 and 1 mM EDTA. RNA concentration was quantified using a Tecan Infinite M1000 Pro microplate reader (Quant-iT RNA Assay Kit; Ingenium Biotech), and each well was individually diluted to 1 pmol / μL. RNA was stored at -80°C.
[0334] Example 2: Chemical modification and screening of small molecule fragments
[0335] Fragments were obtained from Maybridge in the form of a fragment screening library, a subset of its Ro3 diversity fragment library, containing 1500 compounds dissolved in 50 mM DMSO. Most of these compounds adhered to the “rule of three” for fragment compounds: a molecular weight <300 Da, ≤3 hydrogen bond donors and ≤3 hydrogen bond acceptors, and ClogP ≤3.0. All compounds used for ITC, except those listed in Example 5, were purchased from Millipore-Sigma and used without further purification. Screening experiments were performed in 25 μL of 96-well plates on a Tecan Freedom Evo-150 liquid handling system equipped with an 8-channel air-displacement pipette arm, disposable filter tips, a robotic manipulator arm, and an EchoTherm RIC20 remote-controlled heated / cooled dry bath (Torrey Pines Scientific). Liquid handling system programs for screening were available upon request.
[0336] For the first fragment ligand screening, 5 pmol of RNA per well was diluted to 19.6 μL in RNase-free water on a 4°C cooling block. The plate was then heated at 95°C for 2 minutes, followed by immediate rapid cooling at 4°C for 5 minutes. 19.6 μL of 2× folding buffer (final concentration 50 mM, HEPES pH 8.0, 200 mM potassium acetate, and 10 mM MgCl2) was added to each well, and the plate was incubated at 37°C for 30 minutes. For the second fragment ligand screening, 24.3 μL of folded RNA per well was added to 2.7 μL of DMSO containing the primary binding fragment, to a final concentration of 10×K. d The final concentration was determined, and the sample was incubated at 37°C for 10 minutes. To combine the target RNA with the fragment, 24.3 μL of RNA solution or RNA plus the primary binding fragment was added to the well containing 2.7 μL of 10× screening fragment (in DMSO to produce a final fragment concentration of 1 mM). The solution was thoroughly mixed by pipetting and incubated at 37°C for 10 minutes. For SHAPE detection, 22.5 μL of RNA-fragment solution from each well of the screening plate was added to DMSO containing 2.5 μL of 10× SHAPE reagent on a 37°C heating block, and rapidly mixed by pipetting to achieve a uniform distribution of SHAPE reagent and RNA. After the appropriate reaction time, the sample was placed on ice. For the first fragment screening, 1-methyl-7-nitroindosanhydride (1M7) was used as the SHAPE reagent at a final concentration of 10 mM, and the reaction was carried out for 5 minutes. For the second fragment screening, 5-nitroindosanhydride (5NIA) was used. 40As the SHAPE reagent, the final concentration was 25 mM, and the reaction time was 15 minutes. Excess fragments, solvents, and hydrolyzed SHAPE reagent were removed using AutoScreen-A 96-well plates (GE Healthcare Life Sciences), and 5 μL of modified RNA from each well of the 96-well plate was pooled into a single sample for sequencing library preparation.
[0337] Each screening consists of a 19-fragment test plate, two plates containing distributions of positive (fragment 2, final concentration 1 mM) and negative (solvent, DMSO) controls, and a negative SHAPE control plate treated with solvent (DMSO) instead of SHAPE reagent. For hit validation experiments, the well positions for each hit fragment are altered to control well placement and RNA barcoding effects. Plate diagrams for primary and secondary screening are also available.
[0338] Once the test fragments have been screened, statistical tests are performed to identify differences in the modification rates of a given nucleotide. Specifically, the screening analysis requires a statistical comparison of the modification rates of a given nucleotide in the presence and absence of the fragment. For each nucleotide, the number of modifications in a given reaction is a Poisson process with known variance; therefore, the statistical significance of the difference in modification rates between two observed samples can be determined by performing two Poisson count comparison tests. 31 In other words, if m1 modifications of the tested nucleotides are counted in n1 reads of sample 1 and m2 modifications are counted in n2 reads of sample 2, then the null hypothesis predicts that, of all counted modifications (m1+m2), the proportion of modifications in sample 1 will be p1 = n1 / (n1+n2). The Z-test for this hypothesis is:
[0339]
[0340] Z = min(|Z p |,|Z n |)
[0341] If the Z-value exceeds the specified significance threshold, the tested nucleotide is considered to be statistically significant due to the presence of the test fragment.
[0342] Next, for each fragment, a Z-test must be performed on a large number of nucleotides including the RNA sequence, increasing the likelihood of false positives. While the number of false positive assignments for the SHAPE reactivity of each nucleotide can be minimized by increasing the Z-significance threshold, this approach reduces the sensitivity of the screening (meaning it reduces the ability to detect weaker binding ligands). To reduce the number of Z-tests performed, such tests are applied only to nucleotides in the region of interest, rather than all nucleotides in the RNA screening construct. For the dengue motif of RNA, the region of interest is positions 59–110; for the TPP motif, the region of interest is positions 100–199. The number of Z-tests is further reduced by omitting nucleotides with low modification rates in both samples. The threshold for considering nucleotides with low modification rates is set to 25% of the plate average modification rate, calculated over all nucleotides in all 96 wells of a given plate. The Z-test is performed only on nucleotides whose modification rate exceeds this 25% threshold in at least one of the two comparison samples.
[0343] Ideally, the only difference between the conditions in two comparison samples is the presence of a fragment in one sample but its absence in the other. Cross-testing negative control samples can be used to measure the prevalence of uncontrolled factors that may introduce cross-sample variability in nucleotide modification rates. For example, if the Z-significance threshold is set to 2.7, theoretically, in the absence of any such factors, a Z-test applied to negative control (fragment-free) sample pairs should identify differentially reactive nucleotides with a probability of P = 0.0035. However, when the Z-test was applied to negative control sample pairs randomly selected from 587 negative control samples tested in the primary screening, the actual probability was 90-fold higher, with P = 0.32. Therefore, in the absence of a fragment, there is statistically significant variability in SHAPE reactivity at individual nucleotides.
[0344] While most replicas have substantially the same spectrum, a considerable number of replicas exhibit different spectra; some have coefficients of determination as low as 0.85. Applying the Z-test to different negative control samples yields numerous instances where nucleotides are misclassified as differentially reactive. To avoid this result, each sample is compared to the five most highly correlated negative control samples. Applying the Z-test to such selectivity pairs of negative controls with a Z-significance threshold of 2.7 identifies differentially reactive nucleotides with a probability of P = 0.067.
[0345] This probability is approximately 20 times higher than the theoretical P = 0.0035, indicating variability in sample processing. Some of this variability is scaled equally across the reactivity of all nucleotides in all RNA in the sample. This variability can be removed by reducing the overall reactivity in more reactive samples to match the overall reactivity in less reactive samples. This scaling is performed by: (i) calculating the ratio of the modification rate of each nucleotide in the RNA sequence in the more reactive sample to the modification rate in the less reactive sample, and (ii) dividing the modification rate of all nucleotides in the more reactive sample by the median of the ratio obtained in step (i). Such scaling of the negative control well pairs with maximized reactivity reduces the probability of finding nucleotide hits to P = 0.030, which is 9 times higher than the theoretical probability. Therefore, false positives in fragment identification will occur, as actually occurs in all high-throughput screening assays, and actual fragment hits from non-ligand variants will be distinguished by repeated SHAPE validation and direct ligand binding measurements using ITC.
[0346] Because effective ligands are expected to influence the modification rate of multiple nucleotides in the target RNA, a fragment is considered a hit only if the number of nucleotides with different reactivity from the negative control exceeds a defined threshold set to 2. Secondly, when looking for a relatively robust effect of the fragment on RNA, small relative differences in nucleotide reactivity, even if statistically significant, are excluded from the total count of differentially reactive nucleotides. In practice, the minimum acceptable difference is set at 20% of the mean.
[0347] |r1–r2| / (r1+r2) / 2=0.2,
[0348] Here, r1 and r2 represent the nucleotide modification rates in the two samples. Third, the given sample is tested against five negative control samples that are most highly correlated with it. All five tests must detect changes in the test sample relative to the negative control samples.
[0349] Finally, the sensitivity and specificity of the screening were controlled by selecting a Z-significance threshold. Evaluations of samples containing the fragment and all negative control samples were performed at multiple Z-significance threshold settings. For each such setting, the false positive score (FPF) was calculated as the score of the altered negative control sample found, and the ligand score (LF) was estimated by subtracting the FPF from the score of the altered sample containing the fragment. The balance between LF and FPF was quantified by their ratio LF / FPF. The optimal balance (LF / FPF ≈ 1.3) for TPP riboswitching RNA was achieved at a Z-significance threshold ranging from 2.5 to 2.7, where 0.022 > FPF > 0.014. For dengue pseudoknots, the optimal balance (LF / FPF ≈ 4) was achieved at a Z-significance threshold ranging from 2.5 to 2.65, where 0.007 > FPF > 0.005.
[0350] Example 3: Library preparation and sequencing
[0351] Reverse transcription was performed on 100 μL of pooled modified RNA. 6 μL of reverse transcription primers were added to 71 μL of pooled RNA to achieve a final primer concentration of 150 nM, and the sample was incubated at 65 °C for 5 minutes, then placed on ice. To this solution, 6 μL of 10× first-strand buffer (500 mM Tris pH 8.0, 750 mM KCl), 4 μL of 0.4 M DTT, 8 μL of dNTP mixture (10 mM each), and 15 μL of 500 mM MnCl2 were added, and the solution was incubated at 42 °C for 2 minutes, followed by the addition of 8 μL of SuperScript II reverse transcriptase (Agencourt). The reaction was incubated at 42 °C for 3 hours, then heat-inactivated at 70 °C for 10 minutes, and then placed on ice. The resulting cDNA product was purified (using Agencourt RNAClean beads; Beckman Coulter), eluted to RNase-free water, and stored at -20 °C. The sequence of the reverse transcription primer is 5′-CGGGC TTCGGTCCGG TTC-3′ (SEQ ID NO:3).
[0352] A two-step PCR reaction was used to amplify DNA, and the necessary TruSeq adaptor was added to prepare a DNA library for sequencing. 24DNA was amplified by PCR using a 200 μM dNTP mixture (New England Biolabs), 500 nM forward primers, 500 nM reverse primers, 1 ng cDNA or double-stranded DNA template, 20% (v / v) Q5 reaction buffer (New England Biolabs), and 0.02 U / μL Q5 hot-start high-fidelity polymerase (New England Biolabs). Excess unincorporated dNTPs and primers were removed by affinity purification (Agencourt AmpureXP magnetic beads; Beckman Coulter; sample to bead ratio 0.7:1). The DNA library was quantified on a Qubit fluorometer (Ingencourt) (Qubit dsDNA high-sensitivity assay kit; Ingencourt) to check the quality of the DNA library (Bioanalyzer 2100 on-chip electrophoresis system; Agilent Technologies), and sequenced on an Illumina NextSeq 550 high-throughput sequencer.
[0353] The amplicon-specific forward primers for preparing the SHAPE-MaP library are 5′-CCCTA CACGA CGCTC TTCCGATCTN NNNN G GCCTT CGGGC CAAGG A -3′(SEQ ID NO:4). The amplicon-specific reverse primers for preparing the SHAPE-MaP library are 5′-GACTG GAGTT CAGAC GTGTG CTCTT CCGAT CTNNN NNTT G AACCG GACCG AAGCC CGATT T -3′(SEQ ID NO:5). Sequences overlapping with the RNA selection construct are underlined.
[0354] Example 4: Isothermal titration calorimetry
[0355] ITC experiments were performed using the Microcal PEAQ-ITC automated instrument (MalvernAnalytical) under RNase-free conditions. 41 The in vitro transcribed RNA was exchanged for folding buffer containing 100 mM CHES, pH 8.0, 200 mM potassium acetate, and 3 mM MgCl2 using centrifugation (Amicon Ultra centrifuge filter, 10K MWCO, Millipore-Sigma). Ligands were dissolved in the same buffer at 10–20 times the desired experimental RNA concentration (to minimize the heat of mixing when adding ligands to RNA). RNA concentration was quantified (Nanodrop UV-VIS spectrophotometer; Thermo Fisher Scientific) by diluting the RNA concentration in the buffer to the expected Kc. dThe RNA was diluted 1-10 times and requantified to confirm the final experimental RNA concentration. The RNA diluted in folding buffer was heated at 65°C for 5 minutes, placed on ice for 5 minutes, and folded at 37°C for 15 minutes. If needed, the primary binding ligand (e.g., 2) was pre-bound to the RNA by adding 0.1 volume at 10 times the desired final concentration of the binding ligand, followed by incubation at room temperature for 10 minutes.
[0356] Each ITC experiment involved two runs: one to titrate the ligand into RNA (experimental trace), and another to titrate the same ligand into buffer (control trace). ITC experiments were performed using the following parameters: cell temperature 25°C, reference power 8 μcal / s, agitation speed 750 RPM, high feedback mode, initial injection of 0.2 μL, followed by 19 injections of 2 μL each. Each injection took 4 seconds to complete, with a 180-second interval between injections.
[0357] ITC data were analyzed using MicroCalPEAQ-ITC analysis software (Marvin analysis). First, the baseline of each injection peak was manually adjusted to address any incorrectly chosen injection endpoints. Second, the control trace was subtracted from the experimental trace using point-to-point subtraction. Third, a least-squares regression line was fitted to the data using the Levenberg-Marquardt algorithm. For weakly binding ligands (>500 μM), N was manually set to 1.0 to achieve a low c-value curve fit.
[0358] Example 5: Testing the chemical synthesis of compounds 35, 36, 37, 38, 39 and 40.
[0359]
[0360] Compound 35: 3-C-linked hydroxamic acid 35 is prepared by reacting carboxylic acid S19 with a mixed anhydride intermediate and an aqueous hydroxylamine. Acid S19 is obtained by treating quinoxaline-6-amine with a cyclized anhydride dihydrofuran-2,5-dione.
[0361]
[0362] Compound 36: The 2-C-linked analog 36 is obtained by reacting the corresponding ester S20 with an in-situ formed hydroxylamine. Ester S20 is prepared by the Michael addition of quinoxaloline-6-amine with ethyl acrylate.
[0363]
[0364] Compound 37: The Buchwald-Hartwig reaction was used to synthesize intermediates S21 and S22. The protecting group (Boc) was removed with diethyl ether containing HCl, followed by further treatment with Na2CO3 to obtain 37.
[0365]
[0366] Compound 38: The formation of an imine and subsequent reduction of quinoxaline-6-carboxaldehyde and diamine with sodium borohydride to obtain 38.
[0367]
[0368] Compound 39: Using S N The imine formation of quinoxaline-6-ylmethylamine hydrochloride and aldehyde S23 prepared by Ar reaction and subsequent reduction with sodium borohydride yields intermediate S24, which is then deprotected by (Boc) HCl to give 39.
[0369]
[0370] Compound 40: Less constrained analog 40 was prepared by two Buchwald-Hartwig reactions with 3,5-dibromopyridine followed by (Boc) deprotection with HCl.
[0371] Example 6: X-ray crystallography
[0372] To evaluate whether structural variant 2 was a good candidate for binding to the TPP riboswitcher, compound 17 was investigated in X-ray crystallography. The TPP riboswitcher RNA was prepared via in vitro transcription as described. 27 TPP riboswitching RNA (0.2 mM) and RNA-17 (2 mM) were heated at 60 °C for 3 min in a buffer containing 50 mM potassium acetate (pH 6.8) and 5 mM MgCl2, rapidly cooled on crushed ice, and incubated at 4 °C for 30 min prior to crystallization. For crystallization, 1.0 μL of the RNA-17 complex was mixed with a 1.0 μL reservoir solution containing 0.1 M sodium acetate (pH 4.8), 0.35 M ammonium acetate, and 28% (v / v) PEG4000. Crystallization was carried out over 2 weeks at 291 K by hanging drop vapor diffusion. The crystals were cryoprotected in a stock solution supplemented with 15% glycerol before rapid freezing in liquid nitrogen. Data were obtained at NSLS-II (Brookhaven National Laboratory). Data were collected on a 17-ID-2 (FMX) beamline at the specified wavelength. Data were processed using an HKL200043. The structure was resolved by molecular substitution using Phenix44 and 2GDI riboswitching RNA structures. 27 The structure was improved in Phenix. In later stages of the improvement, organic ligands, water molecules, and ions were added based on the electron density maps of Fo-Fc and 2Fo-Fc.
[0373] The results showed that compound 17 binds to the TPP riboswitch in a manner similar to the thiamine moiety of the TPP ligand, thereby stacking between G42 and A43 in the J3 / 2 junction (Figure 3). 27,28 17 forms three hydrogen bonds with RNA: one each to the ribose at G40 and the Watson-Crick face, and one to the ribose at G19. The local RNA structure shows a significant change compared to RNA complexed with the native TPP ligand. In the 17-binding structure, G72 is flipped into the binding site containing the pyrophosphate portion of the TPP ligand. This binding pattern is consistent with previous work that visualized the inverted G72 orientation of fragments bound in the thiamine subsite of the riboswitching binding notch. 17,34 Consistent with SAR analysis, the orientation of the C-6 substituent appears to be relatively unaffected by its interaction with RNA, suggesting that this vector would be a good candidate for fragment purification.
[0374] By referencing and incorporating into the sequence list
[0375] The material in the appended sequence list is hereby incorporated herein by reference in its entirety. The accompanying file, named Sequence List 39397600002_ST25, was created on August 5, 2020, and is 4KB in size.
[0376] Government support
[0377] This invention was carried out with government support under license numbers GM098662 and AI068462 granted by the National Institutes of Health (NIH). The government enjoys certain rights in this invention.
[0378] References
[0379] 1. Hajduk, PJ, Huth, JR and Tse, C. Predicting protein druggability. Drug Discovery Today, 10, 1675-1682 (2005).
[0380] 2. Vukovic, S. and Huggins, DJ. Quantitative metrics for drug-target ligandability. Drug Discovery Today 23, 1258-1266 (2018).
[0381] 3. Batey, RT, Rambo, RP and Doudna, JA. Tertiary Motifs in RNA Structure and Folding. Angew. Chem. Int. Ed. 38, 2326-2343 (1999).
[0382] 4. Warner, KD, Hajdin, CE, and Weeks, KM. Principles for targeting RNA with drug-like small molecules. Nature Reviews Drug Discovery, 17, 547-558 (2018).
[0383] 5. Sharp, The Centrality of RNA. Cell 136, 577-580 (2009).
[0384] 6. Kozak, M. Regulation of translation via mRNA structure in prokaryotes and eukaryotes. Gene 361, 13-37 (2005).
[0385] 7. Corbino, KA, Sherlock, ME, McCown, PJ, Breaker, RR, and Stav, S. Riboswitch diversity and distribution. RNA 23, 995-1011 (2017).
[0386] 8. Cech, TR and Steitz, JA. The noncoding RNA revolution - Trashing old rules to forge new ones. Cell 157, 77-94 (2014).
[0387] 9. Parsons, C., Slack, FJ, Zhang, WC, Adams, BD, and Walker, L. Targeting noncoding RNAs in disease. Journal of Clinical Research (J. Clin. Invest.) 127, 761-771 (2017).
[0388] 10. Matsui, M. and Corey, DR. Non-coding RNAs as drug targets. Nature Reviews Drug Discovery 16, 167-179 (2017).
[0389] 11. Guan, L. and Disney, MD. Recent advances in developing small molecules targeting RNA. ACS Chemical Biology 7, 73-86 (2012).
[0390] 12. The Emerging Role of RNA as a Therapeutic Target for Small Molecules. Cell Chem. Biology 23, 1077-1090 (2016).
[0391] 13. Murray, CW and Rees, DC. The rise of fragment-based drug discovery. Nature Chemistry, 1, 187-92 (2009).
[0392] 14. Doak, BC, Norton, RS, and Scanlon, MJ. The ways and means of fragment-based drug design. Pharmacology and Therapeutics, 167, 28-37 (2016).
[0393] 15. Cressina, E., Chen, L., Abell, C., Leeper, FJ, and Smith, AG. Fragment screening against the thiaminepyrophosphate riboswitch thiM. Chem.Sci. 2, 157-165 (2011).
[0394] 16. Moumné, R., Catala, M., Larue, V., Micouin, L. and Tisné, C. Fragment-based design of small RNAbinders: Promising developments and contribution of NMR. Biochimie 94, 1607-1619 (2012).
[0395] 17. Warner, KD et al. Validating fragment-based drug discovery for biological RNAs: Lead fragments bind and remodel the TPP riboswitch specifically. Chemical Biology 21, 591-595 (2014).
[0396] 18. Zeiger, M. et al. Fragment-based search for small molecule inhibitors of HIV-1 Tat-TAR. Bioorganic Med. Chem. Lett. 24, 5576-5580 (2014).
[0397] 19. Bottini, A. et al. Targeting Influenza AVirus RNA Promoter. Chemical Biology and Drug Design 86, 663-673 (2015).
[0398] 20. Hunter, CA and Anderson, HL. What is cooperation? (Applied Chemistry International Edition) 48, 7488-7499 (2009).
[0399] 21. Ichihara, O., Barker, J., Law, RJ and Whittaker, M. Compound design by fragment-linking. Molecular Informatics 30, 298-306 (2011).
[0400] 22. Zeller, MJ, Li, K., Aubé, J. and Weeks, KMTPP. Multisite ligand recognition and cooperation in the TPPriboswitch RNA. Prep. (2019).
[0401] 23. Siegfried, NA, Busan, S., Rice, GM, Nelson, JAE, and Weeks, KM. RNA motif discovery by SHAPE and mutational profiling (SHAPE-MaP). Nature Methods 11, 959-65 (2014).
[0402] 24. Smola, MJ, Rice, GM, Busan, S., Siegfried, NA, and Weeks, KM. Selective 2'-hydroxyl acylation analyzed by primer extension and mutational profiling (SHAPE-MaP) for direct, versatile, and accurate RNA structure analysis. Nat. Protoc. 10, 1643-1669 (2015).
[0403] 25. Merino, EJ, Wilkinson, KA, Coughlan, JL, and Weeks, KM. RNA structure analysis at single nucleotide resolution by selective 2'-hydroxyl acylation and primer extension (SHAPE). Journal of the American Chemical Society (J. Am. Chem. Soc.) 127, 4223-4231 (2005).
[0404] 26. Liu, Z.-Y. et al. Novel cis-acting element within the capsid-coding region enhances flavivirus viral-RNA replication by regulating genome cyclization. Journal of Virology 87, 6804-18 (2013).
[0405] 27. Serganov, A., Polonskaia, A., Phan, AT, Breaker, RR, and Patel, DJ. Structural basis for gene regulation by athiamine pyrophosphate-sensing riboswitch. Nature 441, 1167-1171 (2006).
[0406] 28. Edwards, TE and Ferré-D'Amaré, AR. Crystal structures of the thi-box riboswitch bound to thiamine pyrophosphate analogs reveal adaptive RNA-small molecule recognition. Structure 14, 1459-68 (2006).
[0407] 29. Thore, S., Frick, C. and Ban, N. Structural basis of thiamine pyrophosphate analogues binding to the eukaryotic riboswitch. Journal of the American Chemical Society 130, 8116-8117 (2008).
[0408] 30. Busan, S. and Weeks, KM. Accurate detection of chemical modifications in RNA by mutational profiling (MaP) with ShapeMapper 2. RNA 24, 143-148 (2018).
[0409] 31. Woolson, R. Statistical Methods for the Analysis of Biomedical Data. (John Wiley & Sons, 1987).
[0410] 32. Jhoti, H., Williams, G., Rees, DC, and Murray, CW. The 'rule of three' for fragment-based drug discovery: Where are we now? Nature Reviews Drug Discovery 12, 644 (2013).
[0411] 33. Chen, L. et al. used thiamine pyrophosphate analogues to detect riboswitch-ligand interactions. Organic and Biomolecular Chemistry, 10, 5924-5931 (2012).
[0412] 34. Warner, KD and Ferré-D'Amaré, AR. Crystallographic analysis of TPPriboswitch binding by small-molecule ligands discovered through fragment-based drug discovery approaches. Methods Enzymol. 549, 221-233 (2014).
[0413] 35. Codd, R. Traversing the coordination chemistry and chemical biology of hydroxamic acids. Coord. Chem. Rev. 252, 1387-1408 (2008).
[0414] 36. Jencks, W.P. On the attribution and additivity of binding energies. Proceedings of the National Academy of Sciences of the United States of America (Proc. Natl. Acad. Sci. USA) 78, 4046-4050 (1981).
[0415] 37. Olejniczak, ET et al. Stromelysin inhibitors designed from weakly bound fragments: Effects of linking and cooperativity. Journal of the American Chemical Society 119, 5828-5832 (1997).
[0416] 38. Borsi, V., Calderone, V., Fragai, M., Luchinat, C., and Sarti, N. Entropic contribution to the linking coefficient in fragment-based drug design: A case study. Journal of Medicinal Chemistry, 53, 4285-4289 (2010).
[0417] 39. Reuter, JS and Mathews, DHRNA structure: software for RNA secondary structure prediction and analysis. BMC Bioinformatics 11, 129 (2010).
[0418] 40. Busan, S., Weidmann, CA, Sengupta, A., and Weeks, K. Guidelines for SHAPE Reagent Choice and Detection Strategy for RNA Structure Probing Studies. Biochemistry 58, 2655-2664 (2019).
[0419] 41. Gilbert, SD and Batey, RT. Monitoring RNA-ligand interactions using isothermal titration calorimetry. Methods in Molecular Biology, 540, 97-114 (2009).
[0420] 42. Turnbull, WB Dispersion Leads to Failure? Studying low affinity fragments of ligands by ITC. Microcal Application Notes (2005).
[0421] 43. Otwinowski, Z. and Minor, W. Processing of X-ray diffraction data collected in oscillation mode. Enzymatic Methods (1997). doi:10.1016 / S0076-6879(97)76066-X.
[0422] 44. Liebschner, D. et al. Macromolecular structure determination using X-rays, neutrons and electrons: recent developments in Phenix. Acta Crystallography, 75, 861-877 (2019).
[0423] 45. Hajduk, PJ et al. Discovery of potent nonpeptide inhibitors of stromelysin using SAR by NMR. Journal of the American Chemical Society 119, 5818-5827 (1997).
[0424] 46. Howard, N. et al. Application of fragment screening and fragment linking to the discovery of novel thrombin inhibitors. Journal of Medicinal Chemistry 49, 1346-1355 (2006).
[0425] 47. Barker, JJ et al. Discovery of a novel Hsp90 inhibitor by fragment linking. ChemMedChem 5, 1697-1700 (2010).
[0426] 48. H. et al. discovered potent, selective, and structurally novel Dot1L inhibitors using a fragment linking approach. ACS Med. Chem. Lett. 8, 338-343 (2017).
[0427] 49. Hung, AW, et al. Application of fragment growing and fragment linking to the discovery of inhibitors of mycobacterium tuberculosis pantothenate synthase. Applied Chemistry International Edition 48, 8452-8456 (2009).
[0428] 50. Jordan, JB et al., Using 19F NMR Spectroscopy to Obtain Highly Potent and Selective Inhibitors of β-Secretase. Journal of Medicinal Chemistry, 59, 3732-3749 (2016).
[0429] 51. Maly, DJ, Choong, IC, and Ellman, JA. Combinatorial target-guided ligand assembly: Identification of potent subtype-selective c-Src inhibitors. Proceedings of the National Academy of Sciences (PNAS) 97, 2419-2424 (2000).
[0430] 52. Shuker, SB, Hajduk, PJ, Meadows, RP, and Fesik, SW. Discovering High-Affinity Ligands for Proteins: SAR by NMR. Science (80-.). 274, 1531-1534 (1996).
[0431] 53. Mondal, M. et al. Fragment Linking and Optimization of Inhibitors of the Aspartic Protease Endothiapepsin: Fragment-Based Drug Design Facilitated by Dynamic Combinatorial Chemistry. Angewandte Chemie International Edition 55, 9422-9426 (2016).
[0432] 54. Swayze, EE et al. SAR by MS: Aligand-based technique for drug lead discovery against structured RNA targets. Journal of Medicinal Chemistry 45, 3816-3819 (2002).
Claims
1. A compound having the structure of formula (III): in L is Where q and r are independently selected from the integers 0, 1, 2, and 3; and A is Where X6 is N, and X4, X5 and X7 are CR3, where R3 is -H; m is 1 or 2; and W is –O or –NR4, where R4 is selected from -H, -CO (C1-C6 alkyl), substituted or unsubstituted C1-C6 alkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted cycloalkyl, -CO (aryl), -CO (heteroaryl), and -CO (cycloalkyl); Or its pharmaceutically acceptable salt.
2. The compound according to claim 1, wherein q and r are independently selected from integers 0, 1, and 2.
3. The compound according to claim 1 or 2, wherein q and r are 0 or 1.
4. The compound according to claim 1 or 2, wherein q and r are 1.
5. The compound according to claim 1 or 2, wherein q is 1 and r is 0.
6. The compound according to any one of claims 1 to 5, wherein W is selected from -NH, -O and -N(C1-C6 alkyl)2.
7. The compound according to any one of claims 1 to 6, wherein W is -NH.
8. The compound according to claim 7, wherein A is 9. A composition comprising the compound according to any one of claims 1 to 8, wherein the compound is contained in a pharmaceutically acceptable carrier.
10. A medicament for treating a disease or condition associated with dysfunction of RNA expression, comprising a therapeutically effective amount of the compound according to any one of claims 1 to 8 or the composition according to claim 9. in, Because the compound binds to RNA, administration of the drug reduces protein expression. The disease or condition is selected from hereditary diseases, degenerative diseases, cancer, diabetes, autoimmune diseases, cardiovascular diseases, coagulation disorders, eye diseases, infectious diseases, and diseases caused by mutations in one or more genes.
11. A method for preparing the compound according to any one of claims 1 to 8, the method comprising: In the presence of a Pd catalyst, the IV' fragment is brought into contact with the V-1 fragment. Wherein, formula IV' is Equation V-1 is Where X is a halogen selected from F, Br, Cl and I; X6 is N, and X4, X5 and X7 are CR3, where R3 is -H; m is 1 or 2; and W is -O or -NR4, where R4 is (C1-C6) alkyl or -H.
Citation Information
Patent Citations
Detection of chemical modifications in nucleic acids
US10240188B2
Regulation of gene expression by aptamer-mediated modulation of alternative splicing
US20180010146A1
Polylactide-drug mixtures
US3773919A
Aptamer-mediated regulation of gene expression
US6949379B2
High-throughput RNA structure analysis
US8318424B2