Divalent nucleic acid ligands and their use
Patent Information
- Application Number
- JP2024135571
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-02-14
- Filing Date
- 2024-08-15
- Publication Date
- 2026-09-03
- Estimated Expiration
- 2038-12-21
Smart Images

Figure 0007914961000077 
Figure 0007914961000078 
Figure 0007914961000079
Abstract
Description
[Technical Field]
[0001] Complaint regarding federal government subsidies This invention was made with government support under authorization number R21NS098102 from the National Institutes of Health and authorization number CHE1039870 from the National Science Foundation. The government has certain rights to this invention.
[0002] Cross-reference of related applications This application claims the benefit of the concurrently pending U.S. Provisional Patent Application No. 62 / 708,783, filed on 21 December 2017, and the U.S. Provisional Patent Application No. 62 / 710,262, filed on 14 February 2018, both of which are incorporated herein by reference in their entirety.
[0003] The sequence listing related to this application has been submitted electronically via EFS-Web and is incorporated in its entirety into the specification by reference. The name of the text file containing the sequence listing is 6526_1807597_ST25.txt. The size of this text file is 444 bytes, and it was created on December 21, 2018.
[0004] 1. Field of Invention This specification describes compositions for binding nucleic acids using nucleic acid and nucleic acid oligomer compositions. It also provides methods for treating expanded repeat diseases, such as PolyQ diseases including Huntington's disease. [Background technology]
[0005] 2. Description of related technologies RNA-repeated expansion is commonly seen in neuromuscular disorders (or diseases). One example is Huntington's disease (HD), an autosomal dominant disorder that affects muscle coordination and leads to behavioral changes, cognitive decline, and dementia. HD typically develops in middle age and results in death within 10 to 20 years of onset. This is due to the expansion of the CAG repeat in the first exon of the Huntington (htt) gene from the normal range of 6 to 29 to the pathogenic range of 40 to 180. exp The length of the CAG repeat is inversely proportional to the age of onset. Htt is eccentrically expressed but primarily in CNS neurons and encodes a 348 kDa protein with diverse physiological roles, including embryonic development and neuroprotective effects. The CAG repeat is translated into a polyglutamine (polyQ) sequence in the N-terminal region of Htt. Despite the vast amount of information on molecular dysfunction and clinical symptoms, CAG exp The precise mechanisms that cause HD are not yet fully understood; however, emerging evidence suggests that HD is a multivariate disorder.
[0006] Loss of protein function may contribute to the pathogenesis of HD, but CAG expHeterozygous and homozygous patients with the htt allele exhibit similar clinical features, making it unlikely to be the primary cause. Furthermore, individuals have been observed who do not exhibit the abnormal phenotype despite a 50% reduction in normal Htt protein levels due to the deletion of one of the htt alleles. Similarly, heterozygous mice for the htt null mutation do not exhibit the clinical features of HD, while homozygous mice die during early embryonic development. These findings highlight the importance of Htt in embryonic development but are not significant in the pathogenesis of HD. Newly emerging evidence points to adverse function acquisition as a more plausible cause of the disease. This suggestion is supported by observations that CAG repeats exceeding similar thresholds (30-40 units) in unrelated genes are the cause of several neuromuscular disorders, including HD, DRPLA (dentateburubral-pallidoluysian atrophy), SBMA (spinal and bulbar muscular atrophy), SCA1 (spinocerebellar ataxia type 1), SCA2 (spinocerebellar ataxia type 2), SCA3 (spinocerebellar ataxia type 3 or Machado-Joseph disease), SCA6 (spinocerebellar ataxia type 6), SCA7 (spinocerebellar ataxia type 7), and SCA17 (spinocerebellar ataxia type 17). These disorders, referred to as polyQ disorders, share many characteristics with HD, including subcortical and cortical atrophy, as well as nuclear aggregates.
[0007] Three main cytotoxic mechanisms have been proposed for HD and other polyQ diseases. The first is polyQ toxicity. PolyQ proteins tend to aggregate and interact with other polyQ-containing proteins to form large amyloid-like structures. The binding and sequestration of these key proteins leads to the loss of their physiological function, resulting in abnormal regulation of a cascade of molecular and cellular events. The second is toxic-gain of RNA function. During transcription, rCAG exp It adopts an incomplete hairpin structure that segregates muscleblind-like protein 1 (MBNLl), alternative RNA splicing regulators, and other major proteins. While normal rCAG repeats can also select hairpin motifs, they are selected in a different sequence configuration than the aforementioned extended version.exp When MBNL1 is associated with MBNL1, a complex is obtained that is trapped in the nucleus as a nuclear foci, preventing its transport to the cytoplasm for Htt protein production. The gene transcripts that are misspliced as a result of MBNL1 loss are diverse. A third pathogenesis mechanism is protein toxicity. It has been shown that rCAG repeats exceeding a certain length (>42 units), in the absence of an ATG start codon, are translated via repeat-associated non-ATG (RAN) translation, leading to the production of toxic polyQ and polyalanine (polyA) proteins along with polyserine (polyS). The latter two mechanisms are further supported by the finding that expression of long untranslated rCAG repeats is harmful in animal models, and the abnormal behavioral phenotypes are similar to those of HD. [Overview of the project] [Problems that the invention aims to solve]
[0008] In general, these findings indicate that HD is a multivariate disorder resulting from the acquisition of toxicity in RNA and protein function. Therefore, one promising strategy for treating HD is to selectively target the elongated transcript, because this interferes with all three disease pathways. Described herein are divalent nucleic acid ligands for the recognition of rCAG repeats and other RNA-repeat sequences. [Means for solving the problem]
[0009] Summary of the Invention In one embodiment, a gene recognition reagent is provided. The reagent comprises a nucleic acid or nucleic acid analog skeleton containing three or more ribose, deoxyribose, or nucleic acid analog skeleton residues, and a divalent nucleic acid base sequence bound to the skeleton residues, wherein the divalent nucleic acid base sequence is bound to a unit target sequence of an extended repeat of a repeat expansion disease on two nucleic acid strands, or to one or more sequential repeats of the unit target sequence. In another embodiment, a composition is provided comprising the gene recognition reagent and a pharmaceutically acceptable carrier.
[0010] In another embodiment, a method is also provided for binding nucleic acids containing elongated repeats associated with repeat elongation disorders. The method comprises a nucleic acid or nucleic acid analog skeleton comprising three or more ribose, deoxyribose, or nucleic acid analog skeleton residues, a divalent nucleic acid base sequence bound to the skeleton residues, wherein the divalent nucleic acid base sequence is bound to a unit target sequence of the elongated repeat of the repeat elongation disorder on two nucleic acid strands, or to one or more sequential repeats of the unit target sequence.
[0011] In yet another embodiment, a method is provided for identifying the presence of nucleic acids containing elongated repeats associated with repeat elongation disease in a sample obtained from a patient. The method includes contacting the nucleic acid sample obtained from the patient with a gene recognition reagent comprising three or more ribose, deoxyribose, or nucleic acid analog skeletons comprising three or more ribose, deoxyribose, or nucleic acid analog skeleton residues, a divalent nucleic acid base sequence bound to the skeleton residues, wherein the divalent nucleic acid base sequence is bound to a unit target sequence of the elongated repeats of repeat elongation disease on two nucleic acid strands, or to one or more sequential repeats of the unit target sequence; determining whether binding and linking of the gene recognition reagent occurs to indicate that the sample contains nucleic acids containing elongated repeats associated with repeat elongation disease; and optionally treating the patient for the repeat elongation disease. The nucleic acid or nucleic acid analog skeleton comprises a first end and a second end, further comprising a first linking group bound to the first end of the skeleton, and a second linking group bound to the second end, which is non-covalently bound to or self-ligated to the first linking group.
[0012] In yet another embodiment, a method is provided for knocking down the expression of a gene containing an elongated repeat associated with repeat elongation disease in cells. The method comprises contacting a nucleic acid, such as RNA containing the elongated repeat, with a gene recognition reagent comprising three or more ribose, deoxyribose, or nucleic acid analog backbone residues, a nucleic acid or nucleic acid analog backbone, a divalent nucleic acid sequence bound to the backbone residues, wherein the divalent nucleic acid sequence is bound to a unit target sequence of the elongated repeat of repeat elongation disease on two nucleic acid strands, or to one or more sequential repeats of the unit target sequence. [Brief explanation of the drawing]
[0013] [Figure 1]Figure 1(A) (Figure 1, Panel (A)) Designs of JB-MPγPNA ligands LG1, LG2, and LG3 for targeting the rCAGexp-hairpin structure, K: L-lysine, arrow: N-terminus, (B) Representative binding mode of LG2 in preferred antiparallel direction (N-terminus facing the 3' end of the Watson strand), (C) H-binding interactions of E, I, and F with each CG, AA, and GC base pair of RNA, (D) Chemical structure of the MPγPNA backbone with covalently bonded J bases, (E) Assumptions of γPNA, γPNA-DNA duplex, and PNA [as PNA-DNA duplex] confirmed by NMR and X-ray. [Figure 2] Figure 2, panels A-F show exemplary structures of nucleic acid analogs. [Figure 3] Figure 3 shows various examples of amino acid side chains. [Figure 4] Figure 4 shows an exemplary nucleic acid base structure. [Figure 5A] Figures 5A to 5C show exemplary divalent nucleic acid base structures, where R represents a nucleotide, nucleotide analog backbone monomer, or residue. [Figure 5B] Figures 5A to 5C show exemplary divalent nucleic acid base structures, where R represents a nucleotide, nucleotide analog backbone monomer, or residue. [Figure 5C] Figures 5A to 5C show exemplary divalent nucleic acid base structures, where R represents a nucleotide, nucleotide analog backbone monomer, or residue. [Figure 6]Results of MD simulations of LG1 binding to an RNA duplex containing the sequence r(CAG)4 / r(CAG)4. LG1: NH2-EIF-H, γ side chain is a methyl group. (A) Surface diagrams of the binding complex with the RNA duplex containing four distinct LG1 ligands and four CAG repeats before (t=0) and after (t=100ns) simulation. (B) H-binding interactions of the CEG, AJA, and GFC triads after simulation. (C) Number of H-bindings per triad between RNA and J bases of LG1, (D) radius of rotation of the LG1-RNA complex, and (E) mean square deviation (RMSD) of the LG1-RNA complex relative to the initial structure. [Figure 7] Figure 7 shows the structures of E, I, and F JB-MPγPNA monomers. [Figure 8] Figures 8 and 9 show the composite schemes of E(15) and F(26), respectively. [Figure 9] Figures 8 and 9 show the composite schemes of E(15) and F(26), respectively. [Figure 10] Figure 10 shows a scheme for coupling E(15), F(26), and I(27) to the MPγPNA skeleton. [Figure 11] Figure 11 shows the model RNA target selected to link the studies described in the examples (SEQ ID NO: 1). [Figure 12] Figure 12 shows the CD titration of LG2 in R11A. The concentration of R11A was 0.5 μM, and the concentrations of LG2 were X: 0, 1.5, 2.5, 3.5, 4.5, 5.5, 6.5, 8.5, and 11.5 μM. After adding LG2 to R11A, the mixture was incubated at 37°C for 30 minutes, and then data were acquired at 25°C. [Figure 13] Figure 13 shows the CD spectra of LG2 and (A) mismatched R11U, (B) single-stranded WS and CS, and (C) double-stranded HP including a single binding site, and (D) CD spectra of LG2P and R11A. The CD spectra of RNA (solid line) and the CD spectra of the [RNA + ligand] mixture (dashed line) are shown. [Figure 14]Figure 14 shows the effects of ligand orientation (LG2P), mismatch R11U, and single-strand targeting (WS+CS). The inset shows the fluorescence spectra of the corresponding samples after the addition of 2.5 μM pentamidine. [Figure 15] Figure 15 shows the NMR titration of LG2 in R11A. The concentration of R11A was 0.1 mM. Aliquots of LG2 at the indicated molar ratios were added to R11A, incubated at 37°C for 15 minutes, and then data were collected. Solvent-exchangeable protons are shown with dashed lines, and non-exchangeable protons are shown with vertical solid lines. [Figure 16] Figure 16 schematically illustrates native chemical ligation as described in the examples. [Figure 17] Figure 17 shows the reaction of the parent compound M after the addition of 4MP at 37°C. (A) A schematic diagram of the reaction sequence. The M# and M## intermediates are first formed, then converted to the reactive LG2N* intermediate, and finally to the cyclic product cLG2N. (B) MALDI-TOF MS spectra of M after the addition of 4MP and incubation at 37°C for 0 min, 15 min, and 60 min. The samples were prepared in 0.1×PBS buffer. [Figure 18] Figure 18 shows the MALDI-TOF spectra of (A) [LG2N+R11A], (B) [LG2N+R11U], and (C) LG2N alone after incubation with 2ME at 37°C for 16 hours. Prior to MS analysis, the samples were restored with 4 mM guanidinium chloride and denatured at 95°C for 5 minutes. The concentrations of LG2N, R11A, R11U, and 2ME were 22, 1, 1, and 500 μM, respectively. No RNA molecules were observed in the MALDI-TOF spectra in the positive mode. The Y-axis in (B) and (C) is the same as in (A) and has been omitted for space saving. [Figure 19]Figure 19 shows a competitive binding assay. The concentrations of R24A, R127A, and R96U were 100, 18, and 22 nM, respectively, each containing a 1.1 μM binding site. The concentrations of LG2N in lanes 2-6 and 8 were 11, 22, 44, 88, 88, and 88 μM, respectively, and the concentrations of 4MP and TCEP were 500 and 100 μM, respectively. The ratio of binding sites to ligands was as shown. The samples were prepared in 0.1 × PBS buffer, incubated at 37°C for 24 hours, separated by 2% agarose gel, and stained with SYBR-Gold. [Figure 20] Figure 20 shows the gel shift assay at physiologically simulated ionic intensities. Samples were prepared in the same manner as shown in Figure 9, except that the buffer was PS (10 mM NaPi, 150 mM KCl, 2 mM MgCl2 pH 7.4). The concentrations of R24A, R127A, and R96U were 100, 18, and 22 nM, respectively, and each contained 1.1 μM of binding site. The concentrations of LG2N in lanes 2-6 and 8 were 11, 22, 44, 88, 88, and 88 μM, respectively, and the concentrations of 4MP and TCEP were 500 and 100 μM, respectively. The ratios of binding sites to ligands were as shown. After incubation at 37°C for 24 hours, the samples were separated by 2% agarose gel and stained with SYBR-Gold. [Figure 21] Figure 21 shows a general method for synthesizing the ligands described herein. [Modes for carrying out the invention]
[0014] (Detailed description of the invention) The use of numerical values within the various ranges specified in this application is indicated as an approximation, as if preceded by "approximately" both the minimum and maximum values within the defined range, unless otherwise explicitly stated. In this way, even with slight variations above and below the defined range, substantially identical results can be achieved to those obtained with values within the defined range. Furthermore, unless otherwise indicated, the disclosure of a range is also intended as a continuous range encompassing all values between the minimum and maximum values. Where used herein, the singular forms ("a" and "an") refer to one or more.
[0015] As used herein, the term “comprising” is open-ended and may be synonymous with “including,” “containing,” or “characterized by.” The term “consisting essentially of” limits the claims to materials or processes that do not substantially affect the specified materials or processes or one or more fundamental and novel features of the invention described in the claims. The term “consisting of” excludes any elements, processes, or components not specified in the claims. As used herein, embodiments “comprising” one or more elements or processes also include, but are not limited to, embodiments “consisting essentially of” and “consisting of” those specified elements or processes.
[0016] Provided herein are compositions and methods for binding target sequences in nucleic acids, for example, for binding disease-related repeat extensions associated with repeat extensions of nucleic acid sequences. Figure 1 is a schematic diagram illustrating a non-limiting example of the binding of recognition reagents (gene recognition reagents, ligands) described herein, targeting CAG repeats in RNA hairpin structures, as seen in HD and other PolyQ diseases. Cooperative binding of modules to adjacent modules is facilitated by terminal self-ligating groups or π-stacking aromatic groups, for example, exemplary self-ligating groups described in the following examples and more extensively disclosed in U.S. Patent Application Publication 2016 / 0083433, which is incorporated herein by reference. In summary, self-ligating groups, π-stacking groups, and similar groups that covalently or non-covalently bind adjacent recognition reagents on nucleic acid strands or between two nucleic acid strands in recognition reagents containing divalent nucleic acid bases are collectively referred to herein as “linking groups.” Multiple recognition reagents bind (hybridize) to the template nucleic acid by Watson-Crick or Watson-Crick-like cooperative base pairing. In cells, the template nucleic acid is an RNA or DNA molecule, but in vitro, the template nucleic acid can be any RNA or DNA, and similarly, it can be a modified nucleic acid or nucleic acid analog. Recognition reagents that bind to sufficiently close adjacent sequences, for example, on the template nucleic acid, self-ligate or π-π stack, and thus ligate to essentially form longer oligomers or polymers. Further details are presented below.
[0017] As used herein, “patient” refers to mammals, including primates (e.g., humans, non-human primates, e.g., monkeys, and chimpanzees), non-primates (e.g., cattle, pigs, camels, llamas, horses, goats, rabbits, sheep, hamsters, guinea pigs, cats, dogs, rats, mice, horses, and whales), or animals such as birds (e.g., ducks or geese).
[0018] As used herein, the terms “to treat” or “to cure” refer to a beneficial or desired outcome, such as improvement of one or more symptoms of a disease. The terms “to treat” or “to cure” also include, but are not limited to, the reduction or improvement of one or more symptoms of PolyQ repeat elongation diseases, such as Huntington’s disease. “Cure” may also mean extended survival compared to the expected survival in the absence of treatment.
[0019] In relation to disease markers or symptoms, “lower” means a clinically relevant and / or statistically significant decrease at such a level. Such decrease may be, for example, at least 10%, at least 20%, at least 30%, or at least 40% or more, to a level acceptable as within the normal range for individuals without such disorder, or to a level below the detection level of the assay. In certain embodiments, such decrease may also be referred to as normalization of the level, which is a decrease to a level acceptable as within the normal range for individuals without such disorder. In certain embodiments, such decrease is normalization of the level of signs or symptoms of the disease, a decrease in the difference between the target level of the signs of the disease and the normal level of the signs of the disease (for example, to the upper limit of normal if the level with respect to the subject must be decreased to reach a normal value, and to the lower limit of normal if the value with respect to the subject must be increased to reach a normal level). In one embodiment, the method includes clinically significant inhibition of mRNA expression of polyQ repeat elongation disorders, such as Huntington's disease, as indicated by clinically significant results after treatment of the subject with a recognition reagent such as those described herein.
[0020] As used herein, “therapeutic effective dose” is intended to include the amount of the recognition reagent described herein that, when administered to a subject with a disease (for example, by alleviating, improving, or maintaining one or more symptoms of an existing disease), is sufficient to result in the treatment of the disease. The “therapeutic effective dose” may vary depending on the recognition reagent (drug), the method of drug administration, the disease and its severity, as well as the subject's medical history, age, weight, family history, genetic makeup, the type of prior or concomitant treatment, if any, and other individual characteristics of the subject to be treated. In relation to the gene recognition reagents described herein, the exemplary dose range per unit is in the range of 0.1 pg to 1 mg.
[0021] The “therapeutic effective dose” also includes the amount of an agent that produces several desired local or systemic effects in a reasonable benefit / risk ratio applicable to any treatment. The recognition reagent agent used in the methods described herein may be administered in an amount sufficient to produce a reasonable benefit / risk ratio applicable to such treatment.
[0022] As used herein, the term “pharmaceutically acceptable carrier” means a pharmaceutically acceptable material, composition, or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, talc, magnesium stearate, calcium stearate or zinc stearate, or stearic acid) or solvent encapsulation material, that is involved in transporting or delivering the compound of interest from one organ or part of the body to another organ or part of the body. Each carrier must be “acceptable” in the sense that it is compatible with the other components of the formulation and is not harmful to the body to be treated. Some examples of materials that can act as pharmaceutically acceptable carriers include: (1) sugars, e.g., lactose, glucose, and sucrose; (2) starches, e.g., corn starch and potato starch; (3) cellulose and cellulose derivatives, e.g., sodium carboxymethylcellulose, ethylcellulose, and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; and (7) lubricants, e.g., magnesium stearate. (1) state, sodium lauryl sulfate and talc, (8) excipients, e.g., cocoa butter and suppository wax, (9) oils, e.g., peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil and soybean oil, (10) glycols, e.g., propylene glycol, (11) polyols, e.g., glycerin, sorbitol, mannitol and polyethylene glycol, (12) esters, e.g., ethyl oleate and ethyl laurate, (13) agar, (14) buffering agent These include, for example, magnesium hydroxide and aluminum hydroxide, (15) alginic acid, (16) pyrogen-free water, (17) isotonic saline, (18) Ringer's solution, (19) ethyl alcohol, (20) pH buffer, (21) polyesters, polycarbonates and / or polyanhydrides, (22) fillers, such as polypeptides and amino acids, (23) serum components, such as serum albumin, HDL and LDL, and (22) other non-toxic, suitable substances used in pharmaceutical formulations.
[0023] As used herein, the terms “cell” and “cells” refer to any type of cell from any animal, including but not limited to rats, mice, monkeys, and humans. For example, but not limited to, a cell could be a progenitor cell, such as a stem cell, or a differentiated cell, such as an endothelial cell or a smooth muscle cell.
[0024] "Expression" or "gene expression" means the overall flow of information from a gene for the production of a gene product (typically a protein, optionally post-translationally modified or functional / structural RNA) (but not limited to, a functional genetic unit for producing a gene product such as RNA or a protein in a cell, or other expression systems encoded on nucleic acids, including transcription promoters and other cis-acting elements, e.g., response elements and / or enhancers, typically expressible sequences encoding proteins (open reading frames or ORFs) or functional / structural RNA, and polyadenylated sequences). "Expression of a gene under transcriptional regulation" of a specified sequence, or "regulated" by a specified sequence, means gene expression from a gene containing the specified sequence operably linked to the gene (typically functionally linked in cis). The specified sequence may be all or part of a transcription element (but not limited to promoters, enhancers, and response elements) that may regulate and / or influence gene transcription, either entirely or partially. The “gene for expression” of the indicated gene product is a gene that can express the product of the indicated gene when placed in a suitable environment, that is, when the cell is transformed, transfected, transduced, etc., and subjected to conditions suitable for expression. In the case of a constitutive promoter, “suitable conditions” means that the gene typically only needs to be introduced into a host cell. In the case of an inductive promoter, “suitable conditions” means that an effective amount of each inducer to cause gene expression is administered to the expression system (e.g., cells).
[0025] As used herein, the terms “knockdown” or “to knock down” mean that the expression of one or more genes in an organism is significantly reduced, typically with respect to functional genes, such as to a therapeutically effective degree. Gene knockdown also includes complete gene silencing. As used herein, “gene silencing” means that gene expression is essentially completely prevented. Knockdown and gene silencing can occur at either the transcriptional or translational stage. The use of the described recognition reagents for targeting RNA in cells, such as mRNA, can modify gene expression by knocking down or silencing one or more genes at the post-transcriptional or translational stage.
[0026] As used herein, the term “nucleic acid” refers to deoxyribonucleic acid (DNA) and ribonucleic acid (RNA). Nucleic acid analogs include, for example, but are not limited to, 2'-O-methyl-substituted RNA, locked nucleic acids, unlocked nucleic acids, triazole-bound DNA, peptide nucleic acids, morpholino oligomers, dideoxynucleotide oligomers, glycol nucleic acids, threose nucleic acids, and any combination thereof containing one or more ribonucleotides or deoxyribonucleotide residues. In this specification, with respect to nucleic acids and nucleic acid analogs, “nucleic acid” and “oligonucleotide,” which is a short, single-stranded structure composed of nucleotides, are used interchangeably. Oligonucleotides may be referred to by the nomenclature “-mer,” by the length of the chain (i.e., the number of nucleotides). For example, an oligonucleotide with 22 nucleotides would be referred to as a 22-mer.
[0027] A "nucleic acid analog" is a composition containing a sequence of nucleic acid bases arranged on a substrate such as a polymer backbone, which can bind DNA and / or RNA by hybridization via Watson-Crick hydrogen bond base pairing or Watson-Crick-like hydrogen bond base pairing. Non-limiting examples of common nucleic acid analogs include peptide nucleic acids, such as γPNA, morpholino nucleic acids, phosphorothioates, locked nucleic acids (2'-O-4'-C-methylene crosslinks including oxy, thio, or amino types), unlocked nucleic acids (C2'-C3' bonds cleaved), 2'-O-methyl-substituted RNA, threose nucleic acids, glycol nucleic acids, and the like.
[0028] A structurally pre-organized nucleic acid analog is a nucleic acid analog having a skeleton (pre-organized skeleton) that forms only either a right-handed or left-handed helix, depending on the structure of the nucleic acid skeleton. As shown herein, an example of a structurally pre-organized nucleic acid analog is γPNA, which has a chiral center at the γ carbon and forms only either a right-handed or left-handed helix, depending on and due to the chirality of the group at the γ carbon. Similarly, locked nucleic acids include ribose having a bridge between the 2' oxygen and 4' carbon that "locks" the ribose into a 3'-end (north) conformation.
[0029] In the context of this disclosure, “nucleotide” refers to a monomer comprising at least one nucleic acid base and a skeletal element (backbone) that is ribose or deoxyribose in nucleic acids such as RNA or DNA. “Nucleotides” also typically include reactive groups that enable polymerization under certain conditions. In natural DNA and RNA, these reactive groups are the 5' phosphate group and the 3' hydroxyl group. For the chemical synthesis of nucleic acids and nucleic acid analogs, the bases and skeletal monomers may include modifying groups such as blocked amines, as is known in the art. “Nucleotide residue” refers to a single nucleotide incorporated into an oligonucleotide or polynucleotide. Similarly, “nucleic acid base residue” refers to a nucleic acid base incorporated into a nucleotide or nucleic acid or nucleic acid analog. “Gene recognition reagent” broadly refers to a nucleic acid or nucleic acid analog comprising a sequence of nucleic acid bases that can hybridize (bind) to a complementary nucleic acid or nucleic acid analog sequence on a nucleic acid by cooperative base pairing, e.g., Watson-Crick base pairing or Watson-Crick-like base pairing.
[0030] More specifically, a nucleotide, in the case of RNA, DNA, or a nucleic acid analog, has structure A and B, where A is the backbone monomer and B is a nucleic acid base as described herein. The backbone monomer may be any suitable nucleic acid backbone monomer, such as ribose triphosphate or deoxyribose triphosphate, or a monomer of a nucleic acid analog such as a peptide nucleic acid (PNA), e.g., gamma PNA (γPNA). For example, the backbone monomer may be ribose monophosphate, diphosphate, or triphosphate, or deoxyribose monophosphate, diphosphate, or triphosphate, e.g., 5'-monophosphate, diphosphate, or triphosphate of ribose or deoxyribose. The backbone monomer includes both the structural "residue" component, such as ribose in RNA, and any active group that is modified when the monomer is bonded together, e.g., the 5'-triphosphate group and the 3'-hydroxyl group of a ribonucleotide that are modified to leave a phosphodiester bond when polymerized to RNA. Similar to PNA, the C-terminal carboxyl group and N-terminal amine active group of the N-(2-aminoethyl)glycine backbone monomer condense during polymerization to leave a peptide (amide) bond. In another embodiment, the active group is a phosphoramidite group useful for the synthesis of phosphoramidite oligomers, as is widely known in the art. The nucleotide also optionally comprises one or more protecting groups known in the art and described herein, such as 4,4'-dimethoxytrityl (DMT). Many additional methods for preparing synthetic gene recognition reagents are known, and these methods depend on the backbone structure and the specific chemical properties of the base addition process. The determination of which active group to utilize for the bonding of the nucleotide monomer, which group to protect at the base, and the steps required in the preparation of the oligomer are well within the scope of chemical technology and the capabilities of those skilled in the art in the field of nucleic acid and nucleic acid analog oligomer synthesis.
[0031] Non-limiting examples of common nucleic acid analogs include peptide nucleic acids, e.g., γPNA, phosphorothioates (e.g., Figure 2(A)), locked nucleic acids (2'-O-4'-C-methylene bridges including oxy, thio, or amino forms, e.g., Figure 2(B)), unlocked nucleic acids (C2'-C3' bond cleaved, e.g., Figure 2(C)), 2'-O-methyl-substituted RNA, morpholino nucleic acids (e.g., Figure 2(D)), threose nucleic acids (e.g., Figure 2(E)), glycol nucleic acids (e.g., Figure 2(F), showing R and S forms), etc. Figures 2(A-F) show monomer structures for various examples of nucleic acid analogs. Each of Figures 2(A-F) shows two monomer residues incorporated into a longer chain, indicated by a dashed line. The incorporated monomer is referred to herein as a “residue,” and the portion of the nucleic acid or nucleic acid analog excluding the nucleic acid base is referred to as the “backbone” of the nucleic acid or nucleic acid analog. For example, in the case of RNA, the exemplary nucleic acid base is adenine, the corresponding monomer is adenosine triphosphate, and the incorporated residue is an adenosine monophosphate residue. In the case of RNA, the “backbone” consists of ribose subunits linked by phosphate, and therefore the backbone monomer is ribose triphosphate before incorporation and ribose monophosphate residue after incorporation. Similar to γPNA, locked nucleic acids (Figure 2(B)) are pre-organized stereochemically.
[0032] A "part" is a portion of a molecule, and as a class, it includes "residues," which are the parts of a compound or monomer that remain in a larger molecule, such as a polymer chain, after the compound or monomer has been incorporated into a larger molecule, such as nucleotides or polypeptides incorporated into nucleic acids, or amino acids incorporated into proteins.
[0033] The term "polymer composition" refers to a composition containing one or more polymers. One class of "polymers" includes, but is not limited to, homopolymers, heteropolymers, copolymers, block polymers, and block copolymers, and may be natural or synthetic. Homopolymers contain one type of component or monomer, while copolymers contain two or more types of monomers. An "oligomer" is a polymer containing a small number of monomers, such as 3 to 100 monomer residues. Therefore, the term "polymer" includes oligomers. The terms "nucleic acid" and "nucleic acid analog" include nucleic acids, as well as nucleic acid polymers and oligomers.
[0034] A polymer "contains" a specified monomer or "is derived from" a specified monomer if the monomer is incorporated into the polymer. Therefore, the monomers included in or incorporated into a polymer are not identical to the monomers before their incorporation into the polymer, in that, during the polymerization process, at least certain linking groups are incorporated into the polymer backbone, or certain groups are removed. A polymer is said to contain a particular type of bond if that type is present in the polymer. The incorporated monomers are "residues." Typical monomers of nucleic acids or nucleic acid analogs are called nucleotides.
[0035] "Non-reactive" means that, in relation to chemical moieties such as molecules, compounds, compositions, groups, parts, or ions, the part does not substantially react with other chemical moieties for its intended use. Non-reactive moieties are selected so as not to interfere with, or only minimally interfering with, the intended use of the part, part, or group as a recognition reagent. In relation to linker moieties as described herein, they are non-reactive in that they do not interfere with the binding of the recognition reagent to the target template and do not interfere with the linking of the recognition reagent on the target template.
[0036] As used herein, “alkyl” means, for example, a linear, branched, or cyclic hydrocarbon group containing 1 to about 20 carbon atoms, for example, but not limited to C1~3 , C 1~6 , C 1~10 refers to a group, including but not limited to linear and branched alkyl groups, such as methyl, ethyl, propyl, butyl, pentyl, hexyl, heptyl, octyl, nonyl, decyl, undecyl, dodecyl, and the like. The term "substituted alkyl" refers to alkyl substituted at 1 or more, for example, 1, 2, 3, 4, 5, or even 6 positions, and the substituents are bonded at any available atom to form a stable compound having the substitution as described herein. The term "optionally substituted alkyl" refers to alkyl or substituted alkyl. "Halogen", "halide" and "halo" refer to -F, -Cl, -Br and / or -I. "Alkylene" and "substituted alkylene" refer to divalent alkyl and divalent substituted alkyl respectively, including but not limited to methylene, ethylene, trimethylene, tetramethylene, pentamethylene, hexamethylene, heptamethylene, octamethylene, nonamethylene or decamethylene. "Optionally substituted alkylene" refers to alkylene or substituted alkylene.
[0037] "Alkene or alkenyl" refers to a linear, branched or cyclic hydrocarbyl group containing, for example, 2 to about 20 carbon atoms, which has one or more, for example 1, 2, 3, 4 or 5 carbon-carbon double bonds, including but not limited to C 1~3 , C 1~6 , C 1~10This refers to a group. "Substitutable alkene" refers to an alkene substituted in one or more positions, for example, at positions 1, 2, 3, 4, or 5, where the substituent can be attached to any available atom to produce a stable compound having the substitutions described herein. "Optionally substituted alkene" refers to an alkene or substituted alkene. Similarly, "alkenylene" refers to a divalent alkene. Examples of alkenylenes include, but are not limited to, ethenylene (-CH=CH-) and all its stereoisomers and conformational isomeric forms. "Substitutable alkenylene" refers to a divalent substituted alkene. "Optionally substituted alkenylene" refers to an alkenylene or substituted alkenylene.
[0038] "Alkyne" or "alkynyl" refers to a linear or branched unsaturated hydrocarbon having the indicated number of carbon atoms and at least one triple bond. Examples of (C2-C8) alkynyl groups include, but are not limited to, acetylene, propyne, 1-butyne, 2-butyne, 1-pentine, 2-pentine, 1-hexine, 2-hexine, 3-hexine, 1-heptine, 2-heptine, 3-heptine, 1-octyne, 2-octyne, 3-octyne, and 4-octyne. The alkynyl group may be unsubstituted or optionally substituted with one or more substituents as described below herein. The term "alkynylene" refers to a divalent alkyne. Examples of alkynylenes include, but are not limited to, ethynylene and propynylene. "Substitutable alkynylene" refers to a divalent substituted alkyne.
[0039] The term "alkoxy" refers to an -O-alkyl group having the indicated number of carbon atoms. For example, (C1-C6) alkoxy groups include -O-methyl (methoxy), -O-ethyl (ethoxy), -O-propyl (propoxy), -O-isopropyl (isopropoxy), -O-butyl (butoxy), -O-sec-butyl (sec-butoxy), -O-tert-butyl (tert-butoxy), -O-pentyl (pentoxy), -O-isopentyl (isopentoxy), -O-neopentyl (neopentoxy), -O-hexyl (hexyloxy), -O-isohexyl (isohexyloxy), and -O-neohexyl (neohexyloxy). "Hydroxyalkyl" refers to an alkyl group in which one or more hydrogen atoms of an alkyl group are substituted with an -OH group (C1-C6). 10 ) refers to alkyl groups. Examples of hydroxyalkyl groups include, but are not limited to, -CH2OH, -CH2CH2OH, -CH2CH2CH2OH, -CH2CH2CH2CH2OH, -CH2CH2CH2CH2CH2OH, -CH2CH2CH2CH2CH2CH2OH, and their branched forms. The term "ether" or "oxygen ether" refers to alkyl groups in which one or more alkyl carbon atoms are substituted with -O- groups. The term ether includes -CH2-(OCH2-CH2) where P1 is a protecting group, -H, or (C1-C10)alkyl. q OP1 compounds are included. Exemplary ethers include polyethylene glycol, diethyl ether, and methylhexyl ether.
[0040] "Heteroatoms" refer to N, O, P, and S. Compounds containing an N or S atom can optionally be oxidized to the corresponding N oxide, sulfoxide, or sulfone compound. "Heterosubstituted" means an organic compound in any embodiment described herein in which one or more carbon atoms are substituted with N, O, P, or S.
[0041] "Aryl" refers to an aromatic ring system, such as phenyl or naphthyl, either alone or in combination. "Aryl" also includes aromatic ring systems that may optionally be condensed with a cycloalkyl ring. "Substitutable aryl" is an aryl group independently substituted with one or more substituents bonded to any available atom to form a stable compound, the substituents as described herein. "Optionally substituted aryl" refers to an aryl or substituted aryl group. "Arylene" means a divalent aryl group, and "substituted arylene" refers to a divalent substituted aryl group. "Optionally substituted arylene" refers to an arylene or substituted arylene. As used herein, the terms "polycyclic aryl group" and related terms such as "polycyclic aromatic group" mean a group consisting of at least two condensed aromatic rings. "Heteroaryl" or "hetero-substituted aryl" refers to an aryl group substituted with one or more heteroatoms such as N, O, P, and / or S.
[0042] "Cycloalkyl" refers to a monocyclic, bicyclic, tricyclic, or polycyclic, 3- to 14-membered ring system that is saturated, unsaturated, or aromatic. Cycloalkyls may be bonded via any atom. Cycloalkyls also refer to fused rings in which a cycloalkyl is fused with an aryl or heteroaryl ring. Representative examples of cycloalkyls, but not limited to, include cyclopropyl, cyclobutyl, cyclopentyl, and cyclohexyl. Cycloalkyls may be unsubstituted or optionally substituted with one or more substituents as described herein. "Cycloalkylene" refers to a divalent cycloalkyl. The term "optionally substituted cycloalkylene" refers to a cycloalkylene substituted with one, two, or three substituents bonded via any atom available to produce a stable compound, the substituents as described herein.
[0043] "Carboxyl" or "carboxyl (carboxylic)" means, where indicated, a group having the number of carbon atoms indicated, with a -C(O)OH group at its terminus, and therefore having the structure -RC(O)OH, where R is a divalent organic group including linear, branched, or cyclic hydrocarbons. These non-limiting examples include C 1~8 Carboxyl groups include, for example, ethanolic, propanoic, 2-methylpropanoic, butanoic, 2,2-dimethylpropanoic, and pentanoic groups. "Amine" or "amino" refers to a group having the number of carbon atoms indicated, with an -NH2 group at the end, and thus having the structure -R-NH2, where R is an unsubstituted or substituted divalent organic group containing a linear, branched, or cyclic hydrocarbon and optionally one or more heteroatoms.
[0044] The above combined terms refer to preferred combinations of the above, for example, arylalkenyl, arylalkynyl, heteroarylalkyl, heteroarylalkenyl, heteroarylalkynyl, heterocyclylalkyl, heterocyclylalkenyl, heterocyclylalkynyl, aryl, heteroaryl, heterocyclyl, cycloalkyl, cycloalkenyl, alkylarylalkyl, alkylarylalkenyl, alkylarylalkynyl, alkenylarylalkyl, alkenylarylalkenyl, alkenylarylalkynyl, alkenylarylalkynyl, alkynylarylalkyl, alkynylarylalkenyl, alkynylarylalkynyl, alkylheteroarylalkyl, alkylheteroarylalkenyl, alkylheteroarylalkynyl, alkenyl hetero This refers to arylalkyls, alkenyl heteroarylalkenyls, alkenyl heteroarylalkynyls, alkynyl heteroarylalkyls, alkynyl heteroarylalkenyls, alkynyl heteroarylalkynyls, alkyl heterocyclylalkyls, alkyl heterocyclylalkenyls, alkyl hererocyclylalkynyls, alkenyl heterocyclylalkyls, alkenyl heterocyclylalkenyls, alkenyl heterocyclylalkynyls, alkynyl heterocyclylalkynyls, alkynyl heterocyclylalkynyls, alkylaryls, alkenylaryls, alkynylaryls, alkyl heteroaryls, alkenyl heteroaryls, and alkynylhereroaryls. For example, "arylalkylene" refers to a divalent alkylene in which one or more hydrogen atoms in the alkylene group are substituted with an aryl group, such as a (C3-C8) aryl group. Examples of (C3-C8)aryl-(C1-C6)alkylene groups include, but are not limited to, 1-phenylbutylene, phenyl-2-butylene, l-phenyl-2-methylpropylene, phenylmethylene, phenylpropylene, and naphthylethylene. The term "(C3-C8)cycloalkyl-(C1-C6)alkylene" refers to a divalent alkylene in which one or more hydrogen atoms in the C1-C6 alkylene group are substituted with a (C3-C8)cycloalkyl group.Examples of (C3-C8)cycloalkyl-(C1-C6)alkylene groups include, but are not limited to, 1-cycloproyl(proyl)butylene, cycloproyl(2)butylene, cyclopentyl-1-phenyl-2-methylpropylene, cyclobutylmethylene, and cyclohexylpropylene.
[0045] "Amino acid" refers to the structure H2N-C(R)-C(O)OH, where R is a side chain, e.g., an amino acid side chain. "Amino acid residue" refers to the remainder of an amino acid when incorporated into an amino acid chain, for example, when incorporated into the recognition reagents disclosed herein, having structures such as -NH-C(R)-C(O)-, H2N-C(R)-C(O)- (when it is the N-terminus of a polypeptide), or -NH-C(R)-C(O)OH (when it is the C-terminus of a polypeptide). "Amino acid side chain" refers to the side chain of an amino acid, including proteogenic or non-proteogenic amino acids. Amino acids have the following structure: [ka] It has, Here, R is the amino acid side chain. An example of an amino acid side chain that is not limited is shown in Figure 3. Glycine (H2N-CH2-C(O)OH) does not have a side chain.
[0046] "Peptide nucleic acids" refer to nucleic acid analogs, or DNA or RNA mimics, in which the sugar phosphodiester backbone of DNA or RNA is replaced with an N-(2-aminoethyl)glycine unit. Gamma-PNA (γPNA) has the following structure: [ka] An oligomer or polymer of γ-modified N-(2-aminoethyl)glycine monomer, where at least one of R1 or R2 bonded to the γ carbon is not hydrogen, so the γ carbon is a chiral center. If R1 and R2 are hydrogen (N-(2-aminoethyl)-glycine skeleton), or are identical, there is no such chirality with respect to the γ carbon. Incorporated PNA or γPNA monomer, [ka] In this specification, the remaining structure after integration into the oligomer is referred to as a PNA or γPNA "residue," and each residue has an R group that is identical or different to its base (nucleic acid base), such as adenine, guanine, cytosine, thymine, and uracil base, or other base, such as the monovalent and divalent bases described herein. Thus, the order of the bases on the PNA is its "sequence," similar to DNA or RNA. The sequence of nucleic acid bases in a nucleic acid or nucleic acid analog oligomer or polymer, such as a PNA or γPNA oligomer or polymer, binds to the complementary sequence of adenine, guanine, cytosine, thymine, and / or uracil residues in the nucleic acid or nucleic acid analog strand by nucleic acid base pairing in a Watson-Crick or Watson-Crick-like manner, essentially similar to double-stranded DNA or RNA.
[0047] The solubility and / or bioavailability can be increased by adding a "guanidine" or "guanidinium" group to the recognition reagent. Since PNA is produced in a similar manner to synthetic peptides, a simple way to add a guanidine group is to add one or more terminal arginine (Arg) residues to the N-terminus and / or C-terminus of the PNA, for example, the γPNA recognition reagent. Similarly, arginine side groups, [ka] , or a guanidine-containing portion, for example, [ka] Here, n is, for example, in the range of 1 to 5, but is not limited thereto, or its salt can be attached to the recognition reagent skeleton described herein. The guanidine-containing group is a group containing a guanidine moiety and may have fewer than 100 atoms, fewer than 50 atoms, for example, fewer than 30 atoms. In one embodiment, the guanidine-containing group has the structure: [ka] It has, Here, L is a linker according to any embodiment described herein, for example, a non-reactive aliphatic hydrocarbyl linker, such as a methylene, ethylene, trimethylene, tetramethylene, or pentamethylene linker. In some embodiments, the guanidine-containing group is structured as follows: [ka] It has, Here, n is between 1 and 5, and for example, the guanidine group may be arginine.
[0048] "Nucleic acid bases" include not only the major nucleic acid bases, namely adenine, guanine, thymine, cytosine, and uracil, but also modified purine and pyrimidine bases, such as, but not limited to, hypoxanthine, xanthene, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine, and 5-hydroxymethylcytosine. Figures 4 and 5A-5C also represent non-limiting examples of nucleic acid bases, including monovalent nucleic acid bases (e.g., adenine, cytosine, guanine, thymine, or uracil, which bind to one strand of a nucleic acid or nucleic acid analog), divalent nucleic acid bases (e.g., JB1-JB16 as described herein) which bind simultaneously to complementary nucleic acid bases on two strands of a nucleic acid, and "clamp" nucleic acid bases such as "G-clamps" which bind to complementary nucleic acid bases with increased strength. Further purines, purine-like, pyrimidine, and pyrimidine-like nucleic acid bases are known in the art, for example, disclosed in U.S. Patents 8,053,212, 8,389,703, and 8,653,254. For the divalent nucleic acid bases JB1–JB16 shown in Figure 5A, Table A shows the specificity of the various nucleic acid bases. In particular, the JB1–JB4 series bind complementary bases (CG, GC, AT, and TA), while JB5–JB16 bind mismatches and can therefore be used to bind two strands of matched and / or mismatched bases. Divalent nucleic acid bases are described in more detail in U.S. Patent Publication 20160083434A1 and International Publication WO / 2018 / 058091, both of which are incorporated herein by reference.
[0049] [Table A]
[0050] As used herein, "JB# nucleic acid base," for example, "JB4 nucleic acid base," refers to all nucleic acid bases having the JB4 designation representing C / G (bonding G / C), including JB4, JB4b, JB4c, JB4d, and JB4e. Exemplary γPNA structures that are not terminally modified with an aryl group in the manner described herein, but which may be terminally modified as described herein, are disclosed in International Publication No. 2012 / 138955, which is incorporated herein by reference. As shown herein, miniPEG γPNA contains one or more groups including a poly(ethylene glycol) (PEG, polyoxyethylene) moiety.
[0051] Complementarity refers to the ability of polynucleotides (nucleic acids) to hybridize (bond) with each other to form interchain base pairs. Base pairs are formed by hydrogen bonds between nucleotide units in antiparallel polynucleotide chains. Complementary polynucleotide chains can base pair (hybridize or bond) in the Watson-Crick manner (e.g., A to T, A to U, C to G) or in any other manner that allows for duplex formation. When using RNA instead of DNA, uracil, not thymine, is the complementary base to adenosine. Two sequences containing complementary sequences can hybridize if they form a duplex under specific conditions such as water, saline (e.g., physiological saline or 0.9% w / v saline) or phosphate-buffered saline), or under other stringency conditions such as 0.1×SSC (sodium citrate in saline) to 10×SSC, for example, but not limited to these, where 1×SSC is 0.15M NaCl and 0.015M sodium citrate in water. Hybridization of complementary sequences is influenced by, for example, salt concentration and temperature, and the melting temperature (T) increases with increasing mismatch and stringency. m) decreases. Perfectly matched sequences are considered "perfectly complementary," but one sequence (e.g., the target sequence in mRNA) may be longer than the other, as in the case of the small recognition reagents described herein for much longer target sequences to which they are linked, such as mRNA containing repeat extensions. Two complementary strands of nucleic acid are joined in antiparallel directions, with one strand oriented from 5' to 3' and the other from 3' to 5'. Both parallel and antiparallel orientations are possible with PNA, but antiparallel joining is preferred for γPNA.
[0052] According to one aspect of the invention, a gene recognition reagent is provided. In some aspects, the recognition reagent is self-concatenating, meaning that it contains a linking group, i.e., a terminal group, which is either covalently linked when hybridizing to an adjacent sequence on a target nucleic acid, or, in the case of non-covalent bonding, for example, by π-stacking. The recognition reagent contains three or more adjacent nucleic acid or nucleic acid analog monomer residues, for example, 3 to 10 or 3 to 8, which have terminal groups containing a sequence of nucleic acid bases for hybridizing (binding) to a repeat elongation sequence associated with repeat elongation disease, for example, by self-ligating or π-stacking, to target a double-stranded RNA hairpin sequence. In the case of polyQ disease, rCAG exp The double-stranded hairpin sequence formed from the sequence is the target sequence, and the reagent has a sequence of divalent nucleic acid bases selected from the following, EIF or (EIF) n (where n is an integer greater than or equal to 2, for example, 2, 3, 4, or 5), for example, EIFEIF, IFE, or (IFE) n (where n is an integer greater than or equal to 2, for example, 2, 3, 4, or 5), for example, IFEIFE, or FEI, or (FEI) n(where n is an integer greater than or equal to 2, e.g., 2, 3, 4, or 5), for example, FEIFEI, where E is a divalent nucleic acid base, e.g., JB3 or JB3b, bound C / G (Watson chain / Crick chain), F is a divalent nucleic acid base, e.g., JB4, JB4b, JB4c, JB4d, or JB4e, bound G / C, and I is a divalent nucleic acid base, e.g., JB6 or JB6b, bound A / A, so that it can bind to both hybridized strands of the hairpin formed from the CAG repeat, as shown in Figure 1.
[0053] The recognition reagents described herein combine the characteristics of small molecules, such as low molecular weight, ease of large-scale production, low production cost, cell membrane permeability, and desired drug kinetics, with sequence-specific recognition of oligonucleotides by Watson-Crick base pairing. The linking of oligomer recognition reagents has been demonstrated for both self-ligating and π-stacking linking methods and is described according to examples, which are intended to be illustrative rather than limiting in all embodiments. Therefore, in detailed implementations, the present invention is capable of many modifications that those skilled in the art can derive from the descriptions contained herein.
[0054] Examples of applications for the recognition reagents described herein include the treatment of genetic disorders involving small sequence repeat extensions, such as those listed in Table B.
[0055] [Table B]
[0056] Based on Table B, exemplary sequences of divalent nucleic acid bases of recognition reagents targeting the described gene products, including sequences in the 5' to 3' direction relative to the sense strand, are presented in Table C. Note that not all repeat sequences form a hairpin structure under normal conditions, but they can be induced into a triplex "hairpin" structure by the listed recognition reagents. Furthermore, the sequences listed in Table C are merely illustrative, and other sequences are expected to form triplex structures depending on the alignment of the folding sequences in the natural hairpin structure or the hairpin structure induced by the recognition reagent. Also, due to the repeat characteristics of the sequences, in the case of a three-base repeat, three different frameshifts are useful for each sequence (Table C, (GAA)). n (See the third column of the diagram), in the case of a tetranucleotide repeat, four different frameshifts are useful. Therefore, in the case of a trinucleotide sequence, the ability to rearrange the duplex binding alignment and shift the frame of the recognition reagent within each alignment (e.g., rCAG illustrated in Figure 1) is useful. exp The alignments within the sequence (EIF, IFE, and FEI) result in nine possible rearrangements of the unit recognition reagent, which can be repeated in the recognition reagent. For example, the sequence (GAA) n In this case, there is no folding alignment expected to form a hairpin duplex under normal conditions, but the three different sequences are expected to form a triplex structure. In Table C, the only repeat sequence that shows all sorts is (GAA) nThe sequence is shown, and for other repeats, only illustrative sequences are presented. In particular, for the sake of repeat structures, any of the sequences presented in Table C may be repeated 2 to 5 times, for example, G / AA / AA / G, also referred to as G / AA / AA / GG / AA / AA / G, and 13-6-14 (JB13 series-JB6 series-JB14 series) also includes 13-6-14-13-6-14. In Table C, nucleic acid base sequences are first identified by the target nucleic acid base to which they bind (for example, JB4 binds to G / C), and also by the "JB" reference number, for example "4", which refers to "JB4", and includes JB4, JB4b, JB4c, JB4d, and JB4e.
[0057] [Table C-1] [Table C-2]
[0058] In Table C, the unit recognition reagent target sequence (bound base) is listed as "B1 / B2," showing each base pair of the two strands bound by a single divalent nucleic acid base of the recognition reagent. When bound by the divalent nucleic acid base of the recognition reagent, B1 is the base from the first strand and B2 is the base from the second strand. Therefore, the sequence (CAG) n In this case, the unit target sequence is C / GA / AG / C, referring to the antiparallel sequence 5'-CAG-3', and the sequence, (CCG) n In this case, the sequence of the second alignment of the target sequence is C / GG / CC / C, which refers to the 5'-CGC-3' alignment that is antiparallel to 5'-CCG-3', while the sequence C / GC / CG / C refers to the antiparallel sequence of 5'-CCG-3'.
[0059] In relation to the binding of the recognition reagent, the "unit target sequence" is the shortest repeat sequence found in the repeat nucleic acid base sequence within the target nucleotide sequence, or a repeat double-stranded nucleic acid base sequence, including mismatches, which is aligned when a double-stranded sequence of the same or different nucleic acid molecule is bound in a triplex structure by a divalent recognition reagent as described herein.
[0060] In one embodiment, a recognition reagent is provided, comprising a peptide nucleic acid or γ-peptide nucleic acid skeleton having a first and second terminus and prepared from 3 or more, for example, 3 to 10, or 3, 4, 5, 6, 7, 8, 9, or 10, or 3 to 8 PNA or γ-PNA skeleton residues, a sequence of nucleic acid bases successively bound or ligated to multiple nucleic acid or nucleic acid analog skeleton residues, an -SH, -OH, or -SS-Lg portion at the first terminus of the skeleton, and an -C(O)-S-Lg or -C(O)-O-Lg portion at the second terminus of the skeleton (the leaving group Lg is linked to the skeleton by an ester or thioester bond), wherein the sequence of nucleic acid bases is complementary to a unit target sequence or two or more sequential repeats of a unit target sequence within a target sequence in a nucleic acid (two or more direct repeats of a specific sequence without gaps), so that when multiple recognition modules bind to adjacent sequences of the target sequence, they ligate to each other. In one embodiment, Lg is biocompatible and / or non-toxic. Non-limiting examples of Lg include substituted or unsubstituted (C1-C8) alkyl groups, substituted or unsubstituted (C3-C8) aryl groups, (C3-C8) aryl (C1-C6) alkylene groups, (C1-C8) carboxyl groups, groups that may be optionally substituted in amino acid side chains, or guanidine-containing groups, or [ka] It includes, Here, o is 1 to 20, each of R6 is independently an amino acid side chain, and R7 is either -OH or -NH2.
[0061] In another embodiment, a recognition reagent is provided, comprising a nucleic acid or nucleic acid analog skeleton having a first and a second end, prepared from 3 or more, for example, 3 to 10, or 3, 4, 5, 6, 7, 8, 9, or 10, or 3 to 8 nucleic acid or nucleic acid analog skeleton residues, optionally pre-organized stereochemically; a sequence of nucleic acid bases successively bound or linked to the plurality of nucleic acid or nucleic acid analog skeleton residues; a first aryl moiety linked to the first end of the nucleic acid or nucleic acid skeleton; and a second aryl moiety, which may optionally be identical to the first aryl moiety, and is bound to the second end of the nucleic acid or nucleic acid skeleton. The aryl moiety is independently a 2-5 ring condensed polycyclic aromatic moiety, for example, a substituted or unsubstituted aryl or heteroaryl moiety having a 2-5 condensed ring, for example, but not limited to, unsubstituted or substituted pentalene, indene, naphthalene, azulene, heptalene, biphenylene, as-indacene, s-indacene, acenaphthylene, fluorene, phenalene, phenanthrene, anthracene, fluorantene, acephenanthrylene, aceanthrylene, triphenylene, pyrene, chrysene, naphthacene / tetracene, pleiadene, picene, or perylene, which may be optionally substituted with one or more heteroatoms such as O, N, and / or S, for example, xanthene, riboflavin (vitamin B2), mangosteen, or mangiferin, which may be the same or different, and in some embodiments are the same. However, if the recognition reagent hybridizes with an adjacent sequence of the target nucleic acid, each shall be stackable with the aryl portion of the adjacent recognition reagent. The recognition reagent can bind to both strands of a hairpin formed from extended repeats associated with repeat extension diseases such as polyQ disease, as shown in Table B. In one embodiment, as described above, the sequence EIF or (EIF) n (where n is an integer greater than or equal to 2, for example, 2, 3, 4, or 5), for example EIFEIF, IFE, or (IFE) n (where n is an integer greater than or equal to 2, for example, 2, 3, 4, or 5), for example, IFEIFE, or FEI, or (FEI) n(where n is an integer greater than or equal to 2, e.g., 2, 3, 4, or 5), for example, to be used when joining CAG repeats that have FEIFEI.
[0062] In one aspect, the recognition reagent is self-ligating and structured: [ka] It has, Here, X is S or O, n is an integer from 1 to 6, m is an integer from 0 to 4, and R1 and R2 are each independently H, a guanidine-containing group, for example, [ka] Here, n = 1, 2, 3, 4, or 5, amino acid side chains, for example, [ka] Linear or branched (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C1-C8) hydroxyalkyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, (C3-C8) cycloalkyl(C1-C6) alkylene, which may be optionally substituted with ethylene glycol units containing 1 to 50 ethylene glycol moieties, -CH2-(OCH2-CH2) q OP1, -CH2-(OCH2-CH2) q -NHP1, -CH2-(SCH2-CH2) q -SP1, -CH2-(OCH2-CH2) r -OH, -CH2-(OCH2-CH2) r -NH2, -CH2-(OCH2-CH2) r -NHC(NH)NH2, or -CH2-(OCH2-CH2) r -SS[CH2CH2] sNHC(NH)NH2, where P1 is H, (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, or (C3-C8) cycloalkyl(C1-C6) alkylene, q is an integer from 0 to 50, and r and s are each an integer from 1 to 50. R3 is H or a leaving group, which in one embodiment is biocompatible and / or non-toxic and may be, for example, a substituted or unsubstituted (C1-C8) alkyl, a substituted or unsubstituted (C3-C8) aryl, a (C3-C8) aryl (C1-C6) alkylene (C1-C8) carboxy, a guanidine-containing group which may be optionally substituted with an amino acid side chain, or [ka] Here, o is 1 to 20, each of R6 is independently an amino acid side chain, R7 is -OH or -NH2, and R4 is (C1 to C 10 ) Divalent hydrocarbons or those substituted with one or more N or O moieties, such as -O-, -OH, -C(O)-, -NH-, -NH2, -C(O)NH- (C1~C 10 ) is a divalent hydrocarbon, R5 is -OH, -SH or a disulfide protecting group, and each of R is independently a nucleic acid base, and each of the multiple recognition modules produces a nucleic acid base sequence complementary to the target sequence of nucleic acid bases in the target nucleic acid, such that each of the multiple recognition modules binds to the target sequence of nucleic acid bases on the template nucleic acid and ligates with each other. In one embodiment, R5 has the structure -SH, -OH, or -SS-R8, where R8 is one or more amino acid residues, amino acid side chains, linear, branched or heterosubstituted (C1~C8)alkyl, (C2~C8)alkenyl, (C2~C8)alkynyl, (C1~C8)hydroxyalkyl, (C3~C8)aryl, (C3~C8)cycloalkyl, (C3~C8)aryl(C1~C6)alkylene, or (C3~C8)cycloalkyl(C1~C6)alkylene. In one embodiment, R1 or R2 is an amino acid side chain or a guanidine-containing group, for example, [ka] Here, n=1, 2, 3, 4, or 5. In one embodiment, R1 and R2 are different, either R1 is H and R2 is not H, or R2 is H and R1 is not H. For binding to native nucleic acids, e.g., RNA or DNA, R1 may be H and R2 is not H, thereby forming a "right-handed" L-γPNA. For example, a "left-handed" D-γPNA where R2 is H and R1 is not H will not bind to native nucleic acids. In one embodiment, the recognition reagent is one or more guanidine-containing groups, e.g., [ka] This is replaced by n=1, 2, 3, 4, or 5.
[0063] In one embodiment, R1 or R2 is -(OCH2-CH2) q OP1, -(OCH2-CH2) q -NHP1, -(SCH2-CH2) q -SP1, -(OCH2-CH2)r-OH, -(OCH2-CH2) r -NH2, -(OCH2-CH2) r -NHC(NH)NH2, or -(OCH2-CH2) r -SS[CH2CH2] s It is an (C1-C6) alkyl substituted with NHC(NH)NH2, where P1 is H, (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, or (C3-C8) cycloalkyl(C1-C6) alkylene, where q is an integer from 0 to 50, r is an integer from 1 to 50, and s is an integer from 1 to 50.
[0064] Recognition reagents with an aryl group at the end have the ability to link with the aryl group of the adjacent hybridizing recognition reagent as a result of non-covalent bonding, e.g., π-stacking. Therefore, in another embodiment, the structure: [ka] A recognition reagent (recognition module) having the following is provided: Here, R is an independent nucleic acid base, and each of R can be the same or a different nucleic acid base. n is an integer in the range of 1 to 6, for example, 1, 2, 3, 4, 5, or 6. Each B is independently a ribose-5-phosphate residue, a deoxyribose-5-phosphate residue, or a nucleic acid analog backbone residue, and in one embodiment, is a backbone residue of a structurally pre-organized nucleic acid analog, such as γPNA or LNA. L is independently a linker, for example, a non-reactive linker or a non-reactive, non-large linker, and each of L may be the same or different, Each Ar is independently a 2-5 ring condensed polycyclic aromatic moiety, for example, a substituted or unsubstituted aryl or heteroaryl moiety having a 2-5 condensed ring, for example, but not limited to, unsubstituted or substituted pentalene, indene, naphthalene, azulene, heptalene, biphenylene, as-indacene, s-indacene, acenaphthylene, fluorene, phenalene, phenanthrene, anthracene, fluorantene, acephenanthrylene, aceanthrylene, triphenylene, pyrene, chrysene, naphthacene / tetracene, pleiadene, picene, or perylene, which may be optionally substituted with one or more heteroatoms such as O, N, and / or S, for example, xanthene, riboflavin (vitamin B2), mangosteen, or mangiferin, which may be the same or different, and in some embodiments are the same. In some embodiments, when the recognition reagent hybridizes with adjacent sequences of the target nucleic acid, each Ar stacks with the Ar group of the adjacent recognition reagent.
[0065] A portion of a compound, such as an aryl portion or a nucleic acid base, is said to be covalently bonded to the recognition reagent skeleton and therefore "linked" to the skeleton. Depending on the chemical properties used to prepare the compound, the bond may be direct or mediated through a "linker," which is a portion covalently bonded to two other portions or groups. In one embodiment, a terminal aromatic (aryl) group is linked to the recognition reagent via a linker. The linker is a non-reactive portion that links the aromatic group to the skeleton of the recognition reagent, and in some embodiments, it consists of 1 to 10 carbon atoms (C1 to C1C1), which may be optionally substituted with heteroatoms, such as N, S, or O. 10 ), or a non-reactive bond, such as an amide bond (peptide bond) formed by reacting an amine with a carboxyl group. C1~C 10 Examples of alkylenes include linear or branched alkylene (divalent) moieties, which may optionally contain a cyclic moiety, such as methylene, ethylene, trimethylene, tetramethylene, pentamethylene, hexamethylene, heptamethylene, octamethylene, nonamethylene, or decamethylene moieties (i.e., -CH2-[CH2]) which may optionally contain an amide bond. n -where n=1~9). The linker is not bulky in that it substantially does not sterically shield or otherwise interfere with the binding of the recognition reagent to the target nucleic acid, nor does it interfere with the linking of the recognition reagent on the target nucleic acid. The linker has residual components resulting from the linking of aromatic groups and the skeleton of the recognition reagent, for example, [ka] Or, in a non-limiting example, due to binding of an acetate-substituted aryl compound (A), such as pyrene-1-acetic acid, to a Dab (n=1), Orn (n=2), or Lys (n=3) residue, [ka] That is the case.
[0066] In a further embodiment, the linker or linking group is an organic moiety that connects to two parts of a compound, for example, an organic moiety that is covalently bonded to two parts of a compound, for example, an aromatic group connected to the backbone of a recognition reagent, a nucleic acid base connected to a nucleic acid or nucleic acid analog backbone, and / or a guanidium group connected to a recognition reagent.Linkers are typically directly bonded or units such as atoms like oxygen or sulfur, C(O), C(O)NH, SO, SO2, SO2NH, or atomic chains, for example, but not limited to substituted or unsubstituted alkyls, substituted or unsubstituted alkenyls, substituted or unsubstituted alkynyls, arylalkyls, arylalkenyls, arylalkynyls, heteroarylalkyls, heteroarylalkenyls, heteroarylalkynyls, heterocyclylalkyls, heterocyclylalkenyls, heterocyclylalkynyls, aryls, heteroaryls, heterocyclyl, cycloalkyls, cycloalkenyls, alkylarylalkyls, alkylarylalkenyls, alkylarylalkynyls, alkenylarylalkyls, alkenylarylalkenyls, alkenylarylalkynyls, alkenylarylalkynyls, alkynylarylalkyls, alkynylarylalkenyls, alkynylarylalkynyls, alkylheteroarylalkyls, alkylheteroarylalkenyls, alkylheteroarylalkynyls, alkenyl This includes teloarylalkyl, alkenyl heteroarylalkenyl, alkenyl heteroarylalkynyl, alkynyl heteroarylalkyl, alkynyl heteroarylalkenyl, alkynyl heteroarylalkynyl, alkyl heterocyclylalkyl, alkyl heterocyclylalkenyl, alkyl herero cyclylalkynyl, alkenyl heterocyclylalkyl, alkenyl heterocyclylalkenyl, alkenyl heterocyclylalkynyl, alkynyl heterocyclylalkyl, alkynyl heterocyclylalkenyl, alkynyl heterocyclylalkynyl, alkylaryl, alkenylaryl, alkynylaryl, alkyl heteroaryl, alkenyl heteroaryl, and alkynyl herero aryl, wherein one or more carbon atoms, for example, methylene or methylidine (-CH=), may optionally have a heteroatom, for example, O, S, or N, a substituted or unsubstituted aryl, a substituted or unsubstituted heteroaryl, or a substituted or unsubstituted heterocycle in the middle or at the end.In one embodiment, the linker contains about 5 to 25 atoms, for example 5 to 20, 5 to 10, for example 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 atoms, or a total of 1 to 10 atoms, for example 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 C and heteroatoms, for example O, P, N, or S atoms.
[0067] With regard to binding to PNA, for example γPNA, suitable and available linkers include, for example, those using well-known peptide synthesis chemistry to add an amino acid to the recognition reagent, reacting the amine with a carboxyl group to form an amide bond, wherein the amino acid may be pre-modified with a chemical moiety, such as an aryl moiety or a guanidine group, by adding an aryl-modified amino acid to link the pyrenearyl moiety to the recognition reagent, as shown in the examples below, or by using arginine to provide a guanidine group. Linking to non-peptide nucleic acid analogs can be achieved by linking the amine-modified aromatic moiety to the recognition reagent, for example via a terminal phosphate, using any suitable linking chemistry, such as carbodiimide chemistry, for example EDC (EDAC, 1-ethyl-3-[3-dimethylaminopropyl]carbodiimide hydrochloride), as is widely known.
[0068] When linking an aromatic group to the backbone of a recognition reagent, the linker is of an appropriate size or length to position the aromatic group in a position for π-π stacking between the linking of the recognition reagent on the target nucleic acid sequence as described herein.
[0069] In a further embodiment, the recognition reagent has the following structure: [ka] It has, and here, Each of R is independently a nucleic acid base, and each of R can be the same or different nucleic acid base. n is an integer in the range of 1 to 6, for example, 1, 2, 3, 4, 5, or 6. R1 and R2 are each independently H, a guanidine-containing group, for example, [ka] Here, n = 1, 2, 3, 4, or 5, amino acid side chains, for example, [ka] , unsubstituted or substituted (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C1-C8) hydroxyalkyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, or (C3-C8) cycloalkyl(C1-C6) alkylene, -CH2-(OCH2-CH2) q OP1, -CH2-(OCH2-CH2) q -NHP1, -CH2-(OCH2-CH2-O) q -SP1, -CH2-(SCH2-CH2) q -SP1, -CH2-(OCH2-CH2) r -OH, -CH2-(OCH2-CH2) r -NH2, -CH2-(OCH2-CH2) r -NHC(NH)NH2, or -CH2-(OCH2-CH) r -SS[CH2CH2] s NHC(NH)NH2, where P1 is H, (C1~C8)alkyl, (C2~C8)alkenyl, (C2~C8)alkynyl, (C3~C8)aryl, (C3~C8)cycloalkyl, (C3~C8)aryl(C1~C6)alkylene, or (C3~C8)cycloalkyl(C1~C6)alkylene, q is an integer from 0 to 50, r is an integer from 1 to 50, s is an integer from 1 to 50, and in one embodiment, R1 and R2 are different, R1 is H and R2 is not H, or R2 is H and R1 is not H. R 10 or R 11 On the other hand, and R 12 , R 13 , or R 14One of these is -L-R3, where each of R3 is independently a 2-5 ring condensed polycyclic aromatic moiety, for example, a substituted or unsubstituted aryl or heteroaryl moiety having a 2-5 condensed ring, for example, but not limited to unsubstituted or substituted pentalene, indene, naphthalene, azulene, heptalene, biphenylene, as-indacene, s-indacene, acenaphthylene, fluorene, phenalene, phenanthrene, anthracene, fluoranthenatene, acephenanthrylene, aceanthr Len, triphenylene, pyrene, chrysene, naphthacene / tetracene, pleiaden, picene, or perylene, which may be optionally substituted with one or more heteroatoms such as O, N, and / or S, for example, xanthene, riboflavin (vitamin B2), mangosteen, or mangiferin, which may be the same or different, and in some embodiments may be the same, and in some cases, when the recognition reagent hybridizes with an adjacent sequence of the target nucleic acid, it stacks with the R3 group of the adjacent recognition reagent, Here, L is a linker, for example, an unreactive linker or an unreactive, non-large linker, and each of L may be the same or different, an amino acid residue, or a substituted or unsubstituted alkyl, substituted or unsubstituted alkenyl, substituted or unsubstituted alkynyl, arylalkyl, arylalkenyl, arylalkynyl, heteroarylalkyl, heteroarylalkenyl, heteroarylalkynyl, heterocyclylalkyl, heterocyclylalkenyl, heterocyclylalkynyl, aryl, heteroaryl, heterocyclyl, cycloalkyl, cycloalkenyl, alkylarylalkyl, alkylarylalkenyl, alkylarylalkynyl, alkenylarylalkyl, alkenylarylalkenyl, alkenylarylalkynyl, alkynylarylalkyl, alkynylarylalkenyl, alkynylarylalkynyl, alkylheteroarylalkyl, alkylheteroarylalkenyl, alkylheteroarylalkynyl, alkenylheteroarylalk It may also contain alkenyl heteroaryl alkenyl, alkenyl heteroaryl alkynyl, alkynyl heteroaryl alkyl, alkynyl heteroaryl alkenyl, alkynyl heteroaryl alkynyl, alkyl heterocyclyl alkyl, alkyl heterocyclyl alkenyl, alkyl herero cyclyl alkynyl, alkenyl heterocyclyl alkyl, alkenyl heterocyclyl alkenyl, alkenyl heterocyclyl alkynyl, alkynyl heterocyclyl alkyl, alkynyl heterocyclyl alkenyl, alkynyl heterocyclyl alkynyl, alkylaryl, alkenyl aryl, alkynyl aryl, alkyl heteroaryl, alkenyl heteroaryl, and alkynyl herero aryl, and one or more carbon atoms, for example, methylene or methylidine (-CH=), may optionally have heteroatoms, for example, O, S, or N, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, or substituted or unsubstituted heterocycles interrupted or at the terminal. And, R 10 , R 11 , R 12 , R 13 , and R 14The remainder is independently H, one or more adjacent amino acid residues, a guanidine-containing group, an amino acid side chain, a linear or branched (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C1-C8) hydroxyalkyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, (C3-C8) cycloalkyl(C1-C6) alkylene, which may be optionally substituted with ethylene glycol units containing 1 to 50 ethylene glycol moieties, -CH2-(OCH2-CH2) q OP1, -CH2-(OCH2-CH2) q -NHP1, -CH2-(SCH2-CH2) q -SP1, -CH2-(OCH2-CH2) r -OH, -CH2-(OCH2-CH2) r -NH2, -CH2-(OCH2-CH2) r -NHC(NH)NH2, or -CH2-(OCH2-CH2) r -SS[CH2CH2] s NHC(NH)NH2, where P1 is H, (C1~C8)alkyl, (C2~C8)alkenyl, (C2~C8)alkynyl, (C3~C8)aryl, (C3~C8)cycloalkyl, (C3~C8)aryl(C1~C6)alkylene, or (C3~C8)cycloalkyl(C1~C6)alkylene, q is an integer from 0 to 50, r is an integer from 1 to 50, and s is an integer from 1 to 50.
[0070] In one embodiment, R4 and R7 are -L-R3, and in another embodiment, R4 and R7 are -L-R3 and R11 and R14 are Arg. In one embodiment, the linker contains about 5 to 25 atoms, e.g., 5 to 20, 5 to 10, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 atoms, or a total of 1 to 10, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 C and heteroatoms, e.g., O, P, N, or S atoms. In another embodiment, one or more R1, R2, R 10 , R 11 , R12 , R 13 , or R 14 is -(OCH2-CH2) q OP1, -(OCH2-CH2) q -NHP1, -(SCH2-CH2) q -SP1, -(OCH2-CH2) r -OH, -(OCH2-CH2) r -NH2, -(OCH2-CH2) r -NHC(NH)NH2, or -(OCH2-CH2) r -SS[CH2CH2] s It is an (C1-C6) alkyl substituted with NHC(NH)NH2, where P1 is H, (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, or (C3-C8) cycloalkyl(C1-C6) alkylene, q is an integer from 0 to 50, r is an integer from 1 to 50, and s is an integer from 1 to 50.
[0071] In yet another embodiment, the recognition reagent comprises a PNA skeleton, and therefore its structure: [ka] It has, and here, n is an integer in the range of 1 to 8, including 1, 2, 3, 4, 5, 6, 7, or 8. m is an integer in the range of 1 to 5, for example, 1 to 3, including 1, 2, 3, 4, or 5. R2 is a guanidine-containing group, for example [ka] Here, n = 1, 2, 3, 4, or 5, amino acid side chains, for example, [ka] , unsubstituted or substituted (C1~C8)alkyl, (C2~C8)alkenyl, (C2~C8)alkynyl, (C1~C8)hydroxyalkyl, (C3~C8)aryl, (C3~C8)cycloalkyl, (C3~C8)aryl(C1~C6)alkylene, or (C3~C8)cycloalkyl(C1~C6)alkylene, -CH2-(OCH2-CH2) q OP1, -CH2-(OCH2-CH2) q -NHP1, -CH2-(OCH2-CH2-O) q -SP1, -CH2-(SCH2-CH2) q -SP1, -CH2-(OCH2-CH2) r -OH, -CH2-(OCH2-CH2) r -NH2, -CH2-(OCH2-CH2) r -NHC(NH)NH2, or -CH2-(OCH2-CH2) r -S-S[CH2CH2] s NHC(NH)NH2, wherein P1 is H, (C1~C8)alkyl, (C2~C8)alkenyl, (C2~C8)alkynyl, (C3~C8)aryl, (C3~C8)cycloalkyl, (C3~C8)aryl(C1~C6)alkylene or (C3~C8)cycloalkyl(C1~C6)alkylene, q is an integer from 0 to 50, r and s are each independently an integer from 1 to 50, R3 is an unsubstituted fused-ring polycyclic aromatic moiety, for example, pentalene, indene, naphthalene, azulene, heptalene, biphenylene, as-indacene, s-indacene, acenaphthylene, fluorene, phenalene, phenanthrene, anthracene, fluoranthene, acephenanthrylene, aceanthrylene, triphenylene, pyrene, chrysene, naphthacene / tetracene, pleiadene, picene, or perylene, R 11 , R 13 , and R 14 are each independently H, a guanidine-containing group, for example, [Chemical Formula] Here, n = 1, 2, 3, 4, or 5, It is an amino acid side chain, or one or more adjacent amino acid residues, for example, one or more Arg residues. In one embodiment, R3 is pyrene.
[0072] In another embodiment, R2 is -CH2-(OCH2-CH2) r -OH, where r is an integer in the range of 1 to 50, for example, 1 to 10, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, and in Example 2 it is 1. In another embodiment, R2, R is 1 or greater. 11 , R 13 , or R 14 is -(OCH2-CH2) q OP1, -(OCH2-CH) q -NHP1, -(SCH2-CH2) q -SP1, -(OCH2-CH2) r -OH, -(OCH2-CH2) r -NH2, -(OCH2-CH2) r -NHC(NH)NH2, or -(OCH2-CH2) r -SS[CH2CH2] s It is an (C1-C6) alkyl substituted with NHC(NH)NH2, where P1 is H, (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, or (C3-C8) cycloalkyl(C1-C6) alkylene, q is an integer from 0 to 50, r is an integer from 1 to 50, and s is an integer from 1 to 50.
[0073] Some PNA-based recognition reagents exhibit no chirality, in one example, the γ carbon (to which R1 and R2 are bound) is oriented either R2 is H and R1 is not H (R), or the γ carbon is oriented, for example, R1 is H and R2 is not H (S). Furthermore, in one example, if present, R 10 , R 11 , R 12 , R 13 , and / or R 14At this position, one or more, or all, of the chiral amino acid residues may be L-amino acids. In another example, if present, R 10 , R 11 , R 12 , R 13 , and / or R 14 At this position, one or more or all of the chiral amino acid residues may be D-amino acids.
[0074] We also provide any of the aforementioned pharmaceutically acceptable salts.
[0075] Any pharmaceutically acceptable salt of any of the compounds described herein may also be used in the methods described herein. The pharmaceutically acceptable salt forms of the compounds described herein may be prepared by conventional methods known in the pharmaceutical field and may be included as veterinarily acceptable salts. For example, if a compound contains a carboxyl group, a suitable salt therefor may be formed by reacting the compound with a suitable base to provide a corresponding base addition salt. Non-limiting examples include alkali metal hydroxides, such as alkaline earth metal hydroxides like potassium hydroxide, sodium hydroxide and lithium hydroxide, barium hydroxide and calcium hydroxide, alkali metal alkoxides such as potassium ethanolate and sodium propanolate, and various organic bases such as piperidine, diethanolamine, and N-methylglutamine.
[0076] Non-limiting examples of pharmaceutically acceptable basic salts include aluminum, ammonium, calcium, copper, ferric, ferrous, lithium, magnesium, manganese, manganese, potassium, sodium, and zinc salts. Examples of salts derived from pharmaceutically acceptable, non-toxic organic bases include, but are not limited to, primary, secondary, and tertiary amines, substituted amines including naturally occurring substituted amines, cyclic amines, and basic ion exchange resins, such as salts of arginine, betaine, caffeine, chloroprocaine, choline, N,N'-dibenzylethylenediamine (benzathine), dicyclohexylamine, diethanolamine, diethylamine, 2-diethylaminoethanol, 2-dimethylaminoethanol, ethanolamine, ethylenediamine, N-ethylmorpholine, N-ethylpiperidine, glucamine, glucosamine, histidine, hydravamin, isopropylamine, lidocaine, lysine, meglumine, N-methyl-D-glucamine, morpholine, piperazine, piperidine, polyamine resins, procaine, purine, theobromine, triethanolamine, triethylamine, trimethylamine, tripropylamine, and tris-(hydroxymethyl)-methylamine (tromethamine).
[0077] Examples of pharmacokinetically acceptable salts that are not limited include acetate, adipine, alginate, arginate, aspartate, benzoate, besylate (benzenesulfonate), bisulfate, bisulfite, bromide, butyrate, camphorate, camphorsulfonate, caprylate, chloride, chlorobenzoate, citrate, cyclopentanepropionate, digluconate, dihydrogen phosphate, dinitrobenzoate, dodecyl sulfate, ethanesulfonate, fumarate, galacterate, galacturonate, glucoheptanate, gluconate, glutamate, and glyceroline. Examples include salts, hemisuccinate, hemisulfate, heptanoate, hexanoate, hippurate, hydrochloride, hydrobromide, hydroiodide, 2-hydroxyethanesulfonate, iodide, isethionate, isobutyrate, lactate, lactobionate, malate, maleate, malonate, mandelate, metaphosphate, methanesulfonate, methylbenzoate, monohydrogen phosphate, 2-naphthalenesulfonate, nicotinate, nitrate, oxalate, oleate, pamoate, pectinate, persulfate, phenylacetate, 3-phenylpropionate, phosphate, phosphonate, and phthalate.
[0078] Compound salt forms are also considered pharmaceutically acceptable salts. General non-limiting examples of compound salt forms include bicarbonate tartrate, diacetate, difumarate, dimeglumine, diphosphate, disodium, and trihydrochloride.
[0079] Therefore, as used herein, “pharmaceutically acceptable salt” is intended to mean the active moiety (drug) containing the salt form of any compound described herein. The salt form preferably imparts improved and / or desired pharmacokinetic properties to the compounds described herein.
[0080] The compositions described herein can be administered by any effective route. Examples of delivery routes include, but are not limited to, topical delivery (e.g., on the skin, by inhalation, by enema, intraocular, intraaural, intra-ocular, intra-aural, intra-oral, intra-gastric tube, and transrectal), and parenteral delivery (e.g., intravenous, intra-arterial, intramuscular, intracardiac, subcutaneous, intraosseous, intradermal, intramedullary, intraperitoneal, transdermal, ion implantation, transmucosal, epidural, and intravitreous), with oral, intravenous, intramuscular, and transdermal approaches being preferred in many cases. Preferred forms of administration may include single-dose or multi-dose vials or other containers, such as medical syringes, containing compositions with active components useful for treating recurrent elongation disorders, as described herein.
[0081] The drug regimen can be adjusted to provide the desired optimal response (e.g., a therapeutic or prophylactic response). For example, a single bolus may be administered, multiple fractional doses may be administered over a long period of time, or the composition may be administered continuously or in a pulsed manner, with the dose or partial dose being proportionally reduced or increased at regular intervals, for example, every 10, 15, 20, 30, 45, 60, 90, or 120 minutes, every 2 to 12 hours daily, or every other day, as indicated by the requirements of the treatment situation. In some cases, it may be particularly advantageous to formulate the composition, such as parenteral or inhalation compositions, in unit dose form for ease of administration and uniformity of dosage. The specification of unit dose form is determined and directly depends on (a) the unique characteristics of the active compound and the specific therapeutic or prophylactic effect to be achieved, and (b) the inherent limitations in the technique of formulating such active compounds for the treatment of individual hypersensitivity.
[0082] Useful forms of administration include intravenous, intramuscular, or intraperitoneal solutions, oral tablets or liquids, topical ointments or creams, and transdermal devices (e.g., patches). In one embodiment, the compound is a sterile solution containing an active portion (drug or compound) and a solvent, such as water, saline, lactated Ringer's solution, or phosphate-buffered saline (PBS). Further excipients, such as polyethylene glycol, emulsifiers, salts, and buffers, may be included in the solution.
[0083] Therapeutic / pharmaceutical compositions are prepared according to the acceptable dispensing procedures described, for example, in Remington: The Science and Practice of Pharmacy, 21st edition, ed. Paul Beringer et al., Lippincott, Williams & Wilkins, Baltimore, MD Easton, Pa. (2005) (see Chapters 37, 39, 41, 42, and 45 for examples of powder, liquid, parenteral, intravenous, and oral solid formulations and methods for preparing such formulations).
[0084] The drug regimen can be adjusted to provide the desired optimal response (e.g., a therapeutic or prophylactic response). For example, a single bolus may be administered, multiple fractional doses may be administered, or the composition may be administered continuously or in a pulsed manner, in doses or partial doses at regular intervals, for example, every 10, 15, 20, 30, 45, 60, 90, or 120 minutes, every 2 to 12 hours daily, or every other day, with proportional decreases or increases as desired depending on the severity of the treatment situation. In some cases, compositions such as parenteral or inhalation compositions may be particularly advantageous to be prescribed in unit dose forms for ease of administration and uniformity of dosage. The details of the unit dose form are determined and directly depend on (a) the unique characteristics of the active compound and the specific therapeutic or prophylactic effect to be achieved, and (b) the limitations inherent in the technique of formulating such active compounds for the treatment of hypersensitivity in individuals.
[0085] Useful forms of administration include intravenous, intramuscular, or intraperitoneal solutions, oral tablets or liquids, topical ointments or creams, and transdermal devices (e.g., patches). In one embodiment, the compound is a sterile solution containing an active portion (drug or compound) and a solvent such as water, saline, lactated Ringer's solution, or phosphate-buffered saline (PBS). Additional excipients such as polyethylene glycol, emulsifiers, salts, and buffers may be included in the solution.
[0086] In addition to a gene recognition reagent, a method is also provided for conjugating a nucleic acid containing an extended repeat associated with a repeat extension disorder, the method comprising contacting the nucleic acid containing the extended repeat with any embodiment of the gene recognition reagent described herein. In one embodiment, the method is carried out in vitro or ex vivo. In another embodiment, the method is carried out in vivo by administering a suitable formulation containing the gene recognition reagent to a patient in an amount effective to knock down the expression of a gene having an extended repeat or to treat the patient.
[0087] In another embodiment, (CAG) in a sample obtained from a patient n A method is provided for identifying the presence of nucleic acids containing extended repeats associated with repeat extension disorders. The method involves contacting a nucleic acid sample obtained from a patient with a gene recognition reagent in any form as described herein. The binding and ligation of the gene recognition reagent then occurs in the presence of the target sequence in the nucleic acid in the sample, so the binding and ligation indicate the presence of the target sequence in the nucleic acid in the sample. The binding and ligation of the gene recognition reagent can be detected and / or quantified using any effective method. The presence of nucleic acids containing extended repeats in the target sequence is associated with repeat extension disorders, and in this respect, the patient can be treated for repeat extension disorders by, for example, administration to the patient in a pharmaceutical dosage form containing the gene recognition reagent described herein.
[0088] Thus, a method is provided for treating patients with diseases associated with extended repeats of nucleotide sequences. Examples of diseases are listed in Table B, or polyQ diseases, such as Huntington's disease, DRPLA (dentatorubral-pallidoluysian atrophy), SBMA (spinal and bulbar muscular atrophy), SCA1 (spinocerebellar ataxia type 1), SCA2 (spinocerebellar ataxia type 2), SCA3 (spinocerebellar ataxia type 3 or Machado-Joseph disease), SCA6 (spinocerebellar ataxia type 6), SCA7 (spinocerebellar ataxia type 7), and SCA17 (spinocerebellar ataxia type 17). For polyQ diseases, such as Huntington's disease, (CAG) n The treatment involves administering to the patient a composition or pharmaceutical dosage form containing a gene recognition reagent, such as those described herein, selected to bind characteristic repetitive sequences of specific diseases. The composition in pharmaceutical dosage form is administered to the patient in an effective amount and dosage regimen for treating the disease. Accordingly, the use of compositions containing gene recognition reagents as described herein is also provided for the treatment of diseases associated with extended repeats of nucleotide sequences, such as PolyQ diseases, e.g., Huntington's disease, DRPLA (dentatorubral-pallidoluysian atrophy), SBMA (spinal and bulbar muscular atrophy), SCA1 (spinocerebellar ataxia type 1), SCA2 (spinocerebellar ataxia type 2), SCA3 (spinocerebellar ataxia type 3 or Machado Joseph disease), SCA6 (spinocerebellar ataxia type 6), SCA7 (spinocerebellar ataxia type 7), and SCA17 (spinocerebellar ataxia type 17).
[0089] In some embodiments, the disclosure provides a compound comprising a) a series of peptide nucleic acid (PNA) units, i) a first unit, ii) a final unit, and iii) at least one intermediate unit between the first unit and the final unit, wherein each unit in the series of units comprises A) a skeletal portion, where 1) the skeletal portion of the first unit is covalently bonded to the skeletal portion of one other unit, 2) the skeletal portion of the final unit is covalently bonded to the skeletal portion of one other unit, and 3) the skeletal portion of each intermediate unit is covalently bonded to the skeletal portions of two other units, and B) a divalent nucleic acid base covalently bonded to the skeletal portion, b) a first aryl portion covalently bonded to the first unit, and c) a final aryl portion covalently bonded to the final unit.
[0090] In some embodiments, the disclosure provides a compound comprising a) a series of PNA units, i) a first unit, ii) a final unit, and iii) at least one intermediate unit between the first unit and the final unit, wherein each unit in the series of units comprises A) a skeletal portion, where 1) the skeletal portion of the first unit is covalently bonded to the skeletal portion of one other unit, 2) the skeletal portion of the final unit is covalently bonded to the skeletal portion of one other unit, and 3) the skeletal portion of each intermediate unit is covalently bonded to the skeletal portions of two other units, and B) a series of units comprising a divalent nucleic acid base covalently bonded to the skeletal portion, and b) two aryl portions, one covalently bonded to the first unit and the other covalently bonded to the final unit.
[0091] In some embodiments, the disclosure provides compounds in which a first aryl portion is covalently bonded to a first unit by a first linker portion, and a final aryl portion is covalently bonded to a final unit by a final linker portion. In some embodiments, the disclosure provides compounds in which the first linker portion and the final linker portion each independently contain a guanidine group. In some embodiments, the disclosure provides compounds in which the first linker portion and the final linker portion each independently contain an amino acid residue. In some embodiments, the disclosure provides compounds in which the first linker portion and the final linker portion each independently contain three guanidine-containing amino acid residues.
[0092] In some embodiments, the disclosure provides compounds in which the first linker moiety and the final linker moiety each independently comprise three adjacent arginine residues. In some embodiments, the disclosure provides compounds in which the series of units comprises 3 to 8 units. In some embodiments, the disclosure provides compounds in which the series of units comprises γ-PNA. In some embodiments, the disclosure provides compounds in which the first aryl moiety and the final aryl moiety each independently comprise a 2-5 ring condensed polycyclic aromatic moiety, such as a polyaromatic hydrocarbon, such as pyrene. In some embodiments, the disclosure provides compounds in which the first aryl moiety and the final aryl moiety are identical. In some embodiments, the disclosure provides compounds further comprising ethylene glycol units, such as diethylene glycol units.
[0093] In some embodiments, the disclosure provides compounds that present divalent nucleic acid bases in a complementary order to a target nucleic acid sequence. In some embodiments, the disclosure provides compounds in which the target nucleic acid sequence is associated with repeat elongation disorders. In some embodiments, the disclosure provides compounds in which the repeat elongation disorder is myotonic dystrophy type 1 (DM1) or myotonic dystrophy type 2 (DM2). In some embodiments, the disclosure provides compounds in which, when two compounds hybridize with a nucleic acid, the first aryl portion of one compound and the final aryl portion of the other compound stack.
[0094] In some embodiments, the disclosure is of formula:H- L Arg-CAGCAG- L Arg-NH2(P1), H- L Arg- L Dab(Pyr)-CAGCAG- L Orn(Pyr)- L Arg-NH2(P2), H- L Arg- L Orn(Pyr)-CAGCAG- L Orn(Pyr)- L Arg-NH2(P3), H- L Arg- L Lys(Pyr)-CAGCAG- L Lys(Pyr)- L Arg-NH2(P4), H- L Arg- L Lys(Pyr)-CATCAG- L Lys(Pyr)- L Arg-NH2(P5), or H- L Arg- L Lys(Pyr)-CTGCTG- L Lys(Pyr)- LWe provide compounds of Arg-NH2(P6), where Orn is ornithine, Dab is diaminobutyric acid, and Pyr is a carboxyl-functionalized aromatic compound, such as pyrene, including pyrene-1-carboxylic acid, pyrene-2-carboxylic acid, pyrene-4-carboxylic acid, pyrene-1-acetic acid, pyrene-2-acetic acid, or pyrene-4-acetic acid.
[0095] In some embodiments, the disclosure provides a method for binding nucleic acids, the method comprising contacting a nucleic acid with a compound of the disclosure, the compound binding to the nucleic acid upon contact. In some embodiments, the disclosure provides a method for knocking down mRNA expression in cells, the method comprising contacting cells with a compound of the disclosure, the compound knocking down mRNA expression in cells upon binding to a DNA sequence in the cell corresponding to the mRNA.
[0096] In some embodiments, the disclosure provides a gene recognition reagent comprising a nucleic acid or nucleic acid analog skeleton having a first end and a second end and having 3 to 8 ribose-5-phosphate, deoxyribose-5-phosphate, or nucleic acid analog skeleton residues; a divalent nucleic acid base linked to a plurality of ribose-5-phosphate, deoxyribose-5-phosphate, or nucleic acid analog skeleton residues, which may be the same or different; a first aryl moiety linked to the first end of the nucleic acid or nucleic acid analog skeleton by a linker; and a second aryl moiety, which may optionally be the same as the first aryl moiety, linked to the second end of the nucleic acid or nucleic acid analog skeleton by a linker.
[0097] In some embodiments, the Disclosure provides a detection method comprising: a) hybridizing a first probe nucleic acid comprising one or more divalent nucleic acid bases, a first end connected to a first emitter portion, and a second end connected to a second emitter portion, with a first repeat portion of a target nucleic acid; and b) hybridizing a second probe nucleic acid comprising one or more divalent nucleic acid bases, a first end connected to a third emitter portion, and a second end connected to a fourth emitter portion, with a second repeat portion of the target nucleic acid, wherein i) the first and second repeat portions of the target nucleic acid are associated with repeat extension defects; ii) when the first and second probe nucleic acids are bound to the target, a first or second emitter portion is obtained that is very close to a third or fourth emitter portion; and iii) the presence of a first or second emitter portion very close to a third or fourth emitter portion causes a change in the emission wavelength of the very close emitter portion.
[0098] In some embodiments, the detection method further includes detecting changes in the emission wavelength of very close emitter portions. In some embodiments, hybridizing the first probe nucleic acid to the first repeat portion increases the affinity of the second probe nucleic acid to the second repeat portion. In some embodiments, the presence of the first or second emitter portion adjacent to the third or fourth emitter portion induces a π-π stacking interaction between the first or second emitter portion and the third or fourth emitter portion. In some embodiments, the target nucleic acid is obtained directly from a biological sample. [Examples]
[0099] Example 1 NMR, X-ray, and biochemical studies revealed that the rCAG-hairpin structure is relatively dynamic compared to canonical RNA duplexes, with the internal AA ridges exhibiting large-amplitude motion. Such molecular scaffolds, containing internal AA mismatches in all two canonical GC / CG base pairs (Figure 1(A)), analogous to "deep holes" in a road, provide different viable receptor-like binding sites for exogenous ligands. Relatively short nucleic acid ligands targeting the rCAG-hairpin structure were developed (Figure 1(B)). Using Janus bases (J bases, i.e., JB), E, I, and F (Figure 1(C)), these can form divalent H-bond interactions with nucleic acid bases on both strands of the RNA double helix, possessing a structurally pre-organized MPyPNA backbone (Figure 1(D and E)). The J bases described herein total 16 and are part of a larger set of bifacial nucleic acid recognition elements designed to bind to all 16 possible RNA base pair combinations (Figures 5A–5C). They differ from other J bases, including the "Janus wedge," in that they are uniform in shape, size, and chemical functionality, and therefore, in combination with the modular format, can recognize any combination of RNA base pairs.
[0100] Molecular mechanics (MD) simulation To evaluate the feasibility of the bivalent recognition design concept, MD simulations were performed on ligand LG1 bound to an RNA duplex containing four rCAG repeats. The computer modeling was simplified by excluding the C-terminal lysine residue and substituting the MP side chain with a methyl group (MeγPNA). CEO, AIA, and GFC triads were constructed, optimized using the 6-31 base system, and grafted onto each RNA and MeγPNA backbone. The structures of the bound RNA-LG1-RNA complexes were generated using the Ambertool NAB module. The final structures were solvated with water molecules and ions to minimize energy and simulated for 100 ns. The resulting complexes remained fairly stable throughout the simulation, with the four distinct LG1 ligands tightly fitted between the two RNA strands (Figure 6(A)). The number of H bonds remained constant throughout the simulation, at 5 for each of the CEO and GFC triads and 4 for AIA (Figure 6(B)). However, attempts to simulate binding with fewer than 4 LG1 ligands failed. The complex unraveled, and some formed twisted structures due to the instability of the terminal base pairs (Figure 6(D and E)).
[0101] Synthesis of chemical components and ligands J bases E and F were synthesized together with the corresponding JB-MPγPNA monomers 1-3 (Figure 7). J base I was not prepared as it is commercially available. E and F were synthesized according to Scheme 1 (Figure 8) and Scheme 2 (Figure 9), respectively. The Boc protection / deprotection sequences in compounds 5-9 were found to be preferable for improving the chemical yield of the NBS reaction and suppressing side reactions in the subsequent condensation and cyclization steps. Further protection of 12 was found to be preferable for the Stille coupling to proceed smoothly. This was achieved using anhydrous Boc and carried out at high temperature. Despite the apparent bulkiness, there were no problems in coupling E to the MPγPNA skeleton or in assembling the corresponding monomers on the MBHA resin. The core structure of F, 4-amino-2-(methylthio)pyridine-5-carbonitrile (17) was prepared according to a published protocol. Subsequently, the cyano group was converted to amidine, followed by Boc protection to obtain 21. Attempts to completely protect the extracyclic amines with a large excess of Boc anhydride under various conditions were unsuccessful. This resulted in the formation of multiple spots on the TLC with varying numbers of Boc groups, which were difficult to separate by column chromatography. To address this challenge, Boc protection was performed in two steps. 22 was subjected to oxidation and hydrolysis, followed by alkylation and hydrolysis to obtain F. Once prepared, E(15), F(26), and I(27) were coupled to the MPγPNA skeleton (Figure 10, Scheme 3). Removal of the Alloc protecting group yielded the desired monomers 1-3. Ligands LG1, LG2, and LG3, along with LG2P which is oriented opposite (parallel) to LG2, were prepared on MBHA resin using PAL linker and HBTU as coupling reagents. LG2P was included as a negative control because the parent ligand LG2 showed the highest likelihood of binding the rCAG repeat. Lysine residues were incorporated at the C-terminus to improve water solubility. After the final monomer coupling was complete, the ligand was cleaved from the resin, precipitated with diethyl ether, purified by RP-HPLC, and confirmed by MALDI-TOF mass spectrometry (Table D).
[0102]
Table D
Chem
[0103] Target Selection and Sample Preparation A series of model RNA targets were selected for binding studies (Figure 11). The R11X series contains 24 rCXG repeats (Figure 11(A)), which, when adopting a secondary hairpin structure, provided 11 binding sites for ligands (Figure 11(B)). R11A contains a perfect match (A-A), while R11U and R11C contained U-U and C-C mismatches, respectively, in the ligand binding region. R11G (X=G) was not investigated, since it has been shown to form G-quadruplex as well as hairpin structures, which may confound the interpretation of experimental results. WS (Watson strand) and CS (Crick strand) (Figure 11(C)), single-stranded RNA targets with the LG2 binding site underlined, and HP (Figure 11(D)), a double-stranded hairpin containing a single ligand binding site, were also investigated. Samples were prepared by incubating pre-annealed RNA with each ligand in 0.1×PBS buffer (1 mM NaPi, 13.7 mM NaCl, 0.27 mM KCl, pH 7.4) at 37°C. This particular buffer was selected based on our initial screen, where we found that R11A adopts a stable hairpin structure and is capable of ligand binding.
[0104] Spectroscopic Identification of Ligand Binding The binding properties of the ligands were determined using a combination of UV-Vis and circular dichroism (CD). UV-Vis measurements revealed that all three ligands could bind to R11A. Evidence of their interaction could be gathered from the pale and deep color shifts in the absorption of [R11A + ligand] at 252 and 330 nm. The UV absorption in the 275–375 nm regime corresponds to the π–π* transition of the E base. This finding was supported by CD data revealing a significant increase and redshift in the signal at approximately 270 nm. Of this series, LG2 showed the greatest difference in CD amplitude. CD titration of LG2 with R11A reached a saturation point in an 11:1 ratio, consistent with the predicted number of ligand-binding sites on R11A (Figure 12). We hypothesized that R11A adopts a uniform hairpin structure because it is annealed by heat before ligand addition; however, other combinations of intramolecular and intermolecular folding are also possible. In contrast, no significant differences were observed in the CD signal after incubation of LG2 mismatches R11U (Figure 13(A)) and R11C (data not shown), single-stranded WS and CS (Figure 13(B)), double-stranded HP including a single binding site (Figure 13(C)), or mismatch orientation LG2P and R11A (Figure 13(D)). Since the results for R11U and R11C mismatches were almost identical in all respects, all subsequent discussions regarding base mismatches will focus on R11U. In summary, these results indicate that ligand-RNA interactions occur in a sequence-specific and orientation-specific manner, with double-stranded RNA targets preferred over single-stranded RNA targets, and those with multiple consecutive binding sites being more advantageous than isolated ones.
[0105] Confirmation of ligand binding by fluorescence measurement To further verify the findings of UV-Vis and CD, the fluorescence signal of the ligand with or without the RNA target was measured after excitation at 330 nm (λmax of E base). Significant fluorescence quenching of LG3 was observed compared to LG1 and LG2. Such a dramatic decrease in fluorescence signal can be attributed to photoinduced electron transfer quenching, which is expected to be most efficient for LG3 since the fluorophore (E) is stacked between two other J bases. This interpretation is consistent with the observation (data not shown) that the fluorescence signal of LG3 is completely restored by heating. For all three ligands, the luminescence signal was further decreased by the addition of R11A, which was the most dramatic for LG2, with a 60% decrease for LG2 compared to 35% and 33% for LG1 and LG3, respectively. Similar fluorescence quenching patterns were observed upon incubation of LG2 with R11U, [WS+CS], HP, and LG2P with R11A (Figure 14), albeit not as dramatic as the fluorescence quenching pattern of LG2 and R11A. Similar observations have been made for related classes of bicyclic pyrimidine analogs, indicating that the fluorescence yields of these fused-ring fluorophores are highly sensitive to the local environment. To test the hypothesis that non-specific binding of ligands to RNA can lead to fluorescence quenching, each RNA target was incubated with pentamidine followed by addition of LG2, and vice versa. The data showed that the fluorescence signals of LG2 with R11U and [WS+CS], as well as the fluorescence signal of LG2P with R11A, were fully restored, except for the case of LG2 with R11A, where the fluorescence signal remained suppressed (Figure 14, inset). This result is consistent with LG2 binding to R11A through a mode different from that of pentamidine, probably via defined divalent H-bonding interactions (Figures 1B and C).
[0106] NMR titration To gain insight into the binding mode of LG2, a series of multinuclear and multidimensional NMR experiments were performed using LG2, R11A, and combinations. Samples were prepared in the same 0.1× PBS buffer as before, containing H2O:D2O in a 9:1 volume ratio. NOESY data were collected at mixing times of 300, 200, and 100 ms to obtain interproton distances, while COSY experiments were performed to map the proton spin system of each residue. The H-binding interaction between LG2 and R11A was expected to result in linewidth expansion and downfield chemical shift of the iminoproton signal of the canonical GC / CG pair due to base pair opening, as well as the emergence of a novel set of iminoproton signals due to the formation of H-bindings between the ligand and RNA. Consistent with the CD data, 1 ¹H-NMR experiments revealed that R11A adopts a stable hairpin structure in 0.1×PBS buffer, as evidenced by its sharp iminoproton signal at 12.35 ppm. Variable temperature measurements further supported this chemical shift assignment, showing that the peak intensity gradually decreased with increasing temperature, while the peak intensity of the non-exchangeable proton remained nearly constant. Partial assignment of nucleic acid base protons was performed by NOESY experiments. As expected, the addition of LG2 gradually expanded and downfield-shifted the iminoproton signal of R11A (Figure 15). Furthermore, the chemical shifts of C4-NH and G2-NH were significantly affected, indicating their interaction with their ligands. Despite efforts, due to degeneracy in the repeating sequence, the assignment, and therefore monitoring, of the A6-NH chemical shift during titration of R11A with LG2 was not possible. Nevertheless, the results indicate that GC / CG base pair opening is mediated by the ligand. However, as expected for H-bond formation between ligand and RNA, no novel iminoproton signals were observed in the 10–20 ppm regime.
[0107] Template-mediated native chemical ligation (NCL) Weak and transient interactions between ligands and RNA, as observed with LG2 and R11A, are unlikely to result in significant biological responses. To further improve the binding affinity of LG2, a second LG2 derivative, LG2N (Figure 16(A)), was synthesized, containing a C-terminal thioester and an N-terminal cystine. This ligand was prepared using a second-generation N-acylurea linker. The dual-function probe design utilizes the sterically pre-organization of MPγPNA to prevent the two functional groups from spontaneously reacting with each other upon reduction of the disulfide bond. The cystine group provided better control over probe handling. Previous studies have shown that ligands of the same length containing all native nucleic acid bases have a reduced (acyclic) half-life of approximately 1 hour at physiological temperature (37°C). A similar half-life was expected for LG2N*. Figure 16(B) shows the reaction pathway of LG2N after reduction of the disulfide bond and the expected chemical state of LG2N in a reducing intracellular environment. In the presence of an RNA target, the ligand was expected to form a transient divalent H-bond interaction with the nucleic acid bases of the adjacent RNA target as a result of intermolecular base stacking (step 2). Upon cleavage of the disulfide bond, the ligand undergoes template-mediated NCL to form a ligated product that more firmly binds to the RNA template (step 3). In the absence of an RNA target, the ligand self-inactivates by undergoing an intramolecular NCL reaction to form a cyclic product (step 4).
[0108] To verify the prediction that LG2N remains stable for a certain period after the reduction of the disulfide bond before undergoing a cyclization reaction, LG2N *The reaction progress of the parent compound M induced in situ was monitored (Figure 17(A)). For quantification of the reaction product, intramolecular NCL was chosen over HPLC and other analytical techniques because it has a relatively short timescale and can be achieved in less than 5 minutes using the former. Attempts were made to quench the reaction with an electrophile before HPLC analysis; however, such efforts were unsuccessful due to the rapid intramolecular reaction of the ligand. 4-mercaptophenol (4MP) and 2-mercaptoethanol (2ME) were investigated as possible reducing agents. Similar dynamic profiles were obtained for parent compound M with both; however, the former resulted in an intermediate with a mass-to-charge ratio (m / z) that overlapped with the hydrolyzed parent compound, making it difficult to differentiate them from each other. For this reason, 4MP was chosen for reaction monitoring and gel shift assays using MALDI-TOF, while 2ME was used in subsequent melt experiments due to its clarity in the nucleic acid base absorption region. Our data revealed that upon addition of 4MP to the parent compound M, two reaction species corresponding to transesterification (M#) and C-terminal thioester (M##) were initially formed (Figure 17(A)). After persisting for <15 minutes, reactive LG2N * It was converted to an intermediate and finally to a cyclic product (cLG2N) (Figure 17(B)). LG2N * It was stable at physiological temperature. It was present as the main product at 1 hour of reduction, constituting approximately 75% of the total products in the mixture, with the remaining 25% being mainly cLG2N. LG2N * It has a half-life of approximately 3 hours, which is almost three times longer than its natural counterpart. LG2N * We hypothesized that the chemical stability of the E base is due to the enlarged aromatic ring size, which leads to better base stacking interactions and, consequently, less structural flexibility of the ligand.
[0109] Next, UV melting experiments were conducted to determine the effect of template-mediated NCL on the thermal stability of RNA. Samples were prepared in 0.1×PBS buffer with a ratio of LG2N to R11A binding sites of 2:1, incubated at 37°C for 16 hours, and then subjected to UV melting analysis. As expected, in the absence of a reducing agent, LG2N was found to be the melting transition of R11A (T m The effect on ) is minimal, and at most approximately +1°C ΔT m The same was true for LG2, which does not contain either the C-terminal thioester or the N-terminal cysteine. However, the results were similar regardless of whether the incubation of [LG2N+R11A] in the presence of 2ME was performed at ambient temperature or 37°C. m A significant increase was observed in both cases in the 59–68°C range. The derivative of the melting curve showed a broad S-shaped distribution, suggesting the presence of various ligation products formed and bound to the RNA template. This finding was predicted for template-mediated synthesis. m Despite their similarities, the two melt curves exhibit different patterns, with the one incubated at high temperatures showing inverse light absorption in the 25–50°C range. This melt behavior is characteristic of hydrophobic / aromatic interactions, a phenomenon previously observed in thermophilic foldomers and other aromatic systems, such as perylene and pyrene. We hypothesized that the inverse intensity distribution is due to the interaction of the J bases of the ligated product, which were formed in greater quantities at 37°C than at ambient temperature. Upon heating, the hydrophobic interaction became more pronounced up to a certain point (approximately 50°C), beyond which repulsion occurred. The enhanced thermal stability of R11A is consistent with the formation of NCL products and their binding to RNA templates.
[0110] R11A T mTo confirm that the enhancement in the sample was due to template-mediated ligation and the binding of the resulting product to the RNA template, MALDI-TOF analysis was performed on UV-melted samples, followed by heating. The formation of the ligated product is evident from the appearance of a new peak with a gradually increasing m / z value at a step size of approximately 1319 Daltons, depending on the ligand mass (Figure 18(A)). However, these ligated products were not observed with mismatched R11U (Figure 18(B)) or LG2N alone without a target (Figure 18(C)). Similarly, there was no evidence that the ligated product was formed with single-stranded [WS+CS] or double-stranded HP (data not shown). This result is consistent with the occurrence of template-mediated synthesis mediated by the R11A hairpin structure.
[0111] Selective binding of linked ligands LG2N is rCAG exp A competitive binding assay was performed to determine whether it could be distinguished from normal repeats. r(CAG) within the normal (R24A, n=24, same as R11A) and pathogenic (R127A, n=127) ranges were tested. n Repeat mismatch (CUG) 96Two RNA targets were used, along with a control (R96U). When each hairpin motif was adopted, these RNA transcripts provided 11, 62, and 47 binding sites for LG2N. Two sets of samples were prepared, one in 0.1×PBS and the other in a physiologically relevant buffer (10 mM NaPi, 150 mM KCl, 2 mM MgCl2, pH 7.4). In both sets, equimolar mixtures of R24A (100 nM chain concentration, 1.1 μM binding site) and R127A (18 nM chain concentration, 1.1 μM binding site) were incubated with various concentrations of LG2N, 4MP, and TCEP at 37°C for 24 hours. The inclusion of TCEP ensured complete reduction of the disulfide bonds. The reaction mixtures were analyzed on agarose gels and stained with SYBR-Gold for visualization. The investigation in Figure 19 revealed that LG2N preferentially binds to R127A rather than R24A, as evidenced by the faster disappearance rate of the top two bands compared to the bottom band in lanes 2-5. Such binding events occurred only in the presence of a reducing agent and with perfectly matched RNA targets, because no evidence of binding was observed in the absence of 4MP and TCEP (comparing lane 6 and lane 1) or with mismatched R96U (comparing lane 8 and lane 7). As expected for ligand-RNA complexation, no shifted band formation was observed. This suggests that SYBRGold was unable to intercalate the RNA-ligand complex. This observation was not surprising, given that ligand binding is known to cause dramatic changes in the conformation of RNA (Figure 6). In contrast, no evidence of ligand binding occurring at physiologically relevant ionic strengths was observed (Figure 20), meaning that under such conditions, the RNA hairpin could not mediate template-mediated ligand oligomerization.
[0112] Over 20 neuromuscular disorders, including HD, myotonic dystrophy type 1 (DM1) and type 2 (DM2), spinocerebellar ataxia (SCA), fragility X syndrome (FXS), Friedreich's ataxia (FRDA), and subpopulations of amyotrophic lateral sclerosis (ALS, or Lou Gehrig's disease), are manifestations of unstable repeat elongation. While elongation in the coding region of a gene can lead to altered protein function, elongation occurring in non-coding regions can cause disease without interfering with the acquisition of toxic RNA function, leading to the unintended production of harmful polypeptides by protein sequencing and / or RAN translation. Many of these hereditary disorders are autosomal dominant, requiring elongation (or mutation) in only one allele to cause the disease. Therefore, one way to interfere with such disease pathways is to target the affected allele or the corresponding gene transcript. Between the two, the latter offers greater accessibility to exogenous molecules and greater recognition specificity to divalent ligands, such as JB-MPyPNA, due to its tendency to adopt an incomplete hairpin structure.
[0113] Due to the elongated genetic penetrance of rCAG repeats in HD, and, but not limited to, other neurodegenerative disorders including MJD, SBMA, DRPLA, and SCA, we have focused our efforts on rCAG repeats. By using molecules that selectively bind to mutant htt transcripts and interfere with the disease pathway, we can treat not only HD but also other related genetic disorders. The nucleic acid ligands described herein differ from conventional antisense drugs in several ways. First, they bind to the CAG-RNA hairpin motif via divalent H-bond interactions. Second, they are relatively small in size, the length of a single triplet repeat unit. Therefore, they are easier to chemically synthesize, structurally modify, and scale up by solution-phase chemistry rather than solid-phase chemistry. Third, in terms of molecular weight on the outer edge of small molecules, such “miramolecule” systems generally have more favorable pharmacokinetic properties than typical antisense drugs. Fourth, J-base recognition is more sequence-specific and selective. Typically, single-base mismatches that occur on one side of natural nucleic acid bases occur on both sides of J bases, making them more specific and selective, and thus preferred over single-stranded RNA targets due to the large difference in binding free energy between the two resulting products. However, unlike small molecules or riboswitches, J base recognition follows a set of predetermined rules via divalent Watson-Crick H-binding interactions, rather than a more diverse combination of attractive forces. Furthermore, JB-ligand designs are modular. Therefore, ligands can be prepared and modified to bind to any sequence of RNA repeats. This allows for further improvement of recognition selectivity by adjusting the ligand binding affinity and the ionic strength of the buffer, so weak and transient interactions between JB-ligands and rCAG repeats of normal length at low salt concentrations are desirable in the initial design stage.
[0114] material and method UV melting analysis All UV melting samples were prepared by mixing the ligand with the RNA target in 0.1×PBS buffer at the indicated concentrations, annealing by incubation at 90°C for 5 minutes, and then gradually cooling to room temperature. UV melting curves were collected using an Agilent Cary UV-Vis300 spectrometer with a thermoelectrically controlled multi-cell holder. UV melting spectra were collected at 260 nm by monitoring UV absorption at a rate of 1°C per minute for both heating experiments (from 25°C to 95°C) and cooling experiments (from 95°C to 25°C). The cooling and heating curves were nearly identical, indicating that the hybridization process is reversible. The recorded spectra were smoothed using a 20-point adjacent averaging algorithm. The melting temperature of the complex was determined using the first derivative of the melting curve.
[0115] CD analysis Samples were prepared in 0.1×PBS buffer. All spectra represent the average of at least 15 scans collected at a rate of 100 nm / min between 200 and 375 nm in a 1 cm path length cuvette at 25°C. The CD spectrum from the buffer solution was subtracted from the sample spectrum and then smoothed by a 5-point neighboring average algorithm. Steady-state fluorescence measurements. All steady-state fluorescence samples were prepared by mixing the ligand with RNA in 0.1 l×PBS buffer at the indicated concentrations, annealed by incubation at 90°C for 5 minutes, and then gradually cooled to 37°C. After incubation of the samples at 37°C for 1 hour, measurements were taken. Steady-state fluorescence data were collected at 25°C using a Cary Eclipse Fluorescence spectrometer. (ε ex =330nm, ε em (=340~600nm).
[0116] Competitive binding assay I purchased R24A from IDT. R127A[r(CAG) 127 ] and R96U[r(CUG)96 The preparations were made as described above. R24A, R126A, and R96U were prepared in 0.1× PBS buffer and annealed by heating at 90°C for 5 minutes, followed by gradual cooling to room temperature. The ligands and RNAs were mixed at the indicated concentrations and incubated at 37°C for 24 hours. The samples were then loaded onto a 2% agarose gel using 1× Tris-borate buffer and separated by electrophoresis at 100V for 25 minutes. The gels were stained with SYBR-Gold and visualized with a UV-Transilluminator.
[0117] MD simulation, ligand synthesis, and monomer synthesis The procedure is presented in the attached supplemental information incorporated herein.
[0118] general technology All starting materials and chemicals were purchased from commercial suppliers and used without further purification, unless otherwise specified. All reactions were carried out under anhydrous conditions using dry solvents and a nitrogen atmosphere, unless otherwise specified. Solvent removal in vacuum refers to distillation using a rotary evaporator attached to a high-vacuum pump. Products obtained as solids, syrups, or liquids were dried under high vacuum. Analytical thin-layer chromatography was performed on pre-coated silica plates (60F-254, 0.25 mm thick), and compounds were visualized by UV light, staining with ninhydrin or iodine, or by heat as a developing agent. NMR spectra were recorded at either 300 or 500 MHz using CDCl3, DMSO, and D2O as solvents and TMS as an internal standard. The following abbreviations were used to describe multiplicity: s=singlet, d=doublet, t=triplet, q=quadruplet, m=multiplet, ddt=triplet-doublet-doublet, dt=triplet-doublet, td=doublet-triplet, ABq=ABquadruplet, dd=doublet-doublet, tt=triplet-triplet, br=broad wtc. High-resolution mass spectrometry (HRMS) was performed using an ESI-ESI-TOF mass spectrometer and a DART analyzer. Low-resolution mass spectrometry (LRMS) was performed using an ESI-Ion Trap-MS mass spectrometer. Unless otherwise specified, oligomer purification by RP-HPLC was performed using a C18 column (15 cm × 2.1 mm, 5 μm particles) with a flow rate of 1.0 mL / min and a linear elution gradient (Method A) from 95% H2O (0.01% TFA) to 95% MeCN (0.01% TFA) over 40 minutes. MALDI spectra were measured using a TOF / TOF spectrometer. Absorption profiles were recorded using a UV-Vis spectrophotometer. Circular dichroism (CD) spectra were recorded using a spectropolarimeter. Fluorescence (emission) was recorded using a fluorescence spectrophotometer.
[0119] MD Simulation The C-terminal lysine residue was removed from γPNA, and the MP side chain was substituted with a methyl group (MeγPNA, published X-ray crystal structure, PDB-ID3PA0). The CEG, AIA, and GFC triads were constructed using Chimera 1 and set to HF / 6-31G with a Gaussian distribution. * The structure was optimized according to the criteria. The helical structure of the bound RNA-HD1-RNA complex was generated from a triad using the Ambertool NAB module, and a modified MeγPNA backbone from the crystal structure was grafted onto the helical structure (initial structure as shown in Figure 2A). Four such RNA-HD1-RNA complexes were obtained, with the RNA being a duplex containing four rCAG repeats, and the number of HD1 ligands bound to the RNA varying from 1 to 4. Each RNA-(HD1) n - The RNA complex was solvated with TIP3P4 water molecules, and ions were added to neutralize the charge. The system was energy-minimized and then heated to 300 K under harmonic suppression of 25 kcal / mol / Å with respect to the atoms of the RNA and HD1 ligand. The suppression was gradually released in a series of six short simulations, and finally, unsuppressed NPT simulations were performed for 100 ns at 300 K and a pressure of 1 bar using a Nose-Hoover thermostat and a Parinello-Rahman barostat, respectively. Electrostatic interactions were handled using the particle mesh Ewald method, and all simulations were performed using the χOL3-modified GROMACS Amber99-parmbsc0 force field for RNA, and the general Amber force field was used to generate the force field parameters for the HD1 ligand.
[0120] Resin Loading [(MP)(PAL)]
[0121] (1) MP loading. 1 g of MBHA resin (1 mmol / g, Peptide International, RMB-2100-PI) was soaked in DCM in a reaction vessel for 1 hour, then washed with DCM (3×) and 5% DIEA in DCM (3×). See Figure 21. After confirming that amine groups were neutralized by the Kaiser test (blue), the following solutions were sequentially added into a 15 mL canonical tube: 3.5 mL of NMP, 450 μL of solution A, 460 μL of solution B, and 550 μL of solution C. The mixture was vortexed for 10 seconds, left to stand for 3 minutes, then added to the resin in the reaction vessel. The reaction was allowed to proceed for 1 hour with gentle shaking. The reaction mixture was discharged under positive air pressure, and the resin was washed with DMF (3×), DCM (3×), 5% DIEA in DCM (1×), and DCM (3×). The resin was capped by adding a capping solution to the reaction vessel, and the vessel was gently stirred for 45 minutes (2×). After discarding the final capping solution, the resin was washed with DCM (3×), and completion of capping was confirmed by the Kaiser test (pale yellow).
[0122] Solution A: 43 mg of Fmoc-MP (0.200 M) in 500 μL of NMP Solution B: 87 μL of DIEA (0.500 M) in 913 μL of pyridine Solution C: 55 mg of HATU (0.201 M) in 750 μL of NMP Capping solution (Ac₂O / NMP / pyridine: 1 / 12 / 2): 2 mL of Ac₂O, 4 mL of NMP, and 4 mL of pyridine
[0123] (2) Loading of PAL and Lys (1 g of resin from the preceding section). The Fmoc protecting group was removed by treating the resin with 20% piperidine solution in DMF (2×) for 7 minutes each, followed by washing with DMF (3×) and DCM (3×). After confirming successful removal of Fmoc by the Kaiser test (blue), a coupling solution was prepared by mixing the following materials: 3 mL of 0.2 M monomer solution, 1.5 mL of 0.52 M DIEA solution, and 1.5 mL of 0.39 M HBTU solution. The mixture was activated for 3 minutes and then added to the resin. The reaction was allowed to proceed for 30 minutes with gentle stirring of the reaction vessel. Completion of the reaction was confirmed by the Kaiser test (pale yellow). Unreacted amines were capped by washing the resin with DMF (3×), 5% DIEA solution (1×), and DCM (1×), and the reaction vessel was gently stirred for 30 minutes. After removing the capping solution, the resin was washed with 20% piperidine (2×), DMF (2×), and DCM (2×).
[0124] Monomer solution: 3 mL of NMP contains 303 mg of Fmoc-PAL (0.200 M) and 3 mL of NMP contains 281 mg of Fmoc-Lys(Boc)-OH (0.200 M). DIEA solution: 364 μL of DIEA (0.52M) in 3.638 mL of DMF. HBTU solution: 740 mg of HBTU (0.39 M) in 5 mL of DMF. Capping solution: (Ac2O / NMP / Pyridine: 1 / 25 / 25 vol)
[0125] Fmoc deprotection The Fmoc protecting group was removed by treating the resin with a 20% piperidine solution in DMF (0.5 mL per 50 mg of resin) (2×) for 7 minutes each, followed by washing with DMF (3×) and DCM (3×). Fmoc deprotection was confirmed by the Kaiser test (blue).
[0126] Monomer coupling After confirming that Fmoc had been successfully removed by the Kaiser test (blue), the resin was washed with DMF (4×) and DCM (5×). The coupling solution was prepared by mixing the following materials: 150 μL of 0.2 M monomer solution, 75 μL of 0.52 M DIEA solution, and 75 μL of 0.39 M HBTU solution. The mixture was activated for 10 minutes. After washing the resin with pyridine (1×), the monomer solution was added to the resin. The reaction was allowed to proceed for 2-4 hours with gentle stirring of the reaction vessel. Completion of the reaction was confirmed by the Kaiser test (pale yellow). The resin was washed with DMF (3x) and DCM (3x).
[0127] Capping. Unreacted amines were capped with a freshly prepared capping solution. The capping solution (1:25:25, acetic anhydride:NMP:Py) was added to the resin, and the reaction vessel was stirred for 4 minutes. The resin was washed with DMF (4×) and DCM (5×).
[0128] Cutting After the final Fmoc deprotection step, the resin was washed with DMF (5×) and DCM (8×). The resin was dried under vacuum for 15 minutes. A freshly prepared 95% TFA:5% m-cresol (0.6 mL per 50 mg of resin) was added to the dried resin, and the reaction vessel was held in standby mode. After 1 hour at room temperature, the TFA solution was collected in a canonical centrifuge tube. The cutting solution (0.5 mL per 50 mg of resin) was added to the resin again, and the reaction vessel was left to stand for 30 minutes. The TFA solution was then combined with the previously collected solution.
[0129] Precipitation Cold-dried diethyl ether (-60°C, 14 mL) was added to the collected TFA solution and shaken. Precipitation occurred within 30 minutes. The precipitated oligomer was collected by centrifugation and washed with cold diethyl ether (2×).
[0130] purification The crude oligomer was dissolved in 0.5 mL of 95% water, 5% acetonitrile, and 0.1% TFA. The resulting solution was purified by reverse-phase analysis column.
[0131] Example 2 - Monomer Synthesis
[0132] [ka]
[0133] Monomer E(1) Palladium tetrakiss (triphenylphosphine) (28.22 mg, 24.42 μmol) was added at room temperature to a mixture of phenylsilane (0.12 mL, 0.98 mmol) and allyl monomer 29a (0.60 g, 0.49 mmol) in anhydrous dichloromethane (10 mL). Complete conversion of the starting materials to the corresponding acid was observed by TLC over 15 hours (eluent: EA:Hex(SO:SO), Rf 0.6 for 29a, Rf 0.0 for product (dragging)). Silica gel was added to the reaction mixture at room temperature, and the solvent was removed using a rotary evaporator. The silica gel absorbed crude product was purified by flash silica gel chromatography to obtain the desired product. Yield: 0.44 g, 7 S%. 1 H NMR (S00.13 MHz, DMSO-d6, rotamer): δ 1.08 / 1.09(s / s,9H),1.25-1.28(m,18H),1.37-1.38(m,18H),1.41 / 1.6S(s / s,9H),3.12-3.64(m,13H),3.77-3.90(m,2H), 3.90-4.00(m,2H),3.99-4.4.34(m,4H),7.21-7.49(m,5H),7.54-7.79(m,2H),7.89(d,J=7.2 Hz,2H),8.55 / 8.64(s / s,1H), 13C NMR (125.77 MHz, DMSO-d6, rotamer): δ 27.2-27.6(18C,tBu),35.8,46.7,54.9,59.7,60.6 / 60.6(2C),65.6 / 65.6,69.6 / 69 .6,69.8 / 69.9,70.4 / 70.4,72.2 / 72.2,82.1(3C),82.2,83.2,83.4,86.2,108.1,12 0.1(2C),125.2,127.0(4C),127.6(4C),140.7(2C),143.8 / 143.8,147.9,149.2 / 14 9.3,149.3,149.5,149.7,155.0 / 155.9,156.3 / 156.5,162.3,166.6 / 166.6,HRMS:[C 61 H 85 N7O 17 Calculated value of m / z for [+H]: 1188.6080, measured value: 1188.6071.
[0134] [ka]
[0135] Monomer I(2) The above procedure for compound 1 was used to prepare this compound, starting with compound 29b (3.10 g, 4.39 mmol) as the starting material. TLC (eluent: MeOH:DCM (5.95), Rr=0.2 for 29b, Rf=0.0 for 2). Yield: 2.22 g, 76%. 1 H NMR(300.13 MHz,DMSO-d6):δ 1.09(bs,9H),3.11-3.52(m,15H),3.87-3.96(m,2H),4.24(dd,J=13.4,5.4 Hz,4H),7.19-7.44(m,SH),7.69(d,J=7.5 Hz,2H),7.89(d,J=7.4 Hz,2H),10.71-10.75(m,1H),11.06(bs,1H),12.46(bs,1H), 13C NMR (126 MHz, DMSO-d6, rotamer): δ 27.8(3C),47.2 / 47.2,61.1(3C),65.6 / 66.0,70.1(2C),70.3 / 70.4,70.9(3C),72.7,107.8 ,120.6-128.1(10C),139.7 / 140.1,141.2(2C),144.3,151.8,156.3,164.6,170.9,HRMS:[C 34 H 42 N4O 10 Calculated value of m / z for [+H]: 667.2979, measured value: 667.2988.
[0136] [ka]
[0137] Monomer F(3) The above procedure for compound 1 was used to prepare this compound, starting with compound 29c (2.50 g, 2.00 mmol) as the starting material. TLC (eluent: EA:Hex (50:50), Rf=0.6 for 29c, Rf=0.0 for 3 (dragging)). Yield: 1.94 g, 80%. 1 H NMR (300.13 MHz, DMSO-d6, rotamer): δ 1.08(s,9H),1.35(s,9H),1.36(s,9H),1.41 / 1.42(s / s,18H),3.20-3.62(m,20H),3.7 7-4.15(m,4H),4.18-4.34(m,4H),4.82-5.35(m,2H),7.17-7.46(m,5H),7.70(d,J=7.0 Hz,2H),7.88(d,J=7.5 Hz,2H),8.93 / 8.95(s / s,1H), 13C NMR (125.77 MHz, DMSO-d6, rotamer): δ 27.3-27.5(18C,tBu),46.7 / 46.7,60.6(2C),60.7(2C),69.6(2C),69.7,69.9,70.4 (3C),72.2,82.1,82.1,82.9,82.9,83.3,83.4,108.0,110.8,118.4,120.1(2C),12 3.1,125.2 / 125.3,127.0(4C),127.6(4C),140.7(2C),143.8,146.9,149.5,156.6,157.1,162.4 / 162.4.HRMS:[C60H85N7O19+H] m / z calculated value: 1208.5980, measured value: 1208.5959.
[0138] [ka]
[0139] (4-chloropyridine-2-yl)-carbamate tert-butyl ester (5) In an ice bath, 556 mL of a stirred solution of 2-amino-4-chloropyridine (1) (50.00 g, 0.39 mol) in anhydrous THF was mixed with Et3N (81.37 mL, 0.58 mol) over 10 minutes in an inert atmosphere. After stirring for another 10 minutes at the same temperature, Boc2O (84.88 g, 0.39 mol) was added dropwise to THF (100 mL) over 30 minutes, followed by the addition of a catalytic amount of DMAP (4.75 g, 0.04 mol). After stirring overnight at room temperature, the reaction mixture was diluted with ethyl acetate (200 mL). The combined organic layer was filtered through a sintering funnel, the filtrate was washed with a saturated solution of NaHCO3, and the mixture was dried over Na2SO4. The solvent was removed using a rotary evaporator to obtain the pure product. Yield: 67.59 g, 76%. 1 H NMR (300.13 MHz, CDCl3): δ 1.54(s,9H),6.95(dd,J=5.5,1.8 Hz,1H),8.12(d,J=1.9 Hz,1H),8.24(d,J=5.5 Hz,1H),9.96(s,1H), 13C NMR(125.77 MHz,CDCl3):δ 28.5(3C),81.5,112.9,118.7,146.2,148.3,152.6,153.8,ESI-MS:[C 10 H 13 The calculated m / z value for ClN2O2+H is 229.1, and the measured value is 229.2.
[0140] [ka]
[0141] 2-((tert-butoxycarbonyl)amino)-4-chloronicotinate ethyl(6) In a dry ice-acetone bath (-78°C), 534 mL of a stirred solution of (4-chloropyridine-2-yl)carbamate tert-butyl ester (30.50 g, 0.13 mol) in anhydrous THF was mixed with n-BuLi (24.86 mL, 0.27 mol, 11 M) over 1 hour under an inert atmosphere. After stirring at -78°C for a further 30 minutes, ClCO2Et (13.93 mL, 0.15 mol) was added dropwise over 10 minutes. The reaction was checked by TLC, and after the starting materials were completely consumed, the reaction mixture was slowly quenched with 1% HCl (200 mL). The resulting reaction mixture was transferred to a separatory funnel containing 10% HCl (50 mL) and EA (200 mL). The crude product was extracted from the aqueous layer with EA, and the combined organic layer was dried over Na2SO4. The resulting solvent was removed using a rotary evaporator, and the crude substance was purified by column chromatography to obtain ethyl 2-((tert-butoxycarbonyl)amino)-4-chloronicotinate as a pale yellow solid. Yield: 29.68 g, 74%. 1 H NMR(500.13 MHz,CDCl3):δ 1.40(t,J=7.2 Hz,3H),1.50(s,9H),4.41(q,J7.2 Hz,2H),7.08(d,J=5.2 Hz,1H),8.31(d,J=5.2 Hz,1H),8.62(s,1H), 13C NMR(125.77 MHz,CDCl3):δ 14.1,28.3(3C),62.3,81.8,120.8,144.7,149.9,150.0,151.3,151.8,164.9,HRMS:[C 13 H 17 The calculated m / z value for ClN2O4+Na is 323.0774, and the measured value is 323.0777.
[0142] [ka]
[0143] 2-amino-4-chloronicotinate ethyl(7) Ethyl 2-((tert-butoxycarbonyl)amino)-4-chloronicotinate (29.00 g, 0.10 mol) was added slowly at room temperature to a mixture of concentrated HCl (52.60 mL, 0.58 mol) and dichloromethane (64 mL). After 2 hours at room temperature, the completion of the starting material was monitored by TLC, and the reaction mixture was then cooled to 0°C and neutralized with 2N NaOH solution. The reaction mixture was transferred to a separatory funnel containing 100 mL of dichloromethane, and the compound was extracted from the aqueous layer into the organic layer. The combined dichloromethane layer was dried over anhydrous Na2SO4 and concentrated. The crude material was purified by silica gel column chromatography to obtain pure ethyl 2-amino-4-chloronicotinate as a white solid. Yield: 18.38 g, 95%. 1 H NMR(500.13 MHz,CDCl3):δ 1.41(t,J=7.1Hz,3H),4.42(q,J=7.1 Hz,2H),6.13(bs,2H),6.68(d,J=5.3 Hz,1H),7.97(d,J=5.4 Hz,1H), 13 ¹³C NMR (125.77 MHz, CDCl3): δ 14.2, 61.8, 108.3, 115.9, 145.6, 151.2, 159.9, 166.4; HRMS: Calculated m / z value for [C8H9ClN2O2+H]: 201.0431, Measured value: 201.0422.
[0144] [ka]
[0145] Ethyl 2-amino-4-chloro-5-bromonicotinate (8) To a solution of 7 (18.00 g, 0.09 mol) in anhydrous DMF, NBS (18.36 g, 0.10 mol) was added in several portions at room temperature. TLC analysis of the resulting solution revealed that 8 was completely consumed after 4 hours. The resulting mixture was quenched with water (100 mL) and extracted with ethyl acetate (EA, 2 × 100 mL). The combined organic layers were washed with brine, dried over anhydrous sodium sulfate (Na2SO4), filtered, and concentrated. The resulting residue was purified by silica gel column chromatography (100-200 mesh) to obtain only ethyl 2-amino-4-chloro-5-bromonicotinate (8) as a pale yellow solid. Yield: 22.07 g, 88%. 1 H NMR (500.13 MHz, CDCl3): δ 1.41 (t, J=7.1 Hz, 3H), 4.43 (q, J=7.1 Hz, 2H), 5.86 (s, 2H), 8.24 (s, 1H), 13 ¹³C NMR (125.77 MHz, CDCl3): δ 14.2, 62.3, 109.7, 110.4, 144.4, 152.8, 158.0, 165.9. HRMS: Calculated m / z value for [C8H8BrClN2O2+H]: 278.9536, measured value: 278.9541.
[0146] [ka]
[0147] 2-((tert-butoxycarbonyl)amino)-4-chloro-5-bromonicotinate ethyl(9) Diisopropylethylamine (DIEA, 27.42 mL, 157.41 mmol) was added to a chilled solution of 8 (20.00 g, 71.55 mmol) in anhydrous dichloromethane (CH2Cl2), followed by the addition of triphosgene (7.64 g, 25.76 mmol). After 1 hour at 0°C, tert-butanol (33.96 mL, 357.76 mmol) was added over 5 minutes, and the ice bath was removed. The progress of the reaction was monitored by TLC, and after 24 hours at room temperature, the starting material was completely consumed. The resulting mixture was quenched with water and extracted with dichloromethane (CH2Cl2, 2 × 100 mL). The collected organic layer was washed with brine and a saturated solution of sodium bicarbonate (NaHCO3), dried over anhydrous sodium sulfate, filtered, and concentrated. The resulting crude material was purified by silica gel column chromatography (100-200 mesh) to obtain 2-((tert-butoxycarbonyl)amino)-4-chloro-5-bromonicotinate ethyl (5) as a white solid. Yield: 13.58 g, 50%. 1 H NMR (500.13 MHz, CDCl3): δ 1.41 (t, J=7.2 Hz, 3H), 1.50 (s, 9H), 4.42 (q, J7.2 Hz, 2H), 8.37 (s, 1H), 8.57 (s, 1H), 13 C NMR(125.77 MHz,CDCl3):δ 14.1,28.3(3C),62.6,82.2,113.9,116.3,119.9,144.0,149.5,151.7,164.2, HRMS:[C 13 H 16 The calculated m / z value for [BrClN2O4+H] is 379.0060, and the measured value is 379.0063.
[0148] [ka]
[0149] 5-Bromo-2-((tert-butoxycarbonyl)amino)-4-chloronicotinic acid (10) 41.41 mL of 1 M sodium hydroxide (118.53 mmol) was added over 15 minutes at 0°C to a mixture of 9 (10.00 g, 26.34 mmol) ethanol:THF (3:1, 76 mL). After monitoring by TLC to ensure complete consumption of the starting material, the reaction mixture was concentrated. The resulting crude product was transferred to a separatory funnel containing water (200 mL), extracted with EA (1 × 100 mL), and the organic layer was discarded. 10% HCl (200 mL) and EA (100 mL) were carefully added to the aqueous layer (pH approximately 3-5). The compound was extracted into the EA layer, dried over anhydrous sodium sulfate, and concentrated to obtain pure 5-bromo-2-((tert-butoxycarbonyl)amino)-4-chloronicotinic acid (10) as a pale yellow solid. Yield: 6.02 g, 65%. 1 H NMR (500.13 MHz, DMSO-d6): δ 1.43 (s, 9H), 8.69 (s, 1H), 9.88 (s, 1H), 13 C NMR(125.77 MHz,DMSO-d6):δ 27.9(3C),80.1,116.3,124.5,141.2,148.8,150.8,152.6,164.1,HRMS:[C 11 H 12 The calculated m / z value for [BrClN2O4+H] is 350.9747, and the measured value is 350.9712.
[0150] [ka]
[0151] (11) Purification of this compound by column chromatography was difficult. Therefore, we proceeded to the next step without purification or analysis. HRMS: m / z[C 17 H 23 Calculated value for [BrClN5O5+H]: 492.0649, measured value: 492.0655.
[0152] [ka]
[0153] Di-tert-butyl(8-bromo-4-oxo-3,4-dihydropyrido[4,3-d]pyrimidine-2,5-diyl)dicarbamate(12) To a 10 (5.00 g, 11.07 mmol) ice-cold solution in anhydrous dimethylformamide (DMF, 28 mL), N-methylmorpholine (NMM, 1.64 mL, 14.94 mmol) was added, followed by chloroethyl formate (1.63 mL, 13.84 mmol). After 1 hour at 0°C, Boc-guanidine (2.38 g, 14.94 mmol) was added all at once, and the ice bath was removed. The progress of the reaction was monitored by TLC, and after 24 hours at room temperature, the starting materials were completely consumed. The resulting mixture was concentrated to dryness to obtain a crude Boc-guanidine adduct, to which water was added and stirred at room temperature for 10 minutes. The resulting solid was filtered through filter paper and dried overnight under vacuum. To the resulting solid ice-cold solution in anhydrous DMF, NaH (60%, 1.77 g, 44.28 mmol) was carefully added. After 24 hours, the DMF was removed under vacuum, and water was carefully added. The resulting solid was collected by filtration and dried under vacuum. The obtained crude material was purified by silica gel column chromatography (100-200 mesh) to obtain compound 12 as a yellow solid. Yield: 3.64 g, 72% (over two steps). 1 H NMR (500.13 MHz, DMSO-d6): δ 1.48 (s, 9H), 1.50 (s, 9H), 8.54 (s, 1H), 10.85 (s, 1H), 11.58 (s, 1H), 11.78 (s, 1H), 13 C NMR(125.77 MHz,DMSO-d6):δ 27.7(3C),27.8(3C),80.3,83.2,103.2,109.2,149.0,150.4,152.0,152.6,153.7,154.2,161.7,HRMS:[C 17 H 22 The calculated m / z value for [BrN5O5+H] is 456.0883, and the measured value is 456.0890.
[0154] [ka]
[0155] (8-bromo-4-(tert-butoxy)pyrido[4,3-d]pyrimidine-2,5-diyl)bis((tert-butoxycarbonyl)-carbamate)di-tert-butyl(13) To a solution of 12 (3.10 g, 6.79 mmol) in anhydrous THF, DIEA (3.50 mL, 20.38 mmol) and DMAP (0.17 g, 1.36 mmol) were added, followed by the addition of Boc2O (5.93 g, 27.18 mmol) at room temperature. After 4 hours at room temperature, Boc2O (2.97 g, 13.59 mmol) was added, and the reaction mixture was heated to 50°C. The consumption of the starting material was monitored by TLC, and after 24 hours, the reaction mixture was diluted with a saturated solution of NaHCO3 and EA. The organic layer was separated from the aqueous layer and dried over sodium sulfate. The solvent was removed under vacuum, and the resulting crude material was purified by silica gel column chromatography to obtain pure 13 as a white solid. Yield: 3.15 g, 65%. 1 H NMR (500.13 MHz, DMSO-d6): δ 1.27 (s, 18H), 1.45 (s, 18H), 1.65 (s, 9H), 9.06 (s, 1H), 13 C NMR(125.77 MHz,DMSO-d6):δ 27.2(6C),27.4(6C),27.5(3C),82.7(2C),83.8(2C),87.6,109.7,118.2,148.7,148.8(2C),149.6(2C),150.8,155.0,156.4,166.5, HRMS:[C 31 H 46 The calculated m / z value for [BrN5O9+H] is 712.2557, and the measured value is 712.2540.
[0156] [ka]
[0157] (8-allyl-4-(tert-butoxy)pyridol-4,3-dipyrimidine-2,5-diyl)bis((tert-butoxycarbonyl)carbamate)di-tert-butyl(14) To a stirred and degassed solution of 13 (3.00 g, 4.21 mmol) in anhydrous toluene (21 mL), allyl tributyltin (2.79 g, 8.42 mmol), followed by tetrakistriphenylphosphine palladium (0) (0.49 g, 0.42 mmol), was added at room temperature. The resulting solution was flushed with argon to create an inert atmosphere, and the solution was perfused at 125°C for 15 hours. The resulting mixture was cooled to room temperature, silica gel was added, the solvent was removed under vacuum, and the slurry was loaded into a silica gel column chromatograph to obtain 14 as a pale yellow solid. Yield: 1.70 g, 60%. 1 H NMR(500.13 MHz,CDCl3):δ 1.32(s,18H),1.47(s,18H),1.71(s,9H),3.78(dq,J=6.6,1.3 Hz,2H),4.94-5.15(m,2H),6.09(ddt,J=16.7,10.1,6.4 Hz,1H),8.49(s,1H), 13 C NMR(125.77 MHz, CDCl3):δ 27.8(6C),27.9(6C),28.2(3C),32.4,82.4(2C),83.3(2C),86.3,108.8,116.5, 131.4,136.0,148.1(2C),149.8(2C),150.3(2C),155.6,156.8,167.1,HRMS:[C 34 H 51 The calculated value of the m / z for [N5O9+H] is 674.3765, and the measured value is 674.3746.
[0158] [ka]
[0159] 2-(2,5-bis(bis(tert-butoxycarbonyl)amino)-4-(tert-butoxy)pyrido[4,3-d]pyrimidine-8-yl)acetic acid (15,E) To a chilled, stirred solution of 14 (1.00 g, 1.48 mmol) in CH3CN:CCl4:H2O (23:23:34 mL), NaLO4 (2.54 g, 11.87 mmol) was added, followed by ruthenium(III) chloride hydrate (0.07 g, 0.30 mmol). After 2 hours at 0°C, the reaction mixture was filtered through a Celite bed and concentrated to 3 / 4 of its volume under vacuum at 35°C. Water (20 mL) and EA (20 mL) were added to this mixture at 25°C. The organic layer was separated from the aqueous layer and dried over sodium sulfate. The resulting solvent was removed under vacuum, and the crude material was purified by silica gel column chromatography to obtain 15(E) as a light brown solid. Yield: 0.51 g 50%. 1 H NMR (300.13 MHz, CDCl3): δ 1.34 (s, 18H), 1.56 (s, 18H), 1.71 (s, 9H), 4.03 (s, 2H), 8.62 (s, 1H), 13 C NMR(125.77 MHz, CDCl3):δ 28.0(6C),28.1(6C),28.3(3C),37.1,83.2(2C),85.1(2C),87.7,109.0,124 .5,149.6,149.9(2C),150.1,155.9(2C),156.5(2C),167.4,170.4,HRMS:[C 33 H 49 N5O 11 Calculated value of m / z for [+H]: 692.3507, measured value: 692.3491.
[0160] [ka]
[0161] 4-amino-2-(methylthio)pyrimidine-5-carbonitrili(17) A mixture of 2-(ethoxymethylene)malononitrile (100.00 g, 0.82 mol), S-methylisothiourea hemisulfate (73.81 g, 0.82 mol), and triethylamine (370.92 mL, 2.66 mol) in anhydrous ethanol (819 mL) was stirred at room temperature. After two days, the solvent was removed under vacuum, and the resulting residue was washed with water to obtain the desired compound as a white solid. Yield: 122.48 g, 90%. 1 H NMR (500.13 MHz, DMSO-d6): δ 2.45 (s, 3H), 7.87 (s, 2H), 8.43 (s, 1H), 13 ¹³C NMR (125.77 MHz, DMSO-d6): δ 13.4, 85.3, 115.6, 160.5, 161.4, 174.6; ESI-MS: Calculated m / z value for [C6H6N4S+H]: 167.0, Measured value: 167.1.
[0162] [ka]
[0163] 4-amino-N'-hydroxy-2-(methylthio)pyrimidine-5-carboxyimidoamide (18) A mixture of 17 (50.00 g, 0.30 mol), hydroxylamine hydrochloride (24.04 g, 0.35 mol), and triethylamine (83.92 mL, 0.60 mol) was refluxed in methanol (602 mL) at 75°C. After two days, the reaction mixture was cooled to room temperature, the precipitated solid was collected by filtration, and washed with methanol and hexane to obtain 18 as a pale yellow solid. Yield: 46.15 g, 77%. 1 H NMR (500.13 MHz, DMSO-d6): δ 2.43 (s, 3H), 5.96 (s, 2H), 7.59 (s, 1H), 8.10 (s, 1H), 8.35 (s, 1H), 9.87 (s, 1H), 13 ¹³C NMR (125.77 MHz, DMSO-d6): δ 13.2, 103.6, 149.9, 152.8, 159.6, 169.3; ESI-MS: Calculated m / z value for [C6H9N5OS+H] 200.1, measured value: 200.0.
[0164] [ka]
[0165] (E)-N'-acetoxy-4-amino-2-(methylthio)pyrimidine-5-carboxyimidoamide(19) To a solution of 18 (31.00 g, 0.16 mol) in glacial acetic acid, acetic anhydride (51.48 mL, 0.55 mol) was added at room temperature. TLC analysis of the resulting solution revealed that 18 was completely consumed after 15 hours. The acetic acid was removed under vacuum, and saturated sodium bicarbonate solution and EA (200 mL each) were added to the crude substance (pH -8 to -9). The aqueous layer was extracted twice with EA, and the organic layer was dried over sodium sulfate. EA was removed under vacuum to obtain pure 19 as a white solid. Yield: 28.02 g, 90%. 1 H NMR (500.13 MHz, DMSO-d6): δ 2.15 (s, 3H), 2.45 (s, 3H), 6.89 (s, 2H), 7.78 (s, 1H), 7.96 (s, 1H), 8.40 (s, 1H), 13 C NMR(125.77 MHz,DMSO-d6):δ 13.3,19.5,102.4,154.5,154.6,159.8,167.9,171.3, HRMS:[C8H 11 The calculated value of the m / z for [N5O2S+H] is 242.0712, and the measured value is 242.0718.
[0166] [ka]
[0167] 4-amino-2-(methylthio)pyrimidine-5-carboxyimidoamide (20) A mixture of 19 (25.00 g, 0.10 mol), formic acid (97.73 mL, 2.59 mol), and palladium on charcoal (10%, 5.51 g, 5.18 mmol) in methanol (130 mL) was refluxed at 80°C. After 15 hours, the reaction mixture was cooled to room temperature, and the solvent was removed under vacuum to obtain the desired product as formate. Yield of formate of product: 29.96 g, 80%. (For NMR analysis: To the resulting crude substance (1 g), methanol (10 mL) and TEA (5 mL, pH -9~10) were added and stirred for 15 minutes. The resulting solid was collected by filtration and washed with hexane.) 1 H NMR (500.13 MHz, DMSO-d6): δ 2.43 (s, 3H), 5.96 (s, 2H), 7.59 (s, 1H), 8.11 (s, 1H), 8.36 (s, 1H), 9.87 (s, 1H), 13 ¹³C NMR (125.77 MHz, DMSO-d6): δ 13.3, 103.6, 149.9, 152.9, 159.6, 169.4; HRMS: Calculated m / z value for [C6H9N5S+H] 184.0657, measured value 184.0664.
[0168] [ka]
[0169] ((4-amino-2-(methylthio)pyrimidine-5-yl)(imino)methyl)carbamate tert-butyl(21) Sodium bicarbonate was added in several portions (34.51 mL, 0.41 mol) to a chilled, stirred solution of 20 (22.00 g, 68.47 mmol) of formate in THF:water (100:125 mL). When the pH of the resulting solution reached 8-9, 17.19 g, 78.74 mmol of Boc-anhydrous (25 mL, 78.74 mmol) was slowly added in 25 mL of THF. TLC analysis of the resulting solution revealed that 20 was completely consumed after 24 hours. The reaction mixture was transferred to a separatory funnel containing 100 mL of water and 200 mL of EA. The aqueous layer was extracted using EA (2×), and the combined EA layers were dried over sodium sulfate. The EA was removed under vacuum, and the mixture was purified by silica gel column chromatography to obtain pure 21 as a white solid. Yield: 15.71 g, 81%. 1 H NMR (300.13 MHz, CDCl3): δ 1.51 (s, 9H), 2.50 (s, 3H), 7.61 (s, 2H), 8.42 (s, 1H), 13 C NMR(125.77 MHz,CDCl3):δ 13.9,28.1(3C),79.6,104.6,155.2(2C),161.1(2C),173.9,HRMS:[C 11 H 17 The calculated value of the m / z for [N5O2S+H] is 284.1181, and the measured value is 284.1189.
[0170] [ka]
[0171] (Z)-((4-(bis(tert-butoxycarbonyl)amino)-2-(methylthio)pyrimidine-5-yl)((tert-butoxycarbonyl)imino)methyl)(tert-butoxycarbonyl)carbamate tert-butyl(22) To a frozen, stirred solution of 21 (15.00 g, 52.94 mmol) in 106 mL of THF, DIPEA was added in one batch (46.11 mL, 264.69 mmol), followed by DMAP (1.29 g, 10.59 mmol), and then Boc-anhydrous (57.77 g, 264.69 mmol) in 25 mL of THF (added slowly). The reaction was monitored by TLC and it was found that 21 was completely consumed after 24 hours. The reaction mixture was transferred to a separatory funnel containing 200 mL of water and 100 mL of EA. The mixture was extracted with EA (2×), and the combined EA layer was dried over sodium sulfate. EA was removed under vacuum, and the crude mixture was purified by silica gel column chromatography to obtain pure 22 as a white solid. Yield: 28.96 g, 80%. 1 H NMR (500.13 MHz, CDCl3): δ 1.44(s,l8H),1.45(s,18H),1.51(s,9H),2.61(s,3H),8.69(s,1H), 13 C NMR(125.77 MHz, CDCl3):δ 14.6,28.0(6C),28.1(6C),28.1(3C),83.2,83.7(2C),84.9(2C),119.2,148.3,150.1,157.3(2C),158.8,159.5(2C),160.4,176.1,HRMS:[C 31 H50N5O 10 The calculated value for [S+H] in m / z is 684.3278, and the measured value is 684.3275.
[0172] [ka]
[0173] Compound (23) mCPBA was added in several portions (10.60 g, 61.42 mmol) to an ice-cold, stirred solution of 22 (15.00 g, 21.94 mmol) in 55 mL of DCM. The reaction was analyzed by TLC and it was found that 22 was completely consumed after 3 hours. The reaction mixture was transferred to a separatory funnel containing 200 mL of saturated sodium bicarbonate solution and 100 mL of DCM. The mixture was extracted with DCM (2×) and the combined DCM layer was dried over sodium sulfate. The DCM was removed under vacuum and the crude mixture was used in the next step without further purification.
[0174] [ka]
[0175] (Z)-((4-(bis(tert-butoxycarbonyl)amino)-2-oxo-1,2-dihydropyrimidine-5-yl)((tert-butoxycarbonyl)imino)methyl)(tert-butoxycarbonyl)carbamate tert-butyl(24) To a chilled, stirred solution of crude 23 (15.00 g) in 55 mL of THF, 2N NaOH (54.84 mL, 109.68 mmol) was slowly added. The consumption of the starting material was monitored by TLC. After 30 minutes, the reaction mixture was transferred to a separatory funnel containing 200 mL of 1% HCl and 100 mL of EA. The aqueous layer was extracted with EA (2×), and the combined EA layer was dried over sodium sulfate. The EA was removed under vacuum, and the crude material was purified by silica gel column chromatography. Yield: 10.75 g, 75% over two steps. 1 H NMR (300.13 MHz, DMSO-d6): δ 1.37 (s, 18H), 1.39 (s, 18H), 1.41 (s, 9H), 8.65 (s, 1H), 12.80 (s, 1H), 13 C NMR(125.77 MHz,DMSO-d6):δ 27.4(15C),82.0,82.9(2C),83.5(2C),108.0,145.2,147.5(2C),149.1(2C),152.8,155.0,157.0,162.9,HRMS:[C 30 H47 N5O 11 The calculated value for m / z at [+H] is 654.3350, and the measured value is 654.3359.
[0176] [ka]
[0177] (Z)-2-(4-(bis(tert-butoxycarbonyl)amino)-2-oxo-5-(N,N,N'-tris(tert-butoxycarbonyl)carbamimidoyl)pyrimidine-1(2H)-yl)benzyl acetate(25) Compound 24 (10.00 g, 15.30 mmol) and K2CO3 (4.23 g, 30.59 mmol) were mixed in anhydrous DMF (75 mL) to which 2-benzyl bromoacetate (2.42 mL, 15.30 mmol) was added dropwise over 10 minutes. The reaction was monitored by TLC and found to be complete after 4 hours. The reaction mixture was filtered through a sintering funnel, and the DMF was evaporated under vacuum. Water was added to the resulting crude substance and stirred for 5 minutes. The resulting precipitate was collected by filtration and washed with water and hexane to obtain pure compound 25 as a white solid. (Note: If the compound does not precipitate sufficiently, extraction is an alternative method to obtain the compound, and if necessary, short-bed column filtration can be performed to obtain the pure compound). Yield: 10.92 g, 89%. 1 H NMR (500.13 MHz, CDCl3): δ 1.46(s,18H),1.48(s,18H),1.50(s,9H),4.75(s,2H),5.22(s,2H),7.27-7.47(m,5H),7.95(s,1H), 13 C NMR(125.77 MHz, CDCl3):δ 27.8(6C),27.9(6C),28.0(3C),51.0,68.2,83.0,83.7(2C),85.2(2C),110.6,128.6(3C),1 28.7(2C),128.8,134.5,148.2,149.5(2C),152.0(2C),154.4,157.1,163.9,166.1,HRMS:[C39 H 55 N5O 13 The calculated value for m / z at +H is 802.3875, and the measured value is 802.3882.
[0178] [ka]
[0179] (Z)-2-(4-(bis(tert-butoxycarbonyl)amino)-2-oxo-5-(N,N,N'-tris(tert-butoxycarbonyl)carbamimidoyl)pyrimidioyl-1(2H)-yl)acetic acid (26,F) Palladium (10% loading, 0.66 g, 0.62 mmol) on charcoal was carefully added under argon balloon pressure to a solution of 25 (5.00 g, 6.24 mmol) in EA (50 mL) under argon balloon pressure. After 5 minutes, the argon balloon was replaced with a hydrogen balloon, and the reaction mixture was stirred for a further 1 hour under the same atmosphere. Complete depletion of the starting materials was determined by TLC. The resulting reaction mixture was carefully filtered through a Celite bed, and the filtrate was evaporated in a rotary evaporator to obtain pure compound 26 as a white solid. Yield: 3.99 g, 90%. 1 H NMR (500.13MHz, DMSO-d6): δ1.36(s,18H), 1.41(s,18H), 1.42(s,9H), 4.85(s,2H), 9.10(s,1H), 13.45(s,1H), 13 C NMR(125.77MHz,DMSO-d6):0 27.4(6C), 27.4(6C), 27.5(3C), 50.9, 82.2, 83.0(2C), 83.3(2C), 108.3, 1 45.1, 146.9(2C), 149.4(2C), 154.1, 156.1, 157.0, 162.5, 168.4, HRMS:[C 32 H 49 N5O 13 Calculated m / z value for +H: 712.3405, measured value: 712.3386.
[0180] [ka]
[0181] 2-(2,4-dioxo-1,2,3,4-tetrahydropyrimidine-5-yl)acetic acid (27) This compound was purchased from a commercially available source.
[0182] [ka]
[0183] (R)-(2-((((9H-fluoren-9-yl)methoxy)carbonyl)amino)-3-(2-(2-(tert-butoxy)ethoxy)ethoxy)propyl)glycinate allyl(28) Dess Martin periodinane (9.96 g, 23.49 mmol) was added in several portions to a saturated DCM:water (28 mL) cooled solution of N-Fmoc(O-MP(t-butyl)-L-serinol (5.00 g, 10.93 mmol) in water (28 mL). The reaction mixture was stirred for a further 30 minutes at the same temperature and then at room temperature for 1 hour to completely convert the alcohol to the aldehyde. The reaction mixture was diluted with ether, a saturated solution of sodium bicarbonate, and 10% sodium thiosulfate, and stirred for 5 minutes. The entire reaction mixture was transferred to a separatory funnel, and the organic compound was extracted into ether (2x). The ether layer was dried over sodium sulfate and removed using a rotary evaporator to obtain the crude aldehyde. Further purification of the aldehyde was performed. It was used in the next step. To the aldehyde in anhydrous DCM (44 mL), allyl glycinate PTSA salt (6.28 g, 21.86 mmol) and DIEA (5.90 mL, 33.89 mmol) were added and reacted for 1 hour. NaB(OAc)3H (5.79 g, 27.33 mmol) was added to the reaction mixture all at once. After stirring overnight at room temperature, the reaction mixture was diluted with DCM (100 mL) and a saturated solution of sodium bicarbonate. The reaction mixture was transferred to a separatory funnel and the layers were separated from the aqueous layer. The combined organic layers were dried over sodium sulfate and removed using a rotary evaporator, and the crude mixture was purified by flash silica gel column chromatography. Yield: 4.85 g, 85%. 1H NMR (500.13 MHz, CDCl3): δ 1.19(s,9H),2.77-2.92(m,2H),3.49-3.69(m,13H),3.88(bs,1H),4.23(t,J=6.9 Hz,1H),4.39(d,J=7.5 Hz,2H),4.64(dt,J=5.9,1.4 Hz,2H),5.23-5.36(m,2H),5.73(d,J=8.4 Hz,1H),5.92(ddt,J=17.3,10.4,5.8 Hz,1H),7.32(td,J=7.5,1.2 Hz,2H),7.36-7.43(m,2H),7.60-7.68(m,2H),7.74-7.79(m,2H), 13 C NMR (125.77 MHz, CDCl3): δ 27.5(3C),47.3,50.3,50.5,50.6,50.7,61.2,65.5,66.7,70.5,70.8,71.2,71.4,73.1,118.7,119.9, 120.0,120.0,125.2,127.0,127.6,127.7,131.8,141.3,144.0,144.0,156.4,171.8,174.3, ESI-MS:[C 31 H 42 N2O 71 +H]のm / zのcalculated value: 555.3, measured value: 555.3.
[0184]
change
[0185] アリルモノマー(29a) To a mixture of 15(E) (0.50 g, 0.72 mmol) and DIEA (163.67 μL, 0.94 mmol) in anhydrous DMF (6 mL), HBTU (0.32 g, 0.83 mmol) was added in one step at room temperature. After 10 minutes at room temperature, skeleton 28 (0.60 g, 1.08 mmol) in DMF (4 mL) was added to the reaction mixture. The progress of the reaction was monitored by TLC (eluent:EA:Hex (50:50), skeleton Rf=0.2, product Rf=0.6, E Rf=0.1). After the starting materials were completely consumed, the DMF was removed under vacuum at 45°C. To the resulting crude product, 1% HCl (20 mL) and EA (20 mL) were added, and the organic layer (2x) was separated from the aqueous layer using a separatory funnel. The combined organic layers were dried and removed, and the resulting crude material was purified by flash silica gel column chromatography. Yield: 0.62 g, 70%. 1 H NMR (500.13 MHz, DMSO-d6, rotamer): δ 1.08 / 1.09(s,9H),1.211.33(m,18H),1.36 / 1.38(s,18H),1.41 / 1.65(s,9H),3.36-3.70(m,10H),3.82-4.27(m,8 H),4.32(bs,2H),4.41-4.71(m,2H),5.13-5.41(m,2H),5.81-6.01(m,1H),7.30-7.57(m,5H),7.71 / 7.73(d,J=9.7 Hz,2H),7.89 / 7.99(d,J=7.5 Hz, 2H), 8.56 / 8.60(s, 1H), 13 C NMR (125.77 MHz, DMSO-d6, rotamer): δ 27.2-27.6(18C),46.7,59.7,60.6 / 60.6,64.7,65.3,65.5,69.6,69.6,69.8,70. 0,70.4,70.4,72.2,81.8-86.3(6C),108.1,117.7,118.2,120.l(2C),125.1,127. 0(2C),127.3,127.6(2C),127.8 / 127.9,132.1,132.3,140.7,143.8,147.9,149. 3,149.5,149.6,155.0 / 155.1,155.9,156.5,166.6,168.8,169.2,170.2,HRMS:[C64 H 89 N7O 17 Calculated value of m / z for [+H]: 1228.6393, measured value: 1228.6398.
[0186] [ka]
[0187] Allyl monomer (29b) This compound was prepared using compound 27 (1.00 g, 5.88 mmol) and skeleton 28 (4.08 g, 7.35 mmol) as starting materials, following the procedure described above for compound 29a. TLC (eluent: 5% MeOH:DCM (5:95), Rf=0.6 for skeleton, Rf=0.2 for product, Rf=0.0 for 27). Yield: 3.5 g, 84%. 1 H NMR (500.13 MHz, DMSO-d6, rotamer): δ 1.10 / 1.10(s,9H),3.30-3.59(m,14H),3.744.13(m,2H),4.174.42(m,4H),4.60(dd,J=33.4,4.8 Hz,2H),5.00-5.44(m,2H),5.82-6.01(m,1H),7.10-7.55(m,6H),7.69(d,J=7.4 Hz,2H),7.89(d,J=7.5 Hz,2H),10.75(bs,1H),11.08(s,1H), 13 C NMR (125.77 MHz, DMSO-d6, rotamer) δ 27.2(3C),46.7,60.6(2C),64.7,65.2,65.5,69.6,69.8,69.9,70.4(2C),72.2,107.2,117.7,118.1,120.1(2C),125.2,127.0(2C),1 27.3,127.6(2C),132.2 / 132.3,139.2,139.7,140.7,143.8 / 143.8,151.2 / 151.3,155.8,164.1 / 164.1,168.9,169.4,170.6,HRMS:[C 37 H 46 N4O 10 Calculated value of m / z for [+H]: 707.3292, measured value: 707.3290.
[0188] [ka]
[0189] Allyl monomer (29c) This compound was prepared using compound 26 (3.50 g, 4.92 mmol) and skeleton 28 (3.41 g, 6.15 mmol) as starting materials, following the procedure described above for compound 29a. TLC (eluent: EA:Hex(40:60), Rf=0.1 for skeleton, Rf=0.7 for product, Rf=0.1 for 26). Yield: 4.80 g, 78%. 1 H NMR(300 MHz,DMSO-d6):δ 1.09(bs,9H),1.34 / 13.6(s,18H),1.41 / 1.42(s,27H),3.35-3.59(m,14H),4.07-4.49(m,4H),4.61(dd,J=37.6,5.4 Hz,3H),4.97-5.42(m,3H),5.77-6.04(m,1H),7.32(t,J=7.5 Hz,2H),7.50-7.57(m,3H),7.89(d,J=7.5 Hz,2H),7.99(d,J=8.4 Hz, 2H), 8.94 / 8.96(s, 1H), 13 ¹³C NMR (125.77 MHz, DMSO-d6, rotational isomer): δ 27.2-27.5(18C),46.7,60.2 / 60.7(2C),64.7,65.5,69.5,69.6,69.8,69. 9,70.4 / 70.4(2C),72.2,82.2-83.4(6C),108.1,117.7,120.1(2C),124.4, 125.1,127.0(2C),127.4,127.6(2C),127.8,132.2(2C),140.7(2C),143. 8(2C),145.1,146.9(2C),149.5(2C),157.0,162.5,166.5,168.4,HRMS:[C 63 H 89 N7O 19 The calculated value for +H is dm / z: 1248.6291, and the measured value is 1248.6256.
[0190] Non-limiting aspects of the present invention are described in the following numbered clauses.
[0191] Clause 1: A gene recognition reagent comprising a nucleic acid or nucleic acid analog skeleton, and comprising three or more ribose, deoxyribose, or nucleic acid analog skeleton residues, and a sequence of divalent nucleic acid bases bound to the skeleton residues, wherein the sequence of divalent nucleic acid bases binds to a unit target sequence or one or more sequential repeats of a unit target sequence of an extended repeat of a repeat extension disease on two nucleic acid strands.
[0192] Clause 2: The gene recognition reagent according to Clause 1, wherein the nucleic acid or nucleic acid analog skeleton comprises a first end and a second end, further comprising a first linking group attached to the first end of the skeleton, and a second linking group attached to the second end, which is non-covalently bonded to or self-ligated with the first linking group.
[0193] Clause 3: The gene recognition reagent according to Clause 1 or 2, wherein the sequence of the divalent nucleic acid bases binds to a unit target sequence or two or more sequential repeats of the unit target sequence of the extended repeats of the repeat extension disease on each of the two strands when the two strands are aligned in antiparallel orientation.
[0194] Clause 4: The unit target sequence is on both strands of the (GAA) n (CGG) n (CCG) n (CAG) n (CTG) n (CCTG) n (ATTCT) n , or (GGGGCC) n A gene recognition reagent according to any one of the clauses 1 to 3, formed by sequential repeats of or sequences complementary to either of the above.
[0195] Clause 5: A gene recognition reagent according to any one of Clauses 1 to 4, wherein the unit target sequence is C / GA / AG / C.
[0196] Clause 6: A gene recognition reagent according to any one of Clauses 1 to 5, wherein the nucleic acid base sequence is in the order of JB3 nucleic acid base, JB6 nucleic acid base, and JB4 nucleic acid base, JB6 nucleic acid base, JB4 nucleic acid base, and JB3 nucleic acid base, or JB4 nucleic acid base, JB3 nucleic acid base, and JB6 nucleic acid base, or two or more sequential repeats of the above.
[0197] Clause 7: The gene recognition reagent according to Clause 6, having the nucleic acid base sequence in the order of JB3 or JB3b, JB6 or JB6b, and JB4, JB4b, JB4c, JB4d, or JB4e, JB6 or JB6b, JB4, JB4b, JB4c, JB4d, or JB4e, and JB3 or JB3b, or JB4, JB4b, JB4c, JB4d, or JB4e, JB3 or JB3b, and JB6 or JB6b, or two or more sequential repeats of the above.
[0198] Clause 8: A gene recognition reagent according to any one of Clauses 1 to 5, wherein the nucleic acid base sequence is in the order of JB3, JB6, and JB4, JB6, JB4, and JB3, or JB4, JB3, and JB6, or two or more sequential repeats thereof.
[0199] Clause 9: A gene recognition reagent according to any one of Clauses 1 to 3, having the nucleic acid base sequence of the unit recognition reagent sequence listed in Table C.
[0200] Clause 10: A gene recognition reagent according to any one of Clauses 1 to 9, wherein both strands are of the same nucleic acid molecule.
[0201] Clause 11: A gene recognition reagent according to any one of Clauses 1 to 10, wherein both strands are RNA.
[0202] Clause 12: A gene recognition reagent as described in any one of Clauses 1 to 11, consisting of a sequence of 3 to 8 nucleic acid bases, such as 3, 4, 5, 6, 7, or 8 nucleic acid bases.
[0203] Clause 13: A gene recognition reagent according to any one of Clauses 1 to 12, wherein the skeletal residue is a nucleic acid analog skeletal residue.
[0204] Clause 14: The gene recognition reagent according to Clause 13, wherein the skeletal residues are pre-organized in three-dimensional structure.
[0205] Clause 15: The gene recognition reagent according to Clause 14, wherein the skeletal residue is a locked nucleic acid residue.
[0206] Clause 16: The gene recognition reagent according to Clause 14, wherein the skeletal residue is a peptide nucleic acid (PNA) residue.
[0207] Clause 17: The gene recognition reagent according to Clause 14, wherein the skeletal residue is a γPNA residue.
[0208] Clause 18: Structure: [ka]
[0209] Here, X is S or O, n is an integer from 1 to 6, m is an integer from 0 to 4, and R1 and R2 are each independently H, a guanidine-containing group, for example, [ka] Here, n = 1, 2, 3, 4, or 5, amino acid side chains, for example, [ka] Linear or branched (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C1-C8) hydroxyalkyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, (C3-C8) cycloalkyl(C1-C6) alkylene, may be optionally substituted with ethylene glycol units containing 1 to 50 ethylene glycol moieties, -CH2-(OCH2-CH2) qOP1, -CH2-(OCH2-CH2) q -NHP1, -CH2-(SCH2-CH2) q -SP1, -CH2-(OCH2-CH2) r -OH, -CH2-(OCH2-CH2) r -NH2, -CH2-(OCH2-CH2) r -NHC(NH)NH2, or -CH2-(OCH2-CH2) r -SS[CH2CH2] s NHC(NH)NH2, where P1 is H, (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, or (C3-C8) cycloalkyl(C1-C6) alkylene, q is an integer from 0 to 50, and r and s are each an integer from 1 to 50. R3 may be optionally substituted with H, or a leaving group, such as a linear or branched (C1-C8) alkyl, a substituted or unsubstituted (C3-C8) aryl, a (C3-C8) aryl(C1-C6) alkylene, a (C1-C8) carboxy, or an amino acid side chain, or a guanidine-containing group, or [ka] Here, o is 1 to 20, R6 is each an amino acid side chain, and R7 is either -OH or -NH2. R4 is (C1~C 10 ) Divalent hydrocarbons or those substituted with one or more N or O moieties, such as -O-, -OH, -C(O)-, -NH-, -NH2, -C(O)NH- (C1~C 10 ) It is a divalent hydrocarbon, R5 is an -OH, -SH, or disulfide protecting group, and A gene recognition reagent according to any one of Clauses 1 to 14, wherein each of R independently generates a sequence of nucleic acid bases complementary to a unit target sequence of nucleic acid bases in one or more target nucleic acids, and therefore each of the multiple recognition modules is a nucleic acid base that binds to and ligates with the target sequence of nucleic acid bases on a template nucleic acid, or a pharmaceutically acceptable salt thereof.
[0210] Clause 19: R1 or R2 is -(OCH2-CH2) q OP1, -(OCH2-CH2) q -NHP1, -(SCH2-CH2) q -SP1, -(OCH2-CH2) r -OH, -(OCH2-CH2) r -NH2, -(OCH2-CH2) r -NHC(NH)NH2, or -(OCH2-CH2) r -SS[CH2CH2] s A gene recognition reagent as described in Clause 18, wherein P1 is an NHC(NH)NH2-substituted (C1-C6) alkyl, where P1 is H, (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, or (C3-C8) cycloalkyl(C1-C6) alkylene, where q is an integer from 0 to 50, r is an integer from 1 to 50, and s is an integer from 1 to 50.
[0211] The gene recognition reagent according to Clause 20:R5 is -SH, OH, or SS-R8, where R8 is one or more amino acid residues, an amino acid side chain, a guanidine-containing group, an unsubstituted or substituted (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C1-C8) hydroxyalkyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, or (C3-C8) cycloalkyl(C1-C6) alkylene, as described in Clause 18.
[0212] Clause 21: R1 is H and R2 is not H, or R2 is H and R1 is not H, the gene recognition reagent as described in Clause 18.
[0213] Clause 22: The gene recognition reagent according to Clause 18, wherein R1 is H and R2 comprises a polyoxyethylene moiety.
[0214] Clause 23: The gene recognition reagent described in Clause 18, wherein R5 is -SH and X is S.
[0215] Clause 24: The gene recognition reagent according to Clause 18, wherein R5 comprises a lysine residue, an arginine residue, or a guanidine-containing group.
[0216] Clause 25: The gene recognition reagent according to Clause 18, wherein R4 is methylene, R5 is -SH, X is S, and R5 is a hydroxyl-substituted (C1-C8) alkyl, e.g., -CH2-CH2-OH.
[0217] Clause 26: The gene recognition reagent according to any one of Clauses 2 to 14, wherein, when the recognition reagent hybridizes with an adjacent sequence of the target nucleic acid, the first linking group and the second linking group are independently 2-5 ring fused polycyclic aromatic moieties that stack with the aryl moiety of the adjacent recognition reagent.
[0218] Clause 27: The gene recognition reagent according to Clause 26, wherein the 2-5 ring condensed polycyclic aromatic moiety is unsubstituted or substituted pentalene, indene, naphthalene, azulene, heptalene, biphenylene, as-indacene, s-indacene, acenaphthylene, fluorene, phenalene, phenanthrene, anthracene, fluorantene, acephenanthrylene, aceanthrylene, triphenylene, pyrene, chrysene, naphthacene / tetracene, pleiadene, picene, or perylene, which may be optionally substituted with one or more heteroatoms such as O, N, and / or S, for example, xanthene, riboflavin (vitamin B2), mangosteen, or mangiferin.
[0219] Clause 28: The gene recognition reagent according to Clause 26 or 27, wherein the 2-5 ring condensed polycyclic aromatic moiety is identical.
[0220] Clause 29: The recognition reagent has the structure: [ka] It has, Here, each of R is independently a nucleic acid base, and each of R can be the same or different nucleic acid base. n is an integer in the range of 1 to 6, for example, 1, 2, 3, 4, 5, or 6. R1 and R2 are each independently H, a guanidine-containing group, for example, [ka] Here, n = 1, 2, 3, 4, or 5, amino acid side chains, for example, [ka] , unsubstituted or substituted (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C1-C8) hydroxyalkyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, or (C3-C8) cycloalkyl(C1-C6) alkylene, -CH2-(OCH2-CH2) q OP1, -CH2-(OCH2-CH2) q -NHP1, -CH2-(OCH2-CH2-O) q -SP1, -CH2-(SCH2-CH2) q -SP1, -CH2-(OCH2-CH2) r -OH, -CH2-(OCH2-CH2) r -NH2, -CH2-(OCH2-CH2) r -NHC(NH)NH2, or -CH2-(OCH2-CH2) r -SS[CH2CH2] sNHC(NH)NH2, where P1 is H, (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, or (C3-C8) cycloalkyl(C1-C6) alkylene, q is an integer from 0 to 50, and r and s are each an integer from 1 to 50. R 10 or R 11 On the other hand, and R 12 , R 13 , or R 14 One of these is -L-R3, where each of R3 is independently a 2-5 ring condensed polycyclic aromatic moiety, for example, unsubstituted or substituted pentalene, indene, naphthalene, azulene, heptalene, biphenylene, as-indacene, s-indacene, acenaphthylene, fluorene, phenalene, phenanthrene, anthracene, fluorantene, acephenanthrylene, aceanthrylene, triphenylene, pyrene, chrysene, naphthalene / tetracene, pleiadene, picene, or perylene, which may be optionally substituted with one or more heteroatoms such as O, N, and / or S, for example, xanthene, riboflavin (vitamin B2), mangosteen, or mangiferin, each of which can stack with the R3 group of the adjacent recognition reagent if the recognition reagent hybridizes with the adjacent sequence of the target nucleic acid. L is a linker, and R 10 , R 11 , R 12 , R 13 , and R 14 The remaining parts are, each independently, H, one or more consecutive amino acid residues, for example, one or more Arg residues, a guanidine-containing group, for example, [ka] Here, n = 1, 2, 3, 4, or 5, amino acid side chains, for example, [ka] , unsubstituted or substituted (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C1-C8) hydroxyalkyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, or (C3-C8) cycloalkyl(C1-C6) alkylene, -CH2-(OCH2-CH2) q OP1, -CH2-(OCH2-CH2) q -NHP1, -CH2-(OCH2-CH2-O) q -SP1, -CH2-(SCH2-CH2) q -SP1, -CH2-(OCH2-CH2) r -OH, -CH2-(OCH2-CH2) r -NH2, -CH2-(OCH2-CH2) r -NHC(NH)NH2, or -CH2-(OCH2-CH2) r -SS[CH2CH2] s NHC(NH)NH2, where P1 is H, (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, or (C3-C8) cycloalkyl(C1-C6) alkylene, q is an integer from 0 to 50, and r and s are each independently an integer from 1 to 50, or The pharmaceutically acceptable salt of the gene recognition reagent described in Clause 2,
[0221] Clause 30:1 or more R1, R2, R 10 , R 11 , R 12 , R 13 , or R 14 However, -(OCH2-CH2) q OP1, -(OCH2-CH2) q -NHP1, -(SCH2-CH2) q -SP1, -(OCH2-CH2) r -OH, -(OCH2-CH2) r -NH2, -(OCH2-CH2) r -NHC(NH)NH2, or -(OCH2-CH2)r -SS[CH2CH2] s A gene recognition reagent as described in Clause 29, wherein P1 is an NHC(NH)NH2-substituted (C1-C6) alkyl, where P1 is H, (C1-C8) alkyl, (C2-C8) alkenyl, (C2-C8) alkynyl, (C3-C8) aryl, (C3-C8) cycloalkyl, (C3-C8) aryl(C1-C6) alkylene, or (C3-C8) cycloalkyl(C1-C6) alkylene, where q is an integer from 0 to 50, r is an integer from 1 to 50, and s is an integer from 1 to 50.
[0222] Clause 31: The gene recognition reagent described in Clause 29, wherein R4 and R7 are -L-R3.
[0223] Clause 32: The gene recognition reagent according to Clause 29, wherein the linker comprises approximately 5 to 25 atoms, for example 5 to 20, 5 to 10, for example 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 atoms, or a total of 1 to 10, for example 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 C and heteroatoms, for example O, P, N, or S atoms.
[0224] In all cases of Clause 33:R3, the gene recognition reagent described in Clause 29, comprising the same 2-5 ring fused polycyclic aromatic moiety.
[0225] Article 34: Guanidine-containing groups, for example, [ka] A gene recognition reagent as described in any one of Clauses 1 to 33, wherein n = 1, 2, 3, 4, or 5.
[0226] Clause 35: A method for conjugating nucleic acids containing extended repeats associated with repeat extension disorders, comprising contacting the nucleic acid containing the extended repeats with a gene recognition reagent described in any one of Clauses 1 to 32.
[0227] Clause 36: The nucleic acid is on both of the two strands (GAA) n (CGG) n (CCG) n (CAG) n (CTG) n (CCTG) n (ATTCT) n , or (GGGGCC) n The method according to clause 35, comprising a unit target sequence formed by sequential repetition of, or a sequence complementary to either of the above.
[0228] Clause 37: The nucleic acid is on both of the two strands (CAG) n The method according to clause 35, comprising a unit target sequence formed by sequential repetition of or a sequence complementary thereto.
[0229] Clause 38: The method according to any one of Clauses 35 to 37, wherein the nucleic acid is obtained from a patient having a repeat extension disorder.
[0230] Clause 39: A method for identifying the presence of nucleic acids containing extended repeats associated with repeat extension disease in a sample obtained from a patient, comprising contacting the nucleic acid sample obtained from the patient with a gene recognition reagent described in any one of Clauses 2 to 34, determining whether binding and ligation of the gene recognition reagent occurs to indicate that the sample contains nucleic acids containing extended repeats associated with repeat extension disease, and optionally treating the patient for the repeat extension disease.
[0231] Clause 40: A composition comprising a gene recognition reagent described in any one of Clauses 1 to 34 and a pharmaceutically acceptable carrier.
[0232] Clause 41: A method for knocking down the expression of a gene containing an elongated repeat associated with a repeat elongation disease in cells, comprising contacting a nucleic acid such as RNA containing the elongated repeat with a gene recognition reagent described in any one of Clauses 1 to 34.
[0233] The present invention has been described with reference to certain exemplary embodiments, dispersible compositions, and their uses. However, it will be apparent to those skilled in the art that various substitutions, modifications, or combinations of any of the exemplary embodiments can be made without departing from the spirit and scope of the invention. Accordingly, the invention is not limited by the description of the exemplary embodiments, but rather by the claims submitted initially.
Claims
1. An in vitro method for conjugating nucleic acids containing extended repeats associated with a repeat extension disorder, wherein the repeat extension disorder involves (CAG)n repeat extension. The process involves contacting the nucleic acid containing the extended repeat with a gene recognition reagent. The gene recognition reagent is (a) a nucleic acid analog skeleton, and the nucleic acid analog skeleton is (i) Three or more nucleic acid analog backbone residues, (ii) The first end and the second end, (iii) comprising a first linking group bonded to the first end and a second linking group bonded to the second end, (b) a nucleic acid analog skeleton comprising: (a) a sequence of divalent nucleic acid bases bonded to the three or more residues of the nucleic acid analog skeleton, wherein the first linking group comprises an -SH, -OH, or -S-S-Lg moiety, and the second linking group comprises an -C(O)-S-Lg, or -C(O)-O-Lg moiety, where Lg is a leaving group linked to the skeleton by a disulfide (S-S), ester, or thioester bond; and (b) a sequence of divalent nucleic acid bases bonded to the three or more residues of the nucleic acid analog skeleton. Here, the sequence of the divalent nucleic acid base is bound to one or more sequential repeats of the unit target sequence on the two nucleic acid strands, where the one or more sequential repeats of the unit target sequence are extended repeats. The sequence of the aforementioned divalent nucleic acid bases is 3-6-4、 6-4-3、 4-3-6、 4-6-3、 6-3-4, or, It contains a unit recognition reagent sequence of 3-4-6, Here, each of the three independently 【Chemistry 1】 And, Each of the four operates independently. 【Chemistry 2】 And, Each of the six acts independently. 【Transformation 3】 And, Here, R is a nucleic acid analog skeleton residue of the nucleic acid analog skeleton, R 1 However, it is a protecting group or H, A method wherein the nucleic acid is obtained from a patient suffering from a repeat elongation disorder.
2. The in vitro method according to claim 1, wherein the two nucleic acid strands are aligned in antiparallel orientation.
3. The in vitro method according to claim 2, wherein the unit target sequence is C / G-A / A-G / C.
4. The sequence of nucleic acid bases is JB3 nucleic acid base, JB6 nucleic acid base, and JB4 nucleic acid base JB6 nucleic acid base, JB4 nucleic acid base, and JB3 nucleic acid base, JB4 nucleic acid base, JB3 nucleic acid base, and JB6 nucleic acid base, JB3 or JB3b; JB6 or JB6b; and JB4, JB4b, JB4c, JB4d, or JB4e, JB6 or JB6b; JB4, JB4b, JB4c, JB4d, or JB4e; and JB3 or JB3b, JB4, JB4b, JB4c, JB4d, or JB4e; JB3 or JB3b; and JB6 or JB6b, JB3, JB6, and JB4, JB6, JB4, and JB3, or The in vitro method according to claim 1, comprising JB4, JB3, and JB6, or the two or more sequential iterations described above, in that order.
5. The in vitro method according to claim 1, wherein both of the nucleic acid chains are of the same nucleic acid molecule.
6. The in vitro method according to claim 1, wherein both of the nucleic acid strands are RNA.
7. The in vitro method according to claim 1, wherein the nucleic acid analog backbone residues are pre-organized in three-dimensional structure.
8. The in vitro method according to claim 1, wherein the nucleic acid analog backbone residue is a peptide nucleic acid (PNA) residue.
9. The in vitro method according to claim 1, wherein the nucleic acid analog backbone residue is a γPNA residue.
10. The gene recognition reagent has the following structure: 【Chemistry 4】 It has, Here, Each of Y is independently a nucleic acid base of the sequence of the divalent nucleic acid bases, X is either S or O, n is 1, 2, 3, 4, 5, or 6. m is 0, 1, 2, 3, or 4, Each R 1 and R 2 are each independently H, a guanidine-containing group, an amino acid side chain, unsubstituted or substituted (C 1 to C 8 ) alkyl, (C 2 to C 8 ) alkenyl, (C 2 to C 8 ) alkynyl, (C 1 to C 8 ) hydroxyalkyl, (C 3 to C 8 ) aryl, (C 3 to C 8 ) cycloalkyl, (C 3 to C 8 ) aryl (C 1 to C 6 ) alkylene, or (C 3 to C 8 ) cycloalkyl (C 1 to C 6 ) alkylene, -CH 2 -(OCH 2 -CH 2 ) q OP 1 , -CH 2 -(OCH 2 -CH 2 ) q -NHP 1 , -CH 2 -(OCH 2 -CH 2 -O) q -SP 1 , -CH 2 -(SCH 2 -CH 2 ) q -SP 1 , -CH 2 -(OCH 2 -CH 2 ) r -OH, -CH 2 -(OCH 2 -CH 2 ) r -NH 2 , -CH 2 -(OCH 2 -CH 2 ) r -NHC(NH)NH 2 , or -CH 2 - (OCH 2 -CH 2 ) r -S-S[CH 2 CH 2 ] s NHC (NH) NH 2 And here, P 1 H, (C 1 ~C 8 ) alkyl, (C 2 ~C 8 ) Alkenil, (C 2 ~C 8 ) Alkinnil, (C 3 ~C 8 ) Aryl, (C 3 ~C 8 ) Cycloalkyl, (C 3 ~C 8 ) Aryl (C 1 ~C 6 ) Alkylene or (C 3 ~C 8 ) Cycloalkyl (C 1 ~C 6 ) is an alkylene, where q is an integer from 0 to 50, and r and s are each an integer from 1 to 50, R 3 is a leaving group, wherein the leaving group is linear or branched (C 1 to C 8 )alkyl, substituted or unsubstituted (C 3 to C 8 )aryl, (C 3 to C 8 )aryl (C 1 to C 6 )alkylene, (C 1 to C 8 )carboxy (these are optionally substituted with an amino acid side chain), or a guanidine-containing group, or 【Transformation 5】 Includes, Wherein o is 1 to 20, R 6 are each independently an amino acid side chain, and R 7 is -OH or -NH 2 , R 4 However, (C 1 ~C 10 ) Divalent hydrocarbons or (C) substituted with one or more N or O moieties 1 ~C 10 ) It is a divalent hydrocarbon, R 5 However, it is -OH, -SH or -S-S-Lg, where Lg is linear or branched (C 1 ~C 8 ) alkyl, substituted or unsubstituted (C 3 ~C 8 ) Aryl, (C 3 ~C 8 ) Aryl (C 1 ~C 6 ) Alkilen, (C 1 ~C 8 ) Carboxylates (these may be optionally substituted in the amino acid side chains), or guanidine-containing groups, or 【Transformation 6】 Includes, Here, o is between 1 and 20, and R 6 Each of these is independently an amino acid side chain, R 7 is -OH or -NH 2 is, or, The in vitro method according to claim 1, wherein the pharmaceutically acceptable salt is used.
11. R is 1 or greater. 1 and R 2 However, -(OCH 2 -CH 2 ) q OP 1 ,-(OCH 2 -CH 2 ) q - NHP 1 ,-(SCH 2 -CH 2 ) q - SP 1 ,-(OCH 2 -CH 2 ) r -OH, -(OCH 2 -CH 2 ) r -NH 2 ,-(OCH 2 -CH 2 ) r -NHC(NH)NH 2 , or - (OCH 2 -CH 2 ) r -S-S[CH 2 CH 2 ] s NHC (NH) NH 2 Replaced with (C 1 ~C 6 ) is alkyl, and here, P 1 However, H, (C 1 ~C 8 ) alkyl, (C 2 ~C 8 ) Alkenil, (C 2 ~C 8 ) Alkinnil, (C 3 ~C 8 ) Aryl, (C 3 ~C 8 ) Cycloalkyl, (C 3 ~C 8 ) Aryl (C 1 ~C 6 ) Alkylene or (C 3 ~C 8 ) Cycloalkyl (C 1 ~C 6 The in vitro method according to claim 10, wherein the alkylene is an integer from 0 to 50, r is an integer from 1 to 50, and s is an integer from 1 to 50.
12. The guanidine-containing group is 【Transformation 7】 The in vitro method according to claim 10, wherein n = 1, 2, 3, 4, or 5.
13. The one or more N or O portions are -O-, -OH, -C(O)-, -NH-, -NH 2 , or the three in vitro methods according to claim 10, comprising -C(O)NH-.
Citation Information
Patent Citations
Gamma-PNA miniprobes for fluorescent labeling
US20150197793A1
Divalent nucleobase compounds and uses therefor
US20160083434A1