Phenylalanine hydroxylase variants and their use
Patent Information
- Application Number
- JP2026049346
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-06-01
- Filing Date
- 2026-03-24
- Publication Date
- 2026-09-07
AI Technical Summary
Current treatments for phenylketonuria (PKU), such as dietary restrictions and BH4 therapy, face challenges including dietary compliance, neurological issues, nutritional deficiencies, and financial burden, with BH4 therapy being ineffective for a significant portion of patients.
Development of variant phenylalanine hydroxylase polypeptides with specific amino acid substitutions and optimized mRNA encoding these polypeptides to enhance protein expression and activity, delivered via lipid nanoparticles, aiming to reduce phenylalanine levels in the body.
The variant phenylalanine hydroxylase polypeptides and optimized mRNA increase phenylalanine hydroxylase activity, effectively reducing phenylalanine levels and potentially alleviating the symptoms of PKU, including neurological impairments.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims priority based on U.S. Provisional Patent Application No. 63 / 032,875, filed on 1 June 2020, the contents of which are incorporated herein by reference in their entirety.
[0002] Sequence List This application includes an electronically submitted sequence listing in ASCII format, which is incorporated herein by reference in its entirety. The ASCII copy, created on June 1, 2021, is named 45817-0088WO1_SL.txt and is 305,608 bytes in size. [Background technology]
[0003] Phenyletonuria (PKU) is an autosomal recessive congenital metabolic disorder caused by a deficiency in the liver enzyme phenylalanine hydroxylase (PAH). This disorder is found in all ethnic groups, and its incidence varies greatly across the world, with the highest incidence in Northern Europe. PKU is primarily caused by mutations in the PAH gene that reduce catalytic activity, affecting the phenylalanine (Phe) catabolic pathway. The PAH enzyme requires the activity of the cofactor tetrahydrobiopterin (BH4) to convert Phe to tyrosine (Tyr). A deficiency in PAH or either of its cofactors leads to the accumulation of excess phenylalanine. If left untreated, this can result in severe and irreversible intellectual disability. Other clinical features associated with untreated PKU include autistic behaviors, motor deficits, eczematous rashes, and seizures. Behavioral and mental disorders generally develop with age.
[0004] Currently, there is no cure for PKU. Common treatment mainly involves limiting phenylalanine intake to the minimum necessary for normal growth and supplementing with specially formulated medical diets. However, current dietary therapies have at least four major problems: (i) dietary compliance due to the unpalatable taste of the food, (ii) persistent neurological or psychiatric problems and poor quality of life despite early intervention, (iii) the possibility of nutritional deficiencies due to dietary restrictions, and (iv) the financial burden due to the cost of special medical diets and nutritional supplements.
[0005] A group of patients with hyperphenylalaninemia who have mutations in cofactor BH4 are typically treated with a synthetic BH4 analog called sapropterin. PAH-deficient PKU and BH4-deficient PKU can be distinguished by a BH4 loading test. Treatment with synthetic biopterin compounds or sapropterin is generally successful in about 90% of patients with classical PKU (which account for about 50-80% of patients detected by neonatal screening (PAHdb, pahdb.mcgill.ca)) in BH4-responsive PKU patients, but BH4 therapy does not have a beneficial effect. [Overview of the project]
[0006] This disclosure provides a variant phenylalanine hydroxylase polypeptide and a polynucleotide encoding the variant polypeptide. Substitutions of certain amino acid residues in phenylalanine hydroxylase have been found to increase protein expression and / or biological activity.
[0007] In one embodiment, the present invention features a polypeptide (hereinafter referred to as a "variant phenylalanine hydroxylase polypeptide") having an amino acid sequence that is at least 80% identical to the amino acids at positions 118 to 452 of SEQ ID NO: 1, wherein (a) the amino acid sequence includes an amino acid other than methionine substituted at the position corresponding to position 180 of SEQ ID NO: 1, (b) the amino acid sequence includes an amino acid other than lysine substituted at the position corresponding to position 150 of SEQ ID NO: 1, (c) the amino acid sequence includes an amino acid other than glycine substituted at the position corresponding to position 256 of SEQ ID NO: 1, (d) the amino acid sequence includes an amino acid other than asparagine substituted at the position corresponding to position 376 of SEQ ID NO: 1, and (e) the amino acid sequence includes an amino acid other than lysine substituted at the position corresponding to position 361 of SEQ ID NO: 1, and the polypeptide exhibits phenylalanine hydroxylase activity when tetramerized.
[0008] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the methionine residue at position 180 of SEQ ID NO: 1 is substituted with threonine.
[0009] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the methionine residue at position 180 of SEQ ID NO: 1 is substituted with alanine.
[0010] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the methionine residue at position 180 of SEQ ID NO: 1 is substituted with glutamine.
[0011] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the methionine residue at position 180 of SEQ ID NO: 1 is substituted with asparagine.
[0012] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the lysine residue at position 150 of SEQ ID NO: 1 is substituted with threonine.
[0013] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the lysine residue at position 150 of SEQ ID NO: 1 is substituted with arginine.
[0014] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the lysine residue at position 150 of SEQ ID NO: 1 is substituted with proline.
[0015] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the lysine residue at position 150 of SEQ ID NO: 1 is substituted with asparagine.
[0016] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the glycine residue at position 256 of SEQ ID NO: 1 is substituted with alanine.
[0017] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the glycine residue at position 256 of SEQ ID NO: 1 is substituted with valine.
[0018] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the asparagine residue at position 376 of SEQ ID NO: 1 is substituted with glutamic acid.
[0019] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the asparagine residue at position 376 of SEQ ID NO: 1 is substituted with lysine.
[0020] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the lysine residue at position 361 of SEQ ID NO: 1 is substituted with arginine.
[0021] In one embodiment of the variant phenylalanine hydroxylase polypeptide, the lysine residue at position 361 of SEQ ID NO: 1 is substituted with glutamic acid.
[0022] In some embodiments of any of the variant phenylalanine hydroxylase polypeptides described herein, their amino acid sequences, including the specified substitutions, are at least 90% identical to amino acids 118-452 of SEQ ID NO: 1.
[0023] In some embodiments of any of the variant phenylalanine hydroxylase polypeptides described herein, their amino acid sequences, including the identified substitutions, are at least 95% identical to amino acids 118-452 of SEQ ID NO: 1.
[0024] In some embodiments of the variant phenylalanine hydroxylase polypeptides described herein, their amino acid sequences are at least 90% identical to SEQ ID NO: 1, including the specified substitutions.
[0025] In some embodiments of the variant phenylalanine hydroxylase polypeptides described herein, their amino acid sequences are at least 95% identical to SEQ ID NO: 1, including the specified substitutions.
[0026] In some embodiments, the variant phenylalanine hydroxylase polypeptide comprises (or consists of) the amino acid sequence defined in SEQ ID NO: 3.
[0027] In some embodiments, the variant phenylalanine hydroxylase polypeptide comprises (or consists of) the amino acid sequence defined in SEQ ID NO: 4.
[0028] In some embodiments, the variant phenylalanine hydroxylase polypeptide comprises (or consists of) the amino acid sequence defined in SEQ ID NO: 5.
[0029] In some embodiments, the variant phenylalanine hydroxylase polypeptide comprises (or consists of) the amino acid sequence defined in SEQ ID NO: 6.
[0030] In some embodiments, the variant phenylalanine hydroxylase polypeptide comprises (or consists of) the amino acid sequence defined in SEQ ID NO: 7.
[0031] In some embodiments, the variant phenylalanine hydroxylase polypeptide contains (or consists of) the amino acid sequence defined in SEQ ID NO: 8.
[0032] In some embodiments, the variant phenylalanine hydroxylase polypeptide comprises (or consists of) the amino acid sequence defined in SEQ ID NO: 9.
[0033] In some embodiments, the variant phenylalanine hydroxylase polypeptide comprises (or consists of) the amino acid sequence defined in SEQ ID NO: 10.
[0034] In some embodiments, the variant phenylalanine hydroxylase polypeptide comprises (or consists of) the amino acid sequence defined in SEQ ID NO: 11.
[0035] In some embodiments, the variant phenylalanine hydroxylase polypeptide comprises (or consists of) the amino acid sequence defined in SEQ ID NO: 12.
[0036] This disclosure also features polynucleotides comprising an open reading frame (ORF) encoding one of the variant phenylalanine hydroxylase polypeptides described herein.
[0037] This disclosure also features messenger RNA (mRNA) containing polynucleotides, including an ORF encoding one of the variant phenylalanine hydroxylase polypeptides described herein. mRNA pharmaceuticals are particularly well suited to the treatment of PKUs because the technique involves delivering mRNA encoding a variant PAH into a cell, followed by de novo synthesis of a functional variant PAH protein within the target cell. The present invention features the introduction of modified nucleotides into therapeutic mRNA to (1) minimize undesirable immune activation (e.g., innate immune response associated with the introduction of foreign nucleic acids into the body) and (2) optimize the efficiency of mRNA translation into protein. Exemplary embodiments of this disclosure feature the combination of nucleotide modifications to mitigate the innate immune response and the optimization of the sequence, particularly the sequence within the open reading frame (ORF) of therapeutic mRNA encoding a variant PAH, to optimize the sequence, and especially to enhance protein expression.
[0038] In some embodiments of the mRNA described herein, the ORF is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleotide sequence of SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, or SEQ ID NO: 31.
[0039] In some embodiments of the mRNA described herein, the ORF is 100% identical to the nucleotide sequence of SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, or SEQ ID NO: 31.
[0040] In some embodiments of the mRNA described herein, the mRNA includes a 5'UTR containing a nucleic acid sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 55 or SEQ ID NO: 56.
[0041] In some embodiments of the mRNA described herein, the mRNA includes a 3'UTR containing a nucleic acid sequence that is at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 113, SEQ ID NO: 107, SEQ ID NO: 108, SEQ ID NO: 112, SEQ ID NO: 114, SEQ ID NO: 109, SEQ ID NO: 110, or SEQ ID NO: 111.
[0042] In some embodiments of the mRNA described herein, the mRNA comprises the nucleotide sequence of SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 250, or SEQ ID NO: 251.
[0043] In some embodiments of the mRNA described herein, the mRNA includes a 5' terminal cap (e.g., Cap0, Cap1, ARCA, inosine, N1-methyl-guanosine, 2'-fluoro-guanosine, 7-deaza-guanosine, 8-oxo-guanosine, 2-amino-guanosine, LNA-guanosine, 2-azidoguanosine, Cap2, Cap4, 5'methyl G cap, or analogs thereof).
[0044] In some embodiments of the mRNA described herein, the mRNA includes a polyA region (for example, the polyA region is at least about 10 nucleotides long, at least about 20 nucleotides long, at least about 30 nucleotides long, at least about 40 nucleotides long, at least about 50 nucleotides long, at least about 60 nucleotides long, at least about 70 nucleotides long, at least about 80 nucleotides long, at least about 90 nucleotides long, or at least about 100 nucleotides long). In some embodiments, the mRNA includes a nucleotide sequence UCUAGAAAAAAAAAAAAAAAAAAAA (SEQ ID NO: 211) located at the 3' end of the mRNA relative to the polyA region.
[0045] In some embodiments of the mRNA described herein, the mRNA contains an inverted deoxythymidine at its 3' end.
[0046] In some embodiments of the mRNA described herein, all of the mRNA uracil is N1-methylpseudolacil.
[0047] In some embodiments of the mRNA described herein, the ORF is 100% identical to SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, or SEQ ID NO: 31, and the mRNA contains a poly(A) region of at least about 100 nucleotides in length, and all of the uracil in the mRNA is N1-methylpseudolacil.
[0048] In some embodiments of the mRNA described herein, the mRNA comprises a 5' terminal cap containing a guanine cap nucleotide with N7 methylation, the 5' terminal nucleotide of the mRNA contains 2'-O-methyl, the mRNA comprises the nucleotide sequence of SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 250, or SEQ ID NO: 251, the mRNA comprises a poly-A region of at least about 100 nucleotides in length, and all of the uracil in the mRNA is N1-methylpseudolacil. In some embodiments, the mRNA comprises the nucleotide sequence UCUAGAAAAAAAAAAAAAAAAAAAA (SEQ ID NO: 211) located at the 3' end of the mRNA relative to the poly-A region. In some embodiments, the mRNA comprises an inverted deoxythymidine at the 3' end of the mRNA.
[0049] This disclosure also features an expression vector comprising a polynucleotide containing an ORF encoding one of the variant phenylalanine hydroxylase polypeptides described herein.
[0050] In some embodiments of the expression vectors described herein, the ORF is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleotide sequence of SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, or SEQ ID NO: 31.
[0051] In some embodiments of the expression vectors described herein, the ORF is 100% identical to the nucleotide sequence of SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, or SEQ ID NO: 31.
[0052] In some embodiments of the expression vectors described herein, the expression vector is a plasmid, a minimally immunologically defined gene expression (MIDGE) vector, a closed-end linear double-stranded DNA (ceDNA), or a viral vector.
[0053] In some embodiments of the expression vectors described herein, the expression vector includes a transcriptional regulatory element (e.g., a promoter, enhancer, or stop sequence).
[0054] This disclosure also features lipid nanoparticles comprising polynucleotides, mRNAs, or expression vectors as described herein. In some embodiments, the lipid nanoparticles are (i) Compound II, (ii) Cholesterol and (iii) PEG-DMG or Compound I, (i) Compound VI, (ii) Cholesterol and (iii) PEG-DMG or Compound I, (i) Compound II, (ii) DSPC or DOPE, (iii) Cholesterol and (iv) PEG-DMG or Compound I, (i) Compound VI, (ii) DSPC or DOPE, (iii) Cholesterol and (iv) PEG-DMG or Compound I, (i) Compound II, (ii) Cholesterol and (iii) Compound I, or (i) Compound II, (ii) DSPC or DOPE, (iii) Cholesterol, and (iv) Compound I Includes.
[0055] This disclosure also features pharmaceutical compositions comprising polypeptides, polynucleotides, mRNAs, expression vectors, or lipid nanoparticles as described herein.
[0056] This disclosure also features a method for expressing a phenylalanine hydroxylase polypeptide in a human subject that requires the expression of the phenylalanine hydroxylase polypeptide, comprising administering to the human subject an effective amount of the polypeptide described herein, the polynucleotide described herein, the mRNA described herein, the expression vector described herein, the lipid nanoparticles described herein, or the pharmaceutical composition described herein.
[0057] This disclosure also features a method for improving phenylalanine hydroxylase activity in a human subject that requires improvement of phenylalanine hydroxylase activity, the method comprising administering to the human subject an effective amount of a polypeptide, polynucleotide, mRNA, expression vector, lipid nanoparticle, or pharmaceutical composition described herein.
[0058] The disclosure also features a method for treating, preventing or delaying the onset and / or progression of phenylketonuria in a human subject who requires treatment, prevention or delay of the onset and / or progression of phenylketonuria, the method comprising administering to the human subject in an effective amount of a polypeptide, polynucleotide, mRNA, expression vector, lipid nanoparticle, or pharmaceutical composition as described herein.
[0059] The Disclosure also features a method for reducing phenylalanine levels (e.g., phenylalanine levels in blood, plasma, serum, liver, and / or urine) in a human subject that requires such reduction, the method comprising administering to the human subject an effective amount of a polypeptide, polynucleotide, mRNA, expression vector, lipid nanoparticle, or pharmaceutical composition described herein.
[0060] In some embodiments of the methods described herein, tetrahydrobiopterin (BH4), its analogs, salts of tetrahydrobiopterin (BH4), or salts of its analogs are administered to human subjects in combination with the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention (e.g., orally). In some embodiments, the tetrahydrobiopterin (BH4) analog or salt thereof is sapropterin, 6-hydroxymethylpterin (HMP), 6-acetyl-7,7-dimethyl-7,8-dihydropterin (ADDP), or salts thereof. Tetrahydrobiopterin (BH4), its analogs, salts of tetrahydrobiopterin (BH4), or salts of its analogs may be administered simultaneously with, or before or after, the administration of the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention.
[0061] In some embodiments, the subjects are administered the drug approximately once a week, approximately once every two weeks, or approximately once a month.
[0062] In certain embodiments, the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention are administered intravenously. In some cases, the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention are administered in doses of 0.1 mg / kg to 5.0 mg / kg. In some cases, the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention are administered in doses of 0.1 mg / kg to 2.0 mg / kg. In some cases, the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention are administered in doses of 0.1 mg / kg to 1.5 mg / kg. In some cases, the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention are administered in doses of 0.1 mg / kg to 1.0 mg / kg. In some cases, the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention are administered in doses of 0.1 mg / kg to 0.5 mg / kg.
[0063] In certain embodiments, the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention are administered subcutaneously. In some cases, the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention are administered in doses of 0.1 mg / kg to 5.0 mg / kg. In some cases, the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention are administered in doses of 0.1 mg / kg to 2.0 mg / kg. In some cases, the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention are administered in doses of 0.1 mg / kg to 1.5 mg / kg. In some cases, the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention are administered in doses of 0.1 mg / kg to 1.0 mg / kg. In some cases, the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention are administered in doses of 0.1 mg / kg to 0.5 mg / kg.
[0064] In certain embodiments, the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention are administered intravenously and subcutaneously. In some cases, the method of the present invention includes administering the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention intravenously once or more times, followed by administering the polypeptides, polynucleotides, mRNAs, expression vectors, lipid nanoparticles, or pharmaceutical compositions of the present invention subcutaneously once or more times. [Brief explanation of the drawing]
[0065] [Figure 1] This bar graph shows the PAH protein expression levels 96 hours after transfection with a construct encoding a truncated PAH variant, and compares them to the expression levels when transfected with the Δrd PAH construct. [Figure 2] These are Western blot images showing PAH protein expression levels 24 hours and 96 hours after transfection, following transfection with a construct encoding a truncated PAH variant or a Δrd construct. [Figure 3] This bar graph shows the PAH protein expression levels 96 hours after transfection with a construct encoding a truncated PAH variant, compared to the expression levels 24 hours after transfection with a fully-length wild-type PAH construct or a Δrd PAH construct. [Figure 4]This is a Western blot image showing the PAH protein expression level 96 hours after transfection with a construct encoding a truncated or full-length PAH variant, compared to the expression levels 24 and 96 hours after transfection when transfected with a full-length wild-type PAH construct or a Δrd PAH construct. [Figure 5] This bar graph shows the PAH protein expression levels after transfection with a truncated or full-length M180T PAH construct (with or without a stable tail), at 24 and 96 hours after transfection, and compares them to the expression levels after transfection with a full-length wild-type PAH construct or a Δrd PAH construct (with or without a stable tail). [Figure 6] This is a Western blot image showing the PAH protein expression levels after transfection with a truncated or full-length M180T PAH construct (with or without a stable tail), at 24 hours and 96 hours after transfection, and compared to the expression levels after transfection with a full-length wild-type PAH construct or a Δrd PAH construct (with or without a stable tail). [Figure 7] This graph shows PAH activity in homozygous PAHenu2 mice transfected with either the full-length wild-type PAH construct (C1 A100) or the full-length wild-type PAH construct containing a stable tail (C1 idT). The activity is compared with that of control mice injected with PBS, and the blood phenylalanine levels were measured over time. [Figure 8]This graph shows PAH activity in homozygous PAHenu2 mice transfected with a construct encoding a truncated PAH variant or a Δrd PAH construct, and compares it with the activity in PBS-injected control mice by measuring blood phenylalanine levels over time. [Figure 9] This graph shows PAH activity in homozygous PAHenu2 mice transfected with a full-length M180T PAH construct or K150T PAH construct (with or without a stable tail), or a full-length wild-type PAH construct (with or without a stable tail). The activity is compared with that of control mice injected with PBS, and the blood phenylalanine levels were measured over time. [Modes for carrying out the invention]
[0066] The present invention provides a variant phenylalanine hydroxylase polypeptide having a substitution at a predetermined amino acid residue. As disclosed in the accompanying examples, it has been found that the substitution of a predetermined amino acid residue in phenylalanine hydroxylase increases protein expression and / or biological activity.
[0067] Variant phenylalanine hydroxylase (PAH) polypeptide Phenylalanine hydroxylase (PAH, EC1.14.16.1) catalyzes the conversion of L-phenylalanine (L-Phe) to L-tyrosine (L-Tyr) by para-hydroxylation of its aromatic side chain. In mammals, this tetrahydrobiopterin (BH4)-dependent reaction is the first and rate-limiting step in the breakdown of excess L-Phe from dietary proteins, with L-Tyr being further broken down to become a product of the citric acid cycle. PAH consumes approximately 75% of the phenylalanine supply from dietary and physiological protein catabolism.
[0068] PAHs are primarily found in the liver, where they inhibit the neurotoxic effects of hyperphenylalaninemia (HPA) by removing excess L-Phe. However, since L-Phe is also an essential amino acid that makes up proteins, it is important that it is not completely catabolized. To fulfill this dual role of effectively removing excess L-Phe while also maintaining its levels, several regulatory mechanisms have developed in mammalian PAHs.
[0069] Mammalian PAHs are homotetrameric enzymes consisting of 50 kDa subunits. Each PAH subunit comprises an N-terminal regulatory domain (residues 1-110), a central catalytic domain (residues 111-410), and a C-terminal oligomerization domain (residues 411-452). The N-terminal regulatory domain is flexibly bound to the catalytic domain via a hinge region (Arg111-Thr117). The regulatory domain is necessary for the expression of regulatory properties, such as activation by L-Phe, and there is ongoing debate as to whether the regulatory domain contains an allosteric binding site for L-Phe. The catalytic domain contains binding sites for iron, cofactors, and substrates. In the active site, iron binds to two histidines (His285 and His290 in hPAHs) and glutamic acid (Glu330 in hPAHs). The oligomerization domain begins with antiparallel sheets (residues 411-414 and 421-424) responsible for dimerization, followed by a 40 Å helix (residues 428-452) that mediates tetramerization through domain swapping with the other monomer and the formation of antiparallel coiled coils.
[0070] In humans, mutations in the PAH gene lead to phenylketonuria (PKU), with most mutations primarily associated with PAH misfolding and instability. There are over 500 pathogenic mutations (pahdb.mcgill.ca and biopku.org). PAH dysfunction leads to elevated blood L-Phe levels, and metabolites resulting from the transamination of L-Phe to phenylpyruvate appear in the urine. This is characteristic of hyperphenylalaninemia (HPA), the most severe form of which is classical phenylketonuria (PKU), characterized by plasma L-Phe levels exceeding 1,200 μM. L-Phe accumulates, subsequently impairing neurotransmitters in the brain, leading to neurological symptoms including intellectual disability, aimless movement, and depressive symptoms. Therefore, strict management of L-Phe dietary intake is necessary in PKU patients.
[0071] The amino acid sequence of the full-length human PAH is shown in Sequence ID No. 1. The amino acid sequence of the truncated human PAH (referred to as PAH-ΔRD protein, PAHΔRD protein, or ΔrdPAH protein) is shown in Sequence ID No. 2. This truncated form is 336 amino acids long and contains the catalytic domain of the PAH. A considerable number of human PAH mutations have been found in the catalytic domain.
[0072] The present invention discloses a variant PAH polypeptide having substitutions in one or more predetermined amino acid residues of the PAH polypeptide.
[0073] A variant PAH polypeptide, compared to SEQ ID NO: 1, contains amino acid substitutions at (i) the methionine residue at position 180, (ii) the lysine residue at position 150, (iii) the glycine residue at position 256, (iv) the asparagine residue at position 376, or (v) the lysine residue at position 361. Unless otherwise stated, in this specification, when PAH amino acid residues are referred to by positional numbers, they refer to the residue numbers in relation to SEQ ID NO: 1. Additional amino acid substitutions can be introduced into a variant PAH polypeptide, and the variant PAH polypeptide maintains the biological activity of PAH.
[0074] PAH amino acid residues for which substitution is specified can be substituted with non-conserved or conserved amino acid residues.
[0075] Examples of amino acids that can substitute the methionine residue at position 180 include threonine, alanine, glutamine, and asparagine.
[0076] Exemplary amino acids that can substitute the lysine residue at position 150 include threonine, arginine, proline, and asparagine.
[0077] Examples of amino acids that can substitute the glycine residue at position 256 include alanine and valine.
[0078] Exemplary amino acids that can substitute the asparagine residue at position 376 include glutamic acid and lysine.
[0079] Examples of amino acids that can substitute the lysine residue at position 361 include arginine and glutamic acid.
[0080] In certain embodiments, the Disclosure provides a polynucleotide (e.g., RNA, e.g., mRNA) comprising a nucleotide sequence (e.g., an open reading frame (ORF)) encoding a variant PAH polypeptide. In some embodiments, the variant PAH polypeptide includes an amino acid substitution at the methionine residue at position 180. In some embodiments, the variant PAH polypeptide includes an amino acid substitution at the lysine residue at position 150. In some embodiments, the variant PAH polypeptide includes an amino acid substitution at the glycine residue at position 256. In some embodiments, the variant PAH polypeptide includes an amino acid substitution at the asparagine residue at position 376. In some embodiments, the variant PAH polypeptide includes an amino acid substitution at the lysine residue at position 361.
[0081] In some embodiments, the variant PAH polypeptide is a truncated human variant PAH protein (e.g., the protein of SEQ ID NO: 4, SEQ ID NO: 6, SEQ ID NO: 8, SEQ ID NO: 10, or SEQ ID NO: 12). In some embodiments, the variant PAH polypeptide is a full-length human variant PAH protein (e.g., the protein of SEQ ID NO: 3, SEQ ID NO: 5, SEQ ID NO: 7, SEQ ID NO: 9, or SEQ ID NO: 11). In some embodiments, a sequence tag or amino acids can be added to the sequence encoded by the polynucleotide of the present invention (e.g., the N-terminus or C-terminus) for localization, for example. In some embodiments, an amino acid residue located at the carboxyl terminus, amino terminus, or internal region of the polypeptide of the present invention can be optionally deleted to yield a fragment.
[0082] In some embodiments, the variant PAH polypeptide comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to amino acids 118–452 of SEQ ID NO: 1.
[0083] In some embodiments, the variant PAH polypeptide comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to the amino acids of SEQ ID NO: 1.
[0084] In some embodiments, the variant PAH polypeptide comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to the amino acids of SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, or SEQ ID NO: 12.
[0085] In some embodiments, a polynucleotide (e.g., RNA, e.g., mRNA) comprising the nucleotide sequence of the present invention (e.g., ORF) encodes a substitutional variant of a human PAH sequence, the substitutional variant may contain one, two, or three or more substitutions. In some embodiments, the substitutional variant may contain one or more conserved amino acid substitutions. In another embodiment, the variant is an insertional variant. In yet another embodiment, the variant is a deletional variant.
[0086] Protein fragments, functional protein domains, variants, and homologous proteins (orthologs) of PAHs are also within the scope of the PAH polypeptides of this disclosure. Non-limiting examples of polynucleotide-encoded polypeptides of the present invention are shown in SEQ ID NOs: 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12.
[0087] Polynucleotides and open reading frames (ORFs) The present invention features mRNA for use in the treatment or prevention of hyperphenylalaninemia such as PKU. The mRNA featured for use in the present invention is administered to a subject and encodes a variant PAH protein in vivo. Accordingly, the present invention relates to a polynucleotide, e.g., mRNA, comprising an open reading frame consisting of linked nucleosides, which encodes a variant human PAH (e.g., SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11 or SEQ ID NO: 12), its functional fragment, and a fusion protein comprising the variant PAH. In some embodiments, the open reading frame is sequence-optimized.
[0088] In certain embodiments, the present invention provides a polynucleotide (e.g., RNA such as mRNA) comprising a nucleotide sequence (e.g., ORF) encoding one or more variant PAH polypeptides (e.g., SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, or SEQ ID NO: 12).
[0089] In some embodiments, the polynucleotides of the present invention (e.g., RNA, e.g., mRNA), when introduced into cells, increase the intracellular PAH protein expression level and / or detectable PAH enzyme activity level by at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100% compared to, for example, the intracellular PAH protein expression level and / or detectable PAH enzyme activity level before administration of the polynucleotides of the present invention. The PAH protein expression level and / or PAH enzyme activity can be measured according to methods known in the art. In some embodiments, the polynucleotides are introduced into cells in vitro. In some embodiments, the polynucleotides are introduced into cells in vivo.
[0090] In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises a nucleotide sequence (e.g., ORF) encoding a variant human PAH, for example, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, or SEQ ID NO: 12.
[0091] In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises a nucleotide sequence (e.g., ORF) encoding a variant PAH polypeptide, the nucleotide sequence being at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence of SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, or SEQ ID NO: 31.
[0092] In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises a nucleotide sequence (e.g., ORF) encoding a variant PAH polypeptide, wherein the nucleotide sequence has sequence identity with a sequence selected from the group consisting of SEQ ID NOs: 22, 23, 24, 25, 26, 27, 28, 29, 30, or 31 of at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%.
[0093] In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises a nucleotide sequence (e.g., ORF) encoding a variant PAH polypeptide, wherein the nucleotide sequence has sequence identity with a sequence selected from the group consisting of SEQ ID NOs: 22, 23, 24, 25, 26, 27, 28, 29, 30, or 31 of 70%~100%, 75%~100%, 80%~100%, 85%~100%, 70%~95%, 80%~95%, 70%~85%, 75%~90%, 80%~95%, 70%~75%, 75%~80%, 80%~85%, 85%~90%, 90%~95%, or 95%~100%.
[0094] In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises an ORF encoding a variant PAH polypeptide, wherein the polynucleotide includes a nucleic acid sequence having sequence identity with a sequence selected from the group consisting of SEQ ID NOs: 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 of 70%~100%, 75%~100%, 80%~100%, 85%~100%, 70%~95%, 80%~95%, 70%~85%, 75%~90%, 80%~95%, 70%~75%, 75%~80%, 80%~85%, 85%~90%, 90%~95%, or 95%~100%.
[0095] In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises a nucleotide sequence (e.g., ORF) encoding a variant PAH polypeptide, the nucleotide sequence being 70% to 90% identical, 75% to 85% identical, 76% to 84% identical, 77% to 83% identical, 77% to 82% identical, or 78% to 81% identical to the sequence of SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, or SEQ ID NO: 31.
[0096] In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises a nucleotide sequence (e.g., ORF) encoding a variant PAH polypeptide, and further comprises at least one nucleic acid sequence that is a non-coding site, e.g., a microRNA binding site. In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention further comprises a 5'UTR (e.g., SEQ ID NO: 55 or SEQ ID NO: 56) and a 3'UTR (e.g., a 3'UTR selected from the sequences of SEQ ID NO: 113, SEQ ID NO: 107, SEQ ID NO: 108, SEQ ID NO: 112, SEQ ID NO: 114, SEQ ID NO: 109, SEQ ID NO: 110, or SEQ ID NO: 111). In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises a sequence selected from the group consisting of SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, or SEQ ID NO: 31. In further embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) includes a 5' terminal cap (e.g., Cap0, Cap1, ARCA, inosine, N1-methyl-guanosine, 2'-fluoro-guanosine, 7-deaza-guanosine, 8-oxo-guanosine, 2-amino-guanosine, LNA-guanosine, 2-azidoguanosine, Cap2, Cap4, 5'-methyl G cap, or analogs thereof) and a poly-A tail region (e.g., about 100 nucleotides long). In further embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) includes a 3'UTR containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 107 or 108, or a combination thereof. In some embodiments, the mRNA includes a 3'UTR containing one of the nucleic acid sequences: SEQ ID NOs: 113, SEQ ID NOs: 107, SEQ ID NOs: 108, SEQ ID NOs: 112, SEQ ID NOs: 114, SEQ ID NOs: 109, SEQ ID NOs: 110, or SEQ ID NOs: 111. In some embodiments, the mRNA includes a poly-A tail.In some cases, the poly-A tail is 50-150 nucleotides long (sequence number 197), 75-150 nucleotides long (sequence number 198), 85-150 nucleotides long (sequence number 199), 90-150 nucleotides long (sequence number 192), 90-120 nucleotides long (sequence number 193), 90-130 nucleotides long (sequence number 194), or 90-150 nucleotides long (sequence number 192). In some cases, the poly-A tail is 100 nucleotides long (sequence number 195).
[0097] In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises a nucleotide sequence (e.g., ORF) encoding a variant PAH polypeptide, and is single-stranded or double-stranded.
[0098] In some embodiments, the polynucleotide of the present invention, comprising a nucleotide sequence (e.g., ORF) encoding a variant PAH polypeptide, is DNA or RNA. In some embodiments, the polynucleotide of the present invention is RNA. In some embodiments, the polynucleotide of the present invention is mRNA or functions as mRNA. In some embodiments, the mRNA comprises a nucleotide sequence (e.g., ORF) encoding at least one PAH polypeptide, and can be translated in vitro, in vivo, in situ, or ex vivo to produce the encoded PAH polypeptide.
[0099] In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises a sequence-optimized nucleotide sequence (e.g., ORF) encoding a variant PAH polypeptide (e.g., SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, or SEQ ID NO: 31), wherein the polynucleotide comprises at least one chemically modified nucleic acid base, e.g., N1-methylpseudracil or 5-methoxyuracil. In certain embodiments, all uracils in the polynucleotide are N1-methylpseudracil. In other embodiments, all uracils in the polynucleotide are 5-methoxyuracil. In some embodiments, the polynucleotide further comprises miRNA binding sites, e.g., a miRNA binding site that binds to miR-142 and / or a miRNA binding site that binds to miR-126.
[0100] In some embodiments, the polynucleotides (e.g., RNA, e.g., mRNA) disclosed herein are formulated with a delivery agent containing, for example, a compound having formula (I), e.g., any of compounds 1 to 232, e.g., compound II, a compound having formula (III), (IV), (V), or (VI), e.g., any of compounds 233 to 342, e.g., compound VI, or a compound having formula (VIII), e.g., any of compounds 419 to 428, e.g., compound I, or a combination thereof. In some embodiments, the delivery agent contains compound II, DSPC, cholesterol, and compound I or PEG-DMG in a molar ratio of, for example, about 50:10:38.5:1.5. In some embodiments, the delivery agent delivers compound VI, DSPC, cholesterol, and compound I or PEG-DMG, for example, about 30 to about 60 mol% of compound II or VI (or related suitable amino lipids) (e.g., 30 to 40 mol%, 40 to 45 mol%, 45 to 50 mol%, 50 to 55 mol%, or 55 to 60 mol%), and about 5 to about 20 mol% of phospholipids (or related suitable phospholipids or "helper lipids") (e.g., 5 to 10 mol%, 10 to 15 mol%, or 15 mol%). The formulation contains approximately 20 mol% of cholesterol (or related sterols or "non-cationic" lipids), approximately 20 to approximately 50 mol% (e.g., approximately 20 to 30 mol%, 30 to 35 mol%, 35 to 40 mol%, 40 to 45 mol%, or 45 to 50 mol%), and approximately 0.05 to approximately 10 mol% of PEG lipids (or other suitable PEG lipids) (e.g., 0.05 to 1 mol%, 1 to 2 mol%, 2 to 3 mol%, 3 to 4 mol%, 4 to 5 mol%, 5 to 7 mol%, or 7 to 10 mol%). An example of a delivery agent may have a molar ratio of, for example, 47.5:10.5:39.0:3.0 or 50:10:38.5:1.5.In certain cases, exemplary delivery agents have molar ratios such as 47.5:10.5:39.0:3, 47.5:10:39.5:3, 47.5:11:39.5:2, 47.5:10.5:39.5:2.5, 47.5:11:39:2.5, 48.5:10:38.5:3, 48.5:10.5:39:2, 48.5:10.5:38.5:2.5, 48 The ratios can be 0.5:10.5:39.5:1.5, 48.5:10.5:38.0:3, 47:10.5:39.5:3, 47:10:40.5:2.5, 47:11:40:2, 47:10.5:39.5:3, 48:10.5:38.5:3, 48:10:39.5:2.5, 48:11:39:2, or 48:10.5:38.5:3. In some embodiments, the delivery agent comprises compound II or VI, DSPC, cholesterol, and compound I or PEG-DMG in a molar ratio of, for example, about 47.5:10.5:39.0:3.0. In some embodiments, the delivery agent comprises compound II or VI, DSPC, cholesterol, and compound I or PEG-DMG in a molar ratio of, for example, about 50:10:38.5:1.5.
[0101] In some embodiments, the polynucleotide of the present disclosure is mRNA comprising a 5' terminal cap (e.g., Cap1), a 5' UTR (e.g., SEQ ID NO: 55 or SEQ ID NO: 56), an ORF sequence encoding a polypeptide including SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11 or SEQ ID NO: 12, a 3' UTR (e.g., SEQ ID NO: 108, 112, 114 or 111), and a poly-A tail (e.g., about 100 nt in length), wherein all uracil in the polynucleotide is N1-methylpseudolacil. In some embodiments, the delivery agent comprises compound II or compound VI as an ionizable lipid, and PEG-DMG or compound I as a PEG lipid.
[0102] In some embodiments, the polynucleotide of the Disclosure is mRNA comprising a 5' terminal cap (e.g., Cap1), a 5' UTR (e.g., SEQ ID NO: 55 or SEQ ID NO: 56), an ORF sequence selected from the group consisting of SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30 or SEQ ID NO: 31, a 3' UTR (e.g., SEQ ID NO: 113, SEQ ID NO: 107, SEQ ID NO: 108, SEQ ID NO: 112, SEQ ID NO: 114, SEQ ID NO: 109, SEQ ID NO: 110 or SEQ ID NO: 111), and a poly-A tail (e.g., about 100 nt in length), wherein all uracil in the polynucleotide is N1-methylpseudolacil. In some embodiments, the delivery agent comprises compound II or compound VI as an ionizable lipid, and PEG-DMG or compound I as a PEG lipid.
[0103] signal sequence The polynucleotides (e.g., RNA, e.g., mRNA) of the present invention may also include nucleotide sequences encoding additional features that facilitate the transport of the encoded polypeptide to a therapeutically appropriate site. One such feature that assists protein transport is a signal sequence, or targeting sequence. Peptides encoded by these signal sequences are known by various names, including targeting peptides, transit peptides, and signal peptides. In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) includes a nucleotide sequence (e.g., ORF) encoding a signal peptide, which is functionally ligated to the nucleotide sequence encoding the PAH polypeptide described herein.
[0104] In some embodiments, the “signal sequence” is a polynucleotide of approximately 30–210 nucleotides in length, e.g., approximately 45–80 or 15–60 nucleotides, optionally introduced at the 5' end of the coding region of the polypeptide, and the “signal peptide” is a polypeptide of, e.g., approximately 20, 30, 40, 50, 60, or 70 amino acids in length, optionally introduced at its N-terminus. By adding these sequences, the encoded polypeptide is transported to a desired site, such as the endoplasmic reticulum or mitochondria, through one or more targeting pathways. After the protein has been transported to the desired site, some signal peptides are cleaved from the protein, for example, by a signal peptidase.
[0105] In some embodiments, the polynucleotide of the present invention comprises a nucleotide sequence encoding a PAH polypeptide, the nucleotide sequence further comprising a 5' nucleic acid sequence encoding a heterogeneous signal peptide.
[0106] Fusion protein In some embodiments, the polynucleotide of the present invention (e.g., RNA, e.g., mRNA) may contain two or more nucleic acid sequences (e.g., ORFs) encoding the polypeptide of interest. In some embodiments, the polynucleotide of the present invention may contain one ORF encoding a variant PAH polypeptide. However, in some embodiments, the polynucleotide of the present invention may contain two or more ORFs, e.g., a PAH polypeptide (the first polypeptide of interest), a first ORF encoding a functional fragment thereof or a variant thereof, and a second ORF expressing the second polypeptide of interest. In some embodiments, the two or more polypeptides of interest can be genetically fused, i.e., two or more polypeptides may be encoded by the same ORF. In some embodiments, the polynucleotide may contain a nucleic acid sequence encoding a linker (e.g., a peptide linker called G4S (SEQ ID NO: 200) or another linker known in the art) between the two or more polypeptides of interest.
[0107] In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention may contain two, three, or four or more ORFs, each expressing a target polypeptide.
[0108] In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention may include a first nucleic acid sequence (e.g., a first ORF) encoding a PAH polypeptide and a second nucleic acid sequence (e.g., a second ORF) encoding the desired second polypeptide.
[0109] Linker and cleavable peptide In certain embodiments, the mRNA of this disclosure encodes two or more PAH domains (e.g., PAH catalytic domain, PAH tetramerizing domain) or heterogeneous domains, collectively referred to as a multimeric construct. In certain embodiments of the multimeric construct, the mRNA further encodes a linker located between each domain. The linker may be, for example, a cleavable linker or a protease-sensitive linker. In certain embodiments, the linker is selected from the group consisting of F2A linkers, P2A linkers, T2A linkers, E2A linkers, and combinations thereof. This family of self-cleaving peptide linkers (referred to as 2A peptides) has been described in the art (see, for example, Kim, J. Het al. (2011) PLoS ONE 6:e18556). In certain embodiments, the linker is an F2A linker. In certain embodiments, the linker is a linker named GGGS (SEQ ID NO: 201). In a particular embodiment, the linker is a linker named (GGGS)n (SEQ ID NO: 202), where n is 2, 3, 4, or 5. In a particular embodiment, the polymer construct includes three domains having intervening linkers, and has a domain-linker-domain-linker-domain structure, for example, PAH domain-linker-PAH domain.
[0110] In one embodiment, the cleavable linker is an F2A linker (e.g., having the amino acid sequence GSGVKQTLNFDLLKLAGDVESNPGP (SEQ ID NO: 189)). In another embodiment, the cleavable linker is a T2A linker (e.g., having the amino acid sequence GSGEGRGSLLTCGDVEENPGP (SEQ ID NO: 190)), a P2A linker (e.g., having the amino acid sequence GSGATNFSLLLKQAGDVEENPGP (SEQ ID NO: 191)), or an E2A linker (e.g., having the amino acid sequence GSGQCTNYALLKLAGDVESNPGP (SEQ ID NO: 255)). It will be apparent to those skilled in the art that other linkers known in the art may be suitable for use in the constructs of the present invention (e.g., may be encoded by the polynucleotides of the present invention). Similarly, it will be apparent to those skilled in the art that other multi-cistronic constructs may be suitable for use in the present invention. In exemplary embodiments, this construct design also yields nearly equimolar amounts of intrabodies and / or domains encoded by the construct of the present invention.
[0111] In one embodiment, the self-cleaving peptide may, but is not limited to, a 2A peptide. For example, various 2A peptides are known and available in the art, including foot-and-mouth disease virus (FMDV) 2A peptide, equine rhinitis A virus 2A peptide, Thosea asigna virus 2A peptide, and porcine tesiovirus-1 2A peptide, and these peptides may be used. Some viruses use 2A peptides to produce two proteins from one transcript by ribosome skipping, and inhibit normal peptide bonding in the 2A peptide sequence, resulting in the production of two discontinuous proteins from a single translation event. In an unrestricted example, the 2A peptide may have the protein sequence of SEQ ID NO: 191, a fragment thereof, or a variant thereof. In one embodiment, the 2A peptide is cleaved between the last glycine and the last proline. In another unrestricted example, the polynucleotide of the present invention may include a polynucleotide sequence encoding a 2A peptide having the protein sequence of a fragment or variant of SEQ ID NO: 191. An example of a polynucleotide sequence encoding the 2A peptide is GGAAGGGAGCUACUAACUUCAGCCUGCUGAAGCAGGCUGGAGACGUGGAGGAGAACCCUGGACCU (SEQ ID NO: 256). In one exemplary embodiment, the 2A peptide is encoded by the sequence 5'-UCCGGACUCAGAUCCGGGGAUCUCAAAAUUGUCGCUCCUGUCAAACAAACUCUUAACUUUGAUUUACUCAAACUGGCTGGGGAUGUAGAAAGCAAUCCAGGTCCACUC-3' (SEQ ID NO: 257). The polynucleotide sequence of the 2A peptide may be modified or codon-optimized by the methods described herein and / or methods known in the art.
[0112] In one embodiment, this sequence may be used to separate the coding regions of two or more target polypeptides. As a non-limiting example, the sequence encoding the F2A peptide may be located between a first coding region A and a second coding region B (A-F2Apep-B). The presence of the F2A peptide causes a long protein to be cleaved between the glycine and proline at the end of the F2A peptide sequence (NPGP (SEQ ID NO: 260) is cleaved into NPG and P), thereby creating a separate protein A (with 21 amino acids of the F2A peptide attached and terminated with NPG) and a separate protein B (with one amino acid P of the F2A peptide attached). Similarly, with other 2A peptides (P2A, T2A, and E2A), the presence of the peptide in a long protein causes it to be cleaved between the glycine and proline at the end of the 2A peptide sequence (NPGP (SEQ ID NO: 260) is cleaved into NPG and P). Proteins A and B may be the same or different peptides or polypeptides for the same purpose (e.g., a full-length human PAH, or a PAH polypeptide such as a truncated form thereof containing a catalytic domain and a tetramerizing domain of the PAH). In certain embodiments, protein A and protein B are the PAH catalytic domain and the PAH tetramerizing domain (in either order). In certain embodiments, the first and second coding regions encode the PAH catalytic domain and the PAH tetramerizing domain in either order.
[0113] Sequence optimization of nucleotide sequences encoding PAH polypeptides In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention is sequence-optimized. In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises a nucleotide sequence (e.g., ORF) encoding a PAH polypeptide, optionally a nucleotide sequence (e.g., ORF) encoding another target polypeptide, a 5'UTR, a 3'UTR (the 5'UTR or 3'UTR optionally containing at least one microRNA binding site), optionally a nucleotide sequence encoding a linker, a poly-A tail, or a combination thereof, wherein the ORF(or ORF) is sequence-optimized.
[0114] A sequence-optimized nucleotide sequence, such as a codon-optimized mRNA sequence, that encodes a PAH polypeptide is a sequence that contains at least one equivalent nucleic acid base substitution with respect to a reference sequence (e.g., a wild-type nucleotide sequence encoding a PAH polypeptide).
[0115] Sequence-optimized nucleotide sequences can be partially or completely different from their reference sequences. For example, a reference sequence encoding polyserine uniformly encoded by the codon UCU can be optimized by making 100% of its nucleic acid bases substituted (replacing the U at position 1 with A, the C at position 2 with G, and the U at position 3 with C) to obtain a sequence encoding polyserine uniformly encoded by the AGC codon. The sequence identity percentage obtained from a comprehensive pairwise alignment of the reference polyserine nucleic acid sequence and the sequence-optimized polyserine nucleic acid sequence is 0%. However, the protein products obtained from both sequences are 100% identical.
[0116] Several sequence optimization (sometimes called codon optimization) methods are known in the art (discussed in more detail below) and can be useful in obtaining one or more desired results. These results include, for example, matching codon frequencies to specific tissue targets and / or host organisms to ensure correct folding; biasing G / C content to improve mRNA stability or reduce secondary structures; minimizing tandem repeat codon or base sequences that may impair gene construction or expression; customizing transcriptional and translational regulatory regions; inserting or removing protein transport sequences; removing / adding post-translational modification sites (e.g., glycosylation sites) in encoded proteins; adding, removing, or shuffling protein domains; inserting or deleting restriction sites; modifying ribosome binding sites and mRNA degradation sites; adjusting translation rates to ensure correct folding of various domains of a protein; and / or reducing or eliminating problematic secondary structures within polynucleotides. Sequence optimization tools, algorithms, and services are known in the art, and non-exclusive examples include services from GeneArt (Life Technologies), DNA2.0 (Menlo Park, CA), and / or proprietary methods.
[0117] Table 1 shows the codon options for each amino acid. [Table 1]
[0118] In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises a sequence-optimized nucleotide sequence (e.g., ORF) encoding a PAH polypeptide, a functional fragment thereof, or a variant thereof, wherein the PAH polypeptide, its functional fragment, or variant encoded by the sequence-optimized nucleotide sequence has improved properties (compared to, for example, a PAH polypeptide, its functional fragment, or variant encoded by a non-optimized reference nucleotide sequence), such as improved properties relating to expression potency after in vivo administration. Such properties include, but are not limited to, improved nucleic acid stability (e.g., mRNA stability), increased translational potency in target tissues, reduced number of truncated proteins expressed, improved folding of expressed proteins or prevention of misfolding of expressed proteins, reduced toxicity of expression products, reduced cell death caused by expression products, and increased and / or reduced protein aggregation.
[0119] In some embodiments, the sequence-optimized nucleotide sequences (e.g., ORFs) of the present invention are codon-optimized for expression in human subjects and have structural and / or chemical features that avoid one or more problems in the art, such as features useful for optimizing the formulation and delivery of nucleic acid-based pharmaceuticals while maintaining structural and functional integrity, features useful for overcoming expression thresholds, features useful for improving expression rate, half-life and / or protein concentration, features useful for optimizing protein localization, and features useful for avoiding adverse biological responses and / or degradation pathways such as immune responses.
[0120] In some embodiments, the polynucleotide of the present invention is (i) Creating a uridine-modified sequence by replacing at least one codon in a reference nucleotide sequence (e.g., an ORF encoding a PAH polypeptide) with an alternative codon to increase or decrease the uridine content. (ii) Replacing at least one codon in a reference nucleotide sequence (e.g., an ORF encoding a PAH polypeptide) with an alternative codon that has a higher codon frequency in the synonymous codon set. (iii) Increasing the G / C content by substituting at least one codon in the reference nucleotide sequence (e.g., ORF encoding a PAH polypeptide) with an alternative codon, or (iv) A combination of these, Nucleotide sequences whose sequences have been optimized according to a method including (e.g., nucleotide sequences encoding PAH polypeptides (e.g., ORFs), nucleotide sequences encoding another target polypeptide (e.g., ORFs), nucleic acid sequences encoding 5'UTR, 3'UTR, microRNA binding sites, linkers, or a combination of any of these) Includes.
[0121] In some embodiments, the sequence-optimized nucleotide sequence (e.g., an ORF encoding a PAH polypeptide) exhibits improvement in at least one property compared to the reference nucleotide sequence.
[0122] In some embodiments, the sequence optimization method is multiparametric and includes one, two, three, or four or more methods disclosed herein and / or other optimization methods known in the art.
[0123] Features that may be beneficial in some embodiments of the present invention can be encoded by or within a region of a polynucleotide, such region may be upstream (5'), downstream (3'), or within the region encoding the PAH polypeptide. These regions can be introduced into the polynucleotide before and / or after sequence optimization of the protein-coding region or open reading frame (ORF). Examples of such features include, but are not limited to, untranslated regions (UTRs), microRNA sequences, Kozak sequences, oligo(dT) sequences, poly(A) tails, and detectable tags, as well as multicloning sites that may have XbaI recognition sites.
[0124] In some embodiments, the polynucleotide of the present invention comprises a 5'UTR, a 3'UTR, and / or a microRNA binding site. In some embodiments, the polynucleotide comprises two or more 5'UTRs and / or 3'UTRs, which may be the same sequence or different sequences. In some embodiments, the polynucleotide comprises two or more microRNA binding sites, which may be the same sequence or different sequences. Any portion of the 5'UTR, 3'UTR, and / or microRNA binding sites can be sequence-optimized (including none of the portions being optimized), and each portion can independently contain one or more different structural or chemical modifications before and / or after sequence optimization.
[0125] In some embodiments, after optimization, the polynucleotides are reconstituted and introduced into vectors such as plasmids, viruses, cosmids, and artificial chromosomes (but not limited to these). For example, the optimized polynucleotides can be reconstituted and introduced into chemically competent E. coli, yeast, Neurospora, maize, Drosophila, etc., in which high copy number plasmid-like structures or chromosome structures can be produced by the methods described herein.
[0126] Sequence-optimized nucleotide sequence encoding PAH polypeptide In some embodiments, the polynucleotide of the present invention comprises a sequence-optimized nucleotide sequence encoding a variant PAH polypeptide disclosed herein. In some embodiments, the polynucleotide of the present invention comprises an open reading frame (ORF) encoding a PAH polypeptide, wherein the ORF is sequence-optimized.
[0127] Exemplary sequence-optimized nucleotide sequences encoding human variant PAHs are defined as SEQ ID NOs: 22, 23, 24, 25, 26, 27, 28, 29, 30, or 31. In some embodiments, the methods disclosed herein are carried out using sequence-optimized PAH sequences, fragments thereof, and variants thereof.
[0128] In some embodiments, the polynucleotide of the present disclosure, for example, a polynucleotide comprising an mRNA nucleotide sequence encoding a PAH polypeptide, is expressed from the 5' end to the 3' end. (i) A 5' cap as shown herein, for example, Cap 1 and (ii) Sequences shown herein, for example, a 5'UTR such as sequence number 55, (iii) an open reading frame encoding a variant PAH polypeptide, for example, a sequence-optimized nucleic acid sequence encoding PAH (defined as SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30 or SEQ ID NO: 31), (iv) at least one stop codon (if not located at the 5' end of the 3'UTR) (v) Sequences shown herein, for example, a 3'UTR such as sequence number 107 or 108, (vi) The poly-A tail shown above, Includes.
[0129] In some embodiments, the polynucleotide of the present disclosure, for example, a polynucleotide comprising an mRNA nucleotide sequence encoding a PAH polypeptide, is expressed from the 5' end to the 3' end. (i) A 5' cap as shown herein, for example, Cap 1 and (ii) Sequences shown herein, for example, a 5'UTR such as SEQ ID NO: 55 or SEQ ID NO: 56, (iii) an open reading frame encoding a variant PAH polypeptide, for example, a sequence-optimized nucleic acid sequence encoding PAH (defined as SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30 or SEQ ID NO: 31), (iv) at least one stop codon (if not located at the 5' end of the 3'UTR) (v) Sequences shown herein, for example, 3'UTR such as SEQ ID NO: 113, SEQ ID NO: 107, SEQ ID NO: 108, SEQ ID NO: 112, SEQ ID NO: 114, SEQ ID NO: 109, SEQ ID NO: 110 or SEQ ID NO: 111, (vi) The poly-A tail shown above, Includes.
[0130] In certain embodiments, all uracils in the polynucleotide are N1-methylpseudracil (G5). In certain embodiments, all uracils in the polynucleotide are 5-methoxyuracil (G6).
[0131] The sequence-optimized nucleotide sequences disclosed herein differ from the corresponding wild-type nucleotide acid sequences and other known sequence-optimized nucleotide sequences, and for example, these sequence-optimized nucleic acids have unique compositional characteristics.
[0132] In some embodiments, the percentage of uracil or thymine nucleotides in a sequence-optimized nucleotide sequence (e.g., a sequence encoding a PAH polypeptide, a functional fragment thereof, or a variant thereof) is modified (e.g., reduced) compared to the percentage of uracil or thymine nucleotides in a reference wild-type nucleotide sequence. Such sequences are referred to as uracil-modified sequences or thymine-modified sequences. The percentage of uracil or thymine content in a nucleotide sequence can be determined by dividing the number of uracil or thymine in the sequence by the total number of nucleotides and multiplying by 100. In some embodiments, the sequence-optimized nucleotide sequence has a lower uracil or thymine content than the reference wild-type sequence. In some embodiments, the uracil or thymine content in the sequence-optimized nucleotide sequence of the present invention is higher than the uracil or thymine content in the reference wild-type sequence, while retaining beneficial effects, such as increased expression and / or reduced Toll-like receptor (TLR) response, compared to the reference wild-type sequence.
[0133] Methods for optimizing codon usage frequency are known in the art. For example, one or more ORFs among the sequences shown herein may be codon-optimized. In some embodiments, codon optimization may be used to ensure proper folding by matching codon frequencies in the target organism and the host organism, to improve mRNA stability or reduce secondary structures by biasing the GC content, to minimize tandem repeat codon or base sequences that may impair gene construction or expression, to customize transcription and translation regulatory regions, to insert or remove protein transport sequences, to remove / add post-translational modification sites (e.g., glycosylation sites) in the encoded protein, to add, remove or shuffle protein domains, to insert or delete restriction sites, to modify ribosome binding sites and mRNA degradation sites, to adjust the translation rate so that various domains of the protein fold correctly, or to reduce or eliminate problematic secondary structures within a polynucleotide. Tools, algorithms, and services for codon optimization are known in the art, and non-limiting examples include services from GeneArt (Life Technologies), DNA2.0 (Menlo Park CA), and / or proprietary methods. In some embodiments, the open reading frame (ORF) sequence is optimized using an optimization algorithm.
[0134] Modified nucleotide sequences encoding PAH polypeptides In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises a chemically modified nucleic acid base, e.g., chemically modified uracil, e.g., pseudouracil, N1-methylpseudracil, 5-methoxyuracil, etc. In some embodiments, the mRNA is a uracil-modified sequence containing an ORF encoding a PAH polypeptide, and the mRNA comprises a chemically modified nucleic acid base, e.g., chemically modified uracil, e.g., pseudouracil, N1-methylpseudracil, or 5-methoxyuracil.
[0135] In certain embodiments of the present invention, when the modified uracil base is bound to a ribose sugar, as in a polynucleotide, the resulting modified nucleoside or modified nucleotide is called a modified uridine. In some embodiments, the uracil in the polynucleotide is modified uracil by at least about 25%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least 90%, at least 95%, at least 99%, or about 100%. In one embodiment, the uracil in the polynucleotide is modified uracil by at least 95%. In another embodiment, the uracil in the polynucleotide is modified uracil by 100%.
[0136] In embodiments where at least 95% of the uracil in the polynucleotide is modified uracil, the overall uracil content can be adjusted so that the mRNA yields an appropriate level of protein expression without inducing a significant or no immune response. In some embodiments, the uracil content of the ORF is set to the theoretical minimum uracil content (%U) of the corresponding wild-type ORF. TM ) is approximately 100% to 150%, 100% to 110%, 105% to 115%, 110% to 120%, 115% to 125%, 120% to 130%, 125% to 135%, 130% to 140%, 135% to 145%, and 140% to 150%. In another embodiment, the uracil content of the ORF is %U TM It is approximately 121% to 136% or 123% to 134%. In some embodiments, the uracil content of the ORF encoding the PAH polypeptide is %U TM These are approximately 115%, 120%, 125%, 130%, 135%, 140%, 145%, or 150% of the total. In this context, the term "uracil" can refer to modified uracil and / or natural uracil.
[0137] In some embodiments, the uracil content in the ORF of the mRNA encoding the variant PAH polypeptide of the present invention is less than about 30%, less than about 25%, less than about 20%, less than about 15%, or less than 10% of the total nucleic acid base content in the ORF. In some embodiments, the uracil content in the ORF is about 10% to about 20% of the total nucleic acid base content in the ORF. In another embodiment, the uracil content in the ORF is about 10% to about 25% of the total nucleic acid base content in the ORF. In one embodiment, the uracil content in the ORF of the mRNA encoding the PAH polypeptide is less than about 20% of the total nucleic acid base content in the open reading frame. In this context, the term “uracil” may refer to modified uracil and / or natural uracil.
[0138] In further embodiments, ORFs of mRNA encoding a variant PAH polypeptide, having modified uracil and having an adjusted uracil content, exhibit increased cytosine (C), guanine (G), or guanine / cytosine (G / C) content (absolute or relative). In some embodiments, the overall increase in the C, G, or G / C content (absolute or relative) of the ORF is at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 10%, at least about 15%, at least about 20%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or at least about 100% compared to the G / C content (absolute or relative) of the wild-type ORF. In some embodiments, the G content, C content, or G / C content in the ORF is the theoretical value (%G) of the maximum G content, maximum C content, or maximum G / C content of the corresponding wild-type nucleotide sequence encoding the PAH polypeptide. TMX , %C TMX or %G / C TMX) is less than approximately 100%, less than approximately 90%, less than approximately 85%, or less than approximately 80%. In some embodiments, the G content and / or C content (absolute or relative) described herein can be increased by replacing synonymous codons with low G content, C content, or G / C content with synonymous codons with higher G content, C content, or G / C content. In other embodiments, the G content and / or C content (absolute or relative) can be increased by replacing codons ending in U with synonymous codons ending in G or C.
[0139] In further embodiments, the ORF of the mRNA encoding the PAH polypeptide of the present invention contains modified uracil, has a controlled uracil content, and contains fewer uracil pairs (UU) and / or uracil triplets (UUU) and / or urasil adlaplets (UUUU) than the corresponding wild-type nucleotide sequence encoding the PAH polypeptide. In some embodiments, the ORF of the mRNA encoding the PAH polypeptide of the present invention does not contain uracil pairs and / or uracil triplets and / or urasil adlaplets. In some embodiments, the number of uracil pairs and / or uracil triplets and / or uracil adlaplets is reduced below a certain threshold, for example, in the ORF of the mRNA encoding the PAH polypeptide of the present invention, being 1 or less, 2 or less, 3 or less, 4 or less, 5 or less, 6 or less, 7 or less, 8 or less, 9 or less, 10 or less, 11 or less, 12 or less, 13 or less, 14 or less, 15 or less, 16 or less, 17 or less, 18 or less, 19 or less, or 20 or less. In certain embodiments, the ORF of the mRNA encoding the PAH polypeptide of the present invention contains fewer than 20, fewer than 19, fewer than 18, fewer than 17, fewer than 16, fewer than 15, fewer than 14, fewer than 13, fewer than 12, fewer than 11, fewer than 10, fewer than 9, fewer than 8, fewer than 7, fewer than 6, fewer than 5, fewer than 4, fewer than 3, fewer than 2, or fewer than 1 uracil pair and / or uracil triplet other than phenylalanine. In other embodiments, the ORF of the mRNA encoding the PAH polypeptide of the present invention does not contain uracil pairs and / or uracil triplets other than phenylalanine.
[0140] In further embodiments, the ORF of the mRNA encoding the variant PAH polypeptide of the present invention contains modified uracil, with a controlled uracil content, and the uracil-rich clusters are fewer than those in the corresponding wild-type nucleotide sequence encoding the PAH polypeptide. In some embodiments, the ORF of the mRNA encoding the PAH polypeptide of the present invention contains uracil-rich clusters that are shorter in length than the corresponding uracil-rich clusters in the corresponding wild-type nucleotide sequence encoding the PAH polypeptide.
[0141] In further embodiments, less frequent alternative codons are used. At least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or 100% of the codons in the PAH polypeptide coding ORF of the mRNA containing modified uracil are substituted with alternative codons, each of which has a codon frequency lower than that of its substitute codon in the synonymous codon set. The ORF is also adjusted in terms of uracil content as described above. In some embodiments, at least one codon in the ORF of the mRNA encoding the PAH polypeptide of the present invention is substituted with an alternative codon whose codon frequency is lower than that of the substituted codon in the synonymous codon set.
[0142] Methods for modifying polynucleotides This disclosure includes modified polynucleotides, including polynucleotides described herein (e.g., polynucleotides containing a nucleotide sequence encoding a variant PAH polypeptide, e.g., mRNA). These modified polynucleotides may be chemically modified and / or structurally modified. If a polynucleotide of the present invention is chemically modified and / or structurally modified, it may be referred to as a “modified polynucleotide.”
[0143] This disclosure provides modified nucleosides and modified nucleotides of polynucleotides encoding variant PAH polypeptides (e.g., RNA polynucleotides such as mRNA polynucleotides). “Nucleoside” means a compound comprising a sugar molecule (e.g., pentose or ribose) or a derivative thereof in combination with an organic base (e.g., purine or pyrimidine) or a derivative thereof (also referred to herein as “nucleic acid base”). “Nucleotide” means a nucleoside containing a phosphate group. Modified nucleotides can be synthesized by any useful method (e.g., chemical, enzymatic, or recombinant) to include one or more modified nucleosides or unnatural nucleosides. Polynucleotides may contain one or more regions consisting of linked nucleosides. Such regions may have a variety of backbone links. These links may be standard phosphodiester links, in which case the polynucleotide will contain regions consisting of nucleotides.
[0144] The modified polynucleotides disclosed herein may include a variety of distinct modifications. In some embodiments, the modified polynucleotide comprises one or more (optionally different) nucleoside modifications or nucleotide modifications. In some embodiments, the modified polynucleotide introduced into cells may exhibit one or more desired properties, such as improved protein expression, reduced immunogenicity, or reduced intracellular degradation compared to the unmodified polynucleotide.
[0145] In some embodiments, the polynucleotides of the present invention (e.g., polynucleotides comprising a nucleotide sequence encoding a variant PAH polypeptide) are structurally modified. As used herein, “structural” modification refers to a modification in a polynucleotide that involves the insertion, deletion, replication, inversion, or randomization of two or more linked nucleosides without significant chemical modification of the nucleotide itself. Since structural modification always involves the breaking and reformation of chemical bonds, structural modification is a modification of chemical properties, and therefore, structural modification is a chemical modification. However, structural modification results in different nucleotide sequences. For example, the polynucleotide “ATCG” can be chemically modified to “AT-5meC-G”. The same polynucleotide can be structurally modified from “ATCG” to “ATCCCG”. That is, the polynucleotide is structurally modified as a result of inserting the dinucleotide “CC”.
[0146] In some embodiments, therapeutic compositions of the present disclosure include at least one nucleic acid (e.g., RNA) having an open reading frame encoding a variant PAH (e.g., SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, or SEQ ID NO: 31), wherein the nucleic acid includes a nucleotide and / or nucleoside that may be standard (unmodified) or modified, as known in the art. In some embodiments, the nucleotides and nucleosides of the present disclosure include modified nucleotides or modified nucleosides. Such modified nucleotides and modified nucleosides may be naturally occurring modified nucleotides and modified nucleosides or unnaturally occurring modified nucleotides and modified nucleosides. Such modifications may include modifications to the sugar portion, main chain portion or nucleic acid base portion of the nucleotide and / or nucleoside, as recognized in the art.
[0147] In some embodiments, the naturally occurring modified nucleotides or modified nucleotides of this disclosure are those that are generally known or recognized in the art. Such naturally occurring modified nucleotides and non-limiting examples of modified nucleotides can be found, in particular, in the widely recognized MODOMICS database.
[0148] In some embodiments, the non-natural modified nucleotides or nucleosides of this disclosure are those that are generally known or recognized in the art. Non-limiting examples of such non-natural modified nucleotides and modified nucleosides can be found, in particular, in U.S. applications PCT / US2012 / 058519, PCT / US2013 / 075177, PCT / US2014 / 058897, PCT / US2014 / 058891, PCT / US2014 / 070413, PCT / US2015 / 36773, PCT / US2015 / 36759, PCT / US2015 / 36771, or PCT / IB2017 / 051367, all of which are incorporated herein by reference.
[0149] In some embodiments, at least one RNA (e.g., mRNA) of the Disclosure is unmodified and comprises a standard ribonucleotide consisting of adenosine, guanosine, cytosine, and uridine. In some embodiments, the nucleotides and nucleosides of the Disclosure comprise standard nucleoside residues (e.g., A, G, C, or U) as present in transcribed RNA. In some embodiments, the nucleotides and nucleosides of the Disclosure comprise standard deoxyribonucleosides (e.g., dA, dG, dC, or dT) as present in DNA.
[0150] In other words, the nucleic acids of this disclosure (e.g., DNA nucleic acids and RNA nucleic acids such as mRNA nucleic acids) may include standard nucleotides and nucleosides, natural nucleotides and nucleosides, unnatural nucleotides and nucleosides, or any combination thereof.
[0151] The nucleic acids of this disclosure (e.g., DNA nucleic acids and RNA nucleic acids such as mRNA nucleic acids) include, in some embodiments, a variety of (two or more) different types of standard nucleotides and nucleosides, and / or modified nucleotides and modified nucleosides. In some embodiments, a particular region of the nucleic acid includes one or more (arbitrarily different) types of standard nucleotides and nucleosides, and / or modified nucleotides and modified nucleosides.
[0152] In some embodiments, modified RNA nucleic acids (e.g., modified mRNA nucleic acids) introduced into cells or organisms exhibit reduced degradation within the cell or organism compared to unmodified nucleic acids containing standard nucleotides and nucleosides.
[0153] In some embodiments, the modified RNA nucleic acid (e.g., modified mRNA nucleic acid) introduced into a cell or organism may have reduced immunogenicity in the cell or organism compared to unmodified nucleic acids containing standard nucleotides and nucleosides (e.g., reduced spontaneous response).
[0154] Nucleic acids (e.g., RNA nucleic acids such as mRNA nucleic acids) include, in some embodiments, unnaturally modified nucleotides introduced during or after the synthesis of the nucleic acid to achieve a desired function or property. The modifications may be present in internucleotide bonds, purine or pyrimidine bases, or sugars. The modifications may be introduced at the ends of the chain or at other locations within the chain by chemical synthesis or polymerase enzymes. Any region of the nucleic acid may be chemically modified.
[0155] This disclosure provides modified nucleosides and modified nucleotides of nucleic acids (e.g., RNA nucleic acids such as mRNA nucleic acids). “Nucleoside” means a compound comprising a sugar molecule (e.g., pentose or ribose) or a derivative thereof in combination with an organic base (e.g., purine or pyrimidine) or a derivative thereof (also referred to herein as “nucleic acid base”). “Nucleotide” means a nucleoside containing a phosphate group. Modified nucleotides may be synthesized by any useful method (e.g., chemical, enzymatic, or recombinant) to include one or more modified nucleosides or unnatural nucleosides. Nucleic acids may contain one or more regions consisting of linked nucleosides. Such regions may have a variety of backbone links. These links may be standard phosphodiester links, in which case the nucleic acid will contain regions consisting of nucleotides.
[0156] Base pairing of modified nucleotides includes not only standard adenosine-thymine, adenosine-uracil, or guanosine-cytosine base pairs, but also base pairs formed between nucleotides and / or modified nucleotides containing non-standard or modified bases, where hydrogen bonding is possible between non-standard and standard bases, or between two complementary non-standard base structures (e.g., in nucleic acids having at least one chemical modification) through the arrangement of hydrogen bond donors and hydrogen bond acceptors. An example of such non-standard base pairing is the base pairing of inosine of a modified nucleotide with adenine, cytosine, or uracil. Combinations of bases / sugars or linkers may also be introduced into the nucleic acids of this disclosure.
[0157] In some embodiments, the modified nucleic acid bases in nucleic acids (e.g., RNA nucleic acids such as mRNA nucleic acid) include N1-methyl-pseudridine (m1ψ), 1-ethyl-pseudridine (e1ψ), 5-methoxy-uridine (mo5U), 5-methyl-cytidine (m5C), and / or pseudouridine (ψ). In some embodiments, the modified nucleic acid bases in nucleic acids (e.g., RNA nucleic acids such as mRNA nucleic acid) include 5-methoxymethyluridine, 5-methylthiouridine, 1-methoxymethylpseudridine, 5-methylcytidine, and / or 5-methoxycytidine. In some embodiments, the polyribonucleotides of the present invention include chemically modified (but not limited to) combinations of at least two (e.g., two, three, or four or more) of the above modified nucleic acid bases.
[0158] In some embodiments, the RNA nucleic acids of the present disclosure include substitution of one or more or all of the uridine positions of the nucleic acid with N1-methylpseudridine (m1ψ).
[0159] In some embodiments, the RNA nucleic acids of the present disclosure include substitutions to N1-methylpseudridine (m1ψ) at one or more uridine positions of the nucleic acid and substitutions to 5-methylcytidine at one or more cytidine positions of the nucleic acid.
[0160] In some embodiments, the RNA nucleic acids of the present disclosure include substitutions of pseudouridine (ψ) at one or more uridine positions of the nucleic acid.
[0161] In some embodiments, the RNA nucleic acids of the present disclosure include substitutions to pseudouridine (ψ) at one or more uridine positions of the nucleic acid and substitutions to 5-methylcytidine at one or more cytidine positions of the nucleic acid.
[0162] In some embodiments, the RNA nucleic acids of the present disclosure contain uridine at one or more or all of the uridine positions of the nucleic acid.
[0163] In some embodiments, nucleic acids (e.g., RNA nucleic acids such as mRNA nucleic acids) are uniformly modified with certain modifications (e.g., completely modified, modified throughout the entire sequence). For example, a nucleic acid can be uniformly modified with N1-methyl-pseudridine, meaning that all uridine residues in the mRNA sequence are replaced with N1-methyl-pseudridine. Similarly, a nucleic acid can be uniformly modified by replacing any type of nucleoside residue present in its sequence with a modified residue such as those shown above.
[0164] The nucleic acids of the Disclosure may be partially or completely modified along the entire length of the molecule. For example, in the nucleic acids of the Disclosure or in a given sequence region thereof (e.g., mRNA with or without a poly-A tail), one or more or all of the nucleotides, or a given type (e.g., purines or pyrimidines, or one or more or all of A, G, U, and C), may be uniformly modified. In some embodiments, all nucleotides X in the nucleic acids of the Disclosure (or their sequence region) are modified nucleotides, where X may be one of the nucleotides A, G, U, or C, or one of the combinations A+G, A+U, A+C, G+U, G+C, U+C, A+G+U, A+G+C, G+U+C, or A+G+C.
[0165] The nucleic acid contains modified nucleotides in a proportion of approximately 1% to approximately 100%, or any small range thereof (for example, 1% to 20%, 1% to 25%, 1% to 50%, 1% to 60%, 1% to 70%, 1% to 80%, 1% to 90%, 1% to 95%, 10% to 20%, 10% to 25%, 10% to 50%, 10% to 60%, 10% to 70%, 10% to 80%, 10% to 90%). This may include 10%-95%, 10%-100%, 20%-25%, 20%-50%, 20%-60%, 20%-70%, 20%-80%, 20%-90%, 20%-95%, 20%-100%, 50%-60%, 50%-70%, 50%-80%, 50%-90%, 50%-95%, 50%-100%, 70%-80%, 70%-90%, 70%-95%, 70%-100%, 80%-90%, 80%-95%, 80%-100%, 90%-95%, 90%-100%, and 95%-100%. It will be clear that the remaining percentages will all contain unmodified A, G, U, or C.
[0166] The nucleic acid may contain modified nucleotides in a proportion of at least 1%, up to 100%, or any smaller range thereof (e.g., at least 5%, at least 10%, at least 25%, at least 50%, at least 80%, or at least 90%). For example, the nucleic acid may contain modified pyrimidines such as modified uracil or modified cytosine. In some embodiments, at least 5%, at least 10%, at least 25%, at least 50%, at least 80%, at least 90%, or 100% of the uracil in the nucleic acid is replaced with modified uracil (e.g., 5-substituted uracil). The modified uracil may be replaced by a compound having one unique structure, or by multiple compounds having different structures (e.g., two, three, or four or more unique structures). In some embodiments, at least 5%, at least 10%, at least 25%, at least 50%, at least 80%, at least 90%, or 100% of the cytosine in the nucleic acid is replaced with modified cytosine (e.g., 5-substituted cytosine). The modified cytosine may be replaced by a compound having one unique structure, or by multiple compounds having different structures (e.g., two, three, or four or more unique structures).
[0167] Untranslated area (UTR) The untranslated region (UTR) is the portion of the nucleic acid portion of a polynucleotide that is not translated, consisting of the portion before the start codon (5'UTR) and the portion after the stop codon (3'UTR). In some embodiments, the polynucleotide of the present invention (e.g., ribonucleic acid (RNA), e.g., messenger RNA (mRNA)) comprising an open reading frame (ORF) encoding a variant PAH polypeptide further comprises the UTR (e.g., the 5'UTR or a functional fragment thereof, the 3'UTR or a functional fragment thereof, or a combination thereof).
[0168] The UTR (e.g., 5'UTR or 3'UTR) may be of the same or different type as the coding region within the polynucleotide. In some embodiments, the UTR is of the same type as the ORF encoding the variant PAH polypeptide of the present invention. In some embodiments, the UTR is different from the ORF encoding the variant PAH polypeptide of the present invention.
[0169] In some embodiments, the polynucleotide of the present invention comprises two or more 5'UTRs or functional fragments thereof, each having the same or different nucleotide sequences. In some embodiments, the polynucleotide of the present invention comprises two or more 3'UTRs or functional fragments thereof, each having the same or different nucleotide sequences.
[0170] In some embodiments, the 5'UTR or its functional fragment, the 3'UTR or its functional fragment, or a combination thereof, has an optimized sequence.
[0171] In some embodiments, the 5'UTR or a functional fragment thereof, the 3'UTR or a functional fragment thereof, or a combination thereof, contains at least one chemically modified nucleic acid base, such as N1-methylpseudracil or 5-methoxyuracil.
[0172] UTRs may have regulatory properties, such as properties that increase or decrease stability, localization, and / or translation efficiency. Polynucleotides containing UTRs can be administered to cells, tissues, or organisms, and one or more regulatory properties can be measured using conventional methods. In some embodiments, a functional fragment of the 5'UTR contains one or more regulatory properties of the full-length 5'UTR, and a functional fragment of the 3'UTR contains one or more regulatory properties of the full-length 3'UTR.
[0173] The natural 5'UTR has characteristics that play a role in translation initiation. The natural 5'UTR, like the Kozak sequence, has a signature that is widely known to be involved in the process by which ribosomes initiate translation of many genes. The Kozak sequence has a consensus CCR(A / G)CCAUGG (SEQ ID NO: 214), where the R is a purine (adenine or guanine) three bases upstream from the start codon (AUG), followed by another "G" after the start codon. The 5'UTR is also known to form secondary structures involved in the binding of elongation factors.
[0174] By manipulating the characteristics typically found in abundantly expressed genes in a given target organ, the stability of polynucleotides and protein production can be enhanced. For example, polynucleotide expression can be enhanced in hepatocyte lines or in the liver by introducing the 5'UTR of mRNA expressed in the liver, such as albumin, serum amyloid A, apolipoproteins A / B / E, transferrin, alpha-fetoprotein, erythropoietin, or factor VIII. Similarly, improving the expression of a particular mRNA within a tissue using a 5'UTR derived from another tissue-specific mRNA is possible in muscle cells (e.g., MyoD, myosin, myoglobin, myogenin, herculin), endothelial cells (e.g., Tie-1, CD36), myeloid cells (e.g., C / EBP, AML1, G-CSF, GM-CSF, CD11b, MSR, Fr-1, i-NOS), leukocytes (e.g., CD45, CD18), adipose tissue (e.g., CD36, GLUT4, ACRP30, adiponectin), and lung epithelial cells (e.g., SP-A / B / C / D).
[0175] In some embodiments, UTRs are proteins selected from families of transcripts that share a common function, structure, feature, or property. For example, encoded polypeptides may belong to a family of proteins (i.e., a family that shares at least one function, structure, feature, localization, origin, or expression pattern) that are expressed in a particular cell, tissue, or at a specific point in development. Novel polynucleotides can be created by replacing a UTR derived from either the gene or mRNA with any other UTR from the same or a different family of proteins.
[0176] In some embodiments, the 5'UTR and 3'UTR can be of different species. In some embodiments, the 5'UTR can originate from a different species than the 3'UTR. In some embodiments, the 3'UTR can originate from a different species than the 5'UTR.
[0177] Shared international patent application PCT / US2014 / 021522 (publication number WO / 2014 / 164253, which is incorporated herein by reference in its entirety) provides a list of exemplary UTRs that can be used as flanking regions for ORFs in the polynucleotides of the present invention.
[0178] Further exemplary UTRs of this application include globins such as α-globin or β-globin (e.g., claw frog globin, mouse globin, rabbit globin, or human globin), potent Kozak translation initiation signals, CYBA (e.g., human cytochrome b-245α polypeptide), albumin (e.g., human albumin 7), HSD17B4 (hydroxysteroid (17-β) dehydrogenase), viruses (e.g., tobacco etch virus (TEV), Venezuelan encephalitis virus (VEEV), dengue virus, cytomegalovirus (CMV) (e.g., CMV pre-initial 1 (IE1)), hepatitis viruses (e.g., hepatitis B virus), Sindbisvirus, or PA (Barley yellow wilt virus), heat shock proteins (e.g., hsp70), translation initiation factors (e.g., elF4G), glucose transporters (e.g., hGLUT1 (human glucose transporter 1)), actin (e.g., human α-actin or human β-actin), GAPDH, tubulin, histones, citric acid cycle enzymes, topoisomerases (e.g., 5'UTR of the TOP gene lacking the 5'TOP motif (oligopyrimidine tract)), ribosomal proteins Large32 (L32), ribosomal proteins (e.g., human ribosomal proteins such as rps9 or mouse ribosomal proteins), ATP synthases (e.g., ATP5A1 or mitochondrial H +-β-subunit of ATP synthase), growth hormone e (e.g., bovine growth hormone (bGH) or human growth hormone (hGH)), elongation factors (e.g., elongation factor 1α1 (EEF1A1)), manganese superoxide dismutase (MnSOD), muscle cell enhancer factor 2A (MEF2A), β-F1-ATPase, creatine kinase, myoglobin, granulocyte colony-stimulating factor (G-CSF), collagen (e.g., type I collagen, type α2 collagen (Col1A2), type I α1 collagen (Col1A1), collagen VI) Examples include, but are not limited to, one or more 5'UTRs and / or 3'UTRs derived from the nucleic acid sequences of type α2 (Col6A2), type VI α1 (Col6A1), ribophorin (e.g., ribophorin I (RPNI)), low-density lipoprotein receptor-associated protein (e.g., LRP1), cardiotrophin-like cytokine factor (e.g., Nnt1), calreticulin (Calr), procollagen-lysine, 2-oxoglutaric acid 5-dioxygenase 1 (Plod1), and nucleobindin (e.g., Nubc1).
[0179] In some embodiments, the 5'UTR is selected from the group consisting of β-globin 5'UTR, 5'UTR containing a potent Kozak translation initiation signal, cytochrome b-245 α-polypeptide (CYBA) 5'UTR, hydroxysteroid (17-β) dehydrogenase (HSD17B4) 5'UTR, tobacco etch virus (TEV) 5'UTR, Venezuelan encephalitis virus (TEEV) 5'UTR, the 5' proximal open reading frame of rubella virus (RV) RNA encoding a non-structural protein, dengue virus (DEN) 5'UTR, heat shock protein 70 (Hsp70) 5'UTR, eIF4G 5'UTR, GLUT1 5'UTR, functional fragments of these, and combinations thereof.
[0180] In some embodiments, the 3'UTR is selected from the group consisting of β-globin 3'UTR, CYBA 3'UTR, albumin 3'UTR, growth hormone (GH) 3'UTR, VEEV 3'UTR, hepatitis B virus (HBV) 3'UTR, α-globin 3'UTR, DEN 3'UTR, PAV barley yellow wilt virus (BYDV-PAV) 3'UTR, elongation factor 1 α1 (EEF1A1) 3'UTR, manganese superoxide dismutase (MnSOD) 3'UTR, 3'UTR of the β-subunit of mitochondrial H(+)-ATP synthase (β-mRNA), GLUT1 3'UTR, MEF2A 3'UTR, β-F1-ATPase 3'UTR, functional fragments thereof, and combinations thereof.
[0181] Any wild-type UTR derived from any gene or mRNA can be incorporated into the polynucleotide of the present invention. In some embodiments, variant UTRs can be created by modifying the wild-type or native UTR, for example, by changing the orientation or position of the UTR relative to the ORF, or by including or deleting additional nucleotides, or by swapping or transferring nucleotides. In some embodiments, variants of 5'UTR or 3'UTR, such as a variant of the wild-type UTR, or variants in which one or more nucleotides are added to or removed from the end of the UTR can be used.
[0182] In addition, one or more synthetic UTRs can be used in combination with one or more non-synthetic UTRs. See, for example, Mandal and Rossi, Nat. Protoc. 2013 8(3):568-82 (the contents of which are incorporated herein by reference in their entirety).
[0183] UTRs or parts thereof can be positioned in the same orientation as the transcript from which they were selected, or their orientation or position can be altered. That is, 5'UTRs and / or 3'UTRs can be reversed, shortened, lengthened, or combined with one or more other 5'UTRs or 3'UTRs.
[0184] In some embodiments, the polynucleotides of the present invention contain multiple UTRs, for example, 5'UTRs or 3'UTRs in a double, triple, or quadruple configuration. For example, a double UTR contains two copies of the same UTR, in series or substantially in series. For example, a double β-globin 3'UTR can be used (see US2010 / 0129877, the contents of which are incorporated herein by reference in their entirety).
[0185] The polynucleotides of the present invention may contain a combination of features. For example, the ORF may be sandwiched between a 5'UTR containing a potent Kozak translation initiation signal and / or a 3'UTR containing an oligo(dT) sequence for the addition of a polyA tail by template. The 5'UTR may contain a first polynucleotide fragment and a second polynucleotide fragment derived from the same UTR and / or different UTRs (see, for example, US2010 / 0293625, which is incorporated herein by reference in its entirety).
[0186] Other non-UTR sequences can be used as regions or sub-regions within the polynucleotide of the present invention. For example, an intron or a portion of an intron sequence can be incorporated into the polynucleotide of the present invention. Incorporation of an intron sequence can improve protein production and polynucleotide expression levels. In some embodiments, the polynucleotide of the present invention includes an internal ribosome entry site (IRES) in place of or in addition to the UTR (see, for example, Yakubov et al., Biochem. Biophys. Res. Commun. 2010 394(1):189-193 (the contents of which are incorporated herein by reference in their entirety)). In some embodiments, the polynucleotide of the present invention includes an IRES in place of a 5'UTR sequence. In some embodiments, the polynucleotide of the present invention includes an ORF and a viral capsid sequence. In some embodiments, the polynucleotide of the present invention includes a synthetic 5'UTR in combination with a non-synthetic 3'UTR.
[0187] In some embodiments, the UTR may also include at least one translation-enhancer polynucleotide, one translation-enhancer element, or multiple translation-enhancers (collectively referred to as "TEEs," which refer to nucleic acid sequences that increase the amount of polypeptide or protein produced from the polynucleotide). In non-limiting examples, the TEEs may be located between the transcription promoter and the start codon. In some embodiments, the 5' UTR includes TEEs.
[0188] In one embodiment, the TEE is a conserved element in the UTR that can promote the translational activity of nucleic acids, such as cap-dependent translation or cap-independent translation (but not limited to these).
[0189] a. 5' UTR sequence The 5'UTR sequence is important for recruiting ribosomes to mRNA and has been reported to play a role in translation (Hinnebusch A, et al., (2016) Science, 352:6292:1413-6).
[0190] Disclosed in the present invention is a polynucleotide, e.g., mRNA, comprising an open reading frame encoding a variant PAH polypeptide (e.g., SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, or SEQ ID NO: 12), wherein the polynucleotide has a 5'UTR that results in an extension of the half-life, increased expression, and / or improved activity of the polypeptide encoded by the polynucleotide or the polynucleotide itself. In some embodiments, the polynucleotide disclosed herein comprises (a) a 5'UTR (e.g., the 5'UTR shown in Table 2, or a variant or fragment thereof), (b) a coding region containing a stop element (e.g., a region as described herein), and (c) a 3'UTR (e.g., a 3'UTR as described herein), wherein the LNP composition comprises the polynucleotide. In some embodiments, the polynucleotide comprises a 5'UTR containing the sequence shown in Table 2, or a variant or fragment thereof (e.g., a functional variant or functional fragment thereof).
[0191] In some embodiments, polynucleotides having the 5'UTR sequence, or variants or fragments thereof, as shown in Table 2, exhibit an extended half-life, for example, an extension of approximately 1.5 to 20 times. In some embodiments, the extension of the half-life is approximately 1.5 times or more, approximately 2 times or more, approximately 3 times or more, approximately 4 times or more, approximately 5 times or more, approximately 6 times or more, approximately 7 times or more, approximately 8 times or more, approximately 9 times or more, approximately 10 times or more, approximately 11 times or more, approximately 12 times or more, approximately 13 times or more, approximately 14 times or more, approximately 15 times or more, approximately 16 times or more, approximately 17 times or more, approximately 18 times or more, approximately 19 times or more, or 20 times or more. In some embodiments, the extension of the half-life is approximately 1.5 times or more. In some embodiments, the extension of the half-life is approximately 2 times or more. In some embodiments, the extension of the half-life is approximately 3 times or more. In some embodiments, the extension of the half-life is approximately 4 times or more. In one embodiment, the half-life is extended by approximately five times or more.
[0192] In some embodiments, a polynucleotide having the 5'UTR sequence shown in Table 2, or a variant or fragment thereof, enhances the level and / or activity, e.g., production, of the polypeptide encoded by the polynucleotide. In some embodiments, the 5'UTR enhances the level and / or activity, e.g., production, of the polypeptide encoded by the polynucleotide by about 1.5 to 20 times. In some embodiments, the enhancement of level and / or activity is about 1.5 times or more, about 2 times or more, about 3 times or more, about 4 times or more, about 5 times or more, about 6 times or more, about 7 times or more, about 8 times or more, about 9 times or more, about 10 times or more, about 11 times or more, about 12 times or more, about 13 times or more, about 14 times or more, about 15 times or more, about 16 times or more, about 17 times or more, about 18 times or more, about 19 times or more, or about 20 times or more. In some embodiments, the enhancement of level and / or activity is about 1.5 times or more. In some embodiments, the enhancement of level and / or activity is about 2 times or more. In some embodiments, the improvement in level and / or activity is about three times or more. In some embodiments, the improvement in level and / or activity is about four times or more. In some embodiments, the improvement in level and / or activity is about five times or more.
[0193] In one embodiment, the improvement is compared to another similar polynucleotide that does not have a 5'UTR, has a different 5'UTR, or does not have the 5'UTR, variant or fragment thereof, as listed in Table 2.
[0194] In one embodiment, the extension of the half-life of the polynucleotide according to the present invention is measured according to an assay for measuring the half-life of the polynucleotide.
[0195] In one embodiment, the level and / or activity of the polypeptide encoded by the polynucleotide of the present invention, for example, an increase in production, is measured according to an assay that measures the level and / or activity of the polypeptide.
[0196] In one embodiment, the 5'UTR includes a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequences shown in Table 2, or the 5'UTR sequences shown in Table 2, or their variants or fragments. In one embodiment, the 5'UTR includes a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 50, SEQ ID NO: 51, SEQ ID NO: 52, SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, SEQ ID NO: 56, SEQ ID NO: 57, or SEQ ID NO: 58.
[0197] In one embodiment, the 5'UTR includes a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 50. In one embodiment, the 5'UTR includes a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 51. In one embodiment, the 5'UTR includes a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 52. In one embodiment, the 5'UTR includes a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 53. In one embodiment, the 5'UTR includes a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 54. In one embodiment, the 5'UTR includes a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 55. In one embodiment, the 5'UTR includes a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 56. In one embodiment, the 5'UTR includes a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 57. In one embodiment, the 5'UTR includes a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 58.
[0198] In one embodiment, the 5'UTR contains the sequence of sequence number 58. In another embodiment, the 5'UTR consists of the sequence of sequence number 58.
[0199] In one embodiment, the 5'UTR sequence shown in Table 2 has a first nucleotide that is A. In another embodiment, the 5'UTR sequence shown in Table 2 has a first nucleotide that is G. [Table 2-1]
Table 2-2
Table 2-3
[0200] In certain embodiments, the 5’UTR comprises a variant of SEQ ID NO: 50. In certain embodiments, the variant of SEQ ID NO: 50 comprises a nucleic acid sequence of Formula A below, GGAAAUCGCAAAA(N2) X (N3) X CU(N4) X (N5) X CGCGUUAGAUUUCUUUUAGUUUUCUN6N7CAACUAGCAAGCUUUUUGUUCUCGCC(N8CC)x(SEQ ID NO: 59) Wherein, (N2) x is uracil, x is an integer from 0 to 5, for example, x is 3 or 4, (N3) x is guanine, x is an integer from 0 to 1, (N4) x is cytosine, x is an integer from 0 to 1, (N5) x is uracil, x is an integer from 0 to 5, for example, x is 2 or 3, N6 is uracil or cytosine, N7 is uracil or guanine, N8 is adenine or guanine, and x is an integer from 0 to 1.
[0201] In certain embodiments, (N2) x is uracil and x is 0. In certain embodiments, (N2) x is uracil and x is 1. In certain embodiments, (N2) x is uracil and x is 2. In certain embodiments, (N2) xx is uracil, and x is 3. In one embodiment, (N2) x is uracil, and x is 4. In one embodiment, (N2) x x is uracil, and x is 5.
[0202] In one embodiment, (N3) x is guanine, and x is 0. In one embodiment, (N3) x x is guanine, and x is 1.
[0203] In one embodiment, (N4) x x is cytosine, and x is 0. In one embodiment, (N4) x x is cytosine, and x is 1.
[0204] In one embodiment, (N5) x x is uracil, and x is 0. In one embodiment, (N5) x x is uracil, and x is 1. In one embodiment, (N5) x is uracil, and x is 2. In one embodiment, (N5) x x is uracil, and x is 3. In one embodiment, (N5) x x is uracil, and x is 4. In one embodiment, (N5) x x is uracil, and x is 5.
[0205] In one embodiment, N6 is uracil. In another embodiment, N6 is cytosine.
[0206] In some embodiments, N7 is uracil. In other embodiments, N7 is guanine.
[0207] In one embodiment, N8 is adenine and x is 0. In another embodiment, N8 is adenine and x is 1.
[0208] In one embodiment, N8 is guanine and x is 0. In another embodiment, N8 is guanine and x is 1.
[0209] In one embodiment, the 5'UTR includes a variant of sequence number 50. In one embodiment, the variant of sequence number 50 includes a sequence that is at least 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to sequence number 50. In one embodiment, the variant of sequence number 50 includes a sequence that is at least 50% identical to sequence number 50. In one embodiment, the variant of sequence number 50 includes a sequence that is at least 60% identical to sequence number 50. In one embodiment, the variant of sequence number 50 includes a sequence that is at least 70% identical to sequence number 50. In one embodiment, the variant of sequence number 50 includes a sequence that is at least 80% identical to sequence number 50. In one embodiment, the variant of sequence number 50 includes a sequence that is at least 90% identical to sequence number 50. In one embodiment, the variant of sequence number 50 includes a sequence that is at least 95% identical to sequence number 50. In one embodiment, a variant of sequence number 50 includes a sequence that is at least 96% identical to sequence number 50. In another embodiment, a variant of sequence number 50 includes a sequence that is at least 97% identical to sequence number 50. In yet another embodiment, a variant of sequence number 50 includes a sequence that is at least 98% identical to sequence number 50. In yet another embodiment, a variant of sequence number 50 includes a sequence that is at least 99% identical to sequence number 50.
[0210] In one embodiment, the variant of SEQ ID NO: 50 has a uridine content of at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or 80%. In one embodiment, the variant of SEQ ID NO: 50 has a uridine content of at least 5%. In one embodiment, the variant of SEQ ID NO: 50 has a uridine content of at least 10%. In one embodiment, the variant of SEQ ID NO: 50 has a uridine content of at least 20%. In one embodiment, the variant of SEQ ID NO: 50 has a uridine content of at least 30%. In one embodiment, the variant of SEQ ID NO: 50 has a uridine content of at least 40%. In one embodiment, the variant of SEQ ID NO: 50 has a uridine content of at least 50%. In one embodiment, the variant of SEQ ID NO: 50 has a uridine content of at least 60%. In one embodiment, the variant of SEQ ID NO: 50 has a uridine content of at least 70%. In one embodiment, the variant of SEQ ID NO: 50 has a uridine content of at least 80%.
[0211] In one embodiment, the variant of sequence number 50 includes at least 2, 3, 4, 5, 6, or 7 consecutive uridines (e.g., a polyuridine tract). In one embodiment, the polyuridine tract in the variant of sequence number 50 includes at least 1-7 consecutive, 2-7 consecutive, 3-7 consecutive, 4-7 consecutive, 5-7 consecutive, 6-7 consecutive, 1-6 consecutive, 1-5 consecutive, 1-4 consecutive, 1-3 consecutive, 1-2 consecutive, 2-6 consecutive, or 3-5 consecutive uridines. In one embodiment, the polyuridine tract in the variant of sequence number 50 includes 4 consecutive uridines. In one embodiment, the polyuridine tract in the variant of sequence number 50 includes 5 consecutive uridines.
[0212] In one embodiment, the variant of SEQ ID NO: 50 contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 polyuridine tracts. In one embodiment, the variant of SEQ ID NO: 50 contains 3 polyuridine tracts. In one embodiment, the variant of SEQ ID NO: 50 contains 4 polyuridine tracts. In one embodiment, the variant of SEQ ID NO: 50 contains 5 polyuridine tracts.
[0213] In one embodiment, one or more polyuridine tracts are adjacent to different polyuridine tracts. In another embodiment, each of the polyuridine tracts, for example, all of them, are adjacent to one another, and for example, all of the polyuridine tracts are contiguous.
[0214] In one embodiment, one or more polyuridine tracts are separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 2, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, or 60 nucleotides. In another embodiment, each of the polyuridine tracts, for example, is separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 2, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, or 60 nucleotides.
[0215] In one embodiment, the first polyuridine tract and the second polyuridine tract are adjacent to each other.
[0216] In one embodiment, thereafter, for example, the 3rd, 4th, 5th, 6th, 7th, 8th, 9th, or 10th polyuridine tract is separated from the 1st polyuridine tract, the 2nd polyuridine tract, or any of the 3rd and subsequent polyuridine tracts by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 2, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, or 60 nucleotides.
[0217] In one embodiment, the first polyuridine tract is separated from subsequent polyuridine tracts, such as the second, third, fourth, fifth, sixth, seventh, eighth, ninth, or tenth polyuridine tracts, by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 2, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, or 60 nucleotides. In one embodiment, one or more of the second and subsequent polyuridine tracts are adjacent to different polyuridine tracts.
[0218] In one embodiment, the 5'UTR contains a Kozak sequence, for example, a nucleotide sequence called GCCRCC (SEQ ID NO: 79), where R is adenine or guanine. In another embodiment, the Kozak sequence is located at the 3' end of the 5'UTR sequence.
[0219] In one embodiment, a polynucleotide (e.g., mRNA) containing an open reading frame encoding a variant PAH polypeptide (e.g., SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, or SEQ ID NO: 12) and a 5'UTR sequence as disclosed herein is formulated as an LNP. In one embodiment, the LNP composition comprises (i) an ionizable lipid, such as an aminolipid, (ii) a sterol or other structural lipid, (iii) a noncationic helper lipid or phospholipid, and (iv) a PEG-lipid.
[0220] In another embodiment, the LNP composition of the present disclosure is used in a method for treating hyperphenylalaninemia such as PKU in a subject.
[0221] In some embodiments, an LNP composition comprising a polynucleotide disclosed herein, for example, a polynucleotide encoding a variant PAH polypeptide as described herein, can be administered together with additional agents as described herein.
[0222] b.3'UTR sequence The 3'UTR sequence has been shown to affect mRNA translation, half-life, and intracellular localization (Mayr C., Cold Spring Harb Persp Biol 2019 Oct1;11(10):a034728).
[0223] Disclosed in the present invention is a polynucleotide, e.g., mRNA, comprising an open reading frame encoding a variant PAH polypeptide (e.g., SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, or SEQ ID NO: 12), wherein the polynucleotide has a 3'UTR that results in an extension of the half-life, increased expression, and / or improved activity of the polypeptide encoded by the polynucleotide or the polynucleotide itself. In some embodiments, the polynucleotide disclosed herein comprises (a) a 5'UTR (e.g., a 5'UTR as described herein), (b) a coding region containing a stop element (e.g., a region as described herein), and (c) a 3'UTR (e.g., a 3'UTR shown in Table 3, or a variant or fragment thereof), wherein the LNP composition comprises the polynucleotide. In some embodiments, the polynucleotide comprises a 3'UTR containing the sequence shown in Table 3, or a variant or fragment thereof.
[0224] In some embodiments, polynucleotides having the 3'UTR sequence, or variants or fragments thereof, as shown in Table 3, exhibit an extended half-life, for example, by approximately 1.5 to 10 times. In some embodiments, the extension of the half-life is approximately 1.5 times or more, approximately 2 times or more, approximately 3 times or more, approximately 4 times or more, approximately 5 times or more, approximately 6 times or more, approximately 7 times or more, approximately 8 times or more, approximately 9 times or more, or 10 times or more. In some embodiments, the extension of the half-life is approximately 1.5 times or more. In some embodiments, the extension of the half-life is approximately 2 times or more. In some embodiments, the extension of the half-life is approximately 3 times or more. In some embodiments, the extension of the half-life is approximately 4 times or more. In some embodiments, the extension of the half-life is approximately 5 times or more. In some embodiments, the extension of the half-life is approximately 6 times or more. In some embodiments, the extension of the half-life is approximately 7 times or more. In some embodiments, the extension of the half-life is approximately 8 times or more. In one embodiment, the half-life is extended by approximately 9 times or more. In another embodiment, the half-life is extended by approximately 10 times or more.
[0225] In one embodiment, a polynucleotide having the 3'UTR sequence shown in Table 3, or a variant or fragment thereof, can be obtained to have a mean half-life score greater than 10.
[0226] In one embodiment, a polynucleotide having the 3'UTR sequence shown in Table 3, or a variant or fragment thereof, enhances the level and / or activity, e.g., production, of the polypeptide encoded by the polynucleotide.
[0227] In one embodiment, the improvement is compared to another similar polynucleotide that does not have a 3'UTR, has a different 3'UTR, or does not have the 3'UTR of Table 3, or its variant or fragment.
[0228] In one embodiment, the polynucleotide of the present invention includes a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the 3'UTR sequence or a fragment thereof shown in Table 3. In one embodiment, the 3'UTR includes a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 100, SEQ ID NO: 101, SEQ ID NO: 102, SEQ ID NO: 103, SEQ ID NO: 104, SEQ ID NO: 105, SEQ ID NO: 106, SEQ ID NO: 107, SEQ ID NO: 108, SEQ ID NO: 109, SEQ ID NO: 110, SEQ ID NO: 111, SEQ ID NO: 112, SEQ ID NO: 113, SEQ ID NO: 114, or SEQ ID NO: 115.
[0229] In one embodiment, the 3'UTR includes the sequence of sequence number 100, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 100. In another embodiment, the 3'UTR includes the sequence of sequence number 101, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 101. In yet another embodiment, the 3'UTR includes the sequence of sequence number 102, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 102. In one embodiment, the 3'UTR includes the sequence of sequence number 103, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 103. In another embodiment, the 3'UTR includes the sequence of sequence number 104, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 104. In yet another embodiment, the 3'UTR includes the sequence of sequence number 105, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 105. In one embodiment, the 3'UTR includes the sequence of sequence number 106, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 106. In one embodiment, the 3'UTR includes the sequence of sequence number 107, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 107. In one embodiment, the 3'UTR includes the sequence of sequence number 108, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 108. In one embodiment, the 3'UTR includes the sequence of sequence number 109, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 109.In one embodiment, the 3'UTR includes the sequence of sequence number 110, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 110. In another embodiment, the 3'UTR includes the sequence of sequence number 111, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 111. In yet another embodiment, the 3'UTR includes the sequence of sequence number 112, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 112. In one embodiment, the 3'UTR includes the sequence of sequence number 113, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 113. In another embodiment, the 3'UTR includes the sequence of sequence number 114, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 114. In yet another embodiment, the 3'UTR includes the sequence of sequence number 115, or a sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to sequence number 115.
[0230] [Table 3-1] [Table 3-2] [Table 3-3]
[0231] In some embodiments, the 3'UTR includes, for example, a microRNA (miRNA) binding site as described herein, which binds to miRs present in human cells. In some embodiments, the 3'UTR includes a miRNA binding site of sequence number 212, sequence number 174, sequence number 152, or a combination thereof. In some embodiments, the 3'UTR includes multiple miRNA binding sites, for example, two, three, four, five, six, seven, or eight miRNA binding sites. In some embodiments, the multiple miRNA binding sites include the same miRNA binding site or different miRNA binding sites.
[0232] miR122 bs is CAAACACCAUUGUCACACUCCA (sequence number 212).
[0233] miR-142-3p bs is UCCAUAAAGUAGGAAACACUACA (Sequence ID 174).
[0234] miR-126 bs is CGCAUUAUUACUCACGGUACGA (sequence number 152).
[0235] In one embodiment, the present invention discloses a polynucleotide encoding a polypeptide, the polynucleotide comprising (a) a 5'UTR, for example, as described herein, (b) a coding region containing a termination element (for example, a region as described herein), and (c) a 3'UTR (for example, a 3'UTR as described herein).
[0236] In one embodiment, an LNP composition comprising a polynucleotide including an open reading frame encoding a variant PAH polypeptide (e.g., SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, or SEQ ID NO: 12) and a 3'UTR as disclosed herein comprises (i) an ionizable lipid, such as an aminolipid, (ii) a sterol or other structural lipid, (iii) a noncationic helper lipid or phospholipid, and (iv) a PEG-lipid.
[0237] In another embodiment, the LNP composition of the present disclosure is used in a method for treating hyperphenylalaninemia such as PKU in a subject.
[0238] In some embodiments, an LNP composition comprising a polynucleotide disclosed herein, for example, a polynucleotide encoding a variant PAH polypeptide as described herein, can be administered together with additional agents as described herein.
[0239] microRNA (miRNA) binding site The polynucleotides of the present invention may include regulatory elements, such as microRNA (miRNA) binding sites, transcription factor binding sites, structured mRNA sequences and / or structured mRNA motifs, artificial binding sites engineered to act as pseudoreceptors for endogenous nucleic acid-binding molecules, and combinations thereof. In some embodiments, a polynucleotide containing such regulatory elements is referred to as a "sensor sequence" polynucleotide.
[0240] In some embodiments, the polynucleotide of the present invention (e.g., ribonucleic acid (RNA), e.g., messenger RNA (mRNA)) comprises an open reading frame (ORF) encoding the polypeptide of interest, and further comprises one or more miRNA binding sites. By including or incorporating miRNA binding sites, the polynucleotide of the present invention can be regulated based on tissue-specific and / or cell-type-specific expression of native miRNAs, and consequently, the polypeptide encoded from that polynucleotide can be regulated.
[0241] The present invention also provides pharmaceutical compositions and formulations comprising any of the above-mentioned polynucleotides. In some embodiments, the composition or formulation further comprises a delivery agent.
[0242] In some embodiments, the composition or formulation may include a polynucleotide (e.g., RNA, e.g., mRNA) that includes a polynucleotide (e.g., ORF) that has a fairly high degree of sequence identity with the nucleic acid sequence encoding the polypeptide, as disclosed herein. In some embodiments, the polynucleotide may further include a miRNA binding site, e.g., a binding miRNA binding site.
[0243] miRNAs, such as natural miRNAs, are 19-25 nucleotide long non-coding RNAs that downregulate gene expression by either binding to a polynucleotide and reducing its stability or inhibiting its translation. The miRNA sequence occupies a "seed" region, i.e., a sequence within the region between positions 2 and 8 of a mature miRNA. The miRNA seed can occupy positions 2-8 or 2-7 of a mature miRNA.
[0244] MicroRNAs are extracted enzymatically from the region of an RNA transcript, folded in half to form a short hairpin structure, often called pre-miRNA (precursor miRNA). Pre-miRNAs typically have a 2-nucleotide overhang at their 3' end, containing a 3' hydroxyl group and a 5' phosphate group. After processing in the nucleus, this precursor mRNA is transported to the cytoplasm, where it is further processed by dicer (RNase III enzyme) to form a mature microRNA of approximately 22 nucleotides. Subsequently, the mature microRNA is incorporated into a ribonuclear particle to form the RNA-induced silencing complex (RISC), which mediates gene silencing. In the nomenclature recognized in the art, the name of the mature miRNA typically indicates the arm of the pre-miRNA from which the mature miRNA originates. "5p" means that the microRNA originates from the 5th prime arm of the pre-miRNA hairpin, and "3p" means that the microRNA originates from the 3rd prime end of the pre-miRNA hairpin. In this specification, a numbered miR may refer to either of two mature microRNAs (e.g., either a 3p microRNA or a 5p microRNA) resulting from opposing arms of the same premiRNA. Unless specifically defined by the designation of 3p or 5p, all miRs referred to herein are intended to include both 3p and 5p arms / sequences.
[0245] As used herein, the term “microRNA (miRNA or miR) binding site” refers to a sequence within a polynucleotide, for example, within DNA or an RNA transcript (including sequences within the 5'UTR and / or 3'UTR), whose complementarity with the entire miRNA or region is sufficient to interact, associate with, or bind to that miRNA. In some embodiments, the polynucleotide of the present invention, comprising an ORF encoding the polypeptide of interest, further comprises one or more miRNA binding sites. In exemplary embodiments, the 5'UTR and / or 3'UTR of the polynucleotide (e.g., ribonucleic acid (RNA), e.g., messenger RNA (mRNA)) comprises one or more miRNA binding sites.
[0246] A miRNA binding site that is sufficiently complementary to miRNA means that the degree of complementarity is sufficient to promote miRNA-mediated regulation of polynucleotides, for example, miRNA-mediated repression or degradation of polynucleotide translations. In exemplary embodiments of the present invention, a miRNA binding site that is sufficiently complementary to miRNA means that the degree of complementarity is sufficient to promote miRNA-mediated degradation of polynucleotides, for example, miRNA-induced cleavage of mRNA mediated by the miRNA-induced RNA-induced silencing complex (RISC). The miRNA binding site may be complementary to, for example, a miRNA sequence of 19-25 nucleotides, a miRNA sequence of 19-23 nucleotides, or a miRNA sequence of 22 nucleotides. The miRNA binding site may be complementary to only a portion of the miRNA, for example, a portion of a full-length natural miRNA sequence that is less than 1 nucleotide, less than 2 nucleotides, less than 3 nucleotides, or less than 4 nucleotides shorter than the natural miRNA sequence. If the desired regulation is mRNA degradation, then sufficient or complete complementarity (e.g., sufficient or complete complementarity to all or a significant portion of the length of the native miRNA) is preferred.
[0247] In some embodiments, the miRNA binding site includes a sequence that is complementary (e.g., partially complementary or fully complementary) to the miRNA seed sequence. In some embodiments, the miRNA binding site includes a sequence that is fully complementary to the miRNA seed sequence. In some embodiments, the miRNA binding site includes a sequence that is complementary (e.g., partially complementary or fully complementary) to the miRNA sequence. In some embodiments, the miRNA binding site includes a sequence that is fully complementary to the miRNA sequence. In another embodiment, the sequence is not fully complementary. In some embodiments, the miRNA binding site is fully complementary to the miRNA sequence, except for substitutions, terminal additions, and / or truncations of one, two, or three nucleotides.
[0248] In some embodiments, the miRNA binding site is the same length as the corresponding miRNA. In other embodiments, the miRNA binding site is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides shorter than the corresponding miRNA at its 5' end, 3' end, or both. In yet another embodiment, the microRNA binding site is 2 nucleotides shorter than the corresponding microRNA at its 5' end, 3' end, or both. A miRNA binding site shorter than the corresponding miRNA can still degrade mRNA into which one or more miRNA binding sites are incorporated, or prevent that mRNA from being translated.
[0249] In some embodiments, the miRNA binding site binds to a corresponding mature miRNA, which is part of an active RISC containing a dicer. In other embodiments, the miRNA binding site binds to the corresponding miRNA in the RISC, thereby degrading the mRNA containing the miRNA binding site or preventing its translation. In some embodiments, the miRNA binding site is sufficiently complementary to the miRNA so that the RISC complex containing the miRNA cleaves the polynucleotide containing the miRNA binding site. In other embodiments, the miRNA binding site has incomplete complementarity so that the RISC complex containing the miRNA destabilizes the polynucleotide containing the miRNA binding site. In other embodiments, the miRNA binding site has incomplete complementarity so that the RISC complex containing the miRNA represses the transcription of the polynucleotide containing the miRNA binding site.
[0250] In some embodiments, the miRNA binding site has one, two, three, four, five, six, seven, eight, nine, ten, eleven, or twelve mismatches with the corresponding miRNA.
[0251] In some embodiments, the miRNA binding site has at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, or at least 21 consecutive nucleotides complementary to at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, or at least 21 consecutive nucleotides of the corresponding miRNA.
[0252] By manipulating and incorporating one or more miRNA binding sites into the polynucleotide of the present invention, the polynucleotide can be targeted for degradation or reduced translation, provided that the relevant miRNA is available. This reduces off-target effects during polynucleotide delivery. For example, if the polynucleotide of the present invention is not intended to be delivered to a certain tissue or cell, but is delivered to such tissue or cell, manipulating and incorporating one or more miRNA binding sites, which are abundant in that tissue or cell, into the 5'UTR and / or 3'UTR of the polynucleotide allows the miRNA to inhibit the expression of the gene in question. That is, in some embodiments, incorporating one or more miRNA binding sites into the mRNA of the present disclosure reduces the risk of off-target effects during delivery of nucleic acid molecules and / or allows for tissue-specific regulation of the expression of the polypeptide encoded by that mRNA. In yet another embodiment, incorporating one or more miRNA binding sites into the mRNA of the present disclosure modulates the immune response during in vivo delivery of nucleic acids. In further embodiments, the rapid clearance (ABC) of the lipid-containing compounds and lipid-containing compositions described herein from the bloodstream can be regulated by incorporating one or more miRNA binding sites into the mRNA of the disclosed herein.
[0253] Conversely, in order to increase protein expression in a given tissue, miRNA binding sites can be removed from polynucleotide sequences that naturally contain miRNA binding sites. For example, by removing the binding site for a specific miRNA from a polynucleotide, protein expression can be improved in tissues or cells containing that miRNA.
[0254] Regulation of expression in multiple tissues can be achieved through the introduction or removal of one or more miRNA binding sites, for example, one or more distinct miRNA binding sites. The decision to remove or insert a miRNA binding site can be based on the pattern and / or profiling of miRNA expression in tissues and / or cells during development and / or disease. The identification of miRNAs, miRNA binding sites, and their expression patterns and biological roles has been reported (e.g., Bonauer et al., Curr Drug Targets 2010 11:943-949, Anand and Cheresh Curr Opin Hematol 2011 18:171-176, Contreras and Rao Leukemia 2012 26:404-413 (2011 Dec 20.doi:10.1038 / leu.2011.356), Bartel Cell 2009 136:215-233, Landgraf et al, Cell, 2007 129:1401-1414, Gentner and Naldini, Tissue Antigens. 2012 80:393-403 and all of the references therein (each of which is incorporated herein by reference)).
[0255] Tissues in which miRNAs are known to regulate protein expression by regulating mRNA include, but are not limited to, the liver (miR-122), muscle (miR-133, miR-206, miR-208), endothelial cells (miR-17-92, miR-126), myeloid cells (miR-142-3p, miR-142-5p, miR-16, miR-21, miR-223, miR-24, miR-27), adipose tissue (let-7, miR-30c), heart (miR-1d, miR-149), kidney (miR-192, miR-194, miR-204), and lung epithelial cells (let-7, miR-133, miR-126).
[0256] Specifically, miRNAs are known to be differentially expressed in immune cells (also called hematopoietic cells) such as antigen-presenting cells (APCs) (e.g., dendritic cells and macrophages), macrophages, monocytes, B lymphocytes, T lymphocytes, granulocytes, and natural killer cells. Immune cell-specific miRNAs are involved in immune responses to immunogenicity, autoimmunity, infection, and inflammation, as well as undesirable immune responses after gene therapy and tissue / organ transplantation. Immune cell-specific miRNAs also regulate many aspects of hematopoietic cell (immune cell) development, proliferation, differentiation, and apoptosis. For example, miR-142 and miR-146 are expressed only in immune cells, and are particularly abundant in bone marrow dendritic cells. It has been shown that immune responses to polynucleotides can be blocked by adding a miR-142 binding site to the 3'UTR of the polynucleotide, thereby enabling more stable gene transfer within tissues and cells. miR-142 efficiently degrades exogenous polynucleotides in antigen-presenting cells, suppressing the cytotoxic elimination of transductioned cells (e.g., Annoni A et al., blood, 2009, 114, 5152-5161; Brown BD, et al., Nat med. 2006, 12(5), 585-591; Brown BD, et al., blood, 2007, 110(13): 4144-4152 (each of which is incorporated herein by reference)).
[0257] An antigen-mediated immune response refers to an immune response induced by a foreign antigen. When a foreign antigen invades an organism, it is processed by antigen-presenting cells and presented on the surface of those cells. T cells can recognize the presented antigen and induce cytotoxicity to eliminate the cells expressing that antigen.
[0258] By introducing one or more miR-142 binding sites (e.g., one, two, or three) into the 5'UTR and / or 3'UTR of the polynucleotide of the present invention, gene expression in antigen-presenting cells can be selectively suppressed through miR-142-mediated degradation, thereby limiting antigen presentation in antigen-presenting cells (e.g., dendritic cells), and thus inhibiting the antigen-mediated immune response after delivery of the polynucleotide of the present invention. Furthermore, the polynucleotide can be stably expressed in target tissues or target cells without inducing cytotoxic elimination.
[0259] In some embodiments, when both the 3p arm and the 5p arm are abundant (for example, when both miR-142-3p and miR142-5p are abundant in hematopoietic stem cells), it may be beneficial to target the same type of cell with multiple miRs and to incorporate binding sites for the 3p arm and the 5p arm, respectively. That is, in certain embodiments, the polynucleotide of the present invention comprises two or more miR binding sites (e.g., two, three, or four or more) from (i) the group consisting of miR-142, miR-144, miR-150, miR-155, and miR-223 (expressed in many hematopoietic cells), or (ii) the group consisting of miR-142, miR-150, miR-16, and miR-223 (expressed in B cells), or the group consisting of miR-223, miR-451, miR-26a, and miR-16 (expressed in hematopoietic progenitor cells).
[0260] In some embodiments, it may be beneficial to combine various miRs to simultaneously target multiple types of target cells (e.g., miR-142 and miR-126 targeting many hematopoietic cells and endothelial cells). That is, for example, in a particular embodiment, the polynucleotide of the present invention comprises two or more miRNA binding sites (e.g., two, three or four or more), and (i) at least one of its miRs targets hematopoietic cells (e.g., miR-142, miR-144, miR-150, miR-155 or miR-223), and at least one of its miRs targets plasmacytoid dendritic cells, platelets or endothelial cells (e.g., (i) at least one of the miRs targets B cells (e.g., miR-142, miR-150, miR-16 or miR-223), and at least one of the miRs targets plasmacytoid dendritic cells, platelets or endothelial cells (e.g., miR-126), (iii) at least one of the miRs targets hematopoietic progenitor cells (e.g., miR-223, miR-451, miR-26a or (iv) miR-16), at least one of which miRs targets plasmacytoid dendritic cells, platelets or endothelial cells (e.g., miR-126), or (iv) at least one of which miRs targets hematopoietic cells (e.g., miR-142, miR-144, miR-150, miR-155 or miR-223), at least one of which miRs targets B cells (e.g., miR-142, miR-150, miR-16 or miR-223), at least one of which miRs targets plasmacytoid dendritic cells, platelets or endothelial cells (e.g., miR-126), or otherwise combine any of the above four classes of miR binding sites in any possible way (i.e., miRs targeting hematopoietic cells, miRs targeting B cells, miRs targeting hematopoietic progenitor cells and / or miRs targeting plasmacytoid dendritic cells / platelets / endothelial cells).
[0261] In one embodiment, the polynucleotide of the present invention may contain one or more miRNA-binding sequences that bind to one or more miRs expressed in conventional immune cells or any cell that expresses TLR7 and / or TLR8 and secretes pro-inflammatory cytokines and / or chemokines (e.g., immune cells of peripheral lymphoid organs and / or spleen cells and / or endothelial cells) in order to modulate the immune response. It has now been found that incorporating one or more miRs expressed in conventional immune cells or any cell that expresses TLR7 and / or TLR8 and secretes pro-inflammatory cytokines and / or chemokines (e.g., immune cells of peripheral lymphoid organs and / or spleen cells and / or endothelial cells) into mRNA reduces or inhibits the activation of immune cells (e.g., B cell activation as measured by the frequency of activated B cells) and / or the production of cytokines (e.g., production of IL-6, IFN-γ and / or TNFα). Furthermore, it has now been discovered that incorporating one or more miRs expressed in conventional immune cells, or in any cell that expresses TLR7 and / or TLR8 and secretes pro-inflammatory cytokines and / or chemokines (e.g., immune cells and / or spleen cells and / or endothelial cells of peripheral lymphoid organs), into mRNA can reduce or inhibit the anti-drug antibody (ADA) response to the target protein encoded by that mRNA.
[0262] In another embodiment, to regulate the rapid clearance from the blood of polynucleotides delivered with lipid-containing compounds or compositions, the polynucleotides of the present invention may contain one or more miR-binding sequences that bind to one or more miRNAs expressed in conventional immune cells or any cell expressing TLR7 and / or TLR8 and secreting pro-inflammatory cytokines and / or chemokines (e.g., immune cells of peripheral lymphoid organs and / or spleen cells and / or endothelial cells). It has now been found that incorporating one or more miR-binding sites into mRNA reduces or inhibits the rapid clearance (ABC) from the blood of lipid-containing compounds or compositions used to deliver that mRNA. Furthermore, it has now been discovered that when one or more miR binding sites are incorporated into mRNA, serum levels of anti-PEG anti-IgM decrease after administration of a lipid-containing compound or composition containing that mRNA (for example, the rapid production of polyethylene glycol (PEG)-recognizing IgM by B cells is reduced or inhibited), and / or the proliferation and / or activation of plasmacytoid dendritic cells are reduced or inhibited.
[0263] In some embodiments, the miR sequence may correspond to any known microRNA expressed in immune cells, including, but not limited to, those taught in U.S. Patent Publication No. 2005 / 0261218 and U.S. Patent Publication No. 2005 / 0059005 (the contents of which are incorporated herein by reference in their entirety). Non-limiting examples of miRs expressed in immune cells include miRs expressed in spleen cells, myeloid cells, dendritic cells, plasmacytoid dendritic cells, B cells, T cells and / or macrophages. For example, miR-142-3p, miR-142-5p, miR-16, miR-21, miR-223, miR-24, and miR-27 are expressed in myeloid cells, miR-155 is expressed in dendritic cells, B cells, and T cells, miR-146 is upregulated in macrophages by TLR stimulation, and miR-126 is expressed in plasmacytoid dendritic cells. In certain embodiments, miR(s) are expressed abundantly or preferentially in immune cells. For example, miR-142 (miR-142-3p and / or miR-142-5p), miR-126 (miR-126-3p and / or miR-126-5p), miR-146 (miR-146-3p and / or miR-146-5p), and miR-155 (miR-155-3p and / or miR155-5p) are abundantly expressed in immune cells. Since these microRNA sequences are known in the art, those skilled in the art can easily design binding or target sequences that bind to these microRNAs based on Watson-Crick complementarity.
[0264] In one embodiment, the polynucleotide of the present invention contains three copies of the same miRNA binding site. In a particular embodiment, using three copies of the same miRNA binding site can exhibit beneficial properties compared to using a single miRNA binding site.
[0265] In another embodiment, the polynucleotide of the present invention comprises two or more copies (e.g., two, three, or four copies) of binding sites to at least two different miRs expressed in immune cells.
[0266] In another embodiment, the polynucleotide of the present invention comprises at least two miR binding sites for microRNA expressed in immune cells, one of which is for miR-142-3p. In various embodiments, the polynucleotide of the present invention comprises binding sites for miR-142-3p and miR-155 (miR-155-3p or miR-155-5p), miR-142-3p and miR-146 (miR-146-3 or miR-146-5p), or miR-142-3p and miR-126 (miR-126-3p or miR-126-5p).
[0267] In another embodiment, the polynucleotide of the present invention comprises at least two miR binding sites for microRNA expressed in immune cells, one of which is for miR-126-3p. In various embodiments, the polynucleotide of the present invention comprises binding sites for miR-126-3p and miR-155 (miR-155-3p or miR-155-5p), miR-126-3p and miR-146 (miR-146-3p or miR-146-5p), or miR-126-3p and miR-142 (miR-142-3p or miR-142-5p).
[0268] In another embodiment, the polynucleotide of the present invention comprises at least two miR binding sites for microRNA expressed in immune cells, one of which is for miR-142-5p. In various embodiments, the polynucleotide of the present invention comprises binding sites for miR-142-5p and miR-155 (miR-155-3p or miR-155-5p), miR-142-5p and miR-146 (miR-146-3 or miR-146-5p), or miR-142-5p and miR-126 (miR-126-3p or miR-126-5p).
[0269] In yet another embodiment, the polynucleotide of the present invention comprises at least two miR binding sites for microRNA expressed in immune cells, one of which is for miR-155-5p. In various embodiments, the polynucleotide of the present invention comprises binding sites for miR-155-5p and miR-142 (miR-142-3p or miR-142-5p), miR-155-5p and miR-146 (miR-146-3 or miR-146-5p), or miR-155-5p and miR-126 (miR-126-3p or miR-126-5p).
[0270] In some embodiments, the polynucleotide of the present invention includes a miRNA binding site, the miRNA binding site containing one or more nucleotide sequences selected from Table 4 (containing one or more copies of any one or more of the miRNA binding site sequences). In some embodiments, the polynucleotide of the present invention further contains at least one, two, three, four, five, six, seven, eight, nine, or ten or more identical or different miRNA binding sites (including combinations thereof) selected from Table 4.
[0271] In some embodiments, the miRNA binding site binds to or is complementary to miR-142. In some embodiments, miR-142 includes SEQ ID NO: 172. In some embodiments, the miRNA binding site binds to miR-142-3p or miR-142-5p. In some embodiments, the miR-142-3p binding site includes SEQ ID NO: 174. In some embodiments, the miR-142-5p binding site includes SEQ ID NO: 210. In some embodiments, the miRNA binding site includes a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or 100% identical to SEQ ID NO: 174 or SEQ ID NO: 210.
[0272] In some embodiments, the miRNA binding site binds to or is complementary to miR-126. In some embodiments, the miR-126 includes SEQ ID NO: 150. In some embodiments, the miRNA binding site binds to miR-126-3p or miR-126-5p. In some embodiments, the miR-126-3p binding site includes SEQ ID NO: 152. In some embodiments, the miR-126-5p binding site includes SEQ ID NO: 154. In some embodiments, the miRNA binding site includes a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or 100% identical to SEQ ID NO: 152 or SEQ ID NO: 154.
[0273] In one embodiment, the 3'UTR contains two miRNA binding sites, the first miRNA binding site binding to miR-142 and the second miRNA binding site binding to miR-126.
[0274] [Table 4]
[0275] In some embodiments, the miRNA binding site is inserted into the polynucleotide of the present invention at any position on the polynucleotide (e.g., 5'UTR and / or 3'UTR). In some embodiments, the 5'UTR contains the miRNA binding site. In some embodiments, the 3'UTR contains the miRNA binding site. In some embodiments, both the 5'UTR and 3'UTR contain the miRNA binding site. The insertion site on the polynucleotide can be any position on the polynucleotide, as long as the insertion of the miRNA binding site in the polynucleotide does not hinder the translation of the functional polypeptide in the absence of the corresponding miRNA, and in the presence of the miRNA, the insertion of the miRNA binding site in the polynucleotide, and the miRNA binding site binding to the corresponding miRNA, can degrade the polynucleotide or prevent the translation of the polynucleotide.
[0276] In some embodiments, the miRNA binding site is inserted at least about 30 nucleotides downstream from the stop codon of the ORF in the polynucleotide of the present invention, which includes an ORF. In some embodiments, the miRNA binding site is inserted at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, at least about 50 nucleotides, at least about 55 nucleotides, at least about 60 nucleotides, at least about 65 nucleotides, at least about 70 nucleotides, at least about 75 nucleotides, at least about 80 nucleotides, at least about 85 nucleotides, at least about 90 nucleotides, at least about 95 nucleotides, or at least about 100 nucleotides downstream from the stop codon of the ORF of the polynucleotide of the present invention. In some embodiments, the miRNA binding site is inserted approximately 10 to 100 nucleotides downstream, approximately 20 to 90 nucleotides downstream, approximately 30 to 80 nucleotides downstream, approximately 40 to 70 nucleotides downstream, approximately 50 to 60 nucleotides downstream, and approximately 45 to 65 nucleotides downstream from the stop codon of the ORF of the polynucleotide of the present invention.
[0277] In some embodiments, the miRNA binding site is inserted within the 3'UTR immediately following the stop codon in the coding region of the polynucleotide of the present invention, such as mRNA. In some embodiments, if the construct contains multiple copies of the stop codon, the miRNA binding site is inserted immediately following the last stop codon. In some embodiments, the miRNA binding site is inserted further downstream of the stop codon, in which case a 3'UTR base is present between the stop codon and the miRNA binding site(s).
[0278] In some embodiments, one or more miRNA binding sites can be located within the 5'UTR at one or more possible insertion sites.
[0279] In one embodiment, the codon-optimized open reading frame encoding the target polypeptide includes a stop codon, and at least one microRNA binding site is located 1 to 100 nucleotides after the stop codon within the 3'UTR. In another embodiment, the codon-optimized open reading frame encoding the target polypeptide includes a stop codon, and at least one microRNA binding site for miR expressed in immune cells is located 30 to 50 nucleotides after the stop codon within the 3'UTR. In yet another embodiment, the codon-optimized open reading frame encoding the target polypeptide includes a stop codon, and at least one microRNA binding site for miR expressed in immune cells is located at least 50 nucleotides after the stop codon within the 3'UTR. In another embodiment, the codon-optimized open reading frame encoding the polypeptide of interest includes a stop codon, and at least one microRNA binding site for a miR expressed in immune cells is located within the 3'UTR either immediately after the stop codon, 15–20 nucleotides after the stop codon, or 70–80 nucleotides after the stop codon. In another embodiment, the 3'UTR includes two or more miRNA binding sites (e.g., 2–4 miRNA binding sites), and spacer regions (e.g., 10–100 nucleotides, 20–70 nucleotides, or 30–50 nucleotides) may be present between each miRNA binding site. In yet another embodiment, the 3'UTR includes spacer regions between the ends of the miRNA binding site(s) and the poly(A) tail nucleotides. For example, spacer regions of 10-100 nucleotides, 20-70 nucleotides, or 30-50 nucleotides in length can be located between the end of the miRNA binding site(s) and the beginning of the poly(A) tail.
[0280] In one embodiment, the codon-optimized open reading frame encoding the target polypeptide includes a start codon, with at least one microRNA binding site located 1 to 100 nucleotides upstream of the start codon within the 5'UTR. In another embodiment, the codon-optimized open reading frame encoding the target polypeptide includes a start codon, with at least one microRNA binding site for a miR expressed in immune cells located 10 to 50 nucleotides upstream of the start codon within the 5'UTR. In yet another embodiment, the codon-optimized open reading frame encoding the target polypeptide includes a start codon, with at least one microRNA binding site for a miR expressed in immune cells located at least 25 nucleotides upstream of the start codon within the 5'UTR. In another embodiment, the codon-optimized open reading frame encoding the polypeptide of interest includes a start codon and at least one microRNA binding site for a miR expressed in an immune cell is located within the 5'UTR either immediately before the start codon, 15–20 nucleotides before the start codon, or 70–80 nucleotides before the start codon. In another embodiment, the 5'UTR includes two or more miRNA binding sites (e.g., 2–4 miRNA binding sites), and spacer regions (e.g., 10–100 nucleotides, 20–70 nucleotides, or 30–50 nucleotides) may be present between each miRNA binding site.
[0281] In one embodiment, the 3'UTR contains two or more stop codons, and at least one miRNA binding site is located downstream of the stop codons. For example, the 3'UTR may contain one, two, or three stop codons. Non-limiting examples of usable triple stop codons include UGAUAAUAG (SEQ ID NO: 182), UGAUAGUAA (SEQ ID NO: 183), UAAUGAUAG (SEQ ID NO: 184), UGAUAAUAA (SEQ ID NO: 185), UGAUAGUAG (SEQ ID NO: 186), UAAUGAUGA (SEQ ID NO: 187), UAAUAGUAG (SEQ ID NO: 188), UGAUGAUGA (SEQ ID NO: 179), UAAUAAUAA (SEQ ID NO: 180), and UAGUAGUAG (SEQ ID NO: 181). Within the 3'UTR, for example, one, two, three, or four miRNA binding sites, such as miR-142-3p binding sites, can be located directly adjacent to a stop codon(s), or any number of nucleotides downstream from the last stop codon. If the 3'UTR contains multiple miRNA binding sites, these binding sites can be located immediately adjacent to each other (i.e., sequentially) within the construct, or spacer nucleotides can be located between each binding site.
[0282] In one embodiment, the 3'UTR contains three stop codons, with one miR-142-3p binding site located downstream of the third stop codon.
[0283] In one embodiment, the polynucleotide of the present invention comprises a 5'UTR, a codon-optimized open reading frame encoding the polypeptide of interest, a 3'UTR containing at least one miRNA binding site for miRs expressed in immune cells, and a 3' tail region consisting of linked nucleosides. In various embodiments, the 3'UTR contains 1 to 4, at least 2, 1, 2, 3, or 4 miRNA binding sites for miRs expressed in immune cells, preferably miRs that are abundant or preferentially expressed in immune cells.
[0284] In one embodiment, at least one miRNA expressed in an immune cell is a miR-142-3p microRNA binding site. In one embodiment, the miR-142-3p microRNA binding site includes the sequence shown in SEQ ID NO: 174.
[0285] In one embodiment, at least one miRNA expressed in immune cells is a miR-126 microRNA binding site. In one embodiment, the miR-126 binding site is a miR-126-3p binding site. In one embodiment, the miR-126-3p microRNA binding site includes the sequence shown in SEQ ID NO: 152.
[0286] Non-limiting exemplary sequences of miRs to which the microRNA binding site(s)(s)(sequence number 173) of this disclosure can bind include miR-142-3p (sequence number 173), miR-142-5p (sequence number 175), miR-146-3p (sequence number 155), miR-146-5p (sequence number 156), miR-155-3p (sequence number 157), miR-155-5p (sequence number 158), miR-126-3p (sequence number 151), and miR-126-5p (sequence number 158). 3) Examples include miR-16-3p (sequence number 159), miR-16-5p (sequence number 160), miR-21-3p (sequence number 161), miR-21-5p (sequence number 162), miR-223-3p (sequence number 163), miR-223-5p (sequence number 164), miR-24-3p (sequence number 165), miR-24-5p (sequence number 166), miR-27-3p (sequence number 167), and miR-27-5p (sequence number 168). Other suitable miR sequences expressed in immune cells (e.g., abundantly or preferentially expressed in immune cells) are known and available in the art, for example, in the University of Manchester microRNA database miRBase. A binding site to any of the above miRs can be designed based on Watson-Crick complementarity to that miR, typically 100% complementarity, and inserted into the mRNA construct of this disclosure as described herein.
[0287] In another embodiment, the polynucleotide of the present invention (e.g., mRNA, e.g., its 3'UTR) may include at least one miRNA binding site for reducing or inhibiting rapid clearance from the blood, for example, by reducing or inhibiting the production of IgM to PEG by B cells and / or by reducing or inhibiting the proliferation and / or activation of pDCs, and may also include at least one miRNA binding site for regulating the expression of the target protein encoded in tissue.
[0288] Gene regulation by miRNAs can be influenced by the surrounding sequence of the miRNA (including, but not limited to, the species of the surrounding sequence, the type of sequence (e.g., heterologous, homologous, exogenous, endogenous, or artificial), regulatory elements within the surrounding sequence, and / or structural elements within the surrounding sequence). miRNAs can be influenced by the 5'UTR and / or 3'UTR. As a non-limiting example, non-human 3'UTRs can enhance the regulatory effect of a miRNA sequence on the expression of a target polypeptide compared to human 3'UTRs of the same type.
[0289] In one embodiment, other regulatory and / or structural elements of the 5'UTR can influence miRNA-mediated gene regulation. An example of a regulatory and / or structural element is a structured IRES (internal ribosome entry site) within the 5'UTR, which is required for translation elongation factors to initiate protein translation. For miRNA-mediated gene expression, EIF4A2 must bind to this secondary structured element within the 5'UTR (Meijer HA et al., Science, 2013, 340, 82-85 (the entire text is incorporated herein by reference)). The polynucleotide of the present invention may further include this structured 5'UTR to enhance microRNA-mediated gene regulation.
[0290] At least one miRNA binding site can be incorporated into the 3'UTR of the polynucleotide of the present invention. In this context, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten or more miRNA binding sites can be incorporated into the 3'UTR of the polynucleotide of the present invention. For example, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 2, or 1 miRNA binding site can be incorporated into the 3'UTR of the polynucleotide of the present invention. In one embodiment, the miRNA binding sites incorporated into the polynucleotide of the present invention may be the same miRNA site or different miRNA sites. A combination of different miRNA binding sites incorporated into the polynucleotide of the present invention may include a combination in which two or more copies of any one of the different miRNA sites are incorporated. In another embodiment, the miRNA binding sites incorporated into the polynucleotide of the present invention may target the same or different tissues in the body. As a non-limiting example, by introducing a tissue type, cell type, or disease-specific miRNA binding site to the 3'UTR of the polynucleotide of the present invention, the degree of expression in a given type of cell (e.g., myeloid cells, endothelial cells, etc.) can be reduced.
[0291] In one embodiment, the miRNA binding site can be manipulated and incorporated near the 5' end of the 3'UTR of the polynucleotide of the present invention, midway between the 5' and 3' ends of the 3'UTR, and / or near the 3' end of the 3'UTR. As an unrestricted example, the miRNA binding site can be manipulated and incorporated near the 5' end of the 3'UTR, and midway between the 5' and 3' ends of the 3'UTR. As another unrestricted example, the miRNA binding site can be manipulated and incorporated near the 3' end of the 3'UTR, and midway between the 5' and 3' ends of the 3'UTR. As yet another unrestricted example, the miRNA binding site can be manipulated and incorporated near the 5' end of the 3'UTR, and near the 3' end of the 3'UTR.
[0292] In another embodiment, the 3'UTR may contain one, two, three, four, five, six, seven, eight, nine, or ten miRNA binding sites. These miRNA binding sites may be complementary to miRNA, a miRNA seed sequence, and / or miRNA sequences flanking the seed sequence.
[0293] In some embodiments, the expression of the polynucleotide of the present invention can be controlled by incorporating at least one sensor sequence into the polynucleotide and formulating the polynucleotide for administration. As a non-limiting example, the polynucleotide of the present invention can be targeted to tissues or cells by incorporating a miRNA binding site and formulating the polynucleotide in lipid nanoparticles containing ionizable aminolipids (including any of the lipids described herein).
[0294] The polynucleotides of the present invention can be manipulated to achieve further targeted expression in specific tissues, cell types, or biological conditions, based on the miRNA expression patterns in various tissues, cell types, or biological conditions. The polynucleotides of the present invention can be designed by introducing tissue-specific miRNA binding sites to ensure optimal protein expression within tissues, cells, or under specific biological conditions.
[0295] In some embodiments, the polynucleotides of the present invention can be designed to incorporate miRNA binding sites that are 100% identical to known miRNA seed sequences, or less than 100% identical to known miRNA seed sequences. In some embodiments, the polynucleotides of the present invention can be designed to incorporate miRNA binding sites that are at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to known miRNA seed sequences. The miRNA seed sequences can be partially mutated to reduce the miRNA binding affinity, thereby reducing the downmodulation of the polynucleotide. Essentially, the degree of match or mismatch between the miRNA binding site and the miRNA seed can function as a regulatory modulator that more precisely adjusts the ability of the miRNA to regulate protein expression. In addition, mutations in the non-seed region of the miRNA binding site can also affect the ability of the miRNA to regulate protein expression.
[0296] In one embodiment, the miRNA sequence can be incorporated into the stem-loop loop.
[0297] In another embodiment, the miRNA seed sequence can be incorporated into the stem-loop loop, and the miRNA binding site can be incorporated into the 5' or 3' stem of the stem-loop.
[0298] In one embodiment, the polynucleotide of the present invention described herein can be stabilized using a miRNA sequence within the 5'UTR.
[0299] In another embodiment, a miRNA sequence within the 5'UTR of the polynucleotide of the present invention can be used to reduce the reachability of a translation initiation site, such as a start codon (but not limited to one). See, for example, Matsuda et al., PLoS One. 2010 11(5):e15057 (which is incorporated herein by reference in its entirety). In this paper, an antisense-locked nucleic acid (LNA) oligonucleotide and an exon junction complex (EJC) around the start codon (-4 to +37, where A of the AUG codon is +1) were used to reduce the reachability of the first start codon (AUG). Matsuda's paper showed that alteration of the sequence around the start codon by LNA or EJC affects the efficiency, length, and structural stability of the polynucleotide. The polynucleotide of the present invention may include a miRNA sequence near the translation initiation site instead of the LNA sequence or EJC sequence described by Matsuda et al. in order to reduce the reachability of the translation initiation site. The translation initiation site can be located before, after, or within a miRNA sequence. As a non-limiting example, the translation initiation site can be located within a miRNA sequence, such as a seed sequence or binding site.
[0300] In some embodiments, the polynucleotide of the present invention may contain at least one miRNA to suppress antigen presentation by antigen-presenting cells. The miRNA may be a complete miRNA sequence, a miRNA seed sequence, a seedless miRNA sequence, or a combination thereof. As a non-limiting example, the miRNA incorporated into the polynucleotide of the present invention may be specific to the hematopoietic system. As another non-limiting example, the miRNA incorporated into the polynucleotide of the present invention for the purpose of suppressing antigen presentation is miR-142-3p.
[0301] In some embodiments, the polynucleotides of the present invention may contain at least one miRNA to suppress the expression of the encoded polypeptide within a target tissue or cell. As a non-limiting example, the polynucleotides of the present invention may contain at least one miR-142-3p binding site, a miR-142-3p seed sequence, a seedless miR-142-3p binding site, a miR-142-5p binding site, a miR-142-5p seed sequence, a seedless miR-142-5p binding site, a miR-146 binding site, a miR-146 seed sequence, and / or a seedless miR-146 binding site.
[0302] In some embodiments, the polynucleotides of the present invention may contain at least one miRNA binding site within the 3'UTR to selectively degrade mRNA drugs in immune cells and suppress undesirable immunogenic responses caused by drug delivery. In non-limiting examples, the miRNA binding site may increase the instability of the polynucleotides of the present invention in antigen-presenting cells. Non-limiting examples of these miRNAs include miR-142-5p, miR-142-3p, miR-146a-5p, and miR-146-3p.
[0303] In one embodiment, the polynucleotide of the present invention includes at least one miRNA sequence capable of interacting with an RNA-binding protein within the region of the polynucleotide.
[0304] In some embodiments, the polynucleotide (e.g., RNA, e.g., mRNA) of the present invention comprises (i) a sequence-optimized nucleotide sequence (e.g., ORF) encoding a variant PAH, and (ii) a miRNA binding site (e.g., a miRNA binding site that binds to miR-142) and / or a miRNA binding site that binds to miR-126.
[0305] Region with a 5' cap This disclosure also includes polynucleotides that include both a 5' cap and the polynucleotide of the present invention (for example, a polynucleotide containing a nucleotide sequence encoding a variant PAH polypeptide to be expressed).
[0306] The 5' cap structure of natural mRNA is involved in nuclear export, enhances mRNA stability, and binds to mRNA cap-binding proteins (CBPs). CBPs, through association with poly(A)-binding proteins, are involved in the stability and translational capacity of mRNA within cells, forming mature circular mRNA species. The cap also assists in the removal of 5' proximal introns during mRNA splicing.
[0307] In endogenous mRNA molecules, the 5' end can be capped, forming a 5'-ppp-5'- triphosphate bond between the terminal guanosine cap residue and the transcribed sense nucleotide at the 5' end of the mRNA molecule. Next, this 5'-guanylate cap can be methylated to produce an N7-methyl-guanylate residue. The ribose sugars of the end-transcribed nucleotide and / or the second-to-end transcribed nucleotide at the 5' end of the mRNA may also be optionally 2'-O-methylated. By removing the 5'-cap through hydrolysis and cleavage of the guanylate cap structure, nucleic acid molecules such as mRNA can be targeted for degradation.
[0308] In some embodiments, the polynucleotide of the present invention (for example, a polynucleotide comprising a nucleotide sequence encoding a variant PAH polypeptide) incorporates a cap portion.
[0309] In some embodiments, the polynucleotides of the present invention include a non-hydrolyzable cap structure, which prevents the removal of the cap and thus extends the half-life of the mRNA. Since hydrolysis of the cap structure requires cleavage of the 5'-ppp-5' phosphorodiester bond, modified nucleotides can be used in the capping reaction. For example, Vaccinia Capping Enzyme from New England Biolabs (Ipswich, MA) can be used with α-thio-guanosine nucleotides according to the manufacturer's instructions to introduce a phosphorothioate bond at the 5'-ppp-5' cap. Further modified guanosine nucleotides, such as α-methylphosphonate and selenophosphate nucleotides, can be used.
[0310] Further modifications include, but are not limited to, the 2'-O-methylation of the ribose sugar at the 5' end and / or the second nucleotide from the 5' end of the polynucleotide at the 2'-hydroxyl group of the sugar ring (as described above). Multiple distinct 5'-cap structures can be used to provide a 5'-cap for nucleic acid molecules such as polynucleotides that function as mRNA molecules. Cap analogs (hereinafter also referred to herein as synthetic cap analogs, chemical caps, chemical cap analogs, structural cap analogs, or functional cap analogs) differ in chemical structure from natural 5' caps (i.e., endogenous 5' caps, wild-type 5' caps, or physiological 5' caps) while maintaining cap function. Cap analogs can be synthesized chemically (i.e., non-enzymatically) or enzymatically and / or ligated to the polynucleotides of the present invention.
[0311] For example, an anti-reverse cap analog (ARCA) contains two guanines linked by a 5'-5'-triphosphate group, one of which contains an N7 methyl group and a 3'-O-methyl group (i.e., N7,3'-O-dimethyl-guanosine-5'-triphosphate-5'-guanosine(m 7G-3'mppp-G. Synonymous: 3'O-Me-m 7 This can be referred to as G(5')ppp(5')G). The other 3'-O atom (unmodified) of the guanine will be ligated to the 5' terminal nucleotide of the capped polynucleotide. The guanine with methylated N7 and 3'-O will become the terminal portion of the capped polynucleotide.
[0312] Another exemplary cap is mCAP, which is similar to ARCA but has a 2'-O-methyl group on guanosine (i.e., N7,2'-O-dimethyl-guanosine-5'-triphosphate-5'-guanosine, m 7 Gm-ppp-G).
[0313] Another example cap is m 7 It is G-ppp-Gm-A (i.e., N7, guanosine-5'-triphosphate-2'-O-dimethyl-guanosine-adenosine).
[0314] In some embodiments, the cap is a dinucleotide cap analog. In non-limiting examples, the dinucleotide cap analog may be modified with a boranophosphate group or a phosphoroselenoate group at different phosphate positions, such as the dinucleotide cap analog described in U.S. Patent No. 8,519,110 (the contents of which are incorporated herein by reference in their entirety).
[0315] In another embodiment, the cap is a cap analog that is an N7-(4-chlorophenoxyethyl)-substituted dinucleotide form of a cap analog known in the art and / or described herein. Non-limiting examples of cap analogs in N7-(4-chlorophenoxyethyl)-substituted dinucleotide form include N7-(4-chlorophenoxyethyl)-G(5')ppp(5')G and N7-(4-chlorophenoxyethyl)-m 3’-OOne example of a capping analog is G(5')ppp(5')G (see, for example, Kore et al. Bioorganic & Medicinal Chemistry 2013 21:4570-4574 (the contents of which are incorporated herein by reference in their entirety)). In another embodiment, the capping analog of the present invention is a 4-chloro / bromophenoxyethyl analog.
[0316] The polynucleotides of the present invention may also be enzymatically capped after production (either IVT or chemical synthesis) to yield a more faithful 5'-cap structure. As used herein, the term “more faithful” refers to a feature that structurally or functionally closely mirrors or mimics an endogenous or wild-type feature. That is, a “more faithful” feature represents the function and / or structure of an endogenous cell, wild-type cell, native cell, or physiological cell, or surpasses in one or more respects its corresponding endogenous, wild-type, native, or biological feature, more than the synthetic feature or analogue of the prior art. Non-limiting examples of more faithful 5'-cap structures of the present invention include, among other things, structures in which the binding of cap-binding proteins is enhanced, the half-life is extended, the sensitivity to 5' endonucleases is reduced, and / or the removal of the 5' cap is reduced compared to synthetic 5'-cap structures (or wild-type 5'-cap structures, native 5'-cap structures, or biological 5'-cap structures) known in the art. For example, recombinant vaccinia virus capping enzymes and recombinant 2'-O-methyltransferase enzymes can create a canonical 5'-5'-triphosphate bond between the 5' terminal nucleotide of a polynucleotide and a guanine cap nucleotide, where the cap guanine contains N7 methylation and the 5' terminal nucleotide of the mRNA contains 2'-O-methylation. Such a structure is called a Cap1 structure. This cap improves translational capacity and cellular stability, and reduces the activation of pro-inflammatory cytokines in cells compared to other 5' cap analog structures known in the art. Examples of cap structures include, but are not limited to, 7mG(5')ppp(5')N1pN2p(Cap0), 7mG(5')ppp(5')N1mpNp(Cap1), and 7mG(5')-ppp(5')N1mpN2mp(Cap2).
[0317] As a non-limiting example, post-production capping of chimeric polynucleotides can be more efficient, as it can cap nearly 100% of the chimeric polynucleotide. This is in contrast to the approximately 80% capping achieved when a capping analog is linked to the chimeric polynucleotide during the in vitro transcription reaction.
[0318] According to the present invention, the 5'-terminated cap may include an endogenous cap or a cap analog. According to the present invention, the 5'-terminated cap may include a guanine analog. Useful guanine analogs include, but are not limited to, inosine, N1-methyl-guanosine, 2'-fluoro-guanosine, 7-deaza-guanosine, 8-oxo-guanosine, 2-amino-guanosine, LNA-guanosine, and 2-azido-guanosine.
[0319] The present invention also provides exemplary caps, including caps that can be used in co-transcription capping methods during ribonucleic acid (RNA) synthesis using RNA polymerase, such as wild-type RNA polymerase or a variant thereof (e.g., the variants described herein). In one embodiment, the cap can be added when producing RNA in a "one-pot" reaction without requiring a separate capping reaction. That is, in some embodiments, the method involves reacting a polynucleotide template with an RNA polymerase variant, a nucleoside triphosphate, and a cap analog under in vitro transcription reaction conditions for producing an RNA transcript.
[0320] As used herein, the term “cap” includes an inverted G nucleotide, which may contain one or more additional nucleotides at the 3' end of the inverted G nucleotide, for example, one, two or three or more nucleotides at the 3' and 5' UTR of the inverted G nucleotide, for example, at the 5' end of the 5' UTR as described herein.
[0321] An example cap contains the sequence GG, GA, or GGA, where the underlined italicized G is an inverted G nucleotide followed by a 5'-5'-triphosphate group.
[0322] In one embodiment, the cap is a compound of the following formula (I), [ka] or including its stereoisomers, tautomers, or salts, in the formula, [ka] Ring B1 is modified or unmodified guanine. Rings B2 and B3 are, independently, nucleic acid bases or modified nucleic acid bases. X2 is O, S(O) p , NR 24 or CR 25 R 26 And p is 0, 1 or 2, Y0 is either O or CR6R7. Y1 is O, S(O) n , CR6R7 or NR8, where n is 0, 1 or 2, Each --- is either a single bond or none exists, and if each --- is a single bond, Yi is O, S(O) n , CR6R7 or NR8, and if none of the above exist, Y1 is empty, Y2 is (OP(O)R4) m (where m is 0, 1, or 2), or -O-(CR 40 R 41 )u-Q0-(CR 42 R 43 )v-(that Q0 is bonded, O, S(O) r , NR 44 Or CR 45 R 46 (where r is 0, 1, or 2, and u and v are independently 1, 2, 3, or 4), Each R2 and R2' is independently a halo, LNA, or OR3. Each R3 is independently H, C1-C6 alkyl, C2-C6 alkenyl, or C2-C6 alkynyl, and if R3 is C1-C6 alkyl, C2-C6 alkenyl, or C2-C6 alkynyl, it is optionally substituted with one or more OH groups or OC(O)-C1-C6 alkyl groups, one or more OH groups, and one or more C1-C6 alkoxyls. Each R4 and R4 independently represents H, halo, C1-C6 alkyl, OH, SH, SeH, or BH3. - And, Each of R6, R7, and R8 is independently -Q1-T1, where Q1 is a bond or a C1-C3 alkyl linker optionally substituted with one or more halo, cyano, OH, and C1-C6 alkoxy groups, and T1 is H, halo, OH, COOH, cyano, or R s1 And that R s1 These include C1-C3 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxyl, C(O)O-C1-C6 alkyl, C3-C8 cycloalkyl, and C6-C 10 Ariel, NR 31 R 32 , (NR 31 R 32 R 33 ) + , a 4-12 member heterocycloalkyl, or a 5 or 6 member heteroaryl, R s1 These are halo, OH, oxo, C1-C6 alkyl, COOH, C(O)O-C1-C6 alkyl, cyano, C1-C6 alkoxyl, NR 31 R 32 , (NR 31 R 32 R 33 ) + , C3-C8 cycloalkyl, C6-C 10 It is optionally substituted with one or more substituents selected from the group consisting of aryls, 4- to 12-membered heterocycloalkyls, and 5- or 6-membered heteroaryls. R 10 , R 11 , R 12 , R 13 , R 14 and R 15Each of them is independently -Q2-T2, where Q2 is a C1-C3 alkyl linker optionally substituted with one or more of bonding, halo, cyano, OH and C1-C6 alkoxy, and T2 is H, halo, OH, NH2, cyano, NO2, N3, R s2 or OR s2 and R s2 is C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C8 cycloalkyl, C6-C 10 aryl, NHC(O)-C1-C6 alkyl, NR 31 R 32 , (NR 31 R 32 R 33 ), + 4- to 12-member heterocycloalkyl or 5- or 6-member heteroaryl, and R s2 is halo, OH, oxo, C1-C6 alkyl, COOH, C(O)O-C1-C6 alkyl, cyano, C1-C6 alkoxyl, NR 31 R 32 , (NR 31 R 32 R 33 ), + C3-C8 cycloalkyl, C6-C 10 aryl, 4- to 12-member heterocycloalkyl and 5- or 6-member heteroaryl, optionally substituted with one or more substituents selected from the group consisting of, or R 12 is oxo together with R 14 , or R 13 is oxo together with R 15 , R 20 , R 21 , R 22 and R 23 Each of them is independently -Q3-T3, where Q3 is a C1-C3 alkyl linker optionally substituted with one or more of bonding, halo, cyano, OH and C1-C6 alkoxy, and T3 is H, halo, OH, NH2, cyano, NO2, N3, R S3 or OR[[ID=****]] S3 and R S3These include C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C8 cycloalkyl, and C6-C 10 The Rs3 is an aryl, NHC(O)-C1-C6 alkyl, mono-C1-C6 alkylamino, di-C1-C6 alkylamino, 4-12 member heterocycloalkyl or 5 or 6 member heteroaryl, where Rs3 is a halo, OH, oxo, C1-C6 alkyl, COOH, C(O)O-C1-C 6アルキル , cyano, C1-C6 alkoxyl, amino, mono-C1-C6 alkylamino, di-C1-C6 alkylamino, C3-C8 cycloalkyl, C6-C 10 It is optionally substituted with one or more substituents selected from the group consisting of aryls, 4- to 12-membered heterocycloalkyls, and 5- or 6-membered heteroaryls. R 24 , R 25 and R 26 Each of them is independently H or C1-C6 alkyl, R 27 and R 28 Each of these can be H or OR independently. 29 is or R 27 and R 28 They become one, OR 30 -O is formed, each R 29 R is independently H, C1-C6 alkyl, C2-C6 alkenyl or C2-C6 alkynyl, 29 If it is a C1-C6 alkyl, C2-C6 alkenyl, or C2-C6 alkynyl, it is optionally substituted with one or more OH groups or OC(O)-C1-C6 alkyl groups, one or more OH groups, and one or more C1-C6 alkoxyls. R 30 This is a C1-C6 alkylene optionally substituted with one or more halos, OH groups, and C1-C6 alkoxyls. R 31 , R 32 and R 33 Each of these is independently H, C1-C6 alkyl, C3-C8 cycloalkyl, C6-C 10These are aryl groups, 4- to 12-membered heterocycloalkyl groups, or 5- or 6-membered heteroaryl groups. R 40 , R 41 , R 42 and R 43 Each of these independently has one or more OP(O)R 47 R 48 H, halo, OH, cyano, N3, OP(O)R as arbitrarily substituted. 47 R 48 Alternatively, it may be a C1-C6 alkyl group, or one R 41 and one R 43 These, together with the carbon atoms and Q0 bonded to them, form a C4-C 10 Cycloalkyl, 4-14 member heterocycloalkyl, C6-C 10 They form aryl or 5-14 member heteroaryl groups, each of which is cycloalkyl, heterocycloalkyl, phenyl, or 5-6 member heteroaryl, and each is OH, halo, cyano, N3, oxo, or OP(O)R 47 R 48 , optionally substituted with one or more of the following: C1-C6 alkyl, C1-C6 haloalkyl, COOH, C(O)O-C1-C6 alkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, amino, mono-C1-C6 alkylamino, and di-C1-C6 alkylamino. R 44 is an H, C1-C6 alkyl, or amine protecting group. R 45 and R 46 Each of these is independent of H, OP(O)R 47 R 48 , or one or more OP(O)R 47 R 48 It is a C1-C6 alkyl group that is arbitrarily substituted with, R 47 and R 48 Each of these can independently be H, halo, C1-C6 alkyl, OH, SH, SeH, or BH 3- That is the case.
[0323] It should be understood that the capping analogs provided in this invention may include any of the capping analogs described in International Publication No. 2017 / 066797, published on April 20, 2017 (which is incorporated herein by reference in its entirety).
[0324] In some embodiments, the central position of B2 can be a non-ribose molecule such as arabinose.
[0325] In some embodiments, R2 is ethyl-based.
[0326] Therefore, in some embodiments, the cap has the following structure. [ka]
[0327] In another embodiment, the cap has the following structure: [ka]
[0328] In yet another embodiment, the cap has the following structure: [ka]
[0329] In yet another embodiment, the cap has the following structure: [ka]
[0330] In some embodiments, R is an alkyl group (e.g., C1-C6 alkyl). In some embodiments, R is a methyl group (e.g., C1 alkyl). In some embodiments, R is an ethyl group (e.g., C2 alkyl).
[0331] In some embodiments, the cap includes a sequence selected from the sequence GAA, GAC, GAG, GAU, GCA, GCC, GCG, GCU, GGA, GGC, GGG, GGU, GUA, GUC, GUG, and GUU. In some embodiments, the cap includes GAA. In some embodiments, the cap includes GAC. In some embodiments, the cap includes GAG. In some embodiments, the cap includes GAU. In some embodiments, the cap includes GCA. In some embodiments, the cap includes GCC. In some embodiments, the cap includes GCG. In some embodiments, the cap includes GCU. In some embodiments, the cap includes GGA. In some embodiments, the cap includes GGC. In some embodiments, the cap includes GGG. In some embodiments, the cap includes GGU. In some embodiments, the cap includes GUA. In some embodiments, the cap includes GUC. In some embodiments, the cap includes GUG. In some embodiments, the cap includes GUU.
[0332] In some embodiments, the cap is m 7 GpppApA, m 7 GpppApC, m 7 GpppApG, m 7 GpppApU, m 7 GpppCpA, m 7 GpppCpC, m 7 GpppCpG, m 7 GpppCpU, m 7 GpppGpA, m 7 GpppGpC, m 7 GpppGpG, m 7 GpppGpU, m 7 GpppUpA, m 7 GpppUpC, m 7 GpppUpG and m 7 Includes a sequence selected from the sequence GpppUpU.
[0333] In some embodiments, the cap is m 7Includes GpppApA. In some embodiments, the cap is m 7 Includes GpppApC. In some embodiments, the cap is m 7 Includes GpppApG. In some embodiments, the cap is m 7 Includes GpppApU. In some embodiments, the cap is m 7 Includes GpppCpA. In some embodiments, the cap is m 7 Includes GpppCpC. In some embodiments, the cap is m 7 Includes GpppCpG. In some embodiments, the cap is m 7 Includes GpppCpU. In some embodiments, the cap is m 7 Includes GpppGpA. In some embodiments, the cap is m 7 Includes GpppGpC. In some embodiments, the cap is m 7 Includes GpppGpG. In some embodiments, the cap is m 7 Includes GpppGpU. In some embodiments, the cap is m 7 Includes GpppUpA. In some embodiments, the cap is m 7 Includes GpppUpC. In some embodiments, the cap is m 7 Includes GpppUpG. In some embodiments, the cap is m 7 Includes GpppUpU.
[0334] In some embodiments, the cap is m 7 G 3’OMe pppApA, m 7 G 3’OMe pppApC, m 7 G 3’OMe pppApG, m 7 G 3’OMe pppApU, m 7 G 3’OMe pppCpA, m 7 G 3’OMe pppCpC, m 7 G 3’OMe pppCpG, m 7 G 3’OMe pppCpU, m7 G 3’OMe pppGpA, m 7 G 3’OMe pppGpC, m 7 G 3’OMe pppGpG, m 7 G 3’OMe pppGpU, m 7 G 3’OMe pppUpA, m 7 G 3’OMe pppUpC, m 7 G 3’OMe pppUpG and m 7 G 3’OMe Includes sequences selected from the sequence pppUpU.
[0335] In some embodiments, the cap is m 7 G 3’OMe Contains pppApA. In some embodiments, the cap is m 7 G 3’OMe Includes pppApC. In some embodiments, the cap is m 7 G 3’OMe Includes pppApG. In some embodiments, the cap is m 7 G 3’OMe Includes pppApU. In some embodiments, the cap is m 7 G 3’OMe Includes pppCpA. In some embodiments, the cap is m 7 G 3’OMe Includes pppCpC. In some embodiments, the cap is m 7 G 3’OMe Includes pppCpG. In some embodiments, the cap is m 7 G 3’OMe Includes pppCpU. In some embodiments, the cap is m 7 G 3’OMe Contains pppGpA. In some embodiments, the cap is m 7 G 3’OMe Includes pppGpC. In some embodiments, the cap is m 7 G 3’OMe Includes pppGpG. In some embodiments, the cap is m 7 G 3’OMeIncludes pppGpU. In some embodiments, the cap is m 7 G 3’OMe Includes pppUpA. In some embodiments, the cap is m 7 G 3’OMe Includes pppUpC. In some embodiments, the cap is m 7 G 3’OMe Includes pppUpG. In some embodiments, the cap is m 7 G 3’OMe Includes pppUpU.
[0336] In another embodiment, the cap is m 7 G 3’OMe pppA 2’OMe pA, m 7 G 3’OMe pppA 2’OMe pC, m 7 G 3’OMe pppA 2’OMe pG, m 7 G 3’OMe pppA 2’OMe pU, m 7 G 3’OMe pppC 2’OMe pA, m 7 G 3’OMe pppC 2’OMe pC, m 7 G 3’OMe pppC 2’OMe pG, m 7 G 3’OMe pppC 2’OMe pU, m 7 G 3’OMe pppG 2’OMe pA, m 7 G 3’OMe pppG 2’OMe pC, m 7 G 3’OMe pppG 2’OMe pG, m 7 G 3’OMe pppG 2’OMe pU, m 7 G 3’OMe pppU 2’OMe pA, m 7 G 3’OMe pppU 2’OMe pC, m 7 G 3’OMe pppU2’OMe pG and m 7 G 3’OMe pppU 2’OMe Includes an array selected from the array pU.
[0337] In some embodiments, the cap is m 7 G 3’OMe pppA 2’OMe Contains pA. In some embodiments, the cap is m 7 G 3’OMe pppA 2’OMe Includes pC. In some embodiments, the cap is m 7 G 3’OMe pppA 2’OMe Contains pG. In some embodiments, the cap is m 7 G 3’OMe pppA 2’OMe Includes pU. In some embodiments, the cap is m 7 G 3’OMe pppC 2’OMe Contains pA. In some embodiments, the cap is m 7 G 3’OMe pppC 2’OMe Includes pC. In some embodiments, the cap is m 7 G 3’OMe pppC 2’OMe Contains pG. In some embodiments, the cap is m 7 G 3’OMe pppC 2’OMe Includes pU. In some embodiments, the cap is m 7 G 3’OMe pppG 2’OMe Contains pA. In some embodiments, the cap is m 7 G 3’OMe pppG 2’OMe Includes pC. In some embodiments, the cap is m 7 G 3’OMe pppG 2’OMe Contains pG. In some embodiments, the cap is m 7 G 3’OMe pppG 2’OMe Includes pU. In some embodiments, the cap is m 7 G 3’OMe pppU2’OMe Contains pA. In some embodiments, the cap is m 7 G 3’OMe pppU 2’OMe Includes pC. In some embodiments, the cap is m 7 G 3’OMe pppU 2’OMe Contains pG. In some embodiments, the cap is m 7 G 3’OMe pppU 2’OMe Includes pU.
[0338] In yet another embodiment, the cap is m 7 GpppA 2’OMe pA, m 7 GpppA 2’OMe pC, m 7 GpppA 2’OMe pG, m 7 GpppA 2’OMe pU, m 7 GpppC 2’OMe pA, m 7 GpppC 2’OMe pC, m 7 GpppC 2’OMe pG, m 7 GpppC 2’OMe pU, m 7 GpppG 2’OMe pA, m 7 GpppG 2’OMe pC, m 7 GpppG 2’OMe pG, m 7 GpppG 2’OMe pU, m 7 GpppU 2’OMe pA, m 7 GpppU 2’OMe pC, m 7 GpppU 2’OMe pG and m 7 GpppU 2’OMe Includes an array selected from the array pU.
[0339] In some embodiments, the cap is m 7 GpppA 2’OMe Contains pA. In some embodiments, the cap is m 7 GpppA2’OMe Includes pC. In some embodiments, the cap is m 7 GpppA 2’OMe Contains pG. In some embodiments, the cap is m 7 GpppA 2’OMe Includes pU. In some embodiments, the cap is m 7 GpppC 2’OMe Contains pA. In some embodiments, the cap is m 7 GpppC 2’OMe Includes pC. In some embodiments, the cap is m 7 GpppC 2’OMe Contains pG. In some embodiments, the trinucleotide cap is m 7 GpppC 2’OMe Includes pU. In some embodiments, the cap is m 7 GpppG 2’OMe Contains pA. In some embodiments, the cap is m 7 GpppG 2’OMe Includes pC. In some embodiments, the cap is m 7 GpppG 2’OMe Contains pG. In some embodiments, the cap is m 7 GpppG 2’OMe Includes pU. In some embodiments, the cap is m 7 GpppU 2’OMe Contains pA. In some embodiments, the cap is m 7 GpppU 2’OMe Includes pC. In some embodiments, the cap is m 7 GpppU 2’OMe Contains pG. In some embodiments, the cap is m 7 GpppU 2’OMe Includes pU.
[0340] In some embodiments, the cap is m 7 Gpppm 6 A 2’Ome Contains pG. In some embodiments, the cap is m 7 Gpppe 6 A 2’Ome Includes pG.
[0341] In some embodiments, the cap includes GAG. In some embodiments, the cap includes GCG. In some embodiments, the cap includes GUG. In some embodiments, the cap includes GGG.
[0342] In some embodiments, the cap is [ka] It contains one of the following structures.
[0343] In some embodiments, the cap is m7 It comprises GpppN1N2N3, where N1, N2, and N3 are arbitrary (i.e., they may be absent or one or more may be present) and independently a natural nucleoside base, a modified nucleoside base, or a non-natural nucleoside base. In some embodiments, m7 G is further methylated, for example, at the 3' position. In some embodiments, m7 G contains an O-methyl group at the 3' position. In some embodiments, N1, N2, and N3, if present, are optionally and independently adenine, uracil, guanidine, thymine, or cytosine. In some embodiments, one or more (or all) of N1, N2, and N3, if present, are methylated at, for example, the 2' position. In some embodiments, one or more (or all) of N1, N2, and N3, if present, have an O-methyl group at the 2' position.
[0344] In some embodiments, the cap includes the following structure: [ka] In the formula, B1, B2, and B3 are independently a natural nucleoside base, a modified nucleoside base, or a non-natural nucleoside base, and R1, R2, R3, and R4 are independently OH or O-methyl. In some embodiments, R3 is O-methyl and R4 is OH. In some embodiments, R3 and R4 are O-methyl. In some embodiments, R4 is O-methyl. In some embodiments, R1 is OH, R2 is OH, R3 is O-methyl, and R4 is OH. In some embodiments, R1 is OH, R2 is OH, R3 is O-methyl, and R4 is O-methyl. In some embodiments, at least one of R1 and R2 is O-methyl, R3 is O-methyl, and R4 is OH. In some embodiments, at least one of R1 and R2 is O-methyl, R3 is O-methyl, and R4 is O-methyl.
[0345] In some embodiments, B1, B3, and B3 are natural nucleoside bases. In some embodiments, at least one of B1, B2, and B3 is a modified base or a non-natural base. In some embodiments, at least one of B1, B2, and B3 is N6-methyladenine. In some embodiments, B1 is adenine, cytosine, thymine, or uracil. In some embodiments, B1 is adenine, B2 is uracil, and B3 is adenine. In some embodiments, R1 and R2 are OH groups, R3 and R4 are O-methyl groups, B1 is adenine, B2 is uracil, and B3 is adenine.
[0346] In some embodiments, the cap includes a sequence selected from the sequences GAAA, GACA, GAGA, GAUA, GCAA, GCCA, GCGA, GCUA, GGAA, GGCA, GGGA, GGUA, GUCA, and GUUA. In some embodiments, the cap includes a sequence selected from the sequences GAAG, GACG, GAGG, GAUG, GCAG, GCCG, GCGG, GCUG, GGAG, GGCG, GGGG, GGUG, GUCG, GUGG, and GUUG. In some embodiments, the cap includes a sequence selected from the sequences GAAU, GACU, GAGU, GAUU, GCAU, GCCU, GCGU, GCUU, GGAU, GGCU, GGGU, GGUU, GUAU, GUCU, GUGU, and GUUU. In some embodiments, the cap includes a sequence selected from the sequences GAAC, GACC, GAGC, GAUC, GCAC, GCCC, GCGC, GCUC, GGAC, GGCC, GGGC, GGUC, GUAC, GUCC, GUGC, and GUUC.
[0347] In some embodiments, the cap is m 7 G 3’OMe pppApApN, m 7 G 3’OMe pppApCpN, m 7 G 3’OMe pppApGpN, m 7 G 3’OMe pppApUpN, m 7 G 3’OMe pppCpApN, m 7 G 3’OMe pppCpCpN, m 7 G 3’OMe pppCpGpN, m 7 G 3’OMe pppCpUpN, m 7 G 3’OMe pppGpApN, m 7 G 3’OMe pppGpCpN, m 7 G 3’OMe pppGpGpN, m 7 G 3’OMe pppGpUpN, m 7 G 3’OMe pppUpApN, m 7 G3’OMe pppUpCpN, m 7 G 3’OMe pppUpGpN and m 7 G 3’OMe The sequence includes a sequence selected from the sequence pppUpUpN, where N is a natural nucleoside base, a modified nucleoside base, or a non-natural nucleoside base.
[0348] In another embodiment, the cap is m 7 G 3’OMe pppA 2’OMe pApN, m 7 G 3’OMe pppA 2’OMe pCpN, m 7 G 3’OMe pppA 2’OMe pGPN, m 7 G 3’OMe pppA 2’OMe pUpN, m 7 G 3’OMe pppC 2’OMe pApN, m 7 G 3’OMe pppC 2’OMe pCpN, m 7 G 3’OMe pppC 2’OMe pGPN, m 7 G 3’OMe pppC 2’OMe pUpN, m 7 G 3’OMe pppG 2’OMe pApN, m 7 G 3’OMe pppG 2’OMe pCpN, m 7 G 3’OMe pppG 2’OMe pGPN, m 7 G 3’OMe pppG 2’OMe pUpN, m 7 G 3’OMe pppU 2’OMe pApN, m 7 G 3’OMe pppU 2’OMe pCpN, m 7 G 3’OMe pppU 2’OMe pGpN and m 7 G 3’OMe pppU2’OMe The sequence contains a sequence selected from the sequence pUpN, where N is a natural nucleoside base, a modified nucleoside base, or a non-natural nucleoside base.
[0349] In yet another embodiment, the cap is m 7 GpppA 2’OMe pApN, m 7 GpppA 2’OMe pCpN, m 7 GpppA 2’OMe pGPN, m 7 GpppA 2’OMe pUpN, m 7 GpppC 2’OMe pApN, m 7 GpppC 2’OMe pCpN, m 7 GpppC 2’OMe pGPN, m 7 GpppC 2’OMe pUpN, m 7 GpppG 2’OMe pApN, m 7 GpppG 2’OMe pCpN, m 7 GpppG 2’OMe pGPN, m 7 GpppG 2’OMe pUpN, m 7 GpppU 2’OMe pApN, m 7 GpppU 2’OMe pCpN, m 7 GpppU 2’OMe pGpN and m 7 GpppU 2’OMe The sequence contains a sequence selected from the sequence pUpN, where N is a natural nucleoside base, a modified nucleoside base, or a non-natural nucleoside base.
[0350] In another embodiment, the cap is m 7 G 3’OMe pppA 2’OMe pA 2’OMe pN, m 7 G 3’OMe pppA 2’OMe PC 2’OMe pN, m 7 G 3’OMepppA 2’OMe pG 2’OMe pN, m 7 G 3’OMe pppA 2’OMe pU 2’OMe pN, m 7 G 3’OMe pppC 2’OMe pA 2’OMe pN, m 7 G 3’OMe pppC 2’OMe PC 2’OMe pN, m 7 G 3’OMe pppC 2’OMe pG 2’OMe pN, m 7 G 3’OMe pppC 2’OMe pU 2’OMe pN, m 7 G 3’OMe pppG 2’OMe pA 2’OMe pN, m 7 G 3’OMe pppG 2’OMe PC 2’OMe pN, m 7 G 3’OMe pppG 2’OMe pG 2’OMe pN, m 7 G 3’OMe pppG 2’OMe pU 2’OMe pN, m 7 G 3’OMe pppU 2’OMe pA 2’OMe pN, m 7 G 3’OMe pppU 2’OMe PC 2’OMe pN, m 7 G 3’OMe pppU 2’OMe pG 2’OMe pN and m 7 G 3’OMe pppU 2’OMe pU 2’OMe It contains a sequence selected from the sequence pN, where N is a natural nucleoside base, a modified nucleoside base, or a non-natural nucleoside base.
[0351] In yet another embodiment, the cap is m 7 GpppA2’OMe pA 2’OMe pN, m 7 GpppA 2’OMe PC 2’OMe pN, m 7 GpppA 2’OMe pG 2’OMe pN, m 7 GpppA 2’OMe pU 2’OMe pN, m 7 GpppC 2’OMe pA 2’OMe pN, m 7 GpppC 2’OMe PC 2’OMe pN, m 7 GpppC 2’OMe pG 2’OMe pN, m 7 GpppC 2’OMe pU 2’OMe pN, m 7 GpppG 2’OMe pA 2’OMe pN, m 7 GpppG 2’OMe PC 2’OMe pN, m 7 GpppG 2’OMe pG 2’OMe pN, m 7 GpppG 2’OMe pU 2’OMe pN, m 7 GpppU 2’OMe pA 2’OMe pN, m 7 GpppU 2’OMe PC 2’OMe pN, m 7 GpppU 2’OMe pG 2’OMe pN and m 7 GpppU 2’OMe pU 2’OMe It contains a sequence selected from the sequence pN, where N is a natural nucleoside base, a modified nucleoside base, or a non-natural nucleoside base.
[0352] In some embodiments, the cap includes a GGAG. In some embodiments, the cap includes the following structure. [ka]
[0353] Poly A Tail In some embodiments, the polynucleotides of the present disclosure (e.g., polynucleotides comprising a nucleotide sequence encoding a variant PAH polypeptide) further comprise a poly-A tail. In further embodiments, terminal groups of the poly-A tail may be incorporated for stabilization. In another embodiment, the poly-A tail comprises a des-3' hydroxyl tail.
[0354] During RNA processing, a long chain of adenine nucleotides (poly-A tail) can be added to polynucleotides such as mRNA molecules to enhance stability. Immediately after transcription, the 3' end of the transcript can be cleaved to release the 3' hydroxyl group. Subsequently, poly-A polymerase adds the adenine nucleotide chain to the RNA. This process, called polyadenylation, adds a poly-A tail that can be approximately 80 to 250 residues long (including approximately 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 residue lengths). In one embodiment, the poly(A) tail is 100 nucleotides long (SEQ ID NO: 195).
[0355] Poly-A tails can also be attached after the construct has been transported from the nucleus.
[0356] According to this disclosure, terminal groups of the poly(A) tail can be incorporated for stabilization. The polynucleotides of the present invention may include a des-3' hydroxyl tail. The polynucleotides of the present invention may also include structural moieties or 2'-O-methyl modifications as taught by Junjie Li et al. (Current Biology, Vol. 15, 1501-1507, August 23, 2005 (the contents of which are incorporated herein by reference in their entirety)).
[0357] The polynucleotides of the present invention can be designed to encode transcripts having alternative poly(A) tail structures, including histone mRNA. According to Norbury, "Terminal uridylation has also been detected in human replication-dependent histone mRNA. Turnover of these mRNAs is thought to be important in preventing the accumulation of potentially toxic histones after chromosomal DNA replication is complete or inhibited. These mRNAs are identified by the deletion of the 3' poly(A) tail, and instead, their function is carried out by a stable stem-loop structure and its cognitive stem-loop-binding protein (SLBP), which performs the same function as PABP on polyadenylated mRNA" (Norbury, "Cytoplasmic RNA: a case of the tail wagging the dog," Nature Reviews Molecular Cell Biology; AOP, published online 29 August 2013; doi:10.1038 / nrm3645) (the contents of which are incorporated herein by reference in their entirety).
[0358] The specific length of the poly-A tail provides certain advantages of the polynucleotides of the present invention. Generally, when a poly-A tail is present, its length is greater than 30 nucleotides. In another embodiment, the poly-A tail is greater than 35 nucleotides (for example, at least about 35 nucleotides, about 40 nucleotides, about 45 nucleotides, about 50 nucleotides, about 55 nucleotides, about 60 nucleotides, about 70 nucleotides, about 80 nucleotides, about 90 nucleotides, about 100 nucleotides, about 120 nucleotides, about 140 nucleotides, about 160 nucleotides, about 180 nucleotides, about 200 nucleotides, about 250 nucleotides, about 300 nucleotides, about 350 nucleotides, about 400 nucleotides, about 450 nucleotides). Rheotides are approximately 500 nucleotides, 600 nucleotides, 700 nucleotides, 800 nucleotides, 900 nucleotides, 1,000 nucleotides, 1,100 nucleotides, 1,200 nucleotides, 1,300 nucleotides, 1,400 nucleotides, 1,500 nucleotides, 1,600 nucleotides, 1,700 nucleotides, 1,800 nucleotides, 1,900 nucleotides, 2,000 nucleotides, 2,500 nucleotides, and 3,000 nucleotides, or exceed these lengths.
[0359] In some embodiments, the polynucleotide or region of the present invention comprises about 30 to about 3,000 nucleotides (e.g., 30-50 nucleotides, 30-100 nucleotides, 30-250 nucleotides, 30-500 nucleotides, 30-750 nucleotides, 30-1,000 nucleotides, 30-1,500 nucleotides, 30-2,000 nucleotides, 30-2,500 nucleotides, 50-100 nucleotides, 50-250 nucleotides, 50-500 nucleotides, 50-750 nucleotides, 50-1,000 nucleotides, 50-1,500 nucleotides, 50-2,000 nucleotides, 50-2,500 nucleotides, 50-3,000 nucleotides, 100-500 nucleotides, 100-750 nucleotides, 100-1,000 nucleotides). Includes nucleotides (100-1,500 nucleotides, 100-2,000 nucleotides, 100-2,500 nucleotides, 100-3,000 nucleotides, 500-750 nucleotides, 500-1,000 nucleotides, 500-1,500 nucleotides, 500-2,000 nucleotides, 500-2,500 nucleotides, 500-3,000 nucleotides, 1,000-1,500 nucleotides, 1,000-2,000 nucleotides, 1,000-2,500 nucleotides, 1,000-3,000 nucleotides, 1,500-2,000 nucleotides, 1,500-2,500 nucleotides, 1,500-3,000 nucleotides, 2,000-3,000 nucleotides, 2,000-2,500 nucleotides, and 2,500-3,000 nucleotides).
[0360] In some embodiments, the polyA tail is designed for the length of the entire polynucleotide or for the length of a specific region of the polynucleotide. This design may be based on the length of the coding region, the length of a specific feature or region, or the length of the final product expressed from the polynucleotide.
[0361] In this context, the poly(A) tail can be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% longer than the polynucleotide or its feature region. The poly(A) tail can also be designed as a portion of the polynucleotide to which it belongs. In this context, the poly(A) tail can be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% or more of the full length of the construct, the construct region, or the length of the full length of the construct excluding the poly(A) tail. Furthermore, expression can be enhanced by conjugation of the engineered binding site and polynucleotide to the poly(A) binding protein.
[0362] In addition, using the modified nucleotide at the 3' end of the poly(A) tail, multiple separate polynucleotides can be linked together via PABP (poly(A) binding protein) through the 3' end. Transfection experiments can be performed in relevant cell lines, and protein production can be assayed by ELISA 12, 24, 48, 72, and 7 days after transfection.
[0363] In some embodiments, the polynucleotides of the present invention are designed to include a polyAG quartet region. A G-quartet is a circular array of four guanine nucleotides linked by hydrogen bonds, and can be formed by G-rich sequences in both DNA and RNA. In this embodiment, the G-quartet is incorporated at the end of a polyA tail. The resulting polynucleotides are assayed for stability, protein production, and other parameters, including half-lives at various time points. It has been found that the amount of protein produced from mRNA using the polyAG quartet is at least 75% of the amount obtained using only a 120-nucleotide polyA tail (SEQ ID NO: 196).
[0364] In some embodiments, the poly(A) tail contains an alternative nucleoside, such as inverted thymidine. The poly(A) tail containing an alternative nucleoside, such as inverted thymidine, may be produced as described herein. For example, the mRNA construct may be modified by ligation to stabilize the poly(A) tail. Ligation may be performed using 0.5–1.5 mg / mL of mRNA (5':Cap1, 3':A100), 50 mM Tris-HCl (pH 7.5), 10 mM MgCl2, 1 mM TCEP, 1000 units / mL of T4 RNA ligase 1, 1 mM ATP, 20% (w / v) polyethylene glycol 8000, and a molar ratio of 5:1 modified oligo:mRNA. The modified oligonucleotide has the sequence 5'-phosphate-AAAAAAAAAAAAAAAAAAAA-(inverted deoxythymidine (idT) (SEQ ID NO: 209)) (see below). Mix the ligation reaction products and incubate at room temperature (approximately 22°C) for, for example, 4 hours. The mRNA containing the stable tail is purified by, for example, dT purification, reverse-phase purification, hydroxyapatite purification, ultrafiltration into water, and sterile filtration. The obtained mRNA containing the stable tail has an A region at its 3' end, starting with a poly(A) region. 100 It contains the structure -UCUAGAAAAAAAAAAAAAAAAAAAA- inverted deoxythymidine (SEQ ID NO: 258).
[0365] The modified oligo(5'-phosphate-AAAAAAAAAAAAAAAAAAAA-(inverted deoxythymidine)(SEQ ID NO: 209)) used to stabilize the tail is as follows: [ka]
[0366] In some cases, the poly-A tail contains A100-UCUAG-A20-reverse deoxythymidine (SEQ ID NO: 258). In some cases, the poly-A tail consists of A100-UCUAG-A20-reverse deoxythymidine (SEQ ID NO: 258).
[0367] Start codon region The present invention also includes polynucleotides comprising both a start codon region and a polynucleotide described herein (e.g., a polynucleotide comprising a nucleotide sequence encoding a variant PAH polypeptide). In some embodiments, the polynucleotides of the present invention may have a region similar to or functioning like a start codon region.
[0368] In some embodiments, translation of polynucleotides can be initiated with a codon other than the start codon AUG. Translation of polynucleotides can be initiated with alternative start codons such as ACG, AGG, AAG, CTG / CUG, GTG / GUG, ATA / AUA, ATT / AUU, TTG / UUG (see Touriol et al. Biology of the Cell 95(2003)169-178 and Matsuda and Mauro PLoS ONE, 2010 5:11 (the contents of which are incorporated herein by reference)).
[0369] As an unrestricted example, the translation of a polynucleotide may begin with the alternative start codon ACG. As another unrestricted example, the translation of a polynucleotide may begin with the alternative start codon CTG or CUG. As yet another unrestricted example, the translation of a polynucleotide may begin with the alternative start codon GTG or GUG.
[0370] Nucleotides flanking a translation start codon, such as a start codon or an alternative start codon (but not limited to these), are known to affect the translation efficiency, length, and / or structure of a polynucleotide (see, for example, Matsuda and Mauro PLoS ONE, 2010 5:11, which is incorporated herein by reference). By masking any of the nucleotides flanking the translation start codon, the translation start position, translation efficiency, length, and / or structure of a polynucleotide can be altered.
[0371] In some embodiments, a masking agent can be used to mask or shield a start codon or alternative start codon near the start codon to reduce the probability of translation initiation at the masked start codon or alternative start codon. Non-limiting examples of masking agents include antisense-locked nucleic acid (LNA) polynucleotides and exon junction complexes (EJCs) (see, for example, Matsuda and Mauro (PLoS ONE, 2010 5:11) (the entire work is incorporated herein by reference)), which describes LNA polynucleotides and EJCs as masking agents).
[0372] In another embodiment, a masking agent can be used to mask the start codon of a polynucleotide, increasing the likelihood that translation will initiate at an alternative start codon. In some embodiments, a masking agent can be used to mask a first start codon or an alternative start codon, increasing the likelihood that translation will initiate at a start codon or an alternative start codon downstream of the masked start codon or alternative start codon.
[0373] In some embodiments, the start codon or alternative start codon can be located within a region that is perfectly complementary to the miRNA binding site. This region, as well as a masking agent, can help control the translation, length, and / or structure of the polynucleotide. As a non-limiting example, the start codon or alternative start codon can be located in the center of a region perfectly complementary to the miRNA binding site. This start codon or alternative start codon can be located after the 1st, 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, 10th, 11th, 12th, 13th, 14th, 15th, 16th, 17th, 18th, 19th, 20th, or 21st nucleotide.
[0374] In another embodiment, the start codon of a polynucleotide can be removed from its polynucleotide sequence, and translation of the polynucleotide can be initiated with a codon other than the start codon. Translation of the polynucleotide can be initiated with a codon after the removed start codon, or with a downstream start codon or an alternative start codon. In a non-limiting example, the start codon ATG or AUG can be removed as the first three nucleotides of a polynucleotide sequence, and translation can be initiated with a downstream start codon or an alternative start codon. The polynucleotide sequence from which the start codon has been removed may further contain at least one masking agent for the downstream start codon and / or alternative start codon to control, or attempt to control, the initiation of translation, the length of the polynucleotide, and / or the structure of the polynucleotide.
[0375] Combination of mRNA elements Any of the polynucleotides disclosed herein may contain one, two, three, or all of the following elements: (a) a 5'UTR as described herein, for example; (b) a coding region including a termination element (for example, as described herein); (c) a 3'UTR (for example, as described herein); and optionally (d) a 3' stabilization region as described herein, for example. Also disclosed herein are LNP compositions comprising the above polynucleotides.
[0376] In some embodiments, the polynucleotide of the Disclosure comprises (a) a 5'UTR, or a variant or fragment thereof, as described in Table 2, and (b) a coding region comprising a termination element as shown herein. In some embodiments, the polynucleotide further comprises, for example, a cap structure as described herein, or, for example, a poly-A tail as described herein. In some embodiments, the polynucleotide further comprises, for example, a 3' stabilizing region as described herein.
[0377] In some embodiments, the polynucleotides of this disclosure include (a) a 5'UTR, or a variant or fragment thereof, as described in Table 2, and (c) a 3'UTR, or a variant or fragment thereof, as described in Table 3. In some embodiments, the polynucleotides further include, for example, a cap structure as described herein, or, for example, a poly-A tail as described herein. In some embodiments, the polynucleotides further include, for example, a 3' stabilizing region as described herein.
[0378] In some embodiments, the polynucleotide of the Disclosure comprises (c) a 3'UTR, or a variant or fragment thereof, as described in Table 3, and (b) a coding region comprising a termination element as shown herein. In some embodiments, the polynucleotide of the Disclosure comprises a sequence as shown in Table 5. In some embodiments, the polynucleotide further comprises, for example, a cap structure as described herein, or, for example, a poly-A tail as described herein. In some embodiments, the polynucleotide further comprises, for example, a 3' stabilizing region as described herein.
[0379] In some embodiments, the polynucleotide of this disclosure comprises (a) a 5'UTR, or a variant or fragment thereof, as described in Table 2; (b) a coding region comprising a termination element as described herein; and (c) a 3'UTR, or a variant or fragment thereof, as described in Table 3. In some embodiments, the polynucleotide further comprises, for example, a cap structure as described herein, or, for example, a poly-A tail as described herein. In some embodiments, the polynucleotide further comprises, for example, a 3' stabilizing region as described herein.
[0380] [Table 5-1] [Table 5-2]
[0381] Stop codon region The present invention also includes polynucleotides comprising both a stop codon region and a polynucleotide as described herein (e.g., a polynucleotide comprising a nucleotide sequence encoding a variant PAH polypeptide). In some embodiments, the polynucleotide of the present invention may include at least two stop codons prior to the 3' untranslated region (UTR). The stop codons can be selected from TGA, TAA, and TAG in the case of DNA, or from UGA, UAA, and UAG in the case of RNA. In some embodiments, the polynucleotide of the present invention comprises a stop codon called TGA in the case of DNA, or a stop codon called UGA in the case of RNA, and one additional stop codon. In further embodiments, the additional stop codon may be TAA or UAA. In another embodiment, the polynucleotide of the present invention comprises three consecutive stop codons, four stop codons, or five or more consecutive stop codons.
[0382] Stable tail As described herein, mRNA may optionally contain a stable tail to stabilize its poly(A) tail. For example, the mRNA described herein may contain inverted deoxythymidine (idT) as a stable tail. idT can be bound to mRNA, for example, by a ligation reaction using a modified oligonucleotide having the sequence 5'-phosphate-AAAAAAAAAAAAAAAAAAAA-idT (SEQ ID NO: 209), as shown below. [ka]
[0383] WO2017049275 (which is incorporated herein by reference in its entirety) describes exemplary means for binding a stable tail to mRNA.
[0384] Polynucleotides containing mRNA encoding PAH polypeptides In certain embodiments, a polynucleotide of the present disclosure, for example, a polynucleotide comprising an mRNA nucleotide sequence encoding a PAH polypeptide, is configured from the 5' end to the 3' end. (i) The 5' cap shown above, (ii) A 5'UTR like the sequence shown above, (iii) an open reading frame encoding a variant PAH polypeptide, for example, a sequence-optimized nucleic acid sequence encoding a PAH as disclosed herein, (iv) at least one stop codon and (v) A 3'UTR like the sequence shown above, (vi) The poly-A tail shown above, Includes.
[0385] In some embodiments, the polynucleotide further comprises a miRNA-binding site, for example, a miRNA-binding site that binds to miRNA-142. In some embodiments, its 5'UTR comprises a miRNA-binding site. In some embodiments, its 3'UTR comprises a miRNA-binding site.
[0386] In some embodiments, the polynucleotides of the present disclosure include nucleotide sequences encoding polypeptide sequences that are at least 70%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the protein sequence of a variant human PAH (e.g., SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, or SEQ ID NO: 12).
[0387] In some embodiments, the polynucleotide of the present disclosure, for example, a polynucleotide comprising an mRNA nucleotide sequence encoding a polypeptide, comprises (1) a 5' cap as shown above, e.g., CAP1, (2) a 5' UTR, (3) an ORF selected from the group consisting of SEQ ID NOs. 22, SEQ ID NOs. 23, SEQ ID NOs. 24, SEQ ID NOs. 25, SEQ ID NOs. 26, SEQ ID NOs. 27, SEQ ID NOs. 28, SEQ ID NOs. 29, SEQ ID NOs. 30, or SEQ ID NOs. 31, (3) a stop codon, (4) a 3' UTR, and (5) a poly-A tail as shown above, e.g., a poly-A tail of about 100 residues.
[0388] The exemplary variant PAH nucleotide constructs described herein include SEQ ID NOs: 42, 43, 44, 45, 46, 47, 48, 49, 250, and 251 (each containing the 5'UTR, variant PAH nucleotide ORF, and 3'UTR, respectively, from the 5' end to the 3' end).
[0389] In a particular embodiment, in a construct including SEQ ID NOs. 42, 43, 44, 45, 46, 47, 48, 49, 250, and 251, all uracil is replaced with N1-methylpseudolacil.
[0390] In some embodiments, the polynucleotide of the Disclosure, for example, a polynucleotide comprising an mRNA nucleotide sequence encoding a PAH polypeptide, comprises (1) a 5' cap as shown above, e.g., CAP1; (2) a nucleotide sequence selected from the group consisting of SEQ ID NOs. 42, 43, 44, 45, 46, 47, 48, 49, 250, and 251; and (3) a poly-A tail as shown above, e.g., a poly-A tail of about 100 residues. In certain embodiments, all uracil is replaced with N1-methylpseudolacil.
[0391] Method for producing polynucleotides This disclosure also provides a method for producing the polynucleotide (for example, a polynucleotide comprising a nucleotide sequence encoding a variant PAH polypeptide) or its complement.
[0392] In some embodiments, the polynucleotides disclosed herein (e.g., RNA, e.g., mRNA) encoding a variant PAH polypeptide can be constructed by in vitro transcription (IVT). In other embodiments, the polynucleotides disclosed herein (e.g., RNA, e.g., mRNA) encoding a PAH polypeptide can be constructed by chemical synthesis using an oligonucleotide synthesizer.
[0393] In another embodiment, the polynucleotides disclosed herein (e.g., RNA, e.g., mRNA) encoding a variant PAH polypeptide are prepared using host cells. In a particular embodiment, the polynucleotides disclosed herein (e.g., RNA, e.g., mRNA) encoding a PAH polypeptide are prepared by IVT, chemosynthesis, expression in host cells, or a combination of one or more other methods known in the art.
[0394] Natural nucleosides, non-natural nucleosides, or combinations thereof can completely or partially replace natural nucleosides present in candidate nucleotide sequences and can be incorporated into sequence-optimized nucleotide sequences (e.g., RNA, e.g., mRNA) encoding variant PAH polypeptides. Subsequently, the resulting polynucleotide, e.g., mRNA, can be examined for its ability to produce proteins and / or deliver therapeutic outcomes.
[0395] a in vitro transcription / enzymatic synthesis The polynucleotides of the present invention disclosed herein (e.g., polynucleotides comprising a nucleotide sequence encoding a variant PAH polypeptide) can be transcribed using an in vitro transcription (IVT) system. The system typically comprises a transcription buffer, nucleotide triphosphates (NTPs), an RNase inhibitor, and a polymerase. The NTPs can be selected from those described herein (but not limited to), including natural and non-natural NTPs (modified NTPs). The polymerases can be selected from mutant polymerases (but not limited to) such as T7 RNA polymerase, T3 RNA polymerase, and polymerases (but not limited to) that can incorporate the polynucleotides disclosed herein. See U.S. Patent Application Publication No. 20130259923, which is incorporated herein by reference in its entirety.
[0396] Any number of RNA polymerases or variants can be used in the synthesis of the polynucleotides of the present invention. RNA polymerases can be modified by inserting or deleting amino acids in their RNA polymerase sequence. As a non-limiting example, RNA polymerases can be modified to show an improved ability to incorporate 2'-modified nucleotide triphosphates compared to unmodified RNA polymerases (see International Publication No. 2008078180 and U.S. Patent No. 8,101,385, which are incorporated herein by reference in their entirety).
[0397] Variants can be obtained by evolving RNA polymerase, optimizing the amino acid sequence and / or nucleic acid sequence of RNA polymerase, and / or by using other methods known in the art. As a non-limiting example, T7 RNA polymerase variants can be evolved using a continuous directed evolution system defined by Esvelt et al. (Nature 472:499-503 (2011) (which is incorporated herein by reference in its entirety)), in which case T7 RNA polymerase clones include the substitution of lysine to threonine at position 93 (K93T), I4M, A7T, E63V, V64D, A65E, D66Y, T76N, C125R, S128R, A136T, N165S, G175R, H176L, Y178H, F182L, L196F, G198V, D208Y, E222K, S228A, Q239R, T243N, G259D, M267I, G280C, H300R, D351A, A354S, E356 It can encode at least one mutation such as D, L360P, A383V, Y385C, D388Y, S397R, M401T, N410S, K450R, P451T, G452V, E484A, H523L, H524N, G542V, E565K, K577E, K577M, N601S, S684Y, L699I, K713E, N748D, Q754R, E775K, A827V, D851N, or L864F (but not limited to these). As another non-limiting example, T7 RNA polymerase variants can encode at least mutations as described in U.S. Patent Publications 20100120024 and 20070117112 (which are incorporated herein by reference in their entirety). RNA polymerase variants may also include, but are not limited to, substitution variants, conserved amino acid substitution variants, insertion variants, and / or deletion variants.
[0398] In one embodiment, the polynucleotide can be designed to be recognized by wild-type or variant RNA polymerase. In this case, the polynucleotide can be modified to include a resequencing site or region from the wild-type chimeric polynucleotide or the parental chimeric polynucleotide.
[0399] The synthesis of polynucleotides or nucleic acids can be carried out by enzymatic methods using polymerases. Polymerases catalyze the formation of phosphodiester bonds between nucleotides within polynucleotide chains or nucleic acid chains. Currently known DNA polymerases can be classified into various families based on amino acid sequence comparison and crystal structure analysis. The DNA polymerase I (pol I) or A polymerase family (including the Klenow fragment of E. coli, Bacillus DNA polymerase I, Thermus aquaticus (Taq) DNA polymerase, and T7 RNA polymerase and T7 DNA polymerase) is the most studied of these families. Another large family is the DNA polymerase α (pol α) or B polymerase family, which includes replication DNA polymerases of all eukaryotes, as well as polymerases of T4 phage and RB69 phage. These polymerase families employ similar catalytic mechanisms, but differ in substrate specificity, substrate analog uptake efficiency, degree and rate of primer extension, mode of DNA synthesis, exonuclease activity, and sensitivity to inhibitors.
[0400] DNA polymerases are selected based on the optimal reaction conditions required for those polymerases, including reaction temperature, pH, and template and primer concentrations. In some cases, combinations of two or more DNA polymerases are used to achieve the desired DNA fragment size and synthesis efficiency. For example, Cheng et al. used a secondary thermostable DNA polymerase with 3'→5' exonuclease activity, increased pH, glycerol and dimethyl sulfoxide, shortened denaturation time, and increased extension time to efficiently amplify long targets from cloned inserts and human genomic DNA (Cheng et al., PNAS 91:5695-5699 (1994) (the contents of which are incorporated herein by reference)). RNA polymerases derived from bacteriophages T3, T7, and SP6 have been widely used to prepare RNA for biochemical and biophysical research. RNA polymerase, capping enzyme, and poly(A) polymerase are disclosed in concurrently pending International Publication No. 2014 / 028429 (the contents of which are incorporated herein by reference in their entirety).
[0401] In one embodiment, the RNA polymerase that can be used for the synthesis of polynucleotides in the present invention is Syn5 RNA polymerase (see Zhu et al. Nucleic Acids Research 2013, doi:10.1093 / nar / gkt1193 (the entire article is incorporated herein by reference)). Recently, Zhu et al. characterized Syn5 RNA polymerase from the marine cyanophage Syn5 and identified its promoter sequence (see Zhu et al. Nucleic Acids Research 2013 (the entire article is incorporated herein by reference)). Zhu et al. found that Syn5 RNA polymerase catalyzes RNA synthesis over a wide range of temperatures and salinities compared to T7 RNA polymerase. In addition, the stringency of the promoter start nucleotide requirement has been found to be lower in Syn5 RNA polymerase compared to T7 RNA polymerase, making Syn5 RNA polymerase a promising choice for RNA synthesis.
[0402] In one embodiment, Syn5 RNA polymerase can be used for the synthesis of polynucleotides described herein. As a non-limiting example, Syn5 RNA polymerase can be used for the synthesis of polynucleotides that require a precise 3' end.
[0403] In one embodiment, the Syn5 promoter can be used for the synthesis of polynucleotides. As a non-limiting example, the Syn5 promoter may be 5'-ATTGGGCACCCGTAAGGG-3' (SEQ ID NO: 252, as described by Zhu et al. (Nucleic Acids Research 2013)).
[0404] In one embodiment, Syn5 RNA polymerase can be used to synthesize polynucleotides comprising at least one of the chemical modifications described herein and / or known in the art (see, for example, the introduction of pseudo-UTP and 5Me-CTP as described in Zhu et al. Nucleic Acids Research 2013).
[0405] In one embodiment, the polynucleotides described herein can be synthesized using Syn5 RNA polymerase purified using a modified and improved purification procedure described in the literature by Zhu et al. (Nucleic Acids Research 2013).
[0406] Various genetic engineering tools are based on enzymatically amplifying target genes that function as templates. For studying the sequences of individual genes or specific regions of interest, and for other research needs, it is necessary to produce multiple copies of a target gene from small samples of polynucleotides or nucleic acids. Such methods can be applied to the production of polynucleotides according to the present invention.
[0407] For example, polymerase chain reaction (PCR), strand substitution amplification (SDA), nucleic acid sequence-based amplification (NASBA) (also known as transcription-mediated amplification (TMA)), and / or rolling circle amplification (RCA) can be used to produce one or more regions of the polynucleotide of the present invention. Ligase-based assembly of polynucleotides or nucleic acids is also widely used.
[0408] b Chemical synthesis By applying standard methods, an isolated polynucleotide sequence encoding a target isolated polypeptide can be synthesized, such as the polynucleotide of the present invention (e.g., a polynucleotide comprising a nucleotide sequence encoding a variant PAH polypeptide). For example, a single DNA oligomer or RNA oligomer containing a codon-optimized nucleotide sequence encoding a specific isolated polypeptide can be synthesized. In another embodiment, multiple short oligonucleotides encoding a portion of the desired polypeptide can be synthesized and then ligated. In some embodiments, the individual oligonucleotides typically include a 5' overhang or a 3' overhang for complementary assembly.
[0409] The polynucleotides (e.g., RNA, e.g., mRNA) disclosed herein can be chemically synthesized using chemical synthesis methods and possible nucleic acid base substitutions known in the art. See, for example, International Publication Nos. 2014093924, 2013052523, 2013039857, 2012135805, 2013151671, U.S. Patent Application Publication No. 20130115272, or U.S. Patent No. 8999380 or 8710200 (all of which are incorporated herein by reference in their entirety).
[0410] c. Purification of polynucleotides encoding PAH Purification of polynucleotides described herein (e.g., polynucleotides containing nucleotide sequences encoding variant PAH polypeptides) may include, but are not limited to, cleanup, quality assurance, and quality control of polynucleotides. Cleanup may be performed by methods known in the art, such as AGENCOURT® beads (Beckman Coulter Genomics, Danvers, MA), Poly-T beads, LNA® oligo-T capture probes (EXIQON® Inc., Vedbaek, Denmark), or HPLC-based purification methods (but not limited to) such as strong anion exchange HPLC, weak anion exchange HPLC, reverse-phase HPLC (RP-HPLC), and hydrophobic interaction HPLC (HIC-HPLC) (but not limited to these).
[0411] The term "purification," when used in relation to polynucleotides, as in "purified polynucleotides," refers to separation from at least one impurity. As used herein, "impurity" is any substance that makes the other substance unsuitable, impure, or inferior. That is, purified polynucleotides (e.g., DNA and RNA) exist in a form or state different from the form or state found in nature, or in a form or state different from the form or state that existed before the processing or purification method was performed.
[0412] In some embodiments, the purification of the polynucleotide of the present invention (e.g., a polynucleotide comprising a nucleotide sequence encoding a PAH polypeptide) removes impurities, which can reduce or eliminate undesirable immune responses, such as reducing cytokine activity.
[0413] In some embodiments, the polynucleotides of the present invention (e.g., polynucleotides comprising a nucleotide sequence encoding a variant PAH polypeptide) are purified before administration using column chromatography (e.g., strong anion exchange HPLC, weak anion exchange HPLC, reversed-phase HPLC (RP-HPLC), and hydrophobic interaction HPLC (HIC-HPLC) or (LCMS)).
[0414] In some embodiments, polynucleotides of the present invention (e.g., polynucleotides comprising the nucleotide sequence of a variant PAH polypeptide) purified using column chromatography (e.g., strong anion exchange HPLC, weak anion exchange HPLC, reversed-phase HPLC (RP-HPLC, hydrophobic interaction HPLC (HIC-HPLC), or (LCMS)) result in increased expression of the encoded PAH protein compared to expression levels obtained when the same polynucleotides of the present disclosure are purified by different purification methods.
[0415] In some embodiments, the polynucleotides purified by column chromatography (e.g., strong anion exchange HPLC, weak anion exchange HPLC, reversed-phase HPLC (RP-HPLC), hydrophobic interaction HPLC (HIC-HPLC), or (LCMS)) contain nucleotide sequences encoding PAH polypeptides that include one or more point mutations known in the art.
[0416] In some embodiments, when RP-HPLC purified polynucleotides are used, the intracellular PAH protein expression level of those cells increases by, for example, 10–100%, i.e., at least about 10%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 90%, at least about 95%, or at least about 100% when the polynucleotides are introduced into cells.
[0417] In some embodiments, when RP-HPLC purified polynucleotides are used, their intracellular functional PAH activity increases by, for example, 10–100%, i.e., at least about 10%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 90%, at least about 95%, or at least about 100% when the polynucleotides are introduced into cells.
[0418] In some embodiments, when RP-HPLC purified polynucleotides are used, when those polynucleotides are introduced into cells, the detectable PAH activity of those cells increases by, for example, 10–100%, i.e., at least about 10%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 90%, at least about 95%, or at least about 100% compared to the functional PAH activity level in cells before introducing RP-HPLC purified polynucleotides into the cells or after introducing non-RP-HPLC purified polynucleotides into the cells.
[0419] In some embodiments, the purified polynucleotide has a purity of at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100%.
[0420] Quality assurance and / or quality control checks can be performed using methods such as gel electrophoresis, UV absorbance, or analytical HPLC (but not limited to these). In another embodiment, polynucleotides can be sequenced by methods including (but not limited to) reverse transcriptase-PCR.
[0421] d. Quantification of expressed polynucleotides encoding PAH In some embodiments, the polynucleotides of the present invention (e.g., polynucleotides comprising a nucleotide sequence encoding a variant PAH polypeptide), their expression products, and degradation products and metabolites can be quantified according to methods known in the art.
[0422] In some embodiments, the polynucleotides of the present invention can be quantified in exosomes or when derived from one or more bodily fluids. As used herein, “bodily fluids” include peripheral blood, serum, plasma, ascites, urine, cerebrospinal fluid (CSF), sputum, saliva, bone marrow, synovial fluid, aqueous humor, amniotic fluid, earwax, breast milk, bronchoalveolar lavage fluid, semen, prostatic fluid, Cowper's fluid or preejaculatory fluid, sweat, feces, hair, tears, cystic fluid, pleural fluid and ascites, pericardial fluid, lymph, ooze, chyle, bile, interstitial fluid, menstrual blood, pus, sebum, vomit, vaginal secretions, mucosal secretions, fecal water, pancreatic juice, sinus lavage fluid, bronchopulmonary aspirate, blastocoel fluid, and umbilical cord blood. Alternatively, exosomes can be collected from organs selected from the group consisting of the lungs, heart, pancreas, stomach, intestines, bladder, kidneys, ovaries, testes, skin, colon, breasts, prostate, brain, esophagus, liver, and placenta.
[0423] In exosome quantification methods, a sample of 2 mL or less is collected from the target, and exosomes are isolated by size exclusion chromatography, density gradient centrifugation, fractional centrifugation, ultrafiltration with nanomembrane, capture by immunoadsorption, affinity purification, separation by microfluidics, or a combination thereof. During analysis, the level or concentration of polynucleotides can be the expression level, presence, truncation, or alteration of the administered construct. Correlation of these levels with assay results for one or more clinical phenotypes or biomarkers of human diseases is beneficial.
[0424] The assay can be performed using construct-specific probes, cytometry, qRT-PCR, real-time PCR, PCR, flow cytometry, electrophoresis, mass spectrometry, or a combination thereof, while exosomes can be isolated using immunohistochemical methods such as enzyme-linked immunosorbent assay (ELISA). Exosomes can also be isolated by size exclusion chromatography, density gradient centrifugation, fractional centrifugation, ultrafiltration with nanomembranes, capture by immunoadsorption, affinity purification, separation by microfluidics, or a combination thereof.
[0425] These methods give researchers the ability to monitor the levels of residual or delivered polynucleotides in real time. This is possible because the polynucleotides of the present invention differ from their endogenous forms due to structural or chemical modifications.
[0426] In some embodiments, the polynucleotides of the present invention can be quantified using methods such as ultraviolet-visible spectroscopy (UV / Vis) (but not limited to). An example of a UV / Vis spectrometer is the NANODROP® spectrometer (ThermoFisher, Waltham, MA). To determine whether the polynucleotides are of appropriate size, the quantified polynucleotides can be analyzed to confirm that no degradation of the polynucleotides has occurred. Degradation of polynucleotides can be confirmed by HPLC-based purification methods such as agarose gel electrophoresis, strong anion exchange HPLC, weak anion exchange HPLC, reversed-phase HPLC (RP-HPLC), and hydrophobic interaction HPLC (HIC-HPLC) (but not limited to these), as well as by methods such as liquid chromatography-mass spectrometry (LCMS), capillary electrophoresis (CE), and capillary gel electrophoresis (CGE) (but not limited to these).
[0427] DNA vectors for gene delivery Numerous native and recombinant vectors have been developed to deliver genes to cells, including nucleic acid-based gene delivery vectors derived from viruses and bacteria. Several approaches have been developed to deliver a target gene to cells for gene expression using DNA. This disclosure relates to compositions and methods for delivering DNA encoding the variant PAH polypeptide described herein to cells. In some embodiments, a DNA molecule is delivered to a cell to express a transcript encoded by that DNA (e.g., a protein-coding transcript, or a functional nucleic acid such as functional DNA). In some embodiments, a DNA molecule is delivered to a cell to repair or replace a native gene, for example, by recombination. In some embodiments, the DNA molecule acts as a transgene to assist the expression of a native gene, for example, a native gene that has reduced transcription levels or produces abnormal RNA or protein. In some embodiments, the DNA molecule delivered to the cell is functional DNA, i.e., the DNA can exert some biological activity other than simply encoding the mRNA of a protein. For example, the DNA molecule can be folded into a structure that can bind to other molecules and alter their activity; for instance, the DNA could be an aptamer that exerts therapeutic function.
[0428] For example, any DNA molecule capable of transferring a gene into a cell to express a transcript can be incorporated into a delivery vehicle as described herein, such as a lipid nanoparticle. In some embodiments, the DNA molecule may be of natural origin and can be isolated from a natural source, for example. In other embodiments, the DNA molecule is a synthetic molecule, such as a synthetic DNA molecule prepared in vitro. In some embodiments, the DNA molecule is a recombinant molecule.
[0429] The DNA molecule can be double-stranded DNA, single-stranded DNA, or a partially double-stranded DNA molecule, that is, a molecule having a double-stranded portion and a single-stranded portion. In some cases, the DNA molecule may be triple-stranded or partially triple-stranded, that is, having a triple-stranded portion and a double-stranded portion. The DNA molecule may be circular DNA or linear DNA.
[0430] The DNA sequences described herein, for example, DNA vectors, can have a wide variety of features. The DNA sequences described herein, for example, DNA vectors, can include non-coding DNA sequences. For example, a DNA sequence may include at least one gene regulatory element, such as a promoter, enhancer, stop element, polyadenylation signal element, splicing signal element, etc. In some embodiments, the non-coding DNA sequence is an intron. In some embodiments, the non-coding DNA sequence is a transposon. In some embodiments, the DNA sequences described herein may have a non-coding DNA sequence functionally linked to a transcriptionally active gene. In other embodiments, the DNA sequences described herein may have a non-coding DNA sequence that is not linked to a gene; that is, the non-coding DNA does not regulate a gene on the DNA sequence.
[0431] In some embodiments, the DNA sequence, for example, the DNA vector, has at least one transcriptionally active gene, i.e., a gene whose coding sequence can be expressed under intracellular conditions. In some embodiments, the DNA vector includes essential expression regulatory elements necessary for gene expression in the specific intracellular environment into which the DNA vector is introduced. That is, the DNA vectors described herein may include an expression module or expression cassette comprising at least one transcriptionally active gene functionally ligated to at least one transcriptional mediating or regulatory element, such as a promoter, enhancer, stop and polyadenylation signal element, or splicing signal element.
[0432] In some embodiments, the expression module or expression cassette includes transcriptional regulatory elements that express a gene in a wide range of hosts. Various combinations of these elements are known, and specific transcriptional regulatory elements include the SV40 element, as described in Dijkema et al., EMBO J. (1985) 4:761; the transcriptional regulatory element derived from the LTR of Rous sarcoma virus, as described in Gorman et al., Proc. Nat'l Acad. Sci USA (1982) 79:6777; the transcriptional regulatory element derived from the LTR of human cytomegalovirus (CMV), as described in Boshart et al., Cell (1985) 41:521; and the hsp70 promoter (Levy-Holtzman, R. and I. Scheccher (Biochim. Biophys. Acta (1995) 1263:96-98), Presnail, J.K and MAHoy (Exp. Appl. Acarol. (1994) 18:301-308)).
[0433] In some embodiments, at least one of the transcriptional activity genes of the DNA sequence encodes a protein or functional nucleic acid having therapeutic activity in a subject, such as a mammalian subject like a human, e.g., functional DNA. In some embodiments, at least one of the transcriptional activity genes of the DNA sequence encodes a eukaryotic protein or functional nucleic acid, e.g., functional DNA. In some embodiments, at least one of the transcriptional activity genes of the DNA sequence encodes a mammalian protein or functional nucleic acid, e.g., functional DNA. In some embodiments, at least one of the transcriptional activity genes of the DNA sequence encodes a human protein or functional nucleic acid, e.g., functional DNA.
[0434] In some embodiments, the DNA sequence, e.g., DNA vector, includes at least one non-coding sequence, e.g., a promoter, enhancer, stop element, polyadenylation signal element, splicing signal element, and / or an intron that can function as a template when homologous recombination occurs in the host cell genome. In some embodiments, homologous recombination of a non-coding sequence can be used to repair or restore a mutant transcription element in the host cell genome, e.g., a mutant transcription binding site that results in abnormal gene expression. In some embodiments, homologous recombination of a non-coding sequence can repair or restore a mutant transcription element in the host cell that causes a pathological phenotype.
[0435] In some embodiments, the DNA sequences described herein, such as DNA vectors, are triple-stranded DNA sequences induced by the cPPT and CTS regions of a lentivirus. In some embodiments, the DNA sequences described herein have cis-operating sequences cPPT and CTS of a lentivirus, such as HIV, resulting in a DNA sequence having a triple-stranded DNA structure. In some embodiments, the triple-stranded DNA structure induces a high rate of DNA entry into the nucleus of a host cell. In some embodiments, the triple-stranded DNA structure can increase the nuclear translocation rate of the DNA sequence. In some embodiments, the triple-stranded DNA structure can increase the amount of DNA sequence that translocates into the nucleus of a host cell.
[0436] In some embodiments, the triple-stranded DNA sequence can be covalently bound to a target nucleic acid sequence, such as a DNA vector, transgene, or non-coding sequence. In embodiments, the DNA sequence has two or more triple-stranded sequences induced by the cPPT and CTS regions of a lentivirus, for example, the DNA sequence may have two, three, four, five, six, seven, eight, nine, or ten or more triple-stranded sequences.
[0437] In some cases, the DNA sequence can be a double-stranded DNA molecule. In some embodiments, the entire DNA sequence is double-stranded. In some embodiments, only a portion of the DNA sequence is double-stranded. For example, the DNA sequence can be partially double-stranded and partially single-stranded, and can be, for example, a linear double-stranded DNA sequence having single-stranded overhangs at the 5' and / or 3' portions of the sequence. In some embodiments, the DNA sequence is single-stranded.
[0438] In some embodiments, the DNA sequence is a DNA vector. In some embodiments, the DNA vector is a plasmid. Methods for constructing and manipulating DNA plasmids are well known in the art, and there are numerous clinical trials of gene therapies using non-replicating, non-viral plasmid DNA to treat diseases (Hardee et al., Genes, 2017, 8, 65). In some embodiments, the plasmid is a bacterial plasmid. In some embodiments, the plasmid is superhelical. In some embodiments, the plasmid has an open circular topology. In some embodiments, the plasmid is linear.
[0439] In some embodiments, the DNA sequence, e.g., DNA vector, is modified to reduce or minimize the size or length of the molecule; for example, the DNA vector is a plasmid modified to reduce its size. For example, a portion of a bacterial plasmid may be removed to create a miniplasmid. In some embodiments, the DNA sequence, e.g., DNA plasmid, is modified to remove prokaryotic modifications that may induce an innate immune response or transgene silencing; for example, irrelevant portions of the plasmid sequence elements that do not encode the gene of interest can be removed. In some embodiments, removing irrelevant sequence elements from the DNA sequence, e.g., DNA plasmid, improves the safety of the DNA sequence in a mammalian host. In some embodiments, the DNA sequence, e.g., bacterial plasmid, is modified to reduce the number of CpG dinucleotides in the sequence. Unmethylated CpG dinucleotides are more common in bacterial DNA than in mammalian DNA and may induce transgene silencing and / or an immune response in mammalian subjects. In some embodiments, the bacterial plasmid is modified to remove all or part of the bacterial origin of replication (ori).
[0440] In some embodiments, DNA sequences, such as DNA vectors like bacterial plasmids, are modified to confer antibiotic resistance to bacteria and to remove genes that may induce an immune response in mammalian subjects. In some embodiments, the plasmid includes an antibiotic-free system for plasmid selection. For example, the plasmid may include an operator-repressor titration (ORT) system, such as those described in US5,972,708 and Cranenburgh et al., Nucleic Acids Res., 2001, 29, E26 (which is incorporated herein by reference in its entirety). An ORT plasmid (pORT) has an operator sequence that is used to titrate in bacteria via a competitive repressor protein that binds to an endogenous operator sequence upstream of an essential gene encoded on a chromosome. In some embodiments, the plasmid has a conditional origin of replication (COR), i.e., the plasmid is a pCOR plasmid, for example, as described in Sourbrier et al., Gene Therapy, 1999, 6:1482-1488 (the entire work is incorporated herein by reference). In some embodiments, the plasmid is an antibiotic-free plasmid (pFAR), as described in Marie et al., J. Gene Med., 2010, 12:323-332 (the entire work is incorporated herein by reference).
[0441] In some embodiments, the DNA sequence is a minicircle DNA vector, such as those described in, for example, Chen et al., Mol. Ther., 2003, 8(3):495-500, Chen et al., Gene Ther., 2005, 16(1):126-131, Chen et al., Nat. Biotech., 2010, 28(12):1289-1291, U.S. Patent No. 7,897,380, and U.S. / 2016 / 0312230 (which is incorporated herein by reference in its entirety). Minicircle DNA is a minimal circular double-stranded DNA vector, primarily superhelical in structure, containing the eukaryotic gene of interest, and either containing a very short prokaryotic DNA sequence segment or deleting prokaryotic DNA. In minicircle DNA vectors, the deletion of prokaryotic DNA makes the vector less likely to induce inflammation or silencing of gene expression when administered to a subject compared to viral or plasmid vectors. Minicircle DNA vectors also offer enhanced clinical safety because they do not contain bacterial resistance marker genes or origins of replication.
[0442] In some embodiments, the minicircle DNA can sustainably express an exogenous transgene for a longer period after administration to a subject compared to a control vector, such as a bacterial plasmid. In some embodiments, the minicircle DNA sustainably expresses the transgene for at least one week, at least two weeks, at least three weeks, at least one month, at least six weeks, at least two months, at least ten weeks, at least three months, at least four months, at least five months, or at least six months or longer after administration to a subject. In some embodiments, the minicircle DNA sustainably expresses the transgene for at least about twice as long as a control DNA vector after administration to a subject, for example, at least twice, at least three times, at least four times, at least five times, or at least ten times longer.
[0443] In some aspects of this disclosure, minicircle DNA is produced from a parent bacterial plasmid by site-directed recombination in a host cell. The desired transgene is flanked by recombination sites within the parent bacterial plasmid, and most or all of the sequences necessary for plasmid growth in the bacterium (including ori and selection markers (e.g., antibiotic resistance genes)) are positioned outside these recombination sites. The parent bacterial plasmid is then introduced into a host cell, such as an E. coli cell, and site-directed recombination is used to produce minicircles containing the transgene, from which bacterial sequences have been deleted from the parent bacterial plasmid, and miniplasmids containing most or all of the parent plasmid's bacterial DNA (including its ori and selection markers) can be discarded.
[0444] Several recombinase systems for producing minicircles, including wild-type and mutant phage integrases, have been described in the art. In some embodiments, the recombinase recognizes a specific recombination site on the parent plasmid that produces the minicircle DNA. In some embodiments, the recombinase may be, but is not limited to, phage λ integrase, phiC31 (ΦC31) recombinase, Flp recombinase, ParA resolverase, Cre recombinase, R4 integrase, TP901-1 integrase, A118 integrase, ΦFC1 integrase, etc. (see, e.g., Gaspar et al., Expert Opin. Biol. Ther., 2015, 15:353-379; Hardee et al., Genes, 2017, 8, 65; US / 2016 / 0312230). In some embodiments, the site-specific recombination site is the site-specific recombination site of PhiC31. In some embodiments, the site-specific recombination site is a site-specific recombination site of ParA. In some embodiments, the site-specific recombination site is a site-specific recombination site of Cre. In some embodiments, the recombinase is a unidirectional site-specific recombinase.
[0445] In some embodiments, the site-specific recombination sites for producing minicircle DNA are attB and attP, such that the minicircle DNA in the parent bacterial plasmid is flanked by the attB and attP sites. The phiC31 recombinase recognizes these sites and induces recombination that generates the minicircle DNA vector and miniplasmid.
[0446] In some embodiments, minicircle DNA is produced in a microorganism that can amplify the parent vector used to produce minicircle DNA, and can also produce minicircle DNA once the recombinase is expressed. In some embodiments, the microorganism is a bacterium such as Escherichia sp, particularly E. coli, for example, strain ZYCY10P3S2T. In one embodiment, the microorganism that produces minicircle DNA endogenously expresses the recombinase. Alternatively, the microorganism that produces minicircle DNA can be introduced and expressed by introducing the recombinase or a gene encoding the recombinase.
[0447] In some embodiments, during the production of minicircle DNA, the bacterial plasmid DNA introduced into the miniplasmid contains at least one DNA endonuclease site, such as an I-SceI endonuclease site, and site-specific recombination allows the miniplasmid to be degraded by its DNA endonuclease in the host cell once the minicircle DNA has been produced. This makes it even easier to purify the minicircle DNA from the host cell.
[0448] In some embodiments, the DNA molecule, for example, a minicircle, is purified using methods known in the art before being incorporated into a delivery vehicle such as lipid nanoparticles. See Hardee's discussion.
[0449] In some embodiments, the DNA vector is a “minivector” or “microminicircle,” as described, for example, in Hardee et al., Genes, 2017, 8, 65 and Stenler et al., Mol. Ther. Nucleic Acids, 2014, 2:e140 (which is incorporated herein by reference in its entirety). Minivectors are generally smaller than minicircles and encode regulatory RNA, such as shRNA. Generally, the method for constructing minivectors is similar to or identical to the method for constructing minicircles. In some embodiments, the DNA vector is a superhelical minivector, as described in US / 2014 / 0056868 (which is incorporated herein by reference in its entirety).
[0450] In some embodiments, the DNA sequence is the covalent closed circular DNA (cccDNA) of hepatitis B virus (HBV), for example, recombinant HBV cccDNA as described in US / 2017 / 0327797 and Li et al., 2018, Hepatology, 67(1):56-70 (which is incorporated herein by reference in its entirety). HBV is a partially double-stranded DNA virus that can infect human liver cells. Upon infection, cccDNA is formed and maintained in the nucleus of infected cells, where it persists as a stable episome and functions as a template for the transcription of its viral genes. Elimination of intracellular cccDNA is a major obstacle in current treatments for chronic HBV infection, and there is a need for new therapies that directly target cccDNA. However, the discovery of anti-HBV drugs has been hampered by the lack of physiologically relevant in vitro and in vivo models, as existing models have been found to be difficult or inconvenient to use. Developing mouse models for chronic HBV is difficult, and there is a need for cccDNA-based mouse models that can be used in drug discovery for anti-HBV drugs, particularly immunonormal mouse models that can handle the sustained replication of HBV driven by cccDNA.
[0451] In some embodiments, recombinant HBV cccDNA is produced by using known methods for producing minicircle vectors, such as those described in US / 2017 / 0327797. In the parent vector producing minicircle DNA, the full-length HBV genome or a portion thereof is sandwiched between recombination sites, such as attP and attB sites, and recombinant HBV cccDNA is produced by site-specific recombination using recombinases such as phage integrases ΦC31, R4, TP901-1, ΦBT1, Bxb1, RV-1, AA118, U153, and ΦFC1.
[0452] In some embodiments, the DNA vectors described herein may be linear DNA, i.e., the DNA has two defined ends and is not circular.
[0453] In some embodiments, the DNA vector is a double-stranded linear DNA, i.e., linear double-stranded DNA. Any linear double-stranded DNA can be incorporated into the DNA delivery system described herein, for example, lipid nanoparticles. In some embodiments, the DNA vector is a native linear double-stranded DNA, e.g., DNA derived from a bacteriophage or virus, or derived from a native linear double-stranded DNA. In some embodiments, the DNA vector is a synthetic linear double-stranded DNA, e.g., recombinant bacteriophage DNA, recombinant viral DNA, PCR product, DNA fragment prepared by restriction digestion with an endonuclease, or a closed linear DNA molecule, or derived from a synthetic linear double-stranded DNA.
[0454] In some embodiments, the DNA molecule may be a closed-end linear double-stranded DNA ("ceDNA" or "CELiD DNA") as described in WO2017 / 152149 and Li et al., PLoS One, 2013 8(8):e69879 (which is incorporated herein by reference in its entirety). The ceDNA has at least one transgene flanked by asymmetric terminal sequences, e.g., fragmented asymmetric self-complementary sequences, thereby covalently bonded to the asymmetric terminal sequences. In some embodiments, the ceDNA consists of two asymmetric terminal sequences, e.g., a transgene flanked by two fragmented asymmetric self-complementary sequences. In some embodiments, the transgene is flanked by asymmetric terminal sequences at its 5' and 3' ends, respectively. Structurally, ceDNA is a double-stranded linear DNA with nourishing closed ends.
[0455] In some cases, a “disrupted self-complementary sequence” can be a polynucleotide sequence encoding a nucleic acid having a palindromic terminal sequence that is interrupted by one or more non-palindromic polynucleotide regions. Typically, a polynucleotide encoding one or more interrupted palindromic sequences folds in two to form a stem-loop structure. In some embodiments, each self-complementary sequence has a functional terminal splitting site and a rolling-circle replication protein-binding element. In some embodiments, the self-complementary sequence is interrupted by cross-arm sequences that form two longitudinally symmetrical opposing stem-loops, each of which has a stem portion ranging from 5 to 15 base pairs in length and a loop portion having 2 to 5 unpaired deoxyribonucleotides. In some embodiments, the disrupted self-complementary sequence can include three or more cross-arm sequences, e.g., three, four, five, six, seven, eight, nine, or ten or more cross-arm sequences.
[0456] In some embodiments, the fragmented self-complementary sequence is derived from one or more viruses or viral serotypes. In some embodiments, the fragmented self-complementary sequence is derived from a parvovirus. In some embodiments, the fragmented self-complementary sequence is derived from a dependvirus. In some embodiments, the fragmented self-complementary sequence is derived from an adeno-associated virus. In some embodiments, the fragmented self-complementary sequence is derived from an AAV2 serotype. In some embodiments, the fragmented self-complementary sequence is derived from an AAV9 serotype. In some embodiments, the first fragmented self-complementary sequence and the second fragmented self-complementary sequence are derived from the same virus or viral serotype. In some embodiments, the first fragmented self-complementary sequence is derived from a first virus or viral serotype, and the second fragmented self-complementary sequence is derived from a second virus or viral serotype. In some embodiments, the fragmented self-complementary sequences are of different lengths.
[0457] In some embodiments, the fragmented self-complementary sequence is an AAV inverted terminal repeat (ITR) sequence. An AAV ITR sequence can be a sequence of any AAV serotype, including, but not limited to, AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, a non-human primate AAV serotype (e.g., AAVrh.1O), and their variants. In some embodiments, the fragmented self-complementary sequence is an AAV2 ITR or a variant thereof. In some embodiments, the fragmented self-complementary sequence is an AAV5 ITR or a variant thereof. As used herein, a “variant” of an AAV ITR is a polynucleotide having similarity of about 70% to about 99.9% to a wild-type AAV ITR sequence. In some embodiments, the AAV ITR variant is approximately 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to the wild-type AAV ITR. In some embodiments, the AAV ITR variant is a truncated AAV ITR or an AAV ITR with a deletion. In some embodiments, the asymmetric terminal sequence of the ceDNA may be a reverse terminal repeat sequence derived from AAV. In some embodiments, the asymmetric terminal sequence of the ceDNA may be an ITR derived from adeno-associated virus type 2.
[0458] In some embodiments, the DNA sequence is single-stranded linear DNA. Any single-stranded linear DNA can be incorporated into the DNA delivery system described herein, for example, lipid nanoparticles. In some embodiments, the DNA sequence is or is derived from natural single-stranded linear DNA, such as DNA from a bacteriophage or virus. For example, the single-stranded linear DNA may be adeno-associated virus (AAV) DNA, for example, one or more genes derived from adeno-associated virus. AAV has been widely used in gene therapy applications and is well known in the literature.
[0459] In some embodiments, the single-stranded linear DNA can be the DNA of an oncolytic virus. Because oncolytic viruses exhibit selectivity to replicate only in cancer cells, they can be used to infect and kill cancer cells and tumors while minimizing damage to non-cancerous cells and tissues. In some embodiments, the DNA sequence can be all or part of the genome of an oncolytic virus. The oncolytic virus DNA can be incorporated into DNA delivery systems described herein, such as lipid nanoparticles. In some embodiments, the genes of the oncolytic virus DNA can be modified to improve cancer-selective replication, cell lysis, and / or the spread of progeny viruses to the vicinity of cancerous cells, as described, for example, in Seymour and Fisher, Br.J. Cancer, 2016, 114(4):357-361 (which is incorporated herein by reference in its entirety). As just one example, in some embodiments, oncolytic viruses that utilize the host cell's transcription mechanism for replication can be manipulated to promote viral replication by regulating the expression of essential viral genes in a tumor-associated transcription factor-dependent manner, for example, using a tumor-associated promoter. In some embodiments, the genes of oncolytic viral DNA can be modified to encode a “armed” oncolytic virus, so that the oncolytic viral DNA also encodes a transgene encoding an anticancer agent that can be selectively expressed in cancer cells (see Seymour and Fisher, Br.J. Cancer, 2016, 114(4):357-361). In some embodiments, the anticancer agent may be a therapeutic protein or therapeutic nucleic acid. In some embodiments, the therapeutic protein may be a cytokine, chemokine, enzyme, or antibody. In some embodiments, the therapeutic nucleic acid may be mRNA or siRNA.
[0460] In some embodiments, the oncolytic virus is a parvovirus. Parvoviruses are single-stranded DNA viruses that are lytic viruses, meaning they can lyse infected cells. Parvoviruses depend on cellular factors of the host cell that are expressed during the S phase of the cell cycle for viral replication. In some embodiments, parvoviruses can infect cancer cells and kill them, while leaving non-cancerous cells intact or causing minimal damage to them. In some embodiments, the DNA sequence is a parvovirus or derived from a parvovirus, for example, one or more genes derived from a parvovirus, as described in US7,179,456 and EP2579885. In some embodiments, the parvovirus is parvovirus H1, LuIII, mouse microvirus (MMV), mouse parvovirus (MPV), rat microvirus (RMV), rat parvovirus (RPV), and rat virus.
[0461] In some embodiments, the DNA sequence, such as a DNA vector, is not integrated into the genome of the target cells upon administration to the subject; that is, the DNA does not fuse with or covalently bond with the chromosomes within the target cells of the subject. Rather, the DNA sequence is maintained episomal. In some embodiments, the episomal DNA sequence is continuously expressed in the subject. In some embodiments, the episomal DNA sequence is transiently expressed in the subject.
[0462] In some embodiments, the DNA sequence (e.g., a DNA vector) or a portion of the DNA sequence, upon administration to a subject, is integrated into the genome of the target cell, i.e., the DNA fuses or ferries with chromosomes within the target cell of the subject. In some embodiments, the DNA sequence is integrated into the genome of the subject by homologous recombination. In some embodiments, the DNA sequence, e.g., a DNA vector, comprises a coding sequence (e.g., a transgene) and / or a non-coding sequence (e.g., a transcriptional regulatory element), and is integrated into the genome of the target host cell in the subject by homologous recombination. In some embodiments, the DNA sequence comprises a sequence homologous to the DNA of the target cell in the subject so as to promote homologous recombination of at least a portion of the DNA sequence, and by homologous recombination, the coding sequence (e.g., a transgene) and / or a non-coding sequence (e.g., a transcriptional regulatory element) is integrated into the genome of the target cell. In some embodiments, once the DNA sequence is integrated into the genome of the subject, the transgene encoded by the DNA sequence is continuously expressed. In some embodiments, once the DNA sequence is integrated into the genome of the subject, the transgene encoded by the DNA sequence is transiently expressed.
[0463] In some embodiments, when the DNA sequence is incorporated into the target genome, gene expression is restored in the target, for example, to a level close to the wild-type expression level. In some embodiments, when the DNA sequence is incorporated into the target genome, mutant genes in the target are repaired or replaced. In some embodiments, when the DNA sequence is incorporated into the target genome, mutant transcriptional regulatory elements in the target are repaired or replaced. In some embodiments, when the DNA sequence is incorporated into the target genome, pathological phenotypes in the target are treated or prevented.
[0464] Pharmaceutical compositions and preparations The present invention provides pharmaceutical compositions and formulations comprising any of the above-mentioned polynucleotides. In some embodiments, the composition or formulation further comprises a delivery agent.
[0465] In some embodiments, the composition or formulation may include a polynucleotide comprising a sequence-optimized nucleic acid sequence disclosed herein, which encodes a variant PAH polypeptide. In some embodiments, the composition or formulation may include a polynucleotide (e.g., RNA, e.g., mRNA) comprising a sequence-optimized nucleic acid sequence disclosed herein, which has a fairly high degree of sequence identity with the nucleic acid sequence encoding the PAH polypeptide (e.g., ORF). In some embodiments, the polynucleotide may further include a miRNA-binding site that binds to miRNA-binding sites, e.g., miR-126, miR-142, miR-144, miR-146, miR-150, miR-155, miR-16, miR-21, miR-223, miR-24, miR-27, and miR-26a.
[0466] The pharmaceutical composition or formulation may optionally contain one or more additional active substances, such as therapeutically and / or prophylactically active substances. The pharmaceutical composition or formulation of the present invention may be sterile and / or pyrogen-free. General considerations for the preparation and / or manufacture of pharmaceutical formulations are, for example, Remington: The Science and Practice of Pharmacy 21. st See, ed., Lippincott Williams & Wilkins, 2005 (the entire work is incorporated herein by reference). In some embodiments, the composition is administered to a human, a human patient, or a subject. For the purposes of this disclosure, the term “active ingredient” generally refers to the polynucleotide delivered as described herein.
[0467] The formulations and pharmaceutical compositions described herein can be prepared by any method known or to be developed in the field of pharmacology. Generally, such preparation methods include the steps of combining an active ingredient with excipients and / or one or more other auxiliary ingredients, and then, if necessary and / or desirable, dividing, shaping, and / or packaging the product into desired single-dose or multi-dose units.
[0468] The pharmaceutical compositions or formulations described herein may be prepared, packaged and / or sold in bulk, as single doses and / or as multiple single doses. As used herein, “unit dose” means an individual amount of a pharmaceutical composition containing a predetermined amount of the active ingredient. The amount of the active ingredient is approximately equal to the dose of the active ingredient to be administered to the subject and / or a convenient division of that dose, such as half or one-third of such a dose.
[0469] The relative amounts of the active ingredient, pharmaceutically acceptable excipients, and / or any additional ingredients in the pharmaceutical compositions according to this disclosure will vary depending on the individual, physique, and / or condition of the person being treated, as well as the route by which the composition is administered.
[0470] In some embodiments, the compositions and formulations described herein may contain at least one polynucleotide of the present invention. In non-limiting examples, the composition or formulation may contain one, two, three, four, or five polynucleotides of the present invention. In some embodiments, the compositions or formulations described herein may contain two or more types of polynucleotides. In some embodiments, the composition or formulation may contain linear and cyclic polynucleotides. In another embodiment, the composition or formulation may contain cyclic polynucleotides and in vitro transcribed (IVT) polynucleotides. In yet another embodiment, the composition or formulation may contain IVT polynucleotides, chimeric polynucleotides, and cyclic polynucleotides.
[0471] The descriptions of pharmaceutical compositions and formulations provided in this invention primarily pertain to pharmaceutical compositions and formulations suitable for administration to humans. However, those skilled in the art will understand that such compositions are generally suitable for administration to other animals, such as non-human animals, such as non-human mammals.
[0472] The present invention provides pharmaceutical formulations comprising polynucleotides described herein (e.g., polynucleotides comprising a nucleotide sequence encoding a variant PAH polypeptide). The polynucleotides described herein can be formulated with one or more excipients to (1) enhance stability, (2) increase transfection into cells, (3) enable sustained or delayed release (e.g., from a depot formulation of the polynucleotide), (4) alter in vivo distribution (e.g., to direct the polynucleotide to a given type of tissue or cell), (5) increase translation of the encoded protein in vivo, and / or (6) alter the release profile of the encoded protein in vivo. In some embodiments, the pharmaceutical formulation further comprises a delivery agent comprising, for example, a compound having formula (I), e.g., any of compounds 1 to 232, e.g., compound II, a compound having formula (III), (IV), (V), or (VI), e.g., any of compounds 233 to 342, e.g., compound VI, or a compound having formula (VIII), e.g., any of compounds 419 to 428, e.g., compound I, or a combination thereof. In some embodiments, the delivery agent comprises compound II, DSPC, cholesterol, and compound I or PEG-DMG in a molar ratio of, for example, about 50:10:38.5:1.5. In some embodiments, the delivery agent comprises compound II, DSPC, cholesterol, and compound I or PEG-DMG in a molar ratio of, for example, about 47.5:10.5:39.0:3.0. In some embodiments, the delivery agent comprises compound VI, DSPC, cholesterol, and compound I or PEG-DMG in a molar ratio of, for example, about 50:10:38.5:1.5. In some embodiments, the delivery agent comprises compound VI, DSPC, cholesterol, and compound I or PEG-DMG in a molar ratio of, for example, about 47.5:10.5:39.0:3.0.
[0473] Pharmaceutically acceptable excipients, as used herein, include, but are not limited to, any solvent, dispersion medium or other liquid vehicle suitable for a particular desired dosage form, dispersing aids, suspension aids, diluents, granulating agents and / or dispersants, surfactants, isotonic agents, thickeners or emulsifiers, preservatives, binders, lubricants or oils, colorants, sweeteners or flavoring agents, stabilizers, antioxidants, antimicrobial or antifungal agents, osmotic regulators, pH adjusters, buffers, chelating agents, cryoprotective agents and / or volume extenders. Various excipients for compounding pharmaceutical compositions and techniques for preparing such compositions are known in the art (see Remington: The Science and Practice of Pharmacy, 21st Edition, ARGennaro (Lippincott, Williams & Wilkins, Baltimore, MD, 2006 (which is incorporated herein by reference)).
[0474] Examples of diluents include, but are not limited to, calcium carbonate, sodium carbonate, calcium phosphate, calcium hydrogen phosphate, sodium phosphate, lactose, sucrose, cellulose, microcrystalline cellulose, kaolin, mannitol, sorbitol, and / or combinations thereof.
[0475] Examples of granulating and / or dispersing agents include, but are not limited to, starch, pregelatinized starch or microcrystalline starch, alginic acid, guar gum, agar, poly(vinylpyrrolidone)(povidone), cross-linked poly(vinylpyrrolidone)(crospovidone), cellulose, methylcellulose, carboxymethylcellulose, cross-linked sodium carboxymethylcellulose (croscarmellose), magnesium aluminum silicate (VEEGUM®), sodium lauryl sulfate, and / or combinations thereof.
[0476] Examples of surfactants and / or emulsifiers include, but are not limited to, natural emulsifiers (e.g., acacia, agar, alginic acid, sodium alginate, tragacanth, red algae extracts, cholesterol, xanthan gum, pectin, gelatin, egg yolk, casein, lanolin, cholesterol, wax, and lecithin), sorbitan fatty acid esters (e.g., polyoxyethylene sorbitan monooleate [TWEEN® 80], sorbitan monopalmitate [SPAN® 40], glyceryl monooleate, polyoxyethylene esters, polyethylene glycol fatty acid esters (e.g., CREMOPHOR®), polyoxyethylene ethers (e.g., polyoxyethylene lauryl ether [BRIJ® 30]), PLUORINC® F68, POLOXAMER® 188, and / or combinations thereof).
[0477] Examples of binders include, but are not limited to, starch, gelatin, sugars (e.g., sucrose, glucose, dextrose, dextrin, molasses, lactose, lactitol, mannitol), amino acids (e.g., glycine), natural and synthetic gums (e.g., acacia, sodium alginate), ethylcellulose, hydroxyethylcellulose, hydroxypropylmethylcellulose, and combinations thereof.
[0478] Oxidation is a possible degradation pathway for mRNA, particularly liquid mRNA preparations. Antioxidants can be added to the preparation to prevent oxidation. Exemplary antioxidants include, but are not limited to, α-tocopherol, ascorbic acid, ascorbyl palmitate, benzyl alcohol, butylated hydroxyanisole, m-cresol, methionine, butylated hydroxytoluene, monothioglycerol, sodium metabisulfite or potassium metabisulfite, propionic acid, propyl gallate, sodium ascorbate, and combinations thereof.
[0479] Examples of chelating agents include, but are not limited to, ethylenediaminetetraacetic acid (EDTA), citric acid monohydrate, disodium edetate, fumaric acid, malic acid, phosphoric acid, sodium edetate, tartaric acid, trisodium edetate, and combinations thereof.
[0480] Examples of antimicrobial or antifungal agents include, but are not limited to, benzalkonium chloride, benzethonium chloride, methylparaben, ethylparaben, propylparaben, butylparaben, benzoic acid, hydroxybenzoic acid, potassium benzoate or sodium benzoate, potassium sorbate or sodium sorbate, sodium propionate, sorbic acid, and combinations thereof.
[0481] Examples of preservatives include, but are not limited to, vitamin A, vitamin C, vitamin E, beta-carotene, citric acid, ascorbic acid, butylated hydroxyanisole, ethylenediamine, sodium lauryl sulfate (SLS), sodium lauryl ether sulfate (SLES), and combinations thereof.
[0482] In some embodiments, the pH of the polynucleotide solution is maintained between pH 5 and pH 8 to enhance stability. Exemplary buffers for controlling pH include, but are not limited to, sodium phosphate, sodium citrate, sodium succinate, histidine (or histidine-HCl), sodium succinate, sodium carbonate, and / or combinations thereof.
[0483] Examples of lubricants include, but are not limited to, magnesium stearate, calcium stearate, stearic acid, silica, talc, malt, hydrogenated vegetable oil, polyethylene glycol, sodium benzoate, sodium lauryl sulfate, or magnesium lauryl sulfate, and combinations thereof.
[0484] The pharmaceutical compositions or formulations described herein may include cryoprotective agents to stabilize the polynucleotides described herein during freezing. Examples of cryoprotective agents include, but are not limited to, mannitol, sucrose, trehalose, lactose, glycerol, dextrose, and combinations thereof.
[0485] The pharmaceutical compositions or formulations described herein may include a bulking agent in the lyophilized polynucleotide formulation to provide a "pharmaceutically refined" cake, the bulking agent stabilizing the lyophilized polynucleotide during long-term storage (e.g., 36 months). Exemplary bulking agents of the present invention include, but are not limited to, sucrose, trehalose, mannitol, glycine, lactose, raffinose, and combinations thereof.
[0486] In some embodiments, the pharmaceutical composition or formulation further comprises a delivery agent. Examples of delivery agents of this disclosure include, but are not limited to, liposomes, lipid nanoparticles, lipidoids, polymers, lipoplexes, microvesicles, exosomes, peptides, proteins, cells transfected with polynucleotides, hyaluronidases, nanoparticle mimetic bodies, nanotubes, conjugates, and combinations thereof.
[0487] Delivery agent a. Lipid compounds This disclosure provides pharmaceutical compositions having beneficial properties. The lipid compositions described herein may be usefully used in lipid nanoparticle compositions for delivering therapeutic and / or prophylactic agents, such as mRNA, to mammalian cells or organs. For example, the lipids described herein exhibit little to no immunogenicity. For example, the lipid compounds disclosed herein have lower immunogenicity compared to reference lipids (e.g., MC3, KC2, or DLinDMA). For example, formulations containing the lipids disclosed herein and a therapeutic or prophylactic agent, such as mRNA, exhibit an increased therapeutic index compared to corresponding formulations containing a reference lipid (e.g., MC3, KC2, or DLinDMA) and the same therapeutic or prophylactic agent.
[0488] In certain embodiments, this disclosure is, (a) A polynucleotide comprising a nucleotide sequence encoding a variant PAH polypeptide, (b) Delivery agent and, The present invention provides a pharmaceutical composition containing the following:
[0489] Lipid nanoparticle formulations In some embodiments, the nucleic acid of the present invention (e.g., variant PAH mRNA) is formulated in lipid nanoparticles (LNPs). The lipid nanoparticles typically contain, along with the nucleic acid cargo of interest, an ionizable cationic lipid component, a non-cationic lipid component, a sterol component, and a PEG lipid component. The lipid nanoparticles of the present invention can be prepared using components, compositions, and methods generally known in the art. For example, PCT / US2016 / 052352, PCT / US2016 / 068300, PCT / US2017 / 037551, PCT / US2015 / 027400, PCT / US2016 / 047406, PCT / US2016000129, PCT / US2016 / 014280, PCT / US2016 / 014280, PCT / US2017 / 038426, See PCT / US2014 / 027077, PCT / US2014 / 055394, PCT / US2016 / 52117, PCT / US2012 / 069610, PCT / US2017 / 027492, PCT / US2016 / 059575 and PCT / US2016 / 069491 (all of which are incorporated herein by reference in their entirety).
[0490] The nucleic acids of this disclosure (e.g., variant PAH mRNA) are typically formulated in lipid nanoparticles. In some embodiments, the lipid nanoparticles comprise at least one ionizable cationic lipid, at least one non-cationic lipid, at least one sterol, and / or at least one polyethylene glycol (PEG)-modified lipid.
[0491] In some embodiments, the lipid nanoparticles contain ionizable cationic lipids in a molar ratio of 20-60%. For example, the lipid nanoparticles may contain ionizable cationic lipids in molar ratios of 20-50%, 20-40%, 20-30%, 30-60%, 30-50%, 30-40%, 40-60%, 40-50%, or 50-60%. In some embodiments, the lipid nanoparticles contain ionizable cationic lipids in molar ratios of 20%, 30%, 40%, 50%, or 60%.
[0492] In some embodiments, the lipid nanoparticles contain noncationic lipids in a molar ratio of 5 to 25%. For example, the lipid nanoparticles may contain noncationic lipids in molar ratios of 5 to 20%, 5 to 15%, 5 to 10%, 10 to 25%, 10 to 20%, 10 to 25%, 15 to 25%, 15 to 20%, or 20 to 25%. In some embodiments, the lipid nanoparticles contain noncationic lipids in molar ratios of 5%, 10%, 15%, 20%, or 25%.
[0493] In some embodiments, the lipid nanoparticles contain sterols in a molar ratio of 25-55%. For example, the lipid nanoparticles may contain sterols in molar ratios of 25-50%, 25-45%, 25-40%, 25-35%, 25-30%, 30-55%, 30-50%, 30-45%, 30-40%, 30-35%, 35-55%, 35-50%, 35-45%, 35-40%, 40-55%, 40-50%, 40-45%, 45-55%, 45-50%, or 50-55%. In some embodiments, the lipid nanoparticles contain sterols in molar ratios of 25%, 30%, 35%, 40%, 45%, 50%, or 55%.
[0494] In some embodiments, the lipid nanoparticles contain PEG-modified lipids in a molar ratio of 0.5 to 15%. For example, the lipid nanoparticles may contain PEG-modified lipids in molar ratios of 0.5 to 10%, 0.5 to 5%, 1 to 15%, 1 to 10%, 1 to 5%, 2 to 15%, 2 to 10%, 2 to 5%, 5 to 15%, 5 to 10%, or 10 to 15%. In some embodiments, the lipid nanoparticles contain PEG-modified lipids in molar ratios of 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, or 15%.
[0495] In some embodiments, the lipid nanoparticles contain ionizable cationic lipids in a molar ratio of 20-60%, non-cationic lipids in a molar ratio of 5-25%, sterols in a molar ratio of 25-55%, and PEG-modified lipids in a molar ratio of 0.5-15%.
[0496] Ionizable lipids In some embodiments, the ionizable lipids of this disclosure are compounds of the following formula (I), [ka] Or it may contain one or more of its N-oxide, salts or isomers, in the formula, R1 is C 5-30 Alkyl, C 5-20 Alkenil, -R * The group is selected from the group consisting of YR'', -YR'', and -R''M'R'. R2 and R3 are independent of H and C 1-14 Alkyl, C 2-14 Alkenil, -R * YR'', -YR'' and -R * R2 and R3 are selected from the group consisting of OR, or R2 and R3, together with the atoms to which they are bonded, form a heterocycle or a carbon ring. R4 is hydrogen, C 3-6 Carbon ring, -(CH2) n Q, -(CH2) n CHQR, -CHQR, -CQ(R)2 and unsubstituted C 1-6 It is selected from the group consisting of alkyl groups, and its Q is a carbocyclic, heterocyclic, -OR, or -O(CH2) n N(R)2, -C(O)OR, -OC(O)R, -CX3, -CX2H, -CXH2, -CN, -N(R)2, -C(O)N(R)2, -N(R)C(O)R , -N(R)S(O)2R, -N(R)C(O)N(R)2, -N(R)C(S)N(R)2, -N(R)R8, -N(R)S(O)2R8, -O(CH2) nThe values are selected from OR, -N(R)C(=NR9)N(R)2, -N(R)C(=CHR9)N(R)2, -OC(O)N(R)2, -N(R)C(O)OR, -N(OR)C(O)R, -N(OR)S(O)2R, -N(OR)C(O)OR, -N(OR)C(O)N(R)2, -N(OR)C(S)N(R)2, -N(OR)C(=NR9)N(R)2, -N(OR)C(=CHR9)N(R)2, -C(=NR9)N(R)2, -C(=NR9)R, -C(O)N(R)OR, and -C(R)N(R)2C(O)OR, where n is independently selected from 1, 2, 3, 4, and 5. Each R5 operates independently, C 1-3 Alkyl, C 2-3 It is selected from the group consisting of alkenyls and H. Each R6 operates independently, C 1-3 Alkyl, C 2-3 It is selected from the group consisting of alkenyls and H. M and M' are independently selected from -C(O)O-, -OC(O)-, -OC(O)-M''-C(O)O-, -C(O)N(R')-, -N(R')C(O)-, -C(O)-, -C(S)-, -C(S)S-, -SC(S)-, -CH(OH)-, -P(O)(OR')O-, -S(O)2-, -SS-, aryl groups and heteroaryl groups, and their M'' is bonded, C 1-13 Alkyl or C 2-13 It is alkenyl, R7 is C 1-3 Alkyl, C 2-3 It is selected from the group consisting of alkenyls and H. R8 is C 3-6 Selected from the group consisting of carbocyclic and heterocyclic rings, R9 is H, CN, NO2, C 1-6 Alkyl, -OR, -S(O)2R, -S(O)2N(R)2, C 2-6 Alkenil, C 3-6 Selected from the group consisting of carbocyclic and heterocyclic rings, Each R is independent of C 1-3 Alkyl, C 2-3 It is selected from the group consisting of alkenyls and H. Each R' is independent of C 1-18 Alkyl, C 2-18 Alkenil, -R * The group is selected from YR'', -YR'' and H. Each R' is independent of C 3-15 Alkyl and C 3-15 It is selected from a group consisting of alkenyls. Each R * C 1-12 Alkyl and C 2-12 It is selected from a group consisting of alkenyls. Each Y is independent of C 3-6 It is a carbon ring, Each X is independently selected from the group consisting of F, Cl, Br, and I. m is selected from 5, 6, 7, 8, 9, 10, 11, 12 and 13, and R4 is -(CH2) n Q, -(CH2) n If it is CHQR, -CHQR, or -CQ(R)2, then (i) when n is 1, 2, 3, 4, or 5, Q is not -N(R)2, or (ii) when n is 1 or 2, Q is not a 5, 6, or 7-membered heterocycloalkyl.
[0497] In a particular embodiment, a subset of compounds of formula (I) is the compound of formula (IA) below, [ka] Alternatively, it may include its N-oxide, or its salt or isomer, where l is selected from 1, 2, 3, 4, and 5, m is selected from 5, 6, 7, 8, and 9, M1 is a bond or M', and R4 is hydrogen, unsubstituted C 1-3 Alkyl or -(CH2) nQ is OH, -NHC(S)N(R)2, -NHC(O)N(R)2, -N(R)C(O)R, -N(R)S(O)2R, -N(R)R8, -NHC(=NR9)N(R)2, -NHC(=CHR9)N(R)2, -OC(O)N(R)2, -N(R)C(O)OR, heteroaryl or heterocycloalkyl, M and M' are independently selected from -C(O)O-, -OC(O)-, -OC(O)-M''-C(O)O-, -C(O)N(R')-, -P(O)(OR')O-, -SS-, aryl group and heteroaryl group, and R2 and R3 are independently H, C 1-14 Alkyl and C 2-14 It is selected from the group consisting of alkenyls. For example, m is 5, 7, or 9. For example, Q is OH, -NHC(S)N(R)2, or -NHC(O)N(R)2. For example, Q is -N(R)C(O)R, or -N(R)S(O)2R.
[0498] In a particular embodiment, a subset of compounds of formula (I) is the compound of formula (IB) below, [ka] Alternatively, it may include its N-oxide, or its salt or isomer, where all variable groups are as defined herein. For example, m is selected from 5, 6, 7, 8, and 9, and R4 is hydrogen, unsubstituted C 1-3 Alkyl or -(CH2) n Q is OH, -NHC(S)N(R)2, -NHC(O)N(R)2, -N(R)C(O)R, -N(R)S(O)2R, -N(R)R8, -NHC(=NR9)N(R)2, -NHC(=CHR9)N(R)2, -OC(O)N(R)2, -N(R)C(O)OR, heteroaryl or heterocycloalkyl, M and M' are independently selected from -C(O)O-, -OC(O)-, -OC(O)-M''-C(O)O-, -C(O)N(R')-, -P(O)(OR')O-, -SS-, aryl group and heteroaryl group, and R2 and R3 are independently H, C 1-14 Alkyl and C 2-14It is selected from a group consisting of alkenils.
[0499] For example, m is 5, 7, or 9. For example, Q is OH, -NHC(S)N(R)2, or -NHC(O)N(R)2. For example, Q is -N(R)C(O)R, or -N(R)S(O)2R.
[0500] In a particular embodiment, a subset of compounds of formula (I) is the compound of formula (II) below, [ka] Alternatively, it comprises its N-oxide, or its salt or isomer, where l is selected from 1, 2, 3, 4, and 5, M1 is a bond or M', and R4 is hydrogen, unsubstituted C 1-3 Alkyl or -(CH2) n Q is a heteroaryl or heterocycloalkyl group, where n is 2, 3, or 4, and Q is OH, -NHC(S)N(R)2, -NHC(O)N(R)2, -N(R)C(O)R, -N(R)S(O)2R, -N(R)R8, -NHC(=NR9)N(R)2, -NHC(=CHR9)N(R)2, -OC(O)N(R)2, -N(R)C(O)OR, heteroaryl or heterocycloalkyl, M and M' are independently selected from -C(O)O-, -OC(O)-, -OC(O)-M''-C(O)O-, -C(O)N(R')-, -P(O)(OR')O-, -SS-, aryl group and heteroaryl group, and R2 and R3 are independently H, C 1-14 Alkyl and C 2-14 It is selected from a group consisting of alkenils.
[0501] In one embodiment, the compound of formula (I) is the compound of formula (IIa) below, [ka] Alternatively, it may be its N-oxide, a salt thereof, or an isomer thereof, where R4 is as described herein.
[0502] In another embodiment, the compound of formula (I) is the compound of formula (IIb) below, [ka] Alternatively, it may be its N-oxide, a salt thereof, or an isomer thereof, where R4 is as described herein.
[0503] In another embodiment, the compound of formula (I) is the compound of formula (IIc) or (IIe) below, [ka] Alternatively, it may be its N-oxide, a salt thereof, or an isomer thereof, where R4 is as described herein.
[0504] In another embodiment, the compound of formula (I) is the compound of formula (IIf) below, [ka] or its N-oxide, or its salt or isomer, In the formula, M is -C(O)O- or -OC(O)-, and M'' is C 1-6 Alkyl or C 2-6 It is an alkenyl, and R2 and R3 are independently C 5-14 Alkyl and C 5-14 The values are selected from the group consisting of alkenyls, and n is selected from 2, 3, and 4.
[0505] In further embodiments, the compound of formula (I) is the compound of formula (IId) below, [ka] or its N-oxide, or its salt or isomer, where n is 2, 3, or 4, and m, R', R'' and R2-R6 are as described herein. For example, R2 and R3 are independently C 5-14 Alkyl and C 5-14The group may be selected from the group consisting of alkenils.
[0506] In further embodiments, the compound of formula (I) is the compound of formula (IIg) below, [ka] Or its N-oxide, or its salt or isomer, where l is selected from 1, 2, 3, 4 and 5, m is selected from 5, 6, 7, 8 and 9, M1 is a bond or M', M and M' are independently selected from -C(O)O-, -OC(O)-, -OC(O)-M''-C(O)O-, -C(O)N(R')-, -P(O)(OR')O-, -SS-, aryl group and heteroaryl group, and R2 and R3 are independently H, C 1-14 Alkyl and C 2-14 It is selected from the group consisting of alkenils. For example, M'' is C 1-6 Alkyl (for example, C 1-4 (Alkyl) or C 2-6 Alkenyl (for example, C 2-4 (Alkenyl) For example, R2 and R3 are independent of C 5-14 Alkyl and C 5-14 It is selected from a group consisting of alkenils.
[0507] In some embodiments, the ionizable lipid is one or more compounds described in U.S. Patent Applications Nos. 62 / 220,091, 62 / 252,316, 62 / 253,433, 62 / 266,460, 62 / 333,557, 62 / 382,740, 62 / 393,940, 62 / 471,937, 62 / 471,949, 62 / 475,140, and 62 / 475,166, as well as International Application PCT / US2016 / 052352.
[0508] In some embodiments, the ionizable lipid is selected from compounds 1 to 280 described in U.S. Patent Application No. 62 / 475,166.
[0509] In some embodiments, the ionizable lipid is [ka] or its salt.
[0510] In some embodiments, the ionizable lipid is [ka] or its salt.
[0511] In some embodiments, the ionizable lipid is [ka] or its salt.
[0512] In some embodiments, the ionizable lipid is [ka] or its salt.
[0513] The central amine moiety of lipids according to formulas (I), (IA), (IB), (II), (IIa), (IIb), (IIc), (IId), (IIe), (IIf), or (IIg) can be protonated at physiological pH. Therefore, lipids may have a positive charge or a partial positive charge at physiological pH. Such lipids are sometimes called cationic (amino)lipids or ionizable (amino)lipids. Lipids may also be amphoteric molecules, i.e., neutral molecules, possessing both positive and negative charges.
[0514] In some embodiments, the ionizable lipids of this disclosure are compounds of the following formula (III), [ka] or one or more of its salts or isomers, in the formula, W is [ka] And, Ring A is, [ka] And, t is either 1 or 2. A1 and A2 are independently selected from CH or N. Z is either CH2 or absent. When Z is CH2, the dashed lines (1) and (2) each represent a single bond. When Z is absent, neither the dashed lines (1) nor (2) exist. R1, R2, R3, R4, and R5 are independent of C 5-20 Alkyl, C 5-20 Alkenil, -R"MR', -R * YR'', -YR'' and -R * It is selected from the group consisting of "OR". R X1 and R X2 Each of these is independently H or C1-3 alkyl, Each M is independently selected from the group consisting of -C(O)O-, -OC(O)-, -OC(O)O-, -C(O)N(R')-, -N(R')C(O)-, -C(O)-, -C(S)-, -C(S)S-, -SC(S)-, -CH(OH)-, -P(O)(OR')O-, -S(O)2-, -C(O)S-, -SC(O)-, aryl groups, and heteroaryl groups. M * It is a C1-C6 alkyl group, W 1 and W 2 Each of these is independently selected from the groups consisting of -O- and -N(R6)-, Each R6 independently controls H and C 1-5 Selected from the group consisting of alkyl groups, X 1 , X 2 and X 3These are independently bonded, -CH2-, -(CH2)2-, -CHR-, -CHY-, -C(O)-, -C(O)O-, -OC(O)-, -(CH2) n -C(O)-, -C(O)-(CH2) n -,-(CH2) n -C(O)O-, -OC(O)-(CH2) n -,-(CH2) n -OC(O)-, -C(O)O-(CH2) n It is selected from the group consisting of -, -CH(OH)-, -C(S)-, and -CH(SH)-. Each Y is independent of C 3-6 It is a carbon ring, Each R * C 1-12 Alkyl and C 2-12 It is selected from a group consisting of alkenyls. Each R is independent of C 1-3 Alkyl and C 3-6 Selected from the group consisting of carbon rings, Each R' is independent of C 1-12 Alkyl, C 2-12 It is selected from the group consisting of alkenyls and H. Each R' is independent of C 3-12 Alkyl, C 3-12 Alkenyl and -R * Selected from the group consisting of MR', n is an integer between 1 and 6. Ring A [ka] In that case, i)X 1 , X 2 and X 3 At least one of them is not -CH2- and / or ii) At least one of R1, R2, R3, R4, and R5 is -R"MR'.
[0515] In some embodiments, the compound is one of the compounds of the following formulas (IIIa1) to (IIIa8). [ka] [ka]
[0516] In some embodiments, the ionizable lipid is one or more compounds described in U.S. Patent Applications No. 62 / 271,146, No. 62 / 338,474, No. 62 / 413,345, and No. 62 / 519,826, and International Application PCT / US2016 / 068300.
[0517] In some embodiments, the ionizable lipid is selected from compounds 1 to 156 described in U.S. Patent Application No. 62 / 519,826.
[0518] In some embodiments, the ionizable lipid is selected from compounds 1-16, 42-66, 68-76, and 78-156 described in U.S. Patent Application No. 62 / 519,826.
[0519] In some embodiments, the ionizable lipid is [ka] or its salt.
[0520] In some embodiments, the ionizable lipid is [ka] or its salt.
[0521] The central amine moiety of lipids according to formulas (III), (IIIa1), (IIIa2), (IIIa3), (IIIa4), (IIIa5), (IIIa6), (IIIa7), or (IIIa8) can be protonated at physiological pH. Therefore, lipids may have a positive or partial positive charge at physiological pH. Such lipids are sometimes called cationic (amino)lipids or ionizable (amino)lipids. Lipids may also be amphoteric molecules, i.e., neutral molecules, possessing both positive and negative charges.
[0522] Phospholipids The lipid composition of the lipid nanoparticle composition disclosed herein may contain one or more phospholipids, for example, saturated phospholipids or (polyunsaturated) phospholipids, or a combination thereof. Generally, a phospholipid comprises a phospholipid portion and one or more fatty acid portions.
[0523] The phospholipid portion can be selected from a non-limited group, such as phosphatidylcholine, phosphatidylethanolamine, phosphatidylglycerol, phosphatidylserine, phosphatidic acid, 2-lysophosphatidylcholine, and sphingomyelin.
[0524] The fatty acid portion can be selected from a non-limiting group consisting of, for example, lauric acid, myristic acid, myristoleic acid, palmitic acid, palmitoleic acid, stearic acid, oleic acid, linoleic acid, alpha-linolenic acid, erucic acid, phytanic acid, arachidic acid, arachidonic acid, eicosapentaenoic acid, behenic acid, docosapentaenoic acid, and docosahexaenoic acid.
[0525] Certain phospholipids can facilitate fusion to membranes. For example, cationic phospholipids can interact with one or more negatively charged phospholipids in a membrane (e.g., a cell membrane or intracellular membrane). This fusion of phospholipids to the membrane allows one or more elements (e.g., therapeutic agents) of a lipid-containing composition (e.g., LNPs) to pass through the membrane, enabling, for example, delivery of those elements to target tissue.
[0526] Non-natural phospholipid species are also considered, including those that have been modified and substituted (including branching, oxidation, cyclization, and alkynes) from natural species. For example, phospholipids can be functionalized with one or more alkynes (e.g., alkenyl groups in which one or more double bonds are replaced by triple bonds) or crosslinked with one or more alkynes. Under appropriate...
Claims
1. A messenger RNA (mRNA) comprising a polynucleotide including an open reading frame (ORF) encoding a polypeptide containing the amino acid sequence shown in Sequence ID No. 3, wherein the mRNA includes a polyA region, the polyA region is at least 10 nucleotides long, and the mRNA includes a stable tail at its 3' end.
2. The mRNA according to claim 1, wherein the ORF is at least 90% identical to the nucleotide sequence shown in SEQ ID NO:
22.
3. The mRNA according to claim 1, wherein the ORF is at least 97% identical to the nucleotide sequence shown in SEQ ID NO:
22.
4. The mRNA according to claim 1, wherein the ORF is 100% identical to the nucleotide sequence shown in SEQ ID NO:
22.
5. The mRNA according to any one of claims 1 to 4, comprising a 5' UTR containing a nucleic acid sequence that is at least 90% identical to the nucleic acid sequence shown in Sequence ID No.
56.
6. The mRNA according to any one of claims 1 to 4, comprising a 5' UTR containing a nucleic acid sequence that is 100% identical to the nucleic acid sequence shown in Sequence ID No.
56.
7. The mRNA according to any one of claims 1 to 6, comprising a 3'UTR containing a nucleic acid sequence that is at least 90% identical to the nucleic acid sequence shown in Sequence ID No.
108.
8. The mRNA according to any one of claims 1 to 6, comprising a 3' UTR containing a nucleic acid sequence that is 100% identical to the nucleic acid sequence shown in Sequence ID No.
108.
9. The mRNA according to claim 1, comprising the nucleotide sequence shown in Sequence ID No.
42.
10. mRNA according to any one of claims 1 to 9, comprising a 5' terminal cap.
11. The mRNA according to claim 10, wherein the 5' terminal cap includes the caps Cap0, Cap1, ARCA, inosine, N1-methyl-guanosine, 2'-fluoro-guanosine, 7-deaza-guanosine, 8-oxo-guanosine, 2-amino-guanosine, LNA-guanosine, 2-azidoguanosine, Cap2, Cap4, or 5'methyl G.
12. The mRNA according to any one of claims 1 to 11, wherein the polyA region is at least 100 nucleotides long.
13. The mRNA according to any one of claims 1 to 12, comprising the nucleotide sequence UCUAGAAA
14. The mRNA according to any one of claims 1 to 13, wherein all of the uracil in the mRNA is N1-methylpseudolacil.
15. The mRNA according to claim 1, wherein the ORF is 100% identical to the nucleic acid sequence shown in SEQ ID NO: 22, the mRNA includes a poly-A region of at least 100 nucleotides in length, and all of the uracil in the mRNA is N1-methylpseudolacil.
16. The mRNA according to claim 1, wherein the mRNA comprises a 5' terminal cap containing a guanine cap nucleotide containing N7 methylation, the 5' terminal nucleotide of the mRNA comprises 2'-O-methyl, the mRNA comprises the nucleotide sequence shown in SEQ ID NO: 42, the mRNA comprises a poly(A) region of at least 100 nucleotides in length, and all of the uracil in the mRNA is N1-methylpseudolacil.
17. The mRNA according to claim 15 or 16, comprising the nucleotide sequence UCUAGAAA
18. Lipid nanoparticles containing mRNA according to any one of claims 1 to 17.
19. (i) Compound II, (ii) Cholesterol and (iii) PEG-DMG or Compound I, (i) Compound VI, (ii) Cholesterol and (iii) PEG-DMG or Compound I, (i) Compound II, (ii) DSPC or DOPE, (iii) Cholesterol, and (iv) PEG-DMG or Compound I. (i) Compound VI, (ii) DSPC or DOPE, (iii) Cholesterol and (iv) PEG-DMG or Compound I, (i) Compound II, (ii) Cholesterol and (iii) Compound I, or (i) Compound II, (ii) DSPC or DOPE, (iii) Cholesterol, and (iv) Compound I It contains, and compound I, compound II, and compound VI have the following structures: 【Chemistry 1】 【Chemistry 2】 【Transformation 3】 Lipid nanoparticles according to claim 18.
20. A pharmaceutical composition comprising mRNA according to any one of claims 1 to 17.
21. A pharmaceutical composition comprising lipid nanoparticles according to claim 18 or 19.
22. mRNA according to any one of claims 1 to 17, lipid nanoparticles according to claim 18 or 19, or a pharmaceutical composition according to claim 20 or 21, for use in a method for expressing phenylalanine hydroxylase polypeptide.
23. mRNA according to any one of claims 1 to 17, lipid nanoparticles according to claim 18 or 19, or a pharmaceutical composition according to claim 20 or 21, for use in a method for improving phenylalanine hydroxylase activity.
24. mRNA according to any one of claims 1 to 17, lipid nanoparticles according to claim 18 or 19, or a pharmaceutical composition according to claim 20 or 21, for use in a method for treating, preventing or delaying the onset and / or progression of phenylketonuria.
25. mRNA according to any one of claims 1 to 17, lipid nanoparticles according to claim 18 or 19, or a pharmaceutical composition according to claim 20 or 21, for use in a method for reducing phenylalanine levels.
26. The mRNA, lipid nanoparticles, or pharmaceutical composition according to claim 25, wherein the phenylalanine level is the phenylalanine level in blood, plasma, serum, liver, and / or urine.
27. mRNA, lipid nanoparticles, or pharmaceutical composition according to any one of claims 22 to 26 for intravenous administration.
28. mRNA, lipid nanoparticles, or pharmaceutical composition according to any one of claims 22 to 26 for subcutaneous administration.