Cis-regulatory elements of translation and methods using same

EP4720279A2Pending Publication Date: 2026-04-08YALE UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-25
Publication Date
2026-04-08

AI Technical Summary

Technical Problem

Current methods lack effective identification and utilization of cis-regulatory elements in mRNA molecules to modulate translational efficiency, which is crucial for mRNA-based vaccines and protein productions, especially for toxic proteins where controlled expression is necessary.

Method used

Development of a method, NaP-TRAP, to identify and incorporate specific cis-regulatory elements, such as modified nucleobases and sequences, into mRNA molecules to control translational efficiency by capturing actively translating ribosomes and nascent peptides, allowing for the construction of mRNA molecules with desired translational profiles.

Benefits of technology

Enables precise modulation of translational efficiency, increasing or decreasing protein production by up to 50% or more, and adapting to different cellular stages, facilitating the production of therapeutic peptides and vaccines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024031165_05122024_PF_FP_ABST
    Figure US2024031165_05122024_PF_FP_ABST
Patent Text Reader

Abstract

Described herein is a non-natural mRNA molecule, which includes a non-natural cisregulatory element that modulates the translational efficiency of the mRNA molecule. Also described herein is a method for modulating mRNA translational efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CIS-REGULATORY ELEMENTS OF TRANSLATION AND

[0002] METHODS USING SAME

[0003] CROSS-REFERENCE TO RELATED APPLICATIONS

[0004] The present application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application No. 63 / 504,628, filed May 26, 2023, which is incorporated herein by reference in its entirety’.

[0005] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH

[0006] This invention was made with government support under GM122580 awarded by National Institutes of Health. The government has certain rights in the invention.

[0007] SEQUENCE LISTING

[0008] The XML text file named " 047162-7468W01(02270)_Seq Listing. xml" created on May 25, 2024, comprising 497,445 bytes, is hereby incorporated by reference in its entirety.

[0009] BACKGROUND

[0010] The identification of cis-regulatory elements of translation (i.e., specific sequences in mRNA molecules that can modulate translational efficiency of the mRNA molecules - i.e., the rate at which mRNA translates into proteins) would allow the engineering of mRNA molecules that have specific translational efficiencies. For example, high translational efficiency is desirable for mRNA-based vaccines and large-scale protein productions, and decreased translational efficiency is desirable for the delivery of therapeutic nucleic acids which produce proteins that are toxic at higher dosages.

[0011] However, both knowledge on the sequences of cis-regulatory elements and methods of identifying such cis-regulatory elements are unfortunately lacking. Therefore, there is a need for novel methods of identifying cis-regulatory elements of translation, as well as for sequences of specific cis-regulatory elements in various species or cell types. The present study addresses this need. SUMMARY

[0012] In some aspects, the present invention is directed to the following non-limiting embodiments:

[0013] Non-natural mRNA molecule

[0014] In some aspects, the present invention is directed to an mRNA molecule.

[0015] In some embodiments, the mRNA molecule comprises a cis-regulatory element selected from the group consisting of SEQ ID NOs: 1-2353 and 2384, wherein the cis- regulatory element does not naturally exist in the mRNA molecule at a location of the cis- regulatory element.

[0016] In some embodiments, the cis-regulatory element is a cis-regulatory element for the translational machinery' of a mammal, optionally a human.

[0017] In some embodiments, the cis-regulatory element comprises at least one sequence selected from the group consisting of SEQ ID NOs: 1-1 156 and 2384.

[0018] In some embodiments, the cis-regulatory element is a cis-regulatory element for the translational machinery' of a fish, optionally a zebrafish.

[0019] In some embodiments, the cis-regulatory element comprises at least one sequence selected from the group consisting of SEQ ID NOs: 1157-2353.

[0020] In some embodiments, the mRNA molecule further comprises a modified nucleobase that does not naturally exist at the location of the modified nucleobase.

[0021] In some embodiments, the modified nucleobase comprises at least one of N6- methyladenosine, inosine, N 1 -propylpseudouridine. N1 -methoxy methylpseudouridine. Nl- ethylpseudouridine, 5-methoxy cytidine. 5-hydroxyuridine, 5 -carboxy uridine. 5- formyluridine, 5-hydroxy cytidine, 5-hyrdoxymethylcytidine, 5-hydroxymethyluridine, 5- formylcytidine, 5-carboxycytidine, N4-methylcytidine, pseudoisocytidine, 2-thiocytidine, 4- thiocytidine, N1 -methylpseudouridine, pseudouridine, 5-methyluridine, 5-methoxyuridine. dihydrouridine, 5-methylcytidine, 4-thiouridine. 2-thiouridine, uridine-5’-O-(l- thiophosphate), 5-aminoallyluridine, and 4-acetyl cytidine.

[0022] In some embodiments, the mRNA molecule encodes a therapeutic peptide or a therapeutic protein.

[0023] In some embodiments, the therapeutic peptide or the therapeutic protein comprises a vaccine.

[0024] Method of modulating a translational efficiency of mRNA In some aspects, the present invention is directed to a method of modulating a translational efficiency of a messenger RNA (mRNA) molecule in a cell.

[0025] In some embodiments, the method comprises: modifying the mRNA molecule to comprise a cis-regulatory element selected from the group consisting of SEQ ID NOs: 1-2353 and 2384.

[0026] In some embodiments, modifying the mRNA molecule comprises modifying a sequence of a DNA molecule encoding the mRNA molecule.

[0027] In some embodiments, the cis-regulatory element is a cis-regulatory element for the translational machinery' of a mammal, optionally a human.

[0028] In some embodiments, the cis-regulatory element comprises at least one sequence selected from the group consisting of SEQ ID NOs: 1-1156 and 2384.

[0029] In some embodiments, the cis-regulatory element is a cis-regulatory element for the translational machinery of a fish, optionally a zebrafish.

[0030] In some embodiments, the cis-regulator ' element comprises at least one sequence selected from the group consisting of SEQ ID NOs: 1157-2353.

[0031] In some embodiments, the mRNA molecule is further modified to comprise a modified nucleobase.

[0032] In some embodiments, the modified nucleobase comprises at least one of N6- methyladenosine, inosine, N I -propylpseudouridine, N1 -methoxy methylpseudouridine. Nl- ethylpseudouridine, 5-methoxy cytidine. 5-hydroxyuridine, 5 -carboxy uridine. 5- formyl uridine, 5-hydroxy cytidine, 5-hyrdoxymethylcytidine, 5-hydroxymethyluridine, 5- formylcytidine, 5-carboxycytidine, N4-methylcytidine, pseudoisocytidine, 2-thiocytidine, 4- thiocytidine, N1 -methylpseudo uridine, pseudouridine, 5-methyluridine, 5 -methoxy uridine, dihydrouridine, 5-methylcytidine, 4-thiouridine. 2-thiouridine, uridine-5’-O-(l- thiophosphate), 5-aminoallyluridine, or 4-acetylcytidine.

[0033] In some embodiments, the cis-regulatory' element activates the translation of the mRNA molecule, and wherein a translational efficiency of the mRNA molecule is increased by about 50% or more in comparison to an mRNA molecule lacking the cis-regulatory element but otherwise has the same sequence.

[0034] In some embodiments, the cis-regulatory' element represses the translation of the mRNA molecule, and wherein a translational efficiency of the mRNA molecule is decreased by about 30% or more in comparison to an mRNA molecule lacking the cis-regulatory element but otherwise has the same sequence. In some embodiments, the cis-regulatory element activates or represses the translation of the mRNA molecule in the cell at a certain cellular stage.

[0035] In some embodiments, the cis-regulatory element is located in a 5’ untranslated region (5’-UTR), at a region including the start codon of the main open reading frame, within the open reading frame, or in a 3’ untranslated region (3’-UTR).

[0036] BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The following detailed description of exemplary embodiments will be better understood when read in conjunction with the appended drawings. For the purpose of illustrating, non-limiting embodiments are shown in the drawings. It should be understood, how ever, that the instant specification is not limited to the precise arrangements and instrumentalities of the embodiments shown in the drawings.

[0038] Figs. 1A-1C demonstrate that NaP-TRAP is able to translation through immunocapture, in accordance with some embodiments. Fig. 1A: In NaP-TRAP reporters contain an N-terminal FLAG-tag. Ribosomes translating in-frame express the 3xFLAG peptide in their nascent chains. Translation is measured through immunocapture, using anti- FLAG magnetic beads. The immunoprecipitation efficiency of a reporter correlates with mean number of active ribosomes translating the reporter in-frame. Fig. IB: NaP-TRAP captures translation repression by a morpholino at 2 hpf. Fig. 1C: NaP-TRAP captures miR430 mediated translation repression at 4 hpf.

[0039] Figs. 2A-2G demonstrate that NaP-TRAP is able to measure the effect of Kozak strength on translation, in accordance with some embodiments. Fig. 2A: Construction of a library containing six random nucleotides upstream and one nucleotide downstream of the start codon of 3xFLAG-GFP. The NaP-TRAP w as performed at 6 hpf. Sequence logos for the top and bottom 10% of reporters w ere generated. Figs. 2B-2D: A random forest regressor was utilized to predict translation. Reporter Kozak sequences were one-hot encoded. All nucleotide positions were supplied to the model. The model’s prediction explained 86% of the variance of the test set (SPR). The permuted feature importance on the trained RFM w as performed. Nucleotide positions in red correlate negatively with translation, whereas positions in blue correlate positively with translation (logistic regression). Fig. 2E: The translation measurements were compared at 6 hpf to an in silico derived Kozak score based on frequency of Kozak sequences in the transcriptome. Translation measurement and Kozak scores correlated w eakly (SPR 0.3). Figs. 2F-2G: The translation of ten reporters with a Kozak score was measured between 295-305 using luciferase-based assay. The luciferase and NaP-TRAP translation values correlated well. The sequences shown in Figs. 2A and 2F are the following:

[0040] Figs. 3A-3K show an example of using NaP-TRAP to predict the translation of maternal 5‘ UTRs, in accordance with some embodiments. Fig. 3A: A 11,088 reporter synthetic oligo library was designed from the 5' UTRs of 1.750 maternally-supplied transcripts. Endogenous 5’ UTRs were tiled every 25 nucleotides with 124 nucleotide segments. Spike-ins were employed added at Trizol extraction step to normalize translation across experiments. Fig. 3B: The present study observed an increase in translation of 1.8-fold between 2 and 6 hpf. Fig. 3C:The random forest regression model was trained using kmer counts (1-6) and features characterizing uAUGs. Features were filtered prior to training based on their correlation to translation. Figs. 3D-3G: A random forest regression machine learning model was performed to predict translation at 2 and 6 hpf (Figs. 3D-3E). Given the prominent role of uORFs the analysis was repeated on reporters without upstream start codons (Figs. 3F-3G).

[0041] Figs. 4A-4J demonstrate that NaP-TRAP captures cis-regulatory elements encoded in the 5’ UTRs of maternally supplied transcripts, in accordance with some embodiments. Fig. 4A: The present study measured translation at 2 hpf and 6 hpf using NaP-TRAP. Reporters were divided into four groups based on their translation rank at both time-points (repressed, active, early activated, and late activated / orange, blue, green and pink.). Figs. 4B-4E: The fold-change enrichment or depletion was calculated for each tetramer in each group relative to all reporters. A hyper-geometric test was employed to calculate significance (Bonferroni corrected p-value threshold). 4F: The translation of repeats of each tetramer (256 reporters total) at 2 and 6 hpf was measured. Figs. 4G-4J: The reporters were color based on the tetramer enrichment of each group in the maternal 5’ UTR library. Figs. 5A-5H shows examples of the investigation of developmentally dynamic cis- regulatory elements, in accordance with some embodiments. Fig. 5A: Maternal 5' UTR library translation values at 2 and 6 hpf. Reporters containing miR430 seeds labeled in red (GCACUUA, AGCACUU). Fig. 5B: Dual luciferase assay measuring the inhibitory effect of miR430 binding sites in the 5’ UTR. 5xmir430-nanoluc and 5xmir430-shuffled-nanoluc were injected in wildly pe and miR430 knockout embryos. Fig. 5C: Effect of 5' UTR miR430 seeds on mRNA half-life. Fig. 5D: Rank plot comparing features driving translation at 2 and 6 hpf for reporter mRNAs with no uORFs. Fig. 5E: Cartoon illustration NaP-TRAP experiment to measure the effect of ZGA on the dynamic role of C-richness. Figs. 5F-5H: NaP-TRAP experiments measuring U-rich, C-rich and U>C and C>U rich reporters at 2 hpf and 6 hpf.

[0042] Figs. 6A-6G shows the adaption of te NaP-TRAP to human cell lines, in accordance with some embodiments. Fig. 6A: HEK293T were transfected using lipofectamine message max and performed NaP-TRAP at 12 hours post transfection. Fig. 6B: The distribution of translation values in HEK293T cells. Fig. 6C: Pentamers enriched in the bottom 10% of reporters based on translation values. Fig. 6D: Pentamers enriched in the top 10% of reporters based on translation values. Fig. 6E: The effect of uORF number on translation in HEK293T cells. Fig. 6F: Random Forest regression model predicts translation, explaining 79% of the variation observed in the test set. Fig. 6G: Permuted feature importance w as performed on the Random Forest model.

[0043] Figs. 7A-7J: Random Forest regression models predict translation for sv40 (Figs. 7A and 7C), pA (Figs. 7B and 7D) and human (Fig. 7E) NaP-TRAP 5’ UTR maternal libraries, in accordance with some embodiments. Figs. 7F-7J: permuted feature importance was performed on random forest regression models: sv40 (Fig. 7F and 7H), pA (Figs. 7G and 71) and human (Fig. 7J) NaP-TRAP 5’ UTR maternal libraries.

[0044] Figs. 8A-8C demonstrate that NaP-TRAP is able to measure translation through immunocapture, in accordance with some embodiments. Fig. 8A: Schematic detailing the NaP-TRAP method. FLAG-tagged nascent chain complexes of reporter mRNAs are enriched via an anti-FLAG immunoprecipitation. Translation is measured as a ratio of reporter reads in the pulldown relative to the input. Fig. 8B: NaP-TRAP derived translation at 6 hpf (left, N = 3) and immunofluorescence measurements at 24 hpf (right, -MO N = 14, +MO N = 9) in the presence or absence of a translation blocking morpholino (unpaired t-test: ** p < 0.005, **** p < 0.0001). Fig. 8C: NaP-TRAP derived translation values for 3xFLAG-GFP-3xmiR-430 (red) and 3xFLAG-GFP-3xmiR-204 (black) at 2 hpf (N = 4) and 4.3 hpf (N = 4) (unpaired t- test: ** p < 0.01, *** p < 0.005). Schematic representation of repression of translation exercised by miR-430. but not miR-204 at 4.3 hpf.

[0045] Figs. 9A-9G demonstrate that NaP-TRAP is able to measure the effect of Kozak strength on translation, in accordance with some embodiments. Fig. 9A: A schematic detailing the Kozak library containing six random nucleotides upstream and one nucleotide downstream of the start codon of 3xFLAG-GFP (top left). Histogram of NaP-TRAP translation values at 6 hpf (bottom left). Sequence logos for the top and bottom 10% of reporters based on translation (right). Figs. 9B-9C: Cartoon of random forest regression model feature generation through the one-hot-encoding of reporter sequences (Fig. 9B). Scatterplot comparing the model’s prediction to the experimentally derived translation values of a test set of reporters (Fig. 9C; Pearson’s R) (N = 785 reporters). Fig. 9D: Permuted feature importance derived from random forest regression model. Nucleotide positions in purple correlate negatively with translation, whereas positions in blue correlate positively with translation (Spearman Rank Correlation Coefficient). Fig. 9E: Comparison of translation measurements at 6 hpf and an in silico derived Kozak score based on the frequency of Kozak sequences in the transcriptome (Pearson’s R) (N = 2,617 reporters). Figs. 9F-9G: The translation of four reporters with a Kozak score between 295-305 was measured using a dual luciferase-based assay (Fig. 9F). Plot comparing relative luciferase activity' (Nano luciferase / Firefly luciferase), and NaP-TRAP translation values (Fig. 9G) (Pearson's R. N = 3). The sequences shown in Figs. 9A and 9F are the following:

[0046] Figs. 10A-10I show an example of using NaP-TRAP to predict the translation of maternal 5’ UTRs, in accordance with some embodiments. Fig. 10A: A schematic detailing global changes in translation during the matemal-to-zygotic transition in the developing zebrafish embryo (top left). The 5’ UTR library was generated by tiling the 5’ UTRs of zebrafish (bottom left). The NaP-TRAP workflow with the addition of spike-ins at the RNA extraction step (right). Fig. 10B: Comparison of translation values between 2 and 6 hpf of the 5’ UTR library (Mann-Whitney p < 10-100, N = 8,529 reporters). Fig. 10C: Schematic detailing random forest regression model feature selection (k-mer counts 1-6 nt and features characterizing uAUGs). Figs. 10D-10I: Scatterplot comparing each model’s prediction to the experimentally derived translation values of a test set of reporters (N = 2,559 reporters) (Figs. 10D-10E; Pearson's R). The permuted feature importance of features (top 12) with the greatest effect on model performance (blue refers to a positive correlation, purple refers to a negative correlation; Spearman Rank Correlation Coefficient) (Figs. 10F-10G). The correlation of selected features with translation at the indicated timepoints (Figs. 1 OH-101).

[0047] Figs. 11 A-l 1H demonstrate that NaP-TRAP is able to capture general and dynamic 5’ UTRs cis-regulation. in accordance with some embodiments. Fig. 11 A: The translation of the 5’ UTR library at 2 hpf and 6 hpf using NaP-TRAP (Pearson’s R. N = 503 8.529 reporters). Reporters were divided into four groups based on their translation rank at both time-points: repressed, active, active post ZGA, and repressed post ZGA (orange, blue, pink, and green, respectively). Figs. 11B-11E: Fold-change enrichment and depletion for all pentamers in each group from Fig. 11 A relative to the reporter library (hyper-geometric test to calculate significance; Bonferroni corrected p-value threshold p < 5 * 10-6). Fig. 1 IF: Schematic detailing the library of all possible tetramer repeats separated by dinucleotide spacers. Fig. 11G: Translation measurements of the validation library' at 2 and 6 hpf (Pearson’s R, N = 195 reporters). Tetrameric reporters were labeled based on whether their repeat was enriched in the reporter groups described above (Fig. 11 A). Fig. 11H: Cumulative distribution plot comparing the translation (2hpf / 6 hpf) of validation reporters identified as active post ZGA (pink) and repressed post ZGA (green) (Mann- Whitney U-test: p < 10-5). The sequences shown in Fig. 1 IF are the following:

[0048] Figs. 12A-12E demonstrates that zygotic expression of miR-430 inhibits translation of mRNAs with 5’ UTR seed sites, in accordance with some embodiments. Fig. 12A: Cumulative distributions of translation (6 hpf / 2 hpf) for reporters containing miR-430 (GCACUUA, GCACUUU, AGCACUU; N = 212 reporters) or miR-1 (ACAUUCC, CAUUCCA; N = 71 reporters) heptamers, labeled in green and blue respectively. Reporters labeled in gray contain neither seed sequence (N = 8,242 reporters) (Mann- Whitney U test, p < 10-27 miR-430 reporters vs reporters with no seed; Mann- Whitney U test, p < 0.001 miR-1 reporters vs reporters with no seed). Figs. 12B-12C: The effect of number of complementary bases to miR-430 (B, N = 275 reporters) or miR-1 (C, N = 178 reporters) on the translation (6 hpf / 2 hpf) of reporters with seed sequences (Pearson’s R). Figs. 12D-12E: Schematic detailing dual luciferase assay measuring the inhibitory effect of miR-430 binding sites in the 5’ UTR (Fig. 12D). 4xmiR-430-nanoluc and 4xmiR-430-shuffled-nanoluc were injected in wild- type and miR-430 knockout embryos. Relative Luciferase Activity (RLU) values were normalized to each reporter (N = 3, unpaired t-test) (Fig. 12E).

[0049] Figs. 13A-13F show the dynamic effect of C-richness on translation, in accordance with some embodiments. Fig. 13 A: Plot comparing feature rank at 2 and 6 hpf (k-mers 4). At each time point feature rank was determined by calculating the mean difference in translation between non-uAUG reporters enriched and depleted (top and bottom 20%) in each feature (red: repressive at 2 hpf, blue: active at 2 hpf, purple: features that were repressive at 2 hpf and then active at 6 hpf). Features were only included if there was a significant difference in mean translation at each timepoint (Bonferroni corrected t-test). Fig. 13B: Proposed models explaining how the matemal-to-zygotic transition can alter the effect of cis- regulatory elements in the 5‘ UTR. Figs. 13C-13F: U's in two active (U-rich) reporters were mutated to C’s (U > C), whereas C’s in two active post ZGA (C-rich) reporters were mutated to U’s (C > U). Translation values were measured using NaP-TRAP at 2 and 6 hpf. Wild-type reporters are colored gray, whereas mutant reporters are labeled red (N = 3, unpaired t-test).

[0050] Figs. 14A-14J show s the adaptation of NaP-TRAP to human cell lines, in accordance with some embodiments. Fig. 14A: Schematic detailing NaP-TRAP in HEK293T cells. Cells were transfected with the 5’ UTR library’ using lipid nanoparticles. Translation was measured at 12 hours post transfection. Fig. 14B: The distribution of translation values of the 5’ UTR library in HEK293T cells. The top and bottom 10% of reporters based on their translation values are labeled in orange and blue (repressed and active reporters, respectively; N = 7,506 reporters). Figs. 14C-14D: A differential enrichment analysis identified pentamers enriched and depleted in repressed (Fig. 14C) and active (Fig. 14D) reporters, blue and orange respectively (hypergeometric test with a Bonferroni corrected p-value). Fig. 14E: Venn- diagram comparing the pentamers enriched in active reporters in HEK293T cells and zebrafish embryos at 2 and 6 hpf. Fig. 14F: Translation values at 2 hpf in zebrafish compared to translation values in HEK293T cells (Pearson’s R). Reporters were divided into four groups based on their translation rank at each condition: active (blue), repressed (orange), active in zebrafish at 2 hpf (pink) and active in HEK293T cells (green) (N = 7,506 reporters). Figs. 14G-14J: The fold-change enrichment or depletion for all pentamers in each group (active (Fig. 14G), repressed (Fig. 14H), active in zebrafish at 2 hpf (Fig. 141), and active in HEK293T (Fig. 14J)) relative to the reporter library. Significance was calculated using a hyper-geometric test (Bonferroni corrected p-value threshold).

[0051] Figs. 15A-15F demonstrate that the NaP-TRAP method can be adapted to study the translation of multiple ORFs simultaneously in vivo, in accordance with some embodiments. Fig. 15A: NaP-TRAP is an accessible, versatile, and quantitative method that measures the translation of thousands of reporters simultaneously through the immunocapture of FLAG- tagged nascent peptides. Fig. 15B: Through the over-expression of an HA-tagged RNA binding protein (RBP). NaP-TRAP can be employed in conjunction with an RNA immunoprecipitation experiment to measure the effect of RBP recruitment on translation. Figs. 15C-15F: NaP-TRAP quantifies translation in a frame-specific manner. Through the incorporation of additional epitope tags in frames 2,3 or ORFs outside of the main open reading frame, the NaP-TRAP method can be utilized to: (1) measure out-of-frame translation in the main ORF 586 (Fig. 15C). (2) detect internal ribosome entry sites (IRES) sequences in an unbiased manner through the use of a bicistronic 587 reporter (Fig. 15D), (3) identify frameshifting elements (Fig. 15E), and (4) quantify stop codon readthrough (Fig. 15F).

[0052] Fig. 16 shows the comparison of NaP-TRAP to existing translation based MPRA methods, in accordance with some embodiments. NaP-TRAP was developed to quantify the translation of thousands of reporters simultaneously in a frame-specific manner. In contrast to existing methods, NaP-TRAP can be adapted to dynamic model systems (e.g. the developing zebrafish embryo), and does not require specialized equipment or a large amount of source material to measure translation.

[0053] Figs. 17A-17C exemplify using NaP-TRAP to model and predict Kozak Strength in early embry os, in accordance with some embodiments. Fig. 17A: Plot comparing replicate translation values for the Kozak library (Pearson’s R; N = 2,712 reporters). Fig. 17B: The correlation coefficient between the top 12 features identified by the permuted feature importance analysis and translation (blue refers to positive correlation, purple refers to negative correlation; Spearman Rank Correlation Coefficient). Fig. 17C: The predicted translation values of all possible Kozak sequences in early zebrafish embryos (N = 15,364 Kozak sequences) (Random Forest regression model: Figs. 9B-9D). Figs. 18A-18J demonstrate that NaP-TRAP is able to quantitatively measure the translation of zebrafish 5' UTRs, in accordance with some embodiments. Figs. 18A-18B: Comparing the translation values of replicates of the 5’ UTR library with a hard-encoded poly-A at 2 and 6 hpf (Pearson’s R) (Fig. 18A: N = 9,516 reporters; Fig. 18B: N = 8,529 reporters). Figs. 18C-18D: Plot comparing relative luciferase activity (Nano-luciferase / Firefly luciferase) to NaP-TRAP derived translation values of the zebrafish 5’ UTR library validation reporters at 2 (Fig. 18C. N = 3) and 6 hpf (Fig. 18D, N = 4) (Pearson’s R). Figs. 18E-18F: Random Forest regression models were trained on reporters without upstream AUGs. The performance of each model was evaluated by predicting the translation values of a test set of reporters (Pearson’s R) (N = 660 reporters). Figs. 18G-18I: The permuted feature importance of the top 12 features informing the predictive power of each model (Figs. 18G and 18H). The spearman rank correlations between the top 12 features and translation at each timepoint (Figs. 181 and 18J) (blue represents positive correlation, purple represents negative correlation).

[0054] Figs. 19A-19N exemplify using NaP-TRAP to measure the regulatory’ activity of ciselements in an environment with reduced translation, in accordance with some embodiments. Figs. 19A-19B: Replicates of translation values for SV40 zebrafish 5’ UTR library^ (Pearson’s R) (Fig. 19A: N = 9,999 reporters; Fig. 19B: N = 7,442 reporters). Fig. 19C: The distributions of the translation values for the hard encoded pA (60 A) and the sv40 zebrafish 5’ UTR libraries at 2 and 6 hpf (Mann-Whitney U test p < 10-100) at 2 and 6 hpf (N = 7,427 reporters). Figs. 19D-191: Random Forest regression models predict translation at 2 and 6 hpf (N = 2,233 reporters) (Figs. 19D and 19G). The present study utilized a permuted feature importance analysis to identify the top 12 features informing the prediction of the RFM at both timepoints (Figs. 19E and 19H). The present study plotted the Spearman Rank Correlation of each of the selected features with translation (blue is positive and purple is negative) (Figs. 19F and 191). Fig. 19J: NaP-TRAP derived translation values the SV40 zebrafish 5’ UTR library at 2 hpf and 6 hpf. Reporters colored by enrichment group (yellow: repressive, blue: active, pink: active post-ZGA, and green: active pre-ZGA). Figs. 19K-19N: The enrichment for all pentamers in each group relative to all reporters (hyper-geometric test to calculate significance with Bonferroni corrected p-value threshold).

[0055] Figs. 20A-20E illustrate certain aspects of the NaP-TRAP validation library' of tetramer repeats, in accordance with some embodiments. Figs. 20A-20D: Reporters were labeled based on whether their encoded tetramer repeat was enriched in the reporter groups defined in the differential kmer enrichment analysis performed on the 5’ UTR library' (Figs. 11B-11E) (orange: repressive, blue: active, pink: active post ZGA, green: repressed post ZGA, gray: no group). The validation of the active motifs was limited by the length of the motifs as the dinucleotide spacers created the motifs present in the other groups (Fig. 20B). Fig. 20E: Previously labelled reporters were colored gray if they contained tetramers found in other groups (four or more counts).

[0056] Figs. 21A-21L show the comparison of translation in HEK293T cells and the developing zebrafish embryo, in accordance with some embodiments. Figs. 21A-21C: Replicates of translation values for the 5’ UTR library in HEK293T cells (Pearson’s R) (N = 7,584 reporters). Fig. 21D: Venn-diagram comparing the repressive features of HEK293T cells and the zebrafish embry o at 2 and 6 hpf (pentamers enriched in the bottom 10% of reporters). Fig. 21E: NaP-TRAP derived translation values from the zebrafish embryos at 6 hpf and HEK293T cells at 12 hpt. Both libraries had a hard-encoded 60A tail. Reporters colored by enrichment group (orange: repressive, blue: active, pink: active at 6 hpf, and green: active in HEK293T cells) (N = 7,506 reporters). Figs. 21F-21I: The enrichment for all pentamers in each group relative to all pentamers found in the reporter library’. Significance was determined using a hyper-geometric test (Bonferrom corrected p-value threshold). Figs. 21J-21L: Random Forest regression model predicts translation (N = 2,252 reporters) (Fig. 21 J). The permuted feature importance for features with the greatest effect on the model's predictive power (top 12; Fig. 21K). The correlation of said features with translation (Fig. 2 IL) (blue is positive correlation, purple is negative correlation, Spearman Rank Correlation Coefficient).

[0057] Fig. 22 demonstrates that, in the absence of out-of-frame translation, NaP-TRAP and polysome profiling correlate strongly, in accordance with some embodiments. Specifically, the present study performed polysome profiling in HEK 293T cells using the zebrafish 5’ UTR library. The present study observed a strong correlation between NaP-TRAP derived translation values and mean ribosome load (MRL).

[0058] Figs. 23A-23D demonstrate that NaP-TRAP is able to capture translation in multiple frames simultaneously, in accordance with some embodiments. Figs. 23A-23B: To measure translation in multiple frames simultaneously, the present study incorporated a FLAG, HA and MYC tag into frames 0, 1+ and 2+ respectively. The present study measured the translation of library’ consisting of upstream open reading frames in each frame by performing three different pulldowns on each sample using anti-FLAG, anti-HA and anti- MYC beads. Figs. 23C-23D: The present study observed out-of-frame translation (immunocapture) for reporters that contain oORFs in frames 1+ and 2+ An in-frame overlapping ORF or the presence of the main start codon in frame 0 represses translation in frames 1+ and 2+

[0059] Fig. 24 illustrates certain aspects of a luciferase assay, which shows 10-fold increased luciferase expression with RESA mRNA-Vl 5’UTR as compared to mRNA-1273 and BNT16 2b2 5’UTRs, in accordance with some embodiments. Error bars represent SD and statistical significance was calculated using independent t-test (**** p-value < 0.001).

[0060] DETAILED DESCRIPTION

[0061] The following disclosure provides many different embodiments, or examples, for implementing different features of the provided subject matter. Specific examples of components and arrangements are described below7to simplify the present disclosure. These are, of course, merely examples and are not intended to be limiting. For example, the formation of a first feature over or on a second feature in the description that follows may include embodiments in which the first and second features are formed in direct contact, and may also include embodiments in which additional features may be formed between the first and second features, such that the first and second features may not be in direct contact. In addition, the present disclosure may repeat reference numerals and / or letters in the various examples. This repetition is for the purpose of simplicity and clarity and does not in itself dictate a relationship betw een the various embodiments and / or configurations discussed.

[0062] Cis-regulatory elements encoded in the sequence and structure of mRNAs control protein expression. Conventional methods of determine translational efficiency of an mRNA molecule and identifying cis-regulatory elements are based on the immunocapture of ribosomes (such as epitope tagged ribosomes) followed by sequencing the mRNA molecules associated with the captured ribosomes. The rationale of such methods is that mRNA molecules having higher translational activity are being translated by more ribosomes, and vice versa. As such, the ribosome load on the mRNA molecules can show the translational activity of the mRNA molecules and reveal the existence, sequence and / or location of such cis-regulatory elements.

[0063] In the study described herein (“the present study”), it was discovered that ribosome load is not an entirely reliable indication of the translational efficiency. For example, ribosomes associated with an mRNA molecule can be inactive ribosomes, or translating outside of the main open reading frame of the mRNA molecule. Ribosome load cannot distinguish the efficiency of the translation of the main open reading frame from such nonproductive associations.

[0064] The present study further developed, in one aspect, a method of solving the above issues with the conventional methods. Rather than capturing the ribosomes, the method of the invention herein captures the translation complex that includes the mRNA, the ribosome, and the partially translated peptide still associated with the mRNA and the ribosome by capturing the partially translated peptide. This way, all the mRNA captured by the method herein are associated with peptides that are being actively translated from the mRNA molecules.

[0065] Furthermore, using the novel method herein, cis-regulatory elements of translation were identified in zebrafish and human cells. The cis-regulatory elements, some of which result in high translational efficiency, some of which result in low translational efficiency, some of which increase or decrease translational efficiency as time points advance, can be used to construct DNA or mRNA molecules that have desired translational efficiencies or patterns of translational efficiencies.

[0066] Accordingly, in some aspects, the present invention is directed to a method of identifying cis-regulatory elements of mRNA translation.

[0067] In some aspects, the present invention is directed to a method of determining translational efficiency of an mRNA molecule.

[0068] In some aspects, the present invention is directed to a mRNA molecule including a non-natural cis-regulatory elements, or a DNA molecule encoding the same.

[0069] In some aspects, the present invention is directed to a method of modulating translation efficiency of an mRNA molecule or a DNA molecule encoding the same.

[0070] In some embodiments, the present invention is directed to a method of constructing mRNA molecules for producing therapeutic proteins / peptides.

[0071] Definitions

[0072] As used herein, each of the following terms has the meaning associated with it in this section. Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Generally, the nomenclature used herein and the laboratory procedures in animal pharmacology, pharmaceutical science, peptide chemistry, and organic chemistry are those well-known and commonly employed in the art. It should be understood that the order of steps or order for performing certain actions is immaterial, so long as the present teachings remain operable. Any use of section headings is intended to aid reading of the document and is not to be interpreted as limiting; information that is relevant to a section heading may occur within or outside of that particular section. All publications, patents, and patent documents referred to in this document are incorporated by reference herein in their entirety, as though individually incorporated by reference.

[0073] In the application, where an element or component is said to be included in and / or selected from a list of recited elements or components, it should be understood that the element or component can be any one of the recited elements or components and can be selected from a group consisting of two or more of the recited elements or components.

[0074] In the methods described herein, the acts can be carried out in any order, except when a temporal or operational sequence is explicitly recited. Furthermore, specified acts can be carried out concurrently unless explicit claim language recites that they be carried out separately. For example, a claimed act of doing X and a claimed act of doing Y can be conducted simultaneously within a single operation, and the resulting process will fall within the literal scope of the claimed process.

[0075] In this document, the terms "a," "an," or "the" are used to include one or more than one unless the context clearly dictates otherwise. The term "or" is used to refer to a nonexclusive "or" unless otherwise indicated. The statement "at least one of A and B" or "at least one of A or B" has the same meaning as "A, B, or A and B."

[0076] "About" as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ±20% or ±10%, in certain embodiments ±5%, in certain embodiments ±1%, in certain embodiments ±0.1 % from the specified value, as such variations are appropriate to perform the disclosed methods.

[0077] Method of Identifying Cis-Regulatory Elements of mRNA Translation

[0078] In some aspects, the instant specification is directed to a method of identifying cis- regulatory element of translation.

[0079] In some embodiments, the method includes translating a plurality of mRNA molecules, each independently comprising a potential cis-regulatory element, with ribosomes to obtain an initial mixture.

[0080] In some embodiments, the method further includes generating from the initial mixture an enriched complex / mixture comprising mRNA molecules being translated, the ribosomes, and partial translation products being produced by the ribosomes from the mRNA molecules according to an affinity’ of the partial translation products to a binding partner thereof. In some embodiments, the method further includes sequencing the mRNA molecules present in the enriched complex / mixture.

[0081] In some embodiments, the method further includes determining an abundance of each mRNA molecule present in the enriched complex / mixture.

[0082] In some embodiments, the method further includes identifying cis-regulatory elements of translation based on the determined abundance of each of the plurality of mRNA molecules.

[0083] In some embodiments, each mRNA molecule of the plurality of mRNA molecules comprises a common sequence which translates into a peptide motif recognized by the binding partner. The peptide motif is not limited. One of ordinary skill in the art would understand that, as long as a peptide motif can, either by itself or via modifications, specifically interact with a binding partner and therefore enriched or purified by the affinity to the binding partner, the peptide motif is suitable for the purpose herein. Non-limiting examples of peptide motifs include ALFA-tag, AviTag, C-tag, Calmodulin-tag, polyglutamate tag, polyarginine tag, E-tag, FLAG-tag, HA-tag, His-tag, Myc-tag, NE-tag, SBP-tag. Spot-tag. Strep-tag, T7-tag, TC tag, V5 tag, VSV-tag. Xpress tag. and the like. The peptide sequences of such peptide motifs are well known in the art and, as such, nucleic acid sequences that translate into the peptide motifs can also be determined or designed. In some embodiments, the common sequence is a nature sequence in an mRNA molecule and the peptide sequence thereof can be recognized by, for example, an antibody or an antigen binding fragment thereof.

[0084] In some embodiments, the common sequence is to the 3 ’-end of the start codon. In some embodiments, the common sequence is within about 300 nucleotides from the start codon, such as within about 270 nucleotides, within about 240 nucleotides, within about 210 nucleotides, within about 180 nucleotides, within about 150 nucleotides, within about 120 nucleotides, within about 90 nucleotides, within about 60 nucleotides, within about 30 nucleotides, within about 27 nucleotides, within about 24 nucleotides, within about 21 nucleotides, within about 18 nucleotides, within about 15 nucleotides, within about 12 nucleotides, within about 9 nucleotides, within about 6 nucleotides, within about 3 nucleotides from the start codon. In some embodiments, the common sequence starts right after the start codon or include the start codon. In some embodiments, the start codon is the start codon of the main open reading frame. In some embodiments, the start codon is not the start codon of an upstream open reading frame (uORF) or a downstream open reading frame (dORF). In some embodiments, the size of the peptide motif is about 100 amino acid residues or less, such as about 90 amino acid residues or less, about 80 amino acid residues or less, about 70 amino acid residues or less, about 60 amino acid residues or less, about 50 amino acid residues or less, about 40 amino acid residues or less, about 30 amino acid residues or less, about 20 amino acid residues or less, about 15 amino acid residues or less, about 12 amino acid residues or less, about 10 amino acid residues or less, or about 8 amino acid residues or less.

[0085] In some embodiments, apart from the potential cis-regulatory elements, the plurality of the mRNA molecules has a same nucleotide sequence. In some embodiments, the pl urality of the mRNA molecules is engineered from a same nucleic acid molecule, by either inserting the potential cis-regulatory elements into the nucleic acid molecule and / or altering the existing sequences in the nucleic acid molecule to arrive at the potential cis-regulatory elements.

[0086] The choice of the binding partner depends on the peptide motif used for enriching the complex, and is not particularly limited. For example, virtually all unique peptide motifs can be targeted by antibodies or antigen binding fragments thereof that recognizes the peptide motif. In some embodiments, the peptide motif is recognized by a non-antibody binding partner. For example, AviTag can be biotinylation by the enzyme BirA, which can then be recognized by biotin binding proteins such as streptavidin. For example, His-Tag can be recognized by nickel (II) or cobalt (II) ions. For example. SBP-Tag can be recognized by streptavidin. One of ordinary skill in the art would know how to select the binding partner based on the choice of the peptide motif for recognition.

[0087] In some embodiments, generating the enriched complex / mixture includes immobilizing the binding partner to a substrate, a bead, or the like, contacting the translation mixture with the immobilized binding partner, and at least partially removing the unbound translation mixture from the immobilized binding partner. In some embodiments, the immobilized binding partner and the translation complex captured thereon are further washed by, for example, a buffer solution to remove mRNA molecules not associated with the complex.

[0088] In some embodiments, the mRNA molecules present in the enriched complex is sequenced by a high throughput sequencing method. High throughput sequencing methods are described in, for example, Reuter et al. (Mol Cell. 2015 May 21; 58(4): 586-597). In some embodiments, the mRNA molecules present in the enriched complex is converted into a cDNA molecule. In some embodiments, mRNA molecules present in the enriched complex are sequenced directly, such as by a nanopore sequencing method.

[0089] In some embodiments, the abundance of each mRNA molecule present in the enriched complex is determined by comparing abundance of each mRNA molecule with a predetermined baseline. The predetermined baseline can be, for example, an average or a median of the abundance of each mRNA molecules in the same sample or similar types of samples.

[0090] In some embodiments, a potential cis-regulatory element is determined to be cis- regulatory element that enhances translational efficiency if the abundance of mRNA molecules comprising the same is higher than the baseline, such as about 10% higher than the baseline, about 20% higher than the baseline, about 30% higher than the baseline, about 40% higher than the baseline, about 50% higher than the baseline, about 75% higher than the baseline, about 100% higher than the baseline, about 150% higher than the baseline, about 200% higher than the baseline, about 250% higher than the baseline, or about 300% higher than the baseline.

[0091] In some embodiments, a potential cis-regulatory element is determined to be cis- regulatory element that suppresses translational efficiency if the abundance of mRNA molecules comprising the same is lower than the baseline, such as less than about 90% of the baseline, less than about 80% of the baseline, less than about 70% of the baseline, less than about 60% of the baseline, less than about 50% of the baseline, less than about 40% of the baseline, or less than about 30% of the baseline.

[0092] In some embodiments, the method is configured to identify cis-regulatory elements that have different effects on translation efficiency at different time points.

[0093] In some embodiments, the different time points are different time points of cell division, cell differentiation, disease progression, embryo development, aging process, and so forth.

[0094] In some embodiments, the method includes translating the plurality of mRNA molecules at a first time point to obtain a first mixture; translating the plurality of mRNA molecules at a second time point after the first time point to obtain a second mixture; performing the enriching step, the sequencing step, the determination step and the identification step to identify cis-regulatory elements of translation that have different effects on translational efficiency at the first time point and the second time point. In some embodiments, translating the plurality of mRNA molecules comprises introducing the plurality of mRNA molecules into a cell.

[0095] In some embodiments, the cell is from a cell line, a primary cell, a cell in a tissue, a cell in an organ, a stem cell, an embryonic cell, a dividing cell, or a differentiating cell.

[0096] In some embodiments, the cell is a bacterial cell, a plant cell, a eukary otic cell, a vertebrate cell, a mammalian cell, a human cell, or the like.

[0097] Method of Determining Translational Efficiency of mRNA Molecule

[0098] In some embodiments, the instant specification is directed to a method of determining translational efficiency of an mRNA molecule.

[0099] In some embodiments, the method includes translating the mRNA molecules with ribosomes to obtain an initial mixture; generating from the initial mixture an enriched complex / mixture comprising the mRNA molecules being translated, the ribosomes, and partial translation products being produced by the ribosomes according to affinity' of the partial translation products to a binding partner thereof; determining abundance of the mRNA in the enriched complex / mixture; and comparing the abundance with a predetermined baseline.

[0100] In some embodiments, some or all the steps in the method of determining translational efficiency of an mRNA molecule are the same as or similar to those described elsewhere herein, such as in the “Method of Identifying Cis-Regulatory Elements of mRNA Translation” section. mRNA Molecule or DNA Molecule Encoding the Same

[0101] In some aspects, the present invention is directed to an mRNA molecule, or a DNA molecule encoding the same.

[0102] In some embodiments, the mRNA molecule includes a non-natural cis-regulatory elements that does not exist in the mRNA molecule naturally. The term “non-natural cis- regulatory elements" does not indicate that the cis-regulatory elements cannot exist in natural mRNA molecules. Rather, if the mRNA molecules herein includes a “non-natural cis- regulatory element” that exist in some natural mRNA molecules, the cis-regulatory' element does not exist in the natural mRNA molecule.

[0103] In some embodiments, the mRNA molecule includes a modified nucleobase. As used herein, the term “modified nucleobase” refers to nucleobases other than the four common nucleobases adenine (A), cytosine (C), uracil (U), and guanine (G). The term “modified nucleobase” does not exclude naturally occurring nucleobases that are found in non-mRNA molecules. For example, nucleobases like pseudouridine, 5-methylcytosine, Nl- methylpseudouridine and 2’-O-methylated exist in nature. These nucleobases are considered to be examples of modified nucleobases herein.

[0104] In some embodiments, the cis-regulatory element is in the 5’ untranslated region (5’- UTR). In some embodiments, the 5 -UTR includes an upstream open reading frame (uORF).

[0105] In some embodiments, the cis-regulatory element includes some or all the nucleotides of the start codon of the main open reading frame (also referred to herein as the “open reading frame” or the “ORF”), and a portion of the 5’-UTR and / or the main open reading frame.

[0106] In some embodiments, the cis-regulatory’ element is in the main open reading frame.

[0107] In some embodiments, the cis-regulatory elements is in the 3’ untranslated region (3’- UTR).

[0108] In some embodiments, the cis-regulator ’ element is a cis-regulatory element that regulates translation in a mammalian cell, such as a human cell.

[0109] In some embodiments, the cis-regulatory’ element is has a nucleotide sequence set forth in at least one of SEQ ID NOs:l-1156.

[0110] In some embodiments, cis-regulatory’ element is a cis-regulatory’ element that regulates translation in a fish cell, such as a zebrafish cell.

[0111] In some embodiments, the cis-regulatory’ element is has a nucleotide sequence set forth in at least one of SEQ IDs: 1157-2353.

[0112] In some embodiments, the cis-regulatory’ element activates the translation. In some embodiments, the cis-regulatory element activates the translation at a certain cellular stage, such as a certain stage of a cellular process, a cell division, a cell differentiation, a disease progression, an embryo development, an aging process, and so forth. In some embodiments, the inclusion of the cis-regulatory element increase the translation efficiency by about 10% or more, such as about 20% or more, 30% or more, 50% or more, 75% or more, 100% or more, 150% or more, 200% or more, 300% or more. 400% or more, or 500% or more, when compared to an mRNA molecule that does not include the cis-regulatory element but otherwise have the same sequence.

[0113] In some embodiments, the cis-regulatory’ element represses translation. In some embodiments, the cis-regulatory element represses the translation at a certain cellular stage, such as a certain stage of a cellular process, a cell division, a cell differentiation, a disease progression, an embryo development, an aging process, and so forth. In some embodiments, the inclusion of the cis-regulatory element decreases the translation efficiency by about 10% or more, such as about 20% or more, about 30% or more, about 40% or more, about 50% or more, about 60% or more, about 70% or more, about 80% or more, about 85% or more, about 90% or more, about 95% or more, about 98% or more, or about 99% or more, when compared to an mRNA molecule that does not include the cis-regulatory element but otherwise have the same sequence.

[0114] In some embodiments, the increase or decrease of the translation is measured according to the Nascent Peptide Translating Ribosome Affinity Purification (NaP-TRAP) assay according to some embodiments herein.

[0115] In some embodiments, the modified nucleobase is the nucleobases of N6- methyladenosine, inosine, N1 -propylpseudouridine. N 1 -meth oxy methylpseudouridine. Nl- ethylpseudouridine, 5-methoxy cytidine, 5-hydroxyuridine, 5-carboxyuridine, 5- formyluridine, 5-hydroxy cytidine, 5-hyrdoxymethylcytidine, 5-hydroxymethyluridine, 5- formylcytidine, 5-carboxycytidine, N4-methylcytidine, pseudoisocytidine, 2-thiocytidine, 4- thiocytidine, N1 -methylpseudo uridine, pseudouridine, 5-methyluridine, 5 -methoxy uridine, dihydroundine, 5-methylcytidine, 4-thiouridine. 2-thiouridine, uridine-5’-O-(l- thiophosphate), 5-aminoallyluridine, or 4-acetylcytidine.

[0116] In some embodiments, the mRNA molecule encodes a therapeutic peptide or a therapeutic protein.

[0117] In some embodiments, the therapeutic peptide or the therapeutic protein includes a vaccine.

[0118] Method of Modulating Translation Efficiency or Constructing mRNA Molecules

[0119] In some aspects, the present invention is directed to a method of modulating translational efficiency of an mRNA molecule or a DNA molecule encoding the same.

[0120] In some embodiments, the method includes modifying the mRNA molecule, or the DNA molecule encoding the same to include a cis-regulatory element in the mRNA. In some embodiments, the inclusion of the cis-regulatory element increases or decreases the translational efficiency. In some embodiments, the inclusion of the cis-regulatory element increases or decreases the translational efficiency at a certain cellular stage.

[0121] In some aspects, the present invention is directed to a method of constructing an mRNA molecule. In some embodiments, the mRNA molecule is an mRNA molecule for producing a therapeutic peptide or a therapeutic protein. In some embodiments, the method includes incorporating a cis-regulatory element in a template mRNA that encodes the therapeutic peptide or the therapeutic protein.

[0122] In some embodiments, the therapeutic peptide or a therapeutic protein includes a vaccine.

[0123] In some embodiments, the mRNA and / or the cis-regulatory element is the same as or similar to those detailed elsewhere herein, such as the “mRNA Molecule or DNA Molecule Encoding the Same” section.

[0124] Examples

[0125] The instant specification further describes in detail by reference to the following experimental examples. These examples are provided for purposes of illustration only, and are not intended to be limiting unless so specified. Thus, the instant specification should in no way be construed as being limited to the following examples, but rather, should be construed to encompass any and all variations which become evident as a result of the teaching provided herein.

[0126] Example 1: NaP-TRAP, a novel massive parallel reporter assay to quantify translation control

[0127] The cis-regulatory elements encoded in the sequence and structure of an mRNA determine its stability and translational output. While there has been a considerable effort to understand the factors driving transcription, the regulatory framework governing post- transcriptional control remains elusive. This is particularly true in non-steady state conditions (e.g., embryogenesis), where correlations between mRNA abundance and protein production break down.

[0128] According to some non-limiting embodiments, the study described herein (“the present study”) developed a novel massive parallel reporter assay (MPRA) to measure translation through immunocapture, Nascent Peptide Translating Ribosome Affinity Purification (NaP-TRAP). In contrast to other TRAP and polysome profiling based MPRAs, NaP-TRAP only captures actively translating ribosomes in a frame specific manner, resulting in a more accurate quantification of translation.

[0129] The present study employed NaP-TRAP to quantity7the regulatory landscapes of the Kozak sequence and maternally derived 5’ UTRs during early zebrafish embryogenesis, identifying general and developmentally dynamic cis-regulatory elements. The present study identified U-rich motifs as general enhancers of translation, and upstream ORFs and GC-rich motifs as general repressors of translation. Strikingly, the present study reports on a previously uncharacterized translational switch. During the matemal-to-zygotic transition, that pyrimidine-rich motifs shift from a repressor to a prominent activator of translation was observed. The present study also characterized 5’ UTR microRNA mediated translation repression driven by the zygotic expression of miR430. Lastly, the present study also adapted the method to human cell lines.

[0130] Life requires spatial and temporal control of protein expression. The protein output of a given transcript reflects the integration of translation and mRNA stability. While there has been a considerable effort to understand the factors driving mRNA synthesis, maturation, and decay, the regulatory frameworks governing translational control remain elusive. Cis- regulatory elements encoded in the sequence and structure of mRNAs modulate these processes. Although these elements are distributed throughout the transcript, they are often concentrated in the regions up- and dow nstream of the main coding sequence, the 5’ and 3’ untranslated regions (UTRs) respectively. Given that initiation is the rate limiting step of translation, there has been a particular interest in understanding the regulatory elements found in 5’ UTRs. These elements include internal ribosome entry sites (IRESs), G-quadruplexes, iron response elements (IREs), and upstream Open Reading Frames (uORFs).

[0131] To investigate translation control several studies have utilized ribosome profding. While this method quantifies accurately the translation efficiency of individual genes, its capacity to characterize the function of cis-regulatory elements is limited. The generation of ribosome protected fragments decouples the translation measurements of a given mRNA from their cognate untranslated regions. Thus, measurements of translation efficiency often reflect the amalgamation of several isoforms of a given gene each with a unique set of regulatory elements. This process is complicated further by the fact that endogenous transcripts contain multiple regulatory elements that exert differential and often competing effects on translation.

[0132] Massive parallel reporter assays (MPRAs) are particularly suitable to address these challenges. MPRAs measure the abundance and / or translation efficiency of thousands of reporters simultaneously. These assays have improved the understanding of translation control and mRNA stability. Previous works have characterized Kozak strength, uORFs, IRESs, codon optimality, RNA structure, and microRNA binding sites, as well as identified novel motifs driving translation and decay. Despite these insights, conclusions based on MPRAs have been limited by the methods used to measure translation: growth-selection, fluorescent-cell sorting (FACS), and polysome profiling. In MPRAs where reporters are encoded in DNA, measurements of translation can be confounded by additional layers of transcriptional regulation. In polysome fractionation, which quantifies the number of ribosomes on each transcript, inactive ribosomes, ribosomes translating out of frame, or ribosomes translating ORFs outside of the coding sequence can skew the measurement of translation.

[0133] Here, the present study developed NaP-TRAP (Nascent Peptide Translating Ribosome Affinity Purification), a novel method to measure in-frame translation through immunocapture of nascent peptides. More specifically, by including an N-terminal tag in the coding sequences of reporter mRNAs, the present study enriched for reporters in a manner proportional to the number of active ribosomes translating the main ORF in frame. Using this method, the present study quantified the Kozak strength and regulatory potential of 5’ UTRs in the developing zebrafish embryo and human cell lines. The present study demonstrate that the non-limiting NaP-TRAP assay has the capacity to measure the regulatory activity' of 5’UTRs across multiple systems. Using NaP-TRAP the present study identified developmentally controlled translational switches as well as common cis-regulatory elements conserved across vertebrates.

[0134] Example 1-1

[0135] The present study developed NaP-TRAP to measure translation in a frame specific manner. Given the prevalence of upstream open reading frames in 5’ UTRs and the repressed translational state of early embryogenesis, it was concerned that inactive ribosomes or ribosomes translating outside of the main open reading frame would confound the quantification of translation. In addition, the present study wanted to develop a translation based MPRA method that did not require a large amount of source material or specialized equipment.

[0136] It was hypothesized that, contrary to other TRAP methods that tag ribosomes, methods that quantify mRNA molecules associated with ribosomes and tagged partial translation products would produce broader readout of translation.

[0137] In conventional TRAP-based assays, the immunocapture of epitope tagged ribosomes immobilized by cycloheximide treatment enriches reporters in a manner proportional to their mean ribosome load. Like polysome profiling these assays capture all bound ribosomes. To measure frame specific translation, the present study elected to tag the nascent chain complex instead of the ribosome with an epitope. By incorporating a 3x FLAG-tag into the N-terminus of the coding sequence of GFP and performing an anti-FLAG immunoprecipitation, it was reasoned that translation could be measured as a ratio of reporter mRNA levels in the pulldown relative to the input (Fig. 1A).

[0138] Here the present study validated NaP-TRAP using reporters assays of increasing complexity. First, the present study measured the translation of individual reporters containing known cis-regulatory elements. Next, the present study quantified Kozak strength in zebrafish for the first time experimentally, measuring the translation of thousands of reporters simultaneously. Lastly, the present study assessed the regulatory potential of endogenous 5’ UTR sequences in the developing zebrafish embryo and HEK293T cells.

[0139] Example 1-2: NaP-TRAP measures translation via the immunocapture of nascent chains

[0140] Cis-regulatory elements encoded in mRNAs modulate translation efficiency. To measure the regulatory' potential of these elements the present study developed a novel massive parallel reporter assay (MPRA), NaP-TRAP (Nascent Peptide Translating Ribosome Affinity Purification). It was hypothesized that the present study could enrich reporters in a manner proportional to their translation efficiency through the immunocapture of FLAG- tagged nascent chain complexes immobilized by cycloheximide treatment. To evaluate the capacity of NaP-TRAP to measure translation the present study performed tw o reporter assays.

[0141] First, the present study measured the effect of a translation blocking morpholino. To this end, the present study co-injected 3xFLAG-GFP and dsRED mRNAs into single cell zebrafish embryos in the presence or absence of a translation blocking morpholino, targeting the start codon of 3xFLAG-GFP. Using NaP-TRAP the present study observed a 48-fold decrease in translation in the presence of the morpholino at 6hpf (Fig. IB). These results are consistent with the translational repression in morpholino injected embryos observed by the decrease in fluorescence (GFP / dsRED) at 24 hpf (Fig. IB). Next, the present study7tested the capacity7of NaP-TRAP to capture dynamic translation control. In the developing zebrafish embryo, the microRNA miR-430 is one of the first zygotically transcribed genes. At 4.3 hpf the expression of miR-430 results in the translation repression, deadenylation, and eventual decay of targeted mRNAs. To quantify this effect, the present study measured the translation of 3xFLAG-GFP reporters w ith partial complementarity' to tw o different microRNAs in their 3’ UTRs, miR-430 (3xFLAG-GFP-3xmiR430) and miR-204 (3xFLAG-GFP-3xmiR-204) respectively. The present study selected miR-204 target sites as the microRNA is not expressed in the early embryo. Using NaP-TRAP the present study measured the translation of mRNA reporters at 2 and 4.3 hpf, before and after miR-430 expression. Between the two timepoints, the translation of the 3xmiR430 reporter decreased, whereas the translation of the 3xmiR204 reporter increased. These results are consistent with the role of miR-430 in translational repression and the translation ramp observed during the early stages of development. Together these results demonstrate that NaP-TRAP measures the translation of individual reporters targeting general and developmentally dynamic cis-regulatory elements. Further, these experiments highlight the versatility of the method as they demonstrate the capacity of NaP-TRAP to quantify cis-regulation in both the 5’ and 3’ UTRs.

[0142] Example 1-3: Using NaP-TRAP to investigate Kozak strength in the developing zebrafish embryo

[0143] Next, to evaluate the capacity of NaP-TRAP to function as a MPRA, the present study designed a library to quantify the regulatory potential of the Kozak sequence, the present study elected to measure Kozak strength for three reasons: (1) the Kozak sequence is a strong determinant of translation initiation and (2) the Kozak strength is often inferred based on the frequency of these sequences in the genome, (3) Kozak strength has not been measured during development in vivo. To generate a library of diverse Kozak sequences, the present study incorporated six random nucleotides upstream and one random nucleotide downstream of the AUG of the 3xFLAG GFP reporter (Fig. 2A). The present study injected the in vitro transcribed mRNA library into single cell zebrafish embryos and measured the translation using NaP-TRAP at 6 hpf. The present study observed a strong correlation in translation values between replicates (Pearson’s r = 0.8). To identify active and repressive Kozak sequences, the present study generated position weight matrices for the top and bottom 10 % of reporters based on their translation (Fig. 2A). The present study observed that active Kozak sequences were enriched in Cs and As, whereas repressive Kozak sequences were enriched in Us and Gs.

[0144] To examine the effect of nucleotide identity and position on translation the present study employed a random forest regression model (RFM). Model features were generated based on nucleotide identities and positions within the Kozak sequence. Prior to training the model, the present study divided the data into a test and a training set, 20% and 80% of reporters respectively. Using the training set, the present study employed a 5-fold cross validation to optimize model parameters. The model’s prediction explained 86% of the variance observed in the test set data. Given this strong correlation the present study employed the model to predict the translation of all possible Kozaks. To identify the features informing the predictive power of the model, the present study performed a permutated feature importance analysis. Briefly, the present study shuffled the identities of the features of the model one at a time and measured the effect of the change on the predictive power of the model. At 6 hpf bases G, U, G, and C in positions -4, -3, -2 and +4 were the most repressive features, whereas C, A / G, and A in positions -6, -3, +4 were the most activating features.

[0145] Lastly, the present study compared NaP-TRAP translation values to an in silico derived metric (Kozak score). In zebrafish Kozak strength has been previously inferred based on the frequency of a given Kozak sequence within the transcriptome. This metric associates Kozak sequence abundance with increased translation initiation. The experimentally derived translation values here challenge this hypothesis as the present study observ ed a weak correlation between NaP-TRAP and Kozak score (R = 0.35, Fig. 2E). To validate this conclusion, the present study identified sequences that had similar Kozak scores (300 ± 5) yet exhibited different NaP-TRAP derived translation values and measured their translation using a dual luciferase reporter assay (Fig. 2F). The luciferase translation values correlated strongly with NaP-TRAP (Pearson R 0.88). This result not only highlights the importance of measuring Kozak strength experimentally, but also demonstrates that NaP-TRAP derived translation values correlate well with protein production. Taken together, these results demonstrate that NaP-TRAP can be employed as a MPRA. By measuring the translation of thousands of reporters simultaneously, the present study quantified and modeled Kozak strength for the first time in a vertebrate model.

[0146] Example 1-4: Investigating the regulatory potential of maternally supplied 5’ UTRs

[0147] Upon validating NaP-TRAP's capacity to function as an MPRA, the present study investigated the regulatory complexity of endogenous 5’ UTRs during the early stages of embryogenesis. Prior to zygotic genome activation, the translation of maternally supplied mRNAs drives development. Given that initiation is the rate limiting step of translation, 5’ UTRs provide a mechanism for maternally supplied mRNAs to modulate their protein output. To study the cis-regulatory elements encoded in these mRNAs, the present study designed an 11.088-sequence synthetic oligo library (124-nt long) by tiling the 5’-UTRs of 1.725 maternal mRNAs every 25 nucleotides (Fig. 3A). The present study injected the in vitro transcribed library into single cell zebrafish embryos and measured translation at 2 hpf and 6 hpf using NaP-TRAP. Using an internal spike in at the RNA extraction step, allowed the present study to quantify relative changes in translation across developmental stages (Fig. 3A). The present study observed a mean increase in translation between 2 hpf and 6 hpf of 1.8-fold, consistent with gradual de-repression of translation during the first 24 hours of development (Fig. 3B).

[0148] To model translation at each time point, the present study utilized random forest regression models. For each reporter, the present study generated kmer counts (1-6 nucleotides) and characterized upstream open reading frames (uORFs) for a number of features (Fig. 3C). The present study filtered the features based on the strength of their correlation to translation at each timepoint (abs SPR >= 0.05). Model parameters were optimized using a 5-fold cross validation on the training data (80% of reporters). The models predicted translation well at both timepoints, explaining >78% of variance observed in the test set data at 2 hpf and 6 hpf (Fig. 3D). The present study identified Kozak strength and the number of uORFs and out of frame oORFs as the features contributing most significantly to the predictive power of the model at both timepoints (Fig. 3F). Interestingly, features associated with upstream open reading frames exhibited larger feature importance at 6 hpf than at 2 hpf.

[0149] Given the prominent repressive effect of upstream open reading frames on translation, the present study repeated the analysis modeling reporters containing no uORFs or oORFs. Given that this subpopulation of reporters consisted of only about 20% of the original library, the present study trained the models on kmer counts ranging from 1 to 4 nucleotides. At 2 and 6 hpf the models explained 64% and 57% of the observed translation values respectively (Figs. 3G-3H). At 2 hpf the most predictive features were U-rich and UA- rich kmers, whereas at 6 hpf G-rich and UC rich kmers were the most predictive features (Figs. 3I-3J). Interesting the present study observed that the effect of C and UC-rich kmers on translation changed during development. To explore this observation further the present study ranked kmers based on the mean difference in translation between reporters that were enriched or depleted in said kmer, the 80thand 20thquantile respectively. The present study compared the rank of features with a significant difference at both timepoints. The present study observed that C, CC, and UCC rich were some of the most repressive features at 2 hpf yet activated translation at 6 hpf (Fig. 3K), suggesting that there is a developmental switch on the regulatory role of these motifs during development.

[0150] All together these results demonstrate the following: (1) uAUGs are the most prominent repressive elements during early embryogenesis and (2) that 5’UTR sequences can confer dynamic regulation during development.

[0151] Example 1-5: Identifying sequences driving differential translation To identify systematically the sequence elements modulating differential translation during development the present study first divided the reporters into four groups: (1) repressed, (2) active, (3) early activated, and (4) late activated based on their relative translation at 2 and 6 hpf. Next, the present study performed a differential k-mer enrichment analysis in each group relative to the reporter library.

[0152] This analysis revealed that repressed reporters were significantly enriched in upstream ORFs and GC-rich kmers and depleted in U-rich tracks (p >10’5hypergeometric test after Bonferroni correction). Conversely, active reporters were enriched in U-rich pentamers (UUUUU, CUUUU, UUUUA, GUUUU) and depleted in upstream start codons. Early activated reporters were enriched in AG / UG rich sequences (UAGUG, UAUUG, AAGAA, AGACU), as well as overlapping pentamers that were complementary to the seed site of the microRNA miR-430 (GCACU and AGCAC) and depleted in uORFs and pyrimidine repeats. In contrast, late activated reporters were enriched in pyrimidine repeats (CCUCC, UCUCU, CCCUC) and depleted in U-repeats.

[0153] To determine whether the observed regulation is the product of individual elements driving differential translation or the result of a combination of elements, the present study designed a library of repeats for all possible tetramers (Fig. 3H). The present study injected this in vitro transcribed library' into single-cell zebrafish embryos and measured translation using NaP-TRAP at 2 and 6 hpf. Interestingly, the present study identified GUCU as the most active tetramer translating over 2-folds more than the nearest reporter at both timepoints. The present study compared the results of the validation library to the maternal 5’ UTR library by identify ing tetramers enriched in the groups of reporters described above. The present study then plotted the distribution of the validation library, labeling each tetramer repeat based on their group enrichment in the maternal 5 UTR library (Figs. 3I-3M). While the distributions of repressed, early activated, and late activated tetramer reporters largely recapitulated the distributions of their respective groups in the maternal 5’ UTR library, the active tetramers exhibited a more heterogenous distribution. More specifically, U-rich kmers were not the principal driver of translation activation. The present study proposes that this discrepancy is the results of a difference in U repeat length between the respective libraries (Fig. 3 J). In the tetramer library' the longest U-repeat is five nucleotides, whereas U-repeats of 6 to 8 nucleotides are significantly enriched in the active reporters of the maternal 5’ UTR library. Interestingly, the present study once again observed translation repression of reporters with miR430 seeds at 6hpf (CACU and UUGC repeats). Using this differential enrichment analysis, the present study identified general and developmentally dynamic regulators of translation, repressed and active reporters, and early and late activated reporters respectively.

[0154] Example 1-6: Investigating the mechanisms driving developmentally dynamic translation control

[0155] In the developing zebrafish embryo, the pool of factors that regulate mRNA translation and decay changes overtime. Prior to zygotic genome activation at 4.3 hpf maternally supplied mRNAs and proteins drive translation control. Following genome activation zygotically expressed factors begin to reshape the post-transcriptional regulator}' landscape. Given that the microRNA miR430 is one of the first zygotically transcribed factors, the enrichment of miR430 seeds in early activated reporters suggested that ZGA alters the translation of maternal 5’ UTRs. To determine the effect of ZGA on maternally supplied mRNAs, the present study injected the maternal library into single cell zebrafish embryos. The present study soaked the embryos in flavopiridol, a transcription inhibitor or DMSO and measured translation at 2 and 6 hpf.

[0156] Example 1-7 : 5’ UTR miR430 recruitment suppresses translation

[0157] First, given that the microRNA miR430 is one of the first zygotically transcribed genes in the developing zebrafish embryo, the enrichment of miR430 seeds in early activated reporters suggests that genome activation shapes 5’ UTR mediated translation control (Fig. 4A). To validate this hypothesis, the present study constructed two nano-luciferase reporters: 5xmiR430-nanoluc and 5xmiR430-MUT -nanoluc. These reporters were identical with the exception of the fact that the 5’ UTR microRNA binding sites of the 3xmiR430-shuffled- nanoluc reporter were rearranged to prevent miR430 recruitment. The present study injected these reporters into single-cell wild-type and miR430 - / - mutant embryos in the presence of firefly luciferase and measured dual-luciferase activity at 6 hpf. The present study observed a 35% decrease in the ratio of wildtype to mutant luciferase activity for the miR430 reporter relative to the 3xmiR430-shuffled-nanoluc reporter, confirming the hypothesis that the zygotic transcription of miR430 suppresses the translation of mRNAs with seed sites in their 5’ UTRs (Fig. 4B).

[0158] Example 1-8: Zygotic genome activation modulates the effect of C-richness on translation? Second, the present study observed an enrichment of pyrimidine pentamers in the late activated reporters. To investigate this differential regulation further the present study selected two active (U-rich reporters) and two late activated reporters (C-rich reporters) from the maternal 5’ UTR library. While the active reporters were translated ~2 fold more than the late activated reporters at 2 hpf, the translation of all reporters w ere similar at 6 hpf. To determine the role of C-richness in this translation regulation the present study mutated Us in the active reporters to Cs and mutated Cs in the late activated reporters to Us for a total of eight reporters: four wildt pe and four mutants respectively. The present study injected the reporters into single cell zebrafish embryos in the presence and absence of a flavopiridol, a transcription inhibitor, and measured translation at 2 hpf and 6 hpf using NaP-TRAP. The present study observed that the C to U mutants enhanced translation at 2 hpf. whereas the U to C repressed translation at 2 hpf in both conditions. Strikingly, while these mutations did not significantly affect translation at 6 hpf in wild type conditions, the differences in the w ild type and mutant reporters were maintained in the absence of zygotic genome activation. This observation suggests that a zygotic factor is driving the switch in the regulatory’ potential of pyrimidine rich elements. Given that the mutations were distributed randomly throughout each wild-type reporter, the present study proposes that the observed translational switch is driven by C-richness rather than a single motif. That being said, C-richness is not considered a universal activator at 6 hpf. The enrichment of GC rich pentamers in repressed reporters and the fact that the tetramer reporter CCCC was lowly translated at both timepoints suggests that the local context and abundance of Cs within a 5’ UTR are also important.

[0159] All together these results, validate novel developmentally dynamic mechanisms of 5 ’ UTR mediated translation control. Demonstrating the importance of quantifying cis regulation in non-study state systems.

[0160] Example 1-9: NaP-TRAP can be readily adapted to other model systems

[0161] Lastly, to demonstrate the broad applicability of the method the present study- performed NaP-TRAP in HEK293T cells. The present study transfected the in vitro transcribed maternal library into HEK293T cells using lipid nanoparticles and measured translation at 12-hour post transfection using NaP-TRAP. Replicates were strongly correlated (Pearson Correlation Coefficient of 0.92). To identify- motifs modulating translation the present study performed a differential kmer enrichment analysis on active and repressed reporters, the top and bottom 10% of translation values respectively. Repressed reporters were enriched in AUG containing motifs and depleted in U-rich pentamers, whereas active reporters were enriched in pyrimidine and poly U rich pentamers, and depleted in AUG containing motifs (Figs. 6C-6D). Interestingly the human translation values correlated equally well with the zebrafish translation values at 2 hpf and 6 hpf (SPR 0.78 and 0.79 respectively), despite the fact that pyrimidine rich pentamers were only enriched in the active reporters of the human and 6 hpf data.

[0162] The magnitude of the fold-change enrichment and depletion of AUG pentamers in repressed and active reporters respectively suggests that the presence or absence of upstream ORFs are a dominant driver of translation in HEK293T cells. In agreement with this observation, the present study observe a strong negative association between uORF number and translation (Fig. 6E). Lastly, the present study employed a RFM to model translation; the model explained 75% of the experimental data (Fig. 6F). In agreement with the zebrafish data, Kozak strength and number of upstream AUGs were the most predictive features. Interestingly, these features had a greater contribution to the predictive power of the model when compared to the zebrafish feature importance analysis (Fig. 6G).

[0163] All together these results demonstrate that (1) NaP-TRAP is a robust method that can measure translation across multiple model systems, and (2) that upstream AUGs are a dominant driver of 5’ UTR mediated translation control in Hek293T cells.

[0164] Example 1-10

[0165] Here, the present study developed a novel method (NaP-TRAP) and demonstrate its capacity to measure translation quantitatively. The present study use NaP-TRAP to characterize cis-regulatory elements in the 5’ and 3’ UTRs, the strength of the Kozak sequence, and the regulatory landscape of the 5‘ UTRs of maternally supplied mRNAs in the developing zebrafish embryo and HEK293T cells. The present study identified conserv ed and developmentally dynamic cis-regulatory elements as well as characterized global changes to translation associated with early embryogenesis.

[0166] NaP-TRAP marks a significant improvement to translation based MPRAs. In contrast to existing methods NaP-TRAP measures translation an open reading frame specific manner without the potential for transcriptional bias. By enriching for reporters through the immunocapture of FLAG-tagged nascent chain complexes, NaP-TRAP measures active ribosomes translating in frame the tagged ORF. This specificity results in more accurate quantification of translation, as NaP-TRAP derived translation measurements, correlate well with protein output (Fig. 2G). To demonstrate the capacity of NaP-TRAP to quantify the translation of thousands of reporters simultaneously the present study measure Kozak strength for the first time in the developing zebrafish embryo. Interestingly, the present study observe a weak correlation between Kozak strength and the frequency of Kozak sequences in the genome, an in silica derived Kozak score. This observation highlights the importance of measuring Kozak strength experimentally. Novel Kozak sequences that differ from the consensus zebrafish or vertebrate Kozaks have been identified. Further, the present study has utilized the random forest regression model to predict the Kozak strength of all zebrafish Kozak sequences (position -6 to +4). This prediction will improve the annotation of ORFs in the zebrafish transcriptome as well as assist with the development of more efficiently expressed transgenes.

[0167] Development requires dynamic spatial and temporal control of translation. Using NaP-TRAP the present study has identified cis-regulatory elements that differentially regulate translation in development in the 5’ UTR. For example, the present study has shown for the first time that the zygotically expressed microRNA miR430 represses the translation of reporters containing miR430 seeds in the 5’ UTR. This work adds to a growing body of literature demonstrating the capacity of microRNAs to regulate translation outside of canonical 3’ UTR binding sites. The present study also identified a translation switch driven by pyrimidine rich motifs. At 2 hpf pyrimidine rich kmers suppress translation, whereas at 6 hpf these kmers activate translation. The present study proposes two potential mechanisms to explain this novel 5’ UTR mediated translation control.

[0168] First, the expression or loss of a trans-acting factor (TAFs) that binds pyrimidine rich tracts may drive differential translation. During the matemal-to-zygotic transition the pool of regulatory TAFs is dynamic. The present study proposes two potential candidates for activation of pyrimidine tracts: ( 1) a member of the poly -pyrimidine tract-binding protein family (PTBP) or (2) a component of the EIF3 complex. Previous work has demonstrated that both of these factors are recruited to pyrimidine tracts in the 5’ UTR. (1) While PTBPs have been primarily associated with the translation activation of viral IRESs, recent work from Pan et al has shown how the loss of a PTBP2 binding site in the 5‘ UTR of FGF13 represses capdependent translation in humans. Further, in the developing zebrafish the translation efficiency of PTBP la increases dramatically following zygotic genome activation. (2) The EIF3 complex is a potent regulator of 5' UTR mediated translation control, as it is a component of the scanning 40S ribosome. Several works have suggested that differential expression of EIF3 components during development contributes to the translation of mRNAs in a lineage specific manner.

[0169] Second, recruitment of trans-factors depends not only on the presence of a given cis- regulatory element, but rather the composition of the transcriptome and the pool of trans- factors shapes the regulatory potential of a given element. In fact, the present study observed that the relative importance of U-richness depends on the global rate of translation. When translation is low U-richness is the prominent feature predicting translation (sv40 supplement). In contrast, as translation increases the relative importance of U-richness declines, whereas the relative importance of uORFs increases. It is proposed that this change, reflects a change in the availability of the translation initiation machinery. It is speculated that when the supply of ribosomes is limited, the capacity of the 5?UTR to recruit ribosomes drives translation. In contrast, as the ribosome pool increases the effect of ribosome recruitment diminishes.

[0170] This competition model can also be employed to explain the differential effect of pyrimidine tracts on translation. In the early embryo reporters enriched in U-repeats may recruit the limited supply of ribosomes more efficiently than reporters enriched in pyrimidine tracks. As the supply of translation machinery increases the effect of competition diminishes, resulting in the efficient initiation of pyrimidine rich 5’ UTRs. Repressive TAFs may amplify the effect of competition, as in the absence of scanning 40S ribosomes, these factors can be more readily recruited to the 5‘ UTR. Work from Xiang and Bartel investigating the strong correlation between translation efficiency and poly-A tail length exemplifies this model. More specifically, they report that prior to gastrulation the pool of PAPBC1 is limited, driving competition between transcripts for PAPBC1 occupancy. This competition is eliminated following gastrulation as the relative abundance of PAPBC1 increases following zygotic expression resulting in the decoupling of poly-A tail length and translation efficiency.

[0171] NaP-TRAP also identifies general regulators of translation. In both the developing zebrafish embry o and HEK293T cells, the present study observed poly-U repeats greater than five nucleotides as a general activator of translation initiation. The fact that U repeats are activators of translation across timepoints and experimental systems, suggests that this observation may be driven by a component of the canonical translation initiation machinery.

[0172] Example 1-11: Methods

[0173] Zebraflsh maintenance and mating. Wild-type zebrafish embryos were obtained through natural mating of TU-AB strain of mixed ages (5-18 months). Mating pairs were randomly chosen from a pool of 60 males and 60 females allocated for each day of the month. Fish lines were maintained following the International Association for Assessment and Accreditation of Laboratory Animal Care research guidelines and approved by the Yale University7Institutional Animal Care and Use Committee (IACUC).

[0174] NaP-TRAP reporter controls

[0175] To enable nascent chain immunocapture a 3xFLAG tag was incorporated after the first 18 nucleotides of GFP-3xAID* (auxin inducible domain) using infusion cloning. While AID* domains were included in the initial NaP-TRAP vector to enable the future use of an auxin-inducible degron system, this system was not utilized in this study. The vector also included an SP6 promoter sequence and SV40 poly -adenylation signal. NaP-TRAP reporter mRNAs were generated using the mMESSAGE mMACHINE™ SP6 transcription kit from linearized reporter plasmids (NOT1 restriction enzyme).

[0176] To validate NaP-TRAP, two reporter expenments were performed. First, to quantify morpholino mediated translation repression using NaP-TRAP, 100 pg of 3xFLAG-GFP- 3xAID and 75 pg of dsRED were injected into single cell zebrafish embryos in the presence or absence of 250 uM of a morpholino targeting GFP (Primer X). Embryos were collected at 6 and 24 hpf (25 embryos per NaP-TRAP replicate). NaP-TRAP was performed at 6 hpf whereas immunofluorescence was performed at 24 hpf using Microscope and images were quantified using ImageJ. Second, to assess the capacity of NaP-TRAP to measure microRNA mediated repression, two additional reporters were generated: (1) 3xFLAG-GFP-3xmiR430 and (2) 3xFLAG-GFP-3xmiR204. Three binding sites of either miR430 or miR204 were cloned into the 3’ UTR of 3x-FLAG-GFP-3xAID*. Single cell embry os were injected with 20 pg of the 3xFLAG-GFP-3xmiR430 and 3xFLAG-GFP-3xmiR204, and 160 ng of dsRED mRNAs. Twenty-five embryos per replicate were collected and flash frozen in liquid nitrogen at 2 and 4.3 hpf.

[0177] Following NaP-TRAP. cDNA was synthesized using random hexamer and SuperScript III reverse transcriptase following the manufacturer’s protocol. Translation values were determined by qPCR using the Applied Biosystems power SYBR Green PCR master mix. Levels of reporters in the input and pulldown were normalized using the 2'AACtmethod (ref, PMID: 11846609), where the NaP-TRAP reporter is the target and dsRED the control. Translation was calculated as a ratio of fold enrichment in the pulldown relative to the input.

[0178] NaP-TRAP reporter library assembly

[0179] To eliminate excess cytoplasmic 3xFLAG-GFP, a C-terminal PEST domain was incorporated into the NaP-TRAP reporter plasmid using infusion cloning. The 3xFLAG-GFP vector was amplified from the 3xFLAG-GFP-3xAID* plasmid, whereas the PEST domain insert was generated using PCR overlap extension. For the sake of brevity, 3x-FLAG-GFP- PEST is referred to as 3xFLAG-GFP in the text and figures of this section.

[0180] NaP-TRAP reporter libraries were constructed using three different PCR reactions. First, the common coding sequence and 3’ UTR of the reporters were amplified from the 3xFLAG-GFP plasmid, using a forward primer targeting the N-terminus of GFP and reverse primer targeting the 3’ end of the 3’ UTR. For the Kozak library a forward primer targeting dow nstream of the Kozak sequence was used. All PCR amplicons were gel purified (Monarch).

[0181] Second, reporter libraries were amplified from single stranded oligo pools (KAPA. 20 cycles). PCR overlap extension was used to generate double stranded DNA (0. 1-1 ng of template DNA). For the Kozak library' primer XX w as utilized as a reverse primer, whereas the maternal 5’ UTR and the validation libraries utilized primer XX. After 10 cycles a forward primer, containing an SP6 promoter sequence was added to the PCR reaction. Amplicons were purified using the Zymo DNA Clean & Concentrator (DCC) kit in accordance with the manufacturer’s instructions.

[0182] Third, to generate a template for in vitro transcription, a PCR overlap extension was performed between the reporter library’ and the purified 3xFLAG-GFP-PEST amplicons (KAPA HiFi, 20 cycles). To amplify the reporter library' primers targeting the SP6 promoter sequence and the 3’ end of the 3xFLAG-GFP-PEST amplicons w ere added after 10 cycles of overlap extension. To generate reporters w ith a hard encoded poly-A tail, the reverse primer had a 10 or 60 T 3’ overhang. Unless stated explicitly reporter libraries contained a 60A tail. For libraries with an SV40 poly adenylation signal, a reverse primer targeting the 3’ end of the SV40 pA signal (located directly downstream of the 3’ UTR) w as used for the amplification of 3xFLAG-GFP-PEST and the reporter library.

[0183] Lastly, templates for in vitro transcription were gel purified. Reporter mRNAs were generated using the mMESSAGE mMACHINE™ SP6 transcription kit. In vitro transcribed mRNAs were purified using Monarch kit. Random Kozak library

[0184] The variable region of the random Kozak library consisted of an 15 Illumina adaptor followed the 5’ UTR of xenopus beta-globin, seven random nucleotides, six upstream and one dow nstream of a start codon, and the N-terminus of 3xFLAG GFP.

[0185] Maternal 5 ’UTR library

[0186] A custom single stranded DNA oligo pool, consisting of 11,088 oligos, was ordered from Genscript. Each 170 nt oligo contained an Illumina 15 adaptor sequence, a 124- nucleotide variable region, and 22 nt region with homology to the Kozak sequence and N- terminus of 3xFLAG-GFP. The variable region of the maternal 5’ UTR library was generated using a custom script, tiling the 5’ UTRs of 1,775 maternally supplied transcripts and six IRES sequences (human AQP4, human MYT2, human NRF, human XIAP, EMCV and crTMV) in 124 nucleotide segments every 25 nucleotides.

[0187] Kmer validation library

[0188] A custom singled stranded DNA oligo pool, consisting of 256 oligos was ordered from Twist Biosciences. The design of the common regions of library were identical to that of the maternal 5’ UTR library. The variable region (124 nucleotides) consisted of repeats of all possible tetramers. Each repeat occurred 21 times and was separated by two nucleotides. The nucleotides were repeated in a pattern across the variable region (TC, AC, AG, CG, TC). These dinucleotides were selected to prevent the creation of unintended upstream ORFs.

[0189] NaP-TRAP (Nascent Peptide Translating Ribosome Affinity Purification)

[0190] To capture tagged nascent chains of reporter mRNAs, an immunoprecipitation using anti-FLAG magnetic beads was performed. Magnetic beads were purchased from three suppliers during the course study do to supply chain issues. For both the zebrafish and human cell experiments 10 uL of beads (binding capacity of 15 ug of FLAG peptide) were utilized. The magnetic beads were washed with 800 uL of wash buffer three times prior to being added to the lysis solution.

[0191] Frozen embryos and HEK293T cells were lysed in 500 uL of lysis buffer. After 10 minutes at 4 C the lysate was passed through a 25 G needle (5-10 times). Samples were then centrifuged at 16,000g for 5 minutes at 4 C. The supernatants were transferred to a new Eppendorf tube and 2 uL of DNASEI were added. After a 15-minute incubation at 4 C, the samples were diluted to 1 mL using additional lysate buffer and 75 uL of lysate was collected from each sample to serve as an input. Samples were placed on a rotator at 4 C for 2 hrs. Following incubation, the beads were washed with 800 uL of wash buffer three times. The beads (pulldown) and the inputs were then resuspended in 1 mL of Trizol. For the maternal 5’ UTR and tetramer validation libraries, spike in reporters w ere added. Trizol extractions w ere performed in accordance with the manufacture’s protocol. RNA pellets were resuspended in 11 uL of nuclease free H2O.

[0192] Reporter library preparation

[0193] Reverse primers (4 uM total) targeting the N-terminus of 3xFLAG GFP were added to purified input and pulldown RNAs. These primers contained a 6 nt sample barcode and 10 nt Unique Molecular Identifier to allow for demultiplexing and read deduplication as w ell as a 3’ 17 Illumina adaptor sequence. To increase library complexity the primer pairs w ere staggered by 1 nt. Reverse transcription was performed using the Superscript III kit in accordance with manufacturer’s instructions. Reverse transcription reactions were performed at 55 C. cDNA from replicates were pooled and purified using AMPURE XP beads. Illumina 15 and 17 forward and reverse primers containing a 10 nucleotides indexes were utilized to amplify cDNA libraries via PCR (Kappa Polymerase Master Mix). To reduce the number of PCR duplicates 12-18 cycles were utilized. Amplicons were purified by using an Agarose Gel Electrophoresis followed by size selection and gel extraction (Monarch Kit). Libraries were sequenced using NovaSeq 6000.

[0194] Read trimming and translation calculation

[0195] Paired-end reads were trimmed and demultiplexed using ReadKnead (github). Barcodes identifying replicates and UMIs w ere extracted from read two. Common regions w ere trimmed from read one to facilitate accurate mapping. For the Kozak library reporters were counted using a custom python script. Reporter with indels in the Kozak sequence or reporters without an AUG in the appropriate position were eliminated. For the maternal 5’ UTR and validation libraries reads were mapped to a library specific index using Bowtie2 (command here). PCR duplicates were eliminated using UMIs. UMIs were considered identical if they had a Hamming Distance less than 1. Reads for each experiment w ere normalized by dividing the read counts of each reporter by the sum the total number of reads mapped to the spike-ins. In the absence of spike-ins read counts were normalized reads based on the total number of mapped reads per replicate (reads per million, RPM). Differential Motif Enrichment Analysis

[0196] Reporters were ranked based on their translation at 2 and 6 hpf. Using the sum and difference of these rankings across timepoints four groups of reporters were generated: (1) repressed, (2) active, (3) early activating, and (4) late activating. Repressed and active reporters constituted the top and bottom 10% of reporters based on the sum of their ranks at 2 and 6 hpf respectively, whereas the early activated and late activated groups were the top and bottom 10% of groups based on the difference between their ranks at 2 and 6 hpf. A differential motif enrichment analysis was performed on each group. Fold enrichment values were determined by dividing the count of each kmer in the reporter group by the count of the kmer in the library, whereas the significance of the fold-change was determined using a hypergeometric test (Bonferroni corrected p-value).

[0197] Random Forest Regression model

[0198] Random forest regression models were employed to predict translation. For the Kozak library features were generated by one-hot encoding positions -6 to -1 and position +4 of each reporter sequence. In contrast, for the maternal 5’ UTR library Kmer counts (1-8 nucleotides) and uORF features were generated using a custom python script. For the maternal 5’ UTR models features were filtered by calculating a spearman rank correlation coefficient (SPR) between each feature and translation. Features that had a correlation greater than 0.05 or less than -0.05 were included in the Random Forest Model.

[0199] To prevent model overfitting the data were divided randomly into two groups: a test and training set, 20 % and 80 % of the reporters respectively. The training data were divided into five different groups of equal size and then trained the model on four of the groups, using the fifth group as test set to measure the accuracy of the model. The present study trained each model five times using each of the groups as a test set. To quantify model performance, the present study performed a linear regression between the model prediction and the experiment data of the test set. taking the mean prediction across each of the test sets.

[0200] Dual luciferase assay

[0201] Single cell zebrafish embryos were co-injected mRNAs encoding nano and firefly luciferase (0.5 pg of nano I 19.5 pg of firefly). Embryos were collected at 6 hpf and frozen in liquid nitrogen (5 embryos per replicate). Nano and firefly luciferase activity were measured using the Nano-Gio® Dual-Luciferase® Reporter Assay System. Firefly luciferase activity was utilized to normalized nano luciferase measurement across reporters.

[0202] Data Analysis

[0203] All custom scripts were written in Python. Plots were generated using the Matplotlib package. Statistical analysis performed using the scipy package, whereas the random forest analysis was performed using the scikit-leam package. Feature and experimental data were stored using an SQLite database in conjunction with the HDF5 fde format. All scripts can be accessed NaP-TRAP repository on Github.

[0204] Example 2: NaP-TRAP Reveals the Regulatory Grammar in 5'UTR-Mediated Translation Regulation During Zebrafish Development

[0205] Example 2 section is directed to a further study based on that of Example 1. In one non-limiting aspect, the study of Example 2 overlaps with that described in Example 1, albeit includes more extensive data.

[0206] The cis-regulatory elements encoded in a mRNA determine its stability and translational output. While there has been a considerable effort to understand the factors driving mRNA stability , the regulatory frameworks governing translational control remain elusive. The present study has developed a novel massively parallel reporter assay (MPRA) to measure mRNA translation. Nascent Peptide Translating Ribosome Affinity Purification (NaP-TRAP). NaP-TRAP measures translation in a frame specific manner through the immunocapture of epitope tagged nascent peptides of reporter mRNAs. In contrast to existing MPRA methods, NaP-TRAP does not require specialized equipment and is readily adaptable to steady-state and dynamic model systems. The present study has employed NaP-TRAP to quantify Kozak strength and the regulatory landscapes of 5’ UTRs in the developing zebrafish embryo and in human cells, characterizing general and developmentally dynamic cis-regulatory elements. To this end, the present study identifies U-rich motifs as general enhancers, and upstream ORFs and GC-rich motifs as global repressors of translation. The present study also observes a translational switch during the matemal-to-zygotic transition, where C-rich motifs shift from repressors to prominent activators of translation. Conversely, the present study shows that microRNA sites in the 5’ UTR repress translation following the zygotic expression of miR-430.

[0207] Together these results demonstrate that NaP-TRAP is a versatile, accessible, and powerful method to decode the regulatory functions of UTRs across different systems. Example 2-1:

[0208] Life requires spatial and temporal control of protein expression. The protein output of a given transcript reflects the integration of translation and mRNA stability. While there has been a considerable effort to understand the factors driving mRNA synthesis, maturation, and decay, the regulatory frameworks governing translational control remain elusive. Cis- regulatory elements encoded in the sequence and structure of mRNAs modulate these processes. Although these elements are distributed throughout the transcript, they are often concentrated in the regions up- and dow nstream of the main coding sequence, the 5’ and 3’ untranslated regions (UTRs) respectively. Given that initiation is the rate limiting step of translation, there has been a particular interest in understanding the regulatory elements found in 5’ UTRs. These elements include internal ribosome entry’ sites (IRESs), G-quadruplexes, iron response elements (IREs), and upstream Open Reading Frames (uORFs).

[0209] To investigate translation control, several studies have utilized ribosome profiling. While this method accurately quantifies the translation efficiency of individual genes, its capacity to characterize the function of cis-regulatory elements is limited. The generation of ribosome protected fragments decouples the translation measurement of a given mRNA from its cognate untranslated regions. Thus, measurements of translation efficiency often reflect the amalgamation of several isoforms of a given gene, each with a unique set of regulatory elements. This process is complicated further by the fact that endogenous transcripts contain multiple regulatory elements that exert differential and often competing effects on translation.

[0210] Massively parallel reporter assays (MPRAs) are particularly suited to address these challenges. MPRAs measure the abundance and / or translation efficiency of thousands of reporters simultaneously. These assays have improved the understanding of translation control and mRNA stability. Previous works have characterized Kozak strength, uORFs, IRESs, codon optimality, RNA structure, microRNA binding sites, and the effects of variation in human UTRs. These assays have also identified novel sequence motifs driving translation and decay. Despite these insights, conclusions based on MPRAs have been limited by the methods used to measure translation: growth-selection, fluorescent-cell sorting (FACS), polysome profiling, and direct analysis of ribosome targeting (DART) (Fig. 16). In MPRAs where reporters are encoded in DNA, measurements of translation can be confounded by additional layers of transcriptional regulation. In polysome fractionation, which quantifies the number of ribosomes on each transcript, inactive ribosomes, ribosomes translating out of frame, or ribosomes translating ORFs outside of the coding sequence may skew translation measurements. Additionally, the methodological complexity of these assays has reduced the use of translation-based MPRAs across diverse systems.

[0211] Here, the present study develops NaP-TRAP (Nascent Peptide Translating Ribosome Affinity Purification), a novel method to measure in-frame translation through the immunocapture of nascent peptides. More specifically, by including an N-terminal FLAG tag in the coding sequences of reporter mRNAs, the present study enriches for reporters in a manner proportional to the number of active ribosomes translating the main ORF in frame. The present study validates NaP-TRAP using reporter assays of increasing complexity. First, the present study measures the translation of individual reporters containing know n cis- regulatory elements. Next, the present study quantifies Kozak strength in zebrafish bymeasuring the translation of thousands of reporters simultaneously. Lastly, the present study assesses the regulatory potential of endogenous 5’ UTR sequences in the developing zebrafish embryo and HEK293T cells to identify common and developmentally modulated motifs that regulate translation in vertebrates. In doing so the present study demonstrates that NaP-TRAP is an accessible, versatile, and quantitative method, which has the capacity to measure translation control across multiple systems.

[0212] Example 2-2: NaP-TRAP measures translation via the immunocapture of nascent chains

[0213] Cis-regulatory elements encoded in mRNAs modulate translation efficiency. To measure the regulatory' potential of these elements the present study developed a novel massively parallel reporter assay (MPRA), NaP-TRAP (Nascent Peptide Translating Ribosome Affinity Purification). It was reasoned that the present study could enrich for reporters in a manner proportional to their translation efficiency through the immunocapture of FLAG-tagged nascent chain complexes immobilized by cycloheximide treatment (Fig. 8A). To this end, the present study measured translation as a ratio of the amount of a reporter mRNA in the pulldown relative to its input. This approach decouples translation measurements from differences in mRNA abundance.

[0214] To evaluate the capacity of NaP-TRAP to measure translation the present study performed tw o reporter assays. First, the present study measured the effect of a translation blocking morpholino. The present study co-injected 3xFLAG-GFP and dsRED mRNAs into single cell zebrafish embry os in the presence or absence of a translation blocking morpholino, targeting the start codon of 3xFLAG-GFP and performed NaP-TRAP at 6 hpf. To measure the amount of 3xFLAG-GFP reporter mRNA present in the input and pulldown fractions of each anti-FLAG immunoprecipitation, the present study employed reverse transcription- quantitative polymerase chain reaction (RT-qPCR). Given that the dsRED reporter mRNA was neither enriched by the anti-FLAG immunoprecipitation nor targeted by the translation blocking morpholino, the present study utilized the relative abundance of the dsRED reporter mRNA in each fraction to normalize NaP-TRAP derived translation values across experimental conditions. Using NaP-TRAP the present study observed a 48-fold decrease in translation in the presence of the morpholino at 6 hours post fertilization (hpf) (Fig. 8B). These results are consistent with the translational repression in morpholino injected embryos observed by the decrease in fluorescence (GFP / dsRED) at 24 hpf (Fig. 8B).

[0215] Next, the present study tested the capacity of NaP-TRAP to capture dynamic translation control. In the developing zebrafish embryo, the microRNA miR-430 is one of the first zygotically transcribed genes. At 4.3 hpf (hours post fertilization), the expression of miR-430 results in the translation repression, deadenylation, and eventual decay of targeted mRNAs. To quantify this effect, the present study measured the translation of 3xFLAG-GFP reporters with partial complementarity to two different microRNAs in their 3’ UTRs, miR- 430 (3xFLAG-GFP-3xmiR-430) and miR-204 (3xFLAG-GFP-3xmiR-204). The present study selected miR-204 target sites as a control given that this microRNA is not expressed in the early embryo. Using NaP-TRAP the present study measured the translation of mRNA reporters at 2 and 4.3 hpf, before and after miR-430 expression. As described elsewhere herein, the present study utilized the abundance of the dsRED control RNA to normalize translation values across experimental conditions. Between the two timepoints, the translation of the 3xmiR-430 reporter decreased by ~4.3 fold, whereas the translation of the 3xmiR-204 reporter increased by ~2.5 fold (Fig. 8C). These results are consistent with the role of miR- 430 in translational repression and the global increase in translation observed during the early stages of development.

[0216] Together these results demonstrate that NaP-TRAP measures the translation of individual reporters targeting general and developmentally dynamic cis-regulatory elements. Further, these experiments highlight the versatility of the method as they demonstrate the capacity of NaP-TRAP to quantify cis-regulation mediated by the 5’ and 3’ UTRs.

[0217] Example 2-3: Using NaP-TRAP to investigate Kozak strength in the developing zebrafish embryo

[0218] Next, to evaluate the capacity of NaP-TRAP to function as a MPRA, the present study designed a library to quantify the regulatory potential of the Kozak sequence. The present study elected to measure Kozak strength for two reasons: (1) the Kozak sequence is a strong determinant of translation initiation and (2) Kozak strength is often inferred based on the frequency of these sequences in the genome. To generate a library of diverse Kozak sequences, the present study incorporated six random nucleotides upstream and one random nucleotide downstream of the AUG of the 3xFLAG-GFP reporter (Fig. 9A). The present study injected the in vitro transcribed mRNA library’ into single cell zebrafish embryos and measured the translation using NaP-TRAP at 6 hpf. The present study observed a strong correlation in translation values between replicates (Pearson’s r = 0.78, Fig. 17A). To identify active and repressive Kozak sequences, the present study generated position weight matrices for the top and bottom 10 % of reporters based on their translation (Fig. 9A). The present study observed that active Kozak sequences were enriched in C’s and A’s, whereas repressive Kozak sequences were enriched in U’s and G’s.

[0219] To examine the effect of nucleotide identity and position on translation the present study employed a random forest regression model (RFM). Model features were generated based on nucleotide identities and positions within the Kozak sequence (Fig. 9B). Prior to training the model, the present study divided the data into a test and a training set, 30% and 70% of reporters respectively. Using the training set, the present study employed a 5-fold cross validation to optimize model parameters. The model’s prediction explained 73% of the variance observed in the test set data (Fig. 9C). Given this strong correlation the present study employed the model to predict the translation of all possible Kozak sequences (Fig. 17C). To identify' the features informing the predictive power of the model, the present study performed a permutated feature importance analysis. Briefly, the present study shuffled the identities of the features of the model one at a time and measured the effect of the change on the predictive power of the model. At 6 hpf bases G, G, U, and G in positions -5. -4, -3, and - 2 were the most repressive features, whereas C, A / C, and A in positions -4, -3, +4 were the most activating features (Figs. 9D and 17B).

[0220] The present study compared NaP-TRAP translation values to an in silico derived metric, Kozak score. In zebrafish Kozak strength has been previously inferred based on the frequency of a given Kozak sequence within the transcriptome. This metric associates Kozak sequence abundance with increased translation initiation. The experimentally derived translation values challenge this hypothesis as the present study observed a weak correlation between NaP-TRAP and Kozak score (R = 0.31, Fig. 9E). To validate this conclusion, the present study identified sequences that had similar Kozak scores (300 ± 5) yet exhibited different NaP-TRAP derived translation values and measured their translation using a dual luciferase reporter assay (Fig. 9F). The luciferase translation values correlated strongly with NaP-TRAP (Pearson's R 0.88; Fig. 9G). This result not only highlights the importance of measuring Kozak strength experimentally, but also demonstrates that NaP-TRAP derived translation values correlate well with protein production.

[0221] Taken together, these results demonstrate that NaP-TRAP can be employed as a MPRA. By measuring the translation of thousands of reporters simultaneously, the present study quantifies and model Kozak strength in a vertebrate system.

[0222] Example 2-4: Investigating the regulatory potential of endogenous 5’ UTRs

[0223] Upon validating NaP-TRAP's capacity to function as a MPRA, the present study investigated the regulatory complexity of endogenous 5' UTRs during the early stages of embryogenesis. Prior to zygotic genome activation, the translation of maternally supplied mRNAs drives development. Given that initiation is the rate limiting step of translation, 5’ UTRs provide a mechanism for maternally supplied mRNAs to modulate their protein output. To study the cis-regulatory elements encoded in these mRNAs, the present study designed an 11.088-sequence synthetic oligo library (124-nt long) by tiling the 5’-UTRs of 1.725 zebrafish genes every 25 nucleotides (Fig. 10A). As the present study were interested in identify ing specific sequence elements, which modulated differential translation regulation the present study elected to keep the 5’ UTR length and Kozak sequence of reporter mRNAs constant as both of these features correlate with translation efficiency in the developing embryo. The present study injected the in vitro transcribed library into single cell zebrafish embryos and measured translation at 2 hpf and 6 hpf using NaP-TRAP. The addition of an internal spike-in at the RNA extraction step allowed us to quantify relative changes in translation across developmental stages (Fig. 10A). Translation values were strongly correlated across replicates (Pearson’s R >= 0.90; Figs. 18A-18B and 19A-19B) and with protein abundance at both time points (dual luciferase assay; 2 hpf Pearson’s R 0.75, 6 hpf Pearson’s R 0.88, Figs. 18C-18D). The present study observed a mean increase in translation between 2 hpf and 6 hpf of ~1.9 fold, consistent with the gradual de-repression of translation during the first 24 hours of development (Fig. 19B).

[0224] Next, the present study utilized a random forest regression model trained on sequence elements and features characterizing upstream AUGs (uAUGs) to predict translation (Figs. 10C and 19D-19I). These models explained 74% of the variance observed in the test set data at both timepoints (Figs. 10D-10E). Using a permuted feature importance analysis, the present study identified the number of upstream uORFs (uORFs) as well as the Kozak strength of uORFs and out of frame overlapping ORFs (ooORFs) as the features contributing most significantly to the predictive power of the model at both timepoints (Figs. 10F and 10G). When the present study examined the correlation of selected features with translation, the present study observed that U-repeats activated translation and G-repeats and uORFs suppressed translation at both timepoints. Interestingly, the present study also observed that C-rich k-mers suppressed translation at 2 hpf, whereas at 6 hpf C- and CU-rich k-mers activated translation (Figs. 10H and 101). Given the prominent role of upstream open reading frames in the feature importance analysis, the present study repeated the random forest analysis on reporters that contained no upstream start codons (Figs. 18E-18J). In the absence of upstream start codons, the importance and prominence of C- and CU-rich k-mers increase at 6 hpf (Fig. 18 J).

[0225] Altogether these results demonstrate the following: (1) uAUGs are the most prominent repressive elements during early embryogenesis, and (2) 5’ UTR regulatory landscapes are dynamic during development.

[0226] Example 2-5: Identifying sequences driving differential translation

[0227] To identify the sequence elements modulating differential translation during development the present study divided the 5’ UTR reporters into four groups based on their relative translation at 2 and 6 hpf: (1) repressed, (2) active, (3) repressed post ZGA, and (4) active post ZGA (Figs. 11 A and 19J). Next, the present study performed a differential pentamer enrichment analysis on each group relative to the reporter library. This analysis revealed that repressed reporters were significantly enriched in upstream ORFs and GC-rich pentamers and depleted in U-rich tracks (p >10-5 hypergeometric test, Figs. 1 IB and 19K). Conversely, active reporters were enriched in U-rich pentamers (UUUUU, CUUUU. UUUUA, GUUUU) and depleted in upstream start codons (Figs. 11C and 19L). Reporters that were more highly translated after genome activation were enriched in C-rich pentamers (CUCUC, CUCCC, CCAUC, CCUCC) and depleted in U-repeats (Figs. 1 ID and 19M). In contrast, reporters that were repressed post ZGA were enriched in AG / UG-rich sequences (UAGUG, UAUUG. AAGAA, AGACU). as well as sequence motifs complementary to the seed site of the microRNA miR-430 (GCACU and GCACUU) and depleted in uORFs and pyrimidine repeats (Figs. 1 IE and 19N).

[0228] To validate these results, the present study generated a library of 5’ UTR sequences containing all possible tetramer repeats separated by dinucleotide spacers and measured translation using NaP-TRAP at 2 and 6 hpf (Fig. 1 IF). Next, the present study plotted the translation values of the validation library labeling each reporter on the basis of whether its encoded tetramer repeat was enriched in the reporter groups described above (Fig. 11G). The distributions of reporters encoding repressed, repressed post ZGA, and active post ZGA tetramers largely recapitulated the distributions of their respective groups in the 5’ UTR library7(Figs. 20B-20E). Consistent with this, a cumulative analysis of the translation (2 hpf / 6 hpf) of reporters enriched in “active motifs post ZGA'’ revealed a larger increase in translation compared to those enriched in “repressed motifs post ZGA” (Fig. 11H).

[0229] Altogether these results suggest that there are general and dynamic sequence motifs in zebrafish 5’ UTRs that modulate translation during the early stages of development.

[0230] Example 2-6: Complementary sequences to miR-430 in the 5’ UTR suppress translation miR-430 is one of the first zygotically transcribed genes in the developing zebrafish embryo. The enrichment of miR-430 seeds in reporters that were repressed after ZGA suggests that the maternal -to-zygotic transition is shaped by 5’ UTR mediated translation control. To determine whether this repressive effect is specific for miR-430, the present study compared the translation between 2 hpf and 6 hpf for reporters containing seeds for miR-430 or miR-1, a microRNA that is not expressed in the early stages of development. The present study observed a significant decrease in the translation of miR-430 containing reporters (6 vs 2 hpf) relative to the translation of reporters lacking either seed. In contrast, there was a slight increase in translation of miR-1 control reporters (Fig. 12A). To investigate the mechanism driving miR-430 mediated translation repression, the present study measured the degree of complementarity7between the microRNA and the targeted reporter. Translation of the miR- 430 reporters was negatively correlated to miRNA-5’ UTR complementarity (Fig. 12B). Whereas translation of the miR-1 control reporters was not significantly correlated with miRNA-5’ UTR complementarity (Fig. 12C). This observation suggests that 5’ UTR microRNA mediated translation repression is driven by' sequences with homology beyond the seed site of the microRNA.

[0231] To demonstrate further that miR-430 expression drives the translation repression of 5’ UTRs with seed sites, the present study constructed two nano-luciferase reporters: 4xmiR- 430-nanoluc and 4xmiR-430-MUT-nanoluc. The present study shuffled the miR-430 sites in the 4xmiR-430-MUT reporter to prevent miR-430 targeting (GCACUU to GCUCUA). The present study injected these reporter mRNAs together with a firefly luciferase control mRNA into wild-ty pe and miR-430 - / - mutant embryos and measured relative luciferase activity at 6 hpf (Fig. 12D). The present study observed a ~2-fold decrease in the relative luciferase activity of the 4xmiR-430 reporter in the wild-type condition when compared to the mutant. In contrast, the present study observed no significant difference in luciferase activity between the shuffled reporters when comparing the mutant and wild-type embryos (Fig. 12E). This finding suggests that the zygotic expression of microRNA miR-430 represses the translation of mRNAs with 5’ UTR seed sites.

[0232] Taken together these results demonstrate that miRNAs that target sites in the 5‘ UTR can provide significant translational repression in vivo.

[0233] Example 2-7 : The developmentally dynamic role of C-rich motifs

[0234] The present study also observed an enrichment of C-rich pentamers in reporters that were active following ZGA. To explore this observation in greater detail, the present study measured the mean difference in translation between reporters enriched and depleted in each feature (k-mers < 4) that lacked upstream open reading frames. Next, the present study compared the rank of features with a significant difference in translation at 2 hpf and 6 hpf. The present study observed that the k-mers C. CC, CCC, and UCC were some of the most repressive features at 2 hpf yet activated translation at 6 hpf (Fig. 13 A). To investigate this differential regulation further, the present study selected two active (U-rich) and two active post ZGA (C-rich) reporters from the 5’ UTR library. To determine the role of C-rich sequences in this translation regulation the present study mutated U’s in the active reporters to C’s and mutated C’s in the active post ZGA reporters to U’s for a total of eight reporters: four wild-type and four mutants, respectively. The present study injected the reporters into single cell zebrafish embry os and measured translation at 2 hpf and 6 hpf using NaP-TRAP. At 2 hpf the present study observed that mutating C’s to U’s in the C-rich reporters enhanced translation (Fig. 13C), whereas mutating U’s to C’s in U-rich reporters repressed translation (Fig. 13D). In contrast, this effect was largely reduced at 6 hpf (Figs. 13E and 13F). These results support the findings of the differential enrichment analysis and indicate that there is a translational switch driven by C-rich sequences following ZGA (Fig. 13B).

[0235] Altogether, these results validate novel developmentally dynamic mechanisms of 5’ UTR mediated translation control, demonstrating the importance of quantifying cis-regulation in non-steady state systems.

[0236] Example 2-8: NaP-TRAP can be readily adapted to human cells

[0237] To demonstrate the broad applicability of the method herein, the present study adapted NaP-TRAP to human cells. To this end, the present study transfected the in vitro transcribed mRNA 5’ UTR library into HEK293T cells and measured translation at 12 hours post transfection (hpt) using NaP-TRAP (Fig. 14A and 14B). Replicates were strongly correlated (Pearson’s R > 0.92, Figs. 21 A-21C). To identify motifs modulating translation in HEK293T cells the present study performed a differential enrichment analysis on active and repressed reporters. Repressed reporters were enriched in AUG-containing motifs and depleted in U-rich motifs (Fig. 14C), whereas active reporters were enriched in C-rich pentamers and depleted in AUG containing motifs (Fig. 14D), consistent with those observed in zebrafish at 6 hpf (Figs. 14E and 21D).

[0238] Next, the present study performed a differential enrichment analysis comparing the human data to each of the zebrafish timepoints (Figs. 14F-14J and 21E-21I, 2 hpf and 6 hpf, respectively). The present study divided the reporters into four groups: (1) repressed. (2) active, (3) active in zebrafish, and (4) active in HEK293T cells. From these analyses, the present study conclude that uAUGs are general repressors of translation (Figs. 14G and 21F) and U-rich motifs are general activators of translation (Figs. 14H, 21G and 21L). The present study also observed that reporters that are active only in HEK293T cells are enriched in G- rich k-mers, whereas reporters active at 2 hpf or 6 hpf are depleted in these motifs (Figs. 14J and 211). Next, when the present study employed a random forest model to predict translation. The model explained 80% of the variation in the experimental data (Fig. 21J). Consistent with the results observed in zebrafish, Kozak strength and number of upstream AUGs were the most predictive features of translation in HEK293T cells, exhibiting a strong negative correlation with translation (Figs. 21K and 21L).

[0239] Altogether these results demonstrate that (1) NaP-TRAP is a robust method that can measure translation across multiple model systems, (2) upstream AUGs are a dominant driver of 5’ UTR mediated translation control in HEK293T cells, and (3) U-rich motifs are general activators of translation.

[0240] Example 2-9:

[0241] Here, the present study develops a novel MPRA method, NaP-TRAP. and demonstrate its capacity to measure translation quantitatively. The present study uses NaP- TRAP to characterize cis-regulatory elements in the 5’ and 3’ UTRs, the strength of the Kozak sequence, and the regulatory landscape of the 5’ UTRs of mRNAs in the developing zebrafish embryo and HEK293T cells. Using NaP-TRAP the present study has identified general and developmentally dynamic cis-regulatory elements, as well as characterized global changes to translation associated with early embryogenesis. NaP-TRAP is an accessible, versatile, and quantitative MPRA method to measure translation. In contrast to existing approaches NaP-TRAP does not require a large amount of input material or specialized equipment, making the method adaptable to a wide range of model systems, different cell types, and physiological states (Fig. 16). Further, NaP-TRAP measures frame-specific translation as only the nascent chains of ribosomes translating the main ORF are FLAG-tagged and thereby immunocaptured. This approach is particularly important in systems with a low basal level of translation (e.g., the early stages of vertebrate embryogenesis and neurons) and results in translation measurements that reflect protein output (Figs. 9G and 17C-17D). In polysome profiling and other TRAP methods, inactive ribosomes or ribosomes translating outside of the main open reading frame may affect the measurements of translation. By enriching for reporters through the immunocapture of FLAG-tagged nascent chain complexes, NaP-TRAP measures translation in a manner that is proportional to the number of ribosomes actively translating the tagged ORF. Second, by injecting or transfecting in vitro transcribed mRNA reporter libraries, NaP-TRAP eliminates the confounding effects of transcriptional regulation.

[0242] Previous studies have assumed that the frequency of a given Kozak sequence within the genome correlates strongly with its effect on translation. While this hypothesis has been challenged by MPRAs performed in cell culture, to date there is a lack of measurements of the effect of Kozak strength associated with MPRAs that introduce reporters as DNA. When quantifying the regulatory potential of the 5‘ UTR, transcriptional bias may be more pronounced given the region’s proximity to the promoter sequence on translation in zebrafish. To demonstrate the capacity of NaP-TRAP to quantify the translation of thousands of mRNA reporters simultaneously, the present study measured the Kozak strength in zebrafish embryos. In support of this approach, the present study observed a weak correlation between Kozak strength and an in silico derived Kozak score (Fig. 9E). While the approach has identified conserved activators and repressors (-3 A and -31-2 U / G) of translation, the present study has also characterized novel Kozak sequences that differ from the zebrafish or vertebrate consensus Kozaks, including the activating effect of adenosine in the +4 position. Further, the present study has utilized the random forest regression model to predict the Kozak strength of all zebrafish Kozak sequences (positions -6 to +4). This prediction will improve the annotation of Kozak strength in the zebrafish, as well as the capacity to modulate protein output.

[0243] Development requires dynamic spatial and temporal control of translation. Using NaP-TRAP the present study has identified 5’ UTR cis-regulatory elements that differentially regulate translation in development. For example, the present study shows that the zygotically expressed microRNA miR-430 represses the translation of reporters containing miR-430 seeds in the 5’ UTR.

[0244] The present study has also identified a translation switch driven by C-rich motifs. At 2 hpf C-rich k-mers suppress translation, whereas at 6 hpf these k-mers activate translation. The present study proposes two potential mechanisms to explain this novel 5’ UTR mediated translation control (Fig. 13B). First, the expression or loss of a trans-acting factor (TAFs) that binds pyrimidine rich tracts may drive differential translation. Members of the polypyrimidine tract-binding protein family (PTBP) and components of the EIF3 complex have been shown to activate translation by binding pyrimidine tracts in the 5’ UTR. Second, recruitment of trans-factors depends not only on the presence of a given cis-regulatory element, but rather the composition of the transcriptome and the pool of available trans- factors shapes the regulatory potential of a given element. The present study observes that the relative importance of U-rich sequences depends on the global rate of translation. When translation is low, the prevalence of U-rich sequences is a prominent predictor of translation (Figs. 19D-19F). In contrast, as translation increases, the relative importance of U-nch sequences declines, whereas the relative importance of uORFs increases (Fids. 19G-19I). The present study proposes that this change reflects a change in the availability7of the translation initiation machinery. The present study speculates that when the supply of ribosomes is limited, the capacity of the 5?UTR to recruit ribosomes drives translation. In contrast, as the ribosome pool increases, the effect of 5’ UTRs on ribosome recruitment diminishes (Fig. 13B).

[0245] This competition model can also be employed to explain the differential effect of C- rich tracts on translation. In the early embryo, reporters enriched in U-repeats may recruit the limited supply of ribosomes more efficiently than reporters enriched in C’s. As the supply of translation machinery increases, the effect of competition diminishes, resulting in the efficient initiation of C-rich 5’ UTRs (Fig. 13B). Repressive trans-acting factors (TAFs) may amplify the effect of competition, as in the absence of scanning 40S ribosomes, these factors can be more readily recruited to the 5’ UTR.

[0246] NaP-TRAP also identifies conserved regulators of translation. In both the developing zebrafish embryo and HEK293T cells, poly-U repeats activate translation whereas uORFs repress translation.

[0247] Example 2-10: Materials and Methods Zebrafish maintenance and mating

[0248] Wild-type zebrafish embryos were obtained through natural mating of TU-AB strain of mixed ages (5-18 months). Mating pairs were randomly chosen from a pool of 60 males and 60 females allocated for each day of the month.

[0249] Hek293T cells

[0250] HEK293T cells were grown in a media consisting of Dulbecco’s Modified Eagle Medium (DMEM) (ThermoFisher Scientific #10569010), 10% heat inactivated Fetal Bovine Serum (FBS) (ThermoFisher Scientific #16140071), 20 rnM HEPES (ThermoFisher Scientific #15630080). 2 rnM L-Glutamine (ThermoFisher Scientific #25030081), lx Penicillin / Streptomycin (ThermoFisher Scientific #15140122) at 37°C and 5% CO2. mRNA transfections were performed using Lipofectamine™ MessengerMAX™ Transfection Reagent (ThermoFisher Scientific #LMRNA008) in accordance with the manufacture’s protocol.

[0251] NaP-TRAP reporter controls

[0252] To enable nascent chain immunocapture, a 3xFLAG tag was incorporated after the first 18 nucleotides of GFP-3xAID* (auxin inducible domain) using an In-Fusion® HD Cloning kit (Takara #638946) (F: 3xFLAG_inf_fwd; R: AID_inf_rev). While AID* domains were included in the initial NaP-TRAP vector to enable future use of an auxin-inducible degron system, this system was not utilized in this study. The vector also included an SP6 promoter sequence and SV40 poly-adenylation signal. NaP-TRAP reporter and dsRED control (plasmid pCS2+-dsRED) mRNAs were generated using a mMESSAGE mMACHINE™ SP6 transcription kit (Invitrogen™ #AM1340) from linearized reporter plasmids via Notl-HF® (NEB #R3189L) restriction enzyme digest. In vitro transcribed mRNAs were purified using a Monarch® RNA Cleanup Kit (NEB #T2040L) prior to injection.

[0253] To validate NaP-TRAP, two reporter experiments were performed. First, to quantify morpholino mediated translation repression using NaP-TRAP, 100 pg of 3xFLAG-GFP- 3xAID mRNA and 75 pg of dsRED mRNA were injected into single cell zebrafish embryos in the presence or absence of a morpholino targeting the start codon of 3xFLAG-GFP (250 pM, GFP-MO 5 - ACAGCTCCTCGCCCTTGCTCACCAT-3’, SEQ ID NO 2365. Gene Tools LLC). Embryos were collected at 6 and 24 hpf (25 embryos per NaP-TRAP replicate). Embryos collected for NaP-TRAP were flash frozen in liquid nitrogen prior to sample processing. NaP-TRAP was performed at 6 hpf whereas immunofluorescence was measured at 24 hpf. Images were quantified using ImageJ.58 Second, to assess the capacity of NaP- TRAP to measure microRNA mediated repression in the 3’ UTR, two additional reporters were generated: (1) 3xFLAG-GFP-3xmiR-430 and (2) 3xFLAG-GFP-3xmiR-204. Three binding sites of either miR-430 or miR-204 were cloned into the 3’ UTR of 3x-FLAG-GFP- 3xAID*, using an In-Fusion® HD Cloning kit (Takara #638946) (F: FGFP inf fwd; R: 3xmir430_inf_rev and 3xmir204_inf_rev. respectively). Single cell embryos were injected with 20 pg of both the 3xFLAG-GFP-3xmiR-430 and 3xFLAG-GFP-3xmiR-204 mRNAs, as well as 1 0 pg of dsRED mRNA. Twenty-five embryos per replicate were collected and flash frozen in liquid nitrogen at 2 and 4.3 hpf.

[0254] For methods detailing NaP-TRAP and qPCR translation measurements see sections: (1) NaP-TRAP (Nascent Peptide Translating Ribosome Affinity Purification) and (2) NaP- TRAP qPCR analysis, respectively.

[0255] NaP-TRAP reporter library assembly

[0256] To eliminate excess cytoplasmic 3xFLAG-GFP. a C-terminal PEST domain was incorporated into the NaP-TRAP reporter plasmid using an In-Fusion® HD Cloning Kit (Takara #638946). The 3xFLAG-GFP vector was amplified from the 3xFLAG-GFP-3xAID* plasmid (F: GFP ddl inf fwd, R: GFP ddl inf rev), whereas the PEST domain insert was generated using PCR overlap extension (F: PEST fwd, R: PEST rev). For the sake of brevity, 3x-FLAG-GFP-PEST is referred to as 3xFLAG-GFP in the text and figures of this section unless stated otherwise.

[0257] NaP-TRAP reporter libraries w ere constructed using three different PCR reactions: First, the common coding sequence and 3’ UTR of the reporters were amplified from the 3xFLAG-GFP plasmid, using a forward primer targeting the N-terminus of GFP and a reverse primer targeting the 3’ end of the 3’ UTR (F: 3xFLAG_GFP_fwd, pA-R: pcr_II_pA_rev, sv40-R: sv40_rev). For the Kozak library, a forward primer targeting the sequence immediately downstream of the variable Kozak sequence was used (F: ntrapK_GPF_fwd). All PCR amplicons were gel punfied using a Monarch® DNA Gel Extraction Kit (NEB #T1020L).

[0258] Second, reporter libraries w ere amplified from 1 ng of single stranded DNA oligo pools using KAPA HiFi HotStart ReadyMix (Roche #7958935001) for 10-20 cycles of 98°C, 60°C, and 72°C for 15, 20, and 30 seconds, respectively (F: SP6 II adapt, R: GFP-aug rev). For initial Kozak library generation see Random Kozak Library section. Each product was PCR purified using a DNA Clean & Concentrator-5 (Zymo #D4014).

[0259] Third, to generate a template for in vitro transcription, a PCR overlap extension was performed between the reporter library and the purified 3xFLAG-GFP-PEST amplicon using KAPA HiFi HotStart Ready Mix (Roche #7958935001) (0.1-1 ng of template DNA). After 10 cycles of 98°C, 60°C, and 72°C for 15, 20, and 30 seconds respectively, primers targeting the 5‘ SP6 promoter sequence and the 3‘ end of the 3xFLAG-GFP-PEST amplicons were added (F: SP6_II_adapt, pA-R: pCS2_3utr_60A, sv40-R: sv40_rev) followed by an additional 20 cycles at the same conditions. Unless stated explicitly, reporter libraries contained a 60A tail (pA-R: pCS2_3utr_60A). For libraries with an SV40 polyadenylation signal, a reverse primer targeting the 3’ end of the SV40 poly-adenylation signal was used for the amplification of 3xFLAG-GFP-PEST and the assembly of reporter library (R: SV40_rev).

[0260] Lastly, templates for in vitro transcription were gel purified using a Monarch® DNA Gel Extraction Kit (NEB #T1020L). Reporter mRNAs were generated using a mMESSAGE mMACHINE™ SP6 transcription kit (Invitrogen™ #AM1340). In vitro transcribed mRNAs were purified using a Monarch® RNA Cleanup Kit (NEB #T2040L). All libraries were injected into single cell zebrafish embryos at 20 pg per embryo.

[0261] Random Kozak library

[0262] The random Kozak library consisted of an 15 Illumina adaptor (5’- CCCTACACGACGCTCTTCCGATCT-3’, SEQ ID NO:2366) followed by the 5’ UTR of Xenopus beta-globin, seven random nucleotides, six upstream and one dow nstream of the start codon, and the N-terminus of 3xFLAG GFP. The Kozak library was generated by performing a PCR overlap extension using KAPA HiFi HotStart Ready Mix (Roche #7958935001) for 10 cycles of 98°C, 60°C, and 72°C for 15, 20, and 30 seconds respectively (F : kozak ntrap fwd, R: kozak_7_nt_rev, 1 pL of each primer at 100 uM).

[0263] Zebrafish 5 ’UTR library

[0264] A custom single stranded DNA oligo pool consisting of 11.088 oligos was ordered from GenScript (12 K oligo pool). Each 170 nt oligo contained an Illumina 15 (5:- CCCTACACGACGCTCTTCCGATCT-3’, SEQ ID NO:2366) adaptor sequence, a 124- nucleotide variable region, and 22 nt region with homology to the Kozak sequence and N- terminus of 3xFLAG-GFP (5’-GTAAACATGGTGAGCAAGGGCG-3‘, SEQ ID NO:2367). The variable region of the 5’ UTR library was generated using a custom script, tiling the 5’ UTRs of 1,775 maternally supplied genes and six IRES sequences (human AQP4, human MYT2. human NRF. human XIAP, EMCV and crTMV) in 124 nucleotide segments every 25 nucleotides.

[0265] Validation library (T etramer repeats)

[0266] A custom single stranded DNA oligo pool consisting of 256 oligos was ordered from Twist Bioscience as part of a 12 K ohgo pool. The design of the common regions of library were identical to that of the zebrafish 5’ UTR library. The variable region (124 nucleotides) consisted of repeats of all possible tetramers. Each repeat occurred 21 times and was separated by a dinucleotide spacer. The dinucleotide spacers were repeated in a pattern across the variable region (TC, AC. AG, CG). These dinucleotide spacers were selected to prevent the creation of unintended upstream ORFs.

[0267] NaP-TRAP spike-ins

[0268] The design of the spike-in reporters was identical to that of the zebrafish 5' UTR and tetramer validation libraries. Each spike-in reporter contained a 20 nucleotide identifier (see below). NaP-TRAP spike-ins were generated by performing a PCR overlap extension using KAPA HiFi HotStart Ready Mix (Roche #7958935001) (F: ntrap_spl_fwd, ntrap_sp2_fwd, ntrap_sp3_fwd, ntrap_sp4_fwd, ntrap_sp5_fwd: R: ntrap_sp_rev) (1 pL of each primer at 100 pM) for 10 cycles of 98°C, 60°C, and 72°C for 15, 20, and 30 seconds, respectively. Amplicons were gel purified using a Monarch® DNA Gel Extraction Kit (NEB #T1020L). Next, the purified amplicons were cloned into the 3xFLAG-GFP-PEST vector using InFusion® HD Cloning (Takara #638946) (vector F: ntrap_spike_inf_fwd, R: ntrap spike inf rev). To generate spike-in mRNAs, plasmids were amplified using KAPA HiFi HotStart ReadyMix (Roche #7958935001) (F: Sp6-Il-adapt, R: pCS2_3utr_60A) for 20 cycles at 98°C, 60°C, and 72°C for 15, 20, and 30 seconds, respectively. Products were then gel purified with a Monarch® DNA Gel Extraction Kit (NEB #T1020L) and then in vitro transcribed with a mMESSAGE mMACHINE™ SP6 transcription kit (Invitrogen™ #AM1340). mRNAs were purified with a Monarch® RNA Cleanup Kit (NEB #T2040L) prior to use. Spike-ins were pooled and added at the RNA extraction step at concentrations of 1, 5, 25, 50, and 125 fg for the 5’ UTR library'.

[0269] Spike-in #1 TGACGTGGAAGTCGGTCAAG (SEQ ID NO:2368) Spike-in #2 GTCCAGAGACAAAGTCCGGG (SEQ ID NO:2369) Spike-in #3 CACGAGGAGGAACCAGTGAC (SEQ ID NO:2370) Spike-in #4 CTGTTGTTGTGTGAAGGGCG (SEQ ID NO:2371)

[0270] Spike-in #5 GCTCTCGGTCTCGGAAGAAG(SEQ ID NO:2372)

[0271] NaP-TRAP (Nascent Peptide Translating Ribosome Affinity Purification)

[0272] To capture tagged nascent chains of reporter mRNAs, an immunoprecipitation using anti-FLAG magnetic beads was performed. Magnetic beads were purchased from three suppliers during the course of the study due to supply chain issues (ANTI-FLAG® M2 Magnetic Beads, Millipore® #M8823; Anti-Flag Magnetic Beads, BioTools LLC B26102; Pierce™ Anti-DYKDDDDK (SEQ ID NO:2373) Magnetic Agarose, Thermo Fisher Scientific #88836). For the zebrafish and human cell experiments 10 pL and 20 pL of beads (binding capacity’ of >0.8 mg of FLAG peptide / mL) were utilized, respectively. The magnetic beads were washed with 800 pL of wash buffer three times prior to being added to the lysis solution.

[0273] Briefly, frozen embry os and HEK293T cells were lysed in 500 pL of lysis buffer. After 10 minutes at 4°C the lysate was passed through a 25-guage needle (5-10 times). Next, the sample was centrifuged at 16,000 g for 5 minutes at 4°C. The supernatant was transferred to a new Eppendorf tube and 2 pL of DNase I (NEB #M0303L) was added. After a 15 minute incubation at 4°C, the samples were diluted to 1 mL using additional lysis buffer and 75 pL of lysate was collected from each sample to serve as an input. Samples were placed on a rotator at 4°C for 2 hrs. Following incubation and bead capture, the beads were washed with 800 pL of wash buffer three times. The beads (pulldown) and the inputs were then resuspended in 1 mL of Trizol (Invitrogen #15596-018). For the zebrafish 5’ UTR and tetramer validation libraries, spike in reporters were added. RNA extractions were performed in accordance with the manufacture’s protocol. RNA pellets were resuspended in 11 pL of nuclease free H2O.

[0274] Library preparation for Next-generation sequencing

[0275] Reverse transcription primers (4 pM total) targeting the N-terminus of 3xFLAG GFP were added to purified input and pulldown RNAs (5’- GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT(SEQ ID NO:2374)-CTAC-10N UMI-TAAC-6 nt sample barcode-GCGTCATGGTCTTTGTAGTCCTC(SEQ ID NO:2375)- 3’). These primers contained a 6-nucleotide sample barcode and a 10-nucleotide Unique Molecular Identifier (UMI) to allow for demultiplexing and read deduplication, as well as a 3’ 17 Illumina adaptor sequence. To increase library complexity the primer pairs were staggered by 1 nt (e.g., ntrap_RT_bl. l, ntrap_RT_bl.2). Reverse transcription was performed using the Superscript III kit (Invitrogen #18080044) in accordance with manufacturer’s instructions. Reverse transcription reactions were performed at 55°C. cDNA from replicates were pooled and purified by adding AMPure XP Reagent (Beckman Coulter #A63881) at 1.8x the original sample volume. Illumina 15 and 17 forward and reverse primers containing a 10-nucleotide index (see below) were utilized to amplify cDNA libraries via PCR (Kappa Polymerase Master Mix). To reduce the number of PCR duplicates 12-18 cycles were utilized. Amplicons were purified by adding AMPure XP Reagent (Beckman Coulter #A63881) at 0.9x the volume of the PCR reaction. Libraries were sequenced on Illumina NovaSeq 6000 platform.

[0276] Illumina 15 primer:

[0277] 5’- AATGATACGGCGACCACCGAGATCTACAC (SEQ ID NO:2376)-10 nt index- ACACTCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID NO:2377)-3’

[0278] 812 Illumina 17 primer:

[0279] 5’- CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO:2378)-10 nt index- GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO:2379)-3’

[0280] NaP-TRAP-qPCR analysis

[0281] Following NaP-TRAP, cDNA was synthesized using random hexamers (ThermoFisher Scientific #N8080127) and SuperScript™ III Reverse Transcriptase (Invitrogen #18080044) following the manufacturer’s protocol. Translation values were determined by qPCR using the Power SYBR™ Green PCR Master Mix (Applied Biosystems #4367659). Levels of reporters in the input and pulldown were normalized using the 2-AACt method, 59 where the NaP-TRAP reporter was the target and dsRED the control. Translation was calculated as a ratio of fold enrichment in the pulldown relative to the input.

[0282] Primers utilized in qPCR:

[0283] - Translation blocking morpholino (GFP qpcr fwd, GFP qpcr rev)

[0284] - MicroRNA seeds in 3’ UTR (3xmir-430_qpcr_fwd, 3xmir-204_qpcr_fwd, T7_qpcr_rev)

[0285] - U / C rich reporters (igfl ra wt qPCR fwd, igflra mut qPCR fwd, asun wt qPCR fwd, asun mut qPCR fwd, sich21 I wt qPCR fwd, sich211 mut qPCR fwd, atp6v0 wt qPCR fwd. atp6v0 mut qPCR fwd, GFP_qPCR_5p_rev) - dsRED (dsRED qpcr fwd, dsRED qpcr rev)

[0286] NaP-TRAP NanoLuc reporter construction

[0287] The present study substituted NanoLuc luciferase (amplified from a gBlocks™ Gene Fragment, Integrated DNA Technologies IDT) for GFP in the 3xFLAG-GFP-PEST vector using In-Fusion® HD Cloning (Takara #638946) (3xFLAG-PEST vector F: PEST fwd, R: FLAG rev inf; NanoLuc insert F: FLAG-Nluc_inf_fwd, R: Nluc_inf_rev). Note: The PEST domain was included in NanoLuc constructs to reduce the half-life of the NanoLuc protein and thereby improve the sensitivity of the assay. To generate NanoLuc reporters 3xFLAG- NanoLuc-PEST was amplified using KAPA HiFi HotStart Ready Mix (Roche #7958935001) (F: 3xFLAG GFP fwd R: per II pa rev) and gel purified using a Monarch® DNA Gel Extraction Kit (NEB #T1020L).

[0288] Kozak NanoLuc reporters

[0289] To construct Kozak NanoLuc reporters, 3xFLAG-NanoLuc (amplicon from previous section) was amplified with four different forward primers (ntrap_kl_fwd. ntrap_k2_fwd, ntrap_k3_fwd, and ntrap_k4_fwd) and a common reverse primer (pcr_II_pa_rev) using KAPA HiFi HotStart ReadyMix (Roche #7958935001) at 20 cycles of 98°C, 60°C, and 72°C for 15, 20, and 45 seconds, respectively. Products were gel purified using a Monarch® DNA Gel Extraction Kit (NEB #T1020L). To add an SP6 promoter sequence and a 60A hard- encoded tail the purified product (1 ng) was amplified using KAPA HiFi HotStart ReadyMix (Roche #7958935001) at 20 cycles of 98°C, 60°C, and 72°C for 15, 20, and 45 seconds, respectively (F: Sp6-Il-adapt, R: pCS2_3utr_60A). Products were gel purified with a Monarch® DNA Gel Extraction Kit (NEB #T1020L) and then in vitro transcribed with a mMESSAGE mMACHINE™ SP6 transcription kit (Invitrogen™ # AMI 340). mRNAs were purified with a Monarch® RNA Cleanup Kit (NEB #T2040L) prior to injection. Note, the input of one of the replicates of the Kozak reporter K4 w as below the read filter cutoff in the NaP-TRAP Kozak Library experiment.

[0290] Generating NanoLuc reporters using PCR overlap

[0291] The 5’ UTRs of NanoLuc Reporters were generated using PCR overlap (see below7for forward and reverse primers), using KAPA HiFi HotStart ReadyMix (Roche #7958935001) at 10 cycles of 98°C. 60°C, and 72°C for 15, 20, and 30 seconds, respectively (1 pL of each primer at 100 pM). The extension products were then gel purified using a Monarch® DNA Gel Extraction Kit (NEB #T1020L).

[0292] Full-length reporters were constructed by performing a PCR overlap extension of the 5’ UTR products with the NanoLuc amplicon described in the Cloning NanoLuc section using KAPA HiFi HotStart Ready Mix (Roche #7958935001). After 10 cycles of 98°C, 60°C, and 72°C for 15, 20, and 40 seconds, respectively, primers targeting the SP6 promoter sequence (F: Sp6-Il-adapt), and the 3’ end of the nanoLuc amplicon were added (R: pCS2_3utr_60A). The PCR reaction was continued for an additional 20 cycles at the same cycle conditions. Products were gel purified with a Monarch® DNA Gel Extraction Kit (NEB #T1020L) and then in vitro transcribed with ta mMESSAGE mMACHINE™ SP6 transcription kit (Invitrogen™ #AM1340). mRNAs were purified with a Monarch® RNA Cleanup Kit (NEB #T2040L) prior to injection.

[0293] Zebrafish 5’ UTR library validation reporters (slcl6al_fwd, slcl6al_rev, spredl_fwd, spredl_rev, eif4ebp2_fwd, eif4ebp2_rev, fthla fwd, fthla_rev, wdr41_fwd, wdr41_rev, slc38a7_fwd, slc38a7_rev, atp2b4_fwd. atp2b4_rev, atp6v0_wt_fwd, atp6v0_wt_rev, Igflra_wt_fwd, Igflra_wt_rev).

[0294] 4xmiR-430 NanoLuc- (WT-F: 4xmiR-430-fwd, WT-R: 4xmiR-430-rev; MUT-F: 4xmiR-430-MUT-fwd, MUT-R: 4xmiR-430-MUT-rev).

[0295] Dual luciferase assay

[0296] Firefly luciferase was amplified from a pCS2+-Fluc plasmid using KAPA HiFi HotStart Ready Mix (Roche #7958935001) at 20 cycles of 98°C, 60°C, and 72°C for 15, 20, and 50 seconds (F: SP6_ext_fwd, R: SV40_60A_rev). Products were gel purified with a Monarch® DNA Gel Extraction Kit (NEB #T1020L) and then in vitro transcribed with a mMESSAGE mMACHINE™ SP6 transcription kit (Invitrogen™ # AMI 340). mRNAs were purified with a Monarch® RNA Cleanup Kit (NEB #T2040L) prior to injection.

[0297] Single cell zebrafish embryos were co-injected with mRNAs encoding nano and firefly luciferase (0.5 pg of nano / 19.5 pg of firefly). Embryos were collected at either 2 hpf or 6 hpf and frozen in liquid nitrogen (5 embryos per replicate). Nano and firefly luciferase activities were measured using the Nano-Gio® Dual -Luciferase® Reporter Assay System (Promega #N1610). Firefly luciferase activity was utilized to normalize nano luciferase measurement across reporters.

[0298] C and U rich reporters C- and U-rich wild-type and mutant reporters were generated through PCR overlap extension (Primers used: igflra wt fwd, igflra wt rev; igflra mut fwd, igflra mut rev; asun_wt_fwd, asun_wt_rev; asun mut fwd. asun_mut_rev; sich211_wt_fwd, sich21 I t rev: sich211_mut_fwd, sich21 I rnut R atp6v0_wt_fwd, atp6v0_wt_rev; atp6vO_mut_F atp6vO_mut_R) using KAPA HiFi HotStart ReadyMix (Roche #7958935001) at 10 cycles of 98°C. 60°C, and 72°C for 15, 20, and 30 seconds, respectively (1 pL of each primer at 100 pM). The extension products were then gel purified using a using Monarch® DNA Gel Extraction Kit (NEB #T1020L).

[0299] Extension products were cloned into the 3xFLAG-GFP-3xAID* vector using InFusion® HD Cloning (Takara #638946) (vector F: 3xFLAG_GFP_F R: I5_inf_rev). Plasmids were digested with Notl-HF® (NEB #R3189L) and in vitro transcribed with a mMESSAGE mMACHINE™ SP6 transcription kit (Invitrogen™ # AMI 340). mRNAs were purified with a Monarch® RNA Cleanup Kit (NEB #T2040L) prior to injection.

[0300] Quanti fication and Statistical Analysis

[0301] Read trimming and translation calculation

[0302] Reporter libraries were sequenced on an Illumina NovaSeq 6000 platform. Sequencing data were stored using LabxDB.60 Paired-end reads were trimmed and demultiplexed using ReadKnead (github dot com / vejnar / ReadKnead). Barcodes identifying replicates and UMIs were extracted from read two. Common regions were trimmed from read one to facilitate accurate mapping. For the Kozak library, reporters were counted using a custom python script. Reporters with indels in the Kozak sequence or reporters without an AUG in the appropriate position were eliminated. For the zebrafish 5’ UTR and validation libraries, reads were mapped to a library specific index using Bowtie2.61 PCR duplicates were eliminated using UMIs. UMIs were considered identical if they had a Hamming Distance less than 2. Reads for each experiment were normalized by dividing the read counts of each reporter by the sum of the total number of reads mapped to the spike-ins. In the 5’ UTR reporter library (pA and sv40 in zebrafish) spike-in #3 was eliminated from the analysis, because in some of the samples the read counts for spike-in #3 did not correlate with amount of spike-in added. In the absence of spike-ins, read counts were normalized based on the total number of mapped reads per replicate (reads per million, RPM). Translation values were calculated as a ratio of reads in the pulldown relative to reads in the input. Translation values were only included in the downstream analyses if the input contained greater than or equal to 100 unique reads across all replicates. Random Forest Regression models

[0303] Random forest regression models were employed to predict translation (scikit-leam 1.3; RandomForestRegressor).62 For the Kozak library, features were generated by one-hot encoding positions -6 to -1 and position +4 of each reporter sequence. In contrast, for the 5’ UTR library, k-mer counts (1-6 nucleotides) and uORF features were generated using a custom python script. For the 5’ UTR library, features were filtered by calculating the Spearman Rank Correlation Coefficient (SPR) between each feature and translation prior to model training. Features that had a correlation greater than 0.05 or less than -0.05 were included in the random forest model.

[0304] To prevent model overfitting, the data were divided randomly into two groups: a test and training set, comprised of 30% and 70% of the reporters, respectively. To optimize model parameters, the training data were divided into five different groups of equal size and a 5-fold cross-validation was performed. The following parameters were optimized using an exhaustive grid search (n_estimators: 20, 100. 200; max_features: 10, 20 and 30 percent of supplied features; max_depth: 3. 5, 7 and min_samples_split: 2, 4, 8). Bootstrapping was employed to select samples used to train each tree. The predictive power of the model was assessed using the test set. A permuted feature importance analysis was performed to identify the features with the greatest predictive power.

[0305] Differential motif enrichment analysis

[0306] Reporters were ranked based on their translation at 2 and 6 hpf. Using the sum and difference of these rankings across timepoints, four groups of reporters were generated: (1) repressed, (2) active. (3) repressed post ZGA (active in HEK293T cells), and (4) active post ZGA (active in zebrafish). Repressed and active reporters constituted the top and bottom 10% of reporters based on the sum of their ranks at 2 and 6 hpf, respectively, whereas the repress post ZGA and active post-ZGA. were the top and bottom 10% of groups based on the difference between their ranks at 2 and 6 hpf. A differential motif enrichment analysis was performed on each group. Fold enrichment values were determined by dividing the count of each k-mer in the reporter group by the count of the k-mer in the library, whereas the significance of the fold-change was determined using a hypergeometric test (Bonferroni corrected p-value threshold). miRNA complementarity analysis miR-430 (GCACUU) and miR-1 (ACAUUC) seeds were identified in the reporter library. The Vienna RNAcofold program63 was utilized to measure the complementarity between the section of the reporter mRNA (20 nt upstream and 7 nt downstream of the 5’ end of seed site) and the microRNA. For miR-430 the miRNA species with the highest complementary' was selected for downstream analysis. miR-430a: 5 -UAAGUGCUAUUUGUUGGGGUAG-3’ (SEQ ID NO:2380) miR-430b: 5 -AAAGUGCUAUCAAGUUGGGGUAG-37(SEQ ID NO:2381) miR-430c: 5’-UAAGUGCUUCUCUUUGGGGUAG-3?(SEQ ID NO:2382) miR-1-1 / mir-1-2: 5 -UGGAAUGUAAAGAAGUAUGUAU-3’ (SEQ ID NO:2383)

[0307] Feature rank analysis

[0308] To generate the feature rank plot (Fig. 13 A), features (k-mer counts of four nucleotides or fewer) were ranked based on the mean difference in translation between reporters enriched and depleted in the feature, top and bottom 20% respectively, at 2 and 6 hpf Given the prominent effect of upstream open reading frames on translation, reporters with uORFs were excluded from the analysis. Features were only included in the analysis if there w as a significant mean difference in translation at either timepoint (Mann-Whitney U test with Bonferroni corrected p-value threshold).

[0309] Statistical analyses

[0310] The plots and statistical analyses in Figs. 8B, 8C, 9G, 12E and 13C-13F were generated using GraphPad Prism. All other analyses unless otherwise stated w ere performed using custom scripts written in Python 3. Plots were generated using the Matplotlib package. Venn diagrams were generated using Matplot-venn (github dot com / konstantint / matplotlib- venn). Statistical analyses were performed using the SciPy and NumPy packages, whereas the random forest analysis was performed using the scikit-leam package. Feature and experimental data were stored using an SQLite database (https: / / ww w'.sqlite.org / ).

[0311] Detailed Protocol for NaP-TRAP

[0312] Buffers lOx salt buffer - 15 mM Tris 7.4 (ThermoFisher Scientific #612021000, 100 mM NaCl (Sigma- Aldrich # S6546-1L), 10 mM MgCh (Invitrogen™ #AM9530G).

[0313] Bead wash buffer - lx salt buffer, 2 mM DTT (ThermoFisher Scientific #P2325). Lysis buffer - lx salt buffer, 10% triton-X (Sigma- Aldrich #X100-500ML), 2 mM DTT (ThermoFisher Scientific #P2325), 100 pg / mL Cycloheximide (Sigma- Aldrich #O181O-1G), 40 U RNaseOUT™ Recombinant Ribonuclease Inhibitor (ThermoFisher Scientific #10777019), lx cOmplete™ EDTA-free Protease Inhibitor (Roche #11873580001).

[0314] NaP-TRAP wash buffer - Lysis buffer + 400 mM NaCl (Sigma-Aldrich # S6546-1L).

[0315] HEK293T media - Dulbecco's Modified Eagle Medium (DMEM) (ThermoFisher Scientific #10569010), 10% heat inactivated Fetal Bovine Serum (FBS) (ThermoFisher Scientific # 16140071), 20 mM HEPES (ThermoFisher Scientific #15630080), 2 mM L- Glutamine (ThermoFisher Scientific #25030081), lx Penicillin / Streptomycin (ThermoFisher Scientific #15140122).

[0316] DPBS - cycloheximide - Dulbecco's phosphate-buffered saline (DPBS) (ThermoFisher Scientific #14190144), 100 pg / mL cycloheximide (Sigma-Aldrich #01810- 1G).

[0317] Zebrafish injections

[0318] 1. After zebrafish mating, collect embryos and remove their chorion through treatment with Pronase (Sigma-Aldrich #10165921001).

[0319] 2. Inject dechori onated embryos with 20 pg of an in vitro transcribed mRNA reporter library.

[0320] 3. Incubate embryos at 28°C.

[0321] 4. Collect 25-50 embry os per replicate at 2 and 6 hpf Transfer the embryos to a 1.5 mL Eppendorf tube and freeze them in liquid nitrogen (store at -80°C).

[0322] 5. Prior to NaP-TRAP, add 500 mL of lysis buffer to each set of frozen embryos. Vortex the samples to resuspend the frozen pellet (proceed to NaP-TRAP protocol).

[0323] HEK293T RNA transfection

[0324] 1. Coat plates with poly-d-lysine (1 mL per well) (ThermoFisher Scientific # A3890401). Incubate for 1 hour at room temperature. Wash plates with sterile water three times. Allow plates to dry.

[0325] 2. Seed a 6-well plate at 200,000 to 300,000 cells per well (1 mL HEK293T media). Incubate the cells for 24-36 hours prior to transfection.

[0326] 3. For each well mix 3.75 pL of Lipofectamine™ MessengerMAX™ Transfection Reagent (ThermoFisher Scientific # LMRNA008) with 121.25 pL of Opti-MEM™ I Reduced Serum Medium (ThermoFisher Scientific # 31985062). Incubate the solution for 10 minutes at room temperature.

[0327] 4. Dilute 1-2 pg of mRNA in 125 pL of Opti-MEM™ I Reduced Serum Medium (ThermoFisher Scientific # 31985062) and add to the Lipofectamine™ mixture. Gently mix and incubate for another 5 minutes at room temperature.

[0328] 5. Add 250 pL of mRNA-Lipofectamine™ solution to each well drop-wise distributing the solution throughout the plate.

[0329] 6. Incubate for 1 hour at 37°C with 5% CO2.

[0330] 7. Aspirate cells and wash with prewarmed DPBS (ThermoFisher Scientific #14190144). Add new media (1 mL prewarmed)

[0331] 8. Grow the cells for 12 hours at 37°C with 5% CO2.

[0332] 9. Add 100 pg / mL cycloheximide (Sigma- Aldrich #01810-lG) to cells and incubate for 10 minutes at 37°C with 5% CO2.

[0333] 10. Place cells on ice and aspirate media and wash cells with ice cold DPBS + cycloheximide.

[0334] 11. Add 500 pL of lysis buffer to the cells.

[0335] 12. Disrupt cells mechanically with a plastic cell scraper and transfer the lysate to a new tube.

[0336] 13. Incubate for 10 minutes on ice.

[0337] Nascent Peptide Ribosome Affinity) Purification

[0338] 1. Using a magnetic stand (DynaMag™-2 Magnet, Invitrogen™ #1232 ID), wash the anti -FLAG magnetic beads, 10 pL per zebrafish sample and 20 pL per HEK293T sample (binding capacity’ of >0.8 mg of FLAG peptide / mL) with 800 pL of the bead wash buffer three times. Until the samples are ready, leave the beads in the final wash on ice. When working with the anti-FLAG beads it is important to utilize pre-lubricated Eppendorf tubes.

[0339] Note: Magnetic beads were purchased from three suppliers during the course of the study due to supply chain issues (ANTI-FLAG® M2 Magnetic Beads, Millipore® #M8823; Anti-Flag magnetic beads. BioTools LLC #B26102; Pierce™ Anti-DYKDDDDK Magnetic Agarose, ThermoFisher Scientific #88836). Pierce™ Anti-DYKDDDDK Magnetic Agarose is recommend for this experiment.

[0340] 2. Pass lysed cells through a 26 G needle 5-10 times.

[0341] 3. Spin lysed cells at 16,000 g for 5 minutes at 4°C. 4. Transfer the supernatant to a new tube and add 2 pL of DNASE 1 (0.006 U / mL) (NEB #M0303L).

[0342] 5. Incubate samples on ice for 15 minutes.

[0343] 6. Add 500 pL of lysis buffer to bring the total volume of the solution to 1 mL.

[0344] 7. Collect 75 pL of lysate as input (set aside on ice).

[0345] 8. Add supernatant to 10 pL of washed beads from step 1.

[0346] 9. Rotate the beads on a rotor for 2 hours at 4°C.

[0347] 10. Using a magnetic stand (DynaMag™-2 Magnet, Invitrogen™ #1232 ID), remove the supernatant from the beads and wash the beads with 800 pL of the NaP-TRAP wash buffer three times. To limit experimental noise, when adding the NaP-TRAP wash buffer be sure to resuspend the magnetic beads.

[0348] 11. Add 1 mL of Trizol (Invitrogen #15596-018) to the washed beads and the input. (Samples can be stored here at -20°C).

[0349] RNA extraction

[0350] 12. Add in vitro transcribed NaP-TRAP spike-ins to Trizol (Invitrogen #15596-018). The present study utilized 1. 5, 25. 50 and 125 fg for spike-ins #1-5, respectively. This total will need to be adjusted for the experimental system.

[0351] 13. Add 200 pL of chloroform to samples. Vortex and incubate samples for 2-3 minutes at RT.

[0352] 14. Centrifuge samples 12.000 g for 15 minutes at 4°C.

[0353] 15. Transfer the aqueous phase to a fresh 1 .5 ml Eppendorf tube and add 1 .5 pL of GlycoBlue (Invitrogen #AM9516) and 500 pL of Isopropanol. Vortex and incubate samples at RT for 10 minutes.

[0354] 16. Centrifuge samples at 12,000 g for 15 minutes at 4°C.

[0355] 17. Remove supernatant. Wash RNA pellet twice with 1 mL of 75% ethanol. Centrifuge samples 7500 g for 10 minutes at 4°C.

[0356] 18. Allow pellet to dry (5-10 minutes). Resuspend pellet in 11 pL of nuclease free H2O.

[0357] Library Preparation

[0358] 19. Transfer resuspended RNA to PCR tubes. Add 1 pL of 10 mM dNTPs (NEB #N0446S) and 1 pL of 4 pM of barcoded reverse transcription primers to the samples.

[0359] 20. For primer annealing, heat samples to 95°C for 2 minutes and then 65°C for 5 minutes. Place samples on ice for at least 1 minute. 21. Add 7 pL of SuperScript™ III Reverse Transcriptase master mix to each sample (per sample: 4 pL 5xFirst Strand Synthesis Buffer, 1 pL 0. 1 M DTT, 1 pL RNaseOUT, and 1 pL of SuperScript™ III)

[0360] 22. For the reverse transcription reaction, heat samples for 1 hour at 55°C and then 15 minutes at 70°C.

[0361] 23. Add 2 pL of RNase H (NEB # M0297L) to each sample. Heat samples for 15 minutes at 37°C.

[0362] 24. Combine RT barcoded replicates (Note: keep the input and pulldown samples separate).

[0363] 25. Add 1.8x AMPure XP Reagent (Beckman Coulter #A63881) to each sample. Mix the beads by pipetting the samples.

[0364] 26. Incubate the mixture at RT for 15 minutes.

[0365] 27. Using a magnetic stand (DynaMag™-2 Magnet, Invitrogen™ # 12321 D) remove the supernatant and wash the beads twice with 75% ethanol.

[0366] 28. After washing the beads, allow them to dry at RT (~10 minutes).

[0367] 29. Resuspend the beads in 30 pL of nuclease free H2O and incubate for 5 minutes at RT. Using a magnetic stand (DynaMag™-2 Magnet, Invitrogen™ # 12321 D) immobilize the beads and transfer the purified cDNA to a new' tube.

[0368] 30. Amplify library for Next-Generation sequencing. Mix 12.5 pL of 2x KAPA HiFi HotStart ReadyMix (Roche #7958935001), 1 pL of an 15 forward, 1 pL 17 reverse primer. 1- 10.5 pL of cDNA, 0-9.5 pL of nuclease-free H2O. PCR specifications: 95°C 3 minutes, 10-20 cycles of 98°C, 60°C, and 72°C for 15, 20, and 30 seconds, respectively, 72°C 5 minutes (F: 15 10nt_[id]_fwd; R: I7_10nt_[id]_rev). Cycle conditions were optimized by running amplicons of a 20 cycle test reaction on an agarose gel.

[0369] 31. Purify library with AMPure XP Reagent (Beckman Coulter #A63881) at a concentration of 0.9x the sample volume.

[0370] 32. Note: Gel extract for precise library selection and noise reduction. Check sample purify and concentration with a Bioanalyzer prior to sequencing.

[0371] Methods for 5 ’utr 3 ’utr hexamer library half-life calculation

[0372] 5UTR and 3UTR hexamer decay library transfected into HEK293T cells and collected at 6h, 18h, 24h, 36h, 48h, 72h. RNA w as extracted from the collected samples, five spike-ins designed from random 25mer were added, and reverse transcribed using Superscript III kit (Invitrogen). cDNA was size purified using AmPure XP beads (Beckman Coulter) and purified again after amplifying the inserted codon repeats directly using Illumina 5’ and 3’ adaptors fused with reverse transcription barcode region. Libraries were sequenced on Illumina NovaSeq 6000 with paired-end 150bp.

[0373] Sequencing data were stored using LabxDB (Vejnar, C.E. and A.J. Giraldez, Bioinformatics, 2020. 36(16): p. 4530-4531). Barcodes were extracted from the read two. Trimmed reads were mapped using Bowtie2 (Langmead, B. and Salzberg, S.L., Nature Methods. 2012. 9(4):357-9). Counts of constructs were normalized using five spike-ins. Normalized counts of 6h, 18h, 24h, 36h, 48h, 72h were plotted, which follows an exponential decay graph, where the present study can calculate half-lives from graphs.

[0374] Example 3: Translational efficiencies as determined by NaP-TRAP matches those determined by polysome profiling

[0375] Referring to Fig. 22, the present study benchmarked NaP-TRAP against polysome profiling, a well-established translation based MPRA method. To minimize the effect of ribosomes translation out-of-frame, the present study incorporated stop codons early in the coding sequence of 3xFLAG-GFP.

[0376] The present study observed a strong correlation between NaP-TRAP derived translation values and Mean Ribosome load, demonstrating that NaP-TRAP captures ribosomes in a manner proportional to the number of ribosomes translating the main coding sequence.

[0377] It should be noted that in the absence of these out-of-frame stop codons, these translation metrics are likely to diverge. The reason is that the polysome profiling method cannot distinguish translation products produced by frameshifts. NaP-TRAP results, due to the specificity to the sequences of the translation products, filter the translations that involve frameshift (See e.g.. Example 4).

[0378] Example 4: NaP-TRAP is able to capture translation in multiple frames simultaneously

[0379] Referring to Figs. 23A-23D, the present study demonstrated the capacity of NaP- TRAP to measure translation in multi-frames simultaneously.

[0380] Specifically, the present study incorporated Ha-tags in frame +1 and MYC-tags frame +2, in addition to the FLAG tag in frame 0. Translations in all the three frames were detected by detecting HA. MYC and FLAG tags, respectively. This experiment demonstrated the capacity of NaP-TRAP to measure translation in multiple frames simultaneously. Using a library with upstream start codons in each frame, the present study further demonstrated the frame specificity of the method. The present study only observed translation in frames +1 and 2+ when an overlapping open reading frame is present.

[0381] 5 Example 5: Identification of a 5’UTR with high translation potential

[0382] The present study performed NaP-TRAP to test the translation efficiencies of several 5‘ UTRs in HEK293T cells. The present study further selected a 5’UTR that showed high translation potential. The present study further engineered this 5’UTR by introducing base changes based on the understanding of how sequences variations in 5’UTR affect translation.

[0383] 10 The present study tested the translation efficiency of this engineered 5’UTR (RESA mRNA- VI) against the 5’UTRs used in commercially available CoVID vaccines (BNT 162b2 and mRNA-1273). The present study used a luciferase reporter to measure the protein output from these 5’UTRs and observed a 10-fold increased protein output from RESA mRNA-Vl as compared to BNT 162b2 and mRNA-1273 at the time-point tested (Fig. 24).

[0384] 15

[0385] Example 6: Identification of cis-regulatory elements in human and zebrafish cells Using the methods developed herein, the present study further identified cis- regulatory elements that modulates translation in human cells and zebrafish cells.

[0386] 20 Example 6-1: Cis-regulatory elements identified in human cells

[0387] In HEK293T cells, the following sequences were identified:

[0388]

[0389]

[0390]

[0391] Example 6-2: Cis-regulatory elements identified in zebrafish cells

[0392] In zebrafish cells, the following cis-regulatory elements were identified:

[0393]

[0394]

[0395]

[0396]

[0397] Enumerated Embodiments

[0398] Embodiment 1 : An mRNA molecule comprising a cis-regulatory element selected from the group consisting of SEQ ID NOs: 1-2353 and 2384, wherein the cis-regulatory element does not naturally exist in the mRNA molecule at a location of the cis-regulatory' element.

[0399] Embodiment 2: The mRNA molecule of Embodiment 1, wherein the cis-regulatory element is a cis-regulatory element for the translational machinery of a mammal, optionally a human.

[0400] Embodiment 3: The mRNA molecule of Embodiment 2, wherein the cis-regulatory element comprises at least one sequence selected from the group consisting of SEQ ID NOs: 1-1156 and 2384.

[0401] Embodiment 4: The mRNA molecule of Embodiment 1, wherein the cis-regulator ' element is a cis-regulatory element for the translational machinery of a fish, optionally a zebrafish.

[0402] Embodiment 5: The mRNA molecule of Embodiment 4, wherein the cis-regulatory7element comprises at least one sequence selected from the group consisting of SEQ ID NOs: 1157-2353.

[0403] Embodiment 6: The mRNA of any one of Embodiments 1-5, further comprises a modified nucleobase that does not naturally exist at the location of the modified nucleobase.

[0404] Embodiment 7: The mRNA molecule of Embodiment 6, wherein the modified nucleobase comprises at least one of N6-methyladenosine, inosine, N1 -propylpseudouridine, Nl-methoxymethylpseudouridine, N1 -ethylpseudouridine, 5 -methoxy cytidine, 5- hydroxyuridine, 5-carboxyuridine. 5-formyluridine. 5-hydroxy cytidine. 5- hyrdoxymethylcytidine, 5-hydroxymethyluridine, 5-formylcytidine, 5-carboxy cytidine, N4- methylcytidine, pseudoisocytidine, 2-thiocytidine, 4-thiocytidine, N1 -methylpseudouridine, pseudouridine, 5-methyluridine, 5 -methoxy uridine, dihydrouridine. 5-methylcytidine, 4- thiouridine, 2-thiouridine, uridine-5’-O-(l -thiophosphate), 5-aminoallyluridine, and 4- acetylcytidine.

[0405] Embodiment 8: The mRNA molecule of any one of Embodiments 1-7, wherein the mRNA molecule encodes a therapeutic peptide or a therapeutic protein.

[0406] Embodiment 9: The mRNA molecule of Embodiment 8, wherein the therapeutic peptide or the therapeutic protein comprises a vaccine. Embodiment 10: A method of modulating a translational efficiency of a messenger RNA (mRNA) molecule in a cell, the method comprising: modifying the mRNA molecule to comprise a cis-regulatory element selected from the group consisting of SEQ ID NOs: 1-2353 and 2384.

[0407] Embodiment 11 : The method of Embodiment 10, wherein modifying the mRNA molecule comprises modifying a sequence of a DNA molecule encoding the mRNA molecule.

[0408] Embodiment 12: The method of any one of Embodiments 10-11, wherein the cis- regulatory element is a cis-regulatory element for the translational machinery' of a mammal, optionally a human.

[0409] Embodiment 13: The method of Embodiment 12, wherein the cis-regulatory element comprises at least one sequence selected from the group consisting of SEQ ID NOs: 1-1156 and 2384.

[0410] Embodiment 14: The method of any one of Embodiment 10-11, wherein the cis- regulatory element is a cis-regulatory element for the translational machinery of a fish, optionally a zebrafish.

[0411] Embodiment 15: The method of Embodiment 14, wherein the cis-regulatory element comprises at least one sequence selected from the group consisting of SEQ ID NOs: 1157- 2353.

[0412] Embodiment 16: The method of any one of Embodiments 10-15, wherein the mRNA molecule is further modified to comprise a modified nucleobase.

[0413] Embodiment 17 : The method of Embodiment 16, wherein the modified nucleobase comprises at least one of N6-methyladenosine, inosine, N1 -propylpseudouridine. Nl- methoxymethylpseudouridine. N1 -ethylpseudouridine, 5-methoxy cytidine, 5-hydroxyuridine, 5-carboxyuridine, 5-formyluridine, 5-hydroxycytidine, 5-hyrdoxymethylcytidine, 5- hydroxymethyluridine, 5-formylcytidine, 5-carboxy cytidine, N4-methylcytidine, pseudoisocytidine, 2-thiocytidine, 4-thiocytidine, N1 -methylpseudouridine, pseudouridine, 5- methyluridine, 5 -methoxy uridine, dihydrouridine. 5-methylcytidine, 4-thiouridine, 2- thiouridine, uridine-5’-O-(l -thiophosphate). 5-aminoallyluridine. or 4-acetylcytidine.

[0414] Embodiment 18: The method of any one of Embodiments 10-17, wherein the cis- regulatory element activates the translation of the mRNA molecule, and wherein a translational efficiency of the mRNA molecule is increased by about 50% or more in comparison to an mRNA molecule lacking the cis-regulatory element but otherwise has the same sequence. Embodiment 19: The method of any one of Embodiments 10-17, wherein the cis- regulatory element represses the translation of the mRNA molecule, and wherein a translational efficiency of the mRNA molecule is decreased by about 30% or more in comparison to an mRNA molecule lacking the cis-regulatory element but otherwise has the same sequence.

[0415] Embodiment 20: The method of any one of Embodiments 10-19, wherein the cis- regulatory element activates or represses the translation of the mRNA molecule in the cell at a certain cellular stage.

[0416] Embodiment 21: The method of any one of Embodiments 10-20, wherein the cis- regulatory element is located in a 5’ untranslated region (5’-UTR), at a region including the start codon of the main open reading frame, within the open reading frame, or in a 3’ untranslated region (3’-UTR).

[0417] The foregoing outlines features of several embodiments so that those skilled in the art may better understand the aspects of the present disclosure. Those skilled in the art should appreciate that they may readily use the present disclosure as a basis for designing or modifying other processes and structures for earn ing out the same purposes and / or achieving the same advantages of the embodiments introduced herein. Those skilled in the art should also realize that such equivalent constructions do not depart from the spirit and scope of the present disclosure, and that they may make various changes, substitutions, and alterations herein without departing from the spirit and scope of the present disclosure.

Claims

CLAIMSWhat is claimed is:

1. An mRNA molecule comprising a cis-regulatory element selected from the group consisting of SEQ ID NOs: 1 -2353 and 2384, wherein the cis-regulatory element does not naturally exist in the mRNA molecule at a location of the cis-regulatory element.

2. The mRNA molecule of claim 1, wherein the cis-regulatory element is a cis- regulatory element for the translational machinery7of a mammal, optionally a human.

3. The mRNA molecule of claim 2, wherein the cis-regulatory element comprises at least one sequence selected from the group consisting of SEQ ID NOs: 1-1156 and 2384.

4. The mRNA molecule of claim 1, wherein the cis-regulatory element is a cis- regulatory element for the translational machinery' of a fish, optionally a zebrafish.

5. The mRNA molecule of claim 4, wherein the cis-regulatory element comprises at least one sequence selected from the group consisting of SEQ ID NOs: 1157-2353.

6. The mRNA of any one of claims 1-5, further comprises a modified nucleobase that does not naturally exist at the location of the modified nucleobase.

7. The mRNA molecule of claim 6, wherein the modified nucleobase comprises at least one of N6-methyladenosine, inosine, N1 -propylpseudouridine, Nl- methoxymethylpseudouridine, N1 -ethylpseudouridine, 5-methoxycytidine, 5-hydroxyuridine, 5-carboxyuridine, 5-formyluridine, 5-hydroxycytidine, 5 -hyrdoxy methylcytidine, 5- hydroxymethyluridine, 5-formylcytidine, 5-carboxycytidine, N4-methylcytidine, pseudoisocytidine, 2-thiocytidine, 4-thiocytidine, N1 -methylpseudouridine, pseudouridine, 5- methyluridine, 5-methoxyuridine, dihydrouridine, 5-methylcytidine, 4-thiouridine, 2- thiouridine, uridine-5’-O-(l -thiophosphate). 5-aminoallyluridine, and 4-acetylcytidine.

8. The mRNA molecule of any one of claims 1-7, wherein the mRNA molecule encodes a therapeutic peptide or a therapeutic protein.

9. The mRNA molecule of claim 8, wherein the therapeutic peptide or the therapeutic protein comprises a vaccine.

10. A method of modulating a translational efficiency of a messenger RNA (mRNA) molecule in a cell, the method comprising: modifying the mRNA molecule to comprise a cis-regulatory element selected from the group consisting of SEQ ID NOs: 1-2353 and 2384.

11. The method of claim 10, wherein modifying the mRNA molecule comprises modifying a sequence of a DNA molecule encoding the mRNA molecule.

12. The method of any one of claims 10-11, wherein the cis-regulatory element is a cis- regulatory7element for the translational machinery of a mammal, optionally a human.

13. The method of claim 12. wherein the cis-regulatory element comprises at least one sequence selected from the group consisting of SEQ ID NOs: 1-1156 and 2384.

14. The method of any one of claims 10-11, wherein the cis-regulatory element is a cis- regulatory element for the translational machinery of a fish, optionally a zebrafish.

15. The method of claim 14, wherein the cis-regulatory element comprises at least one sequence selected from the group consisting of SEQ ID NOs: 1157-2353.

16. The method of any one of claims 10-15, wherein the mRNA molecule is further modified to comprise a modified nucleobase.

17. The method of claim 16, wherein the modified nucleobase comprises at least one of N6-methyladenosine, inosine, N 1 -propylpseudouridine. N1 -methoxymethylpseudouridine. N1 -ethylpseudouridine, 5 -methoxy cytidine, 5-hydroxyuridine, 5-carboxyuridine, 5- formyluridine, 5-hydroxy cytidine, 5-hyrdoxymethylcytidine, 5-hydroxymethyluridine, 5- formylcytidine, 5-carboxycytidine, N4-methylcytidine, pseudoisocytidine, 2-thiocytidine, 4- thiocytidine, N1 -methylpseudouridine, pseudouridine, 5-methyluridine, 5 -methoxy uridine.dihydrouridine, 5-methylcytidine, 4-thiouridine, 2-thiouridine, uridine-5’-O-(l- thiophosphate), 5 -aminoallyluridine, or 4-acetylcytidine.

18. The method of any one of claims 10-17, wherein the cis-regulatory element activates the translation of the mRNA molecule, and wherein a translational efficiency of the mRNA molecule is increased by about 50% or more in comparison to an mRNA molecule lacking the cis-regulatory element but otherwise has the same sequence.

19. The method of any one of claims 10-17, wherein the cis-regulatory element represses the translation of the mRNA molecule, and wherein a translational efficiency of the mRNA molecule is decreased by about 30% or more in comparison to an mRNA molecule lacking the cis-regulatory element but otherwise has the same sequence.

20. The method of any one of claims 10-19, wherein the cis-regulatory element activates or represses the translation of the mRNA molecule in the cell at a certain cellular stage.

21. The method of any one of claims 10-20, wherein the cis-regulatory element is located in a 5’ untranslated region (5’-UTR), at a region including the start codon of the main open reading frame, within the open reading frame, or in a 3‘ untranslated region (3'-UTR).