A method for MRL prediction and sequence generation based on encoding / decoding structure
By generating 5'UTR sequences with high MRL values using a deep neural network model based on an encoding/decoding structure, the problems of length generalization and prediction error in existing technologies are solved, thereby improving the translation efficiency and R&D efficiency of mRNA vaccines.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WESTGENE BIOPHARMA CO LTD
- Filing Date
- 2025-11-20
- Publication Date
- 2026-05-26
Smart Images

Figure CN122090937A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics, specifically to a method for MRL prediction and sequence generation based on an encoding / decoding structure. Background Technology
[0002] During the COVID-19 pandemic, mRNA technology was used to develop mRNA vaccines to limit the spread of the virus. In the body, mRNA binds to ribosomes and is subsequently translated into proteins. mRNA is generally not transported to the cell nucleus but is directly translated into proteins in the cytoplasm. This allows mRNA to avoid potential insertional mutations, making it safer. Therefore, mRNA therapy is gradually becoming an alternative to protein and peptide therapies for treating various indications, and its development prospects are optimistic. However, current mRNA vaccines still have many shortcomings in translation capabilities, and further improvements are needed to enhance the efficiency of mRNA translation into proteins to achieve sufficient therapeutic efficacy.
[0003] In mRNA vaccines, the 5'UTR sequence, located at the start of mRNA, plays a crucial role in regulating translation and protein expression. The translational capacity of mRNA significantly impacts vaccine efficacy; therefore, designing an efficient 5'UTR sequence ensures sufficient protein production during translation, prompting mRNA vaccines to elicit a robust immune response.
[0004] Existing methods for optimizing 5' UTR sequences include optimizing 5' UTR sequences based on Kozak sequences and screening from 4248 known mRNA sequences, followed by motif modification. However, these methods suffer from limitations: they cannot achieve the desired mRNA secondary structure, and they are costly (requiring screening from a large number of sequences), have long experimental cycles (relying on a series of biological experiments to gradually modify and optimize the sequence), and are ineffective for optimizing and modifying known motifs on natural sequences, particularly for a large number of unknown mRNA UTR motifs. Traditional motif modification methods, which test and optimize from natural sequences, have a limited search range for sequence base combinations and low design efficiency.
[0005] The existing patent CN116453600A discloses a method for constructing an optimized 5'UTR sequence. This method involves obtaining a sequence feature matrix of an initial 5'UTR sequence; then inputting the sequence feature matrix into a trained machine learning model to generate multiple updated 5'UTR sequences; inputting these updated 5'UTR sequences into the trained machine learning model respectively; and using the encoder module and the MRL value prediction module to obtain predicted MRL values corresponding to the multiple updated 5'UTR sequences; finally, determining the optimized 5'UTR sequence from the multiple updated 5'UTR sequences based on the predicted MRL values. The shortcomings of this disclosed optimization method are: it only optimizes 5'UTR sequences on a fixed-length 50nt modified 5'UTR dataset, and cannot generalize to non-fixed-length datasets of 25nt to 100nt; furthermore, it has significant errors in predicting MRL values for sequences of lengths other than 50nt. Therefore, it cannot optimize sequences of lengths other than 50nt, or the optimization results are poor, which does not meet current mRNA applications.
[0006] Therefore, there is an urgent need in this field to develop a more widely applicable method for optimizing 5' UTR sequences in order to improve or solve the technical problems existing in current sequence optimization. Summary of the Invention
[0007] This application is made by the inventors based on their findings regarding the following problems and facts:
[0008] This paper addresses the technical shortcomings of current 5'UTR optimization methods, such as the inability to generalize to non-fixed-length datasets (25 to 100 nt) when training machine learning models on modified 5'UTR datasets for 5'UTR sequence prediction and generation (goal-driven optimization). These shortcomings include the model's limited training on a fixed-length modified 5'UTR dataset (50 nt) and its inability to generalize to non-fixed-length datasets (25 nt to 100 nt). The paper also addresses the error in the model's prediction of sequence MRL values.
[0009] The present invention aims to at least partially solve one of the above-mentioned technical problems or at least provide a useful commercial option.
[0010] Therefore, in a first aspect, the present invention proposes an MRL prediction and sequence generation method based on an encoding / decoding structure; according to an embodiment of the present invention, the method includes:
[0011] Step 1: One-hot encoded sequence;
[0012] Step 2: Build a deep neural network model;
[0013] Step 3: Train the deep neural network model;
[0014] Step 4: Generate sequences with higher MRL values.
[0015] Furthermore, in this invention, the sequence length is 25 nt to 100 nt;
[0016] In this invention, step 1 includes: for sequences with a length of 25nt to 100nt, the non-100nt sequences are filled with 'N' to a length of 100nt using zero-padding.
[0017] In this invention, step 1 includes: using one-hot encoding, encoding the four bases in the sequence: adenine nucleoside (A), cytosine nucleoside (C), guanine nucleoside (G), and uracil nucleoside (U) as follows: , , , At the same time, the filled 'N' is encoded as .
[0018] In this invention, step 1 employs one-hot encoding to encode a sequence of length L into... The matrix; where L is 25~100.
[0019] In this invention, the matrix is a second two-dimensional matrix obtained by encoding the sequence to form a first two-dimensional matrix and then decoding it using a decoder.
[0020] It should be noted that the 5' UTR sequence data described in this application comes from a dataset published in the reference (Sample PJ, Wang B, Reid DW, Presnyak V, McFadyen IJ, Morris DR, Seelig G. Human 5' UTR design and variant effect prediction from a massively parallel translation assay. Nat Biotechnol. 2019 Jul;37(7):803-809.). The dataset can be downloaded from: https: / / www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSE114002.
[0021] According to an embodiment of the present invention, the sequence is selected from at least the top 100,000 sequences with the most total RNA reads.
[0022] According to an embodiment of the present invention, the sequence is selected from the top 200,000 to 500,000 sequences with the most total RNA reads.
[0023] According to an embodiment of the present invention, the sequence is selected from the top 200,000 sequences with the most total RNA reads.
[0024] According to an embodiment of the present invention, the total RNA reads include the sequencing depth of the 5' UTR sequence.
[0025] It should be noted that the total RNA readings were determined by the sequencing depth of the 5' UTR sequence.
[0026] According to an embodiment of the present invention, the sequence length is 20-100 bp.
[0027] According to an embodiment of the present invention, the sequence length is 20-50 bp.
[0028] According to an embodiment of the present invention, the sequence length is 50-100 bp.
[0029] In this invention, the sequence is a 5' UTR sequence, a 3' UTR sequence, or a Poly A sequence;
[0030] In a specific embodiment of the present invention, the sequence is a 5' UTR sequence;
[0031] According to embodiments of the present invention, the length of the 5' UTR sequence affects the translation efficiency of mRNA. In annotated vertebrate mRNA sequences, 75% of mRNA 5' UTR lengths range from 20-100 bp, and 25% range from 100-300 bp. mRNA 5' UTR lengths reaching 1000 bp are extremely rare. The longer the 5' UTR, the more likely it is to form hairpin structures, the presence of which can halt protein synthesis.
[0032] In this invention, in step 2, the model includes three modules: Encoder, PredHead, and Decoder; the skeleton networks of the Encoder and Decoder are 3-layer U-Nets; and the PredHead is an 18-layer ResNet.
[0033] In this invention, in step 2, there is no skip connection between the corresponding layers of the Encoder and Decoder.
[0034] In this invention, step 3, training the deep neural network model, is divided into two stages:
[0035] Phase 1: Supervised Training Phase;
[0036] Phase Two: Self-Supervised Training Phase;
[0037] In this invention, the supervised training phase specifically involves connecting the Encoder and PredHead modules as a predictor for training.
[0038] Furthermore, the supervised training employs a weighted mean-square error (WMSE) loss function to train the predictor;
[0039] In this invention, the WMSE loss function is as follows: Equation 1:
[0040] Formula 1
[0041] In Equation 1,
[0042] For sequence The true normalized MRL value;
[0043] It is the normalized MRL value predicted by the model;
[0044] When calculating the loss, it is a sequence. Additional weights assigned;
[0045] In this invention, normalization is z-score normalization. Calculate as shown in Equation 2:
[0046] Formula 2
[0047] In Equation 2,
[0048] Represents a sequence The original MRL value;
[0049] This represents the average MRL value of all sequences involved in the training.
[0050] This represents the standard deviation of the MRL values for all sequences involved in the training process.
[0051] In this invention, during the supervised training phase, the learning rate (lr) when training the Encoder is: ;
[0052] In this invention, the self-supervised training phase specifically involves connecting the Encoder and Decoder modules as a Reconstructor for training;
[0053] Furthermore, the self-supervised training uses the parameters of the Encoder module trained in the supervised training phase of phase one to initialize the Encoder in the reconstructor;
[0054] In this invention, the self-supervised training is performed by minimizing the binary cross-entropy (BCE) loss function to train the reconstructor.
[0055] In this invention, the loss function that minimizes the binary cross-entropy (BCE) is as follows: Equation 3:
[0056] Formula 3
[0057] Indicates the model's predicted value;
[0058] Indicates the target value;
[0059] Indicates the number of samples;
[0060] In this invention, the learning rate during the self-supervised training phase when training the Encoder is... ;
[0061] In this invention, during the self-supervised training phase, the learning rate during Decoder training is... ;
[0062] In this invention, in step 4, the generation of higher MRL value sequences adopts a gradient ascent sequence optimization method to iteratively generate sequences with high MRL values.
[0063] Furthermore, the gradient ascent sequence optimization method specifically involves: setting the maximum number of iterations to... In the iteration The gradient calculation of Equation 4 is performed in the following manner;
[0064] Formula 4
[0065] In Equation 4,
[0066] Indicates the first The potential vector at the next iteration;
[0067] Indicates the first The potential vector at the next iteration;
[0068] and It is the amplitude coefficient, which is a hyperparameter;
[0069] The mean is The variance is Gaussian distribution;
[0070] In this invention, a setting is provided. , .
[0071] In this invention, the function represented by the PredHead module is: ,but Calculated from Equation 5;
[0072] Formula 5
[0073] According to embodiments of the present invention, the model training method of the present invention can be used to develop a method for constructing optimized 5' UTR sequences, thereby obtaining 5' UTR sequences with higher MRL values. The expression efficiency of the open reading frame (ORF) encoding the target protein or polypeptide constructed using the 5' UTR sequence with higher MRL values is effectively improved. The method of the present invention can be further used in mRNA drug development, which can improve the development speed, reduce the development cost, and improve the effectiveness of mRNA drugs.
[0074] In a second aspect, the present invention provides a 5' UTR sequence.
[0075] According to an embodiment of the present invention, the 5'UTR is obtained by the method described in the first aspect of the present invention.
[0076] According to embodiments of the present invention, the expression efficiency of the ORF encoding the target protein or polypeptide of the mRNA constructed using the 5' UTR sequence obtained by the present invention is effectively improved, which can effectively reduce the dosage of mRNA drugs and increase the dosing interval, thereby improving the bioavailability of mRNA drugs and having good clinical application value.
[0077] In a third aspect, the present invention provides a model structure; according to an embodiment of the present invention, the model structure includes: an Encoder module, a PredHead module, and a Decoder module; the Encoder module and the PredHead module are connected as a Predictor, and the Predictor is trained with a loss function using the original sequence MRL values; the Encoder module and the Decoder module are connected as a Reconstructor, and the Reconstructor is trained with a loss function using the parameters in the Predictor;
[0078] The predictor, together with the reconstructor, optimizes the model's training based on gradient ascent through supervised and self-supervised training phases, iteratively generating sequences with high MRL values.
[0079] According to embodiments of the present invention, the model structure is used to generate a new 5' UTR sequence. The expression efficiency of the ORF encoding the target protein or polypeptide constructed using the newly generated 5' UTR sequence is effectively improved. Therefore, the model structure of the present invention can be used for mRNA drug development, which can improve the development speed, reduce the development cost, and improve the effectiveness of mRNA drugs.
[0080] In a fourth aspect, the present invention provides an apparatus for constructing an optimized 5'UTR sequence. According to an embodiment of the invention, the apparatus includes:
[0081] The feature acquisition unit obtains the MRL value of the initial 5' UTR sequence from the input initial 5' UTR sequence; and encodes and pads the sequence to 100 nt.
[0082] The optimization unit trains the deep neural network model based on the original MRL values of the sequence.
[0083] The generation unit uses gradient ascent iteration to generate 5' UTR sequences with high MRL values;
[0084] The deep neural network model is trained using the method described in the first aspect of this invention.
[0085] According to an embodiment of the present invention, the apparatus is used to obtain an optimized 5' UTR sequence with a high MRL value.
[0086] In a fifth aspect, the present invention provides a computing device. According to an embodiment of the invention, the computing device includes: a memory and a processor;
[0087] The memory is used to store computer programs;
[0088] The processor is configured to execute the computer program to implement the method described in the first aspect of the present invention.
[0089] According to an embodiment of the present invention, the processor runs a program corresponding to the executable program code stored in the memory by reading the executable program code stored in the memory, for implementing a method for model training and constructing optimized 5' UTR sequences.
[0090] In a sixth aspect, the present invention provides a computer-readable storage medium. According to an embodiment of the invention, the storage medium includes computer instructions that, when executed by a computer, cause the computer to perform the method described in the first aspect of the invention.
[0091] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0092] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:
[0093] Figure 1 This is an optimized 5' UTR sequence flowchart according to an embodiment of the present invention;
[0094] Figure 2 This is a diagram of a deep neural network model consisting of three modules: Encoder, PredHead, and Decoder.
[0095] Figure 3 This is the flowchart of the optimization algorithm in the t-th iteration;
[0096] Figure 4 The hEPO expression levels of hEPO mRNA containing 5' UTR before and after optimization in HEK293 cells in a 50nt fixed-length model;
[0097] Figure 5 The hEPO expression level of hEPO mRNA in HEK293 cells, including the 5' UTR before and after optimization, in a 25-100nt variable length model.
[0098] Figure 6 The hEPO expression levels in HEK293 cells are those of hEPO mRNA containing a variable length model of 25-100 nt with optimized receptive field and hEPO mRNA containing the unoptimized 5' UTR sequence.
[0099] Figure 7 This represents the hEPO protein expression level in mouse liver using 5UTRnet-optimized hEPO-mRNA.
[0100] Figure 8 This represents the hEPO protein expression level in mouse spleen after 5UTRnet-optimized hEPO-mRNA;
[0101] Figure 9 This refers to the fluorescence signal intensity generated by the expression of luciferase in mouse liver using 5UTRnet-optimized Fluc-mRNA;
[0102] Figure 10 This is the intensity of the fluorescence signal generated after luciferase was expressed in the spleen of mice using 5UTRnet-optimized Fluc-mRNA. Detailed Implementation
[0103] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0104] Definitions and Explanations
[0105] In this document, unless otherwise stated, the singular forms “a,” “an,” etc., include plural referents (more than one); “a group” or “a plurality” refers to two or more.
[0106] In this document, unless otherwise stated, the terms “first,” “second,” “third,” “fourth,” etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated; features defined with “first,” “second,” etc., may explicitly or implicitly include one or more of the stated features.
[0107] In this paper, the term "UTR" (untranslated region) refers to a segment of an mRNA molecule that is not translated and lies upstream of the start codon and downstream of the stop codon. These regions are transcribed along with the coding regions and are present in mature mRNA. The UTR upstream of the start codon in mRNA is called the 5' UTR.
[0108] In this paper, the term "mean ribosome load" (MRL), determined through polysome profiling experiments, represents the binding affinity of an mRNA sequence to a ribosome, thereby indicating the translation efficiency of the mRNA molecule. MRL serves as the training label for the machine learning model of this invention.
[0109] In this paper, the term "goal-driven optimization" refers to the optimization of the model in a specific direction during each 5' UTR sequence optimization process, which is guided by the input MRL value and noise signal.
[0110] In this article, the term "open reading frame" (ORF) refers to a portion of the genome of an organism, which may be a protein-coding sequence. The ORF in a gene is contained between the start and end coding sequences and is equivalent to "open reading frame," "open reading frame," or "open reading rack."
[0111] In this article, "Encoder" refers to the encoder that extracts features from the input sequence;
[0112] In this paper, "PredHead" is the prediction head that predicts sequence label values based on sequence features;
[0113] In this article, "Decoder" is a decoder that decodes the features of a sequence into a sequence;
[0114] In this paper, "U-Net" is a U-shaped fully convolutional neural network model;
[0115] In this article, "Resnet" stands for Residual Neural Network, a neural network model that addresses the vanishing or exploding gradient problems that occur when the number of model layers increases.
[0116] In this article, "skip connection" refers to a residual connection, also called a skip connection, which adds input data directly to the output of a certain layer of the network.
[0117] The learning rate (lr) is a key hyperparameter in deep learning, determining the speed at which the model updates its weights during training. The size of the learning rate directly affects the model's learning progress; an excessively large learning rate may lead to explosive or oscillating loss values, while an excessively small rate may result in overfitting or slow convergence.
[0118] Gradient ascent sequence optimization is a process that uses gradients to guide iterative optimization of sequences, so that the generation of sequences follows the direction of gradient ascent.
[0119] It should be noted that, unless otherwise stated, single-stranded nucleic acid molecules in this article are written from left to right in the 5' to 3' direction.
[0120] Embodiments of the present invention will now be described in more detail, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the invention, and should not be construed as limiting the invention.
[0121] Example 1: Construction of optimized 5' UTR sequence
[0122] According to an embodiment of the present invention, a 5' UTR sequence is constructed using the method described above.
[0123] like Figure 1 As shown,
[0124] First, existing 5' UTR sequences were selected. For existing 5' UTR sequences of length 25nt to 100nt (SEQ ID NO: 25 to SEQ ID NO: 29), zero-padding was used to pad the sequence length with 'N' to 100nt. Then, one-hot encoding was used to encode the four bases in the above 5' UTR sequences: adenine nucleoside (A), cytosine nucleoside (C), guanine nucleoside (G), and uracil nucleoside (U) as follows: , , , At the same time, the filled 'N' is encoded as The matrices of SEQ ID NO: 1 to SEQ ID NO: 5, with 5'UTR sequences encoded as 100*4, are obtained; the matrices of SEQ ID NO: 25 to SEQ ID NO: 29 and 100*4 are shown in Table 1.
[0125] Second, build a deep neural network model with three modules: Encoder, PredHead, and Decoder, such as... Figure 2 As shown in (a~c), the skeleton network of the Encoder and Decoder is a 3-layer U-Net; the PredHead is an 18-layer ResNet; the Encoder and Decoder modules have the ability to work independently, and there are no residual connections (skip connections) between corresponding layers of the Encoder and Decoder.
[0126] Third, the matrix of the encoded existing 5'UTR sequence is input into the model for supervised training. Specifically, the Encoder and PredHead modules in the model are connected as a predictor, and the predictor is trained using the weighted mean-square error (WMSE) loss function. The learning rate (lr) during Encoder training is... ;
[0127] The WMSE loss function is as follows: Equation 1:
[0128] Formula 1
[0129] In Equation 1,
[0130] For sequence The true normalized MRL value;
[0131] It is the normalized MRL value predicted by the model;
[0132] When calculating the loss, it is a sequence. Additional weights assigned;
[0133] Normalization is achieved through z-score normalization. Calculate as shown in Equation 2:
[0134] Formula 2
[0135] In Equation 2,
[0136] Represents a sequence The original MRL value;
[0137] This represents the average MRL value of all sequences involved in the training.
[0138] This represents the standard deviation of the MRL values for all sequences involved in the training process.
[0139] Fourth, the Encoder and Decoder modules are connected as a reconstructor and trained under self-supervised conditions. Specifically, the parameters of the Encoder module trained in the third supervised training phase are used to initialize the Encoder in the reconstructor, and the reconstructor is trained by minimizing the Binary CrossEntropy (BCE) loss function. The learning rate during Encoder training is... Learning rate during Decoder training ;
[0140] The loss function that minimizes the binary cross-entropy (BCE) is as follows: Equation 3:
[0141] Formula 3
[0142] Indicates the model's predicted value;
[0143] Indicates the target value;
[0144] Indicates the number of samples;
[0145] Fifth, after the deep neural network model is trained, the model uses a gradient ascent-based sequence optimization algorithm to iteratively calculate and generate a 5'UTR sequence with a high MRL value for the input 5'UTR sequence. In the gradient ascent-based sequence optimization algorithm, let the maximum number of iterations be T. In the t-th iteration, The optimization algorithm flowchart is as follows: Figure 3 As shown, the MRL value is calculated in the Encoder and PredHead modules for the input 5'UTR sequence (Table 2).
[0146] In this embodiment, the gradient ascent sequence optimization algorithm, such as Figure 3 The equation, based on the input feature matrix latent vector of the current 5' UTR sequence and the backpropagation latent vector corresponding to the initial (or previous) 5' UTR sequence, calculates the gradient and then predicts the MRL value of the current 5' UTR sequence. The predicted MRL value is compared with the previously predicted MRL value, and iterative gradient ascent is performed to calculate the predicted MRL value of the 5' UTR sequence. When the number of iterations is set to t, the predicted MRL value is calculated... t Greater than Output MRL t To predict and calculate the MRL value, the decoder outputs the 5'UTR sequence constructed by the t-th prediction.
[0147] Table 1
[0148]
[0149] Using the above method, the sequences to be optimized (SEQ ID NO: 1, 2, 3, 4, 9, 10, 11, 12, 17, 18, 19, 20) were input and gradient ascent sequence optimization was performed to construct 5' UTR sequences (SEQ ID NO: 5, 6, 7, 8, 13, 14, 15, 16, 21, 22, 23, 24). The sequences before and after optimization and the MRL values are shown in Table 2 below.
[0150] Table 2
[0151]
[0152] Example 2: Detection of expression efficiency of the newly obtained 5'UTR sequence
[0153] The optimized new 5' UTR sequences 5, 6, 7, 8, 13, 14, 15, 16, 21, 22, 23, and 24 obtained in Example 1 were subjected to expression efficiency testing according to the following steps, and compared with the sequences before optimization SEQ ID NO: 1, 2, 3, 4, 9, 10, 11, 12, 17, 18, 19, and 20:
[0154] 1. hEPO-mRNA preparation
[0155] A plasmid vector expressing hEPO (human erythropoietin) was prepared, containing identical sequences except for the 5' UTR region, along with the target gene region (open reading frame), 3' UTR region, and PolyA tail region. Based on the structural characteristics of the plasmid vector, the plasmid was linearized (digested) with the restriction endonuclease BsaI. The digestion results were identified by gel electrophoresis, and finally, the linearized plasmid was purified by ethanol precipitation to obtain the template plasmid for in vitro mRNA transcription.
[0156] mRNA was synthesized using adenine ribonucleoside, guanine ribonucleoside, cytosine ribonucleoside, and uracil ribonucleoside as raw materials, following an in vitro transcription kit and the manufacturer's instructions. The target gene segment of the mRNA encodes hEPO.
[0157] hEPO-mRNA is a commonly used reporter gene used to examine the expression levels of mRNA sequences containing different 5' UTRs in cells.
[0158] 2. Preparation of hEPO-mRNA formulations
[0159] hEPO-mRNA formulations were prepared using microfluidic technology. SM102, DOPE, Chol, and DMG-PEG2000 were dissolved in anhydrous ethanol at a molar ratio of 50:10:38.5:1.5 to prepare an organic phase, achieving a SM102 concentration of 12 mg / mL. The prepared hEPO-mRNA was dissolved in 50 mM citrate buffer solution (prepared with RNase-free water) to form an aqueous phase, achieving an mRNA concentration of 0.4 mg / mL. The aqueous phase to organic phase volume ratio was controlled at 3:1, and the flow rate was fixed at 12 mL / min. The mixtures were then self-assembled using a microfluidic chip to form lipid nanoparticles loaded with hEPO mRNA. After ultrafiltration with 10 mM pH 6.0 buffer, the hEPO-mRNA formulation was obtained.
[0160] 3. HEK293 cell transfection experiment
[0161] HEK293 cells in logarithmic growth phase were collected, resuspended in culture medium, counted, and the cell density was adjusted to 2 × 10⁻⁶. 5 Cells / mL. Add 0.5 mL of complete culture medium to each well of a 24-well plate, followed by 0.5 mL of cell suspension to achieve a cell density of 1 × 10⁻⁶ cells / mL. 5 Mix 1 µg of hEPO-mRNA containing different 5' UTRs into each well of the 96-well plate and incubate for 18-24 h. After overnight incubation, replace the medium with 0.5 mL of complete culture medium. Add 1 µg of hEPO-mRNA containing different 5' UTRs to each well of the 96-well plate and mix with 2 µL of lipo2000 transfection reagent. Incubate for 30 min, repeating the process in triplicate for each mRNA. After 24 h of administration, centrifuge and collect the supernatant. Analyze the transfection efficacy of different formulations using a human erythropoietin (hEPO) enzyme-linked immunosorbent assay (ELISA) kit.
[0162] The results are as follows:
[0163] (1) In the 50nt fixed-length model, such as Figure 4 As shown: 4 optimized 5' UTRs: SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7 and SEQ ID NO: 8, the average hEPO expression levels of HEK293 cells in each group were 227.7528 mIU / mL, 58.0679 mIU / mL, 110.6355 mIU / mL and 91.6280 mIU / mL, respectively;
[0164] Among the four optimized 5' UTRs that were not optimized using the method of this application, the average hEPO expression levels of HEK293 cells transfected in SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO: 4 were 140.8676 mIU / mL, 40.6756 mIU / mL, 10.0539 mIU / mL, and 5.2658 mIU / mL, respectively.
[0165] (2) In the 25-100nt variable length model, such as Figure 5 As shown: 4 optimized 5' UTRs: SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15 and SEQ ID NO: 16, the average hEPO expression levels of HEK293 cells in each group were 238.3948 mIU / mL, 204.1433 mIU / mL, 187.9048 mIU / mL and 163.1346 mIU / mL, respectively;
[0166] In the four optimized 5' UTRs without using the method of this application, the average hEPO expression levels of HEK293 cells transfected in SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11 and SEQ ID NO: 12 were 22.8332 mIU / mL, 164.6986 mIU / mL, 78.6199 mIU / mL and 128.7483 mIU / mL, respectively.
[0167] (3) 25-100nt variable length model plus receptive field optimization sequence, such as Figure 6 As shown: 4 optimized 5' UTRs: SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23 and SEQ ID NO: 24, the average hEPO expression levels of HEK293 cells in each group were 336.7574 mIU / mL, 281.4328 mIU / mL, 174.7508 mIU / mL and 214.7054 mIU / mL, respectively;
[0168] In the four pre-optimization 5' UTRs (SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, and SEQ ID NO: 20) without using the method of this application, the average hEPO expression levels in HEK293 cells were 93.1556 mIU / mL, 72.0119 mIU / mL, 24.8320 mIU / mL, and 46.3971 mIU / mL, respectively.
[0169] The above data show that the hEPO-mRNA constructed from the 12 optimized 5' UTR sequences (SEQ ID NO: 5, 6, 7, 8, 13, 14, 15, 16, 21, 22, 23, 24) obtained in Example 1 of this invention has a higher hEPO expression efficiency in HEK293 cells than the 5' UTR sequences (SEQ ID NO: 1, 2, 3, 4, 9, 10, 11, 12, 17, 18, 19, 20) before optimization.
[0170] Example 3: Expression effects of hEPO-mRNA containing optimized 5'UTR and NCA-7d 5'UTR in mouse liver and spleen.
[0171] The preparation method of hEPO-mRNA containing NCA-7d 5'UTR and its formulation is the same as in Example 2, except that the target gene segment (open reading frame) is replaced with a nucleic acid sequence expressing LUC.
[0172] Female 7-week-old BALB / c mice were intravenously injected via tail vein with an hEPO-mRNA formulation containing SEQ ID NO: 21, SEQ ID NO: 22, and NCA-7d 5'UTR (10 µg mRNA). Liver and spleen were collected 24 hours after administration. The tissue homogenates were prepared by grinding in liquid nitrogen and centrifuged at 5000 ×g for 5–10 minutes at 2–8°C. The supernatant was collected. The hEPO protein content in the mouse liver and spleen was detected and calculated using the Human EPO (Erythropoietin) ELISA Kit (Elabscience, E-EL-H3640).
[0173] The results are as follows Figures 7-8 In mouse liver and spleen, hEPO-mRNA containing the optimized 5'UTR sequence of this invention showed good in vitro protein expression levels. Compared with hEPO-mRNA containing the NCA-7d 5'UTR sequence, hEPO-mRNA containing the model-optimized high MRL value 5'UTR sequence (SEQ ID NO: 21, SEQ ID NO: 22) of this invention showed higher hEPO protein expression levels.
[0174] Example 4: Expression effects of LUC-mRNA containing optimized 5'UTR and LUC-mRNA containing NCA-7d 5'UTR in mouse liver and spleen.
[0175] Preparation of LUC-mRNA containing optimized 5'UTR and LUC-mRNA containing NCA-7d 5'UTR and its formulations: same as the preparation method of hEPO-mRNA and its formulations in Example 2.
[0176] Balb / c mice were intravenously injected with a LUC-mRNA preparation (10 µg mRNA) containing SEQ ID NO: 21, SEQ ID NO: 22 and NCA-7d 5'UTR. Six hours later, 3 mg of luciferase substrate was injected intraperitoneally. The mice were sacrificed within 15 minutes, and the liver and spleen were harvested. The fluorescence expression in the liver and spleen of the mice was detected by a small animal imaging system spectrometer.
[0177] The results are as follows Figures 9-10 LUC-mRNA containing the optimized 5'UTR sequence of the present invention is well expressed in vivo; compared with LUC-mRNA containing the NCA-7d 5'UTR sequence, LUC-mRNA containing the model-optimized high MRL value 5'UTR sequence (SEQ ID NO: 21, SEQ ID NO: 22) of the present invention has a higher LUC protein expression level.
[0178] It should be noted that the logic and / or steps represented in the flowcharts or otherwise described herein in this application are as follows. For example, a particular sequence of executable instructions for implementing logical functions can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: electrical connections having one or more wires (electronic devices), portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0179] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0180] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0181] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0182] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0183] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.
Claims
1. A method for MRL prediction and sequence generation based on encoding / decoding structure, characterized in that, include: Step 1: One-hot encoded sequence; Step 2: Build a deep neural network model; Step 3: Train the deep neural network model; Step 4: Generate sequences with higher MRL values; In step 1, the sequence length is 25nt to 100nt.
2. The method according to claim 1, characterized in that, In step 1, for sequences with a length of 25nt to 100nt that are not 100nt, the length is filled with 'N' to 100nt using zero padding.
3. The method according to claim 2, characterized in that, In step 1, the four bases in the sequence—adenine nucleoside (A), cytosine nucleoside (C), guanine nucleoside (G), and uracil nucleoside (U)—are encoded as follows: , , , At the same time, the filled 'N' is encoded as .
4. The method according to claim 1, characterized in that, In step 2, the model includes three modules: Encoder, PredHead, and Decoder; the skeleton networks of the Encoder and Decoder are 3-layer U-Nets; and the PredHead is an 18-layer ResNet.
5. The method according to claim 4, characterized in that, In step 2, there is no residual connection between the corresponding layers of the Encoder and Decoder.
6. The method according to claims 1 to 3, characterized in that, The sequence is a 5' UTR sequence, a 3' UTR sequence, or a Ploy A sequence.
7. The method according to claim 1, characterized in that, Step 3 involves training the deep neural network model in two phases: Phase 1: Supervised Training Phase; Phase Two: Self-Supervised Training Phase.
8. The method according to claim 7, characterized in that, The supervised training phase specifically involves connecting the Encoder and PredHead modules as a predictor for training; wherein, the supervised training employs minimizing a weighted mean squared error loss function to train the predictor.
9. The method according to claim 8, characterized in that, During the supervised training phase, the learning rate for training the Encoder is... .
10. The method according to claim 7, characterized in that, The self-supervised training phase specifically involves connecting the Encoder and Decoder modules as a reconstructor for training. The self-supervised training initializes the Encoder in the reconstructor using the parameters of the Encoder module trained in the supervised training phase of phase one; and trains the reconstructor by minimizing the binary cross-entropy loss function.
11. The method according to claim 10, characterized in that, During the self-supervised training phase, the learning rate when training the Encoder is... ; Learning rate during Decoder training .
12. The method according to claim 1, characterized in that, In step 4, the generation of higher MRL value sequences is achieved by using a gradient ascent sequence optimization method to iteratively generate sequences with high MRL values.
13. A 5' UTR sequence, characterized in that, The 5' UTR is obtained by the method described in any one of claims 1 to 12.
14. A model structure, characterized in that, The model structure includes: The system comprises an Encoder module, a PredHead module, and a Decoder module. The Encoder and PredHead modules are connected to form a Predictor, which is trained using a loss function based on the MRL values of the original sequence. The Encoder and Decoder modules are also connected to form a Reconstructor, which is trained using a loss function based on the parameters from the Predictor. The predictor, together with the reconstructor, optimizes the model's training based on gradient ascent through supervised and self-supervised training phases, iteratively generating sequences with high MRL values.
15. A computing device, characterized in that, include: Processor and memory; The memory is used to store computer programs; The processor is configured to execute the computer program to implement the method as described in any one of claims 1 to 12.
16. A computer-readable storage medium, characterized in that, The storage medium includes computer instructions that, when executed by a computer, cause the computer to perform the method as described in any one of claims 1 to 12.