Method and apparatus for optimizing mRNA sequence, mRNA molecule, pharmaceutical composition and uses thereof
By co-optimizing the 5' UTR and CDS of mRNA sequences to maximize scores reflecting translation efficiency and stability, the method addresses the limitations of current mRNA design methods, leading to improved protein yields and mRNA vaccine effectiveness.
Patent Information
- Application Number
- JP2025031528
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-30
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-28
AI Technical Summary
Current mRNA design methods fail to optimize the translation efficiency and stability of mRNA sequences from a holistic perspective, leading to suboptimal protein yields in mRNA vaccines and treatment methods.
A method and apparatus for optimizing mRNA sequences by co-optimizing the 5' untranslated region (UTR) and coding sequences (CDS) to maximize a score reflecting translation initiation efficiency, codon adaptation index, and minimum free energy, thereby enhancing protein synthesis efficiency and stability.
The optimized mRNA sequences result in improved protein yields and enhanced effectiveness of mRNA vaccines and treatment methods by balancing translation efficiency and stability.
Smart Images

Figure 2025087770000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to the field of technologies such as biocomputing. Specifically, it relates to a method and apparatus for optimizing mRNA sequences, an electronic device, a computer-readable storage medium, a computer program product, an mRNA molecule, a pharmaceutical composition, and uses thereof.
Background Art
[0002] Messenger Ribonucleic Acid (mRNA) vaccines and treatment methods have received extensive attention because they have the potential to combat various diseases, including infectious diseases and cancers. The translation efficiency and stability of mRNA sequences are particularly important for the design of mRNA sequences.
[0003] The methods described in this section are not necessarily the methods previously envisioned or adopted. Unless otherwise specified, none of the methods described in this section should be considered to be prior art merely by virtue of being included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to be recognized in the prior art.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present disclosure provides a method and apparatus for optimizing mRNA sequences, an electronic device, a computer-readable storage medium, a computer program product, an mRNA molecule, a pharmaceutical composition, and uses of mRNA.
Means for Solving the Problems
[0005] According to one aspect of the present disclosure, a method for optimizing an mRNA sequence is provided. The method includes obtaining a first mRNA sequence for synthesizing a target protein, where the first mRNA sequence includes a 5' untranslated region sequence and a coding region sequence, and aiming to maximize a first score of the first mRNA sequence, and performing adjustments on the 5' untranslated region sequence and the coding region sequence to obtain an optimized second mRNA sequence for synthesizing the target protein. Here, the first score reflects at least one of the translation initiation efficiency, codon adaptation index, and minimum free energy of the first mRNA sequence.
[0006] According to another aspect of the present disclosure, an apparatus for optimizing an mRNA sequence is provided. The apparatus includes an acquisition unit configured to obtain a first mRNA sequence for synthesizing a target protein, where the first mRNA sequence includes a 5' untranslated region sequence and a coding region sequence, and a processing unit configured to aim to maximize a first score of the first mRNA sequence and perform adjustments on the 5' untranslated region sequence and the coding region sequence to obtain an optimized second mRNA sequence for synthesizing the target protein. Here, the first score reflects at least one of the translation initiation efficiency, codon adaptation index, and minimum free energy of the first mRNA sequence.
[0007] According to another aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processor and a memory storing a computer program. When the computer program is executed by the processor, the processor is caused to execute the above method.
[0008] According to another aspect of the present disclosure, a computer-readable storage medium storing a computer program is provided. When the computer program is executed by a processor, the processor is caused to execute the above method.
[0009] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, which, when executed by a processor, causes the processor to execute the above method.
[0010] According to another aspect of the present disclosure, there is provided an mRNA molecule, the sequence of which is produced by the above method.
[0011] According to another aspect of the present disclosure, there is provided a pharmaceutical composition comprising an mRNA sequence or molecule produced by the above method and a pharmaceutically acceptable adjuvant.
[0012] According to another aspect of the present disclosure, there is provided the use of an mRNA sequence or molecule produced by the above method or the above pharmaceutical composition in the manufacture of a drug or a vaccine.
[0013] According to one or more embodiments of the present disclosure, aiming to maximize the first score of mRNA, by co-optimizing the 5'untranslated region (UTR) and coding sequences (CDS) of the optimized mRNA, from an overall perspective, by realizing targeted optimization of the translation efficiency and stability of mRNA, the final protein yield can be optimized, and the overall effect of mRNA vaccines and treatment methods can be improved.
[0014] It should be understood that the content described in this part is not intended to identify the key points or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily understood from the following description.
[0015] The drawings illustrate embodiments by way of example and constitute a part of the specification, and are used to explain exemplary embodiments of the embodiments together with the description in the text of the specification. The embodiments shown are merely for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to elements that are similar but not necessarily identical.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
DETAILED DESCRIPTION OF THE INVENTION
[0017] In the following description, for purposes of interpretation, specific details are set forth in order to provide an understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these details. As will be recognized by those skilled in the art, the embodiments of the present disclosure described below can be implemented in various ways, such as a process, apparatus, system, device, or method on a tangible computer-readable medium.
[0018] The assemblies or modules shown in the drawings are illustrative of exemplary embodiments of the present disclosure and are intended to avoid confusion of the present disclosure. Further, it should be understood that throughout the discussion, an assembly may be described as a single functional unit, and these functional units may include sub-units, but as would be recognized by those skilled in the art, various assemblies or portions thereof may be divided into separate assemblies or integrated together, for example, included in a single system or assembly. It should be noted that the functions or operations discussed herein may be implemented as an assembly. The assembly may be implemented in software, hardware, or a combination thereof.
[0019] Note that the connections between assemblies or systems in the drawings are not intended to be limited to direct connections. On the contrary, the data between these assemblies may be modified, reformatted, or otherwise changed by an intermediate assembly. Note that more or fewer connections may be used. Further, it should be noted that the terms "coupled", "connected", "communicatively coupled", "interface", "access", or any derivative thereof should be understood to include direct connections, indirect connections via one or more intermediate devices, and wireless connections. Further, it should be noted that any communication, such as a signal, response, reply, confirmation, message, search, etc., may include one or more information exchanges.
[0020] In the specification, the reference to "one or more embodiments", "preferred embodiments", "embodiment", "multiple embodiments", etc. indicates that a particular feature, structure, characteristic, or function described in connection with the embodiment is included in at least one embodiment of the present disclosure and may also be included in multiple embodiments. Note that the above phrases appearing at each location in the specification do not necessarily refer to the same embodiment or multiple embodiments.
[0021] The use of several terms in various places in the specification is for the purpose of explanation and should not be construed as restrictive. A service, function, or resource is not limited to a single service, function, or resource, and the use of these terms may refer to a set of related services, functions, or resources, which may be of a distributed or aggregated type. The terms "include", "comprise", "have", and "contain" should be understood as open terms, and any of the following lists are examples and do not mean being limited to the listed items. A "layer" may include one or more operations. Terms such as "optimal", "optimization", "optimize", etc. refer to the improvement of a result or process and do not require that the specified result or process has already reached an "optimal" or peak state. The use of terms such as memory, database, information bank, data storage, table, hardware, cache memory, etc. may be used in this specification to refer to a system assembly or assembly in which information can be input or recorded in other ways.
[0022] In one or more embodiments, the stop conditions may include: (1) that a set number of iterations have already been executed; (2) that a certain processing time has already been reached; (3) convergence (e.g., the difference between consecutive iterations is less than a first threshold); (4) divergence (e.g., the performance deteriorates); and (5) that an acceptable result has already been reached.
[0023] Those skilled in the art should recognize that: (1) some steps may be optionally executed; (2) the steps may not be limited to the specific order described in this specification; (3) some steps may be executed in a different order; and (4) some steps may be performed simultaneously.
[0024] Any title used in this specification is for organizational purposes only and should not be used to limit the specification or the claims. Each reference / document mentioned in this patent document is hereby incorporated by reference in its entirety into this specification.
[0025] It should be noted that any experiments and results provided in this specification are provided in an illustrative manner and are performed under specific conditions using specific examples. Therefore, none of these experiments and their results should be used to limit the scope of disclosure of this patent document.
[0026] The translation efficiency and stability of mRNA sequences are particularly important for the design of mRNA sequences. Here, the translation efficiency indicates how fast an mRNA sequence can produce proteins, and the stability indicates how long an mRNA sequence can continuously translate proteins within a certain period. Both the translation efficiency and stability determine the number of proteins that an mRNA sequence can generate and ultimately affect the actual effect of mRNA vaccines, drugs, or therapies.
[0027] mRNA design methods in the related art generally focus on the design of a single fragment in mRNA, such as the 5' untranslated region or the coding region, without considering the interactions between fragments, and cannot finely adjust the translation efficiency and stability of mRNA from an overall perspective.
[0028] In response to the above problems, embodiments of the present disclosure provide a method for optimizing mRNA sequences. The method aims to maximize the first score of the mRNA sequence and perform co-optimization on the 5' untranslated region and the coding region. The first score can reflect at least one of the translation initiation efficiency, the codon adaptation index, and the minimum free energy, thereby realizing targeted optimization of the translation efficiency and stability of the mRNA sequence according to design requirements, further optimizing the yield of the final target protein, and improving the overall effect of mRNA vaccines and treatment methods.
[0029] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the drawings.
[0030] According to one aspect of the present disclosure, a method for optimizing a messenger ribonucleotide (mRNA) sequence is provided. FIG. 1 shows a flowchart of an mRNA sequence optimization method 100 according to an embodiment of the present disclosure. As shown in FIG. 1, method 100 includes the following. Step S101, obtaining a first mRNA sequence for synthesizing a target protein, where the first mRNA sequence includes a 5' untranslated region sequence and a coding region sequence; step S102, aiming to maximize a first score of the first mRNA sequence, adjusting the 5' untranslated region sequence and the coding region sequence to obtain an optimized second mRNA sequence for synthesizing the target protein, where the first score reflects at least one of the translation initiation efficiency, codon adaptation index, and minimum free energy of the first mRNA sequence.
[0031] According to an embodiment of the present disclosure, co-optimization is performed on the 5' UTR and CDS aiming to maximize the first score of the mRNA sequence. The first score can reflect at least one of the translation initiation efficiency, codon adaptation index, and minimum free energy, thereby achieving targeted optimization of the translation efficiency and stability of the mRNA sequence according to design requirements, further optimizing the final yield of the target protein, and improving the overall effect of mRNA vaccines and treatment methods.
[0032] In the mRNA sequence, the 5' untranslated region sequence, also called the 5'UTR sequence, is located at the 5' end of the mRNA molecule, that is, it starts after the 5' cap structure and reaches before the coding region. This fragment has a translation regulatory function, that is, the 5'UTR contains regulatory elements such as upstream Open Reading Fragments (uORFs), sub-optimal binding sites (such as GC-rich regions) and regulatory sequences, and this fragment can affect the stability and translation efficiency of mRNA. In addition, this fragment has the function of guaranteeing the mRNA molecule stability, and some sequence elements of this fragment help to protect mRNA from degradation. This fragment can promote ribosome binding, recognize ribosomes, and bind to a specific sequence (such as Kozak sequence) to initiate the translation process. And in mRNA processing, the signal sequence in the 5'UTR has a decisive effect on mRNA splicing and maturation.
[0033] In the mRNA sequence, the coding region sequence, also called the CDS sequence, is located between the 5' untranslated region and the 3' untranslated region of the mRNA molecule. The coding region sequence contains an Open Reading Fragment (ORF), which consists of a series of codons, and each codon corresponds to a specific amino acid, and the sequence of this segment is translated into a protein in the ribosome. The coding region contains all the genetic information necessary for protein synthesis, and generally starts with a start codon (such as AUG) and ends with a stop codon (such as UAA, UAG or UGA).
[0034] The 5' untranslated region and the coding region are extremely important for protein synthesis. The 5' untranslated region is involved in the regulation of mRNA stability and translation efficiency, while the coding region directly determines the amino acid sequence of the protein. Through the co-optimization of the 5' untranslated region and the coding region, the translation efficiency and stability of mRNA can be improved as a whole, and the yield of the final target protein can be further improved.
[0035] In an embodiment of the present disclosure, in step S102, aiming to maximize the first score of the first mRNA sequence, co-adjustment is performed on the 5' untranslated region sequence and the coding region sequence to obtain an optimized second mRNA sequence. The first score can reflect at least one of the translation initiation efficiency (TIE), codon adaptation index (CAI), and minimum free energy (MFE) of the first mRNA sequence. That is, the first score is obtained by calculating based on at least one of the three indicators of the translation initiation efficiency, codon adaptation index, and minimum free energy of the first mRNA sequence.
[0036] The translation initiation efficiency (TIE) is used to evaluate the translation efficiency of an mRNA sequence by evaluating the translation initiation efficiency of ribosomes on the mRNA molecule. The larger the value of TIE, the faster the start of the translation process, indicating a higher translation efficiency of the mRNA sequence. When the first score is obtained by calculating based on TIE, step S102 realizes a directed optimization of the TIE of the mRNA sequence, thereby improving the overall efficiency of protein synthesis and ensuring that the mRNA can produce a gentle and rapid response when entering the cell.
[0037] The codon adaptation index (CAI) is used to evaluate the degree of consistency between the codons in the mRNA sequence and the most frequently used codons in the host cell. The larger the value of CAI, the closer the codons used in the mRNA sequence are to the codons of highly expressed genes in the host cell, thereby having a greater possibility of obtaining higher translation efficiency. When the first score is obtained by calculating based on CAI, step S102 realizes a directed optimization of the CAI of the mRNA sequence and ensures that the mRNA uses the codons preferred by the host translation mechanism, thereby improving the speed and accuracy of protein synthesis.
[0038] The minimum free energy (MFE) is used to evaluate the structural stability of an mRNA molecule by assessing the energy state when the mRNA molecule forms a secondary structure. The smaller the MFE value, the more stable the mRNA structure. A stable mRNA structure improves intracellular stability and half-life by helping to protect the mRNA from degradation. When the first score is obtained by calculating based on the MFE, step S102 can improve the stability of the mRNA, protect the mRNA from degradation, and improve its survival period in the cellular environment by achieving targeted optimization of the MFE of the mRNA sequence. However, a structure that is too stable may inhibit ribosome binding and translation initiation. Therefore, the MFE needs to be balanced with TIE and CAI to achieve the optimal performance of the mRNA.
[0039] As can be understood, since TIE and CAI are positively correlated with the translation efficiency of mRNA and MFE is negatively correlated with the stability of mRNA, in order to optimize the translation efficiency and stability of the mRNA sequence, the first score may be set to be positively correlated with TIE and CAI and negatively correlated with MFE.
[0040] In some embodiments, the first score S may be calculated according to the following formula (1).
[0041] S = λ TIE * TIE + λ CAI * CAI - λ MFE * MFE (1)
[0042] Here, λ TIE 、λ CAI 、λ MFE are the weights of the TIE, CAI, and MFE indicators respectively. The values of λ TIE 、λ CAI 、λ MFE may be set according to the design requirements of the mRNA, thereby achieving the balance and flexible adjustment of the three indicators of TIE, CAI, and MFE, and enabling the generated mRNA sequence to have the required characteristics.
[0043] In some embodiments, by setting the weight of one of the indicators in formula (1) to a fixed value (e.g., 1) and adjusting the weights of the other two indicators, the balance among the three indicators can be realized. For example, by setting the weight of the MFE indicator to 1 and adjusting the weights of TIE and CAI, the balance among TIE, CAI, and MFE can be realized. In this embodiment, formula (1) is simplified to the following formula (2).
[0044] S = λ TIE * TIE + λ CAI * CAI - MFE (2)
[0045] In some embodiments, the first score S may be calculated according to the following formula (3).
[0046] S = λ TIE * L * log(TIE) + λ CAI * L * log(CAI) - MFE (3)
[0047] In the above formula, L is the number of codons included in the coding region array. By introducing L into the TIE term and the CAI term, the values of the TIE term, the CAI term, and the MFE term in formula (3) can be made to have an order similar to each other, thereby facilitating the realization of the balance and flexible adjustment of the three indicators of TIE, CAI, and MFE. By performing logarithmic transformation on TIE and CAI (denoted as log(TIE) and log(CAI) respectively), the multiplication between the internal factors when calculating TIE and CAI can be converted into addition, thereby simplifying the calculation.
[0048] It should be noted that in addition to the 5' untranslated region and the coding region, mRNA further includes other constituent fragments, such as a 5' cap structure, a 3' untranslated region, and a poly(A) tail. The examples of the present disclosure perform co-optimization on the 5' untranslated region and the coding region. Although other fragments in the mRNA are not optimized (preset fragments may be directly adopted), they may be involved in the calculation of the first score. For example, the 3' untranslated region may be related to the value of TIE (for example, the structural characteristics of the 3' untranslated region are considered when calculating TIE), thus affecting the first score S.
[0049] By using three indicators, namely TIE, CAI, and MFE, to perform co-optimization on the 5' untranslated region and the coding region, the optimized second mRNA sequence can balance three important aspects: translation initiation efficiency, translation elongation efficiency (corresponding to CAI), and stability, thereby optimizing the final protein yield.
[0050] The TIE of the mRNA sequence may be obtained, for example, by calculating using the translation initiation efficiency prediction model described below. The CAI of the mRNA sequence may be obtained, for example, by comparing the codon usage of the mRNA sequence with the codon usage of preset highly expressed genes. The MFE of the mRNA sequence may be obtained, for example, by calculating using algorithms such as the thermodynamic perturbation method and the thermodynamic calculus method.
[0051] In some embodiments, for step S101, each constituent fragment of the first mRNA sequence can be obtained respectively, and further each constituent fragment can be spliced to obtain the first mRNA. Specifically, the 5' untranslated region sequence of the first mRNA sequence may be obtained by the following process 200, the coding region sequence of the first mRNA sequence may be obtained by the following process 300, and other constituent fragments in the first mRNA sequence, such as the 3' untranslated region sequence, etc., may adopt preset values.
[0052] Figure 2 shows a flowchart of process 200 for obtaining the 5' untranslated region sequence of the first mRNA sequence for synthesizing a target protein according to an embodiment of the present disclosure. Process 200 may be used to implement step S101 in the above method 100. In some embodiments, as shown in Figure 2, process 200 may include the following. Step S201, obtain a preset untranslated region sequence library, where the untranslated region sequence library includes at least one candidate 5' untranslated region sequence, and each candidate 5' untranslated region sequence among the at least one candidate 5' untranslated region sequences can achieve gene expression. Step S202, determine the 5' untranslated region sequence included in the first mRNA sequence from at least one candidate 5' untranslated region sequence.
[0053] According to the above embodiment, by selecting a known 5' untranslated region sequence that can achieve gene expression as the initial value of the 5' untranslated region sequence in the mRNA sequence, the quality of the 5' untranslated region sequence can be guaranteed, and a better sample can be provided for subsequent further optimization.
[0054] In some embodiments, in step S201, in order to ensure that the 5' untranslated region sequence in the first mRNA sequence is a sequence that can be normally expressed, based on known mRNA databases, such as databases like UTRdb, NCBI (National Center for Biotechnology Information, USA), UTRsite, EMBL (European Molecular Biology Laboratory Database), ENSEMBL, etc., an untranslated region sequence library can be constructed. The candidate 5' untranslated region sequences in the untranslated region sequence library may be natural sequences in the above mRNA databases, or sequences obtained by artificial optimization. By constructing the untranslated region sequence library, the selection range of the 5' untranslated region sequence can be expanded, and a better sample can be provided for subsequent optimization.
[0055] In some embodiments, in step S202, one 5' untranslated region sequence is selected from the untranslated region sequence library constructed in S201 as the 5' untranslated region sequence in the first mRNA sequence. By selecting the 5' untranslated region sequence in the untranslated region sequence library, it can be ensured that the selected 5' untranslated region sequence has normal expression ability and does not adversely affect subsequent optimization.
[0056] FIG. 3 shows a flowchart of a process 300 for obtaining a coding region sequence of a first mRNA sequence for synthesizing a target protein according to an embodiment of the present disclosure. The process 300 may be used to implement step S101 in the above method 100. In some embodiments, as shown in FIG. 3, the process 300 may include the following. Step S301, generating an initial coding region sequence corresponding to the amino acid sequence of the target protein, and step S302, aiming to maximize the second score of the initial coding region sequence, adjusting the initial coding region sequence to obtain a coding region sequence, where the second score reflects the codon adaptation index and / or the minimum free energy of the initial coding region sequence.
[0057] According to the above embodiment, aiming to maximize the second score, adjusting the initial coding region sequence, and enabling the obtained coding region sequence to be translated into the target protein, a better sample is provided for subsequent further optimization by realizing the balance between translation efficiency and stability. In the embodiments of the present disclosure, the target protein may be any one of predetermined proteins. Since the target protein is determined, its amino acid sequence can be obtained.
[0058] In some embodiments, in step S301, based on known information, or by general technical means including but not limited to techniques such as gene cloning and sequencing, transcriptome sequencing, protein sequencing, computational prediction, yeast two-hybrid system, and protein chip, the amino acid sequence of the target protein can be obtained. According to the correspondence rule between amino acids and codons, the codons corresponding to each amino acid in the target protein can be obtained, and further, by splicing the codons corresponding to each amino acid of the target protein, the initial coding region sequence can be obtained. By the above method, for the first mRNA sequence, by providing an accurate initial coding region sequence that can be translated into the target protein, the expression ability of the second mRNA sequence finally generated by optimization can be guaranteed.
[0059] In some embodiments, in step S302, aiming to maximize the second score of the initial coding region sequence in step S301, by adjusting the initial coding region sequence, the optimized coding region sequence that is a component of the first mRNA sequence is obtained. The second score can reflect the codon adaptation index and / or the minimum free energy of the initial coding region sequence. That is, the second score is obtained by calculating based on the codon adaptation index and / or the minimum free energy of the initial coding region sequence.
[0060] In some embodiments, the second score S’ may be calculated according to the following formula (4).
[0061] S’ = -λ MFE * MFE + λ CAI * CAI (4)
[0062] Here, λ MFE 、λ CAI are the weights of the MFE and CAI indicators respectively. λ MFE 、λ CAIThe value of can be set as needed, thereby achieving a balance between the MFE and CAI metrics and enabling flexible adjustment, such that the generated coding region array has the required characteristics.
[0063] In some embodiments, by setting the weight of one of the metrics in formula (4) to a fixed value (e.g., 1) and adjusting the weight of another metric, a balance between the two metrics can be achieved. For example, by setting the weight of the MFE metric to 1 and adjusting the weight of CAI, a balance between MFE and CAI can be achieved. In this embodiment, formula (4) is simplified to the following formula (5).
[0064] S’ = -MFE + λ CAI * CAI (5)
[0065] In some embodiments, the second score S’ may be calculated according to the following formula (6).
[0066] S’ = -MFE + λ CAI * L * log(CAI) (6)
[0067] In the above formula, L is the number of codons included in the coding region array. By introducing L into the CAI term, the values of the CAI term and the MFE term in formula (6) can be made to have a similar order, thereby facilitating the achievement of a balance and flexible adjustment between the two metrics of MFE and CAI. By performing a logarithmic transformation on CAI (denoted as log(CAI)), the multiplication between internal factors when calculating CAI can be converted to addition, thereby simplifying the calculation.
[0068] FIG. 4 shows a flowchart of a process 400 that aims to maximize the first score of a first mRNA sequence according to an embodiment of the present disclosure, and adjusts the 5' untranslated region sequence and the coding region sequence to obtain an optimized second mRNA sequence for synthesizing a target protein. The process 400 may be used to implement step S102 in the method 100 above. In some embodiments, as shown in FIG. 4, the process 400 may include the following. Step S401, perform mutations on the 5' untranslated region sequence and the coding region sequence of the first mRNA sequence to obtain at least one third mRNA sequence; step S402, calculate the first score of each of the at least one third mRNA sequence; step S403, determine the third mRNA sequence with the largest first score as the second mRNA sequence.
[0069] According to the above embodiment, by simultaneously performing mutation adjustment on the 5' untranslated region sequence and the coding region sequence of the first mRNA sequence, an effective and stable second mRNA sequence can be obtained more quickly.
[0070] In some embodiments, in step S401, mutations are simultaneously performed on the 5' untranslated region sequence and the coding region sequence of the first mRNA sequence to obtain at least one third mRNA sequence. The third mRNA sequence is obtained by randomly changing the nucleotides of the 5' untranslated region sequence and the coding region sequence in the first mRNA. Specifically, the 5' untranslated region sequence and the coding region sequence are taken as a whole, and one or more mutations can be performed on them (that is, replace the nucleotides at a randomly selected position with other nucleotides), and one third mRNA sequence can be obtained after each mutation. By performing mutations on the 5' untranslated region sequence and the coding region sequence, new sequences can be explored and obtained, providing more sequence samples for subsequent screening and optimization.
[0071] In some embodiments, in step S402, for each third mRNA sequence, the calculation method of the first score as described above, for example, formulas (1) to (3), is applied to calculate the first score of the third mRNA sequence.
[0072] In some embodiments, in step S403, based on the first score, screening is performed on at least one third mRNA sequence obtained by mutation, and the third mRNA sequence with the highest first score is determined as the optimized second mRNA sequence. The second mRNA sequence has the highest first score, which means that this sequence has the best comprehensive performance among many third mRNA sequences and can achieve a balance of translation initiation efficiency, translation elongation efficiency (corresponding to CAI), and stability.
[0073] In some embodiments, steps S401 to S403 may be repeatedly executed multiple times. The second mRNA sequence obtained by the current iteration (i.e., the optimization result of the current iteration) can be used as the first mRNA sequence of the next iteration (i.e., the optimization starting point of the next iteration), thereby realizing the iterative optimization of the mRNA sequence and obtaining the optimal second mRNA sequence. The repetition of steps S401 to S403 continues until a preset end condition is met. The end condition may be, for example, that the number of repetitions reaches a predetermined number-of-repetitions threshold, the first score of the second mRNA sequence reaches a predetermined first score threshold, the first score of the second mRNA sequence no longer increases significantly (i.e., the first score converges), etc. After the repetition ends, the second mRNA sequence obtained by the last repetition is used as the final mRNA optimization result.
[0074] The above process 400 may be understood as an evolutionary algorithm. In this algorithm, the first score is the fitness value of each third mRNA sequence obtained by mutation.
[0075] FIG. 5 shows a flowchart of another process 500 aimed at maximizing the first score of the first mRNA sequence according to an embodiment of the present disclosure, adjusting the 5' untranslated region sequence and the coding region sequence to obtain an optimized second mRNA sequence for synthesizing a target protein. The process 500 may be used to implement step S102 in the method 100 above. In some embodiments, as shown in FIG. 5, the process 500 may include the following. Step S501, splitting the 5' untranslated region sequence and the coding region sequence of the first mRNA sequence into a translation initiation region sequence and a coding region body sequence, where the translation initiation region sequence includes at least the 5' untranslated region sequence, and the coding region body sequence includes nucleotides in the coding region sequence that are not included in the translation initiation region sequence. Step S502, aiming to maximize the translation initiation efficiency of the first mRNA sequence, adjusting the translation initiation region sequence to obtain a fourth mRNA sequence. Step S503, aiming to maximize the first score of the fourth mRNA sequence, adjusting the coding region body sequence of the fourth mRNA sequence to obtain a second mRNA sequence.
[0076] According to the above embodiment, based on the influence of each constituent fragment of mRNA on the protein translation process, the 5' untranslated region sequence and the coding region sequence are split into two parts, namely the translation initiation region sequence and the coding region body sequence, and these two parts are sequentially optimized according to different optimization goals, thereby realizing more refined and targeted optimization of the translation efficiency and stability of mRNA.
[0077] In some embodiments, in step S501, adjustments are made to the 5' untranslated region sequence and the coding region sequence of the first mRNA, and the 5' untranslated region sequence of the first mRNA and a preset number of nucleotides close to the 5' untranslated region sequence in the coding region sequence constitute the translation initiation region sequence. In some embodiments, the preset number is preferably 30. In the translation process, since it is necessary to have codons on the ribosome for translation, the ribosome occupies a length of about 30 nucleotides on the mRNA. The residence time of the ribosome in the leader region of the coding region may affect the translation efficiency of the mRNA by affecting the subsequent ribosome assembly and translation initiation. Therefore, when setting the translation initiation region sequence, it is necessary to consider the problem of the position occupied by the ribosome on the mRNA and divide the first 30 nucleotides of the coding region sequence into the translation initiation region sequence.
[0078] As described above, the 5' untranslated region sequence and a preset number (for example, 30) of nucleotides close to the 5' untranslated region sequence of the coding region sequence affect ribosome assembly and translation initiation. The two constitute the translation initiation region sequence, and by performing overall optimization and directionally improving the translation initiation efficiency, the translation efficiency of the mRNA sequence can be improved.
[0079] In some embodiments, after obtaining the translation start region array, certain preprocessing can be performed on the translation start region array, and optimization can be performed on the preprocessed translation start region array. The preprocessing operation of the translation start region array may include, for example, recognizing the -3 position at the 5'UTR end, ensuring that this position is a purine (A or G), and making it conform to the Kozak sequence characteristics, which may help improve the efficiency of translation initiation. The preprocessing operation of the translation start region array may further include, for example, analyzing the 5'UTR region and recognizing all possible upstream start codons (uAUGs). For each recognized uAUG, by replacing any one nucleotide in the AUG with another type of nucleotide, the translation start site is blocked, and at the same time, the translation efficiency is improved, the accuracy of the start site of the translation process is ensured, and the generation of the target protein due to translation misalignment is avoided.
[0080] In some embodiments, in step S502, aiming to maximize the translation initiation efficiency of the first mRNA sequence, an adjustment is made to the translation start region array in the first mRNA sequence to obtain a fourth mRNA sequence. As can be understood, although there are differences between the translation start region arrays of the fourth mRNA sequence and the first mRNA sequence, the coding region body sequences of both are the same. According to this embodiment, the fourth mRNA sequence can provide a good basis for subsequent optimization of the coding region body sequence by having the maximized translation initiation efficiency.
[0081] FIG. 6 shows a flowchart of a process 600 that aims to maximize the translation initiation efficiency of the first mRNA sequence according to an embodiment of the present disclosure, and adjusts the translation initiation region sequence to obtain a fourth mRNA sequence. The process 600 may be used to implement the above step S502. In some embodiments, as shown in FIG. 6, the process 600 may include the following. Step S601, performing a mutation on the translation initiation region sequence of the first mRNA sequence to obtain at least one fifth mRNA sequence; step S602, calculating the translation initiation efficiency of each of the at least one fifth mRNA sequence; step S603, determining the fifth mRNA sequence with the highest translation initiation efficiency as the fourth mRNA sequence.
[0082] According to the above embodiment, by effectively increasing the translation initiation efficiency of the fourth mRNA, the translation efficiency of the finally generated second mRNA can be guaranteed.
[0083] In some embodiments, in step S601, by performing multiple mutations on the translation initiation region sequence of the first mRNA sequence, the number of translation initiation region sequences can be enriched, and multiple fifth mRNA sequences can be obtained. By the above method, the sample amount to be optimized can be increased, and a rich sample basis can be provided for subsequent screening of the fifth mRNA sequence with the highest translation initiation efficiency.
[0084] In some embodiments, in step S602, the translation initiation efficiency of each fifth mRNA sequence is calculated.
[0085] FIG. 7 shows a flowchart of a process 700 for calculating the translation initiation efficiency of each of at least one fifth mRNA sequence according to an embodiment of the present disclosure. The process 700 may be used to implement step S602 in the above method 600. In some embodiments, as shown in FIG. 7, the process 700 may include performing the following steps for each fifth mRNA sequence among the at least one fifth mRNA sequence. Step S701, extracting features for predicting the translation initiation efficiency of the fifth mRNA sequence, and step S702, inputting the features into a trained translation initiation efficiency prediction model to obtain the translation initiation efficiency of the fifth mRNA sequence output from the translation initiation efficiency prediction model.
[0086] The translation initiation efficiency may be affected by various factors. According to the above embodiment, by extracting the features of the fifth mRNA sequence and analyzing the features using a trained translation initiation efficiency prediction model to obtain the translation initiation efficiency of the fifth mRNA sequence, the accuracy and generalization of the translation initiation efficiency evaluation can be improved.
[0087] In some embodiments, in step S701, one or more features of the fifth mRNA sequence are extracted and used as the input of the translation initiation efficiency prediction model. In some embodiments, the features for predicting the translation initiation efficiency of the fifth mRNA sequence include at least one of the structural compactness of the translation initiation region (TIR_ddG_pNT), the overall structural compactness (whole_MFE_pNT), the Kozak sequence feature (purime_m3), the upstream start codon (uAUG) and the upstream open reading frame (uORF) sequence feature, and the ribosome residence time in the CDS leader region (CDS_leader_DT).
[0088] According to the above embodiment, by flexibly selecting the feature combination for predicting the translation initiation efficiency, the sequence characteristics and structural features of the translation initiation region of the fifth mRNA sequence can be obtained flexibly and comprehensively, thereby predicting its translation initiation efficiency more accurately.
[0089] The structural compactness of the translation initiation region (TIR_ddG_pNT) represents the free energy change before and after the unfolding of the secondary structure of the translation initiation region (including the 5’UTR and the 5’ leader sequence of the CDS). A low free energy change indicates that the structure is more compact and is generally associated with a low TIE.
[0090] The characteristic of whole structure compactness (whole_MFE_pNT) measures the minimum free energy (MFE) of the entire mRNA sequence (including the 5’UTR, CDS, and 3’UTR) and normalizes it by the sequence length. A high normalized MFE indicates that the overall structure is not very stable, which generally shows a positive correlation with TIE.
[0091] Kozak sequence feature (purime_m3): The presence of a purine (A / G) at the -3 position of the 5’UTR is one mark of the Kozak sequence and can enhance translation initiation. This feature shows a positive correlation with TIE.
[0092] The uAUG and uORF sequence features include the following.
[0093] Upstream open reading frames within the framework (in_frame_uORF): The presence of upstream open reading frames (uORFs) within the same frame as the main open reading frame (ORF) suppresses downstream translation and has a negative impact on TIE.
[0094] Start codons outside the reading frame (out_frame_uAUG): Start codons located upstream of the main open reading frame (ORF) and outside the reading frame show a negative correlation with TIE.
[0095] The feature of the ribosome dwell time in the CDS leader region (CDS_leader_Dwell Time) measures the dwell time of ribosomes in the 5' leader region of the CDS. Since ribosomes occupy approximately 30 nucleotides on the mRNA and a long dwell time may interfere with subsequent ribosome assembly and translation initiation, it shows a negative correlation with TIE.
[0096] By evaluating the above features, the sequence characteristics and structural features of the translation initiation region of the fifth mRNA can be obtained comprehensively and accurately, thereby calculating the translation initiation efficiency more accurately.
[0097] In some embodiments, in step S702, by inputting the features obtained in step S701 into a trained translation initiation efficiency prediction model, the translation initiation efficiency of the fifth mRNA sequence output from the translation initiation efficiency prediction model can be obtained.
[0098] The translation initiation efficiency prediction model may be any machine learning model, including but not limited to regression models, decision tree models, random forest models, neural network models, etc. The translation initiation efficiency prediction model may also be obtained by training with sequence features marked with translation initiation efficiency tags as samples.
[0099] In some exemplary embodiments, a ridge regression model may be employed as a translation start efficiency prediction model. The ridge regression model may be used to predict the logarithmically transformed TIE (i.e., log(TIE)). The model can handle the multicollinearity among features and prevent overfitting through regularization. The training data of the model may be selected from the multimer analysis data of eGFP and the ribosome analysis data of the human genome. The data may be selected, for example, from the National Genomics Data Center. These datasets provide a comprehensive insight into mRNA translation trends. At the same time, to ensure consistency and improve model performance, scaling is performed on each of the above-mentioned input features, and the above features are used as predictor variables to construct a ridge regression model. Ridge regression introduces a penalty term proportional to the square of the magnitude of the coefficients, which avoids over-reliance on any single feature. The model is trained on the collected dataset and uses standard metrics including the mean squared error (MSE) and R 2 score as the loss posterior function of the model to evaluate the performance of the model. Cross-validation is used to evaluate the robustness of the model and fine-tune the hyperparameters of the model.
[0100] After obtaining the translation start efficiency of each fifth mRNA sequence in step S602, step S603 may be executed. In step S603, the fifth mRNA sequence with the highest translation start efficiency is selected and used as the fourth mRNA sequence.
[0101] In some embodiments, steps S601 to S603 may be repeatedly executed multiple times. The fourth mRNA sequence obtained by the current iteration (i.e., the optimization result of the current iteration) can be used as the first mRNA sequence of the next iteration (i.e., the optimization starting point of the next iteration), thereby realizing the iterative optimization of the mRNA sequence and obtaining the optimal fourth mRNA sequence. The repetition of steps S601 to S603 continues until a preset end condition is met. The end condition may be, for example, that the number of repetitions reaches a predetermined number-of-repetitions threshold, the translation initiation efficiency of the fourth mRNA sequence reaches a predetermined translation-initiation efficiency threshold, or the translation initiation efficiency of the fourth mRNA sequence no longer increases significantly (i.e., the translation initiation efficiency converges). After the repetition ends, the fourth mRNA sequence obtained by the last repetition is used as the optimization result for the translation initiation region sequence.
[0102] In some exemplary embodiments, process 600 may be understood as an evolutionary algorithm. The specific operations of the algorithm are as follows.
[0103] Design of the initial population: By mutating the translation initiation region sequence of the first mRNA sequence, an initial population consisting of multiple mRNA sequences is constructed.
[0104] Definition of the fitness function: The translation initiation efficiency (TIE) is used as the fitness function to evaluate the performance of each sequence variant.
[0105] Iterative optimization process: By mimicking the process of natural selection and applying mutation and selection operations, the sequence population is iteratively optimized.
[0106] Mutation: Randomly change the nucleotides in the sequence to explore a new sequence space and obtain multiple fifth mRNA sequences.
[0107] Selection: Based on the TIE evaluation results of each fifth mRNA sequence, the sequence with the largest TIE is selected as the current optimal fourth mRNA sequence for the next generation of iteration.
[0108] End condition: When the specified number of iterations is reached or the array performance no longer increases significantly, stop the iteration.
[0109] In some embodiments, in step S503, on the fourth mRNA sequence for which the optimization of the translation start region obtained in step S502 is completed, aiming to maximize the first score of the fourth mRNA sequence, adjustments are made to the coding region body sequence of the fourth mRNA sequence to obtain an optimized second mRNA sequence.
[0110] FIG. 8 shows a flowchart of a process 800 for obtaining a second mRNA sequence by making adjustments to the coding region body sequence of the fourth mRNA sequence with the goal of maximizing the first score of the fourth mRNA sequence according to an embodiment of the present disclosure. The process 800 may be used to implement step S503 in the above method 500. In some embodiments, as shown in FIG. 8, the process 800 may include the following. Step S801, performing a mutation on the coding region body sequence of the fourth mRNA sequence to obtain at least one sixth mRNA sequence; step S802, calculating the first score of each of the at least one sixth mRNA sequence; step S803, determining the sixth mRNA sequence with the largest first score as the second mRNA sequence.
[0111] According to the above embodiments, the first score can reflect at least one of the indicators of translation initiation efficiency, codon adaptation index, and minimum free energy, thereby enabling targeted optimization of the translation efficiency and stability of the mRNA sequence according to design requirements. In particular, it is optimized and enhanced in the problem of forming base pairings in the translation start region, further optimizing the yield of the final target protein and improving the overall effect of the mRNA vaccine and treatment method.
[0112] In some embodiments, in step S801, by performing multiple mutations on the coding region body sequence of the fourth mRNA sequence, the number of coding region body sequences can be enriched, and multiple sixth mRNA sequences can be obtained. By the above method, the amount of optimization samples can be increased, and a sample basis can be provided for screening the sixth mRNA having the highest first score subsequently.
[0113] In some embodiments, in step S802, the first score of each sixth mRNA is calculated. In this process, by applying the first score calculation formulas (1)-(3) as described above, the first score corresponding to each sixth mRNA can be calculated.
[0114] In some embodiments, in step S803, from the sixth mRNA sequences that have completed scoring in step S802, the sixth mRNA sequence with the highest score is selected as the second mRNA sequence after optimization. The second mRNA sequence has the highest first score, which means that the sequence has the best comprehensive performance among many sixth mRNA sequences and can achieve a balance among three important aspects: translation initiation efficiency, translation elongation efficiency (corresponding to CAI), and stability.
[0115] In some embodiments, in step S803, in response to the codon adaptation index of the sixth mRNA sequence with the largest first score being greater than the threshold, the sixth mRNA sequence is determined as the second mRNA sequence after optimization, where the threshold is determined based on the codon adaptation index of the initial first mRNA sequence (i.e., the first mRNA sequence obtained by step S101). For example, the threshold may be set to the codon adaptation index of the initial first mRNA sequence. According to this embodiment, it can be ensured that the second mRNA sequence after optimization has a translation expression ability not lower than that of the initial first mRNA sequence.
[0116] In some embodiments, in response to the codon adaptation index of the sixth mRNA sequence with the largest first score being below the threshold, the current optimization result may be discarded, that is, steps S801 to S803 may be executed again without using the sixth mRNA sequence as the second mRNA sequence after optimization, and continue until an optimization result with a codon adaptation index greater than the threshold is obtained. In some embodiments, steps S801 to S803 may be repeatedly executed multiple times. The second mRNA sequence obtained by the current iteration (i.e., the optimization result of the current iteration) can be used as the fourth mRNA sequence of the next iteration (i.e., the optimization starting point of the next iteration), thereby realizing iterative optimization of the mRNA sequence and obtaining the optimal second mRNA sequence. The repetition of steps S801 to S803 continues until a preset end condition is met. The end condition may be, for example, that the number of repetitions reaches a predetermined number of repetition thresholds, the first score of the second mRNA sequence reaches a predetermined first score threshold, the first score of the second mRNA sequence no longer increases significantly (i.e., the first score converges), etc. After the repetition ends, the second mRNA sequence obtained by the last repetition is used as the final mRNA optimization result. In some exemplary embodiments, steps S801 to S803 may be considered to be performed by applying an evolutionary algorithm, and the operation is specifically as follows.
[0117] Design of the initial population: Mutate the coding region body sequence of the fourth mRNA sequence to construct an initial population consisting of multiple mRNAs.
[0118] Definition of the fitness function: Use the first score as the fitness function to evaluate the performance of each sequence variant.
[0119] Iterative optimization process: Imitate the process of natural selection and apply mutation and selection operations to iteratively optimize the sequence population.
[0120] Mutation: In the coding region main body sequence, perform mutations at positions corresponding to the translation start region to search for array variants that may improve translation start efficiency and / or reduce the minimum free energy, and obtain a plurality of sixth mRNA sequences.
[0121] Selection: Based on the evaluation results of the first score of each sixth mRNA sequence, select the sequence with the largest first score as the current optimal second mRNA sequence, and perform the next generation of iteration.
[0122] End condition: When the predetermined number of iterations is reached or the first score of the second mRNA sequence no longer increases significantly, stop the iteration.
[0123] According to an embodiment of the present disclosure, there is further provided an apparatus for designing a messenger ribonucleotide (mRNA) sequence. FIG. 9 shows a structural block diagram of a training apparatus for a neural network model for predicting the effect of mutations on protein stability according to an exemplary embodiment of the present disclosure. As shown in FIG. 9, the apparatus 900 includes an acquisition unit 910 configured to acquire a first mRNA sequence for synthesizing a target protein, and a processing unit 920 configured to maximize the first score of the first mRNA sequence and perform adjustments on the 5' untranslated region sequence and the coding region sequence to obtain an optimized second mRNA sequence for synthesizing the target protein.
[0124] As can be understood, the operations of unit 910 to unit 920 in apparatus 900 may refer to the descriptions of steps S101 to S102 in method 100 above, and the descriptions are omitted here.
[0125] In an exemplary embodiment, to analyze the guiding value of the formula for the first score in method 100, the differences in the actual protein yields of different region samples in the distance space were examined. Here, the mRNA sequences and expression data are from the article Kathrin Leppek, Gun Woo Byeon, Wipapat Kladwang, Hannah K Wayment-Steele, Craig H Kerr, Adele F Xu, Do Soon Kim, Ved V Topkar, Christian Choe, Daphna Rothschild, et al. Combinatorial optimization of mRNA structure, stability, and translation for RNA-based therapeutics. Nature communications, 13(1):1536, 2022, in which the expression levels of Nluc reporter genes with different CDS sequences were measured by the activity ratio of the Nluc / Fluc reporter genes. As shown in subfigures A and B in Figure 10, subfigure A shows the expression levels of each sample after 6 hours, and subfigure B shows the expression levels of each sample after 24 hours, where the darker the color of the points, the higher the expression level. As shown in the distributions in subfigures A and B, the samples with the highest protein expression are mainly located in the regions with low MFE (< -350 kcal / mol), high CAI (> 0.75), and moderate TIE (0.35 - 0.42). As shown in subfigures C and D, the figure is a two-dimensional distribution diagram of subfigures A and B respectively, and such a pattern was also observed in subfigures C and D. Samples predicted to have too low TIE showed clear disadvantages in the expression levels at both time points, indicating that sufficient translation efficiency is necessary for effective protein yields. There are also differences in the expression level distribution patterns at 6 hours and 24 hours. Specifically, samples with the highest expression levels at 6 hours tend to have higher TIE, while samples with the highest expression levels at 24 hours tend to have lower MFE.This is in line with the general principle that the short-term expression level is more affected by translation efficiency, and the long-term expression level is more affected by mRNA stability.
[0126] However, in this distribution, the samples with the highest TIE do not show high protein expression levels. This is presumably because these samples also have relatively high MFE values, and the reduced stability affects their sustained expression ability. The drawback between TIE and MFE can be understood as an optimization goal, and the reduction of MFE is due to increasing the compactness of the mRNA structure, which brings obstacles to the binding of ribosomes and other translation factors to mRNA. As shown in Subfigures E and F, Subfigure E is a scatter plot showing the correlation between TIE and Nluc / Fluc activity by selecting samples according to the criteria of MFE < -350 kcal / mol and CAI > 0.75 within 24 hours, and Subfigure F is a scatter plot showing the correlation between TIE and the abundance of YFP expressed in yeast within 24 hours. In Subfigures E and F, when screening samples with too high MFE (> 350 kcal / mol) and too low CAI (< 0.75), a more obvious positive correlation between TIE and protein expression level was observed. Specifically, in the 24-hour expression data, the Spearman correlation between TIE and the Nluc / Fluc activity ratio reached 0.70 (p < 0.05). Therefore, by ensuring the relative optimal values of MFE and CAI and optimizing TIE, it is possible to achieve better protein yields.
[0127] In an exemplary embodiment, two custom parameters λ TIE and λ CAI are introduced into method 100 to balance the relative weights of the three optimization goals, namely TIE, CAI, and MFE. The parameters λ CAI and λ TIE control the weights of CAI and TIE in the optimization process respectively. To enhance the convenience of index adjustment, the algorithm is optimized so that the CAI index of the target sequence is λ TIEEnsure that it is not affected by the value. This means that when CAI is fixed, no matter how TIE changes, the CAI value of the designed array is maintained within a relatively stable range.
[0128] Designing the mRNA sequence of eGFP protein (derived from GenBank: AFA52650.1) demonstrates the ability to regulate method 100 on the mRNA index. Therefore, five TIE parameters (2, 4, 6, 8, 10) and four CAI parameters (2, 4, 6, 8) are set, resulting in a total of 20 parameter combinations. The CAI parameter precisely regulates the CAI value of the target sequence, as shown in Subfigure A of Figure 11. As the CAI value increases, the CAI value of the designed array gradually rises. For each CAI , the variation in the CAI values of the sequences designed with different TIE parameters is very small. The CAI is mainly regulated by the CAI parameter and is hardly affected by TIE , consistent with the design prediction of method 100. When CAI is fixed, method 100 can achieve flexible regulation of the target sequence TIE and MFE index by changing TIE , as shown in Subfigures B and C of Figure 11. As the TIE parameter increases, the TIE value gradually rises, indicating an increase in mRNA translation initiation efficiency, and the MFE value also increases, suggesting a reduction in mRNA structural compactness and thermodynamic stability. As shown in Subfigure D of Figure 11, there is a positive correlation between the TIE and MFE values, indicating a negative correlation between translation initiation efficiency and mRNA structural compactness and thermodynamic stability. In short, by adjusting these two hyperparameters, the optimization goal can be flexibly customized to achieve different balances among these three indicators and meet the diverse requirements in different scenarios.
[0129] The present disclosure further provides an mRNA molecule, and the sequence of the mRNA molecule is produced by the methods, apparatuses, electronic devices or computer program products disclosed herein.
[0130] In an exemplary embodiment, the design of the novel coronavirus (SARS-CoV-2) spike protein and varicella-zoster virus (VZV) antigen (VZV gE protein, UniProtKB / Swiss-Prot: Q9J3M8.1) by the LinearDesign algorithm and method 100 was evaluated, where the amino acid sequences of the novel coronavirus (SARS-CoV-2) spike protein and varicella-zoster virus (VZV) antigen may both be obtained from NCBI (National Center for Biotechnology Information, USA). The LinearDesign algorithm is a conventional mRNA sequence design algorithm. For the SARS-CoV-2 spike protein, as shown in subfigure A of Figure 12, the MFE value of the mRNA designed by LinearDesign was significantly lower than that of the wild type (WT) and commercial vaccine sequences (BNT-162b2 and mRNA-1273) (both commercial vaccine sequences are from: https: / / github.com / NAalytics / Assemblies-of-putative-SARS-CoV2-spike-encoding-mRNA-sequences-for-vaccines-BNT-162b2-and-mRNA-1273 / ), indicating that it was significantly improved in terms of structural compactness and thermodynamic stability. However, it was shown that the TIE value of the sequence designed by LinearDesign was significantly lower, which may reduce the translation efficiency. In contrast, the sequence designed by method 100 had a significantly improved TIE index compared to the sequence designed by LinearDesign, while maintaining similar CAI and MFE values. As shown in subfigure B of Figure 12, a similar pattern was also observed in the design of the VZV antigen. The sequence designed by LinearDesign had a significant advantage in terms of the MFE index compared to the wild type (gE-WT) and the sequence designed by Thermo Fisher's codon optimization tool (gE-Ther). However, they had obvious drawbacks in terms of the TIE index. Method 100 effectively solved this problem by maintaining the advantage of LinearDesign in terms of MFE and significantly improving the performance of the TIE index.This indicates that Method 100 can produce sequences with overall performance that is more suitable in terms of both translation efficiency and stability. In view of the analysis of the relationship between the previous distance space and protein expression levels, such improvements are expected to lead to an increase in protein yield.
[0131] The sequences designed by Method 100 and LinearDesign also have significant differences in secondary structure. Subfigure C in Figure 12 shows the secondary structures of two eGFP mRNA sequences of these two designs, where the eGFP original amino acid sequence is derived from NCBI. Although the two sequences show similar MFE and CAI indicators, the sequence designed by Method 100 has a significantly better TIE indicator. The main structural differences in the start codon region of the two sequences can be observed. In this region, the sequence designed by Method 100 has fewer hairpin structures and base pairings, resulting in a structurally more relaxed configuration. Such a more relaxed structure is generally considered to be advantageous for ribosome binding and scanning in the 5’UTR region, thereby improving translation initiation efficiency. As is also evident when the sequence of the start codon background region is extracted and folded alone, the sequence designed by Method 100 has less secondary structure and higher folding free energy, as shown in Subfigure D in Figure 12. This further supports the technical effect that a relaxed structure in the start codon background region helps to achieve the improved translation initiation efficiency in Method 100.
[0132] In one exemplary embodiment, a large amount of parallel translation assay (MPTA) data was used to analyze the accuracy of the TIE metric for eGFP protein prediction. Samples in this dataset had a fixed CDS and randomly generated 5'UTR sequences, and ribosome loading of each sequence was measured by multimer analysis. Given that the translation elongation efficiency was relatively steady, the ribosome loading value reflected the translation initiation efficiency of each sequence. As shown in Subfigure A of Figure 13, in the first 20,000 samples sorted by read count, the Spearman correlation coefficient between TIE and ribosome loading was 0.83, which was significantly higher than other subcategorized features. Among the subcategorized features, the absolute Spearman correlation coefficient between the features related to upstream AUG codons (uAUG) and ribosome loading was the highest. The presence of uAUG may lead to premature translation initiation and suppress the translation of the main ORF. In the actual mRNA vaccine development scenario, generally, sequences containing uAUG are not used. Therefore, to get closer to the actual situation, further analysis was performed on the correlation between various metrics and ribosome loading in samples without uAUG. As shown in Subfigure B of Figure 13, in these samples, the Spearman correlation between TIE and ribosome loading reached 0.6, exceeding the correlation of other subcategorized metrics. This demonstrated the robustness of TIE as a translation initiation efficiency metric, especially in the background without the uAUG sequence.
[0133] Since the eGFP protein dataset has a fixed CDS, it cannot effectively reflect the impact of the CDS region on mRNA translation initiation efficiency. To address this issue, we further analyzed using ribosome profiling data (GSE35469) from the human PC3 cell line. This dataset contains translation efficiency information for the entire human genome transcriptome, with significant differences in UTR and CDS sequences among transcripts, making it suitable for analyzing the synergistic effect of the 5’UTR and CDS on translation initiation efficiency. According to the article Nicholas T Ingolia, Liana F Lareau, and Jonathan S Weissman. Ribosome profiling of mouse embryonic stem cells reveals the complexity and dynamics of mammalian proteomes. Cell, 147(4):789~802, 2011, there are no significant differences in translation elongation efficiency among different genes, and translation initiation efficiency is the main rate-limiting step in the translation process. Therefore, in such cases, translation efficiency may be mainly regarded as a representative of translation initiation efficiency.
[0134] As shown in Subfigure C of Figure 13, the Spearman correlation coefficients between translation efficiency and various indicators were shown. The correlation between the predicted TIE indicator and the measured translation efficiency (TE) was 0.574, which was superior to other differentiation characteristics. Among other characteristics, the correlation between the minimum free energy per unit length of the entire mRNA strand (whole_MFE_mean) and TE was the highest, reflecting a significant impact of mRNA structural compactness on translation initiation efficiency. A structure that is too compact may inhibit translation initiation. Since ribosomes staying in the CDS leader region can spatially interfere with the assembly of subsequent ribosomes, the ribosome residence time in the CDS leader region (i.e., the ribosome decoding time in the 5’ start region of the CDS, CDS_leader_DT) showed a significant negative correlation with TE. These correlation patterns further demonstrated that translation initiation efficiency is jointly determined by the 5’UTR and CDS regions.
[0135] The present disclosure provides a pharmaceutical composition, which comprises an mRNA sequence produced by the methods, apparatuses, electronic devices or computer program products disclosed herein or an mRNA molecule disclosed herein and a pharmaceutically acceptable adjuvant.
[0136] The pharmaceutical composition of the present disclosure can be prepared by any means known in the art, including but not limited to preparing it into solid, semi-solid or liquid dosage forms such as tablets, capsules, caplets, suspensions, powders, lyophilized preparations, suppositories, eye drops, transdermal patches, orally soluble preparations, sprays, aerosols, etc.
[0137] The pharmaceutical composition may be an immediate release and / or modified release formulation, including delayed release, sustained release, pulsatile release, controlled release, targeted release and programmed release formulations.
[0138] As used herein, "pharmaceutically acceptable adjuvant" refers to a component that is non-toxic to a subject other than the active ingredient in a pharmaceutical composition. Pharmaceutically acceptable adjuvants include, but are not limited to, excipients (e.g., diluents, carriers, etc.) and additives (e.g., stabilizers, preservatives, solubilizers, buffers, etc.). Excipients may include polyvinylpyrrolidone, gelatin, hydroxypropylcellulose (HPC), gum arabic, polyethylene glycol, mannitol, sodium chloride, and sodium citrate. For injectable formulations or other liquid dosage forms, it is preferably water containing at least one or more buffering components, and stabilizers, preservatives, and solubilizers may be employed. For solid dosage forms, any one of various thickeners, fillers, extenders, and carrier additives, such as starch, sugar, cellulose derivatives, fatty acids, etc., may be employed. For topical dosage forms, any one of various creams, ointments, gels, lotions, etc., may be employed. For most drug formulations, based on weight or volume, the inactive ingredients may occupy a majority of the formulation. The present disclosure also covers the possibility of formulating drug formulations by employing any one of various controlled release, sustained release, or extended release formulations and additives to enable delivery of the compounds of the present disclosure over a certain period of time.
[0139] The compounds of the present disclosure may include administration by means such as mucosal administration, oral administration, transdermal administration, inhalation administration, intranasal administration, urethral administration, vaginal administration, and intravenous, subcutaneous, intramuscular, intraperitoneal injection, etc. The adjuvant in the pharmaceutical composition is adapted to its route of administration.
[0140] In some embodiments, the compounds of the present disclosure can be delivered by oral delivery, for example, by tablets or capsules. The compound may be packaged in an enteric protector, preferably such that the compound is not released before delivery of the tablet or capsule to the stomach and optionally before further delivery to a portion of the small intestine.
[0141] In some embodiments, the compounds of the present disclosure can be administered by injection, and drug forms suitable for injection include sterile aqueous solutions or dispersions, and sterile powders for the immediate preparation of sterile injectable solutions or dispersions. In all cases, this form must be sterile and fluid enough to be administered by syringe. The form must be stable under the conditions of manufacture and storage and must be preserved to prevent the contaminating action of microorganisms such as bacteria and fungi. The carrier may be a solvent or dispersion medium containing, for example, water, ethanol, polyols (such as glycerol, propylene glycol or liquid polyethylene glycol), suitable mixtures thereof and vegetable oils.
[0142] Furthermore, therapeutic administration may be performed by injecting sustained-release formulations, for example, formulations that can be subcutaneously injected and include nanospheres / microspheres, liposomes, emulsions, gels, insoluble salts or suspensions.
[0143] In some embodiments, the compounds of the present disclosure may be administered intranasally. The pharmaceutical composition may be in the form of an aqueous solution, for example, a solution containing saline, citrate or other commonly used excipients or preservatives. It may also be in the form of a dry formulation or powder.
[0144] The present disclosure provides the application of the mRNA sequences produced by the methods, devices, electronic devices or computer program products disclosed herein, the mRNA molecules disclosed herein or the pharmaceutical compositions disclosed herein in the manufacture of drugs or vaccines.
[0145] The method can significantly improve the yield and quality of proteins and has important application value in the manufacturing fields such as drugs or vaccines.
[0146] In some embodiments, the drugs disclosed herein include, but are not limited to, mRNA drugs, protein replacement therapy drugs, gene editing drugs, cancer treatment drugs, regenerative medicine drugs, DNA gene therapy agents of viral or non-viral vectors, modifiers of genetically engineered organisms, cell therapy drugs, enzyme replacement therapy drugs, aptamer drugs, microRNA therapy drugs, and ribozyme-based drugs. In some preferred embodiments, the drugs disclosed herein are selected from mRNA drugs, DNA gene therapy agents of viral or non-viral vectors, or modifiers of genetically engineered organisms.
[0147] In some embodiments, the vaccines disclosed herein are selected from mRNA prophylactic vaccines or mRNA therapeutic vaccines.
[0148] The present disclosure provides a method for treating or preventing a disease, comprising administering to a subject in need thereof an effective amount of an mRNA sequence produced by a method, apparatus, electronic device, or computer program product disclosed herein, an mRNA molecule disclosed herein, or a pharmaceutical composition disclosed herein.
[0149] In some embodiments, the diseases disclosed herein include, but are not limited to, infectious diseases, including viral infections such as novel coronavirus and varicella-zoster virus.
[0150] As described herein, "subject" includes animals, such as vertebrates, preferably mammals such as dogs, cats, pigs, cows, sheep, horses, rodents (such as mice, rats or guinea pigs), or primates (such as gorillas, chimpanzees and humans).
[0151] As described herein, "treatment" refers to reducing or ameliorating a disease or disorder (i.e., slowing or arresting the progression of the disease or at least one clinical symptom), or reducing or ameliorating at least one physiological parameter or biomarker associated with the disease or disorder.
[0152] As described herein, an "effective amount" is an amount sufficient to induce an effect such as a desired treatment, prevention, or inhibition, by any of the means described above or any other means known in the art, and is administered in an amount sufficient to provide a benefit or achieve an effect as compared to a corresponding subject not receiving it. The amount is sufficiently low within the scope of reasonable medical judgment to avoid serious side effects. The effective amount is determined according to circumstances such as the selected drug, e.g., mRNA, pharmaceutical composition, vaccine, route of administration, severity of the disease being treated, age, body type, weight, and physical condition of the patient being treated.
[0153] According to an embodiment of the present disclosure, there is further provided an electronic device, which includes at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the method for optimizing the mRNA sequence of the embodiment of the present disclosure is executed by the at least one processor.
[0154] According to an embodiment of the present disclosure, there is further provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the method for optimizing the mRNA sequence of the embodiment of the present disclosure.
[0155] According to an embodiment of the present disclosure, there is further provided a computer program product including computer program instructions, and when the computer program instructions are executed by a processor, the method for optimizing the mRNA sequence of the embodiment of the present disclosure is realized.
[0156] Referring to FIG. 14, a structural block diagram of an electronic device 1400 that can be used as a server or a client of the present disclosure will be described. It is an example of a hardware device applicable to each aspect of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, tablets, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may further represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown in this specification, their connection relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed in this specification.
[0157] As shown in FIG. 14, the electronic device 1400 includes a computing unit 1401, which can execute various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 1402 or a computer program loaded from a storage unit 1408 into a random access memory (RAM) 1403. Various programs and data necessary for the operation of the electronic device 1400 may be further stored in the RAM 1403. The computing unit 1401, the ROM 1402, and the RAM 1403 are connected to each other by a bus 1404. An input / output (I / O) interface 1405 is also connected to the bus 1404.
[0158] The plurality of members in the electronic device 1400 are connected to the I / O interface 1405 and include an input unit 1406, an output unit 1407, a storage unit 1408, and a communication unit 1409. The input unit 1406 may be any type of device capable of inputting information into the electronic device 1400. The input unit 1406 can receive the input numerical or character information and generate key signal inputs related to user settings and / or function controls of the electronic device, and may include, but is not limited to, a mouse, a keyboard, a touch screen, a track board, a track ball, an operation lever, a microphone, and / or a remote control. The output unit 1407 may be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1408 may include, but is not limited to, a magnetic disk and an optical disk. The communication unit 1409 enables the electronic device 1400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth (registered trademark) device, an 802.11 device, a Wi-Fi device, a WiMAX device, a cellular communication device, and / or the like.
[0159] The computing unit 1401 may be various general-purpose and / or special-purpose processing assemblies having processing and computing capabilities. Some examples of the computing unit 1401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1401 executes each method and process described above, for example, method 100. For example, in some embodiments, method 100 may be implemented as a computer software program tangibly included in a machine-readable medium, such as storage unit 1408. In some embodiments, some or all of the computer program may be loaded and / or installed into the electronic device 1400 via the ROM 1402 and / or the communication unit 1409. When the computer program is loaded into the RAM 1403 and executed by the computing unit 501, one or more steps of the method 100 described above can be executed. Optionally, in other embodiments, the computing unit 1401 may be configured to execute method 100 in any other suitable manner (e.g., by firmware).
[0160] The various embodiments of the systems and techniques described above in this specification may be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include the following. Implemented in one or more computer programs, the one or more computer programs may be executed and / or interpreted in a programmable system including at least one programmable processor, the programmable processor may be a dedicated or general-purpose programmable processor, receive data and instructions from a memory system, at least one input device, and at least one output device, and transmit the data and instructions to the memory system, the at least one input device, and the at least one output device.
[0161] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to the processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations defined in the flowchart and / or block diagram are implemented. The program code may be executed entirely by a machine, partially by a machine, partially by a machine as an independent software package and partially by a remote machine, or entirely by a remote machine or server.
[0162] In the context of this disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be either a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium include, but are not limited to, electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fiber, compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.
[0163] To provide for interaction with a user, the systems and techniques described here may be implemented on a computer, which includes a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or trackball), by which the user may provide input to the computer. Other kinds of devices may be further provided for interacting with the user, for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user may be received in any form (including voice input, voice input, tactile input).
[0164] The systems and techniques described herein may be implemented in a computing system that includes a background member (e.g., a data server), a computing system that includes a middleware member (e.g., an application server), a computing system that includes a front-end member (e.g., a user computer having a graphical user interface or a web browser, and a user can interact with embodiments of the systems and techniques described herein through the graphical user interface or the web browser), or a computing system that includes any combination of such background members, middleware members, or front-end members. The members of the system may be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0165] The computer system may include a client and a server. The client and the server are generally far apart from each other and usually interact via a communication network. A computer program having a client-server relationship with each other generates a client-server relationship by running on the corresponding computer. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0166] It should be understood that the steps may be reordered, added, or deleted again using various forms of flow shown above. For example, each step described in this disclosure may be executed in parallel, sequentially, or in a different order, and is not limited herein as long as the technical solutions disclosed in this disclosure can achieve the desired results.
[0167] While the embodiments or examples of the present disclosure have been described with reference to the drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present disclosure is not limited by these embodiments or examples, but is limited only by the scope of the claims after authorization and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalent elements. In addition, each step may be executed in an order different from the order described in the present disclosure. Furthermore, various elements in the embodiments or examples may be combined in various ways. In short, with the evolution of technology, many of the elements described herein may be replaced by equivalent elements that appear after the present disclosure.
Claims
1. 1. A method for optimizing a messenger ribonucleotide (mRNA) sequence comprising: Obtaining a first mRNA sequence for synthesizing a target protein, the first mRNA sequence comprising a 5' untranslated region sequence and a coding region sequence; A method for optimizing a messenger ribonucleotide (mRNA) sequence, comprising: making adjustments to the 5' untranslated region sequence and the coding region sequence to obtain an optimized second mRNA sequence for synthesizing the target protein, with the goal of maximizing a first score of the first mRNA sequence, wherein the first score reflects at least one of the following indicators of the translation initiation efficiency, the codon compatibility index, and the minimum free energy of the first mRNA sequence.
2. The step of obtaining a first mRNA sequence for synthesizing a target protein comprises: Obtaining a predetermined untranslated region sequence library, the untranslated region sequence library including at least one candidate 5' untranslated region sequence, each of the at least one candidate 5' untranslated region sequence being capable of achieving gene expression; determining said 5' untranslated region sequence from said at least one candidate 5' untranslated region sequence.
3. The step of obtaining a first mRNA sequence for synthesizing a target protein comprises: generating an initial coding region sequence corresponding to the amino acid sequence of the target protein; 2. The method of claim 1, further comprising: making an adjustment to the initial coding region sequence to obtain the coding region sequence with the goal of maximizing a second score of the initial coding region sequence, wherein the second score reflects a codon adaptability index and / or a minimum free energy of the initial coding region sequence.
4. said adjusting said 5' untranslated region sequence and said coding region sequence to obtain an optimized second mRNA sequence for synthesizing said target protein with a goal of maximizing said first score of said first mRNA sequence; mutagenizing said 5' untranslated region sequence and said coding region sequence to obtain at least one third mRNA sequence; calculating a first score for each of the at least one third mRNA sequence; and determining the third mRNA sequence having the highest first score as the second mRNA sequence.
5. said adjusting said 5' untranslated region sequence and said coding region sequence to obtain an optimized second mRNA sequence for synthesizing said target protein with a goal of maximizing said first score of said first mRNA sequence; Dividing the 5' non-translated region sequence and the coding region sequence into a translation initiation region sequence and a coding region main sequence, wherein the translation initiation region sequence includes at least the 5' non-translated region sequence, and the coding region main sequence includes nucleotides in the coding region sequence that are not included in the translation initiation region sequence; making adjustments to the translation initiation region sequence to obtain a fourth mRNA sequence with the goal of maximizing translation initiation efficiency of the first mRNA sequence; The method of claim 1, further comprising: making adjustments to a coding region body sequence of the fourth mRNA sequence to obtain the second mRNA sequence, with the goal of maximizing a first score of the fourth mRNA sequence.
6. 6. The method of claim 5, wherein the translation initiation region sequence comprises the 5' untranslated region sequence and a predetermined number of nucleotides proximal to the 5' untranslated region sequence in the coding region sequence.
7. said adjusting said translation initiation region sequence to obtain a fourth mRNA sequence with the goal of maximizing the translation initiation efficiency of said first mRNA sequence; mutagenizing said translation initiation region sequence to obtain at least one fifth mRNA sequence; calculating a translation initiation efficiency for each of the at least one fifth mRNA sequence; and determining the fifth mRNA sequence having the highest translation initiation efficiency as the fourth mRNA sequence.
8. Calculating the translation initiation efficiency of each of the at least one fifth mRNA sequence comprises: For each fifth mRNA sequence of the at least one fifth mRNA sequence, Extracting features for predicting translation initiation efficiency of the fifth mRNA sequence; 8. The method of claim 7, comprising inputting the features into a trained translation initiation efficiency prediction model to obtain a translation initiation efficiency of the fifth mRNA sequence output from the translation initiation efficiency prediction model.
9. The above features include:
9. The method of claim 8, comprising at least one of the following: structural compactness of the translation initiation region sequence, overall structural compactness, Kozak sequence characteristics, upstream start codon and upstream open reading frame sequence characteristics, and ribosome residence time of the coding region sequence leader region.
10. The step of obtaining the second mRNA sequence by adjusting the coding region main sequence of the fourth mRNA sequence with a goal of maximizing the first score of the fourth mRNA sequence includes: mutagenizing the coding region of the fourth mRNA sequence to obtain at least one sixth mRNA sequence; calculating a first score for each of the at least one sixth mRNA sequence; and determining the sixth mRNA sequence having the highest first score as the second mRNA sequence.
11. determining the sixth mRNA sequence having the highest first score as the second mRNA sequence, 11. The method of claim 10, comprising determining the sixth mRNA sequence having the highest first score as the second mRNA sequence in response to the codon adaptability index of the sixth mRNA sequence being greater than a threshold, wherein the threshold is determined based on the codon adaptability index of the first mRNA sequence.
12. An apparatus for optimizing a messenger ribonucleotide (mRNA) sequence, comprising: A capture unit configured to capture a first mRNA sequence for synthesizing a target protein, the first mRNA sequence comprising a 5' untranslated region sequence and a coding region sequence; and a processing unit configured to adjust the 5' untranslated region sequence and the coding region sequence to obtain an optimized second mRNA sequence for synthesizing the target protein, with the goal of maximizing a first score of the first mRNA sequence, wherein the first score reflects at least one of the following indicators of translation initiation efficiency, codon compatibility index, and minimum free energy of the first mRNA sequence.
13. An electronic device, At least one processor; and a memory communicatively coupled to the at least one processor, wherein: The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform a method according to any one of claims 1 to 11.
14. A non-transitory computer readable storage medium having computer instructions stored thereon, the computer instructions being used to cause a computer to perform the method of any one of claims 1 to 11.
15. A computer program product comprising computer program instructions which, when executed by a processor, implements the method according to any one of claims 1 to 11.
16. 16. An mRNA molecule, the sequence of which is produced by a method according to any one of claims 1 to 11, an apparatus according to claim 12, an electronic device according to claim 13 or a computer program product according to claim 15.
17. A pharmaceutical composition comprising the mRNA molecule of claim 16 and a pharma- ceutically acceptable adjuvant.
18. 18. Use of an mRNA molecule according to claim 16 or a pharmaceutical composition according to claim 17 in the manufacture of a drug or a vaccine.
19. The use according to claim 18, wherein the drug is selected from an mRNA drug, a viral or non-viral vector DNA gene therapy agent or a genetically engineered organism modification agent, and the vaccine is selected from an mRNA prophylactic vaccine or an mRNA therapeutic vaccine.
Citation Information
Patent Citations
Method, device and equipment for optimizing 5'untranslated region sequence of messenger ribonucleic acid
CN116168764A