mRNA sequence optimization method, device, mRNA molecule, pharmaceutical composition and use
By jointly optimizing the 5’ non-translational region and coding region of mRNA, using translation initiation efficiency, codon adaptability index and minimum free energy index, the problem that mRNA design methods in the prior art cannot finely adjust the translation efficiency and stability, and improve the effectiveness of mRNA vaccines and treatment methods.
Patent Information
- Application Number
- CN202411390778.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-09-30
AI Technical Summary
The existing mRNA design methods fail to consider the interaction between the 5’ untranslated region and the coding region from an overall perspective, resulting in the inability to finely adjust the translation efficiency and stability, affecting the effectiveness of mRNA vaccines and treatment methods.
By maximizing the first score of mRNA, the 5’ non-translational region and coding region are jointly optimized, and the translation initiation efficiency, codon adaptability index and minimum free energy index are used to achieve targeted optimization of mRNA sequences.
It improves the overall effectiveness of mRNA vaccines and treatment methods, optimizes the yield of target proteins, and ensures the stability and translation efficiency of mRNA in cells.
Smart Images

Figure CN119265214B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as biological computing, and specifically to an mRNA sequence optimization method and device, electronic equipment, computer-readable storage medium, computer program product, mRNA molecule, pharmaceutical composition and use. Background Art
[0002] Messenger ribonucleic acid (mRNA) vaccines and therapeutics have attracted widespread attention due to their potential in combating various diseases, including infectious diseases and cancer. The translation efficiency and stability of mRNA sequences are particularly important for their design.
[0003] The approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, the issues raised in this section should not be considered as having been recognized in any prior art. Summary of the Invention
[0004] The present disclosure provides a method and apparatus for optimizing an mRNA sequence, an electronic device, a computer-readable storage medium, a computer program product, an mRNA molecule, a pharmaceutical composition, and uses of mRNA.
[0005] According to one aspect of the present disclosure, a method for optimizing an mRNA sequence is provided, comprising: obtaining a first mRNA sequence for synthesizing a target protein, wherein the first mRNA sequence includes a 5' untranslated region sequence and a coding region sequence; and adjusting the 5' untranslated region sequence and the coding region sequence with the goal of maximizing a first score of the first mRNA sequence to obtain an optimized second mRNA sequence for synthesizing the target protein, wherein the first score reflects at least one of the following indicators of the first mRNA sequence: translation initiation efficiency, codon adaptability index, and minimum free energy.
[0006] According to another aspect of the present disclosure, an mRNA sequence optimization device is provided, comprising: an acquisition unit configured to acquire a first mRNA sequence for synthesizing a target protein, wherein the first mRNA sequence includes a 5' untranslated region sequence and a coding region sequence; and a processing unit configured to adjust the 5' untranslated region sequence and the coding region sequence with the goal of maximizing a first score of the first mRNA sequence to obtain an optimized second mRNA sequence for synthesizing the target protein, wherein the first score reflects at least one of the following indicators of the first mRNA sequence: translation initiation efficiency, codon adaptability index, and minimum free energy.
[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory on which a computer program is stored, wherein when the computer program is executed by the processor, the processor is caused to perform the above method.
[0008] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor is caused to perform the above method.
[0009] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, causes the processor to perform the above method.
[0010] According to another aspect of the present disclosure, an mRNA molecule is provided, the sequence of which is prepared by the above method.
[0011] According to another aspect of the present disclosure, a pharmaceutical composition is provided, which is composed of the mRNA sequence or molecule prepared by the above method and pharmaceutically acceptable excipients.
[0012] According to another aspect of the present disclosure, there is provided a use of the mRNA sequence or molecule prepared by the above method or the above pharmaceutical composition in preparing a medicine or a vaccine.
[0013] According to one or more embodiments of the present disclosure, by jointly optimizing the 5' untranslated region (UTR) and coding sequence (CDS) of mRNA with the goal of maximizing the first score of mRNA, targeted optimization of the translation efficiency and stability of mRNA is achieved from a holistic perspective, thereby optimizing the final protein yield and improving the overall efficacy of mRNA vaccines and treatment methods.
[0014] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings illustrate exemplary embodiments and constitute a part of the specification. Together with the description of the specification, they serve to explain exemplary implementation of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals designate similar, but not necessarily identical, elements.
[0016] Figure 1A flow chart of a method for optimizing an mRNA sequence according to an embodiment of the present disclosure is shown;
[0017] Figure 2 A flowchart showing a process of obtaining a 5' untranslated region sequence of a first mRNA sequence for synthesizing a target protein according to an embodiment of the present disclosure is shown;
[0018] Figure 3 A flowchart showing a process of obtaining a coding region sequence of a first mRNA sequence for synthesizing a target protein according to an embodiment of the present disclosure is shown;
[0019] Figure 4 A flowchart illustrating a process for adjusting the 5' untranslated region sequence and the coding region sequence to obtain an optimized second mRNA sequence for synthesizing a target protein with the goal of maximizing the first score of the first mRNA sequence according to an embodiment of the present disclosure;
[0020] Figure 5 A flowchart illustrating another process for adjusting the 5' untranslated region sequence and the coding region sequence with the goal of maximizing the first score of the first mRNA sequence to obtain an optimized second mRNA sequence for synthesizing a target protein according to an embodiment of the present disclosure;
[0021] Figure 6 A flowchart illustrating a process of adjusting the translation initiation region sequence to obtain a fourth mRNA sequence with the goal of maximizing the translation initiation efficiency of the first mRNA sequence according to an embodiment of the present disclosure is shown;
[0022] Figure 7 A flowchart showing a process of calculating the translation initiation efficiency of each of at least one fifth mRNA sequence according to an embodiment of the present disclosure is shown;
[0023] Figure 8 A flowchart illustrating a process of adjusting the main sequence of the coding region of the fourth mRNA sequence to obtain a second mRNA sequence with the goal of maximizing the first score of the fourth mRNA sequence according to an embodiment of the present disclosure is shown;
[0024] Figure 9 Shown is a structural block diagram of an mRNA sequence optimization device;
[0025] Figure 10 The distribution of mRNA sequences in the metric space and their corresponding protein expression levels are shown;
[0026] Figure 11 The mediation effect of method 100 on the three indicators of TIE, MFE and CAI is shown;
[0027] Figure 12The comparison of indicators between wild-type, LinearDesign, Method 100, and third-party designed mRNA sequences is shown;
[0028] Figure 13 shows the results of analyzing the accuracy of TIE metrics using massively parallel translation assay (MPTA) data; and
[0029] Figure 14 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0030] In the following description, for purposes of explanation, specific details are set forth to provide an understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these details. Furthermore, those skilled in the art will recognize that the embodiments of the present disclosure described below may be implemented in a variety of ways, such as as a process, apparatus, system, device, or method on a tangible computer-readable medium.
[0031] The components or modules shown in the figures are illustrative of exemplary embodiments of the present disclosure and are intended to avoid obscuring the present disclosure. It should also be understood that throughout the discussion, components may be described as separate functional units that may include sub-units, but those skilled in the art will recognize that various components or portions thereof may be divided into separate groups.
[0032] Components may be integrated together, including, for example, in a single system or component. It should be noted that the functions or operations discussed herein may be implemented as components. Components may be implemented in software, hardware, or a combination thereof.
[0033] Furthermore, the connections between components or systems in the figures are not intended to be limited to direct connections. Instead, the connections between these components
[0034] Data may be modified, reformatted, or otherwise altered by intermediate components. Furthermore, more or fewer connections may be used. It should also be noted that the terms "couple," "connect," "communicatively coupled," "interface," "access," or any derivatives thereof, are understood to include direct connections, indirect connections through one or more intermediate devices, and wireless connections. It should also be noted that any communication, such as a signal, response, reply, confirmation, message, query, or the like, may include one or more information exchanges.
[0035] References in the specification to "one or more embodiments," "preferred embodiments," "embodiments," "multiple embodiments," etc., indicate that a particular feature, structure, characteristic, or function described in connection with the embodiment is included in at least one embodiment of the present disclosure and may be present in more than one embodiment. Furthermore, appearances of the above phrases in various places in the specification are not necessarily all referring to the same embodiment or embodiments.
[0036] The use of certain terms in various places in the specification is for illustrative purposes and should not be construed as limiting. Services, functions, or resources are not limited to a single service, function, or resource; the use of these terms may refer to a group of related services, functions, or resources, which may be distributed or aggregated. The terms "include," "includes," "have," and "contain" should be understood as open-ended terms, and any lists below are examples and are not meant to be limited to the items listed. A "layer" may include one or more operations. The words "optimal," "optimize," "optimize," and the like refer to improvements in results or processes, and do not require that the specified results or processes have reached an "optimal" or peak state. The use of memory, database, repository, data store, table, hardware, cache, and the like may be used herein to refer to system components or components that can input or otherwise record information.
[0037] In one or more embodiments, stopping conditions may include: (1) a set number of iterations have been performed; (2) a certain processing time has been reached; (3) convergence (e.g., the difference between consecutive iterations is less than a first threshold); (4) divergence (e.g., performance deteriorates); (5) an acceptable result has been achieved.
[0038] Those skilled in the art will recognize that: (1) certain steps may be performed optionally; (2) the steps may not be limited to the specific order set forth herein; (3) certain steps may be performed in a different order; and (4) certain steps may be performed simultaneously.
[0039] Any headings used herein are for organizational purposes only and should not be used to limit the scope of the specification or claims. Each reference / document mentioned in this patent document is incorporated herein by reference in its entirety.
[0040] It should be noted that any experiments and results provided herein are provided in an illustrative manner and are performed using specific embodiments under specific conditions; therefore, these experiments and their results should not be used to limit the scope of the disclosure of this patent document.
[0041] The translation efficiency and stability of mRNA sequences are particularly important for their design. Translation efficiency indicates how quickly an mRNA sequence can produce protein, while stability indicates how long the mRNA sequence can continue translating protein. Together, these two factors determine the quantity of protein an mRNA sequence can produce and ultimately influence the effectiveness of an mRNA vaccine, drug, or therapy.
[0042] The mRNA design methods in related technologies usually focus on designing a single fragment in the mRNA, such as the 5' untranslated region or the coding region, without considering the interactions between the fragments and failing to finely adjust the translation efficiency and stability of the mRNA from a holistic perspective.
[0043] To address the above issues, the present disclosure provides an mRNA sequence optimization method. This method aims to maximize the first score of the mRNA sequence and jointly optimizes the 5' untranslated region and the coding region. The first score can reflect at least one of the following indicators: translation initiation efficiency, codon adaptability index, and minimum free energy. This allows for targeted optimization of the translation efficiency and stability of the mRNA sequence according to design requirements, thereby optimizing the final target protein yield and improving the overall efficacy of mRNA vaccines and treatment methods.
[0044] Exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0045] According to one aspect of the present disclosure, a method for optimizing a messenger ribonucleotide (mRNA) sequence is provided. Figure 1 FIG. 1 is a flow chart of a method 100 for optimizing a messenger ribonucleotide (mRNA) sequence according to an embodiment of the present disclosure. Figure 1 As shown, method 100 includes: step S101, obtaining a first mRNA sequence for synthesizing a target protein, wherein the first mRNA sequence includes a 5' untranslated region sequence and a coding region sequence; and step S102, adjusting the 5' untranslated region sequence and the coding region sequence with the goal of maximizing a first score of the first mRNA sequence to obtain an optimized second mRNA sequence for synthesizing the target protein, wherein the first score reflects at least one of the following indicators of the first mRNA sequence: translation initiation efficiency, codon adaptability index, and minimum free energy.
[0046] According to the embodiments of the present disclosure, the 5'UTR and CDS are jointly optimized with the goal of maximizing the first score of the mRNA sequence. The first score can reflect at least one of the following indicators: translation initiation efficiency, codon adaptability index, and minimum free energy. This allows for targeted optimization of the translation efficiency and stability of the mRNA sequence according to design requirements, thereby optimizing the final target protein yield and improving the overall efficacy of mRNA vaccines and treatments.
[0047] In the mRNA sequence, the 5' untranslated region sequence, also known as the 5'UTR sequence, is located at the 5' end of the mRNA molecule, that is, starting from the 5' cap structure and before the coding region. This fragment has the function of regulating translation, that is, the 5'UTR contains regulatory elements, such as upstream open reading frames (uORFs), suboptimal binding sites (such as GC-rich regions) and regulatory sequences. This fragment can affect the stability and translation efficiency of mRNA. In addition, this fragment has the function of ensuring the stability of the mRNA molecule, and certain sequence elements of this fragment help protect mRNA from degradation. This fragment can promote ribosome binding, by recognizing and binding to specific sequences (such as Kozak sequences) by the ribosome to initiate the translation process. And in mRNA processing, the signal sequence in the 5'UTR plays a decisive role in mRNA splicing and maturation.
[0048] Within an mRNA sequence, the coding region, also known as the CDS, is located between the 5' and 3' untranslated regions of the mRNA molecule. The coding region contains an open reading frame (ORF), which consists of a series of codons, each corresponding to a specific amino acid. This sequence is translated into protein in the ribosome. The coding region contains all the genetic information needed to synthesize a protein and typically begins with a start codon (such as AUG) and ends with a stop codon (such as UAA, UAG, or UGA).
[0049] The 5' UTR and coding region are crucial for protein synthesis. The 5' UTR regulates mRNA stability and translation efficiency, while the coding region directly determines the protein's amino acid sequence. By jointly optimizing the 5' UTR and coding region, the overall translation efficiency and stability of mRNA can be improved, thereby increasing the yield of the target protein.
[0050] In an embodiment of the present disclosure, step S102 is to maximize the first score of the first mRNA sequence, and the 5' untranslated region sequence and the coding region sequence are jointly adjusted to obtain the optimized second mRNA sequence. The first score can reflect at least one of the following indicators of the first mRNA sequence: translation initiation efficiency (TIE), codon adaptation index (CAI) and minimum free energy (MFE). That is, the first score is calculated based on at least one of the three indicators of translation initiation efficiency, codon adaptation index and minimum free energy of the first mRNA sequence.
[0051] Translation initiation efficiency (TIE) is used to measure the efficiency of translation initiation by ribosomes on mRNA molecules, thereby measuring the translation efficiency of the mRNA sequence. The larger the TIE value, the faster the translation process is initiated and the higher the translation efficiency of the mRNA sequence. When the first score is calculated based on TIE, step S102 can achieve targeted optimization of TIE for the mRNA sequence, thereby improving the overall efficiency of protein synthesis and ensuring that a robust and rapid response is generated when the mRNA enters the cell.
[0052] The codon adaptability index (CAI) is used to measure the degree of consistency between the codons in the mRNA sequence and the most commonly used codons in the host cell. The larger the CAI value, the closer the codons used in the mRNA sequence are to the codons of highly expressed genes in the host cell, thereby potentially achieving higher translation efficiency. When the first score is calculated based on the CAI, step S102 can achieve targeted optimization of the CAI of the mRNA sequence, ensuring that the mRNA uses the codons preferred by the host translation machinery, thereby improving the rate and accuracy of protein synthesis.
[0053] The minimum free energy (MFE) is used to measure the energy state of an mRNA molecule when forming a secondary structure, thereby measuring the structural stability of the mRNA molecule. The smaller the value of MFE, the more stable the structure of the mRNA. A stable mRNA structure helps protect the mRNA from degradation, thereby improving its stability and half-life in the cell. When the first score is calculated based on the MFE, step S102 can achieve targeted optimization of the MFE of the mRNA sequence, thereby improving the stability of the mRNA, protecting the mRNA from degradation and increasing its survival time in the cellular environment. However, an overly stable structure may hinder ribosome binding and translation initiation. Therefore, MFE needs to be balanced with TIE and CAI to achieve optimal performance of mRNA.
[0054] It can be understood that since TIE and CAI are positively correlated with the translation efficiency of mRNA and MFE is negatively correlated with the stability of mRNA, in order to optimize the translation efficiency and stability of the mRNA sequence, the first score can be set to be positively correlated with TIE and CAI, and negatively correlated with MFE.
[0055] In some embodiments, the first score S can be calculated according to the following formula (1):
[0056] S = λ TIE * TIE + λ CAI * CAI - λ MFE * MFE (1)
[0057] Among them, λ TIE ,λ CAI ,λ MFE are the weights of TIE, CAI, and MFE indicators respectively. TIE ,λ CAI ,λ MFE The value of can be set according to the design requirements of mRNA, so as to achieve the balance and flexible regulation of the three indicators of TIE, CAI, and MFE, so that the generated mRNA sequence has the required characteristics.
[0058] In some embodiments, the weight of one of the indicators in formula (1) can be set to a fixed value (e.g., 1), and the balance among the three indicators can be achieved by adjusting the weights of the other two indicators. For example, the weight of the MFE indicator can be set to 1, and the balance among TIE, CAI, and MFE can be achieved by adjusting the weights of TIE and CAI. In this embodiment, formula (1) is simplified to the following formula (2):
[0059] S = λ TIE * TIE + λ CAI * CAI - MFE (2)
[0060] In some embodiments, the first score S can be calculated according to the following formula (3):
[0061] S = λ TIE * L * log(TIE)+λ CAI * L * log(CAI) - MFE (3)
[0062] In the above formula, L is the number of codons included in the coding region sequence. By introducing L into the TIE and CAI terms, the values of the TIE, CAI, and MFE terms in formula (3) can be made similar in magnitude, thereby facilitating the balance and flexible regulation of the three indicators of TIE, CAI, and MFE. By performing logarithmic transformations on TIE and CAI (expressed as log(TIE) and log(CAI) respectively), the multiplication operations between the internal factors in the calculation of TIE and CAI can be converted into addition operations, thereby simplifying the calculation.
[0063] It should be noted that, in addition to the 5' untranslated region and the coding region, mRNA also includes other constituent fragments, such as a 5' cap structure, a 3' untranslated region and a poly (A) tail. The embodiments disclosed herein jointly optimize the 5' untranslated region and the coding region. Although other fragments in the mRNA are not optimized (pre-set fragments can be directly used), they may participate in the calculation of the first score. For example, the 3' untranslated region may be related to the numerical value of TIE (for example, the structural features of the 3' untranslated region are taken into account when calculating TIE), and therefore will affect the first score S.
[0064] By jointly optimizing the 5' untranslated region and the coding region using the three indicators of TIE, CAI, and MFE, the optimized second mRNA sequence can balance the three key aspects of translation initiation efficiency, translation elongation efficiency (corresponding to CAI), and stability, thereby optimizing the final protein yield.
[0065] The TIE of an mRNA sequence can be calculated, for example, using the translation initiation efficiency prediction model described below. The CAI of an mRNA sequence can be obtained, for example, by comparing the codon usage of the mRNA sequence with the codon usage of a predetermined highly expressed gene. The MFE of an mRNA sequence can be calculated, for example, using algorithms such as the thermodynamic perturbation method and the thermodynamic calculus method.
[0066] In some embodiments, with respect to step S101, each component fragment of the first mRNA sequence can be obtained separately, and then each component fragment can be spliced together to obtain the first mRNA. Specifically, the 5' untranslated region sequence of the first mRNA sequence can be obtained by the following process 200; the coding region sequence of the first mRNA sequence can be obtained by the following process 300; other components of the first mRNA sequence, such as the 3' untranslated region sequence, can use preset values.
[0067] Figure 2FIG2 is a flowchart showing a process 200 for obtaining a 5' untranslated region sequence of a first mRNA sequence for synthesizing a target protein according to an embodiment of the present disclosure. The process 200 can be used to implement step S101 in the above method 100. In some embodiments, as Figure 2 As shown, process 200 may include: step S201, obtaining a preset untranslated region sequence library, wherein the untranslated region sequence library includes at least one candidate 5' untranslated region sequence, and each candidate 5' untranslated region sequence in the at least one candidate 5' untranslated region sequence can achieve gene expression; and step S202, determining the 5' untranslated region sequence included in the first mRNA sequence from the at least one candidate 5' untranslated region sequence.
[0068] According to the above embodiment, selecting a known 5' untranslated region sequence capable of achieving gene expression as the initial value of the 5' untranslated region sequence in the mRNA sequence can ensure the quality of the 5' untranslated region sequence and provide a better sample for subsequent further optimization.
[0069] In some embodiments, in step S201, in order to ensure that the 5' untranslated region sequence in the first mRNA sequence is a sequence that can be normally expressed, a non-translated region sequence library can be constructed based on known mRNA databases, such as UTRdb, NCBI (National Center for Biotechnology Information), UTRsite, EMBL (European Molecular Biology Laboratory Database), ENSEMBL and other databases. The candidate 5' untranslated region sequence in the non-translated region sequence library can be a natural sequence in the above-mentioned mRNA database, or a sequence obtained by artificial optimization. By constructing the non-translated region sequence library, the selection range of the 5' untranslated region sequence can be expanded, providing a better sample for subsequent optimization.
[0070] In some embodiments, in step S202, a 5'-untranslated region sequence is selected from the untranslated region sequence library constructed in S201 as the 5'-untranslated region sequence in the first mRNA sequence. By selecting a 5'-untranslated region sequence from the untranslated region sequence library, it can be ensured that the selected 5'-untranslated region sequence has normal expression capacity and will not adversely affect subsequent optimization.
[0071] Figure 3 FIG. 3 is a flow chart showing a process 300 for obtaining a coding region sequence of a first mRNA sequence for synthesizing a target protein according to an embodiment of the present disclosure. The process 300 can be used to implement step S101 in the above method 100. In some embodiments, as Figure 3As shown, process 300 may include: step S301, generating an initial coding region sequence corresponding to the amino acid sequence of the target protein; and step S302, adjusting the initial coding region sequence with the goal of maximizing the second score of the initial coding region sequence to obtain a coding region sequence, wherein the second score reflects the codon adaptability index and / or minimum free energy of the initial coding region sequence.
[0072] According to the above embodiment, with the goal of maximizing the second score, the initial coding region sequence is adjusted so that the resulting coding region sequence achieves a balance between translation efficiency and stability while being able to be translated into the target protein, thereby providing a better sample for subsequent further optimization. In the embodiments of the present disclosure, the target protein can be any given protein. Since the target protein is determined, its amino acid sequence can be obtained.
[0073] In some embodiments, in step S301, the amino acid sequence of the target protein can be obtained based on known information or through conventional techniques, including but not limited to gene cloning and sequencing, transcriptome sequencing, protein sequencing, computational prediction, yeast two-hybrid systems, and protein chips. Based on the amino acid-codon correspondence rules, the codons corresponding to each amino acid in the target protein can be obtained, and then the codons corresponding to each amino acid of the target protein can be spliced together to obtain the initial coding region sequence. In this way, an accurate initial coding region sequence capable of being translated into the target protein can be provided for the first mRNA sequence, thereby ensuring the expressive power of the ultimately optimized second mRNA sequence.
[0074] In some embodiments, step S302 adjusts the initial coding region sequence with the goal of maximizing the second score of the initial coding region sequence in step S301, thereby obtaining an optimized coding region sequence that is a component of the first mRNA sequence. The second score can reflect the codon adaptability index and / or minimum free energy of the initial coding region sequence. That is, the second score is calculated based on the codon adaptability index and / or minimum free energy of the initial coding region sequence.
[0075] In some embodiments, the second score S' can be calculated according to the following formula (4):
[0076] S' = -λ MFE * MFE + λ CAI * CAI (4)
[0077] Among them, λ MFE ,λ CAI are the weights of MFE and CAI indicators respectively. MFE ,λ CAIThe value of can be set according to requirements, thereby achieving a balance and flexible regulation of the MFE and CAI indicators, so that the generated coding region sequence has the required characteristics.
[0078] In some embodiments, the weight of one of the indicators in formula (4) can be set to a fixed value (e.g., 1), and the balance between the two indicators can be achieved by adjusting the weight of the other indicator. For example, the weight of the MFE indicator can be set to 1, and the balance between MFE and CAI can be achieved by adjusting the weight of CAI. In this embodiment, formula (4) is simplified to the following formula (5):
[0079] S' = - MFE + λ CAI * CAI (5)
[0080] In some embodiments, the second score S' can be calculated according to the following formula (6):
[0081] S' = - MFE + λ CAI * L * log(CAI) (6)
[0082] In the above formula, L is the number of codons in the coding region sequence. By incorporating L into the CAI term, the CAI and MFE values in formula (6) can be made similar in magnitude, facilitating the balance and flexible regulation of the two indicators, MFE and CAI. By performing a logarithmic transformation on the CAI (expressed as log(CAI)), the multiplication operations between the internal factors in the CAI calculation can be converted to addition operations, thereby simplifying the calculation.
[0083] Figure 4 A flowchart of a process 400 for adjusting the 5' untranslated region sequence and the coding region sequence to obtain an optimized second mRNA sequence for synthesizing a target protein according to an embodiment of the present disclosure is shown. Process 400 can be used to implement step S102 in the above method 100. In some embodiments, as Figure 4 As shown, process 400 may include: step S401, obtaining at least one third mRNA sequence by mutating the 5' untranslated region sequence and the coding region sequence of the first mRNA sequence; step S402, calculating a first score for each of the at least one third mRNA sequence; and step S403, determining the third mRNA sequence with the largest first score as the second mRNA sequence.
[0084] According to the above embodiment, by simultaneously performing mutation adjustment on the 5' untranslated region sequence and the coding region sequence of the first mRNA sequence, a highly efficient and stable second mRNA sequence can be obtained more quickly.
[0085] In some embodiments, in step S401, the 5' untranslated region sequence and the coding region sequence of the first mRNA sequence are mutated simultaneously to obtain at least one third mRNA sequence. The third mRNA sequence is obtained by randomly changing the nucleotides of the 5' untranslated region sequence and the coding region sequence in the first mRNA. Specifically, the 5' untranslated region sequence and the coding region sequence can be taken as a whole and mutated once or multiple times (i.e., the nucleotides at a randomly selected position are replaced with other nucleotides), and a third mRNA sequence can be obtained after each mutation. By mutating the 5' untranslated region sequence and the coding region sequence, new sequences can be explored to provide more sequence samples for subsequent screening and optimization.
[0086] In some embodiments, in step S402, for each third mRNA sequence, the first score calculation method as described above, such as formula (1) to formula (3), is applied to calculate the first score of the third mRNA sequence.
[0087] In some embodiments, in step S403, at least one third mRNA sequence obtained by mutation is screened based on the first score, and the third mRNA sequence with the highest first score is determined as the optimized second mRNA sequence. The second mRNA sequence having the highest first score means that the sequence has the best overall performance among the plurality of third mRNA sequences, and can achieve a balance between translation initiation efficiency, translation elongation efficiency (corresponding to CAI), and stability.
[0088] In certain embodiments, steps S401-S403 can be executed repeatedly in a loop, and the second mRNA sequence obtained by the current cycle (i.e., the optimization result of the current cycle) can be used as the first mRNA sequence of the next cycle (i.e., the optimization starting point of the next cycle), so as to achieve iterative optimization of the mRNA sequence and obtain the optimal second mRNA sequence. The cycle of steps S401-S403 is until the preset termination condition is met. The termination condition can be, for example, that the number of cycles reaches a predetermined cycle number threshold, the first score of the second mRNA sequence reaches a predetermined first score threshold, the first score of the second mRNA sequence no longer significantly improves (i.e., the first score converges), etc. After the cycle terminates, the second mRNA sequence obtained by the last cycle is used as the final mRNA optimization result.
[0089] The above process 400 can be understood as an evolutionary algorithm, in which the first score is the fitness value of each third mRNA sequence obtained by mutation.
[0090] Figure 5A flowchart of another process 500 is shown for adjusting the 5' untranslated region sequence and the coding region sequence to obtain an optimized second mRNA sequence for synthesizing a target protein, with the goal of maximizing the first score of the first mRNA sequence according to an embodiment of the present disclosure. Process 500 can be used to implement step S102 in the above method 100. In some embodiments, as Figure 5 As shown, process 500 may include: step S501, splitting the 5' untranslated region sequence and the coding region sequence of the first mRNA sequence into a translation initiation region sequence and a main coding region sequence, wherein the translation initiation region sequence at least includes the 5' untranslated region sequence, and the main coding region sequence includes nucleotides in the coding region sequence that are not included in the translation initiation region sequence; step S502, adjusting the translation initiation region sequence with the goal of maximizing the translation initiation efficiency of the first mRNA sequence to obtain a fourth mRNA sequence; and step S503, adjusting the main coding region sequence of the fourth mRNA sequence with the goal of maximizing the first score of the fourth mRNA sequence to obtain a second mRNA sequence.
[0091] According to the above embodiment, based on the influence of each component fragment of mRNA on the protein translation process, the 5' untranslated region sequence and the coding region sequence are split into two parts: the translation initiation region sequence and the main coding region sequence. These two parts are optimized successively according to different optimization goals, thereby achieving more refined and targeted optimization of the translation efficiency and stability of mRNA.
[0092] In certain embodiments, in step S501, the 5' non-translated region sequence and the coding region sequence of the first mRNA are adjusted, and the nucleotides of the preset number of 5' non-translated region sequences in the 5' non-translated region sequence and the coding region sequence of the first mRNA are composed of a translation initiation region sequence. In certain embodiments, the preset number is preferably 30. Due to the need to apply ribosomes to carry codons for translation during the translation process, ribosomes occupy approximately 30 nucleotide lengths on mRNA. The residence time of ribosomes in the leader region of the coding region may affect the assembly and translation initiation of subsequent ribosomes, thereby affecting the translation efficiency of mRNA. Therefore, when setting the translation initiation region sequence, it is necessary to consider the position problem of ribosomes occupying mRNA, and the first 30 nucleotides of the coding region sequence are divided into the translation initiation region sequence.
[0093] As described above, a predetermined number (e.g., 30) of nucleotides in the 5' untranslated region (UTR) and the coding region (CGR) adjacent to the 5' UTR sequence influence ribosome assembly and translation initiation. Optimizing the translation initiation region sequence by combining the two can improve translation initiation efficiency in a targeted manner, thereby increasing the translation efficiency of the mRNA sequence.
[0094] In certain embodiments, after obtaining the translation initiation region sequence, certain pre-treatment can be carried out to the translation initiation region sequence, and the pre-treated translation initiation region sequence is optimized.The pre-treatment operation of the translation initiation region sequence for example comprises the -3 position of identification 5 ' UTR terminal, guarantees that this position is purine (A or G), so that it meets the Kozak sequence feature, helps to improve the efficient of translation initiation like this.The pre-treatment operation of the translation initiation region sequence for example also comprises analyzing 5 ' UTR zone, identifies all possible upstream start codons (uAUG).For each uAUG identified, by any Nucleotide in AUG being replaced by the Nucleotide of other types, thereby stop it as translation initiation site, further improve translation efficiency, and guarantee that the initiation site of translation process is accurate, avoid causing and cannot generate target protein because of producing translation misplacement.
[0095] In certain embodiments, in step S502, with the translation initiation efficiency of maximizing the first mRNA sequence as a target, the translation initiation region sequence in the first mRNA sequence is adjusted to obtain the 4th mRNA sequence. It is understood that there are differences in the translation initiation region sequence of the 4th mRNA sequence and the first mRNA sequence, but the coding region main sequence of the two is identical. According to this embodiment, the 4th mRNA sequence has the translation initiation efficiency that is maximized, thereby can provide a good basis for the subsequent optimization for the coding region main sequence.
[0096] Figure 6 FIG. 6 is a flow chart showing a process 600 of adjusting the translation initiation region sequence to obtain a fourth mRNA sequence with the goal of maximizing the translation initiation efficiency of the first mRNA sequence according to an embodiment of the present disclosure. Process 600 can be used to implement step S502 above. In some embodiments, Figure 6 As shown, process 600 may include: step S601, obtaining at least one fifth mRNA sequence by mutating the translation initiation region sequence of the first mRNA sequence; step S602, calculating the translation initiation efficiency of each of the at least one fifth mRNA sequence; and step S603, determining the fifth mRNA sequence with the highest translation initiation efficiency as the fourth mRNA sequence.
[0097] According to the above embodiment, the translation initiation efficiency of the fourth mRNA can be effectively improved, thereby ensuring the translation efficiency of the ultimately generated second mRNA.
[0098] In some embodiments, in step S601, the number of translation initiation region sequences can be enriched by multiple mutations in the translation initiation region sequence of the first mRNA sequence, thereby obtaining multiple fifth mRNA sequences. This method can increase the sample size to be optimized, providing a rich sample base for subsequent screening of the fifth mRNA sequence with the highest translation initiation efficiency.
[0099] In some embodiments, in step S602, the translation initiation efficiency of each fifth mRNA sequence is calculated.
[0100] Figure 7 FIG. 7 is a flow chart showing a process 700 for calculating the translation initiation efficiency of at least one fifth mRNA sequence according to an embodiment of the present disclosure. The process 700 can be used to implement step S602 in the above method 600. In some embodiments, Figure 7 As shown, process 700 may include: performing the following steps for each fifth mRNA sequence of at least one fifth mRNA sequence: step S701, extracting features for predicting the translation initiation efficiency of the fifth mRNA sequence; and step S702, inputting the features into a trained translation initiation efficiency prediction model to obtain the translation initiation efficiency of the fifth mRNA sequence output by the translation initiation efficiency prediction model.
[0101] Translation initiation efficiency may be affected by a variety of factors. According to the above embodiment, by extracting the features of the fifth mRNA sequence and analyzing the features using a trained translation initiation efficiency prediction model to obtain the translation initiation efficiency of the fifth mRNA sequence, the accuracy and generalizability of translation initiation efficiency assessment can be improved.
[0102] In some embodiments, in step S701, one or more features of the fifth mRNA sequence are extracted as input to a translation initiation efficiency prediction model. In some embodiments, the features used to predict the translation initiation efficiency of the fifth mRNA sequence include at least one of the following: structural compactness of the translation initiation region (TIR_ddG_pNT), overall structural compactness (whole_MFE_pNT), Kozak sequence features (purime_m3), upstream start codon (uAUG) and upstream open reading frame (uORF) sequence features, and ribosome residence time (CDS_leader_DT) of the CDS leader region.
[0103] According to the above embodiment, by flexibly selecting a feature combination for predicting translation initiation efficiency, the sequence characteristics and structural characteristics of the translation initiation region of the fifth mRNA sequence can be flexibly and comprehensively obtained, thereby more accurately predicting its translation initiation efficiency.
[0104] The structural compactness of the translation initiation region (TIR_ddG_pNT) represents the free energy change of the secondary structure of the translation initiation region (including the 5' UTR and the 5' leader sequence of the CDS) before and after unfolding. A lower free energy change indicates a more compact structure and is generally associated with a lower TIE.
[0105] The feature whole structure compactness (whole_MFE_pNT) measures the minimum free energy (MFE) of the entire mRNA sequence (including 5'UTR, CDS and 3'UTR) and is normalized by sequence length. A higher normalized MFE indicates a less stable overall structure, which is generally positively correlated with TIE.
[0106] Kozak sequence signature (purime_m3): The presence of a purine (A / G) at the -3 position of the 5' UTR is a hallmark of a Kozak sequence that enhances translation initiation. This signature is positively correlated with TIE.
[0107] uAUG and uORF sequence features include:
[0108] In-frame upstream open reading frame (in_frame_uORF): The presence of upstream open reading frames (uORFs) in frame with the main open reading frame (ORF) can inhibit downstream translation and negatively affect TIE.
[0109] Out-of-frame start codons (out_frame_uAUG): Start codons located upstream of the main open reading frame (ORF) and out of reading frame are negatively correlated with TIE.
[0110] The CDS leader dwell time feature measures the time that ribosomes spend in the 5' leader region of the CDS. Since ribosomes occupy approximately 30 nucleotides on an mRNA, a longer dwell time may interfere with subsequent ribosome assembly and translation initiation, and is therefore negatively correlated with TIE.
[0111] By evaluating the above characteristics, the sequence characteristics and structural features of the translation initiation region of the fifth mRNA can be comprehensively and accurately obtained, thereby more accurately calculating the translation initiation efficiency.
[0112] In some embodiments, in step S702, the features obtained in step S701 are input into a trained translation initiation efficiency prediction model to obtain the translation initiation efficiency of the fifth mRNA sequence output by the translation initiation efficiency prediction model.
[0113] The translation initiation efficiency prediction model can be any machine learning model, including but not limited to a regression model, a decision tree model, a random forest model, a neural network model, etc. The translation initiation efficiency prediction model can be trained using sequence features labeled with translation initiation efficiency labels as samples.
[0114] In some exemplary embodiments, a ridge regression model can be used as a translation initiation efficiency prediction model. The ridge regression model can be used to predict logarithmic transformed TIE (i.e., log(TIE)). The model can handle multicollinearity between features and can prevent overfitting through regularization. The training data for the model can be selected from eGFP polymer analysis data and human genome ribosome analysis data, which can be selected from the National Genome Science Data Center, for example. These data sets provide comprehensive insights into the dynamics of mRNA translation. At the same time, in order to ensure consistency and improve model performance, the above features for each input are scaled, and a ridge regression model is constructed using the above features as predictor variables. Ridge regression introduces a penalty term proportional to the square of the coefficient size, which avoids over-reliance on any single feature. The model is trained on the collected data set and uses a combination of mean square error (MSE) and R 2 Standard metrics including the score are used as the model's loss function to evaluate the model's performance. Cross-validation is used to assess the model's robustness and fine-tune the model's hyperparameters.
[0115] After obtaining the translation initiation efficiency of each fifth mRNA sequence in step S602, step S603 may be performed. In step S603, the fifth mRNA sequence with the highest translation initiation efficiency is selected as the fourth mRNA sequence.
[0116] In certain embodiments, step S601-S603 can be executed in a loop repeatedly, and the 4th mRNA sequence (i.e. the optimization result of current circulation) obtained by current circulation can be used as the first mRNA sequence (i.e. the optimization starting point of next circulation) of next circulation, thus realizes the iterative optimization of mRNA sequence, obtains optimal 4th mRNA sequence.The circulation of step S601-S603 is until meeting preset termination condition.Termination condition can for example be that cycle index reaches predetermined cycle index threshold value, the 4th mRNA sequence translation initiation efficiency reaches predetermined translation initiation efficiency threshold value, the 4th mRNA sequence translation initiation efficiency no longer significantly promotes (i.e. translation initiation efficiency converges) etc.After circulation terminates, the 4th mRNA sequence obtained by last circulation is used as the optimization result for translation initiation region sequence.
[0117] In some exemplary embodiments, process 600 can be understood as an evolutionary algorithm. The specific operations of the algorithm are as follows:
[0118] Designing an initial population: mutating the translation initiation region sequence of the first mRNA sequence to construct an initial population consisting of multiple mRNA sequences.
[0119] Define the fitness function: Use translation initiation efficiency (TIE) as the fitness function to evaluate the performance of each sequence variant.
[0120] Iterative optimization process: It iteratively optimizes a population of sequences by applying mutation and selection operations, simulating the process of natural selection.
[0121] Mutation: Randomly changing nucleotides in the sequence to explore new sequence space and obtain multiple fifth mRNA sequences.
[0122] Selection: Based on the TIE evaluation results of each fifth mRNA sequence, the sequence with the largest TIE is selected as the current optimal fourth mRNA sequence for the next generation of iterations.
[0123] Termination condition: When the predetermined number of iterations is reached or the sequence performance no longer improves significantly, the iteration stops.
[0124] In some embodiments, in step S503, based on the fourth mRNA sequence obtained in step S502 after the translation initiation region optimization is completed, the main sequence of the coding region of the fourth mRNA sequence is adjusted with the goal of maximizing the first score of the fourth mRNA sequence to obtain an optimized second mRNA sequence.
[0125] Figure 8 FIG. 8 is a flowchart illustrating a process 800 for adjusting the coding region main sequence of a fourth mRNA sequence to obtain a second mRNA sequence with the goal of maximizing the first score of the fourth mRNA sequence according to an embodiment of the present disclosure. The process 800 can be used to implement step S503 in the above method 500. In some embodiments, as Figure 8 As shown, process 800 may include: step S801, obtaining at least one sixth mRNA sequence by mutating the main sequence of the coding region of the fourth mRNA sequence; step S802, calculating the first score of each of the at least one sixth mRNA sequence; and step S803, determining the sixth mRNA sequence with the largest first score as the second mRNA sequence.
[0126] According to the above embodiment, the first score can reflect at least one indicator among the translation initiation efficiency, codon adaptability index and minimum free energy, so that the translation efficiency and stability of the mRNA sequence can be optimized according to the design requirements, especially the problem of base structure pairing in the translation initiation region is optimized and improved, thereby optimizing the final target protein yield and improving the overall efficacy of mRNA vaccines and treatment methods.
[0127] In some embodiments, in step S801, multiple mutations are performed on the main coding region sequence of the fourth mRNA sequence to enrich the number of main coding region sequences and obtain multiple sixth mRNA sequences. This method can increase the sample size for optimization, providing a sample basis for subsequent screening of the sixth mRNA with the highest first score.
[0128] In some embodiments, in step S802, a first score for each sixth mRNA is calculated. In this process, the first score calculation formulas (1)-(3) shown above can be applied to calculate the first score corresponding to each sixth mRNA.
[0129] In some embodiments, in step S803, the sixth mRNA sequence with the highest score is selected from the sixth mRNA sequences scored in step S802 as the optimized second mRNA sequence. The second mRNA sequence has the highest first score, which means that the sequence has the best overall performance among the plurality of sixth mRNA sequences, and can achieve a balance between the three key aspects of translation initiation efficiency, translation elongation efficiency (corresponding to CAI), and stability.
[0130] In some embodiments, in step S803, in response to the codon adaptability index of the sixth mRNA sequence with the highest first score being greater than a threshold, the sixth mRNA sequence is determined as the optimized second mRNA sequence, wherein the threshold is determined based on the codon adaptability index of the initial first mRNA sequence (i.e., the first mRNA sequence obtained in step S101). For example, the threshold can be set to the codon adaptability index of the initial first mRNA sequence. According to this embodiment, it can be ensured that the optimized second mRNA sequence has a translation expression ability that is no less than that of the initial first mRNA sequence.
[0131] In certain embodiments, in response to the codon adaptability index of the 6th mRNA sequence that is maximum in the first score being less than or equal to a threshold value, this optimization result can be discarded, that is, the 6th mRNA sequence is not used as the second mRNA sequence after optimization, but steps S801-S803 are re-executed, until the optimization result that the codon adaptability index is greater than the threshold value is obtained. In certain embodiments, steps S801-S803 can be cyclically executed multiple times, and the second mRNA sequence obtained by the current cycle (i.e., the optimization result of the current cycle) can be used as the 4th mRNA sequence (i.e., the optimization starting point of the next cycle) of the next cycle, thereby realizing the iterative optimization of the mRNA sequence, obtaining the second optimal mRNA sequence. The circulation of steps S801-S803 is until meeting the preset termination condition. The termination condition can be, for example, that the first scoring of the cycle number reaches a predetermined cycle number threshold, the second mRNA sequence reaches a predetermined first scoring threshold, the second mRNA sequence no longer significantly improves (i.e., the first scoring converges) etc. After the cycle terminates, the second mRNA sequence obtained by the last cycle is used as the final mRNA optimization result. In some exemplary embodiments, steps S801-S803 may be considered as applying an evolutionary algorithm to perform operations, which are specifically as follows:
[0132] Designing an initial population: mutating the main sequence of the coding region of the fourth mRNA sequence to construct an initial population consisting of multiple mRNAs.
[0133] Define the fitness function: Use the first score as the fitness function to evaluate the performance of each sequence variant.
[0134] Iterative optimization process: It iteratively optimizes a population of sequences by applying mutation and selection operations, simulating the process of natural selection.
[0135] Mutation: Mutate the position in the main sequence of the coding region that is paired with the translation initiation region to explore sequence variants that may improve translation initiation efficiency and / or reduce minimum free energy, and obtain multiple sixth mRNA sequences.
[0136] Selection: Based on the evaluation results of the first score of each sixth mRNA sequence, the sequence with the largest first score is selected as the current optimal second mRNA sequence for the next generation of iteration.
[0137] Termination condition: When the predetermined number of iterations is reached or the first score of the second mRNA sequence is no longer significantly improved, the iteration is stopped.
[0138] According to an embodiment of the present disclosure, a device for designing a messenger ribonucleotide (mRNA) sequence is also provided. Figure 9FIG. 5 shows a structural block diagram of a training device for a neural network model for predicting the effect of mutations on protein stability according to an exemplary embodiment of the present disclosure. Figure 9 As shown, the apparatus 900 includes: an acquisition unit 910, configured to acquire a first mRNA sequence for synthesizing a target protein; and a processing unit 920, configured to adjust the 5' untranslated region sequence and the coding region sequence with the goal of maximizing the first score of the first mRNA sequence to obtain an optimized second mRNA sequence for synthesizing the target protein.
[0139] It can be understood that the operations of units 910 to 920 in the apparatus 900 may refer to the above description of steps S101 to S102 in the method 100 and are not described in detail here.
[0140] In an exemplary embodiment, in order to analyze the guiding value of the first scoring formula in method 100, the differences in actual protein production of samples in different regions in the metric space are examined. The mRNA sequence and expression data are from the article Kathrin Leppek, Gun Woo Byeon, Wipapat Kladwang, Hannah K Wayment-Steele, CraigH Kerr, Adele F Xu, Do Soon Kim, Ved V Topkar, Christian Choe, Daphna Rothschild, et al. Combinatorial optimization of mRNA structure, stability, and translation for RNA-based therapeutics. Nature communications, 13(1): 1536, 2022., in which the expression levels of Nluc reporter genes with different CDS sequences are measured by the Nluc / Fluc reporter gene activity ratio. Figure 10As shown in panels A and B, panel A shows the expression levels of each sample after 6 hours, and panel B shows the expression levels of each sample after 24 hours. Darker dots indicate higher expression levels. As shown in the distributions in panels A and B, samples with the highest protein expression are primarily located in regions with low MFE (<-350 kcal / mol), high CAI (>0.75), and moderate TIE (0.35-0.42). This pattern is also observed in panels C and D, which are two-dimensional distribution plots of panels A and B, respectively. Samples with predicted low TIE exhibited significant expression disadvantages at both time points, indicating that sufficient translation efficiency is essential for efficient protein production. The expression level distribution patterns also differed between 6 and 24 hours. Specifically, samples with the highest expression levels at 6 hours tended to have higher TIE, while samples with the highest expression levels at 24 hours tended to have lower MFE. This is consistent with the general principle that short-term expression levels are more influenced by translation efficiency, whereas long-term expression levels are more influenced by mRNA stability.
[0141] However, within this distribution, samples with the highest TIE did not exhibit high protein expression levels. This may be because these samples also had relatively high MFE values, resulting in reduced stability and thus compromising their ability to sustain expression. The trade-off between TIE and MFE as optimization targets is understandable, as a reduced MFE increases mRNA structural compactness, creating a barrier for ribosomes and other translation factors to bind to the mRNA. As shown in panels E and F, panel E shows a scatter plot showing the correlation between TIE and Nluc / Fluc activity over a 24-hour period for samples selected based on an MFE <-350 kcal / mol and a CAI >0.75. Panel F shows a scatter plot showing the correlation between TIE and YFP abundance expressed in yeast over a 24-hour period. In panels E and F, a more pronounced positive correlation between TIE and protein expression levels is observed when samples with excessively high MFE (>350 kcal / mol) and low CAI (<0.75) are filtered out. Specifically, the Spearman correlation between TIE and the Nluc / Fluc activity ratio in the 24-hour expression data reached 0.70 (p < 0.05). Therefore, optimizing TIE may lead to better protein yields by ensuring the relative optimal values of MFE and CAI.
[0142] In an exemplary embodiment, the method 100 introduces two custom parameters λ TIE and λ CAI , to balance the relative weights of the three optimization objectives, namely TIE, CAI and MFE. Parameter λ CAI and λ TIEIn order to enhance the convenience of index adjustment, the optimization algorithm ensures that the CAI index of the target sequence is not affected by λ. TIE This means that once λ is fixed CAI , the CAI value of the designed sequence will remain in a relatively stable range regardless of λ TIE How to change.
[0143] The mRNA sequence of eGFP protein (from GenBank: AFA52650.1) was designed to demonstrate the regulatory ability of method 100 on mRNA indicators. TIE Parameters (2, 4, 6, 8, 10) and four lambdas CAI Parameters (2, 4, 6, 8), there are 20 parameter combinations in total. CAI Parameters precisely adjust the CAI value of the target sequence, such as Figure 11 As shown in subgraph A in . CAI As the value of λ increases, the CAI value of the design sequence gradually increases. CAI value, with different λ TIE The CAI value of the parameter-designed sequence fluctuates very little, indicating that the CAI is mainly determined by λ CAI Parameter adjustment is almost unaffected by λ TIE The effect of λ is consistent with the design expectation of method 100. CAI When fixed, method 100 changes λ TIE Flexible adjustment of the target sequence TIE and MFE indicators is achieved, such as Figure 11 As shown in subgraphs B and C in TIE As the parameters increase, the TIE value gradually increases, indicating an increase in the efficiency of mRNA translation initiation, while the MFE value also increases, suggesting a decrease in the compactness and thermodynamic stability of the mRNA structure. Figure 11 As shown in subfigure D, there is a positive correlation between TIE and MFE values, indicating a negative correlation between translation initiation efficiency and the structural compactness and thermodynamic stability of mRNA. In summary, by adjusting these two hyperparameters, the optimization objective can be flexibly customized to achieve different balances between these three metrics, meeting diverse needs in different scenarios.
[0144] The present disclosure also provides an mRNA molecule, the sequence of which is prepared by the method, apparatus, electronic device or computer program product disclosed herein.
[0145] In an exemplary embodiment, the LinearDesign algorithm and method 100 were evaluated for the design of the novel coronavirus (SARS-CoV-2) spike protein and varicella-zoster virus (VZV) antigen (VZV gE protein, UniProtKB / Swiss-Prot: Q9J3M8.1), wherein the amino acid sequences of the novel coronavirus (SARS-CoV-2) spike protein and varicella-zoster virus (VZV) antigen are available from NCBI (National Center for Biotechnology Information). The LinearDesign algorithm is an existing mRNA sequence design algorithm. For the SARS-CoV-2 spike protein, such as Figure 12 As shown in sub-figure A in , the MFE value of the mRNA designed by LinearDesign is significantly lower than that of the wild type (WT) and commercial vaccine sequences (BNT-162b2 and mRNA-1273) (the commercial vaccine sequences are all from: https: / / github.com / NAalytics / Assemblies-of-putative-SARS-CoV2-spike-encoding-mRNA-sequences-for-vaccines-BNT-162b2-and-mRNA-1273 / ), indicating a significant improvement in structural compactness and thermodynamic stability. However, the TIE value of the sequence designed by LinearDesign is significantly lower, indicating that translation efficiency may be reduced. In contrast, the sequence designed by method 100 has a significantly improved TIE index compared with the sequence designed by LinearDesign, while maintaining similar CAI and MFE values. As Figure 12 As shown in subfigure B in , a similar pattern was observed in the design of VZV antigens. Compared with the wild type (gE-WT) and the sequences designed by ThermoFisher's codon optimization tool (gE-Ther), the sequences designed by LinearDesign have a significant advantage in the MFE metric. However, they have a significant disadvantage in the TIE metric. Method 100 effectively solves this problem by maintaining the advantage of LinearDesign in MFE while significantly improving the performance of the TIE metric. This shows that method 100 is able to produce sequences with superior overall performance in terms of both translation efficiency and stability. Given the previous analysis of the relationship between metric space and protein expression levels, this improvement is expected to lead to increased protein yield.
[0146] There are also significant differences in the secondary structures of the sequences designed by Method 100 and LinearDesign. Figure 12Subfigure C in the figure shows the secondary structures of the two designed eGFP mRNA sequences, where the original amino acid sequence of eGFP is derived from NCBI. Although the two sequences show similar MFE and CAI indicators, the sequence designed by method 100 has a significantly better TIE indicator. The main structural difference between the two sequences can be observed in the start codon region. In this region, the sequence designed by method 100 has fewer hairpin structures and base pairing, resulting in a more relaxed configuration in the structure. This more relaxed structure is generally believed to be conducive to ribosome binding and scanning in the 5'UTR region, thereby improving translation initiation efficiency. When the sequence of the start codon background region was extracted and folded separately, it was also obvious that the sequence designed by method 100 had less secondary structure and higher folding free energy, as shown in Figure 1. Figure 12 This is shown in sub-figure D in . This further supports the technical effect that structural relaxation in the start codon background region helps method 100 achieve improved translation initiation efficiency.
[0147] In one exemplary embodiment, the accuracy of the TIE metric predicted for the eGFP protein was analyzed using massively parallel translation assay (MPTA) data. The samples in this dataset had fixed CDS and randomly generated 5'UTR sequences, and the ribosome load for each sequence was measured using multimer analysis. Given that translation elongation efficiency is relatively constant, the ribosome load value reflects the efficiency of translation initiation for each sequence. Figure 13 As shown in subfigure A in , in the first 20,000 samples sorted by read count, the Spearman correlation coefficient between TIE and ribosome load was 0.83, which was significantly higher than other segmented features. Among the segmented features, the feature associated with the upstream AUG codon (uAUG) had the highest absolute Spearman correlation coefficient with the ribosome load. The presence of uAUG may lead to premature translation initiation and inhibit the translation of the main ORF. In actual mRNA vaccine development scenarios, sequences containing uAUG are generally not used. Therefore, in order to be closer to the real situation, the correlation between various indicators and ribosome load in samples that do not contain uAUG was further analyzed. As shown Figure 13 As shown in panel B, the Spearman correlation between TIE and ribosome loading in these samples reached 0.6, exceeding the correlations of other segmentation metrics. This demonstrates the robustness of TIE as an indicator of translation initiation efficiency, especially in a background without uAUG sequences.
[0148] Since the eGFP protein dataset has a fixed CDS, it cannot effectively reflect the effect of the CDS region on the efficiency of mRNA translation initiation. To solve this problem, we used the ribosome profiling data of the human PC3 cell line (GSE35469) for further analysis. This dataset contains the translation efficiency information of the entire human genome transcripts. There are significant differences in the UTR and CDS sequences between transcripts, which is suitable for analyzing the joint effect of 5'UTR and CDS on translation initiation efficiency. According to the article Nicholas T Ingolia, Liana F Lareau, and Jonathan S Weissman. Ribosome profiling of mouse embryonic stem cells reveals the complexity and dynamics of mammalian proteomes. Cell, 147(4): 789–802, 2011, there is no significant difference in translation elongation efficiency between different genes, and translation initiation efficiency is the main rate-limiting step in the translation process. Therefore, in this case, translation efficiency can be mainly regarded as a representative of translation initiation efficiency.
[0149] like Figure 13 As shown in sub-figure C in , the Spearman correlation coefficient between translation efficiency and various indicators is shown. The correlation between the predicted TIE indicator and the measured translation efficiency (TE) is 0.574, which outperforms other subdivision features. Among other features, the minimum free energy per unit length of the entire mRNA chain (whole_MFE_mean) has the highest correlation with TE, reflecting the significant effect of the compactness of the mRNA structure on the efficiency of translation initiation. An overly compact structure may hinder translation initiation. The ribosome residence time in the CDS leader region (i.e., the ribosome decoding time of the CDS 5' start region, CDS_leader_DT) is significantly negatively correlated with TE, because ribosomes staying in the CDS leader region can spatially hinder the assembly of subsequent ribosomes. These correlation patterns further prove that translation initiation efficiency is jointly determined by the 5'UTR and CDS regions.
[0150] The present disclosure provides a pharmaceutical composition comprising an mRNA sequence prepared by the method, apparatus, electronic device or computer program product disclosed herein or an mRNA molecule disclosed herein and a pharmaceutically acceptable excipient.
[0151] The pharmaceutical compositions of the present disclosure can be formulated in any manner known in the art, including but not limited to tablets, capsules, caplets, suspensions, powders, lyophilized preparations, suppositories, eye drops, skin patches, orally soluble preparations, sprays, aerosols and the like as solid, semi-solid or liquid systems.
[0152] The pharmaceutical composition can be an immediate release and / or modified release formulation, including delayed release, sustained release, pulsed release, controlled release, targeted release, and programmed release formulations.
[0153] Herein, "pharmaceutically acceptable excipient" refers to a component in a pharmaceutical composition that is non-toxic to a subject except for the active ingredient. Pharmaceutically acceptable excipients include, but are not limited to, excipients (e.g., diluents, carriers, etc.) and additives (e.g., stabilizers, preservatives, solubilizers, buffers, etc.). Excipients may include polyvinyl pyrrolidone, gelatin, hydroxypropyl cellulose (HPC), gum arabic, polyethylene glycol, mannitol, sodium chloride, and sodium citrate. For injection preparations or other liquid administration preparations, it is preferred that water contain at least one or more buffer components, and stabilizers, preservatives, and solubilizers may also be used. For solid administration preparations, any one of a variety of thickeners, fillers, extenders, and carrier additives may be used, such as starch, sugar, cellulose derivatives, fatty acids, etc. For topical administration preparations, any one of a variety of creams, ointments, gels, lotions, etc. may be used. For most pharmaceutical preparations, inactive ingredients may account for the majority of the preparation by weight or volume. For pharmaceutical formulations, it is also contemplated that any of a variety of metered-release, slow-release, or sustained-release formulations and additives may be employed so that a dose can be formulated to deliver the disclosed compound over a period of time.
[0154] The compounds of the present disclosure may be administered via mucosal administration, buccal administration, oral administration, transdermal administration, inhalation administration, intranasal administration, urethral administration, vaginal administration, and intravenous, subcutaneous, intramuscular, and intraperitoneal injection. The excipients in the pharmaceutical composition are compatible with the route of administration.
[0155] In some embodiments, the compounds of the present disclosure can be delivered orally, such as tablets or capsules. The compound can be packaged in an enteric protectant, preferably so that the compound is not released until the tablet or capsule is delivered to the stomach and, optionally, further to a portion of the small intestine.
[0156] In some embodiments, the compounds of the present invention can be injected, and the pharmaceutical forms suitable for injection include sterile aqueous solutions or dispersions and sterile powders for the immediate preparation of sterile injectable solutions or dispersions. In all cases, the form must be sterile and its fluidity must allow it to be administered by syringe. The form must be stable under preparation and storage conditions and must be preserved to prevent contamination by microorganisms such as bacteria and fungi. The carrier can be a solvent or dispersion medium containing (for example) water, ethanol, a polyol (for example, glycerol, propylene glycol or liquid polyethylene glycol), a suitable mixture thereof and a vegetable oil.
[0157] Therapeutic administration can also be by injection of sustained-release formulations, such as those allowing subcutaneous injection, including: nanospheres / microspheres, liposomes, emulsions, gels, insoluble salts or suspensions.
[0158] In some embodiments, the compounds of the present disclosure can be administered intranasally. Pharmaceutical compositions can be in the form of aqueous solutions, such as solutions containing saline, citrate, or other commonly used excipients or preservatives. They can also be in the form of dry preparations or powders.
[0159] The present disclosure provides the use of an mRNA sequence prepared by the method, apparatus, electronic device or computer program product disclosed herein, an mRNA molecule disclosed herein or a pharmaceutical composition disclosed herein in the preparation of a drug or vaccine.
[0160] This method significantly improves protein yield and quality, and has important application value in fields such as the preparation of drugs or vaccines.
[0161] In some embodiments, the drugs disclosed herein include, but are not limited to, mRNA drugs, protein replacement therapy drugs, gene editing drugs, cancer treatment drugs, regenerative medicine drugs, DNA gene therapy agents using viral or non-viral vectors, agents for modifying genetically engineered organisms, cell therapy drugs, enzyme replacement therapy drugs, aptamer drugs, microRNA therapy drugs, and ribozyme drugs. In some preferred embodiments, the drugs disclosed herein are selected from mRNA drugs, DNA gene therapy agents using viral or non-viral vectors, or agents for modifying genetically engineered organisms.
[0162] In some embodiments, the vaccine disclosed herein is selected from: an mRNA prophylactic vaccine or an mRNA therapeutic vaccine.
[0163] The present disclosure provides a method for treating or preventing a disease, comprising administering to a subject in need thereof an effective amount of an mRNA sequence prepared by the method, apparatus, electronic device, or computer program product disclosed herein, an mRNA molecule disclosed herein, or a pharmaceutical composition disclosed herein.
[0164] In some embodiments, the diseases disclosed herein include, but are not limited to, infectious diseases, including viral infections such as novel coronavirus and varicella-zoster virus.
[0165] As used herein, "subject" includes animals, such as vertebrates, preferably mammals, such as dogs, cats, pigs, cows, sheep, horses, rodents (e.g., mice, rats, or guinea pigs), or primates (e.g., gorillas, chimpanzees, and humans).
[0166] As used herein, "treating" refers to alleviating or ameliorating a disease or disorder (i.e., slowing or arresting the progression of the disease or at least one clinical symptom); or alleviating or ameliorating at least one physiological parameter or biomarker associated with the disease or disorder.
[0167] As used herein, an "effective amount" is an amount sufficient to induce the desired therapeutic, preventive, or inhibitory effect, such as an amount that, when administered by any of the aforementioned means or any other means known in the art, results in a benefit or effect compared to a corresponding subject not receiving such amount. This amount is sufficiently low within the scope of sound medical judgment to avoid serious side effects. The effective amount will vary depending on the selected drug, such as mRNA, pharmaceutical composition, or vaccine; the route of administration; the severity of the disease being treated; and the age, size, weight, and physical condition of the patient being treated.
[0168] According to an embodiment of the present disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor performs the mRNA sequence optimization method of the embodiment of the present disclosure.
[0169] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to enable a computer to execute the mRNA sequence optimization method of the embodiment of the present disclosure.
[0170] According to an embodiment of the present disclosure, a computer program product is also provided, comprising computer program instructions, which, when executed by a processor, implement the method for mRNA sequence optimization of the embodiment of the present disclosure.
[0171] refer to Figure 14, a block diagram of an electronic device 1400 that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0172] like Figure 14 As shown, electronic device 1400 includes a computing unit 1401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1402 or a computer program loaded from a storage unit 1408 into a random access memory (RAM) 1403. Various programs and data required for the operation of electronic device 1400 can also be stored in RAM 1403. Computing unit 1401, ROM 1402, and RAM 1403 are connected to each other via a bus 1404. An input / output (I / O) interface 1405 is also connected to bus 1404.
[0173] Multiple components in the electronic device 1400 are connected to the I / O interface 1405, including: an input unit 1406, an output unit 1407, a storage unit 1408, and a communication unit 1409. The input unit 1406 can be any type of device that can input information to the electronic device 1400. The input unit 1406 can receive input digital or character information and generate key signal input related to user settings and / or function control of the electronic device, and can include but is not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 1407 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1408 can include but is not limited to a magnetic disk and an optical disk. The communication unit 1409 allows the electronic device 1400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and may include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver and / or a chipset, such as a Bluetooth device, an 802.11 device, a Wi-Fi device, a WiMAX device, a cellular communication device and / or the like.
[0174] The computing unit 1401 may be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 1401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1401 performs the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 may be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 1408. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1400 via ROM 1402 and / or communication unit 1409. When the computer program is loaded into RAM 1403 and executed by the computing unit 501, one or more steps of method 100 described above may be performed. Alternatively, in other embodiments, the computing unit 1401 may be configured to execute the method 100 in any other appropriate manner (eg, by means of firmware).
[0175] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0176] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0177] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0178] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0179] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0180] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0181] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0182] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-mentioned methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present disclosure is not limited by these embodiments or examples, but is only limited by the claims after authorization and their equivalents. Various elements in the embodiments or examples can be omitted or replaced by their equivalents. In addition, the steps can be performed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. It is important that as technology evolves, many of the elements described here can be replaced by equivalent elements that appear after the present disclosure.
Claims
1. A method for optimizing a messenger ribonucleotide (mRNA) sequence, comprising: Obtaining a first mRNA sequence for synthesizing a target protein, wherein the first mRNA sequence includes a 5' untranslated region sequence and a coding region sequence; and With the goal of maximizing the first score of the first mRNA sequence, adjusting the 5' untranslated region sequence and the coding region sequence to obtain an optimized second mRNA sequence for synthesizing the target protein, comprising: Splitting the 5' untranslated region sequence and the coding region sequence into a translation initiation region sequence and a main coding region sequence, wherein the translation initiation region sequence includes the 5' untranslated region sequence and a preset number of nucleotides in the coding region sequence close to the 5' untranslated region sequence, and the main coding region sequence includes nucleotides in the coding region sequence that are not included in the translation initiation region sequence; With the goal of maximizing the translation initiation efficiency of the first mRNA sequence, adjusting the translation initiation region sequence to obtain a fourth mRNA sequence; and With the goal of maximizing the first score of the fourth mRNA sequence, adjusting the main sequence of the coding region of the fourth mRNA sequence to obtain the second mRNA sequence, The first score reflects the following indicators of the first mRNA sequence: translation initiation efficiency, codon adaptability index and minimum free energy. The translation initiation efficiency is used to measure the initiation efficiency of the translation process of ribosomes on mRNA molecules.
2. The method according to claim 1, wherein The obtaining of a first mRNA sequence for synthesizing a target protein comprises: Obtaining a preset untranslated region sequence library, wherein the untranslated region sequence library includes at least one candidate 5' untranslated region sequence, and each candidate 5' untranslated region sequence in the at least one candidate 5' untranslated region sequence can achieve gene expression; and The 5' untranslated region sequence is determined from the at least one candidate 5' untranslated region sequence.
3. The method according to claim 1 or 2, wherein The obtaining of a first mRNA sequence for synthesizing a target protein comprises: generating an initial coding region sequence corresponding to the amino acid sequence of the target protein; and The initial coding region sequence is adjusted with the goal of maximizing the second score of the initial coding region sequence to obtain the coding region sequence, wherein the second score reflects the codon adaptability index and / or minimum free energy of the initial coding region sequence.
4. The method according to claim 1, wherein The step of adjusting the 5' untranslated region sequence and the coding region sequence with the goal of maximizing the first score of the first mRNA sequence to obtain an optimized second mRNA sequence for synthesizing the target protein comprises: At least one third mRNA sequence is obtained by mutating the 5' untranslated region sequence and the coding region sequence; calculating a first score for each of the at least one third mRNA sequence; and The third mRNA sequence with the largest first score is determined as the second mRNA sequence.
5. The method according to claim 1, wherein The step of adjusting the translation initiation region sequence with the goal of maximizing the translation initiation efficiency of the first mRNA sequence to obtain a fourth mRNA sequence comprises: By mutating the translation initiation region sequence, at least one fifth mRNA sequence is obtained; calculating the translation initiation efficiency of each of the at least one fifth mRNA sequences; and The fifth mRNA sequence with the highest translation initiation efficiency is determined as the fourth mRNA sequence.
6. The method according to claim 5, wherein: Calculating the translation initiation efficiency of each of the at least one fifth mRNA sequence comprises: For each fifth mRNA sequence of the at least one fifth mRNA sequence: Extracting features for predicting translation initiation efficiency of the fifth mRNA sequence; and The feature is input into a trained translation initiation efficiency prediction model to obtain the translation initiation efficiency of the fifth mRNA sequence output by the translation initiation efficiency prediction model.
7. The method according to claim 6, wherein: The features include at least one of the following: The structural compactness of the translation initiation region sequence, the overall structural compactness, the Kozak sequence characteristics, the upstream start codon and the upstream open reading frame sequence characteristics and the ribosome residence time of the leading region of the coding region sequence.
8. The method according to claim 1, wherein The step of adjusting the main coding region sequence of the fourth mRNA sequence with the goal of maximizing the first score of the fourth mRNA sequence to obtain the second mRNA sequence comprises: obtaining at least one sixth mRNA sequence by mutating the main sequence of the coding region of the fourth mRNA sequence; calculating a first score for each of the at least one sixth mRNA sequence; and The sixth mRNA sequence with the largest first score is determined as the second mRNA sequence.
9. The method according to claim 8, wherein Determining the sixth mRNA sequence with the largest first score as the second mRNA sequence comprises: In response to the codon adaptability index of the sixth mRNA sequence with the largest first score being greater than a threshold, the sixth mRNA sequence is determined as the second mRNA sequence, wherein the threshold is determined based on the codon adaptability index of the first mRNA sequence.
10. A device for optimizing messenger ribonucleotide (mRNA) sequences, comprising: an acquisition unit configured to acquire a first mRNA sequence for synthesizing a target protein, wherein the first mRNA sequence includes a 5' untranslated region sequence and a coding region sequence; and A processing unit is configured to adjust the 5' untranslated region sequence and the coding region sequence with the goal of maximizing the first score of the first mRNA sequence to obtain an optimized second mRNA sequence for synthesizing the target protein, comprising: Splitting the 5' untranslated region sequence and the coding region sequence into a translation initiation region sequence and a main coding region sequence, wherein the translation initiation region sequence includes the 5' untranslated region sequence and a preset number of nucleotides in the coding region sequence close to the 5' untranslated region sequence, and the main coding region sequence includes nucleotides in the coding region sequence that are not included in the translation initiation region sequence; With the goal of maximizing the translation initiation efficiency of the first mRNA sequence, adjusting the translation initiation region sequence to obtain a fourth mRNA sequence; and With the goal of maximizing the first score of the fourth mRNA sequence, adjusting the main sequence of the coding region of the fourth mRNA sequence to obtain the second mRNA sequence, The first score reflects the following indicators of the first mRNA sequence: translation initiation efficiency, codon adaptability index and minimum free energy. The translation initiation efficiency is used to measure the initiation efficiency of the translation process of ribosomes on mRNA molecules.
11. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 9.
13. A computer program product comprising computer program instructions, wherein: When the computer program instructions are executed by a processor, the method of any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Method, device and equipment for optimizing 5'untranslated region sequence of messenger ribonucleic acid
CN116168764A