Method, device, equipment and medium for determining mRNA sequence
By constructing DFA maps and optimizing the scores of adjacent codons, the problem of insufficient translation efficiency and stability in mRNA sequence design in the prior art is solved, and the effectiveness of mRNA vaccines and drugs is improved.
Patent Information
- Application Number
- CN202410168912.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-05
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2044-02-05
AI Technical Summary
When optimizing mRNA sequences, the prior art ignores the coupling effect of adjacent codons, resulting in the inability to effectively improve translation efficiency and stability, and cannot meet the actual needs of mRNA vaccines and drugs.
By constructing a DFA map of the finite automatic state machine, based on the free energy of the path and the score of adjacent codons, search for the target path, optimize the mRNA sequence design, consider the coupling effect of adjacent codons, and improve translation efficiency and stability.
The translation efficiency and stability of mRNA sequences in designated host cells were achieved, the expression of target proteins was increased, and the effects of mRNA vaccines and drugs were optimized.
Smart Images

Figure CN118038986B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the field of biological computing and other technical fields, and in particular to a method, device, equipment and medium for determining an mRNA sequence. Background Art
[0002] Optimizing the translation efficiency and stability of mRNA (Messenger RiboNucleic Acid) sequences is a key optimization target in fields such as mRNA and DNA (DeoxyriboNucleic Acid) drug development and genetic engineering, and has a wide range of applications. Translation efficiency refers to the rate at which an mRNA sequence can produce protein, while stability refers to the period over which an mRNA sequence can continuously translate protein. Together, these two factors determine the amount of protein an mRNA sequence can produce and ultimately influence the effectiveness of the corresponding mRNA vaccine, drug, or therapy. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device and medium for determining an mRNA sequence.
[0004] According to one aspect of the present disclosure, a method for determining an mRNA sequence is provided, comprising:
[0005] Obtaining a deterministic finite element automatic state machine (DFA) graph of the amino acid sequence of the target protein; wherein each path in the DFA graph is used to represent a candidate mRNA sequence translated into the target protein, and the candidate mRNA sequence consists of a codon for each amino acid in the amino acid sequence;
[0006] Searching for a target path from the DFA graph based on the free energy of each path in the DFA graph and the scores of adjacent codons in each of the paths; wherein the scores are used to evaluate the contribution of the adjacent codons to the translation efficiency and stability of the candidate mRNA sequence in the path;
[0007] The candidate mRNA sequence characterized by the target pathway is taken as the expected mRNA sequence.
[0008] According to another aspect of the present disclosure, there is provided a device for determining an mRNA sequence, comprising:
[0009] an acquisition module, configured to acquire a deterministic finite element automatic state machine (DFA) graph of the amino acid sequence of the target protein; wherein each path in the DFA graph is used to represent a candidate mRNA sequence translated into the target protein, the candidate mRNA sequence consisting of a codon for each amino acid in the amino acid sequence;
[0010] a search module for searching a target path from the DFA graph based on the free energy of each path in the DFA graph and the scores of adjacent codons in each of the paths; wherein the scores are used to evaluate the contribution of the adjacent codons to the translation efficiency and stability of the candidate mRNA sequence in the path;
[0011] The processing module is used to take the candidate mRNA sequence represented by the target pathway as the expected mRNA sequence.
[0012] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0013] at least one processor; and
[0014] a memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method proposed in the above aspect of the present disclosure.
[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium of computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method proposed in the above aspect of the present disclosure.
[0017] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the method proposed in the above aspect of the present disclosure when executed by a processor.
[0018] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0020] Figure 1 This is a schematic diagram of the process of determining the mRNA sequence provided in Example 1 of the present disclosure;
[0021] Figure 2 This is a schematic diagram of the process of determining the mRNA sequence provided in Example 2 of the present disclosure;
[0022] Figure 3 This is a schematic diagram of the process of determining the mRNA sequence provided in Example 3 of the present disclosure;
[0023] Figure 4 This is a schematic diagram of the process of determining the mRNA sequence provided in Example 4 of the present disclosure;
[0024] Figure 5 A schematic diagram of a DFA subgraph of amino acids provided in an embodiment of the present disclosure;
[0025] Figure 6 A schematic diagram of a DFA diagram of an amino acid sequence provided in an embodiment of the present disclosure;
[0026] Figure 7 A schematic diagram of an amino acid provided in an embodiment of the present disclosure forming two incompatible DFAs with the amino acid on the left and the amino acid on the right;
[0027] Figure 8 This is a schematic diagram of the process of the device for determining the mRNA sequence provided in Example 5 of the present disclosure;
[0028] Figure 9 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0029] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0030] Optimizing the translation efficiency and stability of mRNA sequences is a key optimization target in fields such as mRNA and DNA drug development, genetic engineering, and has a wide range of applications. Translation efficiency refers to the rate at which an mRNA sequence can produce protein, which is related to the codon composition of the mRNA sequence. Stability refers to the period over which an mRNA sequence can continuously translate protein, which is related to the folding free energy of the mRNA sequence. Together, these two factors determine the amount of protein an mRNA sequence can produce and ultimately influence the actual effectiveness of the corresponding mRNA vaccine, drug, or therapy.
[0031] The design of the coding region of the mRNA sequence is of great significance for optimizing the expression of genetic information and the synthesis of proteins. When designing and optimizing the coding region of the mRNA sequence, two aspects are mainly considered: optimizing the usage frequency of single codons and reducing the MFE (Minimum Free Energy) of the mRNA sequence. Among them, for the optimization of the usage frequency of single codons, high-frequency codons can be selected to recode the target gene according to the usage preference of codons commonly used in host cells to improve translation efficiency and promote the expression of target proteins. Among them, the commonly used codon frequency optimization index is CAI (Codon Adaptation Index), where the higher the CAI value, the more host-preferred codons the mRNA sequence uses.
[0032] CAI is closely related to the translation efficiency of the mRNA sequence, while MFE is more closely related to the stability of the mRNA sequence. Optimizing MFE can improve the stability of the mRNA sequence and extend its half-life, thereby allowing the target protein to be translated continuously for a longer period of time and increasing overall protein production, which is crucial for the actual effectiveness of biological products such as vaccines.
[0033] In the field of MFE optimization, LinearDesign is a representative method that can efficiently identify the mRNA sequence with the optimal MFE for a given target protein amino acid sequence. LinearDesign also supports the combined optimization of CAI and MFE.
[0034] Currently, coding region design of mRNA sequences primarily focuses on optimizing the frequency of single codon usage. However, experimental evidence shows that the combined frequency of adjacent codons (or adjacent dicodons or codon pairs) significantly and often more significantly influences the translation efficiency of mRNA sequences. Single codon frequency metrics (such as CAI) cannot reflect the frequency of dicodon usage. Therefore, algorithms and tools based on single codon optimization are unable to effectively optimize adjacent codons (or adjacent dicodons or codon pairs).
[0035] Current single codon usage frequency and MFE joint optimization tools also face the same problem. The following uses LinearDesign as an example to analyze its shortcomings:
[0036] 1. Using CAI as the optimization target for the usage frequency of a single codon ignores the coupling effect of adjacent codons on the translation efficiency of the mRNA sequence.
[0037] 2. Methodologically, based on the context-free assumption, a fixed weight is set for each edge of the DFA (Deterministic Finite Automaton). However, the context-free assumption conflicts with the coupling effect of adjacent codons, and the practice of setting fixed weights cannot support dicodon optimization.
[0038] In response to at least one of the above-mentioned problems, the present disclosure provides a method, apparatus, device and medium for determining an mRNA sequence.
[0039] The following describes the mRNA sequence determination method, apparatus, device, and medium of the present invention with reference to the accompanying drawings. Before describing the present invention in detail, for ease of understanding, the following common technical terms are introduced:
[0040] A state can be defined as a path between two nodes and its secondary structure. The path determines the scores of adjacent codons (or adjacent dicodons or codon pairs), while the secondary structure determines the free energy score (which can be derived from the folding rules and free energy parameters defined by the context-free grammar).
[0041] An atomic state is a state that has no lower-level sub-states.
[0042] It should be noted that since there can be different connection paths between two nodes, and each path may also have different secondary structures, there may be many different states between the two nodes. However, in path search, the focus is on the state with the highest score or the highest rating between the two nodes.
[0043] Different parts of a state's path and secondary structure can be defined as substates. The score or rating of a state can be derived by summing the scores or ratings of the substates, adding the additional free energy introduced by the spliced substate, and the scores of the adjacent codons. For example, when the two substates on either side of a split node are spliced together, a new pair of adjacent codons is formed: the last codon of the left substate and the first codon of the right substate. The score of this new pair of adjacent codons can be calculated and added to the score of the spliced state.
[0044] The edge codons between two nodes can be used to define the optimal state as the highest-scoring or highest-scoring state between two nodes, given the edge codons (i.e., the first and last codons) between the pair of nodes. The reason for considering the edge codons is that when joining substates, the scores of adjacent codons need to be calculated based on the edge codons.
[0045] It should be noted that the optimal state defined by the edge codons between two nodes can be obtained by concatenating the optimal substates defined by the edge codons split by the intermediate node between the two nodes. Different splitting nodes lead to different states of splicing. Therefore, different splitting nodes can be enumerated and the one with the highest score after splicing can be selected as the optimal state defined by the edge codons between the two nodes.
[0046] Figure 1 Schematic diagram of the process of determining the mRNA sequence provided in Example 1 of the present disclosure.
[0047] The embodiment of the present disclosure uses the method for determining an mRNA sequence as an example in which the method is configured in an mRNA sequence determination device. The mRNA sequence determination device can be applied to any electronic device so that the electronic device can perform the mRNA sequence determination function.
[0048] Among them, the electronic device can be any device with computing capabilities, such as a computer, mobile terminal, server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, and other hardware devices with various operating systems, touch screens and / or display screens.
[0049] like Figure 1 As shown, the method for determining the mRNA sequence may include the following steps:
[0050] Step S101: Obtain a DFA graph of the amino acid sequence of the protein to be tested.
[0051] In the embodiments of the present disclosure, the target protein can be any given or specified protein.
[0052] In the embodiments of the present disclosure, there is no restriction on the method of obtaining the amino acid sequence of the target protein. For example, the amino acid sequence can be obtained from an existing test set, or the amino acid sequence can be collected online, for example, by using web crawler technology, or the amino acid sequence can be provided by the user, etc. The present disclosure does not impose any restrictions on this.
[0053] In the embodiment of the present disclosure, the DFA graph of the amino acid sequence may include multiple paths, each path is used to represent an mRNA sequence translated into the target protein (referred to as a candidate mRNA sequence in the present disclosure), and each candidate mRNA sequence consists of a codon for each amino acid in the amino acid sequence.
[0054] Among them, an amino acid can correspond to one codon or multiple codons. For example, methionine can correspond to one codon (AUG), valine can correspond to four codons (GUA, GUC, GUG, GUU), leucine can correspond to six codons (UUA, UUG, CUA, CUC, CUG, CUU), and so on. Among them, A, U, G, and C are bases composed of nucleic acids. A represents adenine, U represents uracil, G represents guanine, and C represents cytosine.
[0055] Step S102 : searching for a target path in the DFA graph based on the free energy of each path in the DFA graph and the scores of adjacent codons in each path.
[0056] Among them, the score of adjacent codons (or adjacent dicodons, codon pairs) is used to evaluate the contribution of the adjacent codons to the translation efficiency and stability of the candidate mRNA sequence in the pathway where they are located.
[0057] Among them, there is a positive correlation between the score and the contribution, that is, the higher the score, the greater the contribution, and conversely, the lower the score, the smaller the contribution.
[0058] It should be noted that the adjacent codon score can be any scoring method that considers adjacent codon coupling information, and its specific calculation method and parameters can be changed according to the needs of the specific application scenario. For example, the adjacent codon score can be determined based on the number of occurrences of the adjacent codon in the host genome.
[0059] In the disclosed embodiment, a target path may be searched from the DFA graph based on the free energy of each path in the DFA graph and the scores of adjacent codons in each path in the DFA graph.
[0060] Step S103: taking the candidate mRNA sequence represented by the target pathway as the expected mRNA sequence.
[0061] In the embodiments of the present disclosure, the candidate mRNA sequence characterized by the target pathway can be used as the final desired mRNA sequence.
[0062] The method for determining the mRNA sequence of the embodiment of the present disclosure is based on the scores of adjacent codon pairs in each path in the DFA graph (used to evaluate the contribution of adjacent codons to the translation efficiency and stability of the candidate mRNA sequence in its path) and the free energy of each path. The target path where the desired mRNA sequence is located is searched from each path. This can improve the quality of the path search, thereby obtaining an mRNA sequence with relatively high translation efficiency and stability, improving the translation efficiency and stability of the mRNA sequence in a specified host cell, and increasing the expression of the target protein.
[0063] It should be noted that in the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information are all carried out with the user's consent, and are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0064] In order to clearly illustrate how any embodiment of the present disclosure searches for a target path from a DFA graph based on the free energy of each path in the DFA graph and the scores of adjacent codons in each path, the present disclosure also proposes a method for determining an mRNA sequence.
[0065] Figure 2 Schematic diagram of the flow chart of the method for determining the mRNA sequence provided in Example 2 of the present disclosure.
[0066] like Figure 2 As shown, the method for determining the mRNA sequence may include the following steps:
[0067] Step S201: Obtain a DFA graph of the amino acid sequence of the target protein.
[0068] Each path in the DFA graph is used to represent a candidate mRNA sequence translated into the target protein, and the candidate mRNA sequence consists of a codon for each amino acid in the amino acid sequence.
[0069] For explanation of step S201, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.
[0070] Step S202: storing the first best score of each state in a first hash table; wherein a state is represented by two nodes in a DFA graph and edge codons closest to the two nodes.
[0071] The first best score of each state is determined according to the best scores of the atomic states constituting the state.
[0072] Among them, the optimal score of the atomic state is determined based on the free energy (such as MFE) of at least one first subpath between two nodes in the atomic state and passing through the edge codon in the atomic state and the scores of the adjacent codons in each first subpath, and is determined according to the maximum score among the scores of each first subpath.
[0073] The atomic state may also include two nodes in the DFA graph and the edge codons closest to the two nodes.
[0074] The first subpath refers to a path between two nodes in an atomic state, and the first subpath passes through an edge codon in the atomic state.
[0075] The score is used to evaluate the contribution of adjacent codons to the translation efficiency and stability of the candidate mRNA sequence in its pathway.
[0076] Optionally, each state may include two nodes in the DFA graph (such as q i and q j , where i < j, and i and j are both positive integers), and q i The nearest edge codon is c i (where c i It refers to the DFA graph at position q i Right side and q i The most recent codon), and q j The nearest edge codon is c j (where c j It refers to the DFA graph at position q j Left and q j The most recent codon), and may also include q i and q j The subpath with the maximum score between the two (the subpath passes through c i and c j ) and the scores of the adjacent codons in the subpathway with the maximum score.
[0077] Correspondingly, the atomic state may also include two nodes in the DFA graph, the edge codons closest to the two nodes, and the free energy (such as MFE) of the subpath with the maximum score between the two nodes (referred to as the first subpath in this disclosure, which passes through the edge codons in the atomic state), and the scores of the adjacent codons in the first subpath with the maximum score.
[0078] In the embodiment of the present disclosure, each state may be composed of at least one atomic state, and the first best score of any state may be determined according to the best scores of the atomic states that constitute the state.
[0079] As an example, a state is composed of N atomic states, where N is a positive integer. When N=1, the best score of one atomic state can be used as the first best score of the state.
[0080] When N=2, assume that the two atomic states are atomic state 1 and atomic state 2, where atomic state 1 includes two nodes q in the DFA graph. i and q k (i<k, and i and k are both positive integers), atomic state 2 includes two nodes q in the DFA graph k and q j(k<j), then the atomic state 1 and q k The nearest edge codon and atom state 2 are related to q k The nearest edge codon is used as the adjacent codon, and based on the adjacent codon, the connection between atomic state 1 and atomic state 2 (i.e., the split node q k The score and free energy of the connection point (i.e., the split node q k The score of the connection is calculated based on the score and free energy of the connection, and the score of the connection is added to the best score of atomic state 1 and the best score of atomic state 2 to obtain the first best score of the above state.
[0081] Similarly, when N>2, the score and free energy of the connection between any two atomic states can be calculated. Based on the score and free energy of the connection between any two atomic states, the score of the connection can be calculated, so that the scores of all connections can be accumulated with the best score of each atomic state to obtain the first best score of the above state.
[0082] In the embodiment of the present disclosure, the first best score of each state may be stored in the first hash table.
[0083] In any embodiment of the present disclosure, the number of nodes between two nodes in each state or atomic state may be greater than a predetermined value to ensure that the mRNA folding has no sharp turns.
[0084] The predetermined value is a pre-set value, for example, the predetermined value may be 2, 3, 4, etc.
[0085] Step S203 , storing the first best reverse pointer of each state in the second hash table; wherein the first best reverse pointer is used to point to each atomic state that constitutes the state.
[0086] The first best reverse pointer of each state is used to point to the atomic states that constitute the state.
[0087] In the embodiment of the present disclosure, the first best reverse pointer of each state may be stored in the second hash table.
[0088] Step S204: Search for a target path in the DFA graph according to the first hash table and the second hash table.
[0089] In the embodiment of the present disclosure, a target path may be searched from the DFA graph according to the first hash table and the second hash table.
[0090] In any embodiment of the present disclosure, the target path search method may be, for example:
[0091] 1. Combine each state to obtain at least one combined state.
[0092] 2. Storing the second best score of each combined state in the first hash table, wherein the second best score is determined according to the marginal codons in each state constituting the combined state.
[0093] The second best score of each combined state is determined based on the marginal codons in the individual states that make up the combined state.
[0094] As an example, for any combination state, for each state that makes up the combination state, the edge codons at the split node in the adjacent state can be used as adjacent codons, and based on the adjacent codons, the score and free energy at the split node of the adjacent state are calculated. Based on the score and free energy at the split node, the score at the split node is calculated, and the score at the split node is added to the first best score of each state to obtain the second best score of the combination state.
[0095] 3. Store the second best reverse pointer of each combination state in the second hash table.
[0096] The second best reverse pointer of each combined state is used to point to each state constituting the combined state.
[0097] 4. Determine the target combined state from the combined states in the first hash table.
[0098] As an example, a target combination state can be determined from each combination state based on the second best score of each combination state in the first hash table, wherein the second best score of the target combination state is the highest and the target combination state includes the start node and the end node in the DFA graph.
[0099] Therefore, the second best combination state with the highest score is taken as the final target combination state, which includes the start node and the end node in the DFA graph, so that the expected mRNA sequence is determined based on the target combination state with the highest score, and the mRNA sequence with the best search score or quality can be realized, thereby further improving the translation efficiency and stability of the mRNA sequence in the specified host cell to increase the expression amount of the target protein.
[0100] 5. Determine the target path from the DFA graph based on the second best reverse pointer of the target combination state in the second hash table.
[0101] In the embodiment of the present disclosure, the target path can be determined from the DFA graph based on the second best reverse pointer backtracking of the target combination state in the second hash table, that is, the edges added by each expansion or combination state can be found based on the second best reverse pointer backtracking, and the target path can be obtained by splicing the edges according to their positions.
[0102] In summary, it is possible to achieve the optimization of the mRNA sequence by searching for the optimal path through state expansion (combining the various states) and based on the best score and the best reverse pointer of the expanded state.
[0103] Step S205 , taking the candidate mRNA sequence represented by the target pathway as the expected mRNA sequence.
[0104] For explanation of step S205, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.
[0105] The mRNA sequence determination method of the embodiment of the present disclosure can realize the search for the optimal path based on the best score and the best reverse pointer of each state, thereby improving the effectiveness and accuracy of the path search.
[0106] In order to clearly illustrate how the optimal score of the atomic state is calculated in any embodiment of the present disclosure, the present disclosure also proposes a method for determining the mRNA sequence.
[0107] Figure 3 Schematic diagram of the flow chart of the method for determining the mRNA sequence provided in Example 3 of the present disclosure.
[0108] like Figure 3 As shown, when the scores of adjacent codons include CPAI (Codon Pair Adaptation Index), the best score of any atomic state can be determined by the following steps:
[0109] Step S301 : Analyze the free energy of at least one first subpath between two nodes in an atomic state and passing through an edge codon in the atomic state to obtain the MFE of each first subpath.
[0110] In the embodiment of the present disclosure, for any atomic state, the free energy of each subpath (referred to as the first subpath in the present disclosure) between two nodes in the atomic state and passing through the edge codon in the atomic state can be analyzed to obtain the MFE of each first subpath.
[0111] Step S302 : determining the CPAI of each first subpath according to the first occurrence number of adjacent codons in each first subpath in the host genome.
[0112] In the disclosed embodiments, CPAI is used to evaluate the contribution of adjacent codons to the translation efficiency and stability of the candidate mRNA sequence in its pathway.
[0113] In the disclosed embodiments, the number of occurrences of any two adjacent codons in the host genome can be obtained by pre-statistical analysis.
[0114] In the embodiments of the present disclosure, for any first subpath, the CPAI of the first subpath can be calculated based on the number of occurrences of adjacent codons in the first subpath in the host genome (referred to as the first number of occurrences in the present disclosure), wherein the CPAI is positively correlated with the first number of occurrences.
[0115] In any embodiment of the present disclosure, the CPAI of the first sub-path may be calculated, for example, as follows:
[0116] 1. For any first subpath, obtain a codon pair consisting of codons for adjacent amino acids in an amino acid sequence; wherein the adjacent amino acids are the amino acids corresponding to the adjacent codons in the first subpath.
[0117] 2. Obtain the number of occurrences of each codon pair in the host genome (recorded as the second number of occurrences in this disclosure).
[0118] The number of occurrences of any codon pair in the host genome can also be counted in advance.
[0119] 3. Determine the CPAI of the first subpath based on the maximum number of the second occurrence numbers and the first occurrence number of the adjacent codons in the first subpath.
[0120] For example, the CPAI of the first subpath may be determined based on the ratio of the first number of occurrences of adjacent codons in the first subpath to the maximum number of occurrences.
[0121] As an example, the CPAI of the first sub-path may be: Where, l is the number of codons in the first subpath; c i c i+1 represents the adjacent codon pair at position i and i+1, F(c i c i+1 ) indicates codon pair c i c i+1 The number of occurrences in the host genome (i.e., the first occurrence); a i a i+1 represents the adjacent amino acid pair at position i and i+1, max{F(a i a i+1)} represents the adjacent amino acid pair a at position i and i+1 i a i+1 The number of occurrences of the most frequent codon pair in the host genome (ie, the largest number of occurrences in the second order).
[0122] Therefore, the score of the adjacent codons in the first subpathway (i.e., CPAI) is determined by comprehensively considering the number of occurrences of the adjacent codons in the host genome (i.e., the first number of occurrences) and the number of occurrences of the most frequent codon pairs of the corresponding adjacent amino acids in the host genome (i.e., the maximum number of the second number of occurrences). This can improve the rationality and accuracy of the adjacent codon score calculation.
[0123] Step S303: Determine the score of each first subpath according to the MFE and CPAI of each first subpath.
[0124] In the embodiment of the present disclosure, the score of each first sub-path may be determined according to the MFE and CPAI of each first sub-path.
[0125] In any embodiment of the present disclosure, the score of the first sub-path may be determined, for example, by:
[0126] 1. For any subpath, obtain or count the number of codons on the first subpath.
[0127] 2. According to the number of codons on the first subpath and the set weight balance factor, the MFE and CPAI of the first subpath are weighted to obtain the score of the first subpath.
[0128] As an example, the score of the first subpath may be: -MFE+λllogCPAI, where l is the number of codons in the first subpath and λ is a weight balancing factor.
[0129] Thus, the MFE and CPAI of the first subpath can be weighted based on the weighted summation method to obtain the score of the first subpath, thereby improving the effectiveness and rationality of the score calculation of the first subpath.
[0130] Step S304 : determining the best score of the atomic state according to the maximum score among the scores of the first subpaths.
[0131] In the embodiment of the present disclosure, the maximum score among the scores of the first sub-paths may be used as the best score of the atomic state.
[0132] It should be noted that the optimal score of an atomic state can be obtained by directly calculating its MFE and CPAI; for a state or a combination of states, its optimal score is not directly calculated, but is obtained by adding the added free energy of the splicing node or splicing point and the score of the adjacent codons on the basis of the MFE and CPAI of its substates.
[0133] The method for determining the mRNA sequence of the embodiment of the present disclosure comprehensively calculates the score of each first subpathway by combining the minimum folding free energy MFE and the codon pair adaptability index CPAI of each first subpathway, which can improve the accuracy and rationality of the score calculation, thereby improving the quality of the path search.
[0134] In order to clearly illustrate how to obtain the DFA diagram of the amino acid sequence of the target protein in any embodiment of the present disclosure, the present disclosure also proposes a method for determining the mRNA sequence.
[0135] Figure 4 Schematic diagram of the flow chart of the method for determining the mRNA sequence provided in Example 4 of the present disclosure.
[0136] like Figure 4 As shown, the method for determining the mRNA sequence may include the following steps:
[0137] Step S401 : obtaining a DFA subgraph of multiple amino acids in the amino acid sequence of the target protein.
[0138] The DFA subgraph includes at least one second subpath, and each second subpath is used to indicate a codon of an amino acid.
[0139] In an embodiment of the present disclosure, a DFA subgraph of each amino acid in the amino acid sequence of the target protein may be constructed, wherein the DFA subgraph includes at least one second subpath, and each second subpath is used to indicate a codon for the corresponding amino acid.
[0140] As an example, taking amino acids such as valine or serine as an example, the DFA subgraph of the amino acid can be as follows: Figure 5 As shown, the DFA subgraph of valine includes 4 second subpaths, and the DFA subgraph of valine is used to indicate 4 codons, namely: GUA, GUC, GUG, GUU; the DFA subgraph of serine includes 6 second subpaths, and the DFA subgraph of serine is used to indicate 6 codons, namely: UCA, UCC, UCG, UCU, AGC, AGU.
[0141] Step S402: Generate a DFA graph of the amino acid sequence based on the DFA subgraphs of the plurality of amino acids.
[0142] Each path in the DFA graph is used to represent a candidate mRNA sequence translated into the target protein, and the candidate mRNA sequence consists of a codon for each amino acid in the amino acid sequence.
[0143] In the disclosed embodiment, the DFA subgraphs of multiple amino acids may be sequentially connected head to tail according to the arrangement positions of the multiple amino acids in the amino acid sequence to form a DFA graph of the amino acid sequence.
[0144] As an example, taking an amino acid sequence including two amino acids (methionine and leucine) and a stop codon as an example, the DFA diagram of the amino acid sequence can be as follows: Figure 6 shown.
[0145] Step S403 : searching for a target path in the DFA graph based on the free energy of each path in the DFA graph and the scores of adjacent codons in each path.
[0146] Among them, the score is used to evaluate the contribution of adjacent codons to the translation efficiency and stability of the candidate mRNA sequence in its pathway.
[0147] Step S404: taking the candidate mRNA sequence represented by the target pathway as the expected mRNA sequence.
[0148] The method for determining the mRNA sequence of the embodiment of the present disclosure constructs a DFA subgraph for each amino acid, connects the DFA subgraphs of multiple amino acids head to tail in sequence to form a DFA graph of the amino acid sequence, which can improve the effectiveness of DFA graph acquisition.
[0149] In any embodiment of the present disclosure, the nucleotide sequence of mRNA can be designed to improve the translation efficiency and stability of the mRNA sequence in a specified host cell to increase the expression of the target protein. 3 ) solves the mRNA sequence with the optimal usage frequency of double codons and folding free energy within the time complexity, realizes the joint optimization of the translation efficiency and stability of the mRNA sequence, and ultimately improves the expression of protein.
[0150] As an example, the present disclosure proposes a design method for mRNA sequences that jointly optimizes the scores of MFE and adjacent codons (i.e., adjacent codon pairs) based on the LinearDesign method. Unlike LinearDesign, the present disclosure uses the scores of adjacent codons (i.e., adjacent codon pairs) as one of the optimization targets, wherein the so-called scores of adjacent codons (i.e., adjacent codon pairs) refer to a scoring system that evaluates the effects of the coupling of codons themselves and adjacent codons on the translation efficiency and half-life of the mRNA sequence. The original LinearDesign is not compatible with the optimization of the scores of adjacent codons in terms of algorithm design.
[0151] The input of the present disclosure can be an amino acid sequence composed of amino acids, and the output can be an mRNA sequence that encodes the target protein and has the best score. The score or rating of the present disclosure can be a fusion of the MFE and the scores of adjacent codons (i.e., adjacent codon pairs), and supports coordinating the relative weights of the two scores through a balance factor. Among them, the scores of adjacent codons (i.e., adjacent codon pairs) can be any scoring method that considers adjacent codon coupling information. Its specific calculation method and parameters can be changed according to the needs of specific scenarios.
[0152] For example, we can use LinearDesign to define the mRNA sequence optimization space, that is, use DFA to represent the encoding scheme of all possible mRNA sequences corresponding to the amino acid sequence of a target protein. DFA can be used to define all codon types for each amino acid, such as Figure 5 As shown in . In the DFA subgraph of amino acids, each edge is labeled by a certain nucleotide. The nucleotides on a path from the starting point to the end point together constitute a codon for the amino acid. By traversing all paths, all codons for the amino acid can be obtained. By connecting the DFA subgraphs of each amino acid in the amino acid sequence that constitutes a protein polypeptide chain, the DFA graph of the amino acid sequence of the entire target protein can be constructed, as shown in Figure 6 shown.
[0153] In the DFA graph of the amino acid sequence of a target protein, a path from the starting point to the end point constitutes an encoding scheme for the target protein. By traversing all paths in the DFA graph, all encoding schemes for the target protein can be obtained. The purpose of this disclosure is to determine the encoding scheme with the best score among all encoding schemes for the target protein and use it as the expected mRNA sequence.
[0154] For example, the score of O(N) can be calculated based on the set MFE and the mixed scoring scheme of adjacent codons (i.e., adjacent codon pairs). 3) within the time complexity of . The optimally scored mRNA sequence is found. The calculation of the minimum MFE for all paths between any two nodes in the DFA graph can be performed by intersecting the CFG (Context Free Grammar) with the DFA in LinearDesign. The present disclosure proposes a method that differs from LinearDesign in terms of a hybrid scoring scheme for MFE and adjacent codons (i.e., adjacent codon pairs), as well as a process for finding the optimal MFE and adjacent codon (i.e., adjacent codon pair) scores.
[0155] For example, in terms of mixing the scoring schemes of MFE and adjacent codons (i.e., adjacent codon pairs), the scoring of MFE and adjacent codons (i.e., adjacent codon pairs) can be mixed as shown in the following formula:
[0156] Score=Score MFE +λScore codon_pair ; (1)
[0157] Among them, Score MFE Indicates the score of MFE, Score codon_pair Represents the score of adjacent codons (i.e., adjacent codon pairs), and λ represents the weight balancing factor that coordinates the relative weights of the two. The original LinearDesign scoring scheme considers the integration of MFE scoring and single codon scoring. Among them, the single codon scoring does not consider the coupling effect of adjacent codons and can be regarded as Score codon_pair Therefore, the present disclosure generalizes the original LinearDesign optimization goal at the level of adjacent codon pairs.
[0158] Among them, in terms of searching for the mRNA sequence with the best combined score of MFE and adjacent codons (i.e., adjacent codon pairs), the present disclosure improves the original LinearDesign algorithm so that it can support the optimization of the scores of adjacent codons (i.e., adjacent codon pairs). When the original LinearDesign algorithm realizes the joint optimization of MFE and the score of a single codon, it adopts a method of assigning a fixed weight to each edge in the DFA graph. This is feasible for the optimization of the score of a single codon because each codon can find its own unique edge in the DFA graph, and its score is the weight of its unique edge. For example, in Figure 5In the example, there are two edges between nodes (5,0) and (6,0), belonging to codons UUA and UUG respectively. Therefore, the scores of UUA and UUG can be assigned to the two edges respectively. However, when considering the scores of adjacent codons (i.e., adjacent codon pairs), it is impossible to find an edge in the DFA graph that belongs to each adjacent codon pair, so the method of assigning fixed weights to the edges cannot be used. In this case, if a separate DFA is constructed for each pair of adjacent codons, the DFAs of adjacent codon pairs will be incompatible. Figure 7 shown.
[0159] To address the above issues, the present disclosure does not assign static weights to each edge of the DFA graph. Instead, during the search of the mRNA sequence, the scores of adjacent codons (i.e., adjacent codon pairs) are dynamically calculated based on an external scorer. Generally, the input of the scorer is a pair of adjacent codons, and the output is the score of the adjacent codons, with a time complexity of O(1).
[0160] In terms of the implementation of the optimal mRNA sequence search algorithm, the present disclosure can be based on the grid parsing strategy of LinearDesign, and expands the storage and reading methods of intermediate states to support O(N 3 ) time complexity, and a bottom-up or left-to-right search strategy for the optimal global score. The specific improvement to the storage and reading of intermediate states is to expand the storage of the optimal parsing sequence and its score between two nodes in the DFA graph to the storage of the optimal parsing sequence and its score between two nodes for each edge codon combination. When combining the parsing of two adjacent DFA subgraphs, it is necessary to enumerate the edge codon combinations at the junction and select the sub-parsing combination with the highest score under all edge codon combinations.
[0161] The following describes the application scheme of the present disclosure by taking the optimization of the MFE and the usage frequency of adjacent codons (i.e., adjacent codon pairs) of the mRNA vaccine as an example. The scheme can be implemented according to the following steps:
[0162] Step 1: Read in the amino acid sequence of the target protein.
[0163] Step 2: Based on the correspondence between amino acids and codons, construct an amino acid-level DFA subgraph for each amino acid, such as Figure 5 As shown in Figure 2. This DFA subgraph contains a starting point (0,0), an ending point (3,0), and several intermediate nodes. The edges between the nodes correspond to a nucleotide, and each path from the starting point to the ending point corresponds to a codon for that amino acid.
[0164] Step 3: Connect the DFA subgraphs of each amino acid head to tail to form the DFA graph of the target protein, such as Figure 6 In the DFA graph, each path from the start node of the first amino acid to the end node of the last amino acid corresponds to an encoding method of the target protein.
[0165] Step 4: Use a randomized context-free grammar that can calculate the free energy of each RNA secondary structure to parse the free energy of any path in the DFA graph and obtain the minimum folding free energy (MFE) of each path. For ease of description, this example uses a simple Nussinov-Jacobson model. The grammar of the Nussinov-Jacobson model is defined as follows:
[0166]
[0167]
[0168]
[0169]
[0170] Where S represents grammar, formula (2) is a bifurcation analysis, indicating that the secondary structure of an RNA can be split into two sequential structures; formula (3) is a pairing analysis, indicating that the secondary structure of an RNA is composed of the pairing of bases on both sides and a certain substructure in the middle sequence; formula (4) represents the analysis including unpaired bases, indicating that the secondary structure of an RNA can be generated by the combination of unpaired bases on the edge and another substructure, and can be composed of three unpaired bases; formula (5) indicates that an unpaired base can be composed of A, U, G or C.
[0171] Step 5: Use a scoring model for adjacent codon pairs to evaluate the adaptability of the codon pairs in any path in the DFA graph. In this example, the Codon Pair Adaptability Index (CPAI) can be used as the scoring model, which is defined as follows:
[0172]
[0173] Where L represents the number of amino acids in the amino acid sequence of the target protein, c i c i+1 represents the adjacent codon pair at position i and i+1, F(c i c i+1 ) represents the number of occurrences of the codon pair in the host genome, max{F(a i a i+1)} represents the number of occurrences of the coding mode of the adjacent amino acid pair at positions i and i+1 in the host genome by the most frequent codon pair.
[0174] Step 6: Set a joint score of MFE and codon pair adaptability, and use a weight balance factor to coordinate the relative weights of the two. For example, a scoring method can be defined by formula (7):
[0175] Score=-MFE+λLlogCPAI; (7)
[0176] Step 7: Initialize two hash tables, where the first hash table best is used to store the best score or best rating of each state, (S,q i ,q j ,c i ,c j )->score, representing node q in the DFA graph i To node q j Between, the edge codons are c i and c j The best score of the state when parsing with grammar S; the second hash table back is used to store the best reverse pointer of each state, and stores the pointer to the next sub-parsing state, including the state of the lower parsing, and the edges (nucleotides) added by the combination that do not belong to the lower parsing, (S,q i ,q j ,c i ,c j )->(sub_state_list,nuc_list).
[0177] Step 8: For any subsequence of length 3 in the DFA graph, use formula (4) and formula (5) to calculate the state (i.e., atomic state), MFE, and adjacent codon (i.e., adjacent codon pair) scores of each path of the starting amino acid, and set them to 0. That is, any subsequence of length 3 contains only one codon and no adjacent codons, so the score of the codon pair can be set to 0. In addition, the corresponding state can be stored in the best table, and the best reverse pointer of each state can be set to NULL (i.e., the atomic state has no substates) and stored in the back table.
[0178] Step 9: Calculate the best scoring state and the best reverse pointer for the longer sequence from bottom to top or from left to right, and store the results in the best table and back table. Each time the length is extended, according to the rules of dynamic programming, the combination of the best state in a small range can be searched for the best state in a larger range, and the best reverse pointer points to the best state in the small range used by the combination. When combining two small-range parsing, it is necessary to enumerate the edge codon combinations at their junction and select the sub-parsing combination with the highest score under all edge codon combinations. Continuously iterate the above process until the best parsing state between the start node and the end node in the DFA graph is found.
[0179] Step 10: Backtrack according to the reverse pointer of the global optimal state to find the edges added by each expansion, that is, obtain the corresponding nucleotides, and splice the obtained nucleotides according to the position to obtain the mRNA sequence with the best score.
[0180] In summary, it can make up for the shortcomings of existing products in the optimization of related indicators of adjacent codons, especially the joint optimization of MFE and adjacent codon indicators, which will help improve the translation efficiency of mRNA sequences in specified host cells, increase protein expression yield, and ultimately improve the efficacy of mRNA vaccines, protein replacement therapies and other gene therapies.
[0181] It should be noted that the present disclosure is applicable to various scenarios where the translation yield of a target protein needs to be increased in a specified host cell, including but not limited to the following scenarios:
[0182] 1. mRNA preventive vaccines: The present disclosure can be used to design mRNA vaccines that express antigens of various pathogens, thereby increasing the in vivo expression of antigens and strengthening the protective efficacy of vaccines.
[0183] 2. mRNA therapeutic cancer vaccines: The present disclosure can be used to design mRNA vaccines that express a patient's cancer neoantigens, increase the expression of neoantigens, activate and enhance the body's immune response, and promote the clearance of cancer cells.
[0184] 3. mRNA protein replacement therapy: This disclosure can be used to design mRNA drugs that express various therapeutic proteins, including proteins that are deficient in the body due to genetic defects and aging (for example, insulin and various metabolic enzymes), as well as protein drugs (for example, antibodies). These scenarios often require higher protein yields, and this disclosure will help improve mRNA translation efficiency and protein yield in such therapies.
[0185] 4. DNA gene therapy using viral or non-viral vectors: The present disclosure can be used to design gene coding regions in DNA gene therapy, improve the translation efficiency of the mRNA transcribed therefrom, and ultimately increase the translation yield of the target protein.
[0186] 5. Modification of Genetically Engineered Organisms: This disclosure can be used to modify genetically engineered organisms that produce various proteins, including various plants, animals, and microorganisms. By optimizing the codon composition of their transgenic segments, the expression of target proteins can be enhanced.
[0187] With the above Figures 1 to 7 Corresponding to the method for determining the mRNA sequence provided in the embodiment, the present disclosure further provides a device for determining the mRNA sequence. Figures 1 to 7 The method for determining the mRNA sequence provided in the embodiment corresponds to the method, so the implementation method of the mRNA sequence determination method is also applicable to the mRNA sequence determination device provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.
[0188] Figure 8 This is a schematic diagram of the structure of the device for determining the mRNA sequence provided in Example 5 of the present disclosure.
[0189] like Figure 8 As shown, the mRNA sequence determination device 800 may include: an acquisition module 810 , a search module 820 and a processing module 830 .
[0190] Among them, the acquisition module 810 is used to obtain a deterministic finite element automatic state machine (DFA) graph of the amino acid sequence of the target protein; wherein each path in the DFA graph is used to represent a candidate mRNA sequence translated into the target protein, and the candidate mRNA sequence consists of a codon for each amino acid in the amino acid sequence.
[0191] The search module 820 is used to search for a target path from the DFA graph based on the free energy of each path in the DFA graph and the scores of adjacent codons in each path; wherein the scores are used to evaluate the contribution of adjacent codons to the translation efficiency and stability of the candidate mRNA sequence in the path.
[0192] The processing module 830 is configured to use the candidate mRNA sequence represented by the target pathway as the desired mRNA sequence.
[0193] In a possible implementation of the embodiment of the present disclosure, the search module 820 is used to: store the first best score of each state in a first hash table; wherein the state is represented by two nodes in a DFA graph and the edge codons closest to the two nodes; the first best score is determined based on the best scores of the atomic states that constitute the state, and the best score of the atomic state is determined based on the free energy of at least one first subpath between two nodes in the atomic state and passing through the edge codons in the atomic state and the scores of adjacent codons in each first subpath, and is determined based on the maximum score among the scores of each first subpath; store the first best reverse pointer of each state in a second hash table; wherein the first best reverse pointer is used to point to the atomic states that constitute the state; and search the DFA graph for the target path based on the first hash table and the second hash table.
[0194] In a possible implementation of the embodiment of the present disclosure, the number of nodes between two nodes in a state or an atomic state is greater than a predetermined value.
[0195] In a possible implementation of the embodiment of the present disclosure, the search module 820 is used to: combine states to obtain at least one combined state; store the second best score of each combined state in a first hash table, and store the second best reverse pointer of each combined state in a second hash table; wherein the second best score is determined based on the edge codons in each state constituting the combined state; the second best reverse pointer is used to point to each state constituting the combined state; determine a target combined state from the combined states in the first hash table; and determine a target path from the DFA graph based on the second best reverse pointer of the target combined state in the second hash table.
[0196] In a possible implementation of an embodiment of the present disclosure, the search module 820 is used to: determine a target combination state from each combination state based on the second best score of each combination state in the first hash table; wherein the second best score of the target combination state is the highest, and the target combination state includes the start node and the end node in the DFA graph.
[0197] In one possible implementation of the disclosed embodiment, the scores of adjacent codons include a codon pair adaptability index (CPAI), and the best score of any atomic state is determined using the following module:
[0198] An analysis module is used to analyze the free energy of at least one first subpath between two nodes in any atomic state and passing through an edge codon in any atomic state to obtain a minimum folding free energy MFE of each first subpath;
[0199] A first determination module is used to determine the CPAI of each first subpathway according to a first occurrence number of adjacent codons in each first subpathway in the host genome;
[0200] a second determining module, configured to determine a score of each first subpath according to the MFE and CPAI of each first subpath;
[0201] The third determining module is configured to determine the best score of any atomic state according to the maximum score among the scores of the first subpaths.
[0202] In a possible implementation of the embodiment of the present disclosure, the second determination module is used to: obtain the number of codons on any first subpath for any subpath; weight the MFE and CPAI of any first subpath according to the number and the set weight balance factor to obtain the score of any first subpath.
[0203] In a possible implementation of the embodiments of the present disclosure, the first determination module is used to: for any first subpath, obtain a codon pair consisting of codons for adjacent amino acids in an amino acid sequence; wherein the adjacent amino acids are the amino acids corresponding to the adjacent codons in any first subpath; obtain the second number of occurrences of each codon pair in the host genome; and determine the CPAI of any first subpath based on the maximum number of each second number of occurrences and the first number of occurrences of adjacent codons in any first subpath.
[0204] In a possible implementation of the embodiment of the present disclosure, the acquisition module 810 is used to: obtain a DFA subgraph of multiple amino acids in an amino acid sequence; wherein the DFA subgraph includes at least one second subpath, and each second subpath is used to indicate a codon of an amino acid; and generate a DFA graph of the amino acid sequence based on the DFA subgraphs of multiple amino acids.
[0205] The mRNA sequence determination device of the embodiment of the present disclosure improves the quality of path search by searching for the target path where the desired mRNA sequence is located from each path based on the scores of adjacent codon pairs in each path in the DFA graph (used to evaluate the contribution of adjacent codons to the translation efficiency and stability of the candidate mRNA sequence in its path) and the free energy of each path. This can thereby obtain an mRNA sequence with relatively high translation efficiency and stability, improve the translation efficiency and stability of the mRNA sequence in a specified host cell, and increase the expression of the target protein.
[0206] In order to implement the above embodiments, the present disclosure also provides an electronic device, which may include at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the mRNA sequence determination method proposed in any of the above embodiments of the present disclosure.
[0207] In order to implement the above embodiments, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the mRNA sequence determination method proposed in any of the above embodiments of the present disclosure.
[0208] In order to implement the above embodiments, the present disclosure further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the method for determining the mRNA sequence proposed in any of the above embodiments of the present disclosure.
[0209] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0210] Figure 9 A schematic block diagram of an example electronic device that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0211] like Figure 9 As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 902 or a computer program loaded from a storage unit 908 into a RAM (Random Access Memory) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An I / O (Input / Output) interface 905 is also connected to the bus 904.
[0212] Multiple components in the electronic device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0213] The computing unit 901 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the mRNA sequence determination method described above. For example, in some embodiments, the mRNA sequence determination method described above can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the mRNA sequence determination method described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the above-mentioned mRNA sequence determination method in any other appropriate manner (for example, by means of firmware).
[0214] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0215] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0216] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0217] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0218] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0219] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship is established by computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and poor scalability of traditional physical hosts and VPS services. The server may also be a server in a distributed system or a server integrated with blockchain.
[0220] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0221] According to the technical solution of the embodiments of the present disclosure, by searching for the target path where the desired mRNA sequence is located from each path based on the scores of adjacent codon pairs in each path in the DFA graph (used to evaluate the contribution of adjacent codons to the translation efficiency and stability of the candidate mRNA sequence in its path) and the free energy of each path, the quality of the path search can be improved, thereby obtaining an mRNA sequence with relatively high translation efficiency and stability, improving the translation efficiency and stability of the mRNA sequence in the specified host cell, and increasing the expression of the target protein.
[0222] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions proposed in this disclosure can be achieved. This is not limited herein.
[0223] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for determining a messenger ribonucleic acid (mRNA) sequence, comprising: Obtaining a deterministic finite element automatic state machine (DFA) graph of the amino acid sequence of the target protein; wherein each path in the DFA graph is used to represent a candidate mRNA sequence translated into the target protein, and the candidate mRNA sequence consists of a codon for each amino acid in the amino acid sequence; Storing a first best score for each state in a first hash table; wherein the state is represented by two nodes in the DFA graph and edge codons closest to the two nodes; the first best score is determined based on the best scores of the atomic states constituting the state, the best score of the atomic state is determined based on the free energy of at least one first subpath between two nodes in the atomic state and passing through the edge codon in the atomic state and the scores of adjacent codons in each first subpath, and the score of each first subpath is determined based on the maximum score among the scores of the first subpaths; storing a first best reverse pointer for each state in a second hash table; wherein the first best reverse pointer is used to point to each atomic state constituting the state; Searching for a target path from the DFA graph according to the first hash table and the second hash table; wherein the score is used to evaluate the contribution of the adjacent codons to the translation efficiency and stability of the candidate mRNA sequence in the path; The candidate mRNA sequence characterized by the target pathway is taken as the expected mRNA sequence.
2. The method according to claim 1, wherein The number of nodes between two nodes in the state or atomic state is greater than a predetermined value.
3. The method according to claim 1, wherein The step of searching for a target path from the DFA graph according to the first hash table and the second hash table includes: Combining the states to obtain at least one combined state; Storing the second best score of each of the combined states in the first hash table, and storing the second best reverse pointer of each of the combined states in the second hash table; wherein the second best score is determined based on the marginal codons in each state constituting the combined state; and the second best reverse pointer is used to point to each state constituting the combined state; Determine a target combined state from the combined states in the first hash table; A target path is determined from the DFA graph according to a second best back pointer of the target combined state in the second hash table.
4. The method according to claim 3, wherein: The determining a target combined state from the combined states in the first hash table includes: determining the target combined state from each of the combined states according to the second best score of each of the combined states in the first hash table; The second best score of the target combination state is the highest, and the target combination state includes the start node and the end node in the DFA graph.
5. The method according to claim 1, wherein The scores of the neighboring codons include the codon pair adaptability index CPAI, the best score of any atom state, determined by the following steps: Analyzing the free energy of at least one first subpath between two nodes in any one atomic state and passing through an edge codon in any one atomic state to obtain a minimum folding free energy MFE of each first subpath; Determining the CPAI of each of the first subpathways according to the first occurrence number of adjacent codons in each of the first subpathways in the host genome; determining a score of each first subpath according to the MFE and CPAI of each first subpath; The best score of any one of the atomic states is determined according to the maximum score among the scores of the first subpaths.
6. The method according to claim 5, wherein: The determining, according to the MFE and CPAI of each first subpath, a score of each first subpath includes: For any subpath, obtaining the number of codons on any first subpath; The MFE and CPAI of any one of the first subpaths are weighted according to the quantity and a set weight balancing factor to obtain a score of the any one of the first subpaths.
7. The method according to claim 5, wherein: Determining the CPAI of each of the first subpathways according to the first number of occurrences of adjacent codons in each of the first subpathways in the host genome comprises: For any first subpathway, obtaining a codon pair consisting of codons for adjacent amino acids in the amino acid sequence; wherein the adjacent amino acids are amino acids corresponding to adjacent codons in any first subpathway; Obtaining the second occurrence number of each codon pair in the host genome; The CPAI of any first subpath is determined according to the maximum number of the second occurrence numbers and the first occurrence number of the adjacent codon in any first subpath.
8. The method according to any one of claims 1 to 7, wherein The method of obtaining a deterministic finite element automatic state machine (DFA) graph of the amino acid sequence of the target protein comprises: Obtaining a DFA subgraph of a plurality of amino acids in the amino acid sequence; wherein the DFA subgraph includes at least one second subpath, and each second subpath is used to indicate a codon for the amino acid; A DFA graph of the amino acid sequence is generated according to the DFA subgraphs of the multiple amino acids.
9. A device for determining a messenger ribonucleic acid (mRNA) sequence, comprising: an acquisition module, configured to acquire a deterministic finite element automatic state machine (DFA) graph of the amino acid sequence of the target protein; wherein each path in the DFA graph is used to represent a candidate mRNA sequence translated into the target protein, the candidate mRNA sequence consisting of a codon for each amino acid in the amino acid sequence; a search module for searching a target path from the DFA graph based on the free energy of each path in the DFA graph and the scores of adjacent codons in each of the paths; wherein the scores are used to evaluate the contribution of the adjacent codons to the translation efficiency and stability of the candidate mRNA sequence in the path; a processing module, configured to use the candidate mRNA sequence represented by the target pathway as the expected mRNA sequence; The search module is used to: Storing a first best score for each state in a first hash table; wherein the state is represented by two nodes in the DFA graph and edge codons closest to the two nodes; the first best score is determined based on the best scores of the atomic states constituting the state, the best score of the atomic state is determined based on the free energy of at least one first subpath between two nodes in the atomic state and passing through the edge codon in the atomic state and the scores of adjacent codons in each first subpath, and the score of each first subpath is determined based on the maximum score among the scores of the first subpaths; storing a first best reverse pointer for each state in a second hash table; wherein the first best reverse pointer is used to point to each atomic state constituting the state; A target path is searched from the DFA graph according to the first hash table and the second hash table.
10. The device according to claim 9, wherein The number of nodes between two nodes in the state or atomic state is greater than a predetermined value.
11. The device according to claim 9, wherein The search module is used to: Combining the states to obtain at least one combined state; Storing the second best score of each of the combined states in the first hash table, and storing the second best reverse pointer of each of the combined states in the second hash table; wherein the second best score is determined based on the marginal codons in each state constituting the combined state; and the second best reverse pointer is used to point to each state constituting the combined state; Determine a target combined state from the combined states in the first hash table; A target path is determined from the DFA graph according to a second best back pointer of the target combined state in the second hash table.
12. The device according to claim 11, wherein The search module is used to: determining the target combined state from each of the combined states according to the second best score of each of the combined states in the first hash table; The second best score of the target combination state is the highest, and the target combination state includes the start node and the end node in the DFA graph.
13. The device according to claim 9, wherein The scores of the neighboring codons include the codon pair adaptability index (CPAI), the best score of any atom state, determined using the following module: an analysis module, configured to analyze the free energy of at least one first subpath between two nodes in any one atomic state and passing through an edge codon in any one atomic state, to obtain a minimum folding free energy MFE of each first subpath; a first determining module, configured to determine the CPAI of each of the first subpathways according to a first number of occurrences of adjacent codons in each of the first subpathways in the host genome; a second determining module, configured to determine a score of each of the first subpaths according to the MFE and CPAI of each of the first subpaths; The third determining module is configured to determine the best score of any atomic state according to the maximum score among the scores of the first sub-paths.
14. The device according to claim 13, wherein The second determining module is configured to: For any subpath, obtaining the number of codons on any first subpath; The MFE and CPAI of any one of the first subpaths are weighted according to the quantity and a set weight balancing factor to obtain a score of the any one of the first subpaths.
15. The device according to claim 13, wherein The first determining module is configured to: For any first subpathway, obtaining a codon pair consisting of codons for adjacent amino acids in the amino acid sequence; wherein the adjacent amino acids are amino acids corresponding to adjacent codons in any first subpathway; Obtaining the second occurrence number of each codon pair in the host genome; The CPAI of any first subpath is determined according to the maximum number of the second occurrence numbers and the first occurrence number of the adjacent codon in any first subpath.
16. The device according to any one of claims 9 to 15, wherein: The acquisition module is used to: Obtaining a DFA subgraph of a plurality of amino acids in the amino acid sequence; wherein the DFA subgraph includes at least one second subpath, and each second subpath is used to indicate a codon for the amino acid; A DFA graph of the amino acid sequence is generated according to the DFA subgraphs of the multiple amino acids.
17. An electronic device, wherein: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.
19. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to claims 1-8.
Citation Information
Patent Citations
Method for optimizing genetic codons based on codon-pair usage frequency
CN106650307A
Systems and methods for sequence design
CN114974428A