MRNA sequence generation method and device, model training method and device and electronic equipment

By obtaining the embedded splicing of target protein sequences, species texts and translation efficiency, an mRNA sequence containing coding and untranslated regions is generated, which solves the problem of inaccurate generation and inability to meet regulatory needs in traditional methods, and achieves mRNA sequence generation that is more in line with application scenarios.

CN120412733APending Publication Date: 2025-08-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410149429.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-01
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Traditional mRNA sequence generation methods cannot generate accurate mRNA sequences, which cannot meet the regulatory needs of application scenarios, resulting in poor application effects.

Method used

By obtaining the target protein sequence, species text and translation efficiency, we determine the embedding splicing separately, and call the sequence generation model to generate mRNA sequences, including coding regions and non-translation areas, and adjust the translation efficiency to meet the needs of the application scenario.

Benefits of technology

Generate accurate and complete mRNA sequences, and the translation efficiency meets the regulatory needs of application scenarios and improves the application effect of mRNA sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412733A_ABST
    Figure CN120412733A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an mRNA sequence generation method and device, a model training method and device and electronic equipment. The method comprises the steps that a target protein sequence, a target species text corresponding to the target protein sequence and target translation efficiency corresponding to the target protein sequence are obtained; determining a first target embedding corresponding to the target species text, determining a second target embedding corresponding to the target protein sequence, and determining a third target embedding corresponding to the target translation efficiency; splicing the first target embedding, the second target embedding and the third target embedding to obtain a fourth target embedding; a sequence generation model is called to map the fourth target embedding, a target mRNA sequence is generated, and the target mRNA sequence comprises a first coding region, a first untranslated region and a second untranslated region; according to the embodiment of the invention, the accurate and complete target mRNA sequence can be obtained, the application effect of the target mRNA sequence can be improved, and the method can be widely applied to scenes such as cloud technology, artificial intelligence and smart medical treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and in particular, to an mRNA sequence generation method, a model training method, an apparatus, and an electronic device. Background Art

[0002] Currently, traditional mRNA sequence generation methods usually design coding regions based on the Minimum Free Energy (MFE) and the Codon Adaptation Index (CAI), and cannot obtain accurate mRNA sequences. Moreover, the generated mRNA sequences often cannot meet the regulatory requirements of application scenarios, resulting in poor application effects of the mRNA sequences. Summary of the Invention

[0003] The following is an overview of the subject matter described in detail in the present application. This overview is not intended to limit the scope of protection of the claims.

[0004] Embodiments of the present application provide an mRNA sequence generation method, a model training method, an apparatus, and an electronic device, which can obtain accurate and complete target mRNA sequences and improve the application effects of the target mRNA sequences.

[0005] On the one hand, an embodiment of the present application provides an mRNA sequence generation method, including:

[0006] Obtaining a target protein sequence, a target species text corresponding to the target protein sequence, and a target translation efficiency corresponding to the target protein sequence;

[0007] Determining a first target embedding corresponding to the target species text, determining a second target embedding corresponding to the target protein sequence, and determining a third target embedding corresponding to the target translation efficiency;

[0008] Concatenating the first target embedding, the second target embedding, and the third target embedding to obtain a fourth target embedding;

[0009] Invoking a sequence generation model to map the fourth target embedding to generate a target mRNA sequence, where the target mRNA sequence includes a first coding region, a first untranslated region upstream of the first coding region, and a second untranslated region downstream of the first coding region.

[0010] On the other hand, an embodiment of the present application provides a model training method, including:

[0011] Obtain a sample protein sequence, the sample species text corresponding to the sample protein sequence, the sample translation efficiency corresponding to the sample protein sequence, and the mRNA sequence tag corresponding to the sample protein sequence;

[0012] Determine a first sample embedding corresponding to the sample species text, determine a second sample embedding corresponding to the sample protein sequence, and determine a third sample embedding corresponding to the sample translation efficiency;

[0013] Concatenate the first sample embedding, the second sample embedding, and the third sample embedding to obtain a fourth sample embedding;

[0014] Call a sequence generation model to map the fourth sample embedding to generate a sample mRNA sequence, where the sample mRNA sequence includes a second coding region, a third untranslated region upstream of the second coding region, and a fourth untranslated region downstream of the second coding region;

[0015] Determine a model loss according to the sample mRNA sequence and the mRNA sequence tag, and train the sequence generation model according to the model loss.

[0016] On the other hand, an embodiment of the present application further provides an mRNA sequence generation device, including:

[0017] A first acquisition module, configured to acquire a target protein sequence, the target species text corresponding to the target protein sequence, and the target translation efficiency corresponding to the target protein sequence;

[0018] A first processing module, configured to determine a first target embedding corresponding to the target species text, determine a second target embedding corresponding to the target protein sequence, and determine a third target embedding corresponding to the target translation efficiency;

[0019] A second processing module, configured to concatenate the first target embedding, the second target embedding, and the third target embedding to obtain a fourth target embedding;

[0020] A first generation module, configured to call a sequence generation model to map the fourth target embedding to generate a target mRNA sequence, where the target mRNA sequence includes a first coding region, a first untranslated region upstream of the first coding region, and a second untranslated region downstream of the first coding region.

[0021] Further, the first generation module is specifically configured to:

[0022] Call a sequence generation model to map the fourth target embedding to generate a first candidate mRNA sequence;

[0023] Translate the first candidate mRNA sequence to obtain a first reference protein sequence, and determine a first edit distance between the first reference protein sequence and the target protein sequence;

[0024] When the first edit distance is less than or equal to an edit distance threshold, determine the first candidate mRNA sequence as the target mRNA sequence.

[0025] Furthermore, the above-mentioned first generation module is specifically configured to:

[0026] Determine the cumulative generation number of the first candidate mRNA sequence;

[0027] When the cumulative generation number is less than or equal to a number threshold, translate the first candidate mRNA sequence to obtain a first reference protein sequence.

[0028] Furthermore, the above-mentioned first generation module is specifically configured to:

[0029] When the cumulative generation number is greater than the number threshold, reduce the temperature parameter of the sequence generation model;

[0030] Call the sequence generation model after reducing the temperature parameter to map the fourth target embedding again to generate a second candidate mRNA sequence;

[0031] Translate the second candidate mRNA sequence to obtain a second reference protein sequence, and determine a second edit distance between the second reference protein sequence and the target protein sequence;

[0032] When the second edit distance is less than or equal to the edit distance threshold, determine the second candidate mRNA sequence as the target mRNA sequence.

[0033] Furthermore, the sequence generation model includes a masked self-attention layer and a feed-forward layer, and the fourth target embedding includes a plurality of first word embeddings arranged in sequence. The above-mentioned first generation module is specifically configured to:

[0034] Call the masked self-attention layer to perform self-attention processing on the fourth target embedding to obtain a first attention processing result;

[0035] Call the feed-forward layer to perform transformation processing on the first attention processing result to obtain a first hidden state, where the first hidden state includes first hidden components corresponding to each of the first word embeddings;

[0036] Generate a target word according to the first hidden component corresponding to the first word embedding arranged at the last position of the fourth target embedding;

[0037] Determine the second word embedding corresponding to the target word. After concatenating the second word embedding to the fourth target embedding, call the masked self-attention layer again to perform self-attention processing on the fourth target embedding until the generated target word is a preset end symbol, and concatenate the multiple generated target words in sequence to form a target mRNA sequence.

[0038] Further, the first processing module is specifically configured to:

[0039] Obtain a first description text for describing the target species text, and determine a first text embedding corresponding to the first description text;

[0040] Obtain a target evolutionary map, and determine a first map embedding corresponding to the target species text based on the target evolutionary map, where the edges in the target evolutionary map are evolutionary relationships between nodes, and the target species text is one of the nodes in the target evolutionary map;

[0041] Concatenate the first text embedding and the first map embedding to obtain a first target embedding corresponding to the target species text.

[0042] Further, the first processing module is specifically configured to:

[0043] Determine first node information of the target species text in the target evolutionary map;

[0044] Determine a first map embedding corresponding to the target species text based on the first node information.

[0045] Further, the first node information includes the first node depth of the target species text, the number of first nodes in the subtree connected by the target species text, and the number of second nodes of all nodes connected by the target species text. The first processing module is specifically configured to:

[0046] Embed the target species text to determine a second text embedding corresponding to the target species text;

[0047] Respectively determine first node embeddings corresponding to the first node depth, the first node number, and the second node number, and concatenate each first node embedding with the second text embedding to obtain multiple first concatenated embeddings;

[0048] Weight the multiple first concatenated embeddings to obtain a first map embedding corresponding to the target species text.

[0049] Further, the first processing module is specifically configured to:

[0050] Perform regression on each of the first node embeddings based on the first regression model to obtain the first weight corresponding to each of the first node embeddings;

[0051] Weight the multiple first concatenated embeddings based on the first weight to obtain the first atlas embedding corresponding to the target species text, where the first regression model is jointly trained with the sequence generation model.

[0052] Furthermore, the above first processing module is specifically configured to:

[0053] Obtain a target evolutionary atlas, where the edges in the target evolutionary atlas are the evolutionary relationships between nodes, the evolutionary relationships are determined by gene similarity, and the target species text is one of the nodes in the target evolutionary atlas;

[0054] Determine all the nodes connected to the target species text in the target evolutionary atlas as reference species texts;

[0055] Obtain a second description text for describing the reference species text, and determine a third text embedding corresponding to the second description text;

[0056] Determine the second weight corresponding to each reference species text according to the gene similarity corresponding to each reference species text;

[0057] Weight the multiple third text embeddings based on the second weight to obtain the first target embedding corresponding to the target species text.

[0058] Furthermore, the above second processing module is specifically configured to:

[0059] Obtain a target tissue text corresponding to the target protein sequence, and determine a fifth target embedding corresponding to the target tissue text;

[0060] Concatenate the first target embedding, the fifth target embedding, the second target embedding, and the third target embedding to obtain a fourth target embedding.

[0061] On the other hand, an embodiment of the present application provides a model training device, including:

[0062] A second acquisition module, configured to acquire a sample protein sequence, a sample species text corresponding to the sample protein sequence, a sample translation efficiency corresponding to the sample protein sequence, and an mRNA sequence tag corresponding to the sample protein sequence;

[0063] A third processing module, configured to determine a first sample embedding corresponding to the sample species text, determine a second sample embedding corresponding to the sample protein sequence, and determine a third sample embedding corresponding to the sample translation efficiency;

[0064] A fourth processing module, configured to splice the first sample embedding, the second sample embedding, and the third sample embedding to obtain a fourth sample embedding;

[0065] A second generation module, configured to call a sequence generation model to map the fourth sample embedding to generate a sample mRNA sequence, where the sample mRNA sequence includes a second coding region, a third untranslated region upstream of the second coding region, and a fourth untranslated region downstream of the second coding region;

[0066] A training module, configured to determine a model loss according to the sample mRNA sequence and the mRNA sequence tag, and train the sequence generation model according to the model loss.

[0067] Further, the above fourth processing module is specifically configured to:

[0068] Splice the first sample embedding, the second sample embedding, the third sample embedding, and the mRNA sequence tag to obtain a fourth sample embedding.

[0069] On the other hand, an embodiment of the present application further provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the above mRNA sequence generation method is implemented, or the above model training method is implemented.

[0070] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the above mRNA sequence generation method is implemented, or the above model training method is implemented.

[0071] On the other hand, an embodiment of the present application further provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device implements the above mRNA sequence generation method, or implements the above model training method.

[0072] The embodiments of the present application at least include the following beneficial effects: By obtaining the target protein sequence, the target species text, and the target translation efficiency, then respectively embedding the target protein sequence, the target species text, and the target translation efficiency, determining the first target embedding, the second target embedding, and the third target embedding in sequence, and then splicing the first target embedding, the second target embedding, and the third target embedding to obtain the fourth target embedding, and further mapping the fourth target embedding by calling the sequence generation model to generate the target mRNA sequence. In addition to including the first coding region, the target mRNA sequence also includes the first untranslated region and the second untranslated region, and an accurate and complete target mRNA sequence can be obtained. On this basis, since the fourth target embedding can simultaneously contain the information of the target species text, the target protein sequence, and the target translation efficiency, and the target translation efficiency can be adjusted based on regulatory requirements, the sequence generation model can make the translation efficiency of the target mRNA sequence meet the regulatory requirements of the application scenario when generating the target mRNA sequence corresponding to the target protein sequence of the target species, thereby improving the application effect of the target mRNA sequence and obtaining a target mRNA sequence that better conforms to the application scenario.

[0073] Other features and advantages of the present application will be described in the subsequent specification, and some of them will become obvious from the specification or be understood by implementing the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] The drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification. They are used together with the embodiments of the present application to explain the technical solutions of the present application, and do not constitute a limitation to the technical solutions of the present application.

[0075] Figure 1 It is a schematic diagram of an optional implementation environment provided by the embodiments of the present application;

[0076] Figure 2 It is a schematic diagram of an optional process flow of the mRNA sequence generation method provided by the embodiments of the present application;

[0077] Figure 3 It is a schematic diagram of an optional state of embedding the word segmentation result provided by the embodiments of the present application;

[0078] Figure 4 It is a schematic diagram of an optional structure of the linear model provided by the embodiments of the present application;

[0079] Figure 5 It is a schematic diagram of an optional structure for determining the first edit distance provided by the embodiments of the present application;

[0080] Figure 6An optional structural diagram of the second large language model provided by the embodiments of the present application;

[0081] Figure 7 An optional structural diagram of the sequence generation model provided by the embodiments of the present application;

[0082] Figure 8 An optional flowchart of the model training method provided by the embodiments of the present application;

[0083] Figure 9 An optional training flowchart of the sequence generation model provided by the embodiments of the present application;

[0084] Figure 10 An optional diagram of sequences that are not correctly generated under multiple translation efficiencies provided by the embodiments of the present application;

[0085] Figure 11 An optional column diagram of codon bias provided by the embodiments of the present application;

[0086] Figure 12 An optional column diagram of GC content under multiple translation efficiencies provided by the embodiments of the present application;

[0087] Figure 13 An optional diagram of the minimum free energy under multiple translation efficiencies provided by the embodiments of the present application;

[0088] Figure 14 An optional framework diagram of the mRNA sequence generation method provided by the embodiments of the present application;

[0089] Figure 15 An optional structural diagram of the mRNA sequence generation device provided by the embodiments of the present application;

[0090] Figure 16 An optional structural diagram of the model training device provided by the embodiments of the present application;

[0091] Figure 17 A partial structural block diagram of the terminal provided by the embodiments of the present application;

[0092] Figure 18 A partial structural block diagram of the server provided by the embodiments of the present application. Detailed implementation manners

[0093] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0094] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as target object attribute information or a set of attribute information, the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. Among them, the target object can be a user. In addition, when the embodiments of the present application need to obtain the target object attribute information, a separate permission or separate consent of the target object will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary target object-related data for the normal operation of the embodiments of the present application will be obtained.

[0095] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the functions of that module or unit.

[0096] To facilitate the understanding of the technical solutions provided by the embodiments of the present application, some key terms used in the embodiments of the present application are explained here first:

[0097] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, and can form a resource pool, which can be used as needed, flexibly and conveniently. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system support, which can only be achieved through cloud computing.

[0098] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0099] Machine Learning (ML for short) is an interdisciplinary subject that involves multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0100] Messenger RNA (mRNA) is a single-stranded ribonucleic acid molecule that plays a role in transmitting genetic information in organisms. The main function of mRNA is to carry the genetic information on DNA to the ribosomes in the cytoplasm, and this information is used to guide protein synthesis. Specifically, the base sequence on mRNA (composed of four bases: adenine (A), guanine (G), cytosine (C), and uracil (U)) is complementary to the base sequence on DNA, thereby transcribing the genetic information onto the mRNA molecule.

[0101] Minimum Free Energy (MFE) In the fields of biology and chemistry, the minimum free energy is usually used to predict the most stable structure of biological macromolecules (such as RNA and proteins) under a specific condition. In RNA folding prediction, the minimum free energy (MFE) model is a commonly used method. It predicts the most stable structure of an RNA molecule by finding the secondary structure with the minimum free energy. This method is based on thermodynamic parameters, such as the stacking energy of base pairing, the energy of loop structures, etc. By calculating the free energy values of different structures, the structure with the minimum free energy can be found, thereby predicting the actual structure of the RNA molecule in the organism.

[0102] Codon Adaptation Index (CAI) is an index used to measure the gene expression level and evaluate the relative expression levels of genes among different species. CAI is proposed based on the concept of codon usage bias, that is, in organisms, some codons are used more frequently during translation, thereby increasing protein production and translation efficiency.

[0103] Ribo-seq (Ribosome sequencing, also known as ribosome profiling) is a method based on high-throughput sequencing technology used to study the translation process, translation efficiency, and the regulatory mechanisms of protein synthesis. During the Ribo-seq experiment, specific drugs (such as cycloheximide, tyrosinase inhibitor, etc.) are added to pause the translation process in cells. Then, the mRNA is locally digested by nucleases, but the mRNA fragments covered by ribosomes, that is, ribosome-protected fragments (RPFs), are retained. Then, the results are obtained through RNA extraction, library construction, high-throughput sequencing, and data analysis.

[0104] RNA-seq (RNA sequencing) is a method based on high-throughput sequencing technology used to study the transcriptome, that is, all RNA molecules expressed at a specific time, in a specific tissue, or cell type. RNA-seq can provide quantitative and qualitative RNA information to help researchers understand gene expression patterns, differences, and regulatory mechanisms.

[0105] RPKM (Reads Per Kilobase Million) represents the number of reads from a gene per kilobase length per million reads.

[0106] Currently, traditional mRNA sequence generation methods usually design coding regions based on the Minimum Free Energy (MFE) and the Codon Adaptation Index (CAI), which cannot obtain accurate mRNA sequences. Moreover, the generated mRNA sequences often fail to meet the regulatory requirements of application scenarios, resulting in poor application effects of mRNA sequences.

[0107] Based on this, the embodiments of the present application provide an mRNA sequence generation method, a model training method, a device, and an electronic device, which can obtain accurate and complete target mRNA sequences and improve the application effects of target mRNA sequences.

[0108] Refer to Figure 1 , Figure 1 FIG. 10 is a schematic diagram of an optional implementation environment provided by the embodiments of the present application. The implementation environment includes a terminal 101 and a server 102. Among them, the terminal 101 and the server 102 are connected through a communication network.

[0109] Exemplarily, the server 102 can obtain the target protein sequence sent by the terminal, the target species text corresponding to the target protein sequence, and the target translation efficiency corresponding to the target protein sequence; determine the first target embedding corresponding to the target species text, determine the second target embedding corresponding to the target protein sequence, and determine the third target embedding corresponding to the target translation efficiency; splice the first target embedding, the second target embedding, and the third target embedding to obtain a fourth target embedding; call a sequence generation model to map the fourth target embedding to generate a target mRNA sequence, where the target mRNA sequence includes a first coding region, a first untranslated region located upstream of the first coding region, and a second untranslated region located downstream of the first coding region; the server 102 sends the target mRNA sequence to the terminal 101.

[0110] Server 102 obtains the target protein sequence, the target species text, and the target translation efficiency, then embeds the target protein sequence, the target species text, and the target translation efficiency respectively, determines the first target embedding, the second target embedding, and the third target embedding in sequence, then splices the first target embedding, the second target embedding, and the third target embedding to obtain the fourth target embedding, and further maps the fourth target embedding by calling a sequence generation model to generate a target mRNA sequence. In addition to including the first coding region, the target mRNA sequence also includes a first untranslated region and a second untranslated region, and an accurate and complete target mRNA sequence can be obtained. On this basis, since the fourth target embedding can simultaneously contain the information of the target species text, the target protein sequence, and the target translation efficiency, and the target translation efficiency can be adjusted based on regulatory requirements, the sequence generation model can make the translation efficiency of the target mRNA sequence meet the regulatory requirements of the application scenario when generating the target mRNA sequence corresponding to the target protein sequence of the target species, thereby improving the application effect of the target mRNA sequence and obtaining a target mRNA sequence that better conforms to the application scenario.

[0111] Server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In addition, server 102 can also be a node server in a blockchain network.

[0112] The terminal 101 can be a mobile phone, a computer, an intelligent voice interaction device, an intelligent home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present application do not make limitations here.

[0113] The method provided by the embodiments of the present application can be applied to various scenarios, including but not limited to scenarios such as cloud technology, artificial intelligence, and intelligent healthcare.

[0114] Refer to Figure 2 , Figure 2 FIG. is an optional flowchart of the mRNA sequence generation method provided by the embodiments of the present application. The mRNA sequence generation method can be executed by a server, or can also be executed by a terminal, or can also be executed by a server in cooperation with a terminal. The text type determination method includes but is not limited to the following steps 201 to 204.

[0115] Step 201: Obtain the target protein sequence, the target species text corresponding to the target protein sequence, and the target translation efficiency corresponding to the target protein sequence.

[0116] Among them, the target protein sequence can refer to a sequence containing multiple amino acids, and each amino acid has its corresponding amino acid character. In the target protein sequence, the corresponding amino acid can be represented by the amino acid character. For example, the amino acid character corresponding to alanine is A, the amino acid character corresponding to threonine is T, and other amino acids also have corresponding amino acid characters; the target species text can be the name of the target species. For example, the target species text can be "human", "cat", or "dog", etc. The target species text is used to indicate the target species that has the target protein sequence.

[0117] Among them, translation refers to the process of generating proteins from the genetic information on messenger ribonucleic acid (mRNA) molecules. The translation process is an important part of gene expression, and it occurs in ribosomes in the cytoplasm; in the fields of bioinformatics and molecular biology, translation efficiency (TE) refers to the efficiency of mRNA molecules being converted into proteins under given conditions. Translation efficiency is affected by various factors, including the number and activity of ribosomes, the stability and availability of mRNA, the supply of tRNA, the concentration of amino acids, codon bias, etc. In addition, translation efficiency is also affected by the cell growth stage, environmental conditions, and gene expression regulation mechanisms.

[0118] Based on this, the target protein corresponding to the target protein sequence can be generated by translating the target mRNA molecule. The target translation efficiency is used to characterize the efficiency of the target mRNA molecule being converted into the target protein. The target translation efficiency can be a real number greater than 0. The calculation formula for the target translation efficiency can be:

[0119]

[0120] where TE is the target translation efficiency, and RPKM Ribo-seq is the RPKM value of Ribo-seq, and RPKM RNA-seq is the RPKM value of RNA-seq.

[0121] Step 202: Determine the first target embedding corresponding to the target species text, determine the second target embedding corresponding to the target protein sequence, and determine the third target embedding corresponding to the target translation efficiency.

[0122] Among them, the first target embedding refers to the embedding information that can represent the text of the target species. Optionally, the first target embedding is implemented in the form of an embedding vector. The first target embedding can be determined in various ways. For example, according to a preset mapping relationship, the text of the target species can be mapped to the first target embedding. The text of the target species can be the name of the target species. Assuming that the texts of three target species are "human", "cat", and "dog" respectively, according to the preset mapping relationship, it can be determined that the first target embedding corresponding to "human" is [0, 0, 0, 1], the first target embedding corresponding to "cat" is [0, 0, 1, 0], and the first target embedding corresponding to "dog" is [0, 0, 1, 1]. Here, the dimension of the first target embedding is 4. Generally, the dimension of the first target embedding is large enough so that each species can be mapped to a unique first target embedding. The dimension of the first target embedding can be adjusted according to actual needs, and this is not limited in the embodiments of the present application.

[0123] For another example, the text of the target species can be input into the first embedding model, and based on the first embedding model, the text of the target species is mapped to obtain the first target embedding corresponding to the text of the target species. Therefore, the embodiments of the present application do not limit the specific determination method of the first target embedding.

[0124] Specifically, referring to Figure 3 , Figure 3 is an optional state diagram for embedding the word segmentation result provided by the embodiments of the present application.

[0125] The second target embedding refers to the embedding information that can represent the target protein sequence. Since the target protein sequence usually contains multiple amino acids, and each amino acid has its corresponding amino acid character, in the target protein sequence, the corresponding amino acid can be represented by the amino acid character, and then the target protein sequence can be segmented respectively according to the preset segmentation length to obtain multiple segmentation results, that is, multiple amino acid words. The segmentation length can be preset to 3. Each amino acid word can be a single amino acid character, a combination of two amino acid characters, or a combination of three amino acid characters, or each amino acid character can be directly used as an amino acid word; then, the word embeddings corresponding to each amino acid word can be determined. Specifically, the amino acid word can be input into the second embedding model, and based on the second embedding model, the amino acid word is mapped to obtain the word embedding corresponding to the amino acid word, and then the second target embedding is determined through the word embeddings corresponding to each amino acid word. That is, the second target embedding can be an embedding matrix containing multiple word embeddings. Optionally, the word embedding is implemented in the form of an embedding vector. The dimension of the second target embedding can be adjusted according to actual needs, and this is not limited in the embodiments of the present application. In addition, the embodiments of the present application do not limit the specific determination method of the second target embedding.

[0126] Among them, the third target embedding refers to the embedding information that can represent the target translation efficiency. Optionally, the third target embedding is implemented in the form of an embedding vector. Specifically, the target translation efficiency can be input into the third embedding model, and based on the third embedding model, the target translation efficiency is mapped to obtain the third target embedding corresponding to the target translation efficiency. Generally, the goal of the third embedding model is to learn a compact representation of the input data, so that the distances of similar inputs in the embedding space are closer.

[0127] For example, referring to Figure 4 , Figure 4 is an optional structural schematic diagram of the linear model provided by the embodiments of the present application.

[0128] Among them, the third embedding model can be a linear model. Specifically, the target translation efficiency can be input into the linear model, and based on the linear model, the target translation efficiency is mapped to obtain the third target embedding corresponding to the target translation efficiency. The linear model can include a fully connected layer of a deep learning framework, that is, the mathematical representation form of the linear model can be:

[0129] z = W * x + b

[0130] where x is the target translation efficiency, W is the weight matrix, b is the bias term, and z is the third target embedding. The third target embedding can be a multi-dimensional embedding vector. Here, the linear model needs to optimize the weight matrix and the bias term through a training process. Based on the trained linear model, mapping the target translation efficiency can improve the quality of the third target embedding.

[0131] Again, for example, the third embedding model can include an encoder for non-linear transformation. Specifically, the target translation efficiency can be input into the encoder, and based on the encoder, the target translation efficiency is mapped to obtain the third target embedding corresponding to the target translation efficiency. The dimension of the third target embedding can be adjusted according to actual needs, which is not limited in the embodiments of the present application. In addition, the embodiments of the present application do not limit the specific determination method of the third target embedding.

[0132] Specifically, after determining the first target embedding, the second target embedding, and the third target embedding, the first prompt embeddings corresponding to the first target embedding, the second target embedding, and the third target embedding can be obtained, and then the first target embedding, the second target embedding, and the third target embedding are respectively concatenated with the corresponding first prompt embeddings. The first prompt embedding is used to prompt the factor types of the first target embedding, the second target embedding, and the third target embedding. Specifically, the first prompt embedding can be an embedding vector obtained by embedding the first prompt. The first prompt corresponding to the first target embedding can be <species>, The first prompt corresponding to the second target embedding can be <protein>, The first prompt corresponding to the third target embedding can be <te>。

[0133] Step 203: Concatenate the first target embedding, the second target embedding, and the third target embedding to obtain a fourth target embedding.

[0134] Among them, since the first target embedding is determined by the target species text of the target species, the second target embedding is determined by the target protein sequence, the third target embedding is determined by the target translation efficiency, and the fourth target embedding is obtained by concatenating the first target embedding, the second target embedding, and the third target embedding, the fourth target embedding can contain the information of the target species text, the target protein sequence, and the target translation efficiency at the same time.

[0135] Step 204: Invoke a sequence generation model to map the fourth target embedding to generate a target mRNA sequence.

[0136] Among them, the target mRNA sequence includes a first coding region, a first untranslated region upstream of the first coding region, and a second untranslated region downstream of the first coding region.

[0137] Specifically, the target mRNA sequence is the full-length mRNA. Therefore, the target mRNA sequence will contain a 5'UTR, a coding sequence, and a 3'UTR arranged in sequence. The first coding region refers to the coding sequence (CDS), which is the sequence region on the mRNA molecule that encodes proteins. The CDS contains all the amino acid information of the protein and is stored in the form of codons. Each codon consists of 3 consecutive bases. The CDS starts from the start codon (usually AUG) and ends at the stop codon (such as UAA, UAG, or UGA).

[0138] In addition, the first untranslated region upstream of the first coding region refers to the 5'UTR, and the second untranslated region downstream of the first coding region refers to the 3'UTR. The 5'UTR and 3'UTR will be described in detail below.

[0139] Specifically, the UTR (Untranslated Region) refers to the sequence region on the mRNA molecule that does not encode proteins. The UTR plays an important role in gene expression and translation regulation. The UTR of mRNA is mainly divided into two parts: 5'UTR and 3'UTR. The 5'UTR is located at the 5' end of the mRNA molecule, between the cap structure at the 5' end and the start codon (usually AUG). Therefore, the 5'UTR is located upstream of the first coding region. The 5'UTR plays a key role in the translation initiation process because the 5'UTR contains the ribosome binding site and other regulatory elements, such as the Internal Ribosome Entry Site (IRES). In addition, specific sequences and secondary structures in the 5'UTR may also affect the stability and translation efficiency of mRNA. The 3'UTR is located at the 3' end of the mRNA molecule, between the stop codon (such as UAA, UAG or UGA) and the poly(A) tail at the 3' end. Therefore, the 3'UTR is located downstream of the first coding region. The 3'UTR plays an important role in mRNA stability, transport and translation regulation. Specific sequences in the 3'UTR can serve as binding sites for microRNA (miRNA) and RNA-binding proteins, thereby affecting mRNA degradation, stability and translation efficiency.

[0140] Based on this, by obtaining the target protein sequence, the target species text and the target translation efficiency, and then embedding the target protein sequence, the target species text and the target translation efficiency respectively, the first target embedding, the second target embedding and the third target embedding are determined in sequence. Then, the fourth target embedding is obtained by splicing the first target embedding, the second target embedding and the third target embedding. Furthermore, by calling the sequence generation model to map the fourth target embedding, a target mRNA sequence is generated. In addition to including the first coding region, the target mRNA sequence also includes the first untranslated region and the second untranslated region, and an accurate and complete target mRNA sequence can be obtained. On this basis, since the fourth target embedding can simultaneously contain the information of the target species text, the target protein sequence and the target translation efficiency, and the target translation efficiency can be adjusted based on regulatory requirements, the sequence generation model can make the translation efficiency of the target mRNA sequence meet the regulatory requirements of the application scenario when generating the target mRNA sequence corresponding to the target protein sequence of the target species, thereby improving the application effect of the target mRNA sequence and obtaining a target mRNA sequence that better conforms to the application scenario.

[0141] In a possible implementation, the call sequence generation model maps the fourth target embedding to generate a target mRNA sequence. Specifically, the call sequence generation model maps the fourth target embedding to generate a first candidate mRNA sequence; translates the first candidate mRNA sequence to obtain a first reference protein sequence, and determines the first edit distance between the first reference protein sequence and the target protein sequence; when the first edit distance is less than or equal to the edit distance threshold, the first candidate mRNA sequence is determined as the target mRNA sequence.

[0142] Among them, since the first target embedding is determined by the target species text of the target species, the second target embedding is determined by the target protein sequence, the third target embedding is determined by the target translation efficiency, and the fourth target embedding is obtained by splicing the first target embedding, the second target embedding and the third target embedding, so the fourth target embedding can simultaneously contain the information of the target species text, the target protein sequence and the target translation efficiency. Moreover, the target translation efficiency is adjustable and can be adjusted according to the regulation requirements of the application scenario. By mapping the fourth target embedding through the sequence generation model, a first candidate mRNA sequence corresponding to the fourth target embedding is generated, that is, when the sequence generation model generates the target mRNA sequence corresponding to the target protein sequence of the target species, the translation efficiency of the first candidate mRNA sequence can meet the regulation requirements of the application scenario.

[0143] Among them, the sequence generation model can be a trained neural network model. Since the performance of the sequence generation model is affected by aspects such as the quality of the training data, there is a situation where the first candidate mRNA sequence generated by the sequence generation model is inappropriate. The first candidate mRNA sequence being inappropriate specifically means that the first reference protein sequence translated from the first candidate mRNA sequence is different from the target protein sequence, or the deviation between the first reference protein sequence translated from the first candidate mRNA sequence and the target protein sequence is relatively large. Similarly, the first candidate mRNA sequence being appropriate specifically means that the first reference protein sequence translated from the first candidate mRNA sequence is the same as the target protein sequence, or the deviation between the first reference protein sequence translated from the first candidate mRNA sequence and the target protein sequence is relatively small.

[0144] Under normal circumstances, when the first candidate mRNA sequence is not suitable, a protein sequence that meets the requirements cannot be translated from the first candidate mRNA sequence, and the first candidate mRNA sequence needs to be eliminated; when the first candidate mRNA sequence is suitable, a protein sequence that meets the requirements can be translated from the first candidate mRNA sequence; therefore, after generating the first candidate mRNA sequence, it is necessary to perform a verification process on the first candidate mRNA sequence, eliminate inaccurate first candidate mRNA sequences, and determine the accurate first candidate mRNA sequence as the target mRNA sequence. The process of verifying the first candidate mRNA sequence will be described in detail below.

[0145] During the verification process, refer to Figure 5 , Figure 5 which is an optional structural schematic diagram for determining the first edit distance provided by the embodiments of the present application.

[0146] Among them, first translate the first candidate mRNA sequence to obtain a first reference protein sequence, and then the first edit distance between the first reference protein sequence and the target protein sequence can be calculated. The first edit distance refers to the minimum number of edit operations required to convert the first reference protein sequence into the target protein sequence. Each time an edit operation is performed, the edit operation count is incremented by one. In the first reference protein sequence, one edit operation can be: inserting an amino acid character, deleting an amino acid character, or replacing an amino acid character with another amino acid character;

[0147] Among them, the first edit distance is used to characterize the similarity between the first reference protein sequence and the target protein sequence. When the first edit distance is larger, the similarity between the first reference protein sequence and the target protein sequence is lower. On the contrary, when the first edit distance is smaller, the similarity between the first reference protein sequence and the target protein sequence is higher; in addition to calculating the first edit distance to determine the similarity between the first reference protein sequence and the target protein sequence, other methods can also be used to determine the similarity between the first reference protein sequence and the target protein sequence. The embodiments of the present application do not limit this here.

[0148] Based on this, when the first edit distance is less than or equal to the edit distance threshold, it can be considered that the similarity between the first reference protein sequence and the target protein sequence is relatively high. At this time, the first candidate mRNA sequence can be determined as the target mRNA sequence, which can ensure that the target mRNA sequence is suitable, and a protein sequence that meets the requirements can be translated from the target mRNA sequence.

[0149] Exemplarily, the edit distance threshold can be set to 0. In this case, when the first reference protein sequence and the target protein sequence need to be exactly the same, the first candidate mRNA sequence is determined as the target mRNA sequence. Due to the degeneracy of codons, there are cases where a target protein sequence can be translated from multiple different mRNA sequences. The degeneracy of codons refers to the phenomenon that the same amino acid has two or more codons. Different codons corresponding to the same amino acid can be called synonymous codons. It can be seen that the same protein sequence can be translated from multiple different mRNA sequences. Therefore, the sequence generation model can generate multiple first candidate mRNA sequences that can be determined as the target mRNA sequence, that is, the target mRNA sequence is uncertain. The edit distance threshold can be set according to the actual situation, and the embodiments of the present application do not limit this here.

[0150] Specifically, the target mRNA sequence can be determined in various ways. For example, an optional way to determine the target mRNA sequence is as follows: the sequence generation model can be repeatedly called to generate the first candidate mRNA sequence. Whenever the sequence generation model is called to generate a first candidate mRNA sequence, it is judged whether the first edit distance corresponding to the currently generated first candidate mRNA sequence is less than or equal to the edit distance threshold. When the first edit distance is greater than the edit distance threshold, the sequence generation model is called again to generate a new first candidate mRNA sequence. When the first edit distance is less than or equal to the edit distance threshold, the currently generated first candidate mRNA sequence is determined as the target mRNA sequence, and there is no need to call the sequence generation model again, which is equivalent to the sequence generation model generating any first candidate mRNA sequence that meets the protein translation requirements.

[0151] Another example, another optional way to determine the target mRNA sequence is as follows: the sequence generation model is repeatedly called multiple times to generate multiple first candidate mRNA sequences. Among the generated multiple first candidate mRNA sequences, it is judged whether the first edit distance corresponding to each first candidate mRNA sequence is less than or equal to the edit distance threshold. When the first edit distance is greater than the edit distance threshold, the corresponding first candidate mRNA sequence is excluded. When the first edit distance is less than or equal to the edit distance threshold, the corresponding first candidate mRNA sequence is retained. Finally, all the retained first candidate mRNA sequences are used as the target mRNA sequences.

[0152] In a possible implementation, the first candidate mRNA sequence is determined as the target mRNA sequence. Specifically, the first regulatory sequence and the second regulatory sequence are obtained. When it is determined that the first untranslated region in the first candidate mRNA sequence contains the first regulatory sequence and the second untranslated region in the first candidate mRNA sequence contains the second regulatory sequence, the first candidate mRNA sequence is determined as the target mRNA sequence; otherwise, the sequence generation model is called again to generate the first candidate mRNA sequence. Based on this, both the first regulatory sequence and the second regulatory sequence can be sequences obtained by splicing multiple specific base characters, and can customize the generation of specific first and second untranslated regions, so that the target mRNA sequence can meet the specific regulatory requirements of the application scenario.

[0153] In a possible implementation, the first candidate mRNA sequence is translated to obtain the first reference protein sequence. Specifically, the cumulative generation number of the first candidate mRNA sequence is determined. When the cumulative generation number is less than or equal to the number threshold, the first candidate mRNA sequence is translated to obtain the first reference protein sequence.

[0154] Among them, the sequence generation model can be repeatedly called to generate the first candidate mRNA sequence. Whenever the sequence generation model is called to generate a first candidate mRNA sequence, the cumulative generation number of the first candidate mRNA sequence is incremented by one. Since the first candidate mRNA sequence generated by the sequence generation model is random, the sequence generation model is called again to generate a new first candidate mRNA sequence only when the first reference protein sequence translated from the first candidate mRNA sequence is inappropriate. Therefore, the larger the cumulative generation number, the more difficult it is considered that the sequence generation model is to generate a suitable first candidate mRNA sequence under the current model parameters.

[0155] Based on this, after setting a suitable number threshold, according to the comparison result between the current cumulative generation number and the number threshold, it is determined whether the first candidate mRNA sequence needs to be translated. When the cumulative generation number is less than or equal to the number threshold, it means that the sequence generation model can quickly generate a suitable first candidate mRNA sequence under the current model parameters, and at this time, the currently generated first candidate mRNA sequence needs to be translated.

[0156] In addition, when the cumulative generation count is greater than the count threshold, it means that the sequence generation model cannot quickly generate a suitable first candidate mRNA sequence under the current model parameters. It can be considered that the currently generated first candidate mRNA sequence is inappropriate, and there is no need to spend time translating the currently generated first candidate mRNA sequence. After the model parameters of the sequence generation model are adjusted later, the cumulative generation count is reset, and then a new first candidate mRNA sequence is generated based on the adjusted sequence generation model, which can improve the efficiency of generating a suitable first candidate mRNA sequence, thereby improving the determination efficiency of the target mRNA sequence.

[0157] Specifically, the count threshold can be set according to the actual situation. For example, the count threshold can be set to 10. When the cumulative generation count is less than or equal to 10, the first candidate mRNA sequence is translated to obtain the first reference protein sequence. On the contrary, when the cumulative generation count is greater than 10, there is no need to translate the first candidate mRNA sequence. The specific value of the count threshold in the embodiments of the present application is not limited.

[0158] In a possible implementation manner, the mRNA sequence generation method further includes: when the cumulative generation count is greater than the count threshold, reducing the temperature parameter of the sequence generation model; calling the sequence generation model with the reduced temperature parameter to map the fourth target embedding again to generate a second candidate mRNA sequence; translating the second candidate mRNA sequence to obtain a second reference protein sequence, and determining the second edit distance between the second reference protein sequence and the target protein sequence; when the second edit distance is less than or equal to the edit distance threshold, determining the second candidate mRNA sequence as the target mRNA sequence.

[0159] Among them, the temperature parameter of the sequence generation model can be a hyperparameter. The hyperparameter is not obtained through model training and is usually assigned based on existing experience to configure the hyperparameter for the sequence generation model. The first candidate mRNA sequence can include multiple words. The sequence generation model can sequentially generate each word of the first candidate mRNA sequence. The words generated by the sequence generation model are sampled from the vocabulary based on a probability distribution, and the temperature parameter is used to adjust the word probability distribution.

[0160] Based on this, when the temperature parameter of the sequence generation model is higher, the probability distribution is smoother, which is equivalent to smoothing the initial probabilities of each word in the vocabulary, increasing the generation probability of words with lower initial probabilities, increasing the randomness of the first candidate mRNA sequence, making the first candidate mRNA sequence more diverse, but more likely to generate inappropriate first candidate mRNA sequences; therefore, when the cumulative number of generated sequences is greater than the number threshold, it means that the sequence generation model cannot quickly generate appropriate first candidate mRNA sequences under the current model parameters, and it can be considered that the currently generated first candidate mRNA sequence is inappropriate, and there is no need to spend time translating the currently generated first candidate mRNA sequence. At this time, the temperature parameter of the sequence generation model can be reduced;

[0161] Then, the sequence generation model with the reduced temperature parameter is called to map the fourth target embedding again to generate a second candidate mRNA sequence. When the temperature parameter of the sequence generation model is lower, the probability distribution is sharper, reducing the generation probability of words with lower initial probabilities, making the sequence generation model usually generate words with higher initial probabilities, reducing the randomness of the second candidate mRNA sequence, and being more likely to generate appropriate second candidate mRNA sequences. The appropriateness of the second candidate mRNA sequence specifically means that the similarity between the second reference protein sequence translated from the second candidate mRNA sequence and the target protein sequence is relatively high. Here, the second edit distance between the second reference protein sequence and the target protein sequence is calculated, and the second edit distance is used to represent the similarity between the second reference protein sequence and the target protein sequence. Similar to the judgment rule of the first edit distance, when the second edit distance is less than or equal to the edit distance threshold, the second candidate mRNA sequence can be determined as the target mRNA sequence.

[0162] Specifically, the temperature parameter can be decreased according to a preset temperature adjustment strategy. For example, the temperature adjustment strategy can be configured such that whenever the temperature parameter needs to be decreased, the temperature parameter is gradually decreased according to a preset temperature adjustment amplitude. Exemplarily, assuming that the initial temperature parameter is preset to 1, when the temperature parameter needs to be decreased for the first to ninth times, the temperature adjustment amplitude can be set to 0.1, that is, when the temperature parameter needs to be decreased for the first to ninth times, the temperature parameter is decreased by 0.1 each time. Therefore, after the ninth decrease of the temperature parameter, the temperature parameter is adjusted to 0.1. Since the temperature parameter will not be set to 0, when the temperature parameter needs to be decreased for the tenth time, the temperature adjustment amplitude can be set to 0.09. Therefore, after the tenth decrease of the temperature parameter, the temperature parameter is adjusted to 0.01. After each adjustment of the temperature parameter, it is necessary to call the sequence generation model after decreasing the temperature parameter to generate a second candidate mRNA sequence. The temperature parameter can be continuously adjusted in a similar manner until the second candidate mRNA sequence is determined to be the target mRNA sequence, or until the number of times of adjusting the temperature parameter is equal to a preset number threshold; after the sequence generation model no longer adjusts the temperature parameter, when the cumulative generation number is greater than the number threshold again, the second candidate mRNA sequence corresponding to the smallest second edit distance can be determined as the target mRNA sequence.

[0163] For another example, the temperature adjustment strategy can be configured to determine the current temperature parameter according to a temperature adjustment function, and the temperature adjustment function can be as follows:

[0164] T x = T0 * e -kx

[0165] where T0 is the initial temperature parameter, x is the number of times of decreasing the temperature parameter, T x is the current temperature parameter after x times of decreasing the temperature parameter, k is the attenuation parameter, k > 0. It can be seen that as the number of times of decrease x increases, the temperature parameter T x will gradually decrease, but the adjustment amplitude ΔT of the temperature parameter will decrease as the number of times of decrease x increases. The adjustment amplitude ΔT is the absolute value of the difference between the current temperature parameter and the temperature parameter of the previous round. Exemplarily, assuming that the current temperature parameter is T2 and the temperature parameter of the previous round is T1, the corresponding adjustment amplitude ΔT = |T2 - T1|.

[0166] In addition to the above two temperature adjustment strategies, other temperature adjustment strategies can also be used to decrease the temperature parameter, which is not limited in the embodiments of the present application.

[0167] In a possible implementation, the sequence generation model includes a masked self-attention layer and a feed-forward layer. The fourth target embedding includes a plurality of first word embeddings arranged in sequence. The sequence generation model is called to map the fourth target embedding to generate a target mRNA sequence. Specifically, the masked self-attention layer can be called to perform self-attention processing on the fourth target embedding to obtain a first attention processing result; the feed-forward layer is called to perform transformation processing on the first attention processing result to obtain a first hidden state, where the first hidden state includes first hidden components corresponding to each of the first word embeddings; a target word is generated according to the first hidden component corresponding to the first word embedding arranged at the last position of the fourth target embedding; the second word embedding corresponding to the target word is determined, and after the second word embedding is concatenated to the fourth target embedding, the masked self-attention layer is called again to perform self-attention processing on the fourth target embedding until the generated target word is a preset end symbol, and the multiple target words generated in sequence are concatenated into a target mRNA sequence.

[0168] As can be seen from the above description, the fourth target embedding is obtained by concatenating the first target embedding, the second target embedding, and the third target embedding. The first target embedding can be regarded as the word embedding corresponding to the target species text, and the third target embedding can be regarded as the word embedding corresponding to the target translation efficiency. The single amino acid character or the combination of multiple amino acid characters in the target protein sequence can be used as an amino acid word. Therefore, the second target embedding can be regarded as a combination of the word embeddings corresponding to each amino acid word in the target protein sequence. So the fourth target embedding can be regarded as an input sequence containing multiple word embeddings, and each word embedding in the input sequence can correspond to an input word.

[0169] Among them, the masked self-attention layer can include a multi-head self-attention mechanism and a masked attention mechanism. The multi-head self-attention mechanism refers to introducing multiple attention heads, and each attention head can capture the relationships between different parts of the input sequence in parallel. Therefore, based on the multi-head self-attention mechanism, the masked self-attention layer can effectively capture long-range dependencies in the input sequence. The masked attention mechanism means that there are attention values between the target word embedding at the current position and each word embedding in the input sequence that is before the target word embedding, and there are no attention values between the target word embedding and each word embedding in the input sequence that is after the target word embedding. Therefore, based on the masked attention mechanism, the masked self-attention layer can avoid leaking future information, enabling the sequence generation model to generate the target word at the current position only according to the word embeddings input at the current position and the previous positions. The feed-forward layer, that is, the feed-forward neural network, can perform non-linear transformations on the features at each position, enabling the sequence generation model to capture more complex features and more abstract representations. The feed-forward layer can also integrate the information learned by the masked self-attention layer in different aspects to form a more comprehensive representation.

[0170] Based on this, by calling the masked self-attention layer, the first attention processing result corresponding to the fourth target embedding can be determined. Then, by calling the feed-forward layer, the first hidden state corresponding to the first attention processing result can be determined. The first attention processing result can be a result matrix, and each position of the result matrix refers to the processing representation of the corresponding position in the input sequence. The first hidden state can be a hidden state matrix, and each position in the hidden state matrix refers to the intermediate state of the corresponding position in the input sequence. The first word embedding arranged at the last position of the fourth target embedding refers to the word embedding at the last position of the input sequence. Since the first hidden component corresponding to the word embedding at the last position of the input sequence can contain the information of the last position and the previous positions, the sequence generation model can effectively generate the target word at the next position based on this first hidden component. The target word is the first word in the target mRNA sequence. Then, the second word embedding corresponding to the currently generated target word is concatenated after the fourth target embedding, that is, the second word embedding is used as the word embedding at the last position of the fourth target embedding, and a new target word can be generated again. The target word is the next word in the target mRNA sequence. Repeating the above process is equivalent to repeatedly performing the task of predicting the next word by the sequence generation model until the target word is the end symbol. For example, the end symbol is [EOS]. Then, the multiple target words generated in sequence can be concatenated to obtain the target mRNA sequence.

[0171] Specifically, the sequence generation model can be a large language model. Therefore, the sequence generation model can include multiple stacked decoder layers. For example, there are a total of 12 stacked decoder layers. Stacking two decoder layers means using the output of one decoder layer as the input of the other decoder layer. Any decoder layer can include two sub-layers, which are the masked self-attention layer and the feed-forward layer mentioned above. Each decoder layer enables the sequence generation model to understand the features and relationships at different levels in the input sequence; each sub-layer usually applies residual connection and layer normalization to improve the training effect and model performance. Residual connection can avoid the problem of vanishing gradients, and layer normalization can improve the stability of the training process and the robustness of the model to input data; the sequence generation model can also be other language models, which are not limited in the embodiments of the present application.

[0172] On this basis, the sequence generation model can also include a fully connected layer and a normalization layer. By invoking the fully connected layer, the first hidden component can be transformed into a score vector of the size of the vocabulary. Each dimension in the score vector corresponds to a word in the vocabulary. For example, the vocabulary contains 10,000 words, and the first hidden component is a 512*1 vector. After being processed by the fully connected layer, a 10,000*1 score vector can be obtained. The elements in the score vector are used to indicate the scores of their corresponding words; by invoking the normalization layer, the score vector can be normalized to convert it into a probability distribution, and the probability distribution is used to indicate the probability of each word in the vocabulary being used as the target word.

[0173] The processing process of the masked self-attention layer will be described in detail below.

[0174] Exemplarily, assume that the input sequence includes 3,000 word embeddings, and each word embedding is a 64-dimensional embedding vector, that is, the input sequence is a 3,000 * 64-dimensional input matrix. In any one attention head of the masked self-attention layer, first, multiply the input matrix by a learnable query parameter matrix to obtain a 3,000 * 64-dimensional query matrix, multiply the input matrix by a learnable key parameter matrix to obtain a 3,000 * 64-dimensional key matrix, and multiply the input matrix by a learnable value parameter matrix to obtain a 3,000 * 64-dimensional value matrix. Among them, the query parameter matrix, the key parameter matrix, and the value parameter matrix are all 64 * 64-dimensional matrices, and the query matrix, the key matrix, and the value matrix all contain the hidden states of each word embedding; then, multiply the query matrix by the transposed key matrix to obtain a 3,000 * 3,000-dimensional attention value matrix; then, construct a 3,000 * 3,000-dimensional unidirectional mask matrix. For example, in the unidirectional mask matrix, the positions in the upper right corner of the unidirectional mask matrix can be assigned negative infinity, and the positions in the lower left corner of the unidirectional mask matrix can be assigned zero; then, add the attention value matrix and the unidirectional mask matrix to obtain a 3,000 * 3,000-dimensional attention mask result, and then multiply the attention mask result by the value matrix to obtain a 3,000 * 64-dimensional result matrix; assume that the masked self-attention layer contains one attention head, and use the result matrix as the first attention processing result. Assume that the masked self-attention layer contains multiple attention heads, and weight the result matrices corresponding to each attention head to obtain the first attention processing result.

[0175] It can be seen that the sequence generation model can be the first large language model, can obtain the second prompt embedding, splice the second prompt embedding at the tail of the fourth target embedding. The second prompt embedding is used to prompt the first large language model to generate the target mRNA sequence, and then call the sequence generation model to map the fourth target embedding to generate the target mRNA sequence, which can improve the accuracy of the target mRNA sequence. Specifically, the second prompt embedding can be an embedding vector obtained by embedding the second prompt, and the second prompt can be <mrna>。

[0176] In a possible implementation, the sequence generation model may first splice multiple target words generated in sequence to obtain a first candidate mRNA sequence, and then determine the target mRNA sequence from the generated first candidate mRNA sequence according to the above determination method.

[0177] Specifically, the first target embedding can be determined in various ways. One way to determine the first target embedding will be described in detail below.

[0178] In a possible implementation, to determine the first target embedding corresponding to the target species text, specifically, it may be to obtain a first description text for describing the target species text, and determine the first text embedding corresponding to the first description text; obtain a target evolutionary map, and determine the first map embedding corresponding to the target species text based on the target evolutionary map, where the edges in the target evolutionary map are the evolutionary relationships between nodes, and the target species text is one of the nodes of the target evolutionary map; splice the first text embedding and the first map embedding to obtain the first target embedding corresponding to the target species text.

[0179] Among them, the first description text can be a text describing information such as species characteristics, habits, habitats, etc. Usually, the first description text can be obtained through channels such as academic literature, biological databases, and zoo materials; then natural language processing (NLP) can be used to embed the first description text to determine the first text embedding corresponding to the first description text. Therefore, the first text embedding refers to the embedding information that can represent the first description text, that is, the first text embedding can characterize the similarities and differences between different first description texts. Optionally, the first text embedding is implemented in the form of an embedding vector. Specifically, the first description text can be input into a fourth embedding model, and the first description text is mapped based on the fourth embedding model to obtain the first text embedding corresponding to the first description text. Usually, the goal of the fourth embedding model is to learn a compact representation of the first description text so that the distances of similar first description texts in the embedding space are closer.

[0180] Among them, the target evolutionary graph is a graphical representation describing the species evolution process. The target evolutionary graph is usually presented in the form of a tree diagram. Each node in the target evolutionary graph represents a species, that is, each node in the target evolutionary graph can be species text. The target species text can be any one of the various species texts, and the target species text can be the name of the target species. Each edge in the target evolutionary graph is used to indicate the evolutionary relationship between the two connected nodes. Then, based on the target evolutionary graph, the first graph embedding corresponding to the target species text is determined. The first graph embedding refers to the embedding information that can represent the target species text within the target evolutionary graph, that is, the first graph embedding can characterize the similarities and differences between different target species texts within the target evolutionary graph, making the distances between similar target species texts closer in the embedding space. The first graph embedding can be effectively used for tasks such as analyzing evolutionary relationships, species similarities, and function prediction. Optionally, the first graph embedding is implemented in the form of an embedding vector.

[0181] Under normal circumstances, an accurate target evolutionary graph can be constructed based on the knowledge of disciplines such as morphology, molecular biology, and genetics, thereby improving the accuracy of the first graph embedding. The first graph embedding can better describe the similarities and differences between species; in addition, based on different classification criteria and data sources, different target evolutionary graphs can be constructed.

[0182] Based on this, since the first text embedding can characterize the similarities and differences between different first description texts, and the first graph embedding can characterize the similarities and differences between different target species texts within the target evolutionary graph, and both the first description text and the target species text are texts describing the corresponding species, for the first target embedding obtained by concatenating the first text embedding and the first graph embedding, the first target embedding can effectively characterize the similarities and differences between different species.

[0183] Specifically, the first graph embedding can be concatenated at the end of the first text embedding to obtain a longer first target embedding. The first target embedding can be used as one of the word embeddings of the fourth target embedding. The third target embedding can also be used as one of the word embeddings of the fourth target embedding. The second target embedding can include multiple word embeddings in the fourth target embedding. Each word embedding in the fourth target embedding has the same dimension, and any one of the word embeddings in the fourth target embedding is the embedding vector of the corresponding word in the preset vocabulary.

[0184] Specifically, the first graph embedding can be determined in various ways. Here, one way to determine the first graph embedding is described in detail.

[0185] In a possible implementation, the first graph embedding corresponding to the target species text is determined based on the target evolutionary graph. Specifically, the first node information of the target species text in the target evolutionary graph can be determined, and the first graph embedding corresponding to the target species text is determined based on the first node information.

[0186] Based on this, since the target evolutionary graph is a graphical representation describing the species evolution process, for determining the first node information of the target species text in the target evolutionary graph, the first node information of the target species text is used to characterize the evolution degree of the species corresponding to the target species text. Then, the first node information is embedded to determine the first graph embedding corresponding to the target species text. Therefore, the first graph embedding refers to the embedding information that can represent the first node information. Optionally, the first graph embedding is implemented in the form of an embedding vector. Specifically, the first node information can be input into the fifth embedding model, and the first node information is mapped based on the fifth embedding model to obtain the first graph embedding corresponding to the target species text. Generally, the goal of the fifth embedding model is to learn a compact representation of the first node information, so that the distances of similar first node information in the embedding space are closer.

[0187] In a possible implementation, the first node information includes the first node depth of the target species text, the number of first nodes in the subtree connected by the target species text, and the number of second nodes of all nodes connected by the target species text. To determine the first graph embedding corresponding to the target species text based on the first node information, specifically, the target species text can be embedded to determine the second text embedding corresponding to the target species text. The first node embeddings corresponding to the first node depth, the number of first nodes, and the number of second nodes are determined respectively, and each first node embedding is concatenated with the second text embedding to obtain multiple first concatenated embeddings. The multiple first concatenated embeddings are weighted to obtain the first graph embedding corresponding to the target species text.

[0188] Among them, in the target evolutionary graph, the depth of the first node of the target species text is used to represent the hierarchical depth of the target species text from the root node in the target evolutionary graph. For example, if the target species text is located at the 10th layer in the target evolutionary graph, the depth of the first node of the target species text is 10; the number of the first nodes in the subtree connected by the target species text is used to represent the total number of nodes included in each subtree connected by the target species text in the target evolutionary graph. For example, if the target species text is connected to 3 subtrees and the total number of nodes included in the 3 subtrees is 10, then the number of the first nodes is 10. Generally, closely related species will be located in the same subtree; the number of the second nodes of all the nodes connected by the target species text is used to represent the total number of nodes connected to the target species text in the target evolutionary graph. The nodes connected to the target species text include the ancestor nodes and the descendant nodes of the target species text. For example, if the total number of the ancestor nodes of the target species text is 1 and the total number of the descendant nodes of the target species text is 9, then the number of the second nodes is 10.

[0189] Specifically, the target species text can be the name of the target species. The depth of the first node can characterize the evolutionary timeline of the target species. When the depth of the first node is smaller, the target species diverges earlier. On the contrary, when the depth of the first node is larger, the target species diverges later; the number of the first nodes can characterize the diversity of the target species. When the number of the first nodes is larger, the diversity of the target species is more abundant. On the contrary, when the number of the first nodes is smaller, the diversity of the target species is more scarce; the number of the second nodes can characterize the impact of the target species on the entire target evolutionary graph. When the number of the second nodes is larger, the impact of the target species on the entire target evolutionary graph is greater. On the contrary, when the number of the second nodes is smaller, the impact of the target species on the entire target evolutionary graph is smaller.

[0190] Based on this, by respectively embedding the first node depth, the first node quantity, and the second node quantity, the first node embeddings corresponding to the first node depth, the first node quantity, and the second node quantity can be determined. Therefore, the first node embedding corresponding to the first node depth can indicate the embedding information of the first node depth, the first node embedding corresponding to the first node quantity can indicate the embedding information of the first node quantity, the first node embedding corresponding to the second node quantity can indicate the embedding information of the second node quantity, and by embedding the target species text, the second text embedding corresponding to the target species text can be determined, and the second text embedding can indicate the embedding information of the target species text; then, each first node embedding is respectively concatenated with the second text embedding to obtain a plurality of first concatenated embeddings. For example, the second text embedding is respectively concatenated at the tail of its corresponding first node embedding; then, according to the weights corresponding to the first node depth, the first node quantity, and the second node quantity, the plurality of first concatenated embeddings are weighted, that is, the first concatenated embeddings are weighted and fused to obtain the first graph embedding corresponding to the target species text, and the first graph embedding can effectively represent the evolutionary degree of the species corresponding to the target species text.

[0191] Optionally, each first node embedding is implemented in the form of an embedding vector. The dimension of the first node embedding can be adjusted according to actual needs, which is not limited in this embodiment of the present application. In addition, the specific determination method of the first node embedding is not limited in this embodiment of the present application. For example, the corresponding first node embedding can be determined based on a trained embedding model.

[0192] Optionally, the second text embedding is implemented in the form of an embedding vector. The dimension of the second text embedding can be adjusted according to actual needs, which is not limited in this embodiment of the present application. In addition, the specific determination method of the second text embedding is not limited in this embodiment of the present application. For example, the corresponding second text embedding can be determined based on a trained embedding model.

[0193] In a possible implementation manner, weighting the plurality of first concatenated embeddings to obtain the first graph embedding corresponding to the target species text can specifically be to respectively perform regression on each first node embedding based on a first regression model to obtain the first weight corresponding to each first node embedding; weighting the plurality of first concatenated embeddings based on the first weight to obtain the first graph embedding corresponding to the target species text, where the first regression model is jointly trained with a sequence generation model.

[0194] Based on this, the respective first node embeddings corresponding to the target species text are input into the first regression model simultaneously. Based on the first regression model, the first weights corresponding to the respective first node embeddings can be predicted. The first weights are used to characterize the contribution degree of the corresponding first node embeddings to the first graph embedding. Then, the multiple first concatenated embeddings are weighted based on the first weights. Specifically, each first node embedding is multiplied by the corresponding first weight, and the multiplication results are added together to obtain the first graph embedding, enabling information exchange between different first node embeddings and improving the accuracy of the first graph embedding. Additionally, during the model training phase, the first regression model and the sequence generation model need to be jointly trained so that the first regression model can more accurately predict the first weights, thereby improving the accuracy of the sequence generation model.

[0195] Next, another method for determining the first graph embedding will be described in detail.

[0196] In a possible implementation, the evolutionary relationship is determined by gene similarity. Based on the target evolutionary graph, the first graph embedding corresponding to the target species text is determined. Specifically, multiple nodes connected to the target species text in the target evolutionary graph are all determined as reference species texts, and the corresponding second graph embeddings of each reference species text are determined; according to the gene similarity corresponding to each reference species text, the third weight corresponding to each reference species text is determined; and the second graph embeddings are weighted based on the third weights to obtain the first target embedding corresponding to the target species text.

[0197] Among them, gene similarity refers to the similarity of gene sequences between two species. Gene similarity can be measured by comparing the DNA, RNA, or protein sequences of two species. Therefore, the evolutionary relationship between species can be inferred through gene similarity; each node in the target evolutionary graph can be a species text, the target species text can be any one of the various species texts, and each edge in the target evolutionary graph is used to indicate the evolutionary relationship between the two connected nodes, that is, each edge corresponds to the gene similarity between the two connected nodes; since there is a connection relationship between the node where the reference species text is located and the node where the target species text is located, the node where the reference species text is located belongs to the ancestor node or descendant node of the node where the target species text is located.

[0198] Based on this, the second map embeddings corresponding to each reference species text can be determined first based on the target evolutionary map. Then, since there are differences in the gene similarities between the reference species corresponding to different reference species texts and the target species corresponding to the target species text, the third weight corresponding to the reference species text can be determined first according to the gene similarity corresponding to the reference species text. The gene similarity corresponding to the reference species text refers to the gene similarity corresponding to the edge between the node where the reference species text is located and the node where the target species text is located. Usually, when the gene similarity corresponding to the reference species text is greater, the third weight corresponding to the reference species text is greater; conversely, when the gene similarity corresponding to the reference species text is smaller, the third weight corresponding to the reference species text is smaller. Then, each second map embedding is multiplied by the corresponding third weight respectively, and the multiplication results are added together to obtain the first map embedding, enabling the exchange of information between different second map embeddings and improving the accuracy of the first map embedding.

[0199] Next, another method for determining the first target embedding will be described in detail.

[0200] In a possible implementation, to determine the first target embedding corresponding to the target species text, specifically, the target evolutionary map can be obtained. Among them, the edges in the target evolutionary map are the evolutionary relationships between nodes, and the evolutionary relationships are determined by gene similarities. The target species text is one of the nodes in the target evolutionary map. All the nodes connected to the target species text in the target evolutionary map are determined as reference species texts. The second description text for describing the reference species text is obtained, and the third text embedding corresponding to the second description text is determined. According to the gene similarities corresponding to each reference species text, the second weights corresponding to each reference species text are determined. Based on the second weights, the multiple third text embeddings are weighted to obtain the first target embedding corresponding to the target species text.

[0201] Among them, the gene similarity refers to the similarity of gene sequences between two species. The gene similarity can be measured by comparing the DNA, RNA, or protein sequences of two species. Therefore, the evolutionary relationship between species can be inferred through gene similarity. In addition to gene similarity, other factors usually need to be combined to determine the evolutionary relationship between species, such as morphological characteristics, food chain relationships, etc. Each node in the target evolutionary map can be a species text, the target species text can be any one of the species texts, and each edge in the target evolutionary map is used to indicate the evolutionary relationship between the two connected nodes, that is, each edge corresponds to the gene similarity between the two connected nodes. Since there is a connection relationship between the node where the reference species text is located and the node where the target species text is located, the node where the reference species text is located belongs to the ancestor node or descendant node of the node where the target species text is located.

[0202] Based on this, the species information of the target species text can be inferred from the species information of each reference species text. Specifically, the third text embedding can be obtained by embedding the second description text of the reference species text. The third text embedding can represent the species information of the corresponding reference species text. Then, each third text embedding is weighted to obtain the first target embedding corresponding to the target species text. The first target embedding can represent the species information of the target species text. Assuming that the species information of the reference species text is known species information and the species information of the target species text is unknown species information, the mRNA sequence generation method provided in the embodiments of the present application realizes the prediction of promoting from known species information to unknown species information.

[0203] On this basis, since there are differences in the gene similarity between the reference species corresponding to different reference species texts and the target species corresponding to the target species text, the second weight corresponding to the reference species text can be determined first according to the gene similarity corresponding to the reference species text. The gene similarity corresponding to the reference species text refers to the gene similarity corresponding to the edge between the node where the reference species text is located and the node where the target species text is located. Generally, when the gene similarity corresponding to the reference species text is larger, the second weight corresponding to the reference species text is larger; on the contrary, when the gene similarity corresponding to the reference species text is smaller, the second weight corresponding to the reference species text is smaller. Then, each third text embedding is multiplied by the corresponding second weight, and the multiplication results are added together to obtain the first target embedding, enabling the exchange of information between different third text embeddings and improving the accuracy of the first target embedding.

[0204] In a possible implementation manner, the first target embedding, the second target embedding, and the third target embedding are concatenated to obtain the fourth target embedding. Specifically, the target tissue text corresponding to the target protein sequence can be obtained, and the fifth target embedding corresponding to the target tissue text is determined. The first target embedding, the fifth target embedding, the second target embedding, and the third target embedding are concatenated to obtain the fourth target embedding.

[0205] Among them, the target tissue text can be the name of the target tissue, and the target tissue is one of various tissues. Generally, tissues can include structures such as the heart, lungs, kidneys, liver, etc. Therefore, the target tissue text can be "heart", "lungs", "kidneys", "liver", etc.; embedding the target tissue text can obtain the fifth target embedding corresponding to the target tissue text, and the fifth target embedding refers to the embedding information that can represent the target tissue text. Optionally, the fifth target embedding is implemented in the form of an embedding vector. Specifically, the target tissue text can be input into the sixth embedding model, and based on the sixth embedding model, the target tissue text is mapped to obtain the fifth target embedding corresponding to the target tissue text. Usually, the goal of the sixth embedding model is to learn a compact representation of the input data so that the distances of similar inputs in the embedding space are closer.

[0206] For example, refer to Figure 6 , Figure 6 which is an optional structural schematic diagram of the second large language model provided by the embodiments of the present application.

[0207] Among them, the sixth embedding model can be the second large language model. Taking the target tissue text as "liver" as an example, specifically, the target tissue text can be input into the second large language model, and based on the second large language model, the fifth target embedding corresponding to the target tissue text is generated. The second large language model can be a trained neural network model.

[0208] Specifically, refer to Figure 7 , Figure 7 which is an optional structural schematic diagram of the sequence generation model provided by the embodiments of the present application.

[0209] Since the first target embedding is determined by the target species text of the target species, the fifth target embedding is determined by the target tissue text, the second target embedding is determined by the target protein sequence, the third target embedding is determined by the target translation efficiency, and the fourth target embedding is obtained by splicing the first target embedding, the fifth target embedding, the second target embedding, and the third target embedding, so the fourth target embedding can simultaneously contain the information of the target species text, the target tissue text, the target protein sequence, and the target translation efficiency.

[0210] Based on this, since the target translation efficiency is adjustable and can be adjusted according to the regulation requirements of the application scenario, therefore, subsequently, the fourth target embedding can be mapped through the sequence generation model to generate the first candidate mRNA sequence corresponding to the fourth target embedding, that is, when the sequence generation model generates the target mRNA sequence corresponding to the target protein sequence of the target species in the target tissue, the translation efficiency of the first candidate mRNA sequence can meet the regulation requirements of the application scenario.

[0211] Specifically, the fourth target embedding can be obtained by sequentially concatenating the first target embedding, the fifth target embedding, the second target embedding, and the third target embedding.

[0212] Specifically, the fifth target embedding can also be determined in multiple ways. Here, one way to determine the fifth target embedding will be described in detail.

[0213] In a possible implementation, to determine the fifth target embedding corresponding to the target tissue text, specifically, it can be to obtain the third descriptive text for describing the target tissue text and determine the fourth text embedding corresponding to the third descriptive text; obtain the target relationship graph, and based on the target relationship graph, determine the third graph embedding corresponding to the target tissue text, where the edges in the target relationship graph are the association relationships between nodes, and the target tissue text is one of the nodes in the target relationship graph; concatenate the fourth text embedding and the third graph embedding to obtain the fifth target embedding corresponding to the target tissue text.

[0214] Among them, the third descriptive text can be text describing aspects such as the morphology, structure, function, and location of the tissue. Usually, the third descriptive text can be obtained through channels such as academic literature and biological databases; then natural language processing (NLP) can be used to embed the third descriptive text to determine the fourth text embedding corresponding to the third descriptive text. Therefore, the fourth text embedding refers to the embedding information that can represent the third descriptive text, that is, the fourth text embedding can characterize the similarities and differences between different third descriptive texts. Optionally, the fourth text embedding is implemented in the form of an embedding vector. Specifically, the third descriptive text can be input into the seventh embedding model, and based on the seventh embedding model, the third descriptive text is mapped to obtain the fourth text embedding corresponding to the third descriptive text. Usually, the goal of the seventh embedding model is to learn the compact representation of the third descriptive text so that the distances between similar third descriptive texts in the embedding space are closer.

[0215] Among them, the target relationship graph is a graphical representation describing the association relationships between various organizations. The target relationship graph can be presented in the form of an undirected graph. Each node in the target relationship graph represents an organization, that is, each node in the target relationship graph can be an organization text. The target organization text can be any one of various organization texts, the target organization text can be the name of the target organization, and each edge in the target relationship graph is used to indicate the association relationship between the two connected nodes. Then, based on the target relationship graph, the third graph embedding corresponding to the target organization text is determined. The third graph embedding refers to the embedding information that can represent the target organization text within the target relationship graph, that is, the third graph embedding can characterize the similarities and differences between different target organization texts within the target relationship graph, so that the distances between similar target organization texts in the embedding space are closer. Optionally, the third graph embedding is implemented in the form of an embedding vector.

[0216] Generally, since the tissues of different species may vary, in different species, corresponding relationship graphs can be constructed, and then the target relationship graph can be determined according to the target species text in multiple relationship graphs.

[0217] Based on this, since the fourth text embedding can characterize the similarities and differences between different third description texts, and the third graph embedding can characterize the similarities and differences between different target organization texts within the target relationship graph, and the third description text and the target organization text are both texts describing the corresponding organizations, for the fifth target embedding obtained by splicing the fourth text embedding and the third graph embedding, the fifth target embedding can effectively characterize the similarities and differences between different organizations.

[0218] Specifically, the third graph embedding can be spliced at the end of the fourth text embedding to obtain a longer fifth target embedding. The fifth target embedding can be used as one of the word embeddings of the fourth target embedding, and the fourth target embedding can contain multiple word embeddings with the same number of dimensions.

[0219] In a possible implementation manner, to determine the third graph embedding corresponding to the target organization text based on the target relationship graph, specifically, the second node information of the target organization text in the target relationship graph can be determined; and the third graph embedding corresponding to the target organization text is determined based on the second node information.

[0220] Based on this, since the target relationship graph is a graphical representation describing the organizational evolution process, for determining the second node information of the target organizational text in the target relationship graph, the second node information of the target organizational text is used to characterize the evolution degree of the organization corresponding to the target organizational text; then the second node information is embedded to determine the third graph embedding corresponding to the target organizational text. Therefore, the third graph embedding refers to the embedding information that can represent the second node information. Optionally, the third graph embedding is implemented in the form of an embedding vector. Specifically, the second node information can be input into the eighth embedding model, and the second node information is mapped based on the eighth embedding model to obtain the third graph embedding corresponding to the target organizational text. Generally, the goal of the eighth embedding model is to learn a compact representation of the second node information, so that the distances of similar second node information in the embedding space are closer.

[0221] In a possible implementation manner, the second node information includes the third node quantity of all nodes connected by the target organizational text. Based on the second node information, determining the third graph embedding corresponding to the target organizational text can specifically be to embed the target organizational text to determine the fifth text embedding corresponding to the target organizational text; determine the second node embedding corresponding to the third node quantity, and splice the second node embedding and the fifth text embedding to obtain the third graph embedding corresponding to the target organizational text.

[0222] Among them, in the target relationship graph, the third node quantity is used to represent the total quantity of each node connected to the target organizational text in the target relationship graph. For example, if the total quantity of nodes connected to the target organizational text is 5, then the third node quantity is 5. Specifically, the target organizational text can be the name of the target organization, and the third node quantity can characterize the influence of the target organization on the entire target relationship graph. When the third node quantity is larger, the influence of the target organization on the entire target relationship graph is greater; conversely, when the third node quantity is smaller, the influence of the target organization on the entire target relationship graph is smaller.

[0223] Based on this, embedding the third node quantity can determine the second node embedding corresponding to the third node quantity. Therefore, the second node embedding can indicate the embedding information of the third node quantity, and embedding the target organizational text to determine the fifth text embedding corresponding to the target organizational text, and the fifth text embedding can indicate the embedding information of the target organizational text; then, splicing the second node embedding and the fifth text embedding to obtain the third graph embedding corresponding to the target organizational text. For example, splicing the fifth text embedding at the tail of the second node embedding respectively, the third graph embedding can effectively characterize the association relationship between the target organization corresponding to the target organizational text and other organizations.

[0224] Optionally, the second node embedding is implemented in the form of an embedding vector. The dimension of the second node embedding can be adjusted according to actual needs, which is not limited in the embodiments of the present application. In addition, the embodiments of the present application do not limit the specific determination method of the second node embedding. For example, the corresponding second node embedding can be determined based on the trained embedding model.

[0225] Optionally, the fifth text embedding is implemented in the form of an embedding vector. The dimension of the fifth text embedding can be adjusted according to actual needs, which is not limited in the embodiments of the present application. In addition, the embodiments of the present application do not limit the specific determination method of the fifth text embedding. For example, the corresponding fifth text embedding can be determined based on the trained embedding model.

[0226] Another method for determining the fifth target embedding will be described in detail below.

[0227] In a possible implementation, to determine the fifth target embedding corresponding to the target organization text, specifically, a target relationship graph can be obtained, where the edges in the target relationship graph are the association relationships between nodes, and the association relationships are determined by structural similarities. The target organization text is one of the nodes in the target relationship graph; all the nodes connected to the target organization text in the target relationship graph are determined as reference organization texts; the fourth description text for describing the reference organization texts is obtained, and the sixth text embedding corresponding to the fourth description text is determined; according to the structural similarities corresponding to the respective reference organization texts, the fourth weights corresponding to the respective reference organization texts are determined; and the multiple sixth text embeddings are weighted based on the fourth weights to obtain the fifth target embedding corresponding to the target organization text.

[0228] Among them, the structural similarity refers to the structural similarity between two organizations. The structural similarity can be determined by comparing the morphology, structure, and function of the two organizations. Therefore, the association relationship between organizations can be inferred through the structural similarity; each node in the target relationship graph can be an organization text, the target organization text can be any one of the respective organization texts, and each edge in the target relationship graph is used to indicate the association relationship between the two connected nodes, that is, each edge corresponds to the structural similarity between the two connected nodes; since there is a connection relationship between the node where the reference organization text is located and the node where the target organization text is located, generally, there is a certain structural similarity between the node where the reference organization text is located and the node where the target organization text is located.

[0229] Based on this, the organizational information of the target organizational text can be inferred from the organizational information of each reference organizational text. Specifically, the sixth text embedding can be obtained by embedding the fourth descriptive text of the reference organizational text. The sixth text embedding can represent the organizational information of the corresponding reference organizational text. Then, each sixth text embedding is weighted to obtain the fifth target embedding corresponding to the target organizational text. The fifth target embedding can represent the organizational information of the target organizational text. Assuming that the organizational information of the reference organizational text is known organizational information and the organizational information of the target organizational text is unknown organizational information, the mRNA sequence generation method provided in the embodiments of this application realizes the prediction of promoting from known organizational information to unknown organizational information.

[0230] On this basis, due to the differences in the structural similarities between the reference organizations corresponding to different reference organizational texts and the target organization corresponding to the target organizational text, the fourth weight corresponding to the reference organizational text can be determined first according to the structural similarity corresponding to the reference organizational text. The structural similarity corresponding to the reference organizational text refers to the structural similarity corresponding to the edge between the node where the reference organizational text is located and the node where the target organizational text is located. Generally, when the structural similarity corresponding to the reference organizational text is larger, the fourth weight corresponding to the reference organizational text is larger. On the contrary, when the structural similarity corresponding to the reference organizational text is smaller, the fourth weight corresponding to the reference organizational text is smaller. Then, each sixth text embedding is multiplied by the corresponding fourth weight, and the multiplication results are added together to obtain the fifth target embedding, enabling the exchange of information between different sixth text embeddings and improving the accuracy of the fifth target embedding.

[0231] In a possible implementation manner, the first target embedding, the second target embedding, and the third target embedding are concatenated to obtain the fourth target embedding. Specifically, the environmental parameters corresponding to the target protein sequence can be obtained; the sixth target embedding corresponding to the environmental parameters can be determined; the first target embedding, the second target embedding, the third target embedding, and the sixth target embedding are concatenated to obtain the fourth target embedding.

[0232] Among them, the environmental parameters can be the pH value or temperature value of the environment where the protein corresponding to the target protein sequence is located, etc. The sixth target embedding refers to the embedding information that can represent the environmental parameters. Optionally, the sixth target embedding is implemented in the form of an embedding vector.

[0233] Refer to Figure 8 , Figure 8 which is an optional flowchart of the model training method provided in the embodiments of this application. This model training method can be executed by a server, or can also be executed by a terminal, or can also be executed by a server in cooperation with a terminal. This text type determination method includes but is not limited to the following steps 801 to step 805.

[0234] Step 801: Obtain a sample protein sequence, the sample species text corresponding to the sample protein sequence, the sample translation efficiency corresponding to the sample protein sequence, and the mRNA sequence tag corresponding to the sample protein sequence;

[0235] Step 802: Determine the first sample embedding corresponding to the sample species text, determine the second sample embedding corresponding to the sample protein sequence, and determine the third sample embedding corresponding to the sample translation efficiency;

[0236] Step 803: Concatenate the first sample embedding, the second sample embedding, and the third sample embedding to obtain a fourth sample embedding;

[0237] Step 804: Invoke a sequence generation model to map the fourth sample embedding to generate a sample mRNA sequence, where the sample mRNA sequence includes a second coding region, a third untranslated region upstream of the second coding region, and a fourth untranslated region downstream of the second coding region;

[0238] Step 805: Determine the model loss based on the sample mRNA sequence and the mRNA sequence tag, and train the sequence generation model according to the model loss.

[0239] Herein, the mRNA sequence tag refers to the real mRNA sequence. During model training, a large number of training data sets and the corresponding mRNA sequence tags are usually used. Each training data set contains a sample protein sequence, a sample species text, and a sample translation efficiency.

[0240] Specifically, the server can obtain the training data set and the corresponding mRNA sequence tag from the database. The server can also obtain the training data set and the corresponding mRNA sequence tag sent by the terminal. The server can store the obtained training data set and the corresponding mRNA sequence tag in the sample pool. During model training, the training data set and the corresponding mRNA sequence tag can be randomly selected from the sample pool.

[0241] The above model training method and mRNA sequence generation method are based on the same inventive concept. Therefore, in this model training method, by obtaining a sample protein sequence, a sample species text, a sample translation efficiency, and an mRNA sequence tag, then embedding the sample protein sequence, the sample species text, and the sample translation efficiency respectively, the first sample embedding, the second sample embedding, and the third sample embedding are determined in sequence. Then, the fourth sample embedding is obtained by splicing the first sample embedding, the second sample embedding, and the third sample embedding. Furthermore, by calling the sequence generation model to map the fourth sample embedding, a sample mRNA sequence is generated. In addition to including the second coding region, the sample mRNA sequence also includes a third untranslated region and a fourth untranslated region. Similarly, the mRNA sequence tag also contains the corresponding regions, enabling the sequence generation model to generate an accurate mRNA sequence. On this basis, since the fourth sample embedding can simultaneously contain the information of the sample species text, the sample protein sequence, and the sample translation efficiency, and the sample translation efficiency can be adjusted based on regulatory requirements, therefore, when the sequence generation model generates a sample mRNA sequence corresponding to the sample protein sequence of a sample species, the translation efficiency of the sample mRNA sequence can meet the regulatory requirements of the application scenario.

[0242] For the detailed principles of the above steps 801 to 804, reference can be made to the explanations of steps 201 to 204 above, which will not be elaborated here.

[0243] Specifically, referring to Figure 9 , Figure 9 is a schematic diagram of an optional training process of the sequence generation model provided by an embodiment of the present application.

[0244] In step 805, similar to the target mRNA sequence, the sample mRNA sequence generated by the sequence generation model also includes a second coding region, a third untranslated region upstream of the second coding region, and a fourth untranslated region downstream of the second coding region. In order for the sequence generation model to generate an accurate mRNA sequence, it is necessary to preprocess the mRNA sequence tag.

[0245] Specifically, the mRNA sequence tag includes a third coding region, a fifth untranslated region upstream of the third coding region, and a sixth untranslated region downstream of the third coding region. Then, the third prompt corresponding to each of the third coding region, the fifth untranslated region, and the sixth untranslated region is obtained, and the third coding region, the fifth untranslated region, and the sixth untranslated region are respectively spliced with the corresponding third prompt to obtain the preprocessed mRNA sequence tag.

[0246] Referring again to Figure 3 , since the third coding region, the fifth untranslated region, and the sixth untranslated region usually contain multiple bases, and each base has its corresponding base character. For example, the base character corresponding to adenine in the base is A, the base character corresponding to uracil is U, the base character corresponding to cytosine is abbreviated as C, and the base character corresponding to guanine is G. The third coding region, the fifth untranslated region, and the sixth untranslated region can be segmented respectively according to a preset segmentation length to obtain multiple segmentation results, that is, multiple labeled base words are obtained. The segmentation length can be preset to 3. In the fifth untranslated region and the sixth untranslated region, each labeled base word can be a single base character, a combination of two base characters, or a combination of three base characters. In the third coding region, each labeled base word is a combination of three base characters, that is, each labeled base word is a codon. 3n means that a combination of three base characters corresponds to one amino acid character. Therefore, a word list can be constructed through all possible labeled base words, and the third prompt can also be used as a word in the word list, that is, both the labeled base words and the third prompt are labeled words.

[0247] Based on this, similar to the mRNA sequence tag, the sample mRNA sequence is obtained by splicing multiple sample base words. The sample base words include the segmentation results corresponding to the second coding region, the third untranslated region, and the fourth untranslated region. The sample mRNA sequence will also include the fourth prompts corresponding to the second coding region, the third untranslated region, and the fourth untranslated region respectively. The sample base words and the fourth prompts are both words in the word list. The sequence generation model can generate the predicted probability distribution of each position of the sample mRNA sequence. The predicted probability distribution refers to the predicted probability that each word in the word list is used as the predicted word. The target probability distribution of each position can be determined through the labeled words of each position. The target probability distribution is usually set such that the probability corresponding to the labeled word is 1, and the probability corresponding to other words is 0. Then, in each position of the sample mRNA sequence, the cross-entropy loss is determined according to the corresponding predicted probability distribution and the corresponding target probability distribution. Then, the model loss is determined according to the cross-entropy losses corresponding to all positions. For example, the model loss is obtained by adding all the cross-entropy losses. Then, it can be judged whether the training end condition is met based on the model loss. For example, the training end condition can be that the model loss is less than the loss threshold. When the training end condition is met, the model parameters of the sequence generation model are updated based on the model loss to obtain the updated sequence generation model. Then, the next round of training is performed based on the updated sequence generation model until the training end condition is met, and the trained sequence generation model is obtained, which can ensure the prediction accuracy of the sequence generation model.

[0248] Specifically, both the mRNA sequence tag and the sample mRNA sequence are the full length of mRNA. Therefore, both the mRNA sequence tag and the sample mRNA sequence will contain the 5'UTR, coding sequence, and 3'UTR arranged in sequence. The second coding region and the third coding region both refer to the coding sequence. The third untranslated region and the fifth untranslated region both refer to the 5'UTR. The fourth untranslated region and the sixth untranslated region both refer to the 3'UTR. Therefore, the fourth prompt corresponding to the third untranslated region can be <5UTR>, and the fourth prompt corresponding to the second coding region can be <cds>, The fourth prompt corresponding to the fourth non-translation region can be <3UTR>, the third prompt corresponding to the fifth non-translation region can be <5UTR>, and the third prompt corresponding to the third coding region can be <cds>, The third prompt corresponding to the sixth non-translated region may be <3UTR>.

[0249] Exemplarily, both the predicted probability distribution and the target probability distribution can be implemented in the form of vectors. Each element in the predicted probability distribution is used to indicate the predicted probability of the corresponding word, and each element in the target probability distribution is used to indicate the target probability of the corresponding word. Suppose the vocabulary contains 5 words, namely "AUA", "AUU", "GUU", "GCU", and "UGC". Suppose a labeled word is "AUU", the corresponding target probability distribution can be (0, 1, 0, 0, 0), and the corresponding predicted probability distribution can be (0.2, 0.6, 0.1, 0.07, 0.03). In the target probability distribution and the predicted probability distribution, the first to fifth elements are the probabilities of "AUA", "AUU", "GUU", "GCU", and "UGC" in sequence. Then, the cross-entropy loss at each position can be calculated, and all the cross-entropy losses are added to obtain the model loss.

[0250] In a possible implementation, the first sample embedding, the second sample embedding, and the third sample embedding are concatenated to obtain a fourth sample embedding. Specifically, the first sample embedding, the second sample embedding, the third sample embedding, and the mRNA sequence label can be concatenated to obtain the fourth sample embedding.

[0251] Based on this, during the model training process, the sequence generation model can generate a sample mRNA sequence by performing the task of predicting the next word. Since the fourth sample embedding input to the sequence generation model is obtained by concatenating the first sample embedding, the second sample embedding, the third sample embedding, and the mRNA sequence label, and the mRNA sequence label is equivalent to the true target sequence, using the true target sequence as the input of the sequence generation model enables the sequence generation model to generate the predicted word at the next position of the sample mRNA sequence based on the labeled word at any position of the true target sequence. Therefore, the sequence generation model does not need to wait until the predicted word at a certain position of the sample mRNA sequence is generated before generating the predicted word at the next position of the sample mRNA sequence, and can perform parallel training, improving the training efficiency and making the sequence generation model easier to converge during the training process.

[0252] In another possible implementation, training the sequence generation model based on the fourth sample embedding that does not contain the mRNA sequence label can improve the generalization performance of the sequence generation model.

[0253] In a possible implementation, the first sample embedding, the second sample embedding, the third sample embedding, and the mRNA sequence tag are concatenated to obtain a fourth sample embedding. Specifically, the sample tissue text corresponding to the sample protein sequence can be obtained, and the fifth sample embedding corresponding to the sample tissue text is determined; the first sample embedding, the fifth sample embedding, the second sample embedding, the third sample embedding, and the mRNA sequence tag are concatenated to obtain a fourth sample embedding.

[0254] Based on this, since the sample tissue text can be the name of the sample tissue, the sample tissue is one of multiple tissues, and the fifth sample embedding refers to the embedding information that can represent the sample tissue text, the sequence generation model can make the translation efficiency of the first candidate mRNA sequence meet the regulation requirements of the application scenario when generating the sample mRNA sequence corresponding to the sample protein sequence of the sample species in the sample tissue.

[0255] Specifically, similar to the first sample embedding, the second sample embedding, and the third sample embedding, the mRNA sequence tag can be obtained by embedding the real mRNA sequence. The process of determining the mRNA sequence tag is described in detail below.

[0256] First, the key string is obtained. The key string is the string in the third coding region, where the key string refers to the string with a specific pattern in the third coding region. For example, the key string is the string at the head and the string at the tail in the third coding region. Therefore, a regular expression can be constructed through the key string, and then through regular matching, the third coding region can be accurately matched in the real mRNA sequence, that is, the CDS is determined; then the region upstream of the third coding region is used as the fifth untranslated region, and the region downstream of the third coding region is used as the sixth untranslated region, realizing the accurate determination of the CDS, the 5'UTR upstream of the CDS, and the 3'UTR downstream of the CDS in the real mRNA sequence.

[0257] Then, a regular expression is constructed according to the key string;

[0258] Then, the real mRNA sequence is regularly matched according to the regular expression, and the third coding region is matched in the real mRNA sequence;

[0259] Then, the region upstream of the third coding region in the real mRNA sequence is used as the fifth untranslated region, and the region downstream of the third coding region in the real mRNA sequence is used as the sixth untranslated region;

[0260] Then, based on a preset word segmentation length, word segmentation is respectively performed on the third coding region, the fifth non-translated region, and the sixth non-translated region to obtain their respective corresponding multiple sub-regions. For example, if the word segmentation length is preset to 3, each sub-region of the third coding region contains 3 adjacent base characters in sequence, that is, each sub-region of the third coding region is a codon. In the fifth non-translated region and the sixth non-translated region, except for the last sub-region, other sub-regions contain 3 adjacent base characters in sequence, and the last sub-region can contain a single base character, two base characters, or three base characters. By performing word segmentation on each sequence region respectively, it is possible to avoid destroying the structure where every three bases in the third coding region form a codon;

[0261] Then, embedding is respectively performed on each sub-region to obtain their respective corresponding third word embeddings;

[0262] Then, in the third coding region, the fifth non-translated region, and the sixth non-translated region, the corresponding third word embeddings are concatenated to obtain their respective corresponding sequence region embeddings;

[0263] Then, the third prompts corresponding to the third coding region, the fifth non-translated region, and the sixth non-translated region are obtained respectively;

[0264] Then, each sequence region embedding is concatenated with the corresponding third prompt respectively to obtain multiple sequence region concatenation results, and then the sequence region concatenation results are concatenated to obtain the mRNA sequence tag.

[0265] The test results of the sequence generation model are described in detail below.

[0266] After the sequence generation model is trained, the sequence generation model is tested for sequence generation under multiple TEs through a multi-species multi-tissue data set, and the properties of the mRNA sequences generated by the sequence generation model are evaluated. Specifically, the accuracy rates of the mRNA sequences generated by the sequence generation model on the multi-species multi-tissue data set are shown in Table 1 below:

[0267] Table 1

[0268] TE NTE 0.5 0.6 0.7 0.8 0.9 1.0 Accuracy 94.32% 91.23% 94.32% 94.32% 95.45% 94.32% 95.70%

[0269] Among them, TE refers to translation efficiency, and NTE (Natural TE) refers to the translation efficiency of natural mRNA sequences. It can be seen that the accuracy rates of mRNA sequences generally reach about 95% under multiple TEs.

[0270] Specifically, refer to Figure 10 , Figure 10 This is an optional schematic diagram of the sequences that are not correctly generated under multiple translation efficiencies provided by the embodiments of the present application.

[0271] As can be seen from the figure, under multiple TEs, most of the sequences that are not correctly generated overlap, indicating that there are individual sequences that are difficult to generate no matter how the TE is adjusted. There is also a small part of the sequences that cannot be generated under extreme TEs (for example, 0.5 and 1.0), indicating that the adjustment of the TE cannot exceed the potential of the sequence itself.

[0272] Specifically, referring to Figure 11 , Figure 11 is an optional bar chart showing the codon preference provided by the embodiment of the present application.

[0273] As can be seen from the figure, the mRNA sequences generated by the sequence generation model have similar codon preferences to the natural mRNA sequences, indicating that the mRNA sequences generated by the sequence generation model can adapt to the natural properties of the species.

[0274] Specifically, referring to ​ and ​ , ​ is an optional bar chart showing the GC content at multiple translation efficiencies provided by the embodiment of the present application, ​ is an optional diagram showing the minimum free energy at multiple translation efficiencies provided by the embodiment of the present application.

[0275] Among them, the GC content refers to the percentage of guanine (G) and cytosine (C) in the mRNA sequence.

[0276] As can be seen from the figure, under multiple TEs, the GC contents are relatively close, and the minimum free energy (MFE) is also relatively close. The mRNA sequences generated by the sequence generation model have natural properties.

[0277] The complete process of the mRNA sequence generation method will be described in detail below.

[0278] Referring to ​ , ​ is an optional schematic diagram of the framework of the mRNA sequence generation method provided by the embodiment of the present application. The mRNA sequence generation method can be divided into a training stage and an inference stage.

[0279] The training stage will be described in detail below.

[0280] First, obtain the sample protein sequence, the sample species text corresponding to the sample protein sequence, the sample translation efficiency corresponding to the sample protein sequence, and the mRNA sequence tag corresponding to the sample protein sequence.

[0281] Then, determine the first sample embedding corresponding to the sample species text, determine the second sample embedding corresponding to the sample protein sequence, and determine the third sample embedding corresponding to the sample translation efficiency.

[0282] Then, obtain the sample tissue text corresponding to the sample protein sequence, and determine the fifth sample embedding corresponding to the sample tissue text.

[0283] Then, splice the first sample embedding, the fifth sample embedding, the second sample embedding, the third sample embedding, and the mRNA sequence tag to obtain the fourth sample embedding.

[0284] Then, call the sequence generation model to map the fourth sample embedding to generate a sample mRNA sequence, where the sample mRNA sequence includes a second coding region, a third untranslated region upstream of the second coding region, and a fourth untranslated region downstream of the second coding region.

[0285] Then, determine the model loss according to the sample mRNA sequence and the mRNA sequence tag, and train the sequence generation model according to the model loss.

[0286] The inference stage is described in detail below.

[0287] First, obtain the target protein sequence, the target species text corresponding to the target protein sequence, and the target translation efficiency corresponding to the target protein sequence.

[0288] Then, obtain the first description text for describing the target species text, and determine the first text embedding corresponding to the first description text.

[0289] Then, obtain the target evolutionary map, where the edges in the target evolutionary map are the evolutionary relationships between nodes, and the target species text is one of the nodes in the target evolutionary map.

[0290] Then, determine the first node information of the target species text in the target evolutionary map, where the first node information includes the first node depth of the target species text, the number of first nodes in the subtree connected by the target species text, and the number of second nodes of all nodes connected by the target species text.

[0291] Then, embed the target species text to determine the second text embedding corresponding to the target species text.

[0292] Then, respectively determine the first node embeddings corresponding to the first node depth, the first node number, and the second node number, and splice each first node embedding with the second text embedding to obtain multiple first spliced embeddings.

[0293] Then, based on the first regression model, regression is performed on each first node embedding to obtain the first weight corresponding to each first node embedding.

[0294] Then, the first weight is used to weight multiple first concatenated embeddings to obtain the first atlas embedding corresponding to the target species text, where the first regression model is jointly trained with the sequence generation model.

[0295] Then, the first text embedding and the first atlas embedding are concatenated to obtain the first target embedding corresponding to the target species text.

[0296] Then, the second target embedding corresponding to the target protein sequence is determined, and the third target embedding corresponding to the target translation efficiency is determined.

[0297] Then, the target tissue text corresponding to the target protein sequence is obtained, and the fifth target embedding corresponding to the target tissue text is determined.

[0298] Then, the first target embedding, the fifth target embedding, the second target embedding, and the third target embedding are concatenated to obtain the fourth target embedding.

[0299] Then, the sequence generation model is called to map the fourth target embedding to generate the first candidate mRNA sequence.

[0300] Then, the cumulative generation number of the first candidate mRNA sequence is determined.

[0301] Then, when the cumulative generation number is less than or equal to the number threshold, the first candidate mRNA sequence is translated to obtain the first reference protein sequence; the first edit distance between the first reference protein sequence and the target protein sequence is determined; when the first edit distance is less than or equal to the edit distance threshold, the first candidate mRNA sequence is determined as the target mRNA sequence; where the number threshold can be set to 10 and the edit distance threshold can be set to 0.

[0302] Or, when the cumulative generation number is greater than the number threshold, the temperature parameter of the sequence generation model is reduced; the sequence generation model with the reduced temperature parameter is called to map the fourth target embedding again to generate the second candidate mRNA sequence; the second candidate mRNA sequence is translated to obtain the second reference protein sequence, and the second edit distance between the second reference protein sequence and the target protein sequence is determined; when the second edit distance is less than or equal to the edit distance threshold, the second candidate mRNA sequence is determined as the target mRNA sequence.

[0303] Wherein, the target mRNA sequence includes a first coding region, a first untranslated region upstream of the first coding region, and a second untranslated region downstream of the first coding region.

[0304] Based on this, by obtaining the target protein sequence, the target species text, and the target translation efficiency, and then embedding the target protein sequence, the target species text, and the target translation efficiency respectively, the first target embedding, the second target embedding, and the third target embedding are determined in sequence. Then, the fourth target embedding is obtained by splicing the first target embedding, the second target embedding, and the third target embedding. Furthermore, by calling the sequence generation model to map the fourth target embedding, the target mRNA sequence is generated. In addition to including the first coding region, the target mRNA sequence also includes the first untranslated region and the second untranslated region, and an accurate and complete target mRNA sequence can be obtained. On this basis, since the fourth target embedding can simultaneously contain the information of the target species text, the target protein sequence, and the target translation efficiency, and the target translation efficiency can be adjusted based on regulatory requirements, the sequence generation model can make the translation efficiency of the target mRNA sequence meet the regulatory requirements of the application scenario when generating the target mRNA sequence corresponding to the target protein sequence of the target species, thereby improving the application effect of the target mRNA sequence and obtaining a target mRNA sequence that is more in line with the application scenario.

[0305] It can be seen that the mRNA sequence generation method provided by the embodiments of the present application can be applied to various scenarios.

[0306] For example, in the scenario of mRNA vaccine research and development, the target translation efficiency can be adjusted to a relatively high value. The design of mRNA with high translation efficiency helps to improve the expression level of antigen proteins, thereby enhancing the immune effect of the vaccine. Moreover, by optimizing codon usage, adjusting mRNA structure and stability, etc., efficient and stable vaccine production can be achieved. In addition, mRNA vaccines with high translation efficiency usually have a faster research and development speed and lower production costs, providing important support for dealing with newly emerging pathogens and large-scale vaccination.

[0307] Another example is in the scenario of bioprocess production. The target translation efficiency can be adjusted to a relatively high value. The design of mRNA with high translation efficiency can improve the expression level of recombinant proteins, shorten the production cycle, and reduce production costs. This is of great significance for the production of high-value, low-abundance proteins such as biopharmaceuticals, enzymes, and vaccine antigens. By optimizing the mRNA sequence and structure, researchers can achieve efficient protein production to meet the needs of different application fields.

[0308] For another example, in the field of gene therapy, the target translation efficiency can be adjusted to a relatively high value. The mRNA design with high translation efficiency helps to increase the expression level of the therapeutic protein, thereby improving the therapeutic effect. In addition, by adjusting the translation efficiency, time- and dose-dependent protein expression can also be achieved to meet the dynamic requirements during the treatment process. This is of great value for the research and development of personalized therapies and precision medicine.

[0309] For another example, in the field of synthetic biology, the target translation efficiency can be adjusted to a relatively high value. The mRNA design with high translation efficiency helps to achieve fine regulation of gene expression. By adjusting the translation efficiency, co-expression of different genes under specific conditions can be achieved to realize the design and function of complex biological systems. This provides the possibility for constructing synthetic biological systems and developing new biotechnologies.

[0310] It can be understood that although the steps in each of the above flowcharts are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0311] Referring to ​ , ​ FIG. 13 is an optional structural schematic diagram of an mRNA sequence generation device provided by an embodiment of the present application. The mRNA sequence generation device 1500 includes:

[0312] A first acquisition module 1501, configured to acquire a target protein sequence, a target species text corresponding to the target protein sequence, and a target translation efficiency corresponding to the target protein sequence;

[0313] A first processing module 1502, configured to determine a first target embedding corresponding to the target species text, determine a second target embedding corresponding to the target protein sequence, and determine a third target embedding corresponding to the target translation efficiency;

[0314] A second processing module 1503, configured to splice the first target embedding, the second target embedding, and the third target embedding to obtain a fourth target embedding;

[0315] The first generation module 1504 is configured to call a sequence generation model to map a fourth target embedding to generate a target mRNA sequence, where the target mRNA sequence includes a first coding region, a first untranslated region upstream of the first coding region, and a second untranslated region downstream of the first coding region.

[0316] Furthermore, the above-mentioned first generation module 1504 is specifically configured to:

[0317] Call a sequence generation model to map a fourth target embedding to generate a first candidate mRNA sequence;

[0318] Translate the first candidate mRNA sequence to obtain a first reference protein sequence, and determine a first edit distance between the first reference protein sequence and the target protein sequence; [[ID=NULL]]

[0319] When the first edit distance is less than or equal to an edit distance threshold, determine the first candidate mRNA sequence as the target mRNA sequence.

[0320] Furthermore, the above-mentioned first generation module 1504 is specifically configured to:

[0321] Determine the cumulative generation number of the first candidate mRNA sequence;

[0322] When the cumulative generation number is less than or equal to a number threshold, translate the first candidate mRNA sequence to obtain a first reference protein sequence.

[0323] Furthermore, the above-mentioned first generation module 1504 is specifically configured to:

[0324] When the cumulative generation number is greater than the number threshold, reduce the temperature parameter of the sequence generation model;

[0325] Call the sequence generation model with the reduced temperature parameter to map the fourth target embedding again to generate a second candidate mRNA sequence;

[0326] Translate the second candidate mRNA sequence to obtain a second reference protein sequence, and determine a second edit distance between the second reference protein sequence and the target protein sequence;

[0327] When the second edit distance is less than or equal to the edit distance threshold, determine the second candidate mRNA sequence as the target mRNA sequence.

[0328] Furthermore, the sequence generation model includes a masked self-attention layer and a feed-forward layer, and the fourth target embedding includes a plurality of first word embeddings arranged in sequence. The above-mentioned first generation module 1504 is specifically configured to:

[0329] Call the masked self-attention layer to perform self-attention processing on the fourth target embedding to obtain a first attention processing result;

[0330] Call the feed-forward layer to perform transformation processing on the first attention processing result to obtain a first hidden state, where the first hidden state includes first hidden components corresponding to each first word embedding;

[0331] Generate a target word according to the first hidden component corresponding to the first word embedding arranged at the last position of the fourth target embedding;

[0332] Determine the second word embedding corresponding to the target word, concatenate the second word embedding to the fourth target embedding, and then call the masked self-attention layer to perform self-attention processing on the fourth target embedding until the generated target word is a preset end symbol, and concatenate the multiple target words generated in sequence into a target mRNA sequence.

[0333] Further, the above-mentioned first processing module 1502 is specifically configured to:

[0334] Obtain a first description text for describing the target species text, and determine a first text embedding corresponding to the first description text;

[0335] Obtain a target evolutionary map, and determine a first map embedding corresponding to the target species text based on the target evolutionary map, where the edges in the target evolutionary map are evolutionary relationships between nodes, and the target species text is one of the nodes in the target evolutionary map;

[0336] Concatenate the first text embedding and the first map embedding to obtain a first target embedding corresponding to the target species text.

[0337] Further, the above-mentioned first processing module 1502 is specifically configured to:

[0338] Determine the first node information of the target species text in the target evolutionary map;

[0339] Determine a first map embedding corresponding to the target species text based on the first node information.

[0340] Further, the first node information includes the first node depth of the target species text, the number of first nodes in the subtree connected by the target species text, and the number of second nodes of all nodes connected by the target species text. The above-mentioned first processing module 1502 is specifically configured to:

[0341] Embed the target species text to determine a second text embedding corresponding to the target species text;

[0342] Determine the first node embeddings corresponding to the first node depth, the first node quantity, and the second node quantity respectively, and splice each of the first node embeddings with the second text embedding to obtain a plurality of first spliced embeddings;

[0343] Weight the plurality of first spliced embeddings to obtain the first graph embedding corresponding to the target species text.

[0344] Furthermore, the above-mentioned first processing module 1502 is specifically configured to:

[0345] Perform regression on each of the first node embeddings respectively based on the first regression model to obtain the first weights corresponding to the respective first node embeddings;

[0346] Weight the plurality of first spliced embeddings based on the first weights to obtain the first graph embedding corresponding to the target species text, where the first regression model is jointly trained with the sequence generation model.

[0347] Furthermore, the above-mentioned first processing module 1502 is specifically configured to:

[0348] Obtain a target evolutionary graph, where the edges in the target evolutionary graph are the evolutionary relationships between nodes, the evolutionary relationships are determined by gene similarities, and the target species text is one of the nodes of the target evolutionary graph;

[0349] Determine all the nodes connected to the target species text in the target evolutionary graph as reference species texts;

[0350] Obtain the second description text for describing the reference species text, and determine the third text embedding corresponding to the second description text;

[0351] Determine the second weights corresponding to the respective reference species texts according to the gene similarities corresponding to the respective reference species texts;

[0352] Weight the plurality of third text embeddings based on the second weights to obtain the first target embedding corresponding to the target species text.

[0353] Furthermore, the above-mentioned second processing module 1503 is specifically configured to:

[0354] Obtain the target tissue text corresponding to the target protein sequence, and determine the fifth target embedding corresponding to the target tissue text;

[0355] Splice the first target embedding, the fifth target embedding, the second target embedding, and the third target embedding to obtain a fourth target embedding.

[0356] The above mRNA sequence generation device 1500 and the mRNA sequence generation method are based on the same inventive concept. By obtaining the target protein sequence, the target species text, and the target translation efficiency, then respectively embedding the target protein sequence, the target species text, and the target translation efficiency, the first target embedding, the second target embedding, and the third target embedding are determined in sequence. Then, the fourth target embedding is obtained by splicing the first target embedding, the second target embedding, and the third target embedding. Furthermore, by calling the sequence generation model to map the fourth target embedding, the target mRNA sequence is generated. In addition to including the first coding region, the target mRNA sequence also includes the first untranslated region and the second untranslated region, and an accurate and complete target mRNA sequence can be obtained. On this basis, since the fourth target embedding can simultaneously contain the information of the target species text, the target protein sequence, and the target translation efficiency, and the target translation efficiency can be adjusted based on regulatory requirements, the sequence generation model can make the translation efficiency of the target mRNA sequence meet the regulatory requirements of the application scenario when generating the target mRNA sequence corresponding to the target protein sequence of the target species, thereby improving the application effect of the target mRNA sequence and obtaining a target mRNA sequence that better conforms to the application scenario.

[0357] Refer to ​ , ​ which is an optional structural schematic diagram of the model training device provided by the embodiment of the present application. The model training device 1600 includes:

[0358] A second acquisition module 1601, configured to acquire a sample protein sequence, a sample species text corresponding to the sample protein sequence, a sample translation efficiency corresponding to the sample protein sequence, and an mRNA sequence tag corresponding to the sample protein sequence;

[0359] A third processing module 1602, configured to determine a first sample embedding corresponding to the sample species text, determine a second sample embedding corresponding to the sample protein sequence, and determine a third sample embedding corresponding to the sample translation efficiency;

[0360] A fourth processing module 1603, configured to splice the first sample embedding, the second sample embedding, and the third sample embedding to obtain a fourth sample embedding;

[0361] A second generation module 1604, configured to call a sequence generation model to map the fourth sample embedding to generate a sample mRNA sequence, where the sample mRNA sequence includes a second coding region, a third untranslated region located upstream of the second coding region, and a fourth untranslated region located downstream of the second coding region;

[0362] A training module 1605, configured to determine a model loss according to a sample mRNA sequence and an mRNA sequence tag, and train a sequence generation model according to the model loss.

[0363] Further, the fourth processing module 1603 is specifically configured to:

[0364] Concatenate the first sample embedding, the second sample embedding, the third sample embedding, and the mRNA sequence tag to obtain a fourth sample embedding.

[0365] The above model training device 1600 and the model training method are based on the same inventive concept. By obtaining a sample protein sequence, a sample species text, a sample translation efficiency, and an mRNA sequence tag, then respectively embedding the sample protein sequence, the sample species text, and the sample translation efficiency to sequentially determine a first sample embedding, a second sample embedding, and a third sample embedding, and then obtaining a fourth sample embedding by concatenating the first sample embedding, the second sample embedding, and the third sample embedding. Further, by calling a sequence generation model to map the fourth sample embedding, a sample mRNA sequence is generated. The sample mRNA sequence includes, in addition to the second coding region, a third untranslated region and a fourth untranslated region. Similarly, the mRNA sequence tag also contains corresponding regions, enabling the sequence generation model to generate an accurate mRNA sequence. On this basis, since the fourth sample embedding can simultaneously contain information on the sample species text, the sample protein sequence, and the sample translation efficiency, and the sample translation efficiency can be adjusted based on regulatory requirements, the sequence generation model can, when generating a sample mRNA sequence corresponding to the sample protein sequence of a sample species, enable the translation efficiency of the sample mRNA sequence to meet the regulatory requirements of the application scenario.

[0366] The electronic device provided in an embodiment of the present application for executing the above mRNA sequence generation method or model training method may be a terminal. Referring to ​ , ​ is a partial structural block diagram of the terminal provided in an embodiment of the present application. The terminal includes: a camera component 1710, a memory 1720, an input unit 1730, a display unit 1740, a sensor 1750, an audio circuit 1760, a wireless fidelity (WiFi) module 1770, a processor 1780, and a power supply 1790, etc. Those skilled in the art can understand that ​ the terminal structure shown in

[0367] The camera component 1710 can be used to collect images or videos. Optionally, the camera component 1710 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions.

[0368] The memory 1720 can be used to store software programs and modules. The processor 1780 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the memory 1720.

[0369] The input unit 1730 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the terminal. Specifically, the input unit 1730 may include a touch panel 1731 and other input devices 1732.

[0370] The display unit 1740 can be used to display input information or provided information, as well as various menus of the terminal. The display unit 1740 may include a display panel 1741.

[0371] The audio circuit 1760, the speaker 1761, and the microphone 1762 can provide an audio interface.

[0372] The power supply 1790 can be alternating current, direct current, a disposable battery, or a rechargeable battery.

[0373] The number of sensors 1750 can be one or more. The one or more sensors 1750 include, but are not limited to: an acceleration sensor, a gyroscope sensor, a pressure sensor, an optical sensor, etc. Among them:

[0374] The acceleration sensor can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established by the terminal. For example, the acceleration sensor can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1780 can control the display unit 1740 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor. The acceleration sensor can also be used for game or user motion data collection.

[0375] The gyroscope sensor can detect the body orientation and rotation angle of the terminal. The gyroscope sensor can cooperate with the acceleration sensor to collect the 3D actions of the user on the terminal. Based on the data collected by the gyroscope sensor, the processor 1780 can implement the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0376] The pressure sensor can be disposed on the side frame of the terminal and / or the lower layer of the display unit 1740. When the pressure sensor is disposed on the side frame of the terminal, it can detect the holding signal of the user on the terminal, and the processor 1780 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor. When the pressure sensor is disposed on the lower layer of the display unit 1740, the processor 1780 can control the operable controls on the UI interface according to the pressure operation of the user on the display unit 1740. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0377] The optical sensor is used to collect the ambient light intensity. In one embodiment, the processor 1780 can control the display brightness of the display unit 1740 according to the ambient light intensity collected by the optical sensor. Specifically, when the ambient light intensity is high, the display brightness of the display unit 1740 is increased; when the ambient light intensity is low, the display brightness of the display unit 1740 is decreased. In another embodiment, the processor 1780 can also dynamically adjust the shooting parameters of the camera assembly 1710 according to the ambient light intensity collected by the optical sensor.

[0378] In this embodiment, the processor 1780 included in the terminal can execute the mRNA sequence generation method or the model training method of the previous embodiment.

[0379] The electronic device provided in the embodiment of the present application for executing the above mRNA sequence generation method or model training method can also be a server. Refer to ​ , ​ This is a partial structural block diagram of the server provided by the embodiments of the present application. The server 1800 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1822 (for example, one or more processors) and a memory 1832, and one or more storage media 1830 (for example, one or more mass storage devices) for storing application programs 1842 or data 1844. Among them, the memory 1832 and the storage media 1830 may be transient storage or persistent storage. The program stored in the storage media 1830 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 1800. Further, the central processing unit 1822 may be configured to communicate with the storage media 1830 and execute a series of instruction operations in the storage media 1830 on the server 1800.

[0380] The server 1800 may further include one or more power supplies 1826, one or more wired or wireless network interfaces 1850, one or more input / output interfaces 1858, and / or one or more operating systems 1841, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0381] The central processing unit 1822 in the server 1800 may be used to execute the mRNA sequence generation method or the model training method.

[0382] The embodiments of the present application further provide a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the mRNA sequence generation method or the model training method of the foregoing various embodiments.

[0383] The embodiments of the present application further provide a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the mRNA sequence generation method or the model training method described above.

[0384] In the description of the present application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0385] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c may be single or multiple.

[0386] It should be understood that in the description of the embodiments of the present application, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.

[0387] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be an indirect coupling or communication connection through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0388] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0389] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0390] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0391] It should also be understood that the various embodiments provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0392] The above has specifically described the preferred embodiments of the present application, but the present application is not limited to the above-mentioned embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without violating the spirit of the present application. These equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.< / cds> < / cds> < / mrna> < / te> < / protein> < / species>

Claims

1. A method for generating an mRNA sequence, characterized in that, Including: Obtaining a target protein sequence, a target species text corresponding to the target protein sequence, and a target translation efficiency corresponding to the target protein sequence; Determining a first target embedding corresponding to the target species text, determining a second target embedding corresponding to the target protein sequence, and determining a third target embedding corresponding to the target translation efficiency; Concatenating the first target embedding, the second target embedding, and the third target embedding to obtain a fourth target embedding; Invoking a sequence generation model to map the fourth target embedding to generate a target mRNA sequence, where the target mRNA sequence includes a first coding region, a first untranslated region upstream of the first coding region, and a second untranslated region downstream of the first coding region.

2. The method for generating an mRNA sequence according to claim 1, wherein The invoking the sequence generation model to map the fourth target embedding to generate a target mRNA sequence includes: Invoking the sequence generation model to map the fourth target embedding to generate a first candidate mRNA sequence; Translating the first candidate mRNA sequence to obtain a first reference protein sequence, and determining a first edit distance between the first reference protein sequence and the target protein sequence; When the first edit distance is less than or equal to an edit distance threshold, determining the first candidate mRNA sequence as the target mRNA sequence.

3. The method for generating an mRNA sequence according to claim 2, wherein The translating the first candidate mRNA sequence to obtain a first reference protein sequence includes: Determining the cumulative generation number of the first candidate mRNA sequence; When the cumulative generation number is less than or equal to a number threshold, translating the first candidate mRNA sequence to obtain a first reference protein sequence.

4. The mRNA sequence generation method according to claim 3, wherein The mRNA sequence generation method further includes: When the cumulative generation number is greater than the number threshold, reducing the temperature parameter of the sequence generation model; Invoking the sequence generation model after reducing the temperature parameter to map the fourth target embedding again to generate a second candidate mRNA sequence; Translating the second candidate mRNA sequence to obtain a second reference protein sequence, and determining a second edit distance between the second reference protein sequence and the target protein sequence; When the second edit distance is less than or equal to the edit distance threshold, determining the second candidate mRNA sequence as the target mRNA sequence.

5. The mRNA sequence generation method according to claim 1, characterized in that, The sequence generation model includes a masked self-attention layer and a feed-forward layer, the fourth target embedding includes a plurality of first word embeddings arranged in sequence, and the invoking the sequence generation model to map the fourth target embedding to generate a target mRNA sequence includes: Invoking the masked self-attention layer to perform self-attention processing on the fourth target embedding to obtain a first attention processing result; Invoking the feed-forward layer to perform transformation processing on the first attention processing result to obtain a first hidden state, where the first hidden state includes first hidden components corresponding to each of the first word embeddings; Generating a target word according to the first hidden component corresponding to the first word embedding arranged at the last position of the fourth target embedding. Determine the second word embedding corresponding to the target word. After splicing the second word embedding to the fourth target embedding, call the masked self-attention layer again to perform self-attention processing on the fourth target embedding until the generated target word is a preset end symbol, and splice the multiple generated target words in sequence into a target mRNA sequence.

6. The mRNA sequence generation method according to claim 1, characterized in that, The determination of the first target embedding corresponding to the target species text includes: Obtain the first description text for describing the target species text, and determine the first text embedding corresponding to the first description text; Obtain a target evolutionary map, and determine the first map embedding corresponding to the target species text based on the target evolutionary map, where the edges in the target evolutionary map are the evolutionary relationships between nodes, and the target species text is one of the nodes in the target evolutionary map; Splice the first text embedding and the first map embedding to obtain the first target embedding corresponding to the target species text.

7. The method for generating an mRNA sequence according to claim 6, wherein The determination of the first map embedding corresponding to the target species text based on the target evolutionary map includes: Determine the first node information of the target species text in the target evolutionary map; Determine the first map embedding corresponding to the target species text based on the first node information.

8. The mRNA sequence generation method according to claim 7, wherein The first node information includes the first node depth of the target species text, the number of first nodes in the subtree connected by the target species text, and the number of second nodes of all nodes connected by the target species text. The determination of the first map embedding corresponding to the target species text based on the first node information includes: Embed the target species text to determine the second text embedding corresponding to the target species text; Respectively determine the first node embeddings corresponding to the first node depth, the first node number, and the second node number, and splice each first node embedding with the second text embedding to obtain a plurality of first spliced embeddings; Weight the plurality of first spliced embeddings to obtain the first map embedding corresponding to the target species text.

9. The mRNA sequence generation method according to claim 8, wherein The weighting of the plurality of first spliced embeddings to obtain the first map embedding corresponding to the target species text includes: Respectively perform regression on each first node embedding based on a first regression model to obtain the first weight corresponding to each first node embedding; Weight the plurality of first spliced embeddings based on the first weight to obtain the first map embedding corresponding to the target species text, where the first regression model is jointly trained with the sequence generation model.

10. The method for generating an mRNA sequence according to claim 1, wherein The determination of the first target embedding corresponding to the target species text includes: Obtain a target evolutionary map, where the edges in the target evolutionary map are the evolutionary relationships between nodes, the evolutionary relationship is determined by gene similarity, and the target species text is one of the nodes in the target evolutionary map; Determine all nodes connected to the target species text in the target evolutionary map as reference species texts; Obtain a second description text for describing the reference species text, and determine a third text embedding corresponding to the second description text; Determine a second weight corresponding to each of the reference species texts according to the gene similarity corresponding to each of the reference species texts; Weight the multiple third text embeddings based on the second weight to obtain a first target embedding corresponding to the target species text.

11. The method for generating an mRNA sequence according to claim 1, wherein The step of concatenating the first target embedding, the second target embedding, and the third target embedding to obtain a fourth target embedding includes: Obtain a target tissue text corresponding to the target protein sequence, and determine a fifth target embedding corresponding to the target tissue text; Concatenate the first target embedding, the fifth target embedding, the second target embedding, and the third target embedding to obtain a fourth target embedding.

12. A model training method, characterized in that, It includes: Obtain a sample protein sequence, a sample species text corresponding to the sample protein sequence, a sample translation efficiency corresponding to the sample protein sequence, and an mRNA sequence tag corresponding to the sample protein sequence; Determine a first sample embedding corresponding to the sample species text, determine a second sample embedding corresponding to the sample protein sequence, and determine a third sample embedding corresponding to the sample translation efficiency; Concatenate the first sample embedding, the second sample embedding, and the third sample embedding to obtain a fourth sample embedding; Call a sequence generation model to map the fourth sample embedding to generate a sample mRNA sequence, where the sample mRNA sequence includes a second coding region, a third untranslated region upstream of the second coding region, and a fourth untranslated region downstream of the second coding region; Determine a model loss according to the sample mRNA sequence and the mRNA sequence tag, and train the sequence generation model according to the model loss.

13. The model training method according to claim 12, wherein The step of concatenating the first sample embedding, the second sample embedding, and the third sample embedding to obtain a fourth sample embedding includes: Concatenate the first sample embedding, the second sample embedding, the third sample embedding, and the mRNA sequence tag to obtain a fourth sample embedding.

14. An mRNA sequence generation device, characterized in that, It includes: A first acquisition module, configured to acquire a target protein sequence, a target species text corresponding to the target protein sequence, and a target translation efficiency corresponding to the target protein sequence; A first processing module, configured to determine a first target embedding corresponding to the target species text, determine a second target embedding corresponding to the target protein sequence, and determine a third target embedding corresponding to the target translation efficiency; A second processing module, configured to concatenate the first target embedding, the second target embedding, and the third target embedding to obtain a fourth target embedding; A first generation module, configured to call a sequence generation model to map the fourth target embedding to generate a target mRNA sequence, where the target mRNA sequence includes a first coding region, a first untranslated region upstream of the first coding region, and a second untranslated region downstream of the first coding region.

15. A model training device, characterized in that, It includes: A second acquisition module, configured to acquire a sample protein sequence, a sample species text corresponding to the sample protein sequence, a sample translation efficiency corresponding to the sample protein sequence, and an mRNA sequence tag corresponding to the sample protein sequence; A third processing module, configured to determine a first sample embedding corresponding to the sample species text, determine a second sample embedding corresponding to the sample protein sequence, and determine a third sample embedding corresponding to the sample translation efficiency; A fourth processing module, configured to splice the first sample embedding, the second sample embedding, and the third sample embedding to obtain a fourth sample embedding; A second generation module, configured to call a sequence generation model to map the fourth sample embedding to generate a sample mRNA sequence, where the sample mRNA sequence includes a second coding region, a third untranslated region located upstream of the second coding region, and a fourth untranslated region located downstream of the second coding region; A training module, configured to determine a model loss according to the sample mRNA sequence and the mRNA sequence tag, and train the sequence generation model according to the model loss.

16. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the mRNA sequence generation method according to any one of claims 1 to 11, or implements the model training method according to any one of claims 12 to 13.

17. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the mRNA sequence generation method according to any one of claims 1 to 11, or implements the model training method according to any one of claims 12 to 13.

18. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the mRNA sequence generation method according to any one of claims 1 to 11, or implements the model training method according to any one of claims 12 to 13.

Citation Information

Cited By

  • Generative language model and reinforcement learning-based mRNA 5'untranslated region directional design optimization method

    CN121354678A

  • A method for targeted design optimization of mRNA 5′ untranslated region based on generative language models and reinforcement learning

    CN121354678B